/opt/cloudlinux/venv/lib/python3.11/site-packages/lxml/html/__pycache__
NameSizeModeActions
builder.cpython-311.pyc43060644editdlrm
clean.cpython-311.pyc301400644editdlrm
defs.cpython-311.pyc37810644editdlrm
diff.cpython-311.pyc388760644editdlrm
ElementSoup.cpython-311.pyc7040644editdlrm
formfill.cpython-311.pyc145010644editdlrm
html5parser.cpython-311.pyc101880644editdlrm
soupparser.cpython-311.pyc132800644editdlrm
usedoctest.cpython-311.pyc5500644editdlrm
_diffcommand.cpython-311.pyc45610644editdlrm
_html5builder.cpython-311.pyc61320644editdlrm
_setmixin.cpython-311.pyc27840644editdlrm
__init__.cpython-311.pyc903810644editdlrm
Edit: /opt/cloudlinux/venv/lib/python3.11/site-packages/lxml/html/__pycache__/diff.cpython-311.pyc (38876B)
§ ‡Z½+ÛÉÏ’ãóΗddlmZddlZddlmZddlmZddlZddgZ ddl m Z n#e $r ddl m Z YnwxYw eZn #e$reZYnwxYw en #e$reZYnwxYwd„Zefd „Zd „Zd „Zd „Zd „Zd„Zd„Zd„Zd„ZdGd„Zd„ZGd„d¦«ZGd„d¦«Z Gd„de!¦«Z"d„Z#d„Z$d„Z%d„Z&d„Z'd „Z(Gd!„d"e¦«Z)Gd#„d$e)¦«Z*Gd%„d&e)¦«Z+dHd(„Z,dHd)„Z-ej.d*ej/ej0z¦«Z1ej.d+ej/ej0z¦«Z2ej.d,ej/ej0z¦«Z3d-„Z4ej.d.¦«Z5d/„Z6d0„Z7d1Z8d2Z9d3Z:dGd4„Z;ej.d5ej<¦«Z=d6„Z>ej.d7¦«Z?d8„Z@d9„ZAd:„ZBd;„ZCd<„ZDd=„ZEdGd>„ZFd?„ZGd@„ZHdA„ZIdB„ZJGdC„dDejK¦«ZLeMdEkrddFlmNZNeNjO¦«dSdS)Ié)Úabsolute_importN)Úetree)Úfragment_fromstringÚ html_annotateÚhtmldiff)ÚescapecóJ—dtt|¦«d¦«›d|›d�S)Nz z)Ú html_escapeÚ_unicode)ÚtextÚversions úb/builddir/build/BUILD/cloudlinux-venv-1.0.12/venv/lib64/python3.11/site-packages/lxml/html/diff.pyÚdefault_markuprs/€€å•H˜WÑ%Ô% qÑ)Ô)Ð)Ð)¨4¨4¨4ð 1ð1ócóô—d„|D¦«}|d}|dd…D]}t||¦«|}Œt|¦«}t||¦«}d |¦« ¦«S)a doclist should be ordered from oldest to newest, like:: >>> version1 = 'Hello World' >>> version2 = 'Goodbye World' >>> print(html_annotate([(version1, 'version 1'), ... (version2, 'version 2')])) Goodbye World The documents must be *fragments* (str/UTF8 or unicode), not complete documents The markup argument is a function to markup the spans of words. This function is called like markup('Hello', 'version 2'), and returns HTML. The first argument is text and never includes any markup. The default uses a span with a title: >>> print(default_markup('Some Text', 'by Joe')) Some Text có4—g|]\}}t||¦«‘ŒS©)Útokenize_annotated)Ú.0Údocrs rú z!html_annotate..=s6€ð.ð.ð.Ù!�S˜'õ$ C¨Ñ1Ô1ð.ð.ð.rrr NÚ)Úhtml_annotate_merge_annotationsÚcompress_tokensÚmarkup_serialize_tokensÚjoinÚstrip)ÚdoclistÚmarkupÚ tokenlistÚ cur_tokensÚtokensÚresults rrr#s‘€ð4.ð.Ø%,ð.ñ.ô.€Ià˜1”€JؘA˜B˜B”-ððˆÝ'¨ °FÑ;Ô;Ð;؈ ˆ õ! Ñ,Ô,€Jå $ Z°Ñ 8Ô 8€FØ �7Š7�6‰?Œ?× Ò Ñ "Ô "Ð"rcó@—t|d¬¦«}|D] }||_Œ |S)zFTokenize a document and add an annotation attribute to each token F©Ú include_hrefs)ÚtokenizeÚ annotation)rr)r#Útoks rrrKs3€õ�c¨Ð /Ñ /Ô /€FØð$ð$ˆØ#ˆŒˆØ €Mrcóº—t||¬¦«}| ¦«}|D]2\}}}}}|dkr$|||…} |||…} t| | ¦«Œ3dS)zˆMerge the annotations from tokens_old into tokens_new, when the tokens in the new document already existed in the old document. ©ÚaÚbÚequalN)ÚInsensitiveSequenceMatcherÚ get_opcodesÚcopy_annotations) Ú tokens_oldÚ tokens_newÚsÚcommandsÚcommandÚi1Úi2Új1Új2Úeq_oldÚeq_news rrrSs€õ # Z°:Ð>Ñ>Ô>€AØ�}Š}‰Œ€Hà#+ð-ð-ш��R˜˜RØ �gÒ Ð Ø  2 Ô&ˆFØ  2 Ô&ˆFÝ ˜V VÑ ,Ô ,Ð ,øð -ð-rcóŽ—t|¦«t|¦«ksJ‚t||¦«D]\}}|j|_ŒdS)zN Copy annotations from the tokens listed in src to the tokens in dest N)ÚlenÚzipr))ÚsrcÚdestÚsrc_tokÚdest_toks rr2r2`sU€õ ˆs‰8Œ8•s˜4‘y”yÒ Ð Ð Ð Ý   d™^œ^ð1ð1ш�Ø%Ô0ˆÔÐð1ð1rcóÒ—|dg}|dd…D]R}|djs.|js'|dj|jkrt||¦«Œ=| |¦«ŒS|S)zm Combine adjacent tokens when there is no HTML between the tokens, and they share an annotation rr Néÿÿÿÿ)Ú post_tagsÚpre_tagsr)Úcompress_merge_backÚappend)r#r$r*s rrrhs€€ð �QŒiˆ[€FØ�a�b�bŒzððˆØ�r” Ô$ð Ø” ð à �2ŒJÔ ! S¤^Ò 3Ð 3Ý  ¨Ñ ,Ô ,Ð ,Ð ,à �MŠM˜#Ñ Ô Ð Ð Ø €MrcóL—|d}t|¦«tust|¦«tur| |¦«dSt|¦«}|jr ||jz }||z }t||j|j|j¬¦«}|j|_||d<dS)zY Merge tok into the last element of tokens (modifying the list of tokens in-place). rF©rHrGÚtrailing_whitespaceN)ÚtypeÚtokenrJr rMrHrGr))r#r*Úlastr Úmergeds rrIrIws´€ð �"Œ:€DÝ ˆD�z„z�ÐÐ¥$ s¡)¤)µ5Ð"8Ð"8Ø� Š �cÑÔÐÐÐ嘉~Œ~ˆØ Ô #ð -Ø �DÔ,Ñ ,ˆDØ �‰ ˆÝ�tØ $¤ Ø!$¤Ø+.Ô+BðDñDôDˆð!œOˆÔ؈ˆr‰ ˆ ˆ rc#óÀK—|D]X}|jD]}|V—Œ| ¦«}|||j¦«}|jr ||jz }|V—|jD]}|V—ŒŒYdS)zz Serialize the list of tokens into a list of text chunks, calling markup_func around text to add annotations. N)rHÚhtmlr)rMrG)r#Ú markup_funcrOÚprerSÚposts rrr‰s èè€ð ð ð ˆØ”>ð ð ˆC؈IˆIˆIˆIØ�zŠz‰|Œ|ˆØˆ{˜4 Ô!1Ñ2Ô2ˆØ Ô $ð .Ø �EÔ-Ñ -ˆD؈ ˆ ˆ Ø”Oð ð ˆD؈JˆJˆJˆJð ð ð rcóÊ—t|¦«}t|¦«}t||¦«}d |¦« ¦«}t |¦«S)aŒ Do a diff of the old and new document. The documents are HTML *fragments* (str/UTF8 or unicode), they are not complete documents (i.e., no tag). Returns HTML with and tags added around the appropriate text. Markup is generally ignored, with the markup from new_html preserved, and possibly some markup from old_html (though it is considered acceptable to lose some of the old markup). Only the words in the HTML are diffed. The exception is tags, which are treated like words, and the href attribute of tags, which are noted inside the tag itself when there are changes. r)r(Úhtmldiff_tokensrrÚfixup_ins_del_tags)Úold_htmlÚnew_htmlÚold_html_tokensÚnew_html_tokensr$s rrržsV€õ"˜xÑ(Ô(€OݘxÑ(Ô(€OÝ ˜_¨oÑ >Ô >€FØ �WŠW�V‰_Œ_× "Ò "Ñ $Ô $€FÝ ˜fÑ %Ô %Ð%rcóº—t||¬¦«}| ¦«}g}|D]¡\}}}}} |dkr-| t||| …d¬¦«¦«Œ;|dks|dkr't||| …¦«} t | |¦«|dks|dkr't|||…¦«} t | |¦«Œ¢t |¦«}|S)z] Does a diff on the tokens themselves, returning a list of text chunks (not tokens). r,r/T)r/ÚinsertÚreplaceÚdelete)r0r1ÚextendÚ expand_tokensÚ merge_insertÚ merge_deleteÚcleanup_delete) Ú html1_tokensÚ html2_tokensr5r6r$r7r8r9r:r;Ú ins_tokensÚ del_tokenss rrXrXµs€õ" # \°\ÐBÑBÔB€AØ�}Š}‰Œ€HØ €FØ#+ð -ð -ш��R˜˜RØ �gÒ Ð Ø �MŠM�-¨ °R¸°UÔ(;À4ÐHÑHÔHÑ IÔ IÐ IØ Ø �hÒ Ð  '¨YÒ"6Ð"6Ý& |°B°r°EÔ':Ñ;Ô;ˆJÝ ˜ VÑ ,Ô ,Ð ,Ø �hÒ Ð  '¨YÒ"6Ð"6Ý& |°B°r°EÔ':Ñ;Ô;ˆJÝ ˜ VÑ ,Ô ,Ð ,øõ ˜FÑ #Ô #€Fà €MrFc#óÖK—|D]c}|jD]}|V—Œ|r|js<|jr| ¦«|jzV—n| ¦«V—|jD]}|V—ŒŒddS)zeGiven a list of tokens, return a generator of the chunks of text for the data in the tokens. N)rHÚhide_when_equalrMrSrG)r#r/rOrUrVs rrcrcÛs®èè€ðð ð ˆØ”>ð ð ˆC؈IˆIˆIˆIØð #˜EÔ1ð #ØÔ(ð #Ø—j’j‘l”l UÔ%>Ñ>Ð>Ð>Ð>Ð>à—j’j‘l”lÐ"Ð"Ð"Ø”Oð ð ˆD؈JˆJˆJˆJð ð ð rcó¸—t|¦«\}}}| |¦«|r+|d d¦«s|dxxdz cc<| d¦«|r.|d d¦«r|ddd…|d<| |¦«| d¦«| |¦«dS)z| doc is the already-handled document (as a list of text chunks); here we add ins_chunks to the end of that. rFú zNz )Úsplit_unbalancedrbÚendswithrJ)Ú ins_chunksrÚunbalanced_startÚbalancedÚunbalanced_ends rrdrdêsê€õ 2BÀ*Ñ1MÔ1MÑ.Ð�h Ø‡J‚JÐÑ Ô Ð Ø ð�3�r”7×#Ò# CÑ(Ô(ðð ˆBˆˆŒ�3‰ˆˆ‰Ø‡J‚JˆwÑÔÐØð)�H˜R”L×)Ò)¨#Ñ.Ô.ð)à ”| C R CÔ(ˆ�‰ ؇J‚JˆxÑÔÐØ‡J‚JˆyÑÔÐØ‡J‚Jˆ~ÑÔÐÐÐrcó—eZdZdS)Ú DEL_STARTN©Ú__name__Ú __module__Ú __qualname__rrrrvrvó€€€€€Ø€Drrvcó—eZdZdS)ÚDEL_ENDNrwrrrr}r}r{rr}có—eZdZdZdS)Ú NoDeleteszY Raised when the document no longer contains any pending deletes (DEL_START/DEL_END) N)rxryrzÚ__doc__rrrrrs€€€€€ððððrrcó˜—| t¦«| |¦«| t¦«dS)z¾ Adds the text chunks in del_chunks to the document doc (another list of text chunks) with marker to show it is a delete. cleanup_delete later resolves these markers into tags.N)rJrvrbr})Ú del_chunksrs rrere s@€ð‡J‚J�yÑÔÐØ‡J‚JˆzÑÔÐØ‡J‚J�wÑÔÐÐÐrcó*— t|¦«\}}}n#t$rYnðwxYwt|¦«\}}}t|||¦«t |||¦«|}|r+|d d¦«s|dxxdz cc<| d¦«|r.|d d¦«r|ddd…|d<| |¦«| d¦«| |¦«|}�Œ|S)a¹ Cleans up any DEL_START/DEL_END markers in the document, replacing them with . To do this while keeping the document valid, it may need to drop some tags (either start or end tags). It may also move the del into adjacent tags to try to move it to a similar location where it was originally located (e.g., moving a delete into preceding
tag, if the del looks like (DEL_START, 'Text
', DEL_END)r rFrnzNz )Ú split_deleterroÚlocate_unbalanced_startÚlocate_unbalanced_endrprJrb)ÚchunksÚ pre_deleteraÚ post_deleterrrsrtrs rrfrfsL€ðð Ý.:¸6Ñ.BÔ.BÑ +ˆJ˜    øÝð ð ð à ˆEð øøøõ 6FÀfÑ5MÔ5MÑ2И( Nõ Ð 0°*¸kÑJÔJÐJݘn¨j¸+ÑFÔFÐFØˆØ ð �s˜2”w×'Ò'¨Ñ,Ô,ð à �ˆGˆGŒG�s‰NˆGˆG‰GØ � Š �7ÑÔÐØ ð -˜ œ ×-Ò-¨cÑ2Ô2ð -à# Bœ<¨¨¨Ô,ˆH�R‰LØ � Š �8ÑÔÐØ � Š �9ÑÔÐØ � Š �;ÑÔÐØˆñ7ð8 €Ms ƒ— $£$có.—g}g}g}g}|D�]Z}| d¦«s| |¦«Œ.|ddk}| ¦«d d¦«}|tvr| |¦«Œ†|r˜|rE|dd|kr3| |¦«| ¦«\}}} | ||<ŒÏ|r8| d„|D¦«¦«g}| |¦«�Œ | |¦«�Œ | |t|¦«|f¦«| d¦«�Œ\| d „|D¦«¦«d „|D¦«}|||fS) a]Return (unbalanced_start, balanced, unbalanced_end), where each is a list of text and tag chunks. unbalanced_start is a list of all the tags that are opened, but not closed in this span. Similarly, unbalanced_end is a list of tags that are closed but were not opened. Extracting these might mean some reordering of the chunks.ú/rFcó—g|]\}}}|‘Œ Srr)rÚnameÚposÚtags rrz$split_unbalanced..Ts€ÐBÐBÐB¡n d¨C°˜cÐBÐBÐBrNcó—g|]\}}}|‘Œ Srr)rr�r�Úchunks rrz$split_unbalanced..]s€Ð1Ð1Ð1Ñ#�4˜˜eˆÐ1Ð1Ð1rcó—g|]}|®|‘ŒS©Nr)rr“s rrz$split_unbalanced..^s€ÐAÐAÐA˜%¨uÐ/@�Ð/@Ð/@Ð/@r)Ú startswithrJÚsplitrÚ empty_tagsÚpoprbr?) r‡ÚstartÚendÚ tag_stackrsr“Úendtagr�r�r‘s rroro9sÇ€ð €EØ €CØ€IØ€HØð"ñ"ˆØ×Ò Ñ$Ô$ð Ø �OŠO˜EÑ "Ô "Ð "Ø Ø�q”˜S’ˆØ�{Š{‰}Œ}˜QÔ×%Ò% eÑ,Ô,ˆØ •:Ð Ð Ø �OŠO˜EÑ "Ô "Ð "Ø Ø ð "Øð "˜Y rœ]¨1Ô-°Ò5Ð5Ø—’ Ñ&Ô&Ð&Ø!*§¢¡¤‘��c˜3Ø #�˜‘ � Øð "Ø— ’ ÐBÐB¸ ÐBÑBÔBÑCÔCÐCØ� Ø— ’ ˜5Ñ!Ô!Ð!Ñ!à— ’ ˜5Ñ!Ô!Ð!Ñ!à × Ò ˜d¥C¨¡M¤M°5Ð9Ñ :Ô :Ð :Ø �OŠO˜DÑ !Ô !Ð !Ñ !Ø ‡L‚LØ1Ð1 yÐ1Ñ1Ô1ñ3ô3ð3àAÐA 8ÐAÑAÔA€HØ �(˜CÐ ÐrcóÞ— | t¦«}n#t$rt‚wxYw| t¦«}|d|…||dz|…||dzd…fS)zæ Returns (stuff_before_DEL_START, stuff_inside_DEL_START_END, stuff_after_DEL_END). Returns the first case found (there may be more DEL_STARTs in stuff_after_DEL_END). Raises NoDeletes if there's no DEL_START found. Nr )ÚindexrvÚ ValueErrorrr})r‡r�Úpos2s rr„r„asy€ð Ø�lŠl�9Ñ%Ô%ˆˆøÝ ðððÝˆðøøøà �<Š<�Ñ Ô €DØ �$�3�$Œ<˜  A¡ d  Ô+¨V°D¸±F°G°G¬_Ð <Ð| d¦«| | d¦«¦«nd S�Œ) a° pre_delete and post_delete implicitly point to a place in the document (where the two were split). This moves that point (by popping items from one and pushing them onto the other). It moves the point to try to find a place where unbalanced_start applies. As an example:: >>> unbalanced_start = ['
'] >>> doc = ['

', 'Text', '

', '
', 'More Text', '
'] >>> pre, post = doc[:3], doc[3:] >>> pre, post (['

', 'Text', '

'], ['
', 'More Text', '
']) >>> locate_unbalanced_start(unbalanced_start, pre, post) >>> pre, post (['

', 'Text', '

', '
'], ['More Text', '
']) As you can see, we moved the point so that the dangling
that we found will be effectively replaced by the div in the original document. If this doesn't work out, we just throw away unbalanced_start without doing anything. r rz<>r‹rŒÚinsÚdelzUnexpected delete tag: %rN)r—rrvr–r™rJ)rrrˆr‰ÚfindingÚ finding_nameÚnextr�s rr…r…ms€ð,Øð à ˆEØ" 1Ô%ˆØ—}’}‘” qÔ)×/Ò/°Ñ5Ô5ˆ Øð Ø ˆEؘ1Œ~ˆØ •9Ð Ð  D§O¢O°CÑ$8Ô$8Ð à ˆEØ �Œ7�cŠ>ˆ>à ˆEØ�zŠz‰|Œ|˜AŒ×$Ò$ TÑ*Ô*ˆØ �5Š=ˆ=à ˆEØ�uŠ}ˆ}ˆ}Ø '¨$Ñ .ñŒ}ˆ}à �<Ò Ð Ø × Ò  Ñ #Ô #Ð #Ø × Ò ˜kŸošo¨aÑ0Ô0Ñ 1Ô 1Ð 1Ð 1ð ˆEñ5rcóЗ |sdS|d}| ¦«d d¦«}|sdS|d}|tus| d¦«sdS| ¦«d d¦«}|dks|dkrdS||kr=| ¦«| d| ¦«¦«ndSŒæ) zt like locate_unbalanced_start, except handling end tags and possibly moving the point earlier in the document. r rFrr�ú tag, which takes up visible space just like a word but is only represented in a document by a tag. Nrcó‚—t |t›d|›�|||¬¦«}||_||_||_|S)Nz: rL)rOr¬rNr‘ÚdataÚ html_repr)r­r‘rºr»rHrGrMr®s rr¬ztag_token.__new__èsP€å�mŠm˜C­T¨T¨T°4°4Ð!8Ø%-Ø&/Ø0CðñEôEˆðˆŒØˆŒØ!ˆŒ ؈ rc óh—d|j›d|j›d|j›d|j›d|j›d|j›d� S)Nz tag_token(r°z , html_repr=z , post_tags=z , pre_tags=z, trailing_whitespace=r±)r‘rºr»rHrGrMr³s rr²ztag_token.__repr__ósG€€à ŒHˆHˆHØ ŒIˆIˆIØ ŒNˆNˆNØ ŒMˆMˆMØ ŒNˆNˆNØ Ô $Ð $Ð $ð &ð &rcó—|jSr•)r»r³s rrSztag_token.htmlûs €ØŒ~Ðrr¶)rxryrzr€r¬r²rSrrrr¸r¸âsX€€€€€ð5ð5ð59Ø46ð ð ð ð ð&ð&ð&ðððððrr¸có—eZdZdZdZd„ZdS)Ú href_tokenzh Represents the href in an anchor tag. Unlike other words, we only show the href when it changes. Tcó —d|zS)Nz Link: %srr³s rrSzhref_token.htmls €Ø˜TÑ!Ð!rN)rxryrzr€rlrSrrrr¿r¿þs4€€€€€ð(ð(ð€Oð"ð"ð"ð"ð"rr¿Tcó”—tj|¦«r|}nt|d¬¦«}t|d|¬¦«}t |¦«S)ak Parse the given HTML and returns token objects (words with attached tags). This parses only the content of a page; anything in the head is ignored, and the and elements are themselves optional. The content is then parsed by lxml, which ensures the validity of the resulting parsed document (though lxml may make incorrect guesses when the markup is particular bad). and tags are also eliminated from the document, as that gets confusing. If include_hrefs is true, then the href attribute of tags is included as a special kind of diffable token.T©Úcleanup)Úskip_tagr')rÚ iselementÚ parse_htmlÚ flatten_elÚ fixup_chunks)rSr'Úbody_elr‡s rr(r(sQ€õ „�tÑÔð1؈ˆå˜T¨4Ð0Ñ0Ô0ˆå ˜¨$¸mÐ LÑ LÔ L€Få ˜Ñ Ô ÐrcóF—|rt|¦«}t|d¬¦«S)a Parses an HTML fragment, returning an lxml element. Note that the HTML will be wrapped in a
tag that was not in the original document. If cleanup is true, make sure there's no or , and get rid of any and tags. T)Ú create_parent)Ú cleanup_htmlr)rSrÃs rrÆrÆ s,€ðð"å˜DÑ!Ô!ˆÝ ˜t°4Ð 8Ñ 8Ô 8Ð8rz z zcó—t |¦«}|r|| ¦«d…}t |¦«}|r|d| ¦«…}t  d|¦«}|S)z³ This 'cleans' the HTML, meaning that any page structure is removed (only the contents of are used, if there is any and tags are removed. Nr)Ú_body_reÚsearchr›Ú _end_body_reršÚ _ins_del_reÚsub)rSÚmatchs rrÌrÌ1s|€õ �OŠO˜DÑ !Ô !€EØ ð"Ø�E—I’I‘K”K�L�LÔ!ˆÝ × Ò  Ñ %Ô %€EØ ð$Ø�N�U—[’[‘]”]�NÔ#ˆÝ �?Š?˜2˜tÑ $Ô $€DØ €Krz [ \t\n\r]$cól—t| ¦«¦«}|d|…||d…fS)zP This function takes a word, such as 'test ' and returns ('test',' ') rN)r?Úrstrip)ÚwordÚstripped_lengths rÚsplit_trailing_whitespacerØAs9€õ˜$Ÿ+š+™-œ-Ñ(Ô(€OØ ��/Ð!Ô " D¨Ð)9Ð)9Ô$:Ð :Ð:rc óx—g}d}g}|D�]{}t|t¦«r–|ddkrL|d}t|d¦«\}}td||||¬¦«}g}| |¦«n=|ddkr1|d}t ||d¬ ¦«}g}| |¦«Œ®t |¦«rŒ>ð Ý)BÀ5Ñ)IÔ)IÑ &ˆEÐ&ݘU¨YÐL_Ð`Ñ`Ô`ˆH؈IØ �MŠM˜(Ñ #Ô #Ð #Ð #å ˜%Ñ Ô ð Ø × Ò ˜UÑ #Ô #Ð #Ñ #å ˜Ñ Ô ð Øð 1Ø× Ò  Ñ'Ô'Ð'Ñ'àð9ð9ð9à�x�x   ¨¨¨°°ð8ñ9ô9�xðÔ"×)Ò)¨%Ñ0Ô0Ð0Ñ0à �5à ð/Ý�b 9Ð-Ñ-Ô-Ð.Ð.àˆrŒ Ô×#Ò# IÑ.Ô.Ð.à €Mr) ÚparamrÚÚareaÚbrÚbasefontÚinputÚbaseÚmetaÚlinkÚcol)ÚaddressÚ blockquoteÚcenterÚdirÚdivÚdlÚfieldsetÚformÚh1Úh2Úh3Úh4Úh5Úh6ÚhrÚisindexÚmenuÚnoframesÚnoscriptÚolÚprUÚtableÚul) ÚddÚdtÚframesetÚliÚtbodyÚtdÚtfootÚthÚtheadÚtrc#órK—|sD|jdkr(d| d¦«t|¦«fV—nt|¦«V—|jtvr|jst |¦«s |jsdSt|j¦«}|D]}t|¦«V—Œ|D]}t||¬¦«D]}|V—ŒŒ|jdkr0| d¦«r|rd| d¦«fV—|s;t|¦«V—t|j¦«}|D]}t|¦«V—ŒdSdS)a Takes an lxml element el, and generates all the text chunks for that tag. Each start tag is a chunk, each word is a chunk, and each end tag is a chunk. If skip_tag is true, then the outermost container tag is not returned (just its contents).rÚrANr&r-rÜ) r‘ÚgetÚ start_tagr˜r r?ÚtailÚ split_wordsr rÇÚend_tag)Úelr'rÄÚ start_wordsrÖÚchildÚitemÚ end_wordss rrÇrǬs‰èè€ð ð Ø Œ6�UŠ?ˆ?ؘ"Ÿ&š& ™-œ-­°2©¬Ð7Ð 7Ð 7Ð 7Ð 7å˜B‘-”-Ð Ð Ð Ø „v•ÐРB¤GеC¸±G´GÐÀBÄGÐØˆÝ˜bœgÑ&Ô&€KØð ð ˆÝ˜$ÑÔÐÐÐÐØððˆÝ˜u°MÐBÑBÔBð ð ˆD؈JˆJˆJˆJð à „v�‚}€}˜Ÿš ™œ€}¨M€}Ø�r—v’v˜f‘~”~Ð&Ð&Ð&Ð&Ø ð$Ý�b‰kŒkÐÐÐÝ ¤Ñ(Ô(ˆ Øð $ð $ˆDݘdÑ#Ô#Ð #Ð #Ð #Ð #ð $ð$ð $ð $rz \S+(?:\s+|$)cój—|r| ¦«sgSt |¦«}|S)z_ Splits some text into words. Includes trailing whitespace on each word when appropriate. )rÚsplit_words_reÚfindall)r Úwordss rrrÊs8€ð ð�t—z’z‘|”|ðØˆ å × "Ò " 4Ñ (Ô (€EØ €Lrz ^[ \t\n\r]có„—d|j›d d„|j ¦«D¦«¦«›d�S)z= The text representation of the start tag for a tag. r‹rc óB—g|]\}}d|›dt|d¦«›d�‘ŒS)rnz="Tú")r )rr�Úvalues rrzstart_tag..ÚsG€ð?ð?ð?Ù(˜T 5 5ð(, t t­[¸ÀÑ-EÔ-EÐ-EÐ-EÐFð?ð?ð?rú>)r‘rÚattribÚitems)rs rrrÕs_€€ð Œˆ�—’ð?ð?Ø,.¬I¯OªOÑ,=Ô,=ð?ñ?ô?ñ@ô@ð@ð@ð AðArcór—|jr"t |j¦«rd}nd}d|j›d|›�S)zg The text representation of an end tag for a tag. Includes trailing whitespace when appropriate. rnrr©r!)rÚstart_whitespace_rerÏr‘)rÚextras rrrÝsF€ð „wðÕ&×-Ò-¨b¬gÑ6Ô6ðØˆˆàˆøØœ˜˜  Ð &Ð&rcó.—| d¦« S)Nr‹©r–©r*s rrßrßæs€Ø�~Š~˜cÑ"Ô"Ð "Ð"rcó,—| d¦«S)Nr©r(r)s rráráés€Ø �>Š>˜$Ñ Ô ÐrcóX—| d¦«o| d¦« S)Nr‹r©r(r)s rràràìs(€Ø �>Š>˜#Ñ Ô Ð ; s§~¢~°dÑ';Ô';Ð#;Ð;rcóh—t|d¬¦«}t|¦«t|d¬¦«}|S)z  Given an html string, move any or tags inside of any block-level elements, e.g. transform

word

to

word

FrÂT)Ú skip_outer)rÆÚ_fixup_ins_del_tagsÚserialize_html_fragment)rSrs rrYrYïs;€õ �T 5Ð )Ñ )Ô )€CݘÑÔÐÝ " 3°4Ð 8Ñ 8Ô 8€DØ €Krcó(—t|t¦«r Jd|z¦«‚tj|dt¬¦«}|rQ|| d¦«dzd…}|d| d¦«…}| ¦«S|S)z¨ Serialize a single lxml element as HTML. The serialized form includes the elements tail. If skip_outer is true, then don't serialize the outermost tag z3You should pass in an element, not a string like %rrS)ÚmethodÚencodingr!r Nr‹)rÝÚ basestringrÚtostringr ÚfindÚrfindr)rr-rSs rr/r/øsž€õ ˜"�jÑ)Ô)ðDðDØ=ÀÑBñDôDÐ )å Œ>˜" VµhÐ ?Ñ ?Ô ?€DØðà�D—I’I˜c‘N”N 1Ñ$Ð%Ð%Ô&ˆàÐ$�T—Z’Z ‘_”_Ð$Ô%ˆØ�zŠz‰|Œ|Ðàˆ rcó°—dD]R}| d|z¦«D]7}t|¦«sŒt||¬¦«| ¦«Œ8ŒSdS)z?fixup_ins_del_tags that works on an lxml document in-place )r£r¤zdescendant-or-self::%s)r‘N)ÚxpathÚ_contains_block_level_tagÚ_move_el_inside_blockÚdrop_tag)rr‘rs rr.r. sx€ðððˆØ—)’)Ð4°sÑ:Ñ;Ô;ð ð ˆBÝ,¨RÑ0Ô0ð ØÝ ! "¨#Ð .Ñ .Ô .Ð .Ø �KŠK‰MŒMˆMˆMð  ððrcóp—|jtvs|jtvrdS|D]}t|¦«rdSŒdS)zPTrue if the element contains any block-level elements, like

, , etc. TF)r‘Úblock_level_tagsÚblock_level_container_tagsr9)rrs rr9r9sT€ð „vÕ!Ð!Ð! R¤VÕ/IÐ%IÐ%I؈tØððˆÝ $ UÑ +Ô +ð Ø�4�4ð à ˆ5rcóú—|D]}t|¦«rnTŒtj|¦«}|j|_d|_| t |¦«¦«|g|dd…<dSt |¦«D]»}t|¦«rkt ||¦«|jrStj|¦«}|j|_d|_| |  |¦«dz|¦«Œ|tj|¦«}|  ||¦«|  |¦«Œ¼|jr?tj|¦«}|j|_d|_| d|¦«dSdS)zt helper for _fixup_ins_del_tags; actually takes the etc tags and moves them inside any block-level tags. Nr r) r9rÚElementr rbÚlistr:rr_rŸr`rJ)rr‘rÚ children_tagÚtail_tagÚ child_tagÚtext_tags rr:r:s†€ðð ð ˆÝ $ UÑ +Ô +ð Ø ˆEð õ”} SÑ)Ô)ˆ ØœGˆ ÔØˆŒØ×Ò�D ™HœHÑ%Ô%Ð%Ø�ˆˆ1ˆ1ˆ1‰ØˆÝ�b‘”ð $ð $ˆÝ $ UÑ +Ô +ð $Ý ! %¨Ñ -Ô -Ð -ØŒzð 7Ý œ=¨Ñ-Ô-�Ø %¤ �” Ø!�” Ø— ’ ˜"Ÿ(š( 5™/œ/¨!Ñ+¨XÑ6Ô6Ð6øåœ  cÑ*Ô*ˆIØ �JŠJ�u˜iÑ (Ô (Ð (Ø × Ò ˜UÑ #Ô #Ð #Ð #Ø „wðÝ”= Ñ%Ô%ˆØœˆŒ ؈ŒØ � Š �!�XÑÔÐÐÐð ðrcó—| ¦«}|jpd}|jrUt|¦«s ||jz }n;|djr|dxj|jz c_n|j|d_| |¦«}|rU|dkrd}n ||dz }|€ |jr|xj|z c_n'||_n|jr|xj|z c_n||_| ¦«|||dz…<dS)z© Removes an element, but merges its contents into its place, e.g., given

Hi there!

, if you remove the element you get

Hi there!

rrFrNr )Ú getparentr rr?rŸÚ getchildren)rÚparentr rŸÚpreviouss rÚ_merge_element_contentsrK?s€ð �\Š\‰^Œ^€FØ Œ7ˆ=�b€DØ „wð&Ý�2‰wŒwð &Ø �B”G‰OˆDˆDà�"ŒvŒ{ð &Ø�2”� ” ˜rœwÑ&� ” � à œg��2”” Ø �LŠL˜Ñ Ô €EØ ð%Ø �AŠ:ˆ:؈HˆHà˜e A™g”ˆHØ Ð ØŒ{ð #Ø� ” ˜tÑ#� ” � à"�” � àŒ}ð %Ø� ”  Ñ%� ” � à $�” ØŸNšNÑ,Ô,€Fˆ5��q‘ˆ=ÑÐÐrcó—eZdZdZdZd„ZdS)r0zt Acts like SequenceMatcher, but tries not to find very small equal blocks amidst large spans of changes rÛcóö‡—tt|j¦«t|j¦«¦«}t|j|dz ¦«Štj |¦«}ˆfd„|D¦«S)Nécó<•—g|]}|d‰ks|d°|‘ŒS)rÛr)rrÚ thresholds €rrzBInsensitiveSequenceMatcher.get_matching_blocks..ms=ø€ð ð ð ˜Ø˜”7˜YÒ&Ð&ؘA”wð'ðØ&Ð&Ð&r)Úminr?r.rPÚdifflibÚSequenceMatcherÚget_matching_blocks)r´ÚsizeÚactualrPs @rrTz.InsensitiveSequenceMatcher.get_matching_blocksiswø€Ý•3�t”v‘;”;¥ D¤F¡ ¤ Ñ,Ô,ˆÝ˜œ¨¨q©Ñ1Ô1ˆ ÝÔ(×<Ò<¸TÑBÔBˆð ð ð ð  ð ñ ô ð rN)rxryrzr€rPrTrrrr0r0as4€€€€€ððð €Ið ð ð ð ð rr0Ú__main__)Ú _diffcommand)F)T)PÚ __future__rrRÚlxmlrÚ lxml.htmlrÚreÚ__all__rSrr Ú ImportErrorÚcgiÚunicoder Ú NameErrorÚstrr3rrrrr2rrIrrrXrcrdrvr}Ú Exceptionrrerfror„r…r†rOr¸r¿r(rÆÚcompileÚIÚSrÎrÐrÑrÌÚend_whitespace_rerØrÈr˜r=r>rÇÚUrrr%rrrßráràrYr/r.r9r:rKrSr0rxrXÚmainrrrúrjs^ðð'Ð&Ð&Ð&Ð&Ð&à€€€ØÐÐÐÐÐØ)Ð)Ð)Ð)Ð)Ð)Ø € € € à ˜JÐ '€ð*Ø*Ð*Ð*Ð*Ð*Ð*Ð*øØð*ð*ð*Ø)Ð)Ð)Ð)Ð)Ð)Ð)Ð)ð*øøøðØ€H€HøØðððà€H€H€HðøøøðØ€J€JøØðððà€J€J€Jðøøøð1ð1ð1ð#1ð&#ð&#ð&#ð&#ðPððð -ð -ð -ð1ð1ð1ð ð ð ðððð$ððð*&ð&ð&ð.$ð$ð$ðL ð ð ð ðððð. ð ð ð ð ñ ô ð ð ð ð ð ð ñ ô ð ððððð� ñôððððð%ð%ð%ðN& ð& ð& ðP =ð =ð =ð0ð0ð0ðdððð4'ð'ð'ð'ð'ˆHñ'ô'ð'ðRðððð�ñôðð8"ð"ð"ð"ð"�ñ"ô"ð"ð ð ð ð ð0 9ð 9ð 9ð 9ð ˆ2Œ:�l B¤D¨¬¡IÑ .Ô .€ØˆrŒz˜-¨¬¨b¬d©Ñ3Ô3€ ؈bŒjÐ,¨b¬d°2´4©iÑ8Ô8€ ð ð ð ð�B”J˜}Ñ-Ô-Ðð;ð;ð;ð2ð2ð2ðl#€ ðÐð6 Ðð$ð$ð$ð$ð8�”˜O¨R¬TÑ2Ô2€ðððð!�b”j Ñ/Ô/ÐðAðAðAð'ð'ð'ð#ð#ð#ð ð ð ð<ð<ð<ðððððððð$ððððððððð@ -ð -ð -ðD ð ð ð ð  Ô!8ñ ô ð ð  ˆzÒÐØ&Ð&Ð&Ð&Ð&Ð&Ø€LÔÑÔÐÐÐðÐs- '§ 5´5¹<¼AÁAÁ A Á AÁA