Visitar URL original
feature: Paragraph.text includes hyperlink text · Issue #85 · python-openxml/python-docx · GitHub
Skip to content

feature: Paragraph.text includes hyperlink text #85

Description

@SebasSBM

When getting Document.paragraphs objects, their text method doesn't include hyperlinks in the output. The one with this problem posted his question here:

http://stackoverflow.com/questions/25228106/how-to-extract-text-from-an-existing-docx-file-using-python-docx/25228787#25228787

I've been reading the documentation of python-docx for several hours and didn't find any property or method useful to resolve this. Maybe some class should be created, or some methods should be appended to an existing class to achieve this.

I barely know something about python-docx API. I knew of it's existence trying to help some people in stackoverflow.com with their problems. I don't even know how Windows' DOCX format works (I tried to open it with an hexadecimal editor to try to figure it out, and I just don't get it :P). But I'm skilled with logic problems and I've got good skills with Python scripting. I'd like to help if there's something I can do.

Activity

  1. scanny commented on Aug 10, 2014

    @scanny
    Contributor

    Hi Sebastian,

    A .docx file is a ZIP archive, so it will make a lot more sense to you once you've unzipped it. If you're on a Mac or Linux this will be a good first step:

    $ unzip -l some_document.docx

    Contributors are always welcome, although I expect this project would be a steep learning curve for you. There is a related feature request here #74 you can take a look at and we can talk about what the API might look like for reading and writing hyperlinks if you're still interested.

  2. SebasSBM commented on Aug 16, 2014

    @SebasSBM
    Author

    Thanks for your reply, scanny. I've just read it right now and started researching. The command you posted revealed an inner file structure, so I used the Ubuntu's tool for compressed files and noticed that all of them are XML files. I'm used to XML through I made apps for Android and I used some SOAP webservices, so XML is not new for me. Althrough, you have a point about it may become a steep learning curve, through the XML structure seems to be quite complex.

    Anyway, analyzing it I figured out some things in just less than half an hour: it seems that styles are defined in "styles.xml". There is also a file for the fonts, I just don't get why it seems there are 4 fonts in a test.docx file I created with LibreOffice in which I just used the default font and a hyperlink (which it seems it has it's own style defined), but I don't think extra fonts are relevant for now.

    I've taken a look to the "document.xml" file and noticed a difference between a normal paragraph and a hyperlink paragraph: this would be a normal paragraph structure:

    <w:p>
        <w:pPr>
            <w:pStyle w:val="style0"/>
        </w:pPr>
        <w:r>
            <w:rPr/>
            <w:t>Prueba jajajjajajajaja</w:t>
        </w:r>
    </w:p>
    

    On the other hand, this would be a hyperlink paragraph:

    <w:p>
        <w:pPr>
            <w:pStyle w:val="style0"/>
        </w:pPr>
        <w:hyperlink r:id="rId2">
            <w:r>
                <w:rPr>
                    <w:rStyle w:val="style15"/>
                </w:rPr>
                <w:t>http://www.google.com/</w:t>
            </w:r>
        </w:hyperlink>
    </w:p>
    

    In other words, it seems that the tags <w:hyperlink></w:hyperlink> contain the whole rich text structure that is supposed to be the hyperlink, with an id which would point to the actual URL stored somewhere in the XML file system, I guess. It seems quite interesting, unfortunately, I don't have much spare time lately, because I'm very busy with web developing.

    Anyways, if I ever have some spare time, I'd like to research how your python API reads the paragraphs, and make their .text() method able to recognize the <w:hyperlink> tag as text container. I'll keep you informed if I make any relevant progress.

  3. SebasSBM commented on Aug 16, 2014

    @SebasSBM
    Author

    I think here's the problem: check the class CT_P at master/docx/oxml/text.py . If you take a look at the initial variables (lines 36 and 37) it seems this class (which I suppose it handles <w:p> objects) doesn't handle <w:hyperlink> objects at all. I think that's why text inside hyperlinks are not returned in the text property. I don't know much about the structure of the whole project -not yet-, but I think this is the way to go to resolve the problem.

  4. nagamiki commented on Sep 18, 2014

    @nagamiki

    This problem is very crucial for me, I hope the problem could be solved asap in the coming version.

  5. Brad-Python commented on Nov 9, 2014

    @Brad-Python

    Hi,

    As indicated by SebasSBM above, the difference between a hyperlink and a paragraph is the [w:hyperlink r:id="XXX"] and [/w:hyperlink] tags. A workaround consists in removing those tags from the document's xml code, so that only the code of a standard paragraph remains (easy to do using regular expressions and the re module)

    Here is how I did this:
    I edited C:\Python27\Lib\site-packages\docx\oxml\__init__.py as follows:

    1/ I created a new function that remove hyperlinks by nothing in a xml text:

    def remove_hyperlink_tags(xml):
        import re
        xml = xml.replace("</w:hyperlink>","")
        xml = re.sub('<w:hyperlink[^>]*>',"",xml)
        return xml
    

    2/ I updated the standard parse_xml function as follows:

    def parse_xml(xml):
        """
        """
        root_element = etree.fromstring(remove_hyperlink_tags(xml), oxml_parser)
        return root_element
    

    It worked well for me but I didn't test it much so use at your own risk...

  6. mikkleini commented on Jan 7, 2015

    @mikkleini

    Hi,
    python-docx is a great tool that i just discovered. However, i also faced the missing hyperlink text problem right away.
    Please solve it.

    remove_hyperlink_tags don't work because i get:
    File "C:\Python34\lib\site-packages\docx\oxml__init__.py", line 23, in remove_hyperlink_tags
    xml = xml.replace("/w:hyperlink","")
    TypeError: expected bytes, bytearray or buffer compatible object

  7. mikkleini commented on Jan 7, 2015

    @mikkleini

    Ok, i think i faced the Python 2/3 issue. Here's something which works on Python 3.4:

    def remove_hyperlink_tags(xml):
        import re
        text = xml.decode('utf-8')
        text = text.replace("</w:hyperlink>","")
        text = re.sub('<w:hyperlink[^>]*>', "", text)
        return text.encode('utf-8')
  8. changed the title [-]Getting hyperlinks from paragraphs[/-] [+]feature: Paragraph.text includes hyperlink text[/+] on Feb 13, 2015
  9. funkycode commented on Mar 26, 2015

    @funkycode

    any schedule for that one?

  10. herrera78 commented on Feb 5, 2017

    @herrera78

    I used the code at https://gist.github.com/etienned/7539105#file-extractdocx-py and supplied the path and get hyperlink texts as I need

  11. 2 remaining items

  12. Yakabuff commented on Oct 13, 2020

    @Yakabuff

    @scanny is there any chance you can implement any of these solutions

  13. oroygit commented on Sep 10, 2021

    @oroygit

    Until this functionality is implemented in python-docx, this is the workaround I used to redefine the 'text' property of the docx.text.paragraph.Paragraph class such that it includes hyperlinks.

    Necessary imports:

    from docx.text.paragraph import Paragraph
    import re
    

    First I redefine the text property with:

    Paragraph.text = property(lambda self: GetParagraphText(self))
    

    Every time paragraph.text is called, the function GetParagraphText will be called instead with the instance paragraph of type docx.text.paragraph.Paragraph as parameter.

    The function GetParagraphText is implemented as:

    def GetParagraphText(paragraph)
    
        def GetTag(element):
            return "%s:%s" % (element.prefix, re.match("{.*}(.*)", element.tag).group(1))
    
        text = ''
        runCount = 0
        for child in paragraph._p:
            tag = GetTag(child)
            if tag == "w:r":
                text += paragraph.runs[runCount].text
                runCount += 1
            if tag == "w:hyperlink":
                for subChild in child:
                    if GetTag(subChild) == "w:r":
                        text += subChild.text
        return text
    

    The above implementationis the least intrusive I could think of. It requires no modification to the rest of the code, does not need to create new objects, and re-uses when possible the logic already available in python-docx (e.g., paragraph.runs). Hopefully this will help!

  14. scanny commented on Sep 14, 2021

    @scanny
    Contributor

    Nice job @roydesbois :)

  15. tomking2 commented on Jan 11, 2022

    @tomking2

    Following on from @roydesbois (thanks for the inspiration, a great solution), I needed the requirement to parse down to each individual run so styling and fonts an also be correctly interpreted.

    This also uses the built in qn to get the full qualified name of the tag, thus not requiring custom regex parsing. What's great about this is that calling .text on the paragraph still works as you'd expect, while allowing you to iterate through each run and grab its respective text/styles.

    from docx.oxml.shared import qn
    
    def GetParagraphRuns(paragraph):
        def _get(node, parent):
            for child in node:
                if child.tag == qn('w:r'):
                    yield Run(child, parent)
                if child.tag == qn('w:hyperlink'):
                    yield from _get(child, parent)
        return list(_get(paragraph._element, paragraph))
    
    Paragraph.runs = property(lambda self: GetParagraphRuns(self))
  16. JStooke commented on Feb 16, 2023

    @JStooke

    @tomking2's solution was great but in my case a lot of the links were embedded. "Click Here" was the text returned rather than the link "Click Here" related to. So I tweaked the approach to pull out the link from document.part.rels and added it to the child.text if it differed from the text (i.e. embedded) so "Click Here" becomes "Click Here"[https://www.google.com]

    from docx.text.paragraph import Paragraph
    from docx.text.run import Run
    from docx.oxml.shared import qn
    
    def GetParagraphRuns(paragraph):
        def _get(node, parent, hyperlinkId=None):
            for child in node:
                if child.tag == qn('w:r'):
                    if hyperlinkId:
                        linkToAdd = document.part.rels[hyperlinkId]._target
                        if child.text != linkToAdd:
                            child.text = child.text + f'[{linkToAdd}]'
                    yield Run(child, parent)
                if child.tag == qn('w:hyperlink'):
                    hlid = child.attrib.get(qn('r:id'))
                    yield from _get(child, parent, hlid)
        return list(_get(paragraph._element, paragraph))
    
    Paragraph.runs = property(lambda self: GetParagraphRuns(self))

    I've no doubt there is a better way to do this but it's worked a treat for me :)

  17. oliveslongjohns commented on Mar 3, 2023

    @oliveslongjohns

    Absolutely love JStooke's solution, but noticed a bug (with a super quick fix).

    if child.text != linkToAdd: should be if linkToAdd not in child.text, as the hyperlink will never equal the text, and every time the method is called, it adds the hyperlink an additional time. That makes the full solution:

    from docx.text.paragraph import Paragraph
    from docx.text.run import Run
    from docx.oxml.shared import qn
    
    def GetParagraphRuns(paragraph):
        def _get(node, parent, hyperlinkId=None):
            for child in node:
                if child.tag == qn('w:r'):
                    if hyperlinkId:
                        linkToAdd = document.part.rels[hyperlinkId]._target
                        if linkToAdd not in child.text:
                            child.text = child.text + f'[{linkToAdd}]'
                    yield Run(child, parent)
                if child.tag == qn('w:hyperlink'):
                    hlid = child.attrib.get(qn('r:id'))
                    yield from _get(child, parent, hlid)
        return list(_get(paragraph._element, paragraph))
    
    Paragraph.runs = property(lambda self: GetParagraphRuns(self))

    Also, the document variable is undefined within the method. I solved this by declaring it globally, e.g.

    global document
    document = Document(file)
  18. JStooke commented on Mar 9, 2023

    @JStooke

    Awesome cheers @oliveslongjohns. I'd forgotten to mention i'd already defined the global variable earlier in my use case. Glad its working for you :D

  19. ecestebanjek commented on Jun 6, 2023

    @ecestebanjek

    @oliveslongjohns @JStooke thanks for absolutely great answer, it works very well. However, I'd like to ask if there is a way to adapt code to include in the paragraph.text return the hyperlink working, instead of text_link [ link ].
    I found a function to create a hyperlink, but Idon't really know how to embed:

    `def add_hyperlink(paragraph, url, text, color, underline):

    # This gets access to the document.xml.rels file and gets a new relation id value
    part = paragraph.part
    r_id = part.relate_to(url, docx.opc.constants.RELATIONSHIP_TYPE.HYPERLINK, is_external=True)
    
    # Create the w:hyperlink tag and add needed values
    hyperlink = docx.oxml.shared.OxmlElement('w:hyperlink')
    hyperlink.set(docx.oxml.shared.qn('r:id'), r_id, )
    
    # Create a w:r element
    new_run = docx.oxml.shared.OxmlElement('w:r')
    
    # Create a new w:rPr element
    rPr = docx.oxml.shared.OxmlElement('w:rPr')
    
    # Add color if it is given
    if not color is None:
      c = docx.oxml.shared.OxmlElement('w:color')
      c.set(docx.oxml.shared.qn('w:val'), color)
      rPr.append(c)
    
    # Remove underlining if it is requested
    if not underline:
      u = docx.oxml.shared.OxmlElement('w:u')
      u.set(docx.oxml.shared.qn('w:val'), 'none')
      rPr.append(u)
    
    # Join all the xml elements together add add the required text to the w:r element
    new_run.append(rPr)
    new_run.text = text
    hyperlink.append(new_run)
    
    paragraph._p.append(hyperlink)
    
    return hyperlink`
    
  20. added
    misfeatureDoesn't behave how one might expect
    inner-contentMethods to access all content inside doc, para, run, etc.
    hyperlinkRead and write hyperlinks in paragraph
    on Sep 24, 2023
  21. scanny commented on Oct 2, 2023

    @scanny
    Contributor

    Added in v1.0.0 circa Oct 5, 2023.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    hyperlinkRead and write hyperlinks in paragraphinner-contentMethods to access all content inside doc, para, run, etc.misfeatureDoesn't behave how one might expectshortlisttext

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions