Visitar URL original
翻译扫描档存在重影 / feat (main): supports ocr on scanned document · Issue #19 · PDFMathTranslate/PDFMathTranslate · GitHub
Skip to content

翻译扫描档存在重影 / feat (main): supports ocr on scanned document #19

Description

@jackiehejian

翻译错误

当pdf文件均为图像,而不是可编辑(复制)状态时,翻译完全失败,具体见图

Activity

  1. Byaidu commented on Nov 7, 2024

    @Byaidu
    Member

    图片型的 PDF 文档暂时还没办法翻译,目前主要还是在优化电子书和论文的翻译效果

  2. jackiehejian commented on Nov 7, 2024

    @jackiehejian
    Author

    图片型的 PDF 文档暂时还没办法翻译,目前主要还是在优化电子书和论文的翻译效果

    好的,非常感谢

  3. fireinrain commented on Nov 8, 2024

    @fireinrain

    均为图像有点为难人了,ocr的质量 影响文字的质量 影响翻译的效果

  4. xxsunyxx commented on Nov 19, 2024

    @xxsunyxx

    加一个可选流程paddleOCR,

  5. xxnuo commented on Nov 20, 2024

    @xxnuo
    Contributor

    sayura
    这个模型非常准确,就是对算力的要求会高于 Paddle OCR

  6. Byaidu commented on Nov 20, 2024

    @Byaidu
    Member

    sayura 这个模型非常准确,就是对算力的要求会高于 Paddle OCR

    和 minerU/marker 比较怎么样呀

  7. xxnuo commented on Nov 21, 2024

    @xxnuo
    Contributor

    Owner

    sayura 就是 marker 的作者做的开源多国语言和表格的 OCR 模型😂
    minerU 这个我没有测试,我只测试了 PaddleOCR 高精度模型,Sayura 效果比它好很多,而且支持多国语言效果很好。
    我看 minerU 的 issue,对多国语言的支持好像不佳
    缺点就是 Sayura 对 GPU 显存要求有点高,头疼,不太会量化模型。

  8. changed the title [-]当PDF每一页均为图像时,无法进行翻译[/-] [+]feat (main): supports ocr on scanned document[/+] on Nov 21, 2024
  9. xxnuo commented on Dec 2, 2024

    @xxnuo
    Contributor

    佬们 ocr 的进展如何,我觉得用 paddleocr 撸一个不错,如果已经有佬在做了我就不再造轮子了 @reycn @Byaidu

  10. Byaidu commented on Dec 2, 2024

    @Byaidu
    Member

    佬们 ocr 的进展如何,我觉得用 paddleocr 撸一个不错,如果已经有佬在做了我就不再造轮子了 @reycn @Byaidu

    目前还一点没做…

    如果写好了的话欢迎来贡献代码

  11. hellofinch commented on Dec 6, 2024

    @hellofinch
    Collaborator
    from typing import BinaryIO
    import numpy as np
    import tqdm
    from pymupdf import Document
    from pdfminer.pdfpage import PDFPage
    from pdfminer.pdfinterp import PDFResourceManager
    from pdfminer.pdfdocument import PDFDocument
    from pdfminer.pdfparser import PDFParser
    from pdf2zh.converter import TranslateConverter
    from pdf2zh.pdfinterp import PDFPageInterpreterEx
    from pymupdf import Font
    import numpy as np
    from paddleocr import PaddleOCR
    
    file=""
    
    def extract_text_to_fp(
        inf: BinaryIO,
        pages=None,
        password: str = "",
        debug: bool = False,
        page_count: int = 0,
        vfont: str = "",
        vchar: str = "",
        thread: int = 0,
        doc_en: Document = None,
        model=None,
        lang_in: str = "",
        lang_out: str = "",
        service: str = "",
        resfont: str = "",
        noto: Font = None,
        callback: object = None,
        **kwarg,
    ) -> None:
        ocr = PaddleOCR(use_angle_cls=True, lang="en")
        rsrcmgr = PDFResourceManager()
        layout = {}
        device = TranslateConverter(
            rsrcmgr, vfont, vchar, thread, layout, lang_in, lang_out, service, resfont, noto
        )
    
        assert device is not None
        obj_patch = {}
        interpreter = PDFPageInterpreterEx(rsrcmgr, device, obj_patch)
        if pages:
            total_pages = len(pages)
        else:
            total_pages = page_count
    
        parser = PDFParser(inf)
        doc = PDFDocument(parser, password=password)
        with tqdm.tqdm(
            enumerate(PDFPage.create_pages(doc)),
            total=total_pages,
        ) as progress:
            for pageno, page in progress:
                if pages and (pageno not in pages):
                    continue
                if callback:
                    callback(progress)
                page.pageno = pageno
                pix = doc_en[page.pageno].get_pixmap()
                image = np.fromstring(pix.samples, np.uint8).reshape(
                    pix.height, pix.width, 3
                )[:, :, ::-1]
                page_layout = model.predict(image, imgsz=int(pix.height / 32) * 32)[0]
                # kdtree 是不可能 kdtree 的,不如直接渲染成图片,用空间换时间
                box = np.ones((pix.height, pix.width))
                h, w = box.shape
                result_text=[]
                vcls = ["abandon", "figure", "table", "isolate_formula", "formula_caption"]
                for i, d in enumerate(page_layout.boxes):
                    text=''
                    if not page_layout.names[int(d.cls)] in vcls:
                        x0, y0, x1, y1 = d.xyxy.squeeze()
                        x0, y0, x1, y1 = (
                            np.clip(int(x0 - 1), 0, w - 1),
                            np.clip(int(h - y1 - 1), 0, h - 1),
                            np.clip(int(x1 + 1), 0, w - 1),
                            np.clip(int(h - y0 + 1), 0, h - 1),
                        )
                        box[y0:y1, x0:x1] = i + 2
                        if page_layout.names[int(d.cls)]=="plain text":
                            imagex = image[y0:y1,x0:x1]
                            result = ocr.ocr(imagex, cls=False)
                            for idx in range(len(result)):
                                res = result[idx]
                                for line in res:
                                    text+=line[1][0]
                            result_text.append(text)
                for i, d in enumerate(page_layout.boxes):
                    if page_layout.names[int(d.cls)] in vcls:
                        x0, y0, x1, y1 = d.xyxy.squeeze()
                        x0, y0, x1, y1 = (
                            np.clip(int(x0 - 1), 0, w - 1),
                            np.clip(int(h - y1 - 1), 0, h - 1),
                            np.clip(int(x1 + 1), 0, w - 1),
                            np.clip(int(h - y0 + 1), 0, h - 1),
                        )
                        box[y0:y1, x0:x1] = 0
                layout[page.pageno] = box
                # 新建一个 xref 存放新指令流
                page.page_xref = doc_en.get_new_xref()  # hack 插入页面的新 xref
                doc_en.update_object(page.page_xref, "<<>>")
                doc_en.update_stream(page.page_xref, b"")
                doc_en[page.pageno].set_contents(page.page_xref)
                interpreter.process_page(page)
    
        device.close()
        return obj_patch,result_text
    

    只有一段OCR的内容, 实在是看不懂怎么把OCR出来的结果往后传了。
    :(

  12. changed the title [-]feat (main): supports ocr on scanned document[/-] [+]翻译扫描档存在重影 / feat (main): supports ocr on scanned document[/+] on Dec 13, 2024
  13. pinned this issue on Dec 13, 2024
  14. 14 remaining items

  15. marked 图片类的pdf文档识别 #757 as a duplicate of this issue on Mar 13, 2025
  16. NullYing commented on Mar 25, 2025

    @NullYing

    先关注了MinerU,扫描件的准确读取率挺高的(不是手机拍照);想结合这两个项目看起来还是有点难度

  17. NullYing commented on Mar 25, 2025

    @NullYing

    先关注了MinerU,扫描件的准确读取率挺高的(不是手机拍照);想结合这两个项目看起来还是有点难度

    扫描件可以直接理解为图片,实际上是保持排版的图片翻译功能,可以参考微信的实现,长按图片点翻译可以自动翻译

  18. marked 解决PDF翻译重影问题 #803 as a duplicate of this issue on Mar 26, 2025
  19. marked 翻译完出现重影 #845 as a duplicate of this issue on Apr 16, 2025
  20. anbian123 commented on Apr 16, 2025

    @anbian123

    mathTranslate对于扫描版的pdf文件的翻译效果咋样呢?

  21. awwaawwa commented on Apr 16, 2025

    @awwaawwa
    Collaborator

    mathTranslate对于扫描版的pdf文件的翻译效果咋样呢?

    压根不支持😂

  22. awwaawwa commented on Apr 19, 2025

    @awwaawwa
    Collaborator

    BabelDOC 0.3.17 可以在文字区域底下加个白色背景,来部分支持OCR版PDF文档

  23. one-word commented on Apr 23, 2025

    @one-word

    mark

  24. Jose-Maria-Martins commented on Apr 30, 2025

    @Jose-Maria-Martins

    What about https://ocrmypdf.readthedocs.io/en/latest/. Couldn't it improve detection? And make it work for ocr pdfs?

  25. awwaawwa commented on Apr 30, 2025

    @awwaawwa
    Collaborator

    What about https://ocrmypdf.readthedocs.io/en/latest/. Couldn't it improve detection? And make it work for ocr pdfs?

    #860 thanks

  26. awwaawwa commented on May 6, 2025

    @awwaawwa
    Collaborator

    遇到此问题时,请尝试使用 2.0 预览版 #586 并启用高级选项中的 OCR Workaround 来翻译。

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or requesthelp wantedExtra attention is needed

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions