Repository navigation
翻译扫描档存在重影 / feat (main): supports ocr on scanned document #19
Description
Activity
图片型的 PDF 文档暂时还没办法翻译,目前主要还是在优化电子书和论文的翻译效果
Reacted by Keon Chen, zgshan, jiangqing and DEBRIS-PING图片型的 PDF 文档暂时还没办法翻译,目前主要还是在优化电子书和论文的翻译效果
好的,非常感谢
均为图像有点为难人了,ocr的质量 影响文字的质量 影响翻译的效果
Reacted by z3r0yu, Byaidu, zgshan and zstar加一个可选流程paddleOCR,
sayura
这个模型非常准确,就是对算力的要求会高于 Paddle OCRsayura 这个模型非常准确,就是对算力的要求会高于 Paddle OCR
和 minerU/marker 比较怎么样呀
Owner
sayura 就是 marker 的作者做的开源多国语言和表格的 OCR 模型😂
minerU 这个我没有测试,我只测试了 PaddleOCR 高精度模型,Sayura 效果比它好很多,而且支持多国语言效果很好。
我看 minerU 的 issue,对多国语言的支持好像不佳
缺点就是 Sayura 对 GPU 显存要求有点高,头疼,不太会量化模型。Reacted by Keon Chen, zgshan and Jin Yao- changed the title
[-]当PDF每一页均为图像时,无法进行翻译[/-][+]feat (main): supports ocr on scanned document[/+]on Nov 21, 2024 from typing import BinaryIO import numpy as np import tqdm from pymupdf import Document from pdfminer.pdfpage import PDFPage from pdfminer.pdfinterp import PDFResourceManager from pdfminer.pdfdocument import PDFDocument from pdfminer.pdfparser import PDFParser from pdf2zh.converter import TranslateConverter from pdf2zh.pdfinterp import PDFPageInterpreterEx from pymupdf import Font import numpy as np from paddleocr import PaddleOCR file="" def extract_text_to_fp( inf: BinaryIO, pages=None, password: str = "", debug: bool = False, page_count: int = 0, vfont: str = "", vchar: str = "", thread: int = 0, doc_en: Document = None, model=None, lang_in: str = "", lang_out: str = "", service: str = "", resfont: str = "", noto: Font = None, callback: object = None, **kwarg, ) -> None: ocr = PaddleOCR(use_angle_cls=True, lang="en") rsrcmgr = PDFResourceManager() layout = {} device = TranslateConverter( rsrcmgr, vfont, vchar, thread, layout, lang_in, lang_out, service, resfont, noto ) assert device is not None obj_patch = {} interpreter = PDFPageInterpreterEx(rsrcmgr, device, obj_patch) if pages: total_pages = len(pages) else: total_pages = page_count parser = PDFParser(inf) doc = PDFDocument(parser, password=password) with tqdm.tqdm( enumerate(PDFPage.create_pages(doc)), total=total_pages, ) as progress: for pageno, page in progress: if pages and (pageno not in pages): continue if callback: callback(progress) page.pageno = pageno pix = doc_en[page.pageno].get_pixmap() image = np.fromstring(pix.samples, np.uint8).reshape( pix.height, pix.width, 3 )[:, :, ::-1] page_layout = model.predict(image, imgsz=int(pix.height / 32) * 32)[0] # kdtree 是不可能 kdtree 的,不如直接渲染成图片,用空间换时间 box = np.ones((pix.height, pix.width)) h, w = box.shape result_text=[] vcls = ["abandon", "figure", "table", "isolate_formula", "formula_caption"] for i, d in enumerate(page_layout.boxes): text='' if not page_layout.names[int(d.cls)] in vcls: x0, y0, x1, y1 = d.xyxy.squeeze() x0, y0, x1, y1 = ( np.clip(int(x0 - 1), 0, w - 1), np.clip(int(h - y1 - 1), 0, h - 1), np.clip(int(x1 + 1), 0, w - 1), np.clip(int(h - y0 + 1), 0, h - 1), ) box[y0:y1, x0:x1] = i + 2 if page_layout.names[int(d.cls)]=="plain text": imagex = image[y0:y1,x0:x1] result = ocr.ocr(imagex, cls=False) for idx in range(len(result)): res = result[idx] for line in res: text+=line[1][0] result_text.append(text) for i, d in enumerate(page_layout.boxes): if page_layout.names[int(d.cls)] in vcls: x0, y0, x1, y1 = d.xyxy.squeeze() x0, y0, x1, y1 = ( np.clip(int(x0 - 1), 0, w - 1), np.clip(int(h - y1 - 1), 0, h - 1), np.clip(int(x1 + 1), 0, w - 1), np.clip(int(h - y0 + 1), 0, h - 1), ) box[y0:y1, x0:x1] = 0 layout[page.pageno] = box # 新建一个 xref 存放新指令流 page.page_xref = doc_en.get_new_xref() # hack 插入页面的新 xref doc_en.update_object(page.page_xref, "<<>>") doc_en.update_stream(page.page_xref, b"") doc_en[page.pageno].set_contents(page.page_xref) interpreter.process_page(page) device.close() return obj_patch,result_text只有一段OCR的内容, 实在是看不懂怎么把OCR出来的结果往后传了。
:(- changed the title
[-]feat (main): supports ocr on scanned document[/-][+]翻译扫描档存在重影 / feat (main): supports ocr on scanned document[/+]on Dec 13, 2024 - pinned this issue
on Dec 13, 2024 14 remaining items
- marked Feature Request – Better Text Recognition & OCR for Improved Translations #773 as a duplicate of this issue
on Mar 16, 2025 先关注了MinerU,扫描件的准确读取率挺高的(不是手机拍照);想结合这两个项目看起来还是有点难度
Reacted by awwaawwa and Xuanzhao先关注了MinerU,扫描件的准确读取率挺高的(不是手机拍照);想结合这两个项目看起来还是有点难度
扫描件可以直接理解为图片,实际上是保持排版的图片翻译功能,可以参考微信的实现,长按图片点翻译可以自动翻译
mathTranslate对于扫描版的pdf文件的翻译效果咋样呢?
mathTranslate对于扫描版的pdf文件的翻译效果咋样呢?
压根不支持😂
BabelDOC 0.3.17 可以在文字区域底下加个白色背景,来部分支持OCR版PDF文档
mark
What about https://ocrmypdf.readthedocs.io/en/latest/. Couldn't it improve detection? And make it work for ocr pdfs?
What about https://ocrmypdf.readthedocs.io/en/latest/. Couldn't it improve detection? And make it work for ocr pdfs?
#860 thanks
遇到此问题时,请尝试使用 2.0 预览版 #586 并启用高级选项中的 OCR Workaround 来翻译。
当pdf文件均为图像,而不是可编辑(复制)状态时,翻译完全失败,具体见图