Beatriz Almeida has released papero, an open-source PDF text extraction tool designed for speed and structural accuracy. Unlike many contemporary solutions, papero relies on plain geometry rather than machine learning models, allowing it to run efficiently on a laptop CPU.
Open-Source PDF Parser 'papero' Extracts Structured Data without ML Models
The tool extracts content into multiple formats, including Markdown, JSON, Word, Excel, and HTML. It specifically targets the structural elements necessary for Retrieval-Augmented Generation (RAG) and LLM preprocessing, such as reading order, tables, mathematical formulas (in LaTeX), and figure positions.
Papero provides several integration methods, including a Python library (pip install papero-extract), a Command Line Interface (CLI), and a REST API. It also features a browser-based application powered by pdf.js that allows users to inspect document blocks and export files without the PDF ever leaving the local machine.
For scanned documents, the tool supports OCR via Tesseract. It also utilizes Apache Tika to handle various file formats, including DOCX, PPTX, XLSX, and EPUB. The extractor's performance is notable for complex layouts, such as multi-column arXiv papers, where it reportedly maintains high speed and accuracy for block detection.
Sources
- Lightweight PDF parser with layout, tables, formulas and bounding boxes (Hacker News Frontpage, 2026-10-01)
- 公式サイト