HeadlinesBriefing HeadlinesBriefing.com

Papero: Fast Open-Source PDF Text Extraction API

Hacker News •
×

Papero is a fast, open-source PDF text extraction API that never stores your files. It extracts document structure—reading order, tables, formulas, figures, and block positions—without a heavyweight stack. It runs on CPU only, with no ML models, and works in your browser, Python, or as an API.

Getting structure from a PDF—like column order, table detection, and formula recognition—makes output usable for RAG, search, and LLMs. Papero uses plain geometry, staying fast on a laptop CPU. It handles two- and three-column reading order, real tables (ruled, borderless, LaTeX booktabs), formulas with LaTeX output, and every block's bounding box.

Quick start: `pip install papero-extract` and `from papero_extract import extract`. The browser app lets you load a PDF and export without your file leaving your machine. Options include OCR, image extraction, and batch processing for RAG datasets. CLI commands support extraction to Markdown, JSON, CSV, HTML, and more.

Papero also handles scanned pages via OCR, drops invisible white text, and reads DOCX/PPTX/XLSX/EPUB/HTML through Apache Tika. It includes a fidelity report for batch processing, showing which documents need review.

Source: Hacker News · Summarized by HeadlinesBriefing