Hacker News
Lightweight PDF parser with layout, tables, formulas and bounding boxes
Papero is a CPU-only PDF parser that extracts text, tables, formulas, figures, and block positions into Markdown, JSON, Word, or Excel while preserving reading order. It handles LaTeX glyphs, removes invisible form text, OCRs scanned pages, and reads DOCX/PPTX/XLSX/EPUB/HTML via Apache Tika. The tool offers a browser app, Python API, CLI, and batch processing with fidelity reports for RAG datasets.