Research: PDF ingestion, OCR, and high-precision retrieval for large-scale document analysis #1
Loading…
Reference in a new issue
No description provided.
Delete branch "research/2026-08-12-pdf-ingestion-ocr-retrieval"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Dated research report (2026-08-12) for issue GRA-3: comparison of PDF parsing/OCR solutions (Docling, MinerU, Marker 2, olmOCR) and precision retrieval approaches (PageIndex, ColPali/ColQwen, RAGFlow) for massive technical/legal PDF corpora.
Recommendation — primary: custom Python pipeline: Docling (+ selective VLM OCR for scanned pages) → PageIndex-style tree retrieval, with a lexical/vector pre-filter at corpus scale.
Fallback: self-hosted RAGFlow with Docling/MinerU parsers.
Report file:
docs/research/2026-08-12-pdf-ingestion-ocr-high-precision-retrieval.md— includes executive summary, comparison table, per-candidate analysis, freshness check, data gaps, and full source list.