Matchboxmatchbox

paperless-paddle-ocr

CPU-friendly OCR sidecar that improves Paperless-NGX text extraction using PaddleOCR.

Servicefree

paperless-paddle-ocr is a CPU-friendly sidecar worker that re-runs OCR on documents already stored in a self-hosted Paperless-NGX instance, using PaddleOCR instead of the built-in Tesseract engine, and writes the improved text back via the API. It targets documents where the default OCR struggles, such as difficult scans, handwriting-adjacent text, or non-Latin scripts. It is aimed at Paperless-NGX self-hosters who want more accurate full-text search without needing a GPU.

Categories
document managementOCRself-hosted

Full match profile

Behind the summary, Matchbox keeps a richer profile of paperless-paddle-ocr - the signals our matcher actually reads to decide when to surface it. It stays private; claim the listing to see and control it.

  • Problem & pain-point mapping
  • Who we surface it to (audience fit)
  • What it's a strong alternative to
  • Trust & credibility signals

Something wrong with this listing — dead link, not a real product, wrong info?

Try Matchbox with your own problem

Describe what is not working - we’ll show you whether paperless-paddle-ocr (or something else) actually fits.