← Back to match
pdfmux
Self-healing PDF extraction that flags pages it can't read, and certifies other extractors' output for silent drops.
Desktopfree
pdfmux is an open-source PDF extraction tool that routes pages across multiple extraction backends and OCR, automatically re-extracting pages it judges to be blank, scrambled, or otherwise bad, and flagging any it still can't read rather than dropping them silently. It can also audit another extraction engine's output against the source PDF to find silently dropped pages. It's available as a CLI, Python library, and MCP server.
Categories
Developer ToolsDocument ProcessingAI Tools

