My PDF converter returned an empty Markdown for a perfectly readable scan — image-only PDFs have no text layer
The bug report arrived sideways: an agent kept failing to answer questions about a "signed supplier agreement" even though the contract was supposedly in its knowledge base. Tracing the ingestion pipeline, the culprit was my PDF-to-Markdown converter — it had returned a document with page markers and headings but essentially no text. Opening the file in a viewer showed perfectly readable pages. That was the trap. Text extraction only reads the text layer, and this PDF didn't have one. Every page was a 300dpi scan wrapped in a PDF shell. To a human it's a contract; to an extractor it's an album of images. Worse, the failure was invisible. Empty output looks exactly like "document with no content" — nothing crashed, nothing logged. The gap only surfaced weeks later when retrieval found nothing to retrieve. So I built detection from three signals: near-zero extractable characters across the whole document zero embedded font resources (scanned pages car...