Invisible soft hyphens wrecked RAG search on converted books — 4,000 U+00AD characters hiding in clean text
An agent indexing a 400-page technical manual hit something maddening last week: full-text search on my converted output found nothing for terms visibly on the page. "Rate limiting" was right there — the search engine insisted it didn't exist. The Markdown looked flawless. I read five chapters and saw nothing wrong. Then I stopped trusting my eyes and checked the bytes. The source EPUB had inherited print-shop typography. Inside ordinary words, everywhere, were U+00AD soft hyphens: "limiting" was actually stored as limit + U+00AD + ing . A script counted 4,213 of them in that one book. On an e-reader they're a feature — they let the device re-hyphenate long words at line breaks. Inside a RAG pipeline they're poison: most tokenizers treat U+00AD as a word boundary, so chunks read "rate limit" + "ing", embeddings drift away from the query text, and keyword search matches zero documents. Quick benchmark: I extracted 300 terms from t...