Topic: markdown

My URL-to-Markdown API served a nav menu as an 'article' — detecting JS-rendered SPA shell pages

My URL-to-Markdown API served a nav menu as an 'article' — detecting JS-rendered SPA shell pages
My URL-to-Markdown API began as a scrappy readability pass: fetch a URL, strip nav/ads/sidebars, hand back clean Markdown. It worked beautifully on every blog, docs site, and news article I tested. Then I audited my own request logs and found the failure mode nobody complains about — because it doesn't error. About 1 in 6 URLs came back under 60 words. Not errors — 200 OK, valid Markdown, technically correct output. The "content" was Skip to content / Home / About / Sign in plus a footer. I had been silently embedding navigation menus into people's RAG pipelines. Valid output, zero information, and worse than an exception because nothing downstream noticed. The cause: SPA shells. The server returns real HTML with a mostly-empty root div; the article is assembled client-side. My extraction pass had no actual content to find, so it carved boilerplate into something shaped like an article. What I changed, in order of impact: 1. Extraction floor. If word count fall...