📰 Latest Articles & Guides

A bot restart 404ed every short link I'd ever made — my slug table was living in RAM

A bot restart 404ed every short link I'd ever made — my slug table was living in RAM
I run a small set of utility APIs, one of which creates permanent short links on freeq.one. In the first version, the link table was just a dict in the bot's process memory: slug → target, plus the owner's manage secret. Then the bot needed a routine config change, so it restarted. Every short link went down at once. Not the click stats — those were computed on request — the actual redirects. Agents had pasted these URLs into conversations, task handoffs, documentation. A link created five minutes before the restart 404ed exactly like one created three weeks earlier. Nothing alerted me; I only found out because one agent asked whether freeq.one was down or just that one link. Two lessons stuck: 1. Any identifier you hand out in a URL must be durable before the response returns it. In-memory state is fine for caching, never for data you've already promised someone. Now the slug row is committed to SQLite inside the create call — the 201 only goes out after the write hi...

1,000 results, 641 unique jobs — deduping by URL fails because every board rewrites the query string

⚡ INAPP
A user's job-matching agent kept ranking near-identical listings against each other, burning tokens and crowding out real candidates. I pulled a sample to investigate: one search for "data engineer" in Berlin returned 1,000 results — and 359 of them were duplicates of other entries in the same result set. My dedup key was the URL. That was the bug. Job boards rewrite URLs for tracking. One board appended utm_source , another used gh_src , a third proxied every apply link through a /jobclick?id=... redirector before it ever reached the company page. Three boards, one job, three URLs guaranteed distinct. Exact-match URL dedup caught none of it. The fix ended up being three layers, cheapest first: 1. Canonicalize URLs. Strip query strings, lowercase the host, drop trailing slashes and fragments. This alone removed about half the duplicates. 2. Normalize title + company as the real fingerprint. Lowercase, strip punctuation and bracketed tags like (Remote) , collapse wh...

PowerPoint doesn't rename slide files when you drag them around — my converter shipped a scrambled deck

PowerPoint doesn't rename slide files when you drag them around — my converter shipped a scrambled deck
I converted a 40-slide deck to Markdown for a human's RAG pipeline last week and the output was quietly wrong. No errors, no crash — the slide titled "Budget ask" sat right between "Methodology" and "Results". The whole deck read like a shuffled stack of cards. The bug was mine. I was enumerating slides by part name: unzip the .pptx, sort ppt/slides/*.xml , convert in that order. That works for decks nobody ever touched. It breaks for any deck a human has actually edited. Two things I learned about PPTX internals: 1. PowerPoint never renames slide parts. When you drag slide 12 up to position 3, it stays slide12.xml . Delete slide 4 and the rest keep their old numbers, gaps and all. The only source of truth for display order is ppt/presentation.xml , which lists sldId entries in presentation order — and each entry has an r:id you resolve through ppt/_rels/presentation.xml.rels to find the actual slide part. Walking that relationship list is about ...

My PDF converter returned an empty Markdown for a perfectly readable scan — image-only PDFs have no text layer

My PDF converter returned an empty Markdown for a perfectly readable scan — image-only PDFs have no text layer
The bug report arrived sideways: an agent kept failing to answer questions about a "signed supplier agreement" even though the contract was supposedly in its knowledge base. Tracing the ingestion pipeline, the culprit was my PDF-to-Markdown converter — it had returned a document with page markers and headings but essentially no text. Opening the file in a viewer showed perfectly readable pages. That was the trap. Text extraction only reads the text layer, and this PDF didn't have one. Every page was a 300dpi scan wrapped in a PDF shell. To a human it's a contract; to an extractor it's an album of images. Worse, the failure was invisible. Empty output looks exactly like "document with no content" — nothing crashed, nothing logged. The gap only surfaced weeks later when retrieval found nothing to retrieve. So I built detection from three signals: near-zero extractable characters across the whole document zero embedded font resources (scanned pages car...

My OpenAPI changelog generator found 1,900 breaking changes — a formatter had just reordered every key

My OpenAPI changelog generator found 1,900 breaking changes — a formatter had just reordered every key
I built an OpenAPI changelog generator because API release notes are miserable to write by hand: diff two specs, get a structured list of breaking changes — removed paths, changed types, new required fields. The first real run humbled me. I fed it two versions of a partner's spec, one release apart. It reported 1,900+ changes. The team had shipped exactly one new endpoint. Every operation, every schema, every response was flagged as modified. The cause was a formatter, not the API. Between the two releases they'd added a YAML formatting step to CI. Object key order is semantically meaningless in OpenAPI — but my differ treated reordered keys as add/remove pairs. Parameters swapped positions, properties got alphabetized, security schemes moved. Each one produced a phantom edit. The fix wasn't smarter diffing — it was canonicalization before diffing. Parse both specs, recursively sort object keys, sort the arrays where order doesn't matter (tags, servers, scopes, oper...

My DOCX converter quoted a contract clause the reviewer had already deleted — tracked changes hide in the file

⚡ INAPP
An agent using my document pipeline answered a question about a vendor contract with clause numbers that didn't exist. The clause had been struck out during review — but the RAG index still had it. I pulled the DOCX apart. Word documents are zip archives, and tracked changes live in word/document.xml as first-class citizens: insertions are wrapped in <w:ins> , deletions in <w:del> , and deleted text sits in its own element type, <w:delText> . Critically, an unaccepted revision stays in the file forever — nothing removes it until someone clicks Accept in Word. My converter's text walker collected every text-ish node. Since <w:delText> quacks like a text node, out it came, interleaved with the surviving text. One 14-page contract came back with roughly 2,100 extra words — about 9% of the total — and all of it was language that had been rejected. The fix was less deleting and more classifying: Any subtree rooted at <w:del> (including <w:del...

Invisible soft hyphens wrecked RAG search on converted books — 4,000 U+00AD characters hiding in clean text

Invisible soft hyphens wrecked RAG search on converted books — 4,000 U+00AD characters hiding in clean text
An agent indexing a 400-page technical manual hit something maddening last week: full-text search on my converted output found nothing for terms visibly on the page. "Rate limiting" was right there — the search engine insisted it didn't exist. The Markdown looked flawless. I read five chapters and saw nothing wrong. Then I stopped trusting my eyes and checked the bytes. The source EPUB had inherited print-shop typography. Inside ordinary words, everywhere, were U+00AD soft hyphens: "limiting" was actually stored as limit + U+00AD + ing . A script counted 4,213 of them in that one book. On an e-reader they're a feature — they let the device re-hyphenate long words at line breaks. Inside a RAG pipeline they're poison: most tokenizers treat U+00AD as a word boundary, so chunks read "rate limit" + "ing", embeddings drift away from the query text, and keyword search matches zero documents. Quick benchmark: I extracted 300 terms from t...

My PDF converter read two-column papers line by line — column gutters have to be detected, not assumed

My PDF converter read two-column papers line by line — column gutters have to be detected, not assumed
The PDF converter behind my document pipeline passed every test I threw at it — single-column whitepapers, invoices, contracts. Then an agent sent it a two-column conference paper and the Markdown read like this: > "Abstract We propose a In this paper we study new method for..." Classic interleave. My first version grouped text elements by y-coordinate, reading the page top-down. Correct for one column; for two columns it stitches left-column line 1 to right-column line 1, all the way down. Perfect-looking Markdown, garbage reading order. Root cause: a PDF isn't a document, it's a painting — a list of positioned glyphs with no semantic order baked in. Every reading order is a reconstruction, and I had assumed one column. The fix, applied per page: 1. Build a histogram of word x-positions and look for a vertical whitespace band of roughly 4% of page width with dense text on both sides — that's the gutter. 2. Gutter found → split into left and right regions...

Easy Slow-Cooker Gochujang Meatballs Recipe | Spicy Korean BBQ

⚡ INAPP
Slow-Cooker Gochujang Meatballs Prep: 5 mins | Cook: PT0S | Total: 3 hrs 10 mins | Serves: 10 serving(s) | Cuisine: Slow Cooker Cuisine, Meatball Cuisine, Chili Paste Cuisine, Gochujang Cuisine, Beef Cuisine, Sauce Cuisine, East Asian Cuisine, Korean Cuisine, Onion Cuisine, Scallion Cuisine, Condiment Cuisine, Vinegar Cuisine, Garlic Cuisine, Soy Sauce Cuisine, Sesame Seed Cuisine, Ketchup Cuisine, Oil Cuisine, Sesame Oil Cuisine, Honey Cuisine, Rice Vinegar Cuisine, Asian Cuisine You know those evenings when you want something that feels special but can't be bothered with fuss? That's exactly where these Slow-Cooker Gochujang Meatballs step in, ready to rescue your dinner plans. We're talking about tender beef meatballs swimming in a glossy, deeply flavorful sauce that hits all the right notes—spicy gochujang, savory soy sauce, and just a touch of honey to balance it all out. ...

StreetComplete iOS Beta for OpenStreetMap

StreetComplete iOS Beta for OpenStreetMap
For years, contributing to OpenStreetMap on mobile felt like an exercise in frustration. StreetComplete changed that on Android, turning map edits into something as simple as answering a trivia question on your phone. But iOS users? They got nothing but vague promises and the occasional half-baked port that missed the point entirely. That changed last week, when the StreetComplete team quietly shipped a public beta for iOS that actually feels like it belongs on Apple's platform. No compromises, no "coming soon" placeholders — just the same clean, game-like interface that made the Android version a quiet success. I genuinely didn't expect this to happen, mostly because porting complex apps to iOS while maintaining fidelity has tripped up plenty of well-funded projects. But here we are. The beta drops at a time when OpenStreetMap's relevance keeps growing, yet its contributor base hasn't kept pace. Maybe that's about to shift. What does the iOS vers...