The Friction Was in What I Didn't Say
I built a curated news engine with an agent. Every bug was the system doing exactly what I told it to, colliding with something I never said out loud. A field report on the gap between what you build and what you assume.
Systems are predictable. The assumptions you forget to state are not.
I wanted a small thing: somewhere inside one of my own systems to read the news once a day, curated by interest. Not a firehose. Not another RSS crawler I'd lose interest in. Not one more tab to feel guilty about. A list that something had already thinned out for me, based on what it knows about me and what I actually care about.
What I got, eventually, was exactly that. But the interesting part wasn't the feature. It was that every single bug was the system doing exactly what I told it to, colliding with something I never said out loud.
This is a field report on that gap.

Build on the seams, not the soil
The first decision was the cheapest and the most important: don't invent architecture. The system I already run had every joint this feature needed. I just had to find them.
- Storage? There was already a pattern, a SQLite database driven by a small Python CLI, called from the app's server functions. I cloned it into a
news-db.py. - A daily job? A
launchdcron already ran a handful of tasks every morning. I copied its shape. - "Curated by interest"? An engine turn already triaged my inbox. News scoring is the same move with a different prompt.
I was adding a room by tapping the existing plumbing, not digging new trenches. The whole feature is one seam reused four times: Python CLI → server function → data provider → query hook. Nothing novel. That's the point. Novelty is a cost you pay in bugs, and I wanted to spend my novelty budget elsewhere.
The two decisions that actually mattered
Two choices mattered, and both were about robustness over reach.
RSS first, not scraping. Feeds are a contract, HTML is a rumor. I parse RSS/Atom with the Python standard library, urllib plus xml.etree, zero dependencies. Scraping arbitrary sites would have doubled the surface area and halved the reliability. The door's open for per-source HTML later. I just didn't walk through it on day one.
An agent that's allowed to fail. A local LLM scores each article by interest and writes a one-line rationale. But if the gateway is down, the pipeline doesn't stop. It falls back to recency and shows you the articles anyway. Reading is never blocked by curation. Curation is a luxury layered on top of a system that already works without it.
That's the architecture. It took a while to decide and longer to design. Then the system started teaching me what I'd left unsaid.
Friction lives in the unstated
Here's the honest part. The four moments where a perfectly correct system did something I didn't want, because I'd never told it not to.
1. The switch wired to nothing. I built a toggle: "automatic daily curation." It saved. It flipped. And the cron ignored it completely, because I'd added the flag to the UI and the config and never wired the runner to read it.
A switch connected to nothing is worse than no switch. It's a lie with a nice animation. And it was obvious, to me. But when you're working with an AI (or, in my experience, with another human), your assumptions are obvious only to you. The obvious has to be said.
2. The schema that ate my settings. Then the toggle started reverting on its own. Flip it on, refresh, it's off again. The cause was beautiful in its correctness: the config save was validated against a schema that didn't know the news field existed yet. The validator did its job perfectly. It stripped the field I hadn't declared, every single time.
A whitelist will silently eat anything you forget to add to it. The system wasn't broken. My mental model was. I was editing a value the persistence layer had quietly agreed never to keep.
3. Two limits, colliding. First real run: "385 new · 80 curated." Looks like a bug. It isn't. It's two correct limits meeting.
The ingest had no floor on age, so it pulled the entire backlog of every feed. And I'd only wired up a handful of sources to test with, or it would still be running now. The curator had a per-run cap of 80 to keep cost sane. Both behaved exactly as written, and together they produced nonsense.
The fix wasn't a bigger cap. It was a smaller intake: bound the input, not the processing. A seven-day freshness window at ingest, and dedup so re-runs add only what's genuinely new. Cap the river, not the bucket.
4. The 15-second heartbeat. I opened the network tab and the same request was firing every fifteen seconds, forever, each one spawning a Python process. The dashboard had a global "live mode" default: poll everything, refetch on focus. My news queries inherited it without asking.
But news isn't live data. It changes once a day. A default is an opinion the rest of the system holds about your data, and it was wrong about mine. I made the news queries static and let mutations invalidate them instead. The heartbeat stopped.
The one that wasn't even in the code. One feed, Anthropic's, kept 404-ing. Not my bug, just a dead URL. But it surfaced because I'd built per-source health: last error, consecutive failures, auto-disable after five. The system didn't hide the rot. Observability isn't a feature you bolt on at the end; it's what tells you the difference between "broken" and "fed a bad address." Fail fast, fix fast, but only when the system is honest enough to tell you which is which.
The shape of the polish
Once the foundation stopped lying to me, the rest was subtraction, which, if you've read me before, is basically the only kind of engineering I trust.
- Pagination instead of infinite scroll. Show twenty, not three hundred.
- Curation in small batches instead of one giant turn. Faster, and honest about what's still pending.
- Selectable ordering. Relevance by default, recency when I just want the new stuff. The system had an opinion about order; I gave it back to the reader.
- Filters as dropdowns, not a wall of chips.
None of it is clever. All of it is removal.
What I'm taking with me
The lesson isn't "test more." It's sharper than that.
A predictable system makes your assumptions load-bearing. Every piece behaved exactly as specified. The toggle saved exactly what the schema allowed. The ingest pulled exactly what it was pointed at. The dashboard polled exactly as configured. The friction was never in the machine. It lived in the quiet space between what I built and what I assumed. And in how I was talking to my models.
You don't debug that space by reading the code harder. You debug it by saying the assumption out loud and watching the system disagree with you.
The feature was quick to build. Understanding what I hadn't said took days.
The system did everything right. That's exactly why it was so hard to see what I'd done wrong.