Every knowledge system does the same four things. Almost none of them do the third.

Nine systems, four functions each. The third is the one almost all of them skip.THE SAME FOUR FUNCTIONSCapture01Organise02DistillMISSINGRetrieve04
Nine systems, four functions each. The third is the one almost all of them skip.

Someone ran a comparison of nine different "company brain" setups this week, the internal systems teams build to keep AI agents supplied with organizational context. The tools ranged from a plain git repository to a full knowledge graph with an ingestion pipeline. Different budgets, different engineers, same four jobs underneath: pull in new information, store it in a structure you can query, filter out what's gone stale or wrong, and let people or agents search what's left.

Three of those four jobs showed up working in every system in that comparison. Ingestion was a Zapier hook or a browser extension. Storage was a database, a vault, a folder tree. Search was an embedding index or a full-text query.

Almost none of the nine had a working version of filtering.

Filtering means someone or something decides an entry is wrong now, or redundant, or already answered by three other entries in the system, and removes it or marks it down. It's the function that keeps a knowledge base from turning into a pile. It's also the one every system, corporate or personal, tends to skip, because it takes a judgment call, and judgment calls are expensive to make and harder to automate.

I checked my own save library while reading that comparison. 537 items. 263 still unread. 24 archived. Everything else sits there, read or not, with no mechanism that ever asks whether it still earns its spot. Storage isn't the issue. I have room to spare. Search isn't the issue either. I can find any of the 537 in seconds. The problem is pruning, and I built a bookmarking product and still haven't solved it for myself.

This is where a thread on structuring an Obsidian vault for LLM retrieval gets interesting, and where it falls short. The pitch: give your vault a router, an index, and clear edges between notes, so an agent doesn't have to read a thousand files to answer one question. That's a real fix for a real cost problem. Fewer tokens per query is worth solving. But a router only decides what to retrieve, not what to keep. You can build the sharpest retrieval layer in the world on top of a vault that's 40 percent dead weight, and the router will happily point you to notes you stopped believing months ago.

Read-it-later tools and bookmark managers hit the same wall from a different angle. They're built for the front door: save anything, from anywhere, in one click. Most are decent at search now too, since full-text and semantic search got cheap. Almost none of them ever come back to an old save and ask if it still matters. The result is a pile that grows every year and gets harder to trust, because a chunk of what's in there is outdated, superseded, or saved for a reason nobody remembers anymore. Ask most people what's actually in their bookmarks folder from two years ago and they'll admit they don't know, because nothing ever forced the question.

That's a different failure mode than a company losing track of internal docs, but it comes from the same gap. Both systems treat every saved item as equally worth keeping forever, because nobody built the step that revisits that assumption. A company brain with stale entries gives an agent wrong answers with total confidence. A bookmarks folder with the same problem stops being useful one unread item at a time, until opening it feels like more work than it's worth.

Here's the honest reason nobody builds this: automated pruning is a good way to delete something a person actually needed. Get it wrong once, in the wrong direction, and you've taught someone not to trust the system. So most tools default to keeping everything and calling that safe. A pile nothing ever leaves isn't safe. It just gets harder to trust every year, one unread item at a time.

The fix isn't full automation. It's giving the person or team a fast, low-friction way to make the call themselves at the moment it's cheapest: right when they're already looking at an old save and already have an opinion about it. Is this still true? Have I already acted on it? Is there a newer version of this idea three saves down? Ask that during a two-second interaction and you get real signal. Ask it never, default to keeping everything, and you get exactly the pile that comparison of nine systems flagged as the norm.

Vendors pitch AI as the fix for the first two problems, better ingestion and better search, and it helps with both. AI has barely been pointed at the third one. A model that can summarize 200 saved articles in a second could just as easily flag the 40 that contradict each other, or the 15 that are strict subsets of a save from three weeks earlier. Nobody's shipped that as a default feature, because "delete this for you" is a scarier promise than "search this for you."

That's the part of the four-function breakdown I keep coming back to while working on Saive. Capture and search are close to commodity at this point, in the corporate version and the personal one. The open problem, in both versions, is helping someone decide fast what still belongs in their library and what doesn't.

Save to my Saive library

Opens in a new tab. You'll be asked to sign in if you aren't already.