Skip to content
← All posts
airagretrievalagent-memorycascaderust

Where Retrieval Stops Being Enough for Agent Memory

5 min read

Similarity search fails agent memory in two specific ways. Here's how Cascade fixes both, including the RRF fusion code and why recency is a retrieval channel instead of a filter.

I’ve been building retrieval for agent memory for about two years, and the thing that took me longest to accept is that better search doesn’t fix it.

Here’s the failure that made it click. I asked my own system what I’d decided about a config change, and it handed me a conversation from eight months back. Perfect match, exactly on topic, completely useless. I’d reversed that decision in June, and it gave me both versions with the same confidence and no way to tell them apart.

That’s two separate problems wearing one costume.

Problem one: relevance and recency fight each other

A vector search will happily rank something from last spring above something from yesterday, because it’s answering the question you actually asked, which was “what’s most similar to this.” That isn’t what you wanted. Better embeddings don’t help, because the model is doing its job correctly. The ranking objective is wrong, not the vectors.

The obvious fix is to filter or re-sort by date after retrieval. I tried that first and it’s worse than it sounds. Filter too hard and you throw away the eight-month-old answer that was genuinely the right one. Filter too soft and nothing changes. Either way you’ve bolted a second, unprincipled ranking pass onto the end of a system that already has one.

Problem two: a flat index has no idea what replaced what

Facts get revised, decisions get reversed, and nothing in a cosine score knows that happened. Agents are especially bad here because they’ll act on whichever version came back first and never mention there was another one.

What Cascade actually does

Rather than treat recency as a post-processing step, Cascade makes it a retrieval channel of its own and fuses it with the others. Five channels can contribute:

Channel What it catches Default weight
FTS5 BM25 body Exact identifiers, rare terms, error strings 1.0
Dense ANN Paraphrase, conceptual match 1.0
Curated title/description/tags Human-authored signal 0.8
Document recency What happened lately 0.5
Multi-vec ColBERT Token-level late interaction 1.0

The weights matter more than the channel list. Curated and recency sit deliberately below 1.0 so they nudge the ranking without overriding it. Recency at 0.5 means a fresh document gets a real boost but can still lose to an older document that three other channels agree on. That’s the behavior I wanted and couldn’t get from a filter.

Fusing them is Reciprocal Rank Fusion, which throws away the raw scores and uses only the rank each channel assigned:

pub fn rrf_merge(lists: &[RankedList<'_>], k: f64, top_n: usize) -> Vec<FusedHit> {
    let mut scores: HashMap<i64, (f64, Vec<String>)> = HashMap::new();

    for list in lists {
        if list.hits.is_empty() {
            continue;
        }
        for (zero_based, (chunk_id, _score)) in list.hits.iter().enumerate() {
            let rank = (zero_based + 1) as f64; // 1-based
            let contribution = list.weight / (k + rank);
            let entry = scores.entry(*chunk_id).or_insert_with(|| (0.0, Vec::new()));
            entry.0 += contribution;
            if !entry.1.iter().any(|s| s == list.source) {
                entry.1.push(list.source.to_owned());
            }
        }
    }
    // ...sort by score, truncate to top_n
}

Note the _score on line 8. The underlying relevance score is discarded on purpose. BM25 scores and cosine similarities aren’t on the same scale, aren’t on the same distribution, and there’s no honest normalization between them. Rank is the only thing they share. Throwing away the scores is what lets you add a channel without retuning anything.

k defaults to 60, which comes straight from the Cormack et al. 2009 paper. It’s a smoothing constant. Small k makes rank 1 dominate everything; large k flattens the difference between rank 1 and rank 20. Sixty is high enough that a document ranked third by three channels beats a document ranked first by one. That consensus behavior is the entire reason to fuse in the first place.

The sources_hit vector is a debugging affordance I’d add again. When a result looks wrong, the first question is always which channel put it there, and carrying that through the merge means you can answer it without re-running anything.

The layer underneath

Fusion fixed ranking. It didn’t fix memory, because retrieving the right chunk still doesn’t tell you a chunk was superseded.

Underneath the search I put summary pyramids: rolled-up summaries at turn, day, week, and month granularity, each one embedded and indexed alongside the raw chunks. Ask about last Tuesday and you land on the day layer. Ask how a project evolved over a quarter and you land on the month layer. The rollup is doing compression, but it’s also doing implicit supersession, because the week summary already reflects the reversal that the original turn doesn’t know about.

All of it is indexed against an ltree topic taxonomy, so you can walk the structure instead of only searching by similarity. retrieval.fusion.rrf is a path, not a bag of words, and asking for everything under retrieval is a prefix query rather than a semantic one.

Where this still falls down

The taxonomy is hierarchical, and hierarchy is the wrong shape for the problem it’s being asked to solve. A tree encodes where a fact sits. It cannot encode how facts relate to each other, which is the thing that actually matters once information starts revising itself.

“Decision B superseded decision A” is the clearest example. That’s a typed, directed relationship between two nodes. In a tree you can put both decisions under the same topic path and hope proximity implies something, but the supersession itself is unrepresentable. It’s an edge, and a hierarchy has no edges. Only containment.

A graph-native store with typed edges and proper invalidation addresses both failure modes at once, because supersession stops being metadata you maintain alongside the index and becomes a first-class property of the index. That’s the direction this goes next, and I don’t think retrieval alone gets you there.

The code above is in Cascade, under crates/cascade-rag.