A month after the first scripts, Cascade is a real thing. Multi-account fleet routing, a persistent daemon, hybrid retrieval, and a hook system. It’s the open-source version of patterns that go further inside nSelf and nClaw, but the core holds up on its own.
Three pieces did the work. Here’s what each one actually looks like.
Routing: pick the cheapest model that can still do the job
The naive version of model routing is “use the good model for hard things.” That’s not implementable, because you don’t know a task is hard until you’ve done it. What you can do is classify the kind of task and route on that.
Cascade has three tiers:
pub enum Tier {
/// T1: decisions, architecture gates, CR-C. Default model: Opus.
T1,
/// T2: bulk execution, code gen, QA. Default model: Sonnet.
T2,
/// T3: cheap triage, grep+summarize, taxonomy. Default model: Haiku.
T3,
}
The routing rule is where it gets useful. When several agents can handle a task kind, it takes the lowest tier, not the best one:
if let Some(role) = AgentRole::from_task_kind(task_kind) {
if let Some(spec) = inner.specs.values()
.filter(|s| s.role == role)
.min_by_key(|s| s.tier as u8) // lowest tier wins
{
return Ok(spec.clone());
}
}
// Fall back: prefer Generic, then anything
min_by_key on tier means the router is biased toward cheap by construction. You don’t get expensive-by-default with a cost check bolted on later, which is how most of these systems drift. If a T3 agent is registered for the role, it gets the work, and escalation is an explicit decision rather than an accident.
The fallback chain matters too. Unknown task kind degrades to a Generic agent at the lowest tier rather than erroring. In an agent harness, refusing to route is worse than routing imperfectly, because a stalled agent needs a human and a suboptimal one usually doesn’t.
That structure is where the cost reduction comes from. Not a clever prompt. A default that leans cheap and makes you argue for expensive.
Retrieval: five channels, fused by rank
Lexical search alone misses intent. Dense search alone misses the literal symbol you typed. Everyone knows this. What nobody agrees on is how to combine them when their scores aren’t comparable.
Cascade runs up to five retrieval channels and fuses them with Reciprocal Rank Fusion, which discards the scores entirely and uses only the rank each channel assigned:
| Channel | Catches | Weight |
|---|---|---|
| FTS5 BM25 | Exact identifiers, error strings | 1.0 |
| Dense ANN | Paraphrase, concepts | 1.0 |
| Curated title/tags | Human-authored signal | 0.8 |
| Recency | What happened lately | 0.5 |
| Multi-vec ColBERT | Token-level late interaction | 1.0 |
Recency being a channel at 0.5 rather than a sort or a filter is the design decision I’d defend hardest. A fresh document gets a genuine boost, and still loses to an older document that three other channels agree on. I tried the filter version first and it was strictly worse in both directions: filter hard and you lose the old answer that was right, filter soft and nothing changes.
I wrote up the fusion math and the failure modes separately in Where Retrieval Stops Being Enough.
The other thing I’d do again is make the index degrade instead of failing:
pub enum TierLevel {
/// FTS5 BM25 keyword search only. No embeddings required.
/// Index size: ~text size. RAM: negligible.
#[default]
Minimal,
// ...richer tiers add embeddings, then multi-vector
}
Minimal is the default, and it needs no model, no GPU, and effectively no RAM. Someone who clones the repo gets working search immediately. Embeddings are an upgrade, not a prerequisite. Every tool I’ve built that required a model to do anything useful had a worse adoption curve than the ones that shipped a usable floor.
Memory: the context window is a budget, not a store
Context windows get treated as memory and they aren’t. They’re a per-request budget you re-pay every turn.
Cascade keeps thread state outside the window with rolled-up summaries at turn, day, week, and month granularity. Recalling last week becomes a lookup against the week layer instead of a replay of the full history. The rollup is doing compression, and it’s also doing implicit supersession, because a week summary reflects the reversal that the original turn has no idea happened.
What the month settled
The work that pays off most is in the seams around a good agent, not in replacing it. Claude Code is already strong. What was missing was routing, retrieval, memory, and automation between the pieces.
None of that is machine learning. It’s caching, ranking, cost control, and failure handling. Ordinary systems engineering pointed at a new substrate.
That’s the part I keep seeing underrated. The model is one component, and two decades of building backends transfers almost completely to everything around it. I’ve made that argument properly in From ML Experiments to Production AI; Cascade is the working proof.