Skip to content
← All posts
cascadeaiclaude-codeopen-source

Starting Cascade: Early Lessons Extending Claude Code

3 min read

Cascade began as scripts to paper over my own problems with Claude Code. Four lessons from the first month, including why multi-account routing is a bookkeeping problem and why the daemon changed the whole design.

I use Claude Code every day. It’s genuinely good, and I kept hitting the same walls. One account hits its limit mid-task. Token spend on routine operations adds up faster than you notice. Context resets between sessions, so work that should carry over starts from zero.

Cascade started as a handful of scripts to paper over those gaps for me. Four things I learned in the first month, in the order they hurt.

Multi-account routing is bookkeeping, not intelligence

I expected this to be interesting. It isn’t. It’s a ledger problem.

What you need to know at dispatch time is which account has quota left, at what cost, for which class of task. That’s it. There’s no clever model selection heuristic hiding in there. Once you have an accurate table of current state, the routing decision falls out.

The hard part is keeping the table honest while several sessions run concurrently. Quota is consumed by processes that don’t know about each other, and any state you cache is stale the moment you read it. That’s a distributed-state problem wearing a very small hat, and it’s where all the actual bugs lived.

Which is a pattern I’d see again: in agent tooling, the thing that looks like it needs AI usually needs a correct data structure.

A persistent daemon changes the whole design

This was the decision that reshaped the project.

A chat window can’t own state, because it dies. Once a background process owns session state instead, you get restart survival, resumable interrupted work, and coordination across sessions that a single window could never track. Not as features you build, but as consequences of where the state lives.

Before the daemon I was writing increasingly elaborate mechanisms to reconstruct context on startup. After it, that entire category of code stopped existing. When you find yourself writing a lot of code to rebuild something, the real question is usually why it was allowed to be destroyed.

Most token spend is invisible

A lot of cost hides in routine operations you never think about, because individually they’re trivial and you’re watching the expensive calls.

A transparent proxy that trims obvious overhead saved more than I expected, and it did it without changing how I work. That last part is why it was worth building. An optimization you have to remember to use isn’t an optimization, it’s a chore. This one compounds across every session because it’s on by default and invisible.

Hooks are where to cut the seam

Pre and post hooks on each event turned out to be the right extension point for quota checks, logging, and small automations.

The alternative was forking Claude Code, and forking a tool that’s actively developed means signing up to merge someone else’s changes forever. Hooks let Cascade stay small and let upgrades stay boring. Picking the seam correctly is most of what makes a tool survivable.

Where this actually started

Early Cascade was duct tape. It worked for me and broke for everyone else: hardcoded paths, assumptions about my machine, error handling that consisted of the script exiting.

The month after that was about turning private scripts into something other people could install and trust. That’s a different kind of engineering than making something work once, and it’s the part that makes a tool real. It’s also the part nobody talks about, because “I made it work” is a better story than “I spent three weeks on installation edge cases and error messages.”

Routing tiers, hybrid retrieval, and memory that survives a session all came later. That’s in Cascade Today.