My first real machine learning project was a TensorFlow model in 2016 that predicted server load for a client’s infrastructure. It worked great in a notebook. Getting it into production took three more months.
That gap has defined most of my AI work since, and the thing I keep arguing is that it never closed. It moved.
2016 to 2020: the model was always the easy part
Sentiment analysis on customer feedback. Image classification for a manufacturing client. A recommendation engine for e-commerce. Every one followed the same arc, and it was never the training that hurt.
Data pipelines that had to run reliably at 3am. Model serving that didn’t fall over under concurrency. Drift monitoring, because the model that was accurate in March is quietly wrong by September and nothing tells you. Edge cases the notebook never saw, because the notebook was fed a clean CSV.
That’s 80% of the effort and none of it is machine learning. It’s ETL, capacity planning, observability, and failure handling. The same work I’d been doing since 2004, pointed at a different payload.
What LLMs changed, and what they didn’t
The 2023 wave genuinely changed the front half. Instead of collecting labeled data and training a model per use case, you prompt a foundation model and get something useful immediately. The tradeoff moved from training time and data requirements to inference cost and latency, which for most business applications is the better trade.
What it did not change is everything downstream of “the model produced output.”
If anything that half got harder. A classifier gave you a number you could threshold. An LLM gives you free text that’s usually right, occasionally confidently wrong, and non-deterministic across identical inputs. You can’t unit test it the way you’d test a function. Evaluating it requires building an apparatus that didn’t need to exist before.
The pattern that works in production is narrower than the demos suggest: LLMs for tasks where “pretty good” is acceptable and a human reviews the output. Summarization, code review suggestions, extraction from unstructured text. Cases where the AI does the bulk and a person catches the misses. I’ve shipped several of those and they hold up.
The thing I got wrong
I assumed the orchestration layer would be thin. Route the request, call the API, return the result.
It’s most of the system. Building Cascade, the actual work turned out to be context assembly under a token budget, routing between model tiers so cost doesn’t run away, retrieval good enough that the model gets the right material, memory that survives a session ending, and graceful degradation when a provider is down or rate-limited.
Every one of those is a systems problem with a well-established shape. Context assembly is ranking. Model routing is cost-aware scheduling. Retrieval is search, carrying all the ranking and fusion problems search has had for thirty years. Memory is caching and invalidation. Degradation is circuit breaking.
I’ve written up the specifics elsewhere, in how the routing tiers work and where retrieval stops being enough, but the shape is the point. None of it required me to learn machine learning. It required me to already know distributed systems.
Why I think this matters
There’s a hiring assumption that AI engineering is a specialization you enter from the ML side. My experience is close to the opposite. The scarce skill isn’t understanding transformers. It’s having built enough production systems to know which failure modes are coming and to design for them before they arrive.
The people I’ve watched struggle most with agentic systems are strong ML practitioners who’ve never run anything at 3am. The ones who pick it up fastest are backend engineers who’ve been paged enough times to be appropriately paranoid.
Two decades of building systems turned out to be better preparation for this than a machine learning background would have been. That isn’t what I expected in 2016, and it’s the most useful thing I’ve figured out since.