Architecture ยท Cost and safety

Cascading AI Architecture: Rules, Classic ML, and LLM Escalation

A cascading AI architecture sends each case to the cheapest layer that can handle it: deterministic logic first, then classic machine learning (ML) triage, and a large language model (LLM) only when confidence is low.

Why teams move away from an agent for every task

For a while the instinct was to put an agent in charge of every task. Then the bills and the latency came in, and much of that work could have been handled far more cheaply by classic ML.

The three tiers

  1. Deterministic logicRules and checks that decide what can be decided without a model, including guardrails for out-of-scope queries.
  2. Classic ML triageA conventional model handles the routine cases.
  3. Confidence-gated LLM escalationOnly cases the cheaper layers are not confident about go to an LLM.

The upgrade most teams miss: a teacher-student loop

Most teams stop at escalating to the LLM only when needed. The real improvement is a teacher-student loop, where the LLM's output on hard cases retrains the simpler and cheaper model beneath it. The LLM's job is not only to answer what the cheap tier cannot. It also makes the cheap tier need the LLM less next time.

The trap: treating the confidence threshold as a technical setting

The confidence threshold decides how many cases reach the expensive layer. It is an economic trade-off between cost and risk, and it can build or break user trust. I set it by looking at what a wrong answer costs and what an escalation costs, not by picking a default value.

In clinical settings

In clinical pipelines the cascade also gives me somewhere to put safety: deterministic guardrails for out-of-scope questions, confidence scoring so the system can refuse when it should, and monitoring for drift in production. I also review existing designs, comparing a cascade against an approach where everything goes to an LLM.

Where I present this

The Evolution of the Data Scientist: From Agentic Euphoria to Cascading AI Architecture

hayaData Conference, Israel.

Work with me

Cascading AI architecture is one of the services I offer to early-stage HealthTech teams, either as a focused architecture review or as part of a fractional AI lead engagement.

Frequently asked questions

What is a cascading AI architecture?

A cascading AI architecture sends each case to the cheapest layer that can handle it: deterministic logic first, then classic machine learning (ML) triage, and a large language model (LLM) only when confidence is low.

Why not send every task to an LLM?

Because the cost and latency add up, and much of that work can be handled far more cheaply by classic ML or by plain rules. Sending everything to an LLM also gives up the chance to put guardrails in front of it.

What is a teacher-student loop?

It is a retraining loop in which the LLM's output on hard cases is used to retrain the simpler and cheaper model beneath it. Over time the cheap layer needs the LLM less.

How do I set the confidence threshold?

Treat it as an economic trade-off between the cost of an escalation and the cost of a wrong answer, not as a technical default. It also affects user trust, so I set it by looking at what each kind of error costs in your case.

Is a cascade only for healthcare?

No, the pattern is general. I use it mainly for clinical AI, where cost and safety both matter.

Paying too much for an agent that does everything?

I can review your current design and tell you which cases belong in which tier.

Book a Strategy Call MayaM@MalamudAI.com