I hope this helps for those who are trying to keep LLM costs under control on batch workloads. I have been looking at ai_decide (Beta) as a router rather than as a classifier, and it fits that job well. Sharing the pattern and the open questions, since I have not benchmarked it at scale yet.
THE IDEA
Most rows in a typical batch job (ticket triage, review summaries, document extraction) do not need the big model. A small, fast decision step can look at each row and choose which model should handle it. ai_decide takes text plus a set of questions and returns a choice, a probability or a score, and it is documented as faster and cheaper than a general LLM for decisions. The Databricks announcement lists "routing prompts to the right model" as a use case.
Endpoints (from the supported models page; re-check in your workspace with WorkspaceClient().serving_endpoints.list(), because names change quickly):
- databricks-glm-5-3-flash : cheaper, multimodal, reasoning always on
- databricks-glm-5-3 : the full model, text only, reasoning effort configurable
STEP 1: ask ai_decide for a tier and a confidence
ai_decide(state, questions, options) returns a VARIANT. For a choice question you get the chosen label, per-label probabilities and a confidence.
SELECT id, prompt,
ai_decide(
prompt,
'{
"tier": {
"type": "choice",
"instructions": "Which model tier is needed to answer this request well?",
"criteria": {
"flash": "Short factual lookup, simple extraction, classification, or rewriting. No multi-step reasoning.",
"full": "Multi-step reasoning, ambiguous or conflicting inputs, long synthesis, code generation, or anything where an error is costly."
}
}
}',
map('version', '1.0')
) AS d
FROM prompts_to_process;
Pull the fields out of the VARIANT:
d:response:answers:tier:choice::string -- 'flash' or 'full'
d:response:answers:tier:confidence::double -- how sure the router is
STEP 2: send each row to the model it was routed to
ai_query needs the endpoint name as a constant, so you cannot pass the routed name in as a column value. Two options: (a) a CASE with one ai_query call per branch, or (b) two statements, one filtered on tier = 'flash' and one on tier = 'full', writing into the same results table. I prefer (b): each batch is homogeneous, you can size and monitor each endpoint separately, and you do not depend on how the engine evaluates both branches of a CASE.
CREATE OR REPLACE TABLE routed AS
SELECT id, prompt,
d:response:answers:tier:choice::string AS tier,
d:response:answers:tier:confidence::double AS conf
FROM (SELECT id, prompt, ai_decide(prompt, '<questions json from step 1>', map('version','1.0')) AS d
FROM prompts_to_process);
-- cheap path
INSERT INTO results
SELECT id, 'flash' AS model, ai_query('databricks-glm-5-3-flash', prompt) AS answer
FROM routed WHERE tier = 'flash' AND conf >= 0.7;
-- everything else goes to the full model, including low-confidence rows
INSERT INTO results
SELECT id, 'full' AS model, ai_query('databricks-glm-5-3', prompt) AS answer
FROM routed WHERE tier = 'full' OR conf < 0.7;
The design choice I think matters most: send low-confidence rows to the full model. A router that is unsure should fail expensive, not fail wrong. Tune the 0.7 threshold on your own data.
THINGS TO CHECK BEFORE TRUSTING THIS IN PRODUCTION
1. Label a sample (a few hundred rows) by running both models and comparing. The router only saves money if flash answers are good enough on the rows it keeps. Measure answer quality per tier, not just router accuracy.
2. Watch the cost of the router itself. ai_decide is cheap per call, but on very short prompts it may be a meaningful share of the flash call.
3. Log tier, confidence and model per row so you can see the real flash/full split and recompute savings.
4. ai_decide is Beta, only available in some regions, and not on Databricks SQL Classic, so check your warehouse type and region first.
5. Re-verify endpoint names and pricing. Both are pay-per-token endpoints and the lineup moves fast.
WHAT I HAVE NOT TESTED
I have not benchmarked routing accuracy or the real cost split at scale, so treat the threshold and the criteria wording as a starting point. The criteria text moves results the most, so it is worth iterating on it with a labelled sample.
QUESTIONS FOR THE COMMUNITY
- Has anyone tried ai_decide as a router at volume? What flash/full split did you land on?
- Any experience with score-type questions (a 0 to 2 difficulty score) versus a choice question for routing? I picked choice because it also returns a confidence I can threshold on.
- Are you routing on prompt content alone, or also on metadata such as customer tier or SLA?
References: ai_decide SQL function reference and announcement blog, and the Foundation Model APIs supported models page on docs.databricks.com.