cancel
Showing results for 
Search instead for 
Did you mean: 
Generative AI
Explore discussions on generative artificial intelligence techniques and applications within the Databricks Community. Share ideas, challenges, and breakthroughs in this cutting-edge field.
cancel
Showing results for 
Search instead for 
Did you mean: 

Routing between GLM 5.3 Flash and GLM 5.3 with ai_decide: a cheap router in front of ai_query

DoTA
Valued Contributor II

I hope this helps for those who are trying to keep LLM costs under control on batch workloads. I have been looking at ai_decide (Beta) as a router rather than as a classifier, and it fits that job well. Sharing the pattern and the open questions, since I have not benchmarked it at scale yet.

 

THE IDEA

 

Most rows in a typical batch job (ticket triage, review summaries, document extraction) do not need the big model. A small, fast decision step can look at each row and choose which model should handle it. ai_decide takes text plus a set of questions and returns a choice, a probability or a score, and it is documented as faster and cheaper than a general LLM for decisions. The Databricks announcement lists "routing prompts to the right model" as a use case.

 

Endpoints (from the supported models page; re-check in your workspace with WorkspaceClient().serving_endpoints.list(), because names change quickly):

- databricks-glm-5-3-flash : cheaper, multimodal, reasoning always on

- databricks-glm-5-3 : the full model, text only, reasoning effort configurable

 

STEP 1: ask ai_decide for a tier and a confidence

 

ai_decide(state, questions, options) returns a VARIANT. For a choice question you get the chosen label, per-label probabilities and a confidence.

 

SELECT id, prompt,

  ai_decide(

    prompt,

    '{

      "tier": {

        "type": "choice",

        "instructions": "Which model tier is needed to answer this request well?",

        "criteria": {

          "flash": "Short factual lookup, simple extraction, classification, or rewriting. No multi-step reasoning.",

          "full": "Multi-step reasoning, ambiguous or conflicting inputs, long synthesis, code generation, or anything where an error is costly."

        }

      }

    }',

    map('version', '1.0')

  ) AS d

FROM prompts_to_process;

 

Pull the fields out of the VARIANT:

 

d:response:answers:tier:choice::string -- 'flash' or 'full'

d:response:answers:tier:confidence::double -- how sure the router is

 

STEP 2: send each row to the model it was routed to

 

ai_query needs the endpoint name as a constant, so you cannot pass the routed name in as a column value. Two options: (a) a CASE with one ai_query call per branch, or (b) two statements, one filtered on tier = 'flash' and one on tier = 'full', writing into the same results table. I prefer (b): each batch is homogeneous, you can size and monitor each endpoint separately, and you do not depend on how the engine evaluates both branches of a CASE.

 

CREATE OR REPLACE TABLE routed AS

SELECT id, prompt,

       d:response:answers:tier:choice::string AS tier,

       d:response:answers:tier:confidence::double AS conf

FROM (SELECT id, prompt, ai_decide(prompt, '<questions json from step 1>', map('version','1.0')) AS d

      FROM prompts_to_process);

 

-- cheap path

INSERT INTO results

SELECT id, 'flash' AS model, ai_query('databricks-glm-5-3-flash', prompt) AS answer

FROM routed WHERE tier = 'flash' AND conf >= 0.7;

 

-- everything else goes to the full model, including low-confidence rows

INSERT INTO results

SELECT id, 'full' AS model, ai_query('databricks-glm-5-3', prompt) AS answer

FROM routed WHERE tier = 'full' OR conf < 0.7;

 

The design choice I think matters most: send low-confidence rows to the full model. A router that is unsure should fail expensive, not fail wrong. Tune the 0.7 threshold on your own data.

 

THINGS TO CHECK BEFORE TRUSTING THIS IN PRODUCTION

 

1. Label a sample (a few hundred rows) by running both models and comparing. The router only saves money if flash answers are good enough on the rows it keeps. Measure answer quality per tier, not just router accuracy.

2. Watch the cost of the router itself. ai_decide is cheap per call, but on very short prompts it may be a meaningful share of the flash call.

3. Log tier, confidence and model per row so you can see the real flash/full split and recompute savings.

4. ai_decide is Beta, only available in some regions, and not on Databricks SQL Classic, so check your warehouse type and region first.

5. Re-verify endpoint names and pricing. Both are pay-per-token endpoints and the lineup moves fast.

 

WHAT I HAVE NOT TESTED

 

I have not benchmarked routing accuracy or the real cost split at scale, so treat the threshold and the criteria wording as a starting point. The criteria text moves results the most, so it is worth iterating on it with a labelled sample.

 

QUESTIONS FOR THE COMMUNITY

 

- Has anyone tried ai_decide as a router at volume? What flash/full split did you land on?

- Any experience with score-type questions (a 0 to 2 difficulty score) versus a choice question for routing? I picked choice because it also returns a confidence I can threshold on.

- Are you routing on prompt content alone, or also on metadata such as customer tier or SLA?

 

References: ai_decide SQL function reference and announcement blog, and the Foundation Model APIs supported models page on docs.databricks.com.

2 REPLIES 2

ThiamLee
Contributor

Really useful pattern. I like the “low confidence → full model” approach since it prioritizes answer quality over blindly optimizing cost. Would be interesting to see the actual cost savings and flash/full split once you benchmark it at scale.

mukul1409
Contributor II

Interesting approach. I like using the confidence from the choice output as a routing signal rather than making the router itself another complex LLM workflow.

I'm also interested in the choice vs score approach. For routing, choice + confidence seems easier to operationalize, especially if we log the decision, confidence, selected model, and final outcome for later evaluation.

I haven't tested this pattern hands-on yet, but this gives me a good use case to experiment with ai_decide.

Mukul Chauhan