Hi, some notes from running an internal agent platform on Databricks, in the order where they bit us hardest:
1. Evaluation before deployment is the one that gets skipped. Build a small labeled set (even 50 to 100 real questions with the answer and the source that supports it) and run it through MLflow evaluation on every change to the prompt, the retriever or the model. Without that, every later change is a guess. Keep retrieval metrics (did the right chunk come back) separate from answer quality (was the answer right), otherwise you cannot tell which half broke.
2. Model serving choice is mostly a cost and rate-limit question, not a quality one. Pay-per-token endpoints are fine for prototypes and spiky traffic, but a steady workload usually wants provisioned throughput, and anything user-facing needs a fallback endpoint behind the same AI Gateway route so a single provider throttle does not become an outage. Also check early that the endpoints you plan to use are actually enabled for your account type, since some third-party models are gated differently from open-weight ones.
3. Vector Search and RAG: most of the quality problems we saw were chunking and metadata filtering, not the embedding model. Filter on permissions or tenant metadata at query time instead of post-filtering, and re-run your eval set whenever you re-chunk.
4. Governance: put the tools and data the agent can touch under Unity Catalog and run the agent with the least-privileged identity (a service principal, or on-behalf-of the user where per-user row and column rules must apply). Log requests and responses through AI Gateway inference tables so you have an audit trail from day one.
5. Cost and monitoring after launch: tag every endpoint and app so spend can be attributed to a team, and track cost per successful task, not just cost per token. Retries and long tool loops are where the bill quietly grows.
Biggest lesson overall: the prototype to production gap is mostly evaluation, identity and cost attribution, not model choice. Happy to go deeper on any of these.