- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
4 weeks ago
A few connection options, roughly fastest to slowest for an external agent that already emits SQL:
1. Databricks SQL Connector for Python against a Serverless SQL Warehouse kept warm (auto-stop 5-10 min). Persistent connection, no per-call compute spin-up, no polling - usually the lowest-latency path for "agent generates SQL -> run it -> return rows."
2. Statement Execution API, called correctly: set wait_timeout=30s with on_wait_timeout=CONTINUE so short queries return inline instead of you polling GET /statements/{id} in a loop. Use disposition=EXTERNAL_LINKS + format=ARROW_STREAM above a few MB. Most "the API is slow" cases are warehouse cold start or a fixed-sleep poll loop.
3. Genie Conversation API if you want Databricks to own NL->SQL. Less control over the generated SQL, but you skip building that layer; latency is higher because there is an LLM call in the path.
4. Databricks MCP server (UC functions / Genie space) if your agent framework speaks MCP - same underlying latency as the above.
Latency checklist before changing anything: use Serverless (not classic - 3-5 min cold start); confirm the warehouse is running when the agent fires; time the generated SQL directly in the SQL editor to separate query time from transport; check whether the time is actually in your agent's LLM call. Biggest single win for us was dropping a poll loop for wait_timeout + a warm serverless warehouse.