Are serverless endpoints possible in this Technical Blog post by qianyu?
- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
05-03-2025 10:35 AM
Hi, are serverless endpoints possible for Whisper and Llama in this Technical Blog post by qianyu?
Thanks!
- Labels:
-
GenAI Generation AI
- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
05-05-2025 12:12 AM
Model serving is already serverless enabled. You can even set a budget for this (public preview).
https://docs.databricks.com/aws/en/machine-learning/model-serving/manage-serving-endpoints
https://docs.databricks.com/aws/en/machine-learning/model-serving/foundation-model-overview
Or do you mean something else (like using a databricks serverless sql/eng instance)?
- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
05-05-2025 08:24 AM
Thanks for the clarification and links.
Bottom line, what I am trying to avoid is spinning up AWS resources in the background that will incur ongoing charges until I track them down and terminate them (I am a new Databricks customer and still trying to navigate the billing/cost system). The "Serve this model" UI looked suspiciously like it was going to do this but on second look, maybe not.
I am just wanting to confirm my only costs, Databricks or AWS, for Whisper and Llama in this Technical Blog post will be only for the short duration they will be used. Thanks!
- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
05-06-2025 12:15 AM
https://www.databricks.com/product/pricing/foundation-model-serving
We don´t serve models on databricks, but as far as i can see you pay per input/output tokens (for foundation models).
For classic models:
https://www.databricks.com/product/pricing/model-serving