- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
10-08-2025 12:59 PM
Hey @PiotrM,
Firstly, have you checked the docs out for Managing Model Serving Endpoints?
https://learn.microsoft.com/en-us/azure/databricks/machine-learning/model-serving/manage-serving-end...
I just had a read through. You can certainly set up budgets to monitor them, this can help with preventing costs spiralling! 🙂. I appreciate you've mentioned about the system tables.
This article seems really promising: https://docs.databricks.com/aws/en/ai-gateway/configure-ai-gateway-endpoints 👀🙂... (I'm certain we've got to be onto a winner with this)
If that doesn't quite cut the mustard, perhaps we could also look at the actual token usage per user. Perhaps this can be throttled somehow 🤔.
All the best,
BS