how to reduce scale to zero time in MLFlow Serving
Options
- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
07-28-2024 02:53 AM - edited 07-28-2024 02:57 AM
Hi,
I am deploying MLflow models using Databrick serverless serving but seems servers scale down to 0 only after 30 minute of inactivity. Is there any way to reduce this time?
Also, Is it possible to deploy multiple models under single endpoint. I want to run multiple models in one endpoint to reduce cost like AWS sage maker provides multi-model deployment functionality.
Appreciate any help.
Regards,
Sanjay