how to reduce scale to zero time in MLFlow Serving

sanjay
Valued Contributor II

Hi,

I am deploying MLflow models using Databrick serverless serving but seems servers scale down to 0 only after 30 minute of inactivity. Is there any way to reduce this time?

Also, Is it possible to deploy multiple models under single endpoint. I want to run multiple models in one endpoint to reduce cost like AWS sage maker provides multi-model deployment functionality.

Appreciate any help.

Regards,
Sanjay