It sounds like I need to create the "service wrapper" that will do the pre-processing and fetching of env vars etc.  I'll deploy that using a model serving endpoint, serverless for speed, then each sub model will be on its own compute cluster that scales independently.

Thanks for the great feedback