Comment
New Contributor

Great post! Do you have a recommended set of best practices for clustering queries when doing embedding model evaluation and selection? I've found that without clustering, users have to choose to either use a broad average on their whole dataset, or to look at individual queries to understand their performance (which is hard to scale).

I recently co-authored a blog post with another open source vector database company about my company's python package that helps users to find the most common query patterns with high error rate in their vector database system, and I'd love to give equal attention to Databricks and your community. Would you be open to connecting further about a potential blog post or collaboration between my team at BluelightAI and Databricks? Example post Link