Hi Paolo,

Thanks for sharing your experience! It’s great to hear that the RAG architecture on Databricks is working well in production.

I’m particularly interested in how you approached the chunking strategy for document ingestion into the Vector Search Index database in Databricks. Could you share some insights on how you fine-tuned the chunking process to optimize retrieval performance? Specifically:

  • What chunk sizes worked best for your use case?
  • Did you implement any custom logic for breaking down structured vs. unstructured data?
  • How did you balance retrieval accuracy with performance?

Looking forward to your thoughts!

Thanks

Mantu

Mantu S