Khasim_1
New Contributor III

Hi @sudiptob-DA,

This is a fantastic use case for the Genie Agent. You’ve perfectly illustrated the most important lesson in modern AI-augmented data platforms: the quality of the natural language experience is entirely dependent on the quality of the semantic layer.

I really like how you’ve structured your ingestion strategy—using both Lakeflow Jobs and Spark Declarative Pipelines (SDP) to manage the complexity of eight disparate API sources. It’s a great example of using the right tool for the specific latency requirements of each data stream.

Since you’re using Genie to handle unpredictable, ad-hoc queries, how are you managing the "Golden" layer optimization to balance latency with query accuracy? I’m particularly curious if you found that specific column-level metadata (like descriptions and constraints) played a bigger role in Genie’s accuracy than the data-quality expectations themselves?

Data Architect | 13 Years Domain Expertise | Databricks SA Champion Cohort