<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>topic Re: CosmosGenie — Your Universe, Answered (Genie-Powered App Challenge ) in Community Articles</title>
    <link>https://community.databricks.com/t5/community-articles/cosmosgenie-your-universe-answered-genie-powered-app-challenge/m-p/168035#M1545</link>
    <description>&lt;P&gt;Hi &lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/166008"&gt;@sudiptob-DA&lt;/a&gt;,&lt;/P&gt;&lt;P&gt;This is a fantastic use case for the&amp;nbsp;Genie Agent. You’ve perfectly illustrated the most important lesson in modern AI-augmented data platforms:&amp;nbsp;the quality of the natural language experience is entirely dependent on the quality of the semantic layer.&lt;/P&gt;&lt;P&gt;I really like how you’ve structured your ingestion strategy—using both&amp;nbsp;Lakeflow Jobs&amp;nbsp;and&amp;nbsp;Spark Declarative Pipelines (SDP)&amp;nbsp;to manage the complexity of eight disparate API sources. It’s a great example of using the right tool for the specific latency requirements of each data stream.&lt;/P&gt;&lt;P&gt;Since you’re using&amp;nbsp;Genie&amp;nbsp;to handle unpredictable, ad-hoc queries, how are you managing the "Golden" layer optimization to balance latency with query accuracy? I’m particularly curious if you found that specific column-level metadata (like descriptions and constraints) played a bigger role in Genie’s accuracy than the data-quality expectations themselves?&lt;/P&gt;</description>
    <pubDate>Wed, 09 Sep 2026 05:51:45 GMT</pubDate>
    <dc:creator>Khasim_1</dc:creator>
    <dc:date>2026-09-09T05:51:45Z</dc:date>
    <item>
      <title>CosmosGenie — Your Universe, Answered (Genie-Powered App Challenge )</title>
      <link>https://community.databricks.com/t5/community-articles/cosmosgenie-your-universe-answered-genie-powered-app-challenge/m-p/168002#M1544</link>
      <description>&lt;P&gt;My son asked me one evening whether any asteroids were going to hit Earth. The&lt;BR /&gt;data to answer him exists — NASA publishes it daily, for free — but it lives in&lt;BR /&gt;JSON behind API keys, in units like astronomical units and X-ray flux classes,&lt;BR /&gt;built for people who already know what they're looking for. A curious ten-year-old&lt;BR /&gt;is not that person.&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;CosmosGenie&lt;/STRONG&gt; fixes the interface, not the data. Ask anything about space in plain&lt;BR /&gt;English — asteroids, eclipses, the Moon, planetary line-ups, rocket launches —&lt;BR /&gt;and it queries live NASA, USNO and JPL data and answers in a sentence.&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;**How it's built.** &lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;Eight free public APIs feed a full medallion architecture on&lt;BR /&gt;Databricks Free Edition. Two ingestion paths — standalone notebooks on a Lakeflow&lt;BR /&gt;Job, and a Lakeflow Spark Declarative Pipeline with bronze → silver → gold and&lt;BR /&gt;data-quality expectations — land eight silver tables (the source of record) and&lt;BR /&gt;four gold tables tuned for Genie.&amp;nbsp;&lt;/P&gt;&lt;P&gt;[ Free Public APIs ] [ Databricks Free Edition ]&lt;/P&gt;&lt;P&gt;NASA NeoWs ─────┐ Lakeflow SDP Pipeline&lt;BR /&gt;NASA DONKI ─────┤ ┌─ bronze/ (raw API pull)&lt;BR /&gt;├──► SDP Pipeline ────►├─ silver/ (clean + DQ expectations)&lt;BR /&gt;│ └─ gold/ (business logic, KPIs)&lt;BR /&gt;USNO Moon ─────┐&lt;BR /&gt;JPL / Curated ───┤&lt;BR /&gt;NASA Eclipse ────┼──► Lakeflow Job ────► Delta Tables (cosmos.space.*)&lt;BR /&gt;The Space Devs ──┤&lt;BR /&gt;Spaceflight News ┘&lt;BR /&gt;│&lt;BR /&gt;┌─────────┴──────────┐&lt;BR /&gt;Genie Space&lt;BR /&gt;└─────────┬──────────┘&lt;BR /&gt;Streamlit App&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;&lt;U&gt;&lt;STRONG&gt;**The app** is Streamlit on Databricks Apps with the Genie Agent attached as a&lt;/STRONG&gt;&lt;/U&gt;&lt;BR /&gt;&lt;U&gt;&lt;STRONG&gt;resource: &lt;/STRONG&gt;&lt;/U&gt;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;An aurora theme, a live KPI bar, an interactive 12-month events timeline&lt;BR /&gt;where clicking an event asks a question, and threaded chat that returns prose, the&lt;BR /&gt;generated SQL, a table and a tailored visual.&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;**Why Genie.**&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;Remove it and what's left is four numbers and a timeline — no&lt;BR /&gt;dashboard, no filter panel, no pre-built report. Every ranking and caveat is&lt;BR /&gt;generated live from a question nobody wrote in advance. Almost all the effort went&lt;BR /&gt;into the semantic layer: column comments, instructions and sample questions.&lt;BR /&gt;&lt;BR /&gt;&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;What I learned.&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;The natural-language part is only as good as the semantic layer behind it.&lt;SPAN&gt;The payoff is that once Genie understands the data this well, it reliably handles questions I never anticipated and never wrote an example for. It even explains its own reasoning. The lesson: the model isn't the hard part — describing your data clearly is. Invest there, and the natural-language experience takes care of itself.&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Tue, 08 Sep 2026 20:00:31 GMT</pubDate>
      <guid>https://community.databricks.com/t5/community-articles/cosmosgenie-your-universe-answered-genie-powered-app-challenge/m-p/168002#M1544</guid>
      <dc:creator>sudiptob-DA</dc:creator>
      <dc:date>2026-09-08T20:00:31Z</dc:date>
    </item>
    <item>
      <title>Re: CosmosGenie — Your Universe, Answered (Genie-Powered App Challenge )</title>
      <link>https://community.databricks.com/t5/community-articles/cosmosgenie-your-universe-answered-genie-powered-app-challenge/m-p/168035#M1545</link>
      <description>&lt;P&gt;Hi &lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/166008"&gt;@sudiptob-DA&lt;/a&gt;,&lt;/P&gt;&lt;P&gt;This is a fantastic use case for the&amp;nbsp;Genie Agent. You’ve perfectly illustrated the most important lesson in modern AI-augmented data platforms:&amp;nbsp;the quality of the natural language experience is entirely dependent on the quality of the semantic layer.&lt;/P&gt;&lt;P&gt;I really like how you’ve structured your ingestion strategy—using both&amp;nbsp;Lakeflow Jobs&amp;nbsp;and&amp;nbsp;Spark Declarative Pipelines (SDP)&amp;nbsp;to manage the complexity of eight disparate API sources. It’s a great example of using the right tool for the specific latency requirements of each data stream.&lt;/P&gt;&lt;P&gt;Since you’re using&amp;nbsp;Genie&amp;nbsp;to handle unpredictable, ad-hoc queries, how are you managing the "Golden" layer optimization to balance latency with query accuracy? I’m particularly curious if you found that specific column-level metadata (like descriptions and constraints) played a bigger role in Genie’s accuracy than the data-quality expectations themselves?&lt;/P&gt;</description>
      <pubDate>Wed, 09 Sep 2026 05:51:45 GMT</pubDate>
      <guid>https://community.databricks.com/t5/community-articles/cosmosgenie-your-universe-answered-genie-powered-app-challenge/m-p/168035#M1545</guid>
      <dc:creator>Khasim_1</dc:creator>
      <dc:date>2026-09-09T05:51:45Z</dc:date>
    </item>
  </channel>
</rss>

