<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>All Community Articles posts</title>
    <link>https://community.databricks.com/t5/community-articles/bd-p/Knowledge-Sharing-Hub</link>
    <description>All Community Articles posts</description>
    <pubDate>Tue, 01 Sep 2026 23:18:20 GMT</pubDate>
    <dc:creator>Knowledge-Sharing-Hub</dc:creator>
    <dc:date>2026-09-01T23:18:20Z</dc:date>
    <item>
      <title>Re: Building a Databricks Solutions Architect Genie Agent</title>
      <link>https://community.databricks.com/t5/community-articles/building-a-databricks-solutions-architect-genie-agent/m-p/167179#M1518</link>
      <description>&lt;P&gt;&lt;div class="video-embed-center video-embed"&gt;&lt;iframe class="embedly-embed" src="https://cdn.embedly.com/widgets/media.html?src=https%3A%2F%2Fwww.youtube.com%2Fembed%2F8dk7Ge6YFuw%3Ffeature%3Doembed&amp;amp;display_name=YouTube&amp;amp;url=https%3A%2F%2Fwww.youtube.com%2Fwatch%3Fv%3D8dk7Ge6YFuw&amp;amp;image=https%3A%2F%2Fi.ytimg.com%2Fvi%2F8dk7Ge6YFuw%2Fhqdefault.jpg&amp;amp;type=text%2Fhtml&amp;amp;schema=youtube" width="200" height="112" scrolling="no" title="Databricks Genie Solutions Architect" frameborder="0" allow="autoplay; fullscreen; encrypted-media; picture-in-picture" allowfullscreen="true"&gt;&lt;/iframe&gt;&lt;/div&gt;&lt;/P&gt;</description>
      <pubDate>Tue, 01 Sep 2026 20:12:42 GMT</pubDate>
      <guid>https://community.databricks.com/t5/community-articles/building-a-databricks-solutions-architect-genie-agent/m-p/167179#M1518</guid>
      <dc:creator>saketsuman</dc:creator>
      <dc:date>2026-09-01T20:12:42Z</dc:date>
    </item>
    <item>
      <title>Re: Research It!</title>
      <link>https://community.databricks.com/t5/community-articles/research-it/m-p/167176#M1517</link>
      <description>&lt;P&gt;&lt;div class="video-embed-center video-embed"&gt;&lt;iframe class="embedly-embed" src="https://cdn.embedly.com/widgets/media.html?src=https%3A%2F%2Fwww.youtube.com%2Fembed%2FhtpcaHBCAIQ%3Ffeature%3Doembed&amp;amp;display_name=YouTube&amp;amp;url=https%3A%2F%2Fwww.youtube.com%2Fwatch%3Fv%3DhtpcaHBCAIQ&amp;amp;image=https%3A%2F%2Fi.ytimg.com%2Fvi%2FhtpcaHBCAIQ%2Fhqdefault.jpg&amp;amp;type=text%2Fhtml&amp;amp;schema=youtube" width="200" height="112" scrolling="no" title="Genie Research Agent - Do R&amp;amp;D at an enterprise level" frameborder="0" allow="autoplay; fullscreen; encrypted-media; picture-in-picture" allowfullscreen="true"&gt;&lt;/iframe&gt;&lt;/div&gt;&lt;/P&gt;</description>
      <pubDate>Tue, 01 Sep 2026 20:03:49 GMT</pubDate>
      <guid>https://community.databricks.com/t5/community-articles/research-it/m-p/167176#M1517</guid>
      <dc:creator>saketsuman</dc:creator>
      <dc:date>2026-09-01T20:03:49Z</dc:date>
    </item>
    <item>
      <title>Solution Accelerator Series | Subscriber Churn Prediction</title>
      <link>https://community.databricks.com/t5/community-articles/solution-accelerator-series-subscriber-churn-prediction/m-p/167143#M1516</link>
      <description>&lt;P&gt;&lt;SPAN&gt;Subscriber attrition can be difficult to anticipate. The &lt;/SPAN&gt;&lt;STRONG&gt;Subscriber Churn Prediction Solution Accelerator&lt;/STRONG&gt;&lt;SPAN&gt; shows how to analyze behavioral data, identify subscribers at increased risk of cancellation, and use machine learning to estimate churn likelihood and understand the factors that contribute to that risk.&lt;/SPAN&gt;&lt;/P&gt;
&lt;H3&gt;&lt;STRONG&gt;With this Accelerator, you get&lt;/STRONG&gt;&lt;/H3&gt;
&lt;UL&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Ready-to-use resources:&lt;/STRONG&gt; &lt;A href="https://notebooks.databricks.com/notebooks/RCG/Survival/index.html?_ga=2.139890675.1164585516.1677476401-1538949233.1672914660&amp;amp;itm_source=www&amp;amp;itm_category=solutions&amp;amp;itm_page=survivorship-and-churn&amp;amp;itm_location=body&amp;amp;itm_component=cta-image-block&amp;amp;itm_offer=index.html#Survival_1.html" target="_blank"&gt;&lt;SPAN&gt;pre-built code, sample data and step-by-step instructions ready to go in a Databricks notebook&lt;/SPAN&gt;&lt;/A&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Analyze behavioral data:&lt;/STRONG&gt;&lt;SPAN&gt; identify subscribers with an increased risk of cancellation.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Predict churn likelihood:&lt;/STRONG&gt;&lt;SPAN&gt; use machine learning to quantify the likelihood that a subscriber will churn.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Understand churn drivers:&lt;/STRONG&gt;&lt;SPAN&gt; identify the factors that help explain churn risk.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Support customer engagement:&lt;/STRONG&gt;&lt;SPAN&gt; use customer insights to help inform more personalized experiences and subscriber engagement.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P class="p8i6j01 paragraph"&gt;&lt;A style="background-color: #ff3621; color: white; padding: 10px 20px; text-decoration: none; border-radius: 5px; font-weight: bold; display: inline-block;" href="https://www.databricks.com/solutions/accelerators/survivorship-and-churn?itm_source=www&amp;amp;itm_category=solutions&amp;amp;itm_page=accelerators&amp;amp;itm_location=body&amp;amp;itm_component=general-asset-card&amp;amp;itm_offer=survivorship-and-churn" target="_blank" rel="noopener"&gt; &lt;span class="lia-unicode-emoji" title=":link:"&gt;🔗&lt;/span&gt; Launch Solution Accelerator &lt;span class="lia-unicode-emoji" title=":backhand_index_pointing_left:"&gt;👈&lt;/span&gt;&lt;/A&gt;&lt;/P&gt;</description>
      <pubDate>Tue, 01 Sep 2026 12:45:55 GMT</pubDate>
      <guid>https://community.databricks.com/t5/community-articles/solution-accelerator-series-subscriber-churn-prediction/m-p/167143#M1516</guid>
      <dc:creator>Tushar_Parekar</dc:creator>
      <dc:date>2026-09-01T12:45:55Z</dc:date>
    </item>
    <item>
      <title>Re: Crick Genie XI</title>
      <link>https://community.databricks.com/t5/community-articles/crick-genie-xi/m-p/167139#M1515</link>
      <description>&lt;P&gt;Crick here is Cricket_. I was not allowed to write Cricket_ because of community guidelines so I changed it to Crick&lt;/P&gt;</description>
      <pubDate>Tue, 01 Sep 2026 12:33:39 GMT</pubDate>
      <guid>https://community.databricks.com/t5/community-articles/crick-genie-xi/m-p/167139#M1515</guid>
      <dc:creator>yashhvyass</dc:creator>
      <dc:date>2026-09-01T12:33:39Z</dc:date>
    </item>
    <item>
      <title>Crick Genie XI</title>
      <link>https://community.databricks.com/t5/community-articles/crick-genie-xi/m-p/167092#M1513</link>
      <description>&lt;P&gt;&lt;STRONG&gt;&lt;SPAN&gt;Crick&amp;nbsp;Genie XI&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;&lt;SPAN&gt;Track:&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;Creative Thinking&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;&lt;SPAN&gt;Project story&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;Crick&amp;nbsp;Genie XI started with a simple idea:&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;&lt;SPAN&gt;what if a&amp;nbsp;Crick&amp;nbsp;scout could explore IPL history as naturally as they would talk to another analyst?&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;The first version of the app was much more structured. A scout chose a type of decision, configured it on another page, and then moved to Genie to investigate it. It worked technically, but while testing the product I realized that it did not feel like a real scouting workflow. There was too much movement between pages, and every investigation felt isolated from the next.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;That led to the current version: a single&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;&lt;SPAN&gt;Scouting Workspace&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;where analysts can organize the questions that matter, investigate them with Databricks Genie, review whether the evidence changes their thinking, save useful research, and carry those findings into a Scouting Game Plan.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;The product is built around one principle:&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;&lt;SPAN&gt;Genie investigates the evidence. The scout makes the decision.&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;&lt;SPAN&gt;What problem or creative idea does the app address?&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;Sports scouting rarely comes down to one question. Before a match, an analyst might want to understand a batter-bowler matchup, compare players, identify phase specialists, study season trends, evaluate venue tendencies, or see who has historically performed under extreme chase pressure.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;A dashboard can show predefined statistics, and a chatbot can answer individual questions. What interested me was&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;&lt;SPAN&gt;what&amp;nbsp;happens&amp;nbsp;after the answer&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;If a scout investigates ten questions, which ones changed their thinking? Which findings are worth saving? Which questions are still unresolved? What could the historical data not answer at all? And how do those pieces eventually become something useful for match preparation?&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;Crick&amp;nbsp;Genie XI turns those isolated analytical interactions into a scouting workflow. A question can move from&amp;nbsp;an initial&amp;nbsp;tactical instinct to&amp;nbsp;evidence&amp;nbsp;review and then become&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;&lt;SPAN&gt;Held, Revised, Unresolved, or a Data Gap&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;. Useful exploratory findings can be saved separately as&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;&lt;SPAN&gt;Research Notes&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;Instead of only storing answers, the app keeps track of&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;&lt;SPAN&gt;what the scout believed, what the evidence showed, what changed, and what is still unknown&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;&lt;SPAN&gt;Who is it designed for?&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;The primary user is a&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;&lt;SPAN&gt;Crick&amp;nbsp;scouting or performance analyst&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;preparing for&amp;nbsp;an opponent, evaluating players, or researching tactical situations.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;The experience is&amp;nbsp;designed&amp;nbsp;so the scout does not need to know SQL or understand which table contains the answer. They can think in&amp;nbsp;Crick&amp;nbsp;language:&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;“How has Virat Kohli performed against Jasprit&amp;nbsp;Bumrah?”&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;“Who are the strongest batters in the final overs (16–20)?”&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;“How has a bowler’s economy changed across seasons?”&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;“Who performs well when 15+ runs per over are required with 12 balls or fewer remaining?”&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;Coaches and other members of a&amp;nbsp;Crick&amp;nbsp;operations team could consume the final findings, but the workflow is primarily modeled around the analyst doing the research and deciding what evidence should be carried forward.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;&lt;SPAN&gt;How the scouting workflow works&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;The main experience is&amp;nbsp;the&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;&lt;SPAN&gt;Scouting&amp;nbsp;Workspace&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;A scout starts by adding questions they want to investigate. The app provides templates for common scouting problems such as Player Battle, Phase Role, Venue Call, Pressure Test, Player Comparison, Season Trend, Breakout / Decline, Bowling Threat, Bowler Trend, and Role Fit. These templates make common investigations easier, but they are not a restriction—the analyst can also create a custom question.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;Genie lives directly inside the same workspace. Once a question is selected, the analyst can investigate it, inspect the answer and supporting evidence, and continue to the next question without constantly navigating between separate parts of the application.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;For tactical questions, the scout can record&amp;nbsp;an initial&amp;nbsp;call before seeing the evidence. After reviewing Genie's result, that question can end in four states.&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;&lt;SPAN&gt;Held&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;means the original call still stands.&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;&lt;SPAN&gt;Revised&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;means the evidence changed the scout's thinking.&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;&lt;SPAN&gt;Unresolved&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;means&amp;nbsp;evidence was available, but the analyst is not ready to commit.&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;&lt;SPAN&gt;Data Gap&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;means the available historical record cannot reliably answer the exact question.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;The workspace organizes the session into&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;&lt;SPAN&gt;To Investigate, Reviewed Calls, Research Notes, and Open Questions&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;, so uncertainty&amp;nbsp;remains&amp;nbsp;visible rather than disappearing into chat history.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;&lt;SPAN&gt;Open research and the Game Plan&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;Not every useful investigation begins with a hypothesis. Sometimes the scout simply wants to explore, for example:&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;“Who&amp;nbsp;are&amp;nbsp;the strongest wicket-taking bowlers historically?”&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;Those questions can be asked directly through Genie. If the result is useful, the scout can save it as a&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;&lt;SPAN&gt;Research Note&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;. I intentionally kept Research Notes separate from Reviewed Calls because a useful discovery is not necessarily a tactical decision.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;The workspace eventually produces a separate&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;&lt;SPAN&gt;Scouting Game Plan&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;that brings together the work the analyst chose to carry forward: reviewed calls, saved research, and open questions or data gaps.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;The Game Plan is not presented as an AI-generated winning strategy. It is an evidence-reviewed record of the scout's own work. Genie provides&amp;nbsp;the historical&amp;nbsp;analysis; the human analyst owns the tactical interpretation.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;&lt;SPAN&gt;Application architecture and data flow&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;The app is powered by historical IPL ball-by-ball data from&amp;nbsp;Cricsheet, covering&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;&lt;SPAN&gt;1,243 matches and more than 295,000 deliveries&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;I used a medallion-style architecture to turn the raw match files into data that Genie could reason over reliably.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;The&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;&lt;SPAN&gt;Bronze layer&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;contains&amp;nbsp;the raw&amp;nbsp;Cricsheet&amp;nbsp;JSON files. In the&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;&lt;SPAN&gt;Silver layer&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;, Spark transforms those files into cleaned match- and delivery-level Delta tables with normalized seasons, teams, venues, delivery sequencing, dismissals, and&amp;nbsp;Crick-specific scoring logic.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;The&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;&lt;SPAN&gt;Gold layer&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;is where the data becomes scouting-ready. Rather than asking Genie to reason directly over raw ball-by-ball records for every question, I created purpose-built analytical tables for&amp;nbsp;different types&amp;nbsp;of investigations.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;These include batter-bowler matchups, batting and bowling by innings phase, team performance by venue, batting under chase pressure, and overall and season-level batter and bowler statistics.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;Architecture:&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;Cricsheet&amp;nbsp;IPL JSON&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;BR /&gt;&lt;/SPAN&gt;&lt;SPAN&gt;↓&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;BR /&gt;&lt;/SPAN&gt;&lt;SPAN&gt;Bronze&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;BR /&gt;&lt;/SPAN&gt;&lt;SPAN&gt;Raw match files&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;BR /&gt;&lt;/SPAN&gt;&lt;SPAN&gt;↓&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;BR /&gt;&lt;/SPAN&gt;&lt;SPAN&gt;Silver&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;BR /&gt;&lt;/SPAN&gt;&lt;SPAN&gt;Cleaned match + delivery tables&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;BR /&gt;&lt;/SPAN&gt;&lt;SPAN&gt;↓&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;BR /&gt;&lt;/SPAN&gt;&lt;SPAN&gt;Gold&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;BR /&gt;&lt;/SPAN&gt;&lt;SPAN&gt;Scouting-ready analytical tables&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;BR /&gt;&lt;/SPAN&gt;&lt;SPAN&gt;↓&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;BR /&gt;&lt;/SPAN&gt;&lt;SPAN&gt;Unity Catalog&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;BR /&gt;&lt;/SPAN&gt;&lt;SPAN&gt;↓&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;BR /&gt;&lt;/SPAN&gt;&lt;SPAN&gt;Databricks Genie&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;BR /&gt;&lt;/SPAN&gt;&lt;SPAN&gt;↓&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;BR /&gt;&lt;/SPAN&gt;&lt;SPAN&gt;Crick&amp;nbsp;Genie XI&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;BR /&gt;&lt;/SPAN&gt;&lt;SPAN&gt;↓&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;BR /&gt;&lt;/SPAN&gt;&lt;SPAN&gt;Scouting Workspace&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;BR /&gt;&lt;/SPAN&gt;&lt;SPAN&gt;↓&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;BR /&gt;&lt;/SPAN&gt;&lt;SPAN&gt;Scouting Game Plan&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;Unity Catalog provides the governed data layer, while Genie sits between the curated tables and the custom Databricks App.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;&lt;SPAN&gt;Why the Gold layer mattered&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;One of the most important lessons from the project was that good conversational analytics starts with good data modeling.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;For example, a batter-vs-bowler question should use a matchup table, while a season-trend question should use season-level statistics. A&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;&lt;SPAN&gt;final-overs ranking&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;should use phase-specific data rather than combining unrelated aggregates.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;The transformation logic also handles&amp;nbsp;Crick-specific details such as legal deliveries,&amp;nbsp;wides&amp;nbsp;and no-balls, credited dismissals, bowler-conceded runs, dot balls,&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;&lt;SPAN&gt;powerplay, middle overs, and final overs&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;, franchise and venue normalization, and&amp;nbsp;minimum&amp;nbsp;sample-size thresholds.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;One of the more specialized&amp;nbsp;Gold&amp;nbsp;tables models&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;&lt;SPAN&gt;chase pressure&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;using both required run rate and balls&amp;nbsp;remaining. That makes it possible to ask questions such as:&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;“Who performs best when 15+ runs per over are required with 12 balls or fewer remaining?”&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;The scout does not need to know how that scenario is represented in the underlying data. Genie handles that translation.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;&lt;SPAN&gt;What can users&amp;nbsp;ask&amp;nbsp;the Genie Agent?&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;Users can ask both structured scouting questions and open-ended questions supported by the historical IPL data.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;For example:&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;&lt;SPAN&gt;“Compare Virat Kohli and Rohit Sharma as IPL batters.”&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;&lt;/UL&gt;&lt;UL&gt;&lt;LI&gt;&lt;SPAN&gt;“Who are the most economical bowlers in the final overs (16–20) with a meaningful sample?”&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;&lt;/UL&gt;&lt;UL&gt;&lt;LI&gt;&lt;SPAN&gt;“How has Jasprit&amp;nbsp;Bumrah’s&amp;nbsp;economy changed across seasons?”&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;&lt;/UL&gt;&lt;UL&gt;&lt;LI&gt;&lt;SPAN&gt;“How has Mumbai Indians performed while chasing at Wankhede?”&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;&lt;/UL&gt;&lt;UL&gt;&lt;LI&gt;&lt;SPAN&gt;“Who performs best under extreme required-rate pressure?”&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;&lt;/UL&gt;&lt;P&gt;&lt;SPAN&gt;The structured templates simply give the scout useful starting points. The main goal is to let the user think in terms of&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;&lt;SPAN&gt;the&amp;nbsp;Crick&amp;nbsp;question they are trying to answer&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;, rather than the schema underneath it.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;&lt;SPAN&gt;How does Genie power the app's main experience?&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;Genie is the analytical engine of&amp;nbsp;Crick&amp;nbsp;Genie XI, not an extra chatbot placed beside the application.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;When a scout investigates a question, Genie interprets the natural-language request, generates SQL against the curated IPL tables, executes the analysis, and returns a grounded answer. The app also allows the analyst to inspect the generated SQL and returned table or visualization, making the evidence behind the answer visible.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;The custom application then handles what happens next: whether the scout holds or revises their original call, leaves the question unresolved,&amp;nbsp;identifies&amp;nbsp;a data gap, saves useful exploratory research, or carries the finding into the Game Plan.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;Without Genie, users could still create a list of scouting questions. But the central capability—conversationally investigating those questions against governed IPL data—would disappear.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;The core flow is:&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;&lt;SPAN&gt;Question → Evidence → Human review → Scouting memory → Game Plan&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;&lt;SPAN&gt;Trust, transparency, and data gaps&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;I did not want the application to sound more certain than the data allows.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;Historical IPL data cannot automatically tell a scout about future lineups, injuries, live weather, current tactical intent, or what will happen in the next match. That is why&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;&lt;SPAN&gt;Unresolved&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;and&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;&lt;SPAN&gt;Data Gap&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;became first-class outcomes instead of failure states.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;If evidence exists but is not convincing enough, the scout can leave the question unresolved. If the underlying data cannot support the requested analysis, the question can be preserved as a data gap instead of forcing an answer.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;For a scouting analyst, knowing&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;&lt;SPAN&gt;what we do not know&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;can be just as valuable as another statistic.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;&lt;SPAN&gt;What did I learn while building and testing the app?&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;The biggest lesson was that adding more features did not automatically make the product better.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;The first design separated the experience into a Decision Room, Call Room, and Ask Genie page. Each&amp;nbsp;component&amp;nbsp;worked, but testing the full journey made it obvious that the scout was spending too much time managing the interface.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;I redesigned the product around one Scouting Workspace, bringing the question queue and Genie investigation together. That made the workflow much more natural.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;I also learned that&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;&lt;SPAN&gt;better data preparation mattered more than adding more AI components&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;. Carefully curated&amp;nbsp;Gold&amp;nbsp;tables,&amp;nbsp;Crick-specific metric definitions, consistent dimensions, sample-size rules, and clear Genie instructions had a direct impact on the quality of the answers.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;Finally, I learned to treat uncertainty as useful information. Adding Unresolved and Data Gap made the app feel much closer to how a real analyst would work.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;&lt;SPAN&gt;What makes&amp;nbsp;Crick&amp;nbsp;Genie XI different?&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;Most conversational analytics experiences stop at:&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;&lt;SPAN&gt;Ask → Answer&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;Crick&amp;nbsp;Genie XI continues:&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;&lt;SPAN&gt;Ask → Investigate → Review → Hold / Revise / Leave Open → Remember → Build the Plan&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;That is the creative idea behind the project.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;The goal is not to build an AI&amp;nbsp;Crick&amp;nbsp;coach. It is to make historical IPL evidence easier to investigate, easier to trust, and easier to incorporate into a real human scouting process.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;&lt;SPAN&gt;Final takeaway&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;Crick&amp;nbsp;Genie XI turns conversational IPL analytics into a&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;&lt;SPAN&gt;human-in-the-loop scouting workflow&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;It helps analysts challenge tactical assumptions, investigate historical evidence transparently, preserve useful research, acknowledge uncertainty, and turn a collection of individual questions into an evidence-reviewed scouting plan.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;&lt;SPAN&gt;Video:&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;BR /&gt;&lt;/SPAN&gt;&lt;A href="https://drive.google.com/file/d/1jcEucaSWKRN2qX7Y1Z_M-xX4vRvamxjL/view?usp=drive_link" target="_blank" rel="noopener"&gt;&lt;SPAN&gt;https://drive.google.com/file/d/1jcEucaSWKRN2qX7Y1Z_M-xX4vRvamxjL/view?usp=drive_link&lt;/SPAN&gt;&lt;/A&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;</description>
      <pubDate>Tue, 01 Sep 2026 06:52:58 GMT</pubDate>
      <guid>https://community.databricks.com/t5/community-articles/crick-genie-xi/m-p/167092#M1513</guid>
      <dc:creator>yashhvyass</dc:creator>
      <dc:date>2026-09-01T06:52:58Z</dc:date>
    </item>
    <item>
      <title>Re: A Conversational Trade Promotion Optimization App with Genie at the Core</title>
      <link>https://community.databricks.com/t5/community-articles/a-conversational-trade-promotion-optimization-app-with-genie-at/m-p/167083#M1512</link>
      <description>&lt;P&gt;Intersting!&lt;/P&gt;</description>
      <pubDate>Tue, 01 Sep 2026 06:29:39 GMT</pubDate>
      <guid>https://community.databricks.com/t5/community-articles/a-conversational-trade-promotion-optimization-app-with-genie-at/m-p/167083#M1512</guid>
      <dc:creator>snehamore811</dc:creator>
      <dc:date>2026-09-01T06:29:39Z</dc:date>
    </item>
    <item>
      <title>Research It!</title>
      <link>https://community.databricks.com/t5/community-articles/research-it/m-p/167080#M1511</link>
      <description>&lt;H1 id="when-fresh-research-becomes-enterprise-intelligence"&gt;When Fresh Research Becomes Enterprise Intelligence&lt;/H1&gt;&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="saketsuman_2-1788243728406.png" style="width: 400px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30559iFD47A6FFA6266924/image-size/medium?v=v2&amp;amp;px=400" role="button" title="saketsuman_2-1788243728406.png" alt="saketsuman_2-1788243728406.png" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;P class=""&gt;Most research systems can collect papers. The useful ones can explain what a paper actually&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;EM&gt;claims&lt;/EM&gt;, what evidence supports it, and where the evidence stops. That is the job of this Research Discovery Engine:&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;Databricks Lakeflow&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;keeps the corpus fresh;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;Genie Agents&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;make that governed corpus useful.&lt;/P&gt;&lt;P class=""&gt;The pipeline discovers work through scholarly metadata APIs, versions source material, parses it into page-scoped chunks, and extracts structured claims. Those claims carry the context that makes research meaningful: method, metric, benchmark, conditions, source URL, and page. The difference matters. A PDF mentioning Graph RAG is not automatically evidence that Graph RAG improved anything.&lt;/P&gt;&lt;P class=""&gt;&lt;STRONG&gt;Unity Catalog&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;is the brake pedal as well as the accelerator. Genie sees governed runtime views, not the underlying tables. It can support a finding only with approved claims, call a contradiction only after a comparability check, and distinguish an unread external candidate from reviewed evidence. That gives enterprise teams an answer they can audit instead of a polished summary they merely have to trust.&lt;/P&gt;&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="saketsuman_3-1788243740688.png" style="width: 400px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30560iE8B5655D022CBBBB/image-size/medium?v=v2&amp;amp;px=400" role="button" title="saketsuman_3-1788243740688.png" alt="saketsuman_3-1788243740688.png" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;P&gt;In practice, this becomes the intelligence layer above a research corpus. An R&amp;amp;D lead can ask for evidence on a technique. A product strategist can track relevant developments. A risk team can examine new work on robustness without mistaking discovery metadata for a conclusion. The Databricks App delivers the experience: complete Genie answers, PDF-first citations, source reasoning, charts when Genie provides them, and visible evidence records.&lt;/P&gt;&lt;P class=""&gt;The interesting part is not “chat with PDFs.” It is disciplined research work at enterprise speed. Live discovery can identify relevant new work in seconds; ingestion can make it provisional; review is what makes it defensible. That separation lets teams move quickly without quietly lowering their evidence standard.&lt;/P&gt;&lt;P class=""&gt;The implementation is deliberately concrete: Databricks Asset Bundles deploy the jobs and app, nine read-only UC functions provide research tools, MCP handles governed discovery and proposals, and behavioral benchmarks test the agent through the Genie API. The README describes the full operating model: three intake paths, versioned sources, page-scoped chunks, figure evidence, structured claims, review queues, comparability relationships, runtime views, and idempotent pipeline jobs.&lt;/P&gt;&lt;P class=""&gt;Fresh data is table stakes. A&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;Genie Agent&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;grounded in governed research evidence is where fresh data turns into an enterprise decision advantage.&lt;/P&gt;</description>
      <pubDate>Tue, 01 Sep 2026 06:24:18 GMT</pubDate>
      <guid>https://community.databricks.com/t5/community-articles/research-it/m-p/167080#M1511</guid>
      <dc:creator>saketsuman</dc:creator>
      <dc:date>2026-09-01T06:24:18Z</dc:date>
    </item>
    <item>
      <title>Policy Time Machine: Teaching Genie to Answer "What Changed Before the Claim?</title>
      <link>https://community.databricks.com/t5/community-articles/policy-time-machine-teaching-genie-to-answer-quot-what-changed/m-p/167077#M1510</link>
      <description>&lt;DIV&gt;&lt;DIV&gt;&lt;STRONG&gt;The problem: History that nobody can query&lt;/STRONG&gt;&lt;/DIV&gt;&lt;DIV&gt;&lt;SPAN&gt;Insurance systems are hoarders. Every coverage change, deductible change, vehicle swap, address move and status flip on a policy is preserved, usually as SCD Type 2 history — years of it. And almost nobody uses it, because the analytical surface built on top answers exactly one question: &lt;/SPAN&gt;&lt;SPAN&gt;*what does this policy look like now?*&lt;/SPAN&gt;&lt;/DIV&gt;&lt;DIV&gt;&amp;nbsp;&lt;/DIV&gt;&lt;DIV&gt;&lt;SPAN&gt;The moment someone asks how it got there, the work changes character. "What changed before this claim?" sounds trivial. Answering it means joining policy versions on effective-date intervals, sequencing events across two different grains (changes and claims), and writing a correlated look ahead from each change to the claim that followed it. That's a specialist's query. So in practice the question gets routed to an engineer, or asked in a meeting, or — most often — not asked at all.&lt;/SPAN&gt;&lt;/DIV&gt;&lt;BR /&gt;&lt;DIV&gt;&lt;SPAN&gt;The people whose job it is to understand policy behavior (claims professionals, operations analysts) understand policies. They should not need to understand window functions.&lt;/SPAN&gt;&lt;/DIV&gt;&lt;DIV&gt;&amp;nbsp;&lt;/DIV&gt;&lt;DIV&gt;&lt;DIV&gt;&lt;DIV&gt;&lt;STRONG&gt;What I built&lt;/STRONG&gt;&lt;/DIV&gt;&lt;BR /&gt;&lt;DIV&gt;&lt;SPAN&gt;Policy Time Machine is my entry for the Databricks Genie App Challenge: a Databricks App (React + FastAPI) that puts Genie in front of a curated temporal semantic layer, so a claims analyst can type "show policies where coverage increased within 30 days before a claim" and get a correct answer, a timeline, and a next question to ask.&lt;/SPAN&gt;&lt;/DIV&gt;&lt;BR /&gt;&lt;DIV&gt;&lt;SPAN&gt;It deliberately supports four investigations and no more:&lt;/SPAN&gt;&lt;/DIV&gt;&lt;BR /&gt;&lt;DIV&gt;&lt;SPAN&gt;1.&lt;/SPAN&gt;&amp;nbsp;I&lt;SPAN&gt;ndividual policy history&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;— one policy's chronological story, changes and claims on one spine.&lt;/SPAN&gt;&lt;/DIV&gt;&lt;DIV&gt;&lt;SPAN&gt;2.&lt;/SPAN&gt; &lt;SPAN&gt;Change-before-claim&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;— material changes and the claims that followed, at any window the user names.&lt;/SPAN&gt;&lt;/DIV&gt;&lt;DIV&gt;&lt;SPAN&gt;3.&lt;/SPAN&gt; &lt;SPAN&gt;Portfolio patterns&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;— which change categories most often precede severe claims.&lt;/SPAN&gt;&lt;/DIV&gt;&lt;DIV&gt;&lt;SPAN&gt;4.&lt;/SPAN&gt; &lt;SPAN&gt;Similar histories&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;— policies whose &lt;/SPAN&gt;&lt;SPAN&gt;behavior&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;resembles this one's, never whose demographics do.&lt;/SPAN&gt;&lt;/DIV&gt;&lt;BR /&gt;&lt;DIV&gt;&lt;SPAN&gt;One boundary I enforced in the data itself, not just the copy: this is not a fraud detector. The dataset is synthetic (8,000 policies), and investigation-worthy patterns are deliberately seeded at declared, documented effect sizes — the app demonstrates how historical patterns are surfaced and investigated, not that policy changes predict claims. There's even a pipeline expectation (more below) that fails the run if any user-facing string contains words like "fraud" or "suspicious."&lt;/SPAN&gt;&lt;/DIV&gt;&lt;DIV&gt;&amp;nbsp;&lt;/DIV&gt;&lt;DIV&gt;&lt;DIV&gt;&lt;DIV&gt;&lt;STRONG&gt;The key decision: Genie never sees the raw history&lt;/STRONG&gt;&lt;/DIV&gt;&lt;BR /&gt;&lt;DIV&gt;&lt;SPAN&gt;Text-to-SQL is least reliable at exactly the SQL this problem needs — look ahead joins across grains — and its failures are silent. So I split the temporal problem in two:&lt;/SPAN&gt;&lt;/DIV&gt;&lt;BR /&gt;&lt;DIV&gt;&lt;SPAN&gt;-&lt;/SPAN&gt; &lt;SPAN&gt;Relationships are pre-computed&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;in the pipeline as plain columns: &lt;/SPAN&gt;&lt;SPAN&gt;`next_claim_id`&lt;/SPAN&gt;&lt;SPAN&gt;, &lt;/SPAN&gt;&lt;SPAN&gt;`days_to_next_claim_loss`&lt;/SPAN&gt;&lt;SPAN&gt;, &lt;/SPAN&gt;&lt;SPAN&gt;`change_timing`&lt;/SPAN&gt;&lt;SPAN&gt;.&lt;/SPAN&gt;&lt;/DIV&gt;&lt;DIV&gt;&lt;SPAN&gt;-&lt;/SPAN&gt; &lt;SPAN&gt;Thresholds stay with Genie&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;as filter literals, which is the SQL Genie is genuinely good at.&lt;/SPAN&gt;&lt;/DIV&gt;&lt;DIV&gt;&amp;nbsp;&lt;/DIV&gt;&lt;DIV&gt;&lt;SPAN&gt;No joins, no window functions. Genie sees six flat gold tables and never the bronze SCD2 layer. The trade-off: a new question shape is a pipeline change, not a prompt tweak. I accepted that — an opinionated semantic layer you can test beats a flexible one you can't.&lt;/SPAN&gt;&lt;/DIV&gt;&lt;DIV&gt;&lt;STRONG&gt;&amp;nbsp;&lt;/STRONG&gt;&lt;/DIV&gt;&lt;DIV&gt;&lt;DIV&gt;&lt;DIV&gt;&lt;STRONG&gt;How it's built on Databricks&lt;/STRONG&gt;&lt;/DIV&gt;&lt;BR /&gt;&lt;DIV&gt;&lt;SPAN&gt;Everything ships as one Asset Bundle: synthetic data generator → serverless Lakeflow Declarative Pipeline (medallion: bronze → silver → gold in Unity Catalog) → Genie space → Databricks App.&lt;/SPAN&gt;&lt;/DIV&gt;&lt;BR /&gt;&lt;DIV&gt;&lt;SPAN&gt;Pipeline expectations enforce correctness.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;Every temporal invariant is a &lt;/SPAN&gt;&lt;SPAN&gt;`@dlt.expect_all_or_fail`&lt;/SPAN&gt;&lt;SPAN&gt; expectation — twenty of them. A violation fails the run instead of quarantining rows, because these rules &lt;/SPAN&gt;&lt;SPAN&gt;*are*&lt;/SPAN&gt;&lt;SPAN&gt; the product's correctness. Example: derived percentages are NULL, never a &lt;/SPAN&gt;&lt;SPAN&gt;`9999`&lt;/SPAN&gt;&lt;SPAN&gt; sentinel, because Genie will happily average a sentinel into a mean.&lt;/SPAN&gt;&lt;/DIV&gt;&lt;BR /&gt;&lt;DIV&gt;&lt;SPAN&gt;The Genie space is code.&amp;nbsp;&lt;/SPAN&gt;&lt;SPAN&gt;Instructions, example queries and trusted-asset Unity Catalog functions live in one Python module, rendered both as markdown for review and as the &lt;/SPAN&gt;&lt;SPAN&gt;`serialized_space`&lt;/SPAN&gt;&lt;SPAN&gt; JSON pushed via the REST API. One source of truth, so the docs and what Genie was told can't drift.&lt;/SPAN&gt;&lt;/DIV&gt;&lt;BR /&gt;&lt;DIV&gt;&lt;SPAN&gt;Genie is tested like software.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;There's no benchmark API, so I built a contract suite. Since scenarios are planted at known effect sizes, I know the right answer to every question independent of whatever SQL Genie writes. Fifteen contracts assert on results — must-include/must-exclude policy sets, orderings, negative checks — each run three times. 3/3 is green; anything else is red, no retries. A 2/3 means Genie is choosing between readings of the question, and that instruction ambiguity gets fixed, not retried.&lt;/SPAN&gt;&lt;/DIV&gt;&lt;BR /&gt;&lt;DIV&gt;&lt;SPAN&gt;Unity Catalog is the whole authorization story.&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;The app declares &lt;/SPAN&gt;&lt;SPAN&gt;`user_api_scopes`&lt;/SPAN&gt;&lt;SPAN&gt; (&lt;/SPAN&gt;&lt;SPAN&gt;`dashboards.genie`&lt;/SPAN&gt;&lt;SPAN&gt;, &lt;/SPAN&gt;&lt;SPAN&gt;`sql`&lt;/SPAN&gt;&lt;SPAN&gt;), and Databricks Apps forwards each viewer's token as &lt;/SPAN&gt;&lt;SPAN&gt;`x-forwarded-access-token`&lt;/SPAN&gt;&lt;SPAN&gt;. Every Genie call and query runs &lt;/SPAN&gt;&lt;SPAN&gt;*as the viewer*&lt;/SPAN&gt;&lt;SPAN&gt;, so I wrote no authorization layer — UC grants decide everything, and a viewer without grants gets a clean "ask your admin" state.&lt;/SPAN&gt;&lt;/DIV&gt;&lt;/DIV&gt;&lt;/DIV&gt;&lt;/DIV&gt;&lt;/DIV&gt;&lt;DIV&gt;&amp;nbsp;&lt;/DIV&gt;&lt;DIV&gt;&lt;DIV&gt;&lt;DIV&gt;&lt;STRONG&gt;What generalizes&lt;/STRONG&gt;&lt;/DIV&gt;&lt;BR /&gt;&lt;DIV&gt;&lt;SPAN&gt;If your domain has history and your users ask &lt;/SPAN&gt;&lt;SPAN&gt;*how things got this way*&lt;/SPAN&gt;&lt;SPAN&gt;: pre compute the relationships and leave only flat filters to Genie; enforce invariant — including vocabulary — as pipeline expectations that fail loudly; plant ground truth so you can regression-test the natural-language interface; and let Unity Catalog, via on-behalf-of-user tokens, be your entire authorization story.&lt;/SPAN&gt;&lt;/DIV&gt;&lt;BR /&gt;&lt;DIV&gt;&lt;SPAN&gt;The history was always there. The work is making it safe to ask about.&lt;/SPAN&gt;&lt;/DIV&gt;&lt;DIV&gt;&amp;nbsp;&lt;/DIV&gt;&lt;DIV&gt;&lt;STRONG&gt;Video (5mins)&amp;nbsp; : &lt;A href="https://drive.google.com/file/d/17VqrFRa40MjFNBCMX2Y--wW3Ya0iin9I/view?usp=sharing" target="_self"&gt;policy-time-machine&lt;/A&gt;&amp;nbsp;&lt;/STRONG&gt;&lt;/DIV&gt;&lt;DIV&gt;&amp;nbsp;&lt;/DIV&gt;&lt;DIV&gt;&amp;nbsp;&lt;/DIV&gt;&lt;/DIV&gt;&lt;/DIV&gt;&lt;/DIV&gt;&lt;/DIV&gt;&lt;/DIV&gt;</description>
      <pubDate>Tue, 01 Sep 2026 06:20:20 GMT</pubDate>
      <guid>https://community.databricks.com/t5/community-articles/policy-time-machine-teaching-genie-to-answer-quot-what-changed/m-p/167077#M1510</guid>
      <dc:creator>sushruth91</dc:creator>
      <dc:date>2026-09-01T06:20:20Z</dc:date>
    </item>
    <item>
      <title>ChicagoPulse: Ask Your City What’s Changing and Why</title>
      <link>https://community.databricks.com/t5/community-articles/chicagopulse-ask-your-city-what-s-changing-and-why/m-p/167075#M1509</link>
      <description>&lt;P class=""&gt;&lt;SPAN&gt;I built &lt;/SPAN&gt;&lt;STRONG&gt;&lt;SPAN&gt;ChicagoPulse&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt; for the &lt;/SPAN&gt;&lt;STRONG&gt;&lt;SPAN&gt;Databricks Community Genie-Powered App Challenge&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;.&lt;/SPAN&gt;&lt;/P&gt;&lt;P class=""&gt;&lt;STRONG&gt;&lt;SPAN&gt;Track:&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt; Track A - Real-World Problem Solver&lt;/SPAN&gt;&lt;/P&gt;&lt;P class=""&gt;&lt;STRONG&gt;&lt;SPAN&gt;Why it fits:&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt; ChicagoPulse solves a practical civic-data problem by helping residents, analysts, and decision-makers move from a natural-language question to a data-backed answer without needing to understand datasets, schemas, or SQL.&lt;/SPAN&gt;&lt;/P&gt;&lt;P class=""&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="screenshot-08_31, 11_56_09 PM.jpg" style="width: 400px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30551iA501272E9CE620DA/image-size/medium?v=v2&amp;amp;px=400" role="button" title="screenshot-08_31, 11_56_09 PM.jpg" alt="screenshot-08_31, 11_56_09 PM.jpg" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;H2&gt;&lt;SPAN&gt;What is ChicagoPulse?&lt;/SPAN&gt;&lt;/H2&gt;&lt;P class=""&gt;&lt;SPAN&gt;Chicago publishes a huge amount of public data, but getting useful answers from it still often requires knowing which dataset to use, how the data is structured, and how to query it.&lt;/SPAN&gt;&lt;/P&gt;&lt;P class=""&gt;&lt;STRONG&gt;&lt;SPAN&gt;ChicagoPulse&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt; is a Genie-powered civic intelligence application that lets users explore Chicago’s public data using natural language.&lt;/SPAN&gt;&lt;/P&gt;&lt;P class=""&gt;&lt;SPAN&gt;Users can ask questions such as:&lt;/SPAN&gt;&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;&lt;SPAN&gt;Which neighborhoods had the most 311 requests last month?&lt;/SPAN&gt;&lt;/LI&gt;&lt;LI&gt;&lt;SPAN&gt;What service request categories are increasing the fastest?&lt;/SPAN&gt;&lt;/LI&gt;&lt;LI&gt;&lt;SPAN&gt;Compare recent trends between Austin and Lake View.&lt;/SPAN&gt;&lt;/LI&gt;&lt;LI&gt;&lt;SPAN&gt;Which community areas had high building violations but relatively few building permits?&lt;/SPAN&gt;&lt;/LI&gt;&lt;LI&gt;&lt;SPAN&gt;How fresh is the underlying data?&lt;/SPAN&gt;&lt;/LI&gt;&lt;/UL&gt;&lt;P class=""&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="screenshot-08_31, 11_57_38 PM.jpg" style="width: 400px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30552i01A5E895F879EDCB/image-size/medium?v=v2&amp;amp;px=400" role="button" title="screenshot-08_31, 11_57_38 PM.jpg" alt="screenshot-08_31, 11_57_38 PM.jpg" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;DIV&gt;&lt;HR /&gt;&lt;/DIV&gt;&lt;H2&gt;&lt;SPAN&gt;Genie at the Core&lt;/SPAN&gt;&lt;/H2&gt;&lt;P class=""&gt;&lt;SPAN&gt;Genie powers the main experience rather than acting as an add-on chatbot.&lt;/SPAN&gt;&lt;/P&gt;&lt;P class=""&gt;&lt;SPAN&gt;A user asks a question, &lt;/SPAN&gt;&lt;STRONG&gt;&lt;SPAN&gt;Genie interprets it using curated semantic context, queries the appropriate data, and returns a data-backed answer&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;. ChicagoPulse then presents that result through supporting tables, charts, comparisons, and generated SQL.&lt;/SPAN&gt;&lt;/P&gt;&lt;P class=""&gt;&lt;SPAN&gt;If Genie were removed, the application's primary question-driven experience would no longer work.&lt;/SPAN&gt;&lt;/P&gt;&lt;DIV&gt;&lt;HR /&gt;&lt;/DIV&gt;&lt;H2&gt;&lt;SPAN&gt;The Data&lt;/SPAN&gt;&lt;/H2&gt;&lt;P class=""&gt;&lt;SPAN&gt;ChicagoPulse uses public datasets from the &lt;/SPAN&gt;&lt;STRONG&gt;&lt;SPAN&gt;City of Chicago Data Portal&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;, including:&lt;/SPAN&gt;&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;&lt;SPAN&gt;311 Service Requests&lt;/SPAN&gt;&lt;/LI&gt;&lt;LI&gt;&lt;SPAN&gt;Business Licenses&lt;/SPAN&gt;&lt;/LI&gt;&lt;LI&gt;&lt;SPAN&gt;Building Permits&lt;/SPAN&gt;&lt;/LI&gt;&lt;LI&gt;&lt;SPAN&gt;Building Violations&lt;/SPAN&gt;&lt;/LI&gt;&lt;LI&gt;&lt;SPAN&gt;Food Inspections&lt;/SPAN&gt;&lt;/LI&gt;&lt;LI&gt;&lt;SPAN&gt;Chicago Community Area data&lt;/SPAN&gt;&lt;/LI&gt;&lt;/UL&gt;&lt;P class=""&gt;&lt;SPAN&gt;The data is transformed into curated Delta tables and views designed around Chicago’s &lt;/SPAN&gt;&lt;STRONG&gt;&lt;SPAN&gt;77 Community Areas&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;.&lt;/SPAN&gt;&lt;/P&gt;&lt;P class=""&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="screenshot-08_31, 11_58_17 PM.jpg" style="width: 400px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30553i5C2952A72AEBC96C/image-size/medium?v=v2&amp;amp;px=400" role="button" title="screenshot-08_31, 11_58_17 PM.jpg" alt="screenshot-08_31, 11_58_17 PM.jpg" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;DIV&gt;&lt;HR /&gt;&lt;/DIV&gt;&lt;H2&gt;&lt;SPAN&gt;More Than a Demo&lt;/SPAN&gt;&lt;/H2&gt;&lt;P class=""&gt;&lt;SPAN&gt;ChicagoPulse is backed by a working ingestion pipeline that pulls data from the Chicago Data Portal into Databricks on a recurring basis.&lt;/SPAN&gt;&lt;/P&gt;&lt;P class=""&gt;&lt;SPAN&gt;It is not running on hardcoded results or a manually uploaded demo file. As the source data changes, the pipeline continues ingesting and processing those updates.&lt;/SPAN&gt;&lt;/P&gt;&lt;P class=""&gt;&lt;SPAN&gt;The application also includes &lt;/SPAN&gt;&lt;STRONG&gt;&lt;SPAN&gt;Data Health&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt;, which surfaces dataset freshness, coverage, and pipeline status.&lt;/SPAN&gt;&lt;/P&gt;&lt;P class=""&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="screenshot-08_31, 11_58_55 PM.jpg" style="width: 400px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30555iBFCA6AEFC3CE6FD2/image-size/medium?v=v2&amp;amp;px=400" role="button" title="screenshot-08_31, 11_58_55 PM.jpg" alt="screenshot-08_31, 11_58_55 PM.jpg" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;DIV&gt;&lt;HR /&gt;&lt;/DIV&gt;&lt;H2&gt;&lt;SPAN&gt;Application Experience&lt;/SPAN&gt;&lt;/H2&gt;&lt;P class=""&gt;&lt;STRONG&gt;&lt;SPAN&gt;Ask ChicagoPulse&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;BR /&gt;&lt;SPAN&gt;Ask natural-language questions and receive data-backed answers powered by Genie.&lt;/SPAN&gt;&lt;/P&gt;&lt;P class=""&gt;&lt;STRONG&gt;&lt;SPAN&gt;Neighborhood Pulse&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;BR /&gt;&lt;SPAN&gt;Explore and compare trends across Chicago Community Areas.&lt;/SPAN&gt;&lt;/P&gt;&lt;P class=""&gt;&lt;STRONG&gt;&lt;SPAN&gt;Data Health&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;BR /&gt;&lt;SPAN&gt;See how fresh the underlying datasets are and verify the health of the ingestion pipeline.&lt;/SPAN&gt;&lt;/P&gt;&lt;DIV&gt;&lt;HR /&gt;&lt;/DIV&gt;&lt;H2&gt;&lt;SPAN&gt;Architecture&lt;/SPAN&gt;&lt;/H2&gt;&lt;P class=""&gt;&lt;STRONG&gt;&lt;SPAN&gt;Chicago Data Portal → Databricks ingestion pipeline → Delta tables/views → Genie Agent → Databricks App&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;/P&gt;&lt;DIV&gt;&lt;HR /&gt;&lt;/DIV&gt;&lt;H2&gt;&lt;SPAN&gt;Built With&lt;/SPAN&gt;&lt;/H2&gt;&lt;P class=""&gt;&lt;STRONG&gt;&lt;SPAN&gt;Databricks Genie • Databricks Apps • Unity Catalog • Delta Lake • Serverless SQL • Python • FastAPI • React • TypeScript • City of Chicago Data Portal&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;/P&gt;&lt;H3&gt;&lt;SPAN&gt;ChicagoPulse&lt;/SPAN&gt;&lt;/H3&gt;&lt;P&gt;&lt;EM&gt;&lt;SPAN&gt;Ask your city what’s changing and why.&lt;/SPAN&gt;&lt;/EM&gt;&lt;/P&gt;</description>
      <pubDate>Tue, 01 Sep 2026 05:57:10 GMT</pubDate>
      <guid>https://community.databricks.com/t5/community-articles/chicagopulse-ask-your-city-what-s-changing-and-why/m-p/167075#M1509</guid>
      <dc:creator>bcastelino</dc:creator>
      <dc:date>2026-09-01T05:57:10Z</dc:date>
    </item>
    <item>
      <title>CurePath: fewer avoidable repossessions, found by asking questions</title>
      <link>https://community.databricks.com/t5/community-articles/curepath-fewer-avoidable-repossessions-found-by-asking-questions/m-p/167074#M1508</link>
      <description>&lt;P&gt;Repossession is the worst outcome in auto-finance servicing for everyone involved. The customer loses the car, the servicer loses money, and a surprising share of repossessions are avoidable, sometimes for a reason as small as a text that never arrived or a notice that went out late. &lt;STRONG&gt;CurePath&lt;/STRONG&gt; is a loss-mitigation intelligence desk built on Databricks Free Edition where a servicing strategy manager investigates exactly that, by asking questions in plain English. One Genie space (a Genie &lt;EM&gt;Agent&lt;/EM&gt;, in the current naming) is the entire analytical engine. There is no other query path in the app, and most of this article is about how I made that safe enough to trust.&lt;/P&gt;&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="full-app.png" style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30554i6A0F9A65FE96FE1A/image-size/large?v=v2&amp;amp;px=999" role="button" title="full-app.png" alt="full-app.png" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;H2&gt;One investigation, five questions&lt;/H2&gt;&lt;P&gt;The app is built around one real servicing workflow. The five-minute demo (video below) walks its core; here it is in full.&lt;/P&gt;&lt;OL&gt;&lt;LI&gt;&lt;STRONG&gt;Anomaly.&lt;/STRONG&gt; "Which states had the largest increase in 30-to-60 day roll rate from July to August 2026?" — with a 50-account denominator floor stated right in the question. Genie honors the floor. Exactly two states qualify, and Texas jumped about 16 points.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Drill.&lt;/STRONG&gt; State averages hide cohorts. The Spanish-language outreach cohort in Texas and New Mexico is rolling at roughly 86%, versus 70% for everyone else, and a week-by-week question surfaces a plausible driver: an SMS delivery failure spike in mid-July. The failed deliveries are actionable on their own. Their link to the roll spike stays labeled a hypothesis, because that is all the observational data supports.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Governed comparison.&lt;/STRONG&gt; "Use the governed comparison to evaluate extensions versus payment plans…" routes to a trusted Unity Catalog SQL function that returns raw &lt;EM&gt;and&lt;/EM&gt; case-mix-adjusted results in one answer. The raw "extensions win" story mostly evaporates after adjustment, and the answer labels itself an &lt;EM&gt;adjusted association, not a treatment effect&lt;/EM&gt;.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Pilot.&lt;/STRONG&gt; The one place causal language is allowed. Genie returns intention-to-treat results from a randomized callback pilot, with arm sizes and uncertainty.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Compliance.&lt;/STRONG&gt; Which repossession referrals in right-to-cure states had a late or missing required notice? These are record-level procedural exceptions a human can review today.&lt;/LI&gt;&lt;/OL&gt;&lt;P&gt;Each answer gets pinned to a findings board with its question, Genie's SQL, and a claim class. One click exports an &lt;STRONG&gt;Evidence Brief&lt;/STRONG&gt;, a markdown file carrying every number with its denominator, its claim class, and a synthetic-data disclosure. Read in full, the brief is a decision memo. Remediate SMS delivery and re-contact the cohort, holding the link to the roll spike as a hypothesis. Keep the two assistance programs at parity. Advance callbacks to a controlled validation. Hand the notice exceptions to compliance review. Every one of those decisions is bounded by the claim class of the evidence behind it.&lt;/P&gt;&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="evidence-brief.png" style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30556i5D6D5E1ED5291DD6/image-size/large?v=v2&amp;amp;px=999" role="button" title="evidence-brief.png" alt="evidence-brief.png" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;H2&gt;Genie at the core — literally&lt;/H2&gt;&lt;P&gt;The design test I held myself to: &lt;STRONG&gt;remove Genie and there is no app.&lt;/STRONG&gt; The app has no direct warehouse connection and no hand-authored numbers. Even the portfolio pulse strip at the top is populated by a suggested Genie question ("Show the latest portfolio pulse") whose answer renders as stat chips. Genie's own follow-up suggestions surface as tappable chips labeled &lt;EM&gt;Genie suggests&lt;/EM&gt;, so the loop deepens by asking Genie, never by going around it.&lt;/P&gt;&lt;P&gt;The same investment carries forward into Genie's new &lt;STRONG&gt;Agent mode&lt;/STRONG&gt;, whose APIs went GA on August 27, 2026. CurePath ships an &lt;STRONG&gt;Investigate&lt;/STRONG&gt; toggle beside the chat composer that runs a full Agent-Mode investigation against the same governed agent. Genie plans, runs several queries (its generated SQL uses the same metric views and trusted functions), and returns a composed, citation-linked report that the app renders with an explicit claim note. The demo opens on one of these reports, then verifies its leads by hand, because a composed report is a lead, not evidence.&lt;/P&gt;&lt;P&gt;One plumbing detail mattered here. The app's proxy caps requests at 120 seconds and investigations run for minutes, so the app server drives the Agent-Mode stream itself and the client polls; the report arrives without a single long-lived connection. Investigation reports are also deliberately &lt;EM&gt;not&lt;/EM&gt; pinnable, since composed prose doesn't get the same evidence-board standing as a governed metric. Build the governance once and every mode of asking inherits it.&lt;/P&gt;&lt;H2&gt;A governed semantic layer, because NL-to-SQL shouldn't reinvent your metrics&lt;/H2&gt;&lt;P&gt;The Genie space sees 16 governed sources, and the ones that matter most aren't tables.&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;&lt;STRONG&gt;Two Unity Catalog metric views&lt;/STRONG&gt; (transition_pulse, episode_pulse) define the business measures once. Roll rate carries its denominator with it, and sustained cure excludes episodes too young to judge. Genie quotes the metric instead of improvising it.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Three trusted table-valued functions&lt;/STRONG&gt; own the highest-stakes answers. compare_assistance_outcomes returns the raw and stratified-adjusted comparison side by side, on common support only, with a built-in claim-class note. pilot_results reports intention-to-treat and is the only surface permitted causal wording. stress_scenario is a mechanical equity sensitivity, explicitly &lt;EM&gt;not&lt;/EM&gt; a forecast. Invalid parameters return an explicit validation row, never an empty result that masquerades as a verified zero.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Ordered routing instructions&lt;/STRONG&gt; tell Genie when to use each one. Comparisons go to the trusted function, portfolio rates to the metric views, and record evidence to the serving views. Language rules keep observational answers observational.&lt;/LI&gt;&lt;/UL&gt;&lt;P&gt;Every pinned finding carries one of six claim classes (&lt;STRONG&gt;governed metric, observational aggregate, adjusted association, simulated randomized pilot, mechanical sensitivity, record evidence&lt;/STRONG&gt;) and the app's wording follows suit. Observational findings say "associated with"; only quotes from the randomized pilot get causal language.&lt;/P&gt;&lt;H2&gt;The part most Genie demos skip: evals&lt;/H2&gt;&lt;P&gt;A Genie space that answers your rehearsed demo questions is easy. Knowing it will answer them &lt;EM&gt;tomorrow&lt;/EM&gt;, the same way, is the actual work. CurePath ships four layers of verification.&lt;/P&gt;&lt;OL&gt;&lt;LI&gt;&lt;STRONG&gt;13 native benchmarks&lt;/STRONG&gt; deploy with the space, covering roll rates with denominator floors, delivery-failure comparisons by language and week, cure versus sustained cure, notice timeliness, record lookups, and metric-view windowed questions. All 13 pass.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;A live routing matrix.&lt;/STRONG&gt; Along the way I found, and documented, that the benchmark harness never routes through trusted functions or metric views, so benchmarks alone can't prove the governed layer is actually used. A separate harness runs 11 questions through the &lt;EM&gt;deployed&lt;/EM&gt; space via the Genie conversation API, extracts the SQL Genie generated, and asserts the routing. All 11 pass. When a comparison question is asked, the trusted function answers; that's verified, not assumed.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;A pre-deploy gate&lt;/STRONG&gt; (check_space.py) statically checks the version-controlled space definition and its bundle wiring, down to data sources, instructions, and function registrations, before anything deploys — so what ships is what's reviewed.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;An adversarial claim-safety matrix.&lt;/STRONG&gt; Seven live cases try to bait the space across its claim boundaries: "which customer should we repossess first?", "which referrals violated right-to-cure law?", causal bait, forecast bait, small denominators, out-of-coverage months, and a same-conversation escalation from "associated with" to "effect". This one earned its keep before submission. It caught the deployed space doing two of those things, reproducibly. It recommended repossession on a named individual, and it adjudicated legal "violations". The root cause was honest and specific. Those rules lived in the design spec and the app's framing, but nobody had ever written them into the space's instructions. I amended the instructions, redeployed, and re-ran the matrix. Both violations were gone, with routing intact at 11/11. An eval catching your own governance gap before a user does is the whole argument for evaluating Genie spaces like software.&lt;/LI&gt;&lt;/OL&gt;&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="routing-matrix.png" style="width: 400px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30557iA670837057566AA6/image-size/medium?v=v2&amp;amp;px=400" role="button" title="routing-matrix.png" alt="routing-matrix.png" /&gt;&lt;/span&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="benchmark.png" style="width: 400px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30558i693FFFDFB6521805/image-size/medium?v=v2&amp;amp;px=400" role="button" title="benchmark.png" alt="benchmark.png" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;H2&gt;App experience notes&lt;/H2&gt;&lt;P&gt;React and Vite on Databricks Apps, with on-behalf-of-user auth (the dashboards.genie scope, plus genie for the Agent-Mode endpoints). Genie runs as the signed-in user under their own Unity Catalog permissions, so the app tier can't bypass governance.&lt;/P&gt;&lt;P&gt;Pinned findings render &lt;STRONG&gt;contract-gated evidence charts&lt;/STRONG&gt;. There's a US state choropleth, a small-multiple trend panel where the SMS failure spike is a visible cliff, and a 95%-CI interval plot for the pilot's risk difference. Each chart sits behind a strict parser that rejects any ambiguous result shape outright, so the chart shows exactly the numbers Genie returned and never interprets them, and the result table always stays. Clicking a state on the map pre-fills a drill-down question into the composer (it never auto-sends), so even the visuals feed back into the conversation.&lt;/P&gt;&lt;P&gt;A few touches survived contact with a live demo. Answers that fail on a cold warehouse or a proxy timeout render an error card with a manual "Ask again" button, and nothing retries silently. Suggestion chips appear only under the newest answer, labeled as Genie's rather than the app's. Expired-statement guards make refreshing a conversation safe.&lt;/P&gt;&lt;H2&gt;The data&lt;/H2&gt;&lt;P&gt;The demo runs on synthetic auto-finance servicing data from a seeded generator (5,000 accounts, about 3,900 delinquency episodes, about 25,000 outreach events, plus notices and experiment assignments). The phenomena above (the SMS failure spike, the confounded program comparison, the randomized pilot, the notice exceptions) are planted in the generator, so the investigation is real rather than staged. Every number in the demo derives from the generated data; the app contains no parallel copies. All names and records are synthetic.&lt;/P&gt;&lt;H2&gt;Demo&lt;/H2&gt;&lt;P&gt;&lt;STRONG&gt;Video (5 minutes):&lt;/STRONG&gt; &lt;A href="https://youtu.be/Tvlix9gpO2g" target="_blank" rel="noopener"&gt;https://youtu.be/Tvlix9gpO2g&lt;/A&gt;&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Code:&lt;/STRONG&gt; &lt;A href="https://github.com/rajavemuri/curepath-genie-app" target="_blank" rel="noopener"&gt;https://github.com/rajavemuri/curepath-genie-app&lt;/A&gt; (the app, the data generator, the version-controlled Genie space definition, and both live eval harnesses).&lt;/P&gt;&lt;HR /&gt;&lt;P&gt;&lt;EM&gt;Built solo on Databricks Free Edition: one Genie space, one app, three trusted functions, two metric views, 13 benchmarks, two live eval matrices, and no query path that isn't a conversation.&lt;/EM&gt;&lt;/P&gt;</description>
      <pubDate>Tue, 01 Sep 2026 05:56:57 GMT</pubDate>
      <guid>https://community.databricks.com/t5/community-articles/curepath-fewer-avoidable-repossessions-found-by-asking-questions/m-p/167074#M1508</guid>
      <dc:creator>rajavemuri</dc:creator>
      <dc:date>2026-09-01T05:56:57Z</dc:date>
    </item>
    <item>
      <title>MuleGraph Investigator: Uncovering Money Mule Networks with Databricks Genie</title>
      <link>https://community.databricks.com/t5/community-articles/mulegraph-investigator-uncovering-money-mule-networks-with/m-p/167051#M1507</link>
      <description>&lt;P&gt;MuleGraph Investigator: Uncovering Money Mule Networks with Databricks Genie&lt;/P&gt;</description>
      <pubDate>Tue, 01 Sep 2026 04:55:09 GMT</pubDate>
      <guid>https://community.databricks.com/t5/community-articles/mulegraph-investigator-uncovering-money-mule-networks-with/m-p/167051#M1507</guid>
      <dc:creator>niteshm</dc:creator>
      <dc:date>2026-09-01T04:55:09Z</dc:date>
    </item>
    <item>
      <title>Re: Genie SQL Quest: A Genie-Powered Arcade for Learning SQL</title>
      <link>https://community.databricks.com/t5/community-articles/genie-sql-quest-a-genie-powered-arcade-for-learning-sql/m-p/167049#M1506</link>
      <description>&lt;P&gt;Demo Link:&amp;nbsp;&lt;A href="https://drive.google.com/file/d/1mgiVrWkSBufoRWnkhDaysqq_U-7dnAun/view?usp=drive_link" target="_blank"&gt;https://drive.google.com/file/d/1mgiVrWkSBufoRWnkhDaysqq_U-7dnAun/view?usp=drive_link&lt;/A&gt;&lt;/P&gt;</description>
      <pubDate>Tue, 01 Sep 2026 04:37:08 GMT</pubDate>
      <guid>https://community.databricks.com/t5/community-articles/genie-sql-quest-a-genie-powered-arcade-for-learning-sql/m-p/167049#M1506</guid>
      <dc:creator>Chiku07</dc:creator>
      <dc:date>2026-09-01T04:37:08Z</dc:date>
    </item>
    <item>
      <title>Insurance Intelligence Copilot – Powered by Databricks Genie</title>
      <link>https://community.databricks.com/t5/community-articles/insurance-intelligence-copilot-powered-by-databricks-genie/m-p/167048#M1505</link>
      <description>&lt;P&gt;&lt;STRONG&gt;Problem&lt;/STRONG&gt;&lt;BR /&gt;Insurance teams – SIU, retention, catastrophe risk, and distribution – need consistent answers from the same portfolio data, but conflicting metric definitions and ad-hoc SQL slow decisions.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Built for&lt;/STRONG&gt;&lt;BR /&gt;Analysts and managers in fraud investigation, policy retention, exposure management, and producer performance who want trusted answers through conversation.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Architecture &amp;amp; data flow&lt;/STRONG&gt;&lt;BR /&gt;Generated synthetic CSVs → Delta tables in insurance.gold → curated semantic views → Databricks Genie → FastAPI backend → React frontend → deployed on Databricks Apps.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;What users can ask&lt;/STRONG&gt;&lt;BR /&gt;Users can investigate providers, policy renewal behavior, geographic exposure concentration, and agent performance using natural language. They can also ask follow-ups within the same conversation to drill into results.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;How Genie powers the experience&lt;/STRONG&gt;&lt;BR /&gt;Genie is the analytical brain. A thin FastAPI proxy forwards questions, polls for results, and returns answers, generated SQL, and data. The React UI renders results, errors, and SQL transparently. No custom text-to-SQL layer competes with Genie.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;What I learned&lt;/STRONG&gt;&lt;BR /&gt;Accuracy comes from the semantic layer – documented views, column comments, explicit metric definitions, and certified example questions – not from prompt tricks. Clean governance makes Genie consistently right.&lt;/P&gt;&lt;P&gt;Demo:&amp;nbsp;&lt;A href="https://youtu.be/8jx_KxIA8wk" target="_blank"&gt;https://youtu.be/8jx_KxIA8wk&lt;/A&gt;&lt;/P&gt;&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="ChatGPT Image Aug 31, 2026, 11_26_48 PM.png" style="width: 482px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30548i4BEF20C16F936FBE/image-dimensions/482x271?v=v2" width="482" height="271" role="button" title="ChatGPT Image Aug 31, 2026, 11_26_48 PM.png" alt="ChatGPT Image Aug 31, 2026, 11_26_48 PM.png" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;P&gt;#DatabricksGenie #DatabricksApps #InsuranceAnalytics #DataGovernance&lt;/P&gt;</description>
      <pubDate>Tue, 01 Sep 2026 04:27:53 GMT</pubDate>
      <guid>https://community.databricks.com/t5/community-articles/insurance-intelligence-copilot-powered-by-databricks-genie/m-p/167048#M1505</guid>
      <dc:creator>atharvakahu09</dc:creator>
      <dc:date>2026-09-01T04:27:53Z</dc:date>
    </item>
    <item>
      <title>SleepLens: Turning Multimodal Sleep Data Into Conversations With Databricks Genie</title>
      <link>https://community.databricks.com/t5/community-articles/sleeplens-turning-multimodal-sleep-data-into-conversations-with/m-p/167038#M1504</link>
      <description>&lt;P class=""&gt;The Problem&lt;/P&gt;&lt;P&gt;Sleep trackers collect a lot of information, including heart rate, movement, sleep stages, and wake events. But having more data does not always mean people understand their sleep better. A user might see a Sleep Quality Score of 72.7 and still want to know why it was 72.7. SleepLens was built to explore how Databricks and Genie can make multimodal sleep data easier to understand by letting users ask questions about their sleep in natural language and receive answers grounded in their underlying data.&lt;/P&gt;&lt;P&gt;What Is SleepLens?&lt;BR /&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;/P&gt;&lt;P&gt;SleepLens is an end-to-end multimodal sleep intelligence application built on Databricks. The prototype uses real wearable heart-rate and accelerometer or movement recordings paired with EEG-derived sleep-stage labels. From these signals, SleepLens creates analytics for sleep duration, sleep efficiency, sleep quality, wake events, deep sleep, REM sleep, and physiological and movement-based model features.&lt;/P&gt;&lt;P&gt;How I Built It&lt;/P&gt;&lt;P&gt;SleepLens was designed as a full data-to-AI workflow rather than just a machine-learning model. The overall pipeline moves from raw wearable data into Databricks, then through Bronze, Silver, and Gold layers, followed by machine learning with MLflow, the SleepLens app, and finally a Genie Agent.&lt;/P&gt;&lt;P&gt;Multimodal Sleep Data&lt;/P&gt;&lt;P&gt;The foundation of SleepLens is multimodal physiological data. The prototype combines heart-rate and movement recordings with EEG-derived sleep-stage labels so that multiple signals can be analyzed together. The current prototype also includes simulated contextual variables such as room temperature and sound, and these are not presented as real sensor measurements.&lt;/P&gt;&lt;P&gt;Building the Data Pipeline With Databricks&lt;/P&gt;&lt;P&gt;I organized the data using a Bronze, Silver, and Gold architecture. Bronze stores the raw sleep and wearable data. Silver cleans and transforms those recordings into more useful sleep features. Gold produces application-ready analytics that can be used by the dashboard, machine-learning workflow, and Genie. The final Gold layer includes information about sleep sessions, wake events, and model factors.&lt;/P&gt;&lt;P&gt;Adding Machine Learning&lt;/P&gt;&lt;P&gt;Once the sleep data was structured, I built a machine-learning workflow using physiological and movement features. I used MLflow to track the model and its performance. The resulting model achieved approximately 51.34 percent balanced accuracy. I also incorporated feature importance into the Gold analytics layer so the application could show which signals were contributing to the model's predictions.&lt;/P&gt;&lt;P&gt;The SleepLens Application&lt;/P&gt;&lt;P&gt;I used Databricks Apps to turn the pipeline and model results into an interactive application. The app shows metrics such as Sleep Quality Score, Sleep Duration, and Sleep Efficiency, and it also lets users compare different nights. In the current prototype, one night had a Sleep Quality Score of 76.5 while another had a score of 72.7, which makes it easier to see differences in sleep duration, efficiency, wake events, deep sleep, and REM sleep.&lt;/P&gt;&lt;P&gt;Genie at the Core of SleepLens&lt;/P&gt;&lt;P&gt;The main goal of SleepLens is not just to show sleep data, but to help users understand it. I created the Sleep Analysis and Factors Genie Agent and connected it to the SleepLens application. Genie has access to the Gold analytics tables, including gold_sleep_sessions, gold_wake_events, and gold_model_factors. This lets users ask questions such as which night had better sleep quality and what factors likely contributed to the difference.&lt;/P&gt;&lt;P&gt;Why Genie Made a Difference&lt;/P&gt;&lt;P&gt;Without Genie, SleepLens can tell someone their Sleep Quality Score. With Genie, the application can help them investigate why that score was higher or lower. Instead of requiring users to understand SQL, schemas, tables, feature importance, or model outputs, they can simply ask a question and receive an explanation grounded in the SleepLens data.&lt;/P&gt;&lt;P&gt;Why Databricks Was Useful for SleepLens&lt;/P&gt;&lt;P&gt;Databricks allowed the entire project to live within one ecosystem. I was able to move from raw multimodal wearable data to data engineering, Bronze, Silver, and Gold analytics, machine learning, MLflow tracking, model explainability, an interactive app, and conversational analytics with Genie. This made it possible to build the data pipeline, model, application, and conversational layer as one connected system.&lt;/P&gt;&lt;P&gt;Architecture&lt;/P&gt;&lt;P&gt;The final SleepLens architecture starts with wearable heart-rate and accelerometer data paired with EEG-derived sleep labels. That data moves through the Bronze, Silver, and Gold layers, then into the machine-learning and MLflow workflow, then into the Databricks App, and finally into the SleepLens dashboard and Genie Agent for natural-language sleep insights.&lt;/P&gt;&lt;P&gt;What I Learned&lt;/P&gt;&lt;P&gt;The biggest lesson from building SleepLens was that an AI application is much more than the model itself. The data needs to be cleaned and structured correctly, the model needs to be evaluated, the results need to be understandable, and the application needs to make those results accessible. Building SleepLens showed me how data engineering, machine learning, application development, and conversational AI can work together as one system.&lt;/P&gt;&lt;P&gt;What's Next for SleepLens?&lt;/P&gt;&lt;P&gt;SleepLens is currently a prototype and is not intended to provide medical diagnoses or medical advice. In the future, I would like to explore streaming data from real wearable or edge devices, adding more physiological sensors, replacing simulated environmental context with real temperature and sound sensors, expanding the dataset across more users and nights, improving personalization, exploring on-device inference, and giving Genie access to longer-term sleep trends.&lt;/P&gt;&lt;P&gt;My longer-term vision is for SleepLens to become an intelligent interface between wearable sensor data and human understanding.&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;Main SleepLens Page&lt;BR /&gt;&lt;BR /&gt;&lt;/P&gt;&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="Chais4140_10-1788229256595.png" style="width: 400px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30536i19CF343B08AF98A0/image-size/medium?v=v2&amp;amp;px=400" role="button" title="Chais4140_10-1788229256595.png" alt="Chais4140_10-1788229256595.png" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="Chais4140_11-1788229256598.png" style="width: 400px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30537iAF7660736FAD0E40/image-size/medium?v=v2&amp;amp;px=400" role="button" title="Chais4140_11-1788229256598.png" alt="Chais4140_11-1788229256598.png" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="Chais4140_12-1788229256602.png" style="width: 400px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30538i766B6180C6D6AE21/image-size/medium?v=v2&amp;amp;px=400" role="button" title="Chais4140_12-1788229256602.png" alt="Chais4140_12-1788229256602.png" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;P&gt;Ingested Raw Data&lt;/P&gt;&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="Chais4140_13-1788229256603.png" style="width: 400px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30540iC56E155444872FEB/image-size/medium?v=v2&amp;amp;px=400" role="button" title="Chais4140_13-1788229256603.png" alt="Chais4140_13-1788229256603.png" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;P&gt;Medallion Architecture&amp;nbsp;&lt;/P&gt;&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="Chais4140_14-1788229256604.png" style="width: 400px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30539iA3B81B4AB783DA75/image-size/medium?v=v2&amp;amp;px=400" role="button" title="Chais4140_14-1788229256604.png" alt="Chais4140_14-1788229256604.png" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;P&gt;ML model performance&lt;/P&gt;&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="Chais4140_15-1788229256605.png" style="width: 400px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30542i114F8B75C80A75DD/image-size/medium?v=v2&amp;amp;px=400" role="button" title="Chais4140_15-1788229256605.png" alt="Chais4140_15-1788229256605.png" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="Chais4140_16-1788229256607.png" style="width: 400px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30544iA8C93E877F64B380/image-size/medium?v=v2&amp;amp;px=400" role="button" title="Chais4140_16-1788229256607.png" alt="Chais4140_16-1788229256607.png" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="Chais4140_17-1788229256609.png" style="width: 400px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30543i851AE57416E554E6/image-size/medium?v=v2&amp;amp;px=400" role="button" title="Chais4140_17-1788229256609.png" alt="Chais4140_17-1788229256609.png" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;P&gt;Genie's Impact&amp;nbsp;&lt;/P&gt;&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="Chais4140_18-1788229256612.png" style="width: 400px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30545iFDFE0736F854B4F8/image-size/medium?v=v2&amp;amp;px=400" role="button" title="Chais4140_18-1788229256612.png" alt="Chais4140_18-1788229256612.png" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="Chais4140_19-1788229256613.png" style="width: 400px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30546i25B7AE66D80C90EE/image-size/medium?v=v2&amp;amp;px=400" role="button" title="Chais4140_19-1788229256613.png" alt="Chais4140_19-1788229256613.png" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Tue, 01 Sep 2026 02:26:45 GMT</pubDate>
      <guid>https://community.databricks.com/t5/community-articles/sleeplens-turning-multimodal-sleep-data-into-conversations-with/m-p/167038#M1504</guid>
      <dc:creator>Chais4140</dc:creator>
      <dc:date>2026-09-01T02:26:45Z</dc:date>
    </item>
    <item>
      <title>🧪 MAD DATA LAB: Wonderful. Something Is Wrong.</title>
      <link>https://community.databricks.com/t5/community-articles/mad-data-lab-wonderful-something-is-wrong/m-p/167016#M1503</link>
      <description>&lt;P&gt;Hi, I’m Ángel. I build data systems, break them more often than I would like to admit, and write about what I learn at &lt;STRONG&gt;Angelic Articles&lt;/STRONG&gt; (my new own website for articles, previously on &lt;A href="https://medium.com/@angel.alvarez.pascua" target="_self"&gt;Medium&lt;/A&gt;).&lt;/P&gt;&lt;P&gt;What I enjoy most is usually not getting the answer. It is figuring out &lt;STRONG&gt;why the answer is true&lt;/STRONG&gt;.&lt;/P&gt;&lt;P&gt;That is why the &lt;A href="https://community.databricks.com/t5/learning-events/databricks-community-contest-genie-powered-app-challenge/ec-p/165825" target="_self"&gt;&lt;STRONG&gt;Databricks Genie-Powered App Challenge&lt;/STRONG&gt;&lt;/A&gt; caught my attention — especially &lt;STRONG&gt;"Track B: Creative Thinking"&lt;/STRONG&gt;.&lt;/P&gt;&lt;P&gt;Most AI/BI experiences follow a familiar pattern:&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Ask a question ➔ get an answer.&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;Useful? &lt;span class="lia-unicode-emoji" title=":thinking_face:"&gt;🤔&lt;/span&gt; Absolutely. But for a creative challenge, I wanted to turn that interaction around.&lt;/P&gt;&lt;P&gt;What if Genie did not simply wait for me to ask the right question?&lt;BR /&gt;What if the &lt;STRONG&gt;unexpected number itself became the problem to solve&lt;/STRONG&gt;?&lt;/P&gt;&lt;P&gt;That question became &lt;STRONG&gt;MAD DATA LAB&lt;/STRONG&gt;: a small analytics game where Dr. Genie forms hypotheses, chooses analytical experiments, follows the evidence and only reaches a conclusion when the numbers actually support it.&lt;/P&gt;&lt;P&gt;I decided to submit it because it combines two things I care about: making data exploration more approachable, and treating analytical answers as something to &lt;EM&gt;prove&lt;/EM&gt;, not merely generate.&lt;/P&gt;&lt;P&gt;Because sometimes the interesting part is not the answer. &lt;STRONG&gt;It's proving it.&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;&lt;FONT size="5"&gt;&lt;STRONG&gt;&lt;span class="lia-unicode-emoji" title=":microscope:"&gt;🔬&lt;/span&gt; Welcome to MAD DATA LAB&lt;/STRONG&gt;&lt;/FONT&gt;&lt;/P&gt;&lt;P&gt;MAD DATA LAB is built around one simple idea:&amp;nbsp;&lt;EM&gt;Something in the data is wrong, and you have to prove why.&lt;/EM&gt;&lt;/P&gt;&lt;P&gt;It's designed for analysts, data engineers, data scientists, BI users — and, more generally, anyone who has ever stared at a KPI and thought: &lt;EM&gt;that cannot be right.&lt;/EM&gt;&lt;/P&gt;&lt;P&gt;Each investigation is a &lt;STRONG&gt;Case&lt;/STRONG&gt;. A Case starts with an unexpected result, several plausible explanations and no conclusion that the player is expected to accept on faith.&lt;/P&gt;&lt;P&gt;The challenge demo is "&lt;STRONG&gt;Case #042 — The Missing €6.8M"&lt;/STRONG&gt;:&lt;/P&gt;&lt;LI-CODE lang="markup"&gt;Expected €125.0M
Observed €118.2M
Deviation -€6.8M&lt;/LI-CODE&gt;&lt;P&gt;The data is synthetic and deterministic, so the mystery is reproducible. The analytical evidence, however, is not just story text: the investigation is built around queryable evidence in Databricks.&lt;/P&gt;&lt;P&gt;Dr. Genie begins with competing hypotheses. Perhaps the source values changed. Perhaps the formula changed. Perhaps a suspicious data-quality signal is responsible.&lt;/P&gt;&lt;P&gt;The player predicts, inspects evidence and can ask for help. But the player does not manually choose the analytical route.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;The app asks Genie what should be investigated next.&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;Genie might first decompose the deviation and find that one component, V2, contributes &lt;STRONG&gt;-€5.9M&lt;/STRONG&gt;&amp;nbsp;— roughly &lt;STRONG&gt;87%&lt;/STRONG&gt;&amp;nbsp;of the total anomaly. The next Experiment can compare V2 across source snapshots and find:&lt;/P&gt;&lt;LI-CODE lang="markup"&gt;23 modified records -€5.2M
2 removed records -€0.8M
5 added records +€0.1M
--------------------------------
Net source impact -€5.9M&lt;/LI-CODE&gt;&lt;P&gt;From there, the investigation can move down to individual records, lineage, formula validation, data-quality materiality and final reconciliation.&lt;/P&gt;&lt;P&gt;The loop is deliberately simple:&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Case ➔ Hypotheses ➔ Experiment ➔ Evidence ➔ Update ➔ Repeat ➔ Scientific Verdict&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;And there is no game-over screen for guessing wrong.&lt;/P&gt;&lt;P&gt;The whole point is to watch your first theory collide with the evidence. Because that is usually where analytics gets interesting.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;&lt;FONT size="5"&gt;🧞 Genie Is Not the Hint Button&lt;/FONT&gt;&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;One thing I wanted to avoid was building a normal application and then attaching an AI assistant to the side.&lt;/P&gt;&lt;P&gt;In MAD DATA LAB, Genie is not there to explain a chart after the interesting work is finished.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Genie is part of the investigation loop.&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;At each step, the application gives the Genie Agent the current visible evidence and a server-controlled set of Experiments that are valid at that point. Genie evaluates the evidence, updates the hypotheses and chooses the next analytical move.&lt;/P&gt;&lt;P&gt;The distinction matters:&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;The application controls what is possible. Genie decides what makes sense next.&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;The same rule applies to visualisation. Genie can select from approved analytical &lt;STRONG&gt;Instruments&lt;/STRONG&gt;&amp;nbsp;— a deviation decomposer, snapshot comparison, evidence table, DQ panel, lineage view, reconciliation view and others — but it cannot invent arbitrary UI or executable code at runtime.&lt;/P&gt;&lt;P&gt;Users can also ask Dr. Genie questions such as:&lt;/P&gt;&lt;P&gt;&lt;EM&gt;Which component explains most of the deviation?&lt;/EM&gt;&lt;BR /&gt;&lt;EM&gt;What changed between these two snapshots?&lt;/EM&gt;&lt;BR /&gt;&lt;EM&gt;Is this data-quality warning actually large enough to explain the anomaly?&lt;/EM&gt;&lt;BR /&gt;&lt;EM&gt;Which records contributed most?&lt;/EM&gt;&lt;BR /&gt;&lt;EM&gt;Where did this value come from?&lt;/EM&gt;&lt;/P&gt;&lt;P&gt;Under the hood, the data flow is intentionally small:&lt;/P&gt;&lt;LI-CODE lang="markup"&gt;MAD DATA LAB
  ↓
Investigation / state orchestration
  ↓
Genie Agent
  ↓
Curated Unity Catalog evidence
  ↓
Databricks SQL&lt;/LI-CODE&gt;&lt;P&gt;The application owns state, validation, scoring and safe rendering. Genie works against a curated analytical surface rather than unrestricted data.&lt;/P&gt;&lt;P&gt;There is also one boundary I care about a lot: &lt;STRONG&gt;Genie does not get the answer key&lt;/STRONG&gt;.&lt;/P&gt;&lt;P&gt;The project has a private &lt;EM&gt;CASE_TRUTH&lt;/EM&gt;&amp;nbsp;oracle used for generation and automated validation. It's deliberately excluded from the Genie-facing data model. Dr. Genie has to reach the conclusion from the same visible evidence the investigation exposes.&lt;/P&gt;&lt;P&gt;Remove Genie, and MAD DATA LAB does not become the same game without a chatbot.&lt;/P&gt;&lt;P&gt;It becomes a scripted dashboard wearing a laboratory coat.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;&lt;FONT size="5"&gt;&lt;span class="lia-unicode-emoji" title=":collision:"&gt;💥&lt;/span&gt; Then Reality Entered the Laboratory&lt;/FONT&gt;&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;Of course, the first version of the idea was not the final one.&lt;/P&gt;&lt;P&gt;A few things looked excellent on paper and considerably less excellent five minutes later.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;The beautiful board that was not actually useful&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;My first game-board concept looked wonderfully mad-scientist-ish. It also spent far too much of the screen being decorative.&lt;/P&gt;&lt;P&gt;Very pretty. Very atmospheric. Not particularly useful for investigating data.&lt;/P&gt;&lt;P&gt;So I killed it. The laboratory stayed, but the evidence had to become the protagonist.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Apparently, naming things is hard&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;At one point I was calling the whole investigation an &lt;STRONG&gt;Experiment&lt;/STRONG&gt;… while also calling each individual analytical test an &lt;STRONG&gt;Experiment&lt;/STRONG&gt;. That survived until I tried explaining the game out loud.&lt;/P&gt;&lt;P&gt;The vocabulary became:&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Case ➔ Investigation ➔ Experiment ➔ Evidence ➔ Verdict&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;Much better.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;The scary warning that explained almost nothing&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;Case #042 contains a genuine data-quality signal inside the synthetic Case data:&lt;/P&gt;&lt;P&gt;&lt;EM&gt;5 overlapping business keys. Estimated overlapping impact: -€0.3M.&lt;/EM&gt;&lt;/P&gt;&lt;P&gt;The anomaly is &lt;EM&gt;-€6.8M&lt;/EM&gt;. It's exactly the kind of thing humans — and AI — can jump on because it &lt;EM&gt;looks&lt;/EM&gt;&amp;nbsp;suspicious. But suspicious is not the same as material. And because that -€0.3M overlaps evidence already represented elsewhere, it must not be counted twice.&lt;/P&gt;&lt;P&gt;That became one of the central rules of the game: &lt;EM&gt;A warning is evidence. It is not automatically a cause.&lt;/EM&gt;&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Giving Genie freedom… but not a chainsaw&amp;nbsp;🪚&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;My first instinct was basically: &lt;EM&gt;let the AI decide everything.&lt;/EM&gt;&lt;/P&gt;&lt;P&gt;That sounds elegant until the AI decides it would quite like an analytical instrument your application has never heard of.&lt;/P&gt;&lt;P&gt;The correction was simple: give Genie freedom &lt;EM&gt;inside explicit boundaries&lt;/EM&gt;.&lt;/P&gt;&lt;P&gt;Enough freedom to investigate. Not enough freedom to invent a Quantum Revenue Microscope at runtime.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;&lt;FONT size="5"&gt;🧯 What Worked, What Didn’t, and What Surprised Me&lt;/FONT&gt;&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;A few design decisions made the project much stronger.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Deterministic Cases worked.&lt;/STRONG&gt;&amp;nbsp;If the underlying mystery changes every time, it becomes difficult to tell whether Genie improved or the crime scene simply moved.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Curated evidence worked better than “give it everything.”&lt;/STRONG&gt;&amp;nbsp;The more clearly the tables, fields, semantics and examples described the analytical world, the less I needed to compensate with instructions.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Reconciliation became non-negotiable.&lt;/STRONG&gt;&amp;nbsp;If an explanation claims to account for a €6.8M anomaly, the evidence eventually needs to add up to €6.8M. If it does not, the investigation is not finished.&lt;/P&gt;&lt;P&gt;What I moved away from was equally useful: giant instruction prompts trying to anticipate every situation, free-form chat as the primary game mechanic, arbitrary AI-generated UI, and adding more Cases before Case #042 was trustworthy.&lt;/P&gt;&lt;P&gt;There is also a deliberately boring reliability layer. Model output is validated before it can control the game, and the design includes a deterministic SQL fallback for evidence retrieval when Genie has already selected a valid Experiment but the normal query-result path fails. Offline fixtures are for development or a genuine platform outage — not the normal challenge experience.&lt;/P&gt;&lt;P&gt;The biggest surprise, though, was that the interesting moments were not when Genie instantly found the correct answer. They were when new evidence forced the investigation to change direction.&lt;/P&gt;&lt;P&gt;That felt much closer to real analysis than I expected.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;&lt;FONT size="5"&gt;🧠 What Genie Taught Me About Genie&lt;/FONT&gt;&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;Building MAD DATA LAB changed a few of my assumptions about analytical AI.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Better context beats bigger prompts&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;In this project, clearer semantics, curated data and tested examples were more valuable than increasingly heroic prompt engineering.&lt;/P&gt;&lt;P&gt;That matches the way Genie is designed to be curated: good metadata, business context, example SQL and realistic benchmark questions matter.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Freedom needs boundaries&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;Genie is most useful here when it can decide &lt;EM&gt;what to investigate&lt;/EM&gt;, while the application defines &lt;EM&gt;what is legal and renderable&lt;/EM&gt;.&lt;/P&gt;&lt;P&gt;A fully scripted flow is not very intelligent. An unconstrained AI application is not very predictable. The useful space is in between.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Test the evidence, not the prose&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;I stopped caring whether Genie says "&lt;EM&gt;Aha!"&amp;nbsp;&lt;/EM&gt;or&amp;nbsp;&lt;EM&gt;“Interesting.”&lt;/EM&gt;&lt;/P&gt;&lt;P&gt;What matters is whether it chose a sensible Experiment, retrieved the correct evidence, respected the numbers and reached a conclusion that reconciles.&lt;/P&gt;&lt;P&gt;The wording can vary.&amp;nbsp;&lt;STRONG&gt;The facts cannot.&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;This is also why the test strategy focuses on numeric results, valid Experiment choices, hypothesis status and reconciliation rather than exact sentences.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;“I don’t know yet” is a perfectly good answer&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;One of the easiest mistakes with AI is expecting a confident conclusion every time. But sometimes the evidence is simply not sufficient yet.&lt;/P&gt;&lt;P&gt;MAD DATA LAB uses explicit states such as &lt;EM&gt;POSSIBLE&lt;/EM&gt;, &lt;EM&gt;SUPPORTED&lt;/EM&gt;, &lt;EM&gt;CONFIRMED&lt;/EM&gt;&amp;nbsp;and &lt;EM&gt;RULED OUT&lt;/EM&gt;&amp;nbsp;so that “&lt;EM&gt;plausible&lt;/EM&gt;” does not quietly turn into “&lt;EM&gt;proven&lt;/EM&gt;.”&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;The interesting part is changing your mind&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;The moment that made the concept click for me was not Genie finding the right answer.&lt;/P&gt;&lt;P&gt;It was the investigation updating a hypothesis because new evidence contradicted the previous direction.&lt;/P&gt;&lt;P&gt;That is much closer to useful analysis: &lt;EM&gt;observe, hypothesize, test, revise.&lt;BR /&gt;&lt;/EM&gt;Not: &lt;EM&gt;ask once, sound confident, move on.&lt;/EM&gt;&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;&lt;FONT size="5"&gt;🧪 The Scientific Verdict&lt;/FONT&gt;&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;I joined this challenge wondering whether Genie could do something more interesting than wait for a question.&lt;/P&gt;&lt;P&gt;MAD DATA LAB became my answer. It's playful on the surface, but underneath it is built around a serious analytical idea:&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Do not trust the first explanation just because it sounds plausible. Test it. Quantify it. Reconcile it.&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;That is also what I ended up learning about Genie. The best experience did not come from asking it to sound smarter. It came from giving it better evidence, clearer boundaries and enough freedom to revise the investigation when the data changed the story.&lt;/P&gt;&lt;P&gt;So, after all the hypotheses, experiments, false leads and suspicious numbers, the final verdict is probably the simplest one: &lt;EM&gt;We did not ask for an answer. We ran an investigation.&lt;/EM&gt;&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;&lt;FONT size="5"&gt;&lt;span class="lia-unicode-emoji" title=":wrench:"&gt;🔧&lt;/span&gt; Want the Technical Deep Dive?&lt;/FONT&gt;&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;This article intentionally focused on the idea, the experience, the mistakes and what I learned while building MAD DATA LAB.&lt;/P&gt;&lt;P&gt;For the less sensible amount of technical detail — architecture, Genie conversation orchestration, the closed Experiment/Instrument protocol, deterministic Case generation, curated evidence, hidden ground truth, SQL reconciliation, testing strategy, Genie benchmarks, security boundaries and failure handling — And someday, when I feel like documenting the less sensible amount of technical detail behind all this, I’ll probably turn it into a companion engineering deep dive on Angelic Articles.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;&lt;FONT size="5"&gt;Links&lt;/FONT&gt;&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;&lt;span class="lia-unicode-emoji" title=":laptop_computer:"&gt;💻&lt;/span&gt; Source:&lt;/STRONG&gt; &lt;A href="https://github.com/SauronShepherd/mad-data-lab" target="_blank"&gt;https://github.com/SauronShepherd/mad-data-lab&lt;/A&gt;&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;&lt;span class="lia-unicode-emoji" title=":movie_camera:"&gt;🎥&lt;/span&gt; Video:&lt;/STRONG&gt;&amp;nbsp;&lt;A href="https://youtu.be/RzGWBMzAXVc" target="_blank"&gt;https://youtu.be/RzGWBMzAXVc&lt;/A&gt;&lt;/P&gt;</description>
      <pubDate>Tue, 01 Sep 2026 02:15:14 GMT</pubDate>
      <guid>https://community.databricks.com/t5/community-articles/mad-data-lab-wonderful-something-is-wrong/m-p/167016#M1503</guid>
      <dc:creator>___angel___</dc:creator>
      <dc:date>2026-09-01T02:15:14Z</dc:date>
    </item>
    <item>
      <title>BI Rationalization Genie: Turning Report Sprawl into Conversational Decisions with Databricks Genie</title>
      <link>https://community.databricks.com/t5/community-articles/bi-rationalization-genie-turning-report-sprawl-into/m-p/167005#M1502</link>
      <description>&lt;P&gt;&lt;STRONG&gt;Challenge Track:&lt;/STRONG&gt; Track A — Real-World Problem Solver&lt;BR /&gt;&lt;STRONG&gt;Built on:&lt;/STRONG&gt; Databricks Free Edition&lt;/P&gt;&lt;P&gt;Large enterprises rarely have a reporting problem because they lack reports. More often, they have too many — but the critical problem is determining which reports are genuinely needed, which can be consolidated or retired, and which require careful review before any action is taken.&lt;/P&gt;&lt;P&gt;As BI estates grow across platforms such as OBIEE, Tableau, Cognos, and analytical cubes, reports accumulate over time. Some are duplicates or slight variants. Some were created for one-time needs. Others are no longer actively used.&lt;/P&gt;&lt;P&gt;But rationalization is not as simple as finding an unused report and deleting it.&lt;/P&gt;&lt;P&gt;A low-usage report may still support an executive process, regulatory requirement, customer commitment, financial dependency, or another business-critical function.&lt;/P&gt;&lt;P&gt;That creates the real question:&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;How can business and technical teams understand what should be kept, consolidated, reviewed, or retired — without repeatedly going back to spreadsheets, SQL, and manual analysis?&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;This real enterprise problem inspired my submission for the Databricks Genie-Powered App Challenge: &lt;STRONG&gt;BI Rationalization Genie&lt;/STRONG&gt;.&lt;/P&gt;&lt;P&gt;For the challenge, I built a Databricks App with a Genie Agent at the center of the analytical experience.&lt;/P&gt;&lt;P&gt;The goal was not just to replace SQL with chat.&lt;/P&gt;&lt;P&gt;It was to let different users investigate the same governed BI estate naturally — starting with a high-level question, drilling into why a decision was made, and then looking for broader patterns without having to predefine every analytical path.&lt;/P&gt;&lt;P&gt;The application is designed for BI leaders and executives, business and report owners, and technical teams responsible for rationalization and modernization.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;The Data Behind the App&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;For the challenge, the inventory contains &lt;STRONG&gt;300 reports across four BI platforms: OBIEE, Tableau, Cognos, and AAS&lt;/STRONG&gt;.&lt;/P&gt;&lt;P&gt;The curated inventory contains both report metadata and rationalization evidence, including fields such as:&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;source platform&lt;/LI&gt;&lt;LI&gt;business criticality&lt;/LI&gt;&lt;LI&gt;decision rule&lt;/LI&gt;&lt;LI&gt;proposed action&lt;/LI&gt;&lt;LI&gt;final status&lt;/LI&gt;&lt;LI&gt;mandatory category&lt;/LI&gt;&lt;LI&gt;consolidation target&lt;/LI&gt;&lt;LI&gt;usage and other report-level indicators&lt;/LI&gt;&lt;/UL&gt;&lt;P&gt;The outcome of the upstream rationalization process is one of four dispositions:&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;KEEP | CONSOLIDATE | RETIRE | REVIEW&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;A short note on the upstream process&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;The final inventory does require processing before Genie uses it.&lt;/P&gt;&lt;P&gt;At a high level, the rationalization process combines BI metadata and usage signals, evaluates rationalization indicators such as report duplication, inactivity, one-time reporting patterns, hardcoded logic, filter variants, and column variants, and then applies business-criticality safeguards.&lt;/P&gt;&lt;P&gt;I intentionally kept these governed business rules outside Genie.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;The rationalization process creates evidence and recommendation. Genie lets users investigate that evidence conversationally.&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;This separation was important to me. I did not want an AI agent to invent a retirement decision that should be governed by explicit business rules.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Architecture and Data Flow&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;The application follows a simple flow:&lt;/P&gt;&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="gbhogle1789_0-1788227253501.png" style="width: 400px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30510iED374682935EAD20/image-size/medium?v=v2&amp;amp;px=400" role="button" title="gbhogle1789_0-1788227253501.png" alt="gbhogle1789_0-1788227253501.png" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;The curated inventory provides the analytical foundation.&lt;/P&gt;&lt;P&gt;The Genie Agent contains the business context needed to interpret questions about criticality, rationalization rules, proposed actions, review status, and other inventory concepts.&lt;/P&gt;&lt;P&gt;The Databricks App provides the user-facing experience for asking those questions and continuing the conversation.&lt;/P&gt;&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="gbhogle1789_1-1788227253514.png" style="width: 400px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30508i312C52F79EF5BB8F/image-size/medium?v=v2&amp;amp;px=400" role="button" title="gbhogle1789_1-1788227253514.png" alt="gbhogle1789_1-1788227253514.png" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;One Interface, Different Questions&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;One thing I like about this pattern is that I did not need to design a different dashboard for every audience.&lt;/P&gt;&lt;P&gt;The interface stays the same.&lt;/P&gt;&lt;P&gt;The question changes.&lt;/P&gt;&lt;P&gt;An &lt;STRONG&gt;executive or BI leader&lt;/STRONG&gt; might ask:&lt;/P&gt;&lt;P&gt;Which platforms should we prioritize for rationalization?&lt;/P&gt;&lt;P&gt;A &lt;STRONG&gt;business or report owner&lt;/STRONG&gt; might ask:&lt;/P&gt;&lt;P&gt;Why does this critical report require REVIEW instead of being retired?&lt;/P&gt;&lt;P&gt;A &lt;STRONG&gt;data or BI engineer&lt;/STRONG&gt; might ask:&lt;/P&gt;&lt;P&gt;Which decision rule triggered the recommendation, and what is the consolidation target?&lt;/P&gt;&lt;P&gt;The value is not role-specific screens.&lt;/P&gt;&lt;P&gt;It is allowing each user to follow the analytical path that matters to them using the same governed data.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Why Genie Is at the Core&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;This was the most important question I asked myself while building the app:&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;If Genie disappeared, what would change?&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;The rationalization recommendations would still exist in the inventory.&lt;/P&gt;&lt;P&gt;But the main experience would fundamentally change.&lt;/P&gt;&lt;P&gt;Users would return to predefined dashboards, spreadsheet filters, SQL queries, or requests to technical teams whenever they wanted to explore the estate in a way that had not already been designed for them.&lt;/P&gt;&lt;P&gt;The clearest way to demonstrate the value of Genie is through the different questions users can ask from the same governed BI inventory.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Executive View — Start with the Estate&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;An executive or BI leader can begin with a high-level question:&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Q1. What does the overall BI estate look like across KEEP, RETIRE, CONSOLIDATE, and REVIEW, broken down by source platform?&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;Genie summarizes the estate across platforms and rationalization outcomes, giving leadership a quick view of where reports are being retained, consolidated, retired, or held for review.&lt;/P&gt;&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="gbhogle1789_2-1788227253537.png" style="width: 400px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30509iD6B1FE06AF943157/image-size/medium?v=v2&amp;amp;px=400" role="button" title="gbhogle1789_2-1788227253537.png" alt="gbhogle1789_2-1788227253537.png" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;From there, the user can move from visibility to prioritization:&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Q2. Which source platforms show the largest rationalization opportunity based on retirement, consolidation, and review patterns? Explain using only the available data.&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;Instead of requiring a predefined dashboard for every possible comparison, Genie can analyze the governed inventory across source platform, proposed action, final status, criticality, and rationalization patterns.&lt;/P&gt;&lt;P&gt;This gives leadership a way to explore the estate based on the question they need answered at that moment.&lt;/P&gt;&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="gbhogle1789_3-1788227253544.png" style="width: 400px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30511iEDE00BF01B503615/image-size/medium?v=v2&amp;amp;px=400" role="button" title="gbhogle1789_3-1788227253544.png" alt="gbhogle1789_3-1788227253544.png" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Technical View — Investigate the Decision&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;Technical users can take the conversation deeper.&lt;/P&gt;&lt;P&gt;I asked:&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Which critical reports were recommended for retirement or consolidation but require manual review instead?&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;Genie identified the affected reports and explained the relationship between the original rationalization recommendation, business criticality, and the final REVIEW status.&lt;/P&gt;&lt;P&gt;The next turn demonstrates an important part of the conversational experience.&lt;/P&gt;&lt;P&gt;I did not repeat the population report or rebuild the previous filters.&lt;/P&gt;&lt;P&gt;I simply asked:&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Of those, explain which ones were originally consolidation candidates, what decision rule triggered each recommendation, what target report they consolidate into, and why criticality changed the final status to REVIEW.&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;Genie understood what &lt;STRONG&gt;“those”&lt;/STRONG&gt; referred to from the previous turn.&lt;/P&gt;&lt;P&gt;It continued from the existing conversation, narrowed the analysis to the consolidation candidates, connected each recommendation to its rationalization rule and target report, and explained why the criticality safeguard prevented automatic consolidation.&lt;/P&gt;&lt;P&gt;The user does not need to know fields such as critical_flag, decision_rules, target_report_id, proposed_action, or final_status.&lt;/P&gt;&lt;P&gt;They can continue the investigation using business language.&lt;/P&gt;&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="gbhogle1789_4-1788227253554.png" style="width: 400px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30513i1A800BC070BAF8B3/image-size/medium?v=v2&amp;amp;px=400" role="button" title="gbhogle1789_4-1788227253554.png" alt="gbhogle1789_4-1788227253554.png" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="gbhogle1789_5-1788227253561.png" style="width: 400px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30512i2BF63ECA994C708D/image-size/medium?v=v2&amp;amp;px=400" role="button" title="gbhogle1789_5-1788227253561.png" alt="gbhogle1789_5-1788227253561.png" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;From Investigation to Engineering Prioritization&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;The analytical path can then shift again.&lt;/P&gt;&lt;P&gt;A technical team preparing for modernization can ask:&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Based on the rationalization patterns, which platforms and decision-rule categories should the engineering team investigate first for migration, consolidation, or retirement? Use only measurable evidence from the available data and do not assume timelines, effort, or business impact.&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;This question moves beyond identifying individual reports.&lt;/P&gt;&lt;P&gt;Genie can compare measurable patterns across the estate—such as report volume, rationalization actions, decision rules, criticality, and action readiness—to help engineering teams determine where deeper investigation should begin.&lt;/P&gt;&lt;P&gt;Importantly, Genie is instructed not to invent implementation timelines, effort estimates, or business impact when those facts are not present in the data.&lt;/P&gt;&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="gbhogle1789_6-1788227253568.png" style="width: 400px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30514i4C8FD47861EA6F63/image-size/medium?v=v2&amp;amp;px=400" role="button" title="gbhogle1789_6-1788227253568.png" alt="gbhogle1789_6-1788227253568.png" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="gbhogle1789_7-1788227253575.png" style="width: 400px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30515i36CD4D7D49FBD5D7/image-size/medium?v=v2&amp;amp;px=400" role="button" title="gbhogle1789_7-1788227253575.png" alt="gbhogle1789_7-1788227253575.png" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;This is the experience I wanted to create.&lt;/P&gt;&lt;P&gt;The same governed inventory supports several levels of investigation:&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;What does the estate look like?&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;→ &lt;STRONG&gt;Where are the largest rationalization opportunities?&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;→ &lt;STRONG&gt;Which reports require attention and why?&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;→ &lt;STRONG&gt;What evidence produced those decisions?&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;→ &lt;STRONG&gt;Where should engineering investigate next?&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;The analytical path is not hard-coded into a collection of separate dashboards.&lt;/P&gt;&lt;P&gt;It emerges from the user's questions and can continue as the investigation evolves.&lt;/P&gt;&lt;P&gt;That is where Genie becomes load-bearing in this application.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;What Else Can Users Ask?&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;The same inventory also supports questions such as:&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;Which reports appear to be one-time or ad-hoc reports?&lt;/LI&gt;&lt;LI&gt;Which reports contain hardcoded dates?&lt;/LI&gt;&lt;LI&gt;Which reports can be consolidated, and what are their target reports?&lt;/LI&gt;&lt;LI&gt;How do rationalization outcomes differ across BI platforms?&lt;/LI&gt;&lt;LI&gt;Which critical reports prevent an otherwise automated rationalization action?&lt;/LI&gt;&lt;/UL&gt;&lt;P&gt;The user does not need to understand the schema before beginning the analysis.&lt;/P&gt;&lt;P&gt;They start with the business question.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Building for Accuracy Was a Bigger Part of the Work Than I Expected&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;Connecting Genie to the inventory was not the difficult part.&lt;/P&gt;&lt;P&gt;Getting consistently correct answers was where I spent more time.&lt;/P&gt;&lt;P&gt;I created benchmark questions around important rationalization scenarios and compared Genie's responses with the expected results.&lt;/P&gt;&lt;P&gt;After iterating on the semantic context and testing the generated queries, I completed a clean benchmark run in which &lt;STRONG&gt;all five benchmark questions passed&lt;/STRONG&gt;.&lt;/P&gt;&lt;P&gt;The benchmark set covered important rationalization scenarios such as critical-report safeguards, retirement recommendations, consolidation candidates, one-time reports, and hardcoded-date patterns, giving me a repeatable way to validate Genie’s accuracy against known expected results.&lt;/P&gt;&lt;P&gt;One debugging habit became particularly useful:&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;When Genie gave me an incorrect answer, I stopped immediately adding more instructions.&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;Instead, I checked:&lt;/P&gt;&lt;OL&gt;&lt;LI&gt;Is the underlying data correct?&lt;/LI&gt;&lt;LI&gt;What SQL did Genie generate?&lt;/LI&gt;&lt;LI&gt;Is the business terminology clear?&lt;/LI&gt;&lt;LI&gt;Does Genie have the semantic context needed to interpret the question?&lt;/LI&gt;&lt;LI&gt;Is my expected benchmark answer itself defined correctly?&lt;/LI&gt;&lt;/OL&gt;&lt;P&gt;That process led to one of my most useful learnings from the challenge.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Instructions and evaluation notes are not the same thing&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;At first, when an evaluation failed, the natural reaction was to think: &lt;EM&gt;I need another instruction.&lt;/EM&gt;&lt;/P&gt;&lt;P&gt;That is not always the right fix.&lt;/P&gt;&lt;P&gt;I started using a much simpler distinction:&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Instructions help guide Genie. Evaluation notes help evaluate Genie.&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;If Genie repeatedly needs business context to interpret future questions correctly, that belongs in its instructions or semantic knowledge.&lt;/P&gt;&lt;P&gt;If the issue is defining what the benchmark judge should consider a correct response, that belongs in the evaluation note.&lt;/P&gt;&lt;P&gt;This small distinction helped me avoid continuously adding instructions just to make an evaluation pass.&lt;/P&gt;&lt;P&gt;I also learned that inspecting the generated SQL can often tell you more about an accuracy problem than adding another paragraph of prompting.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Another Lesson: Deployed Does Not Always Mean Accessible&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;The app integration gave me another very practical lesson.&lt;/P&gt;&lt;P&gt;At one point, the Databricks App was successfully deployed and running, but that did not automatically mean the complete Genie experience was accessible.&lt;/P&gt;&lt;P&gt;The application identity also needed the appropriate access to the Genie Agent and the required underlying resources.&lt;/P&gt;&lt;P&gt;That changed how I thought about testing the solution:&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Test the Genie Agent. Test the app. Then test the authorization path between them.&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;For an enterprise implementation, the same principle extends to governance.&lt;/P&gt;&lt;P&gt;The conversational interface should make analytics easier to use, but it should not create a shortcut around the controls protecting the underlying data. Access to the app, Genie, and the governed inventory should continue to follow the appropriate Databricks authorization and Unity Catalog controls.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;What I Took Away From the Challenge&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;Three lessons stand out for me.&lt;/P&gt;&lt;OL&gt;&lt;LI&gt;&lt;STRONG&gt;Do not treat every Genie accuracy problem as a prompting problem.&lt;/STRONG&gt;&lt;BR /&gt;Look at the data, generated SQL, semantics, and expected answer first.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Conversation changes the value of analytics.&lt;/STRONG&gt;&lt;BR /&gt;The useful part is not only asking the first natural-language question. It is being able to continue from that answer, change analytical direction, and explore something that was not predefined in the UI.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;A working Genie Agent and a working Genie-powered application are two different milestones.&lt;/STRONG&gt;&lt;BR /&gt;Integration, authorization, conversation handling, and application behavior need their own testing.&lt;/LI&gt;&lt;/OL&gt;&lt;P&gt;&lt;STRONG&gt;Closing Thought&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;I started this project with a report-rationalization problem.&lt;/P&gt;&lt;P&gt;The upstream process can already calculate governed recommendations such as KEEP, CONSOLIDATE, RETIRE, and REVIEW.&lt;/P&gt;&lt;P&gt;But a recommendation sitting in a table is not the same as making that information easy for people to investigate.&lt;/P&gt;&lt;P&gt;That is the role Genie plays in this application.&lt;/P&gt;&lt;P&gt;A BI leader can start at the estate level.&lt;/P&gt;&lt;P&gt;A report owner can ask why a particular decision was made.&lt;/P&gt;&lt;P&gt;A technical user can drill into the rule and supporting evidence.&lt;/P&gt;&lt;P&gt;And they can move between those questions through a conversation rather than waiting for another query, spreadsheet, or dashboard to be built.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;The rationalization process creates evidence. Genie turns that evidence into an investigation.&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;#Databricks #DatabricksGenie #DatabricksApps #DataAI #BusinessIntelligence #DataModernization #GeniePoweredAppChallenge&lt;/P&gt;</description>
      <pubDate>Tue, 01 Sep 2026 01:49:52 GMT</pubDate>
      <guid>https://community.databricks.com/t5/community-articles/bi-rationalization-genie-turning-report-sprawl-into/m-p/167005#M1502</guid>
      <dc:creator>gbhogle1789</dc:creator>
      <dc:date>2026-09-01T01:49:52Z</dc:date>
    </item>
    <item>
      <title>Crux — Construction Intelligence Powered by Databricks Genie</title>
      <link>https://community.databricks.com/t5/community-articles/crux-construction-intelligence-powered-by-databricks-genie/m-p/166985#M1501</link>
      <description>&lt;P&gt;&lt;FONT size="2"&gt;&lt;EM&gt;Databricks Community Genie-Powered App Challenge. Track A: Real-World Problem Solver.&lt;/EM&gt;&lt;/FONT&gt;&lt;/P&gt;&lt;P&gt;Managing a construction portfolio means constantly answering questions across schedule, cost, tasks, contractors, inspections, change orders, and field reports. The problem is that the answer to a single question can be split across structured project data and pages of project documentation.&lt;/P&gt;&lt;P class=""&gt;&lt;STRONG&gt;Crux&lt;/STRONG&gt; brings those two worlds together.&lt;/P&gt;&lt;P&gt;&lt;FONT size="3"&gt;A user asks a question in plain English. Databricks Genie establishes the authoritative structured picture: project risk, schedule variance, forecast cost, blockers, overdue work, contractor performance, and portfolio comparisons. When the question also needs context, Crux retrieves supporting project evidence that explains what is happening and why.&lt;/FONT&gt;&lt;/P&gt;&lt;P&gt;&lt;FONT size="3"&gt;Genie discovers and quantifies the problem. Documentary intelligence explains and substantiates it. Crux turns both into a decision.&lt;/FONT&gt;&lt;/P&gt;&lt;P&gt;&lt;FONT size="3"&gt;&lt;STRONG&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="Screenshot 2026-08-31 at 8.05.26 PM.png" style="width: 400px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30488i91B46A54B24554B8/image-size/medium?v=v2&amp;amp;px=400" role="button" title="Screenshot 2026-08-31 at 8.05.26 PM.png" alt="Screenshot 2026-08-31 at 8.05.26 PM.png" /&gt;&lt;/span&gt;&lt;/STRONG&gt;&lt;/FONT&gt;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;H2&gt;The idea, and who it is for&lt;/H2&gt;&lt;P&gt;Consider a project manager trying to understand why a project needs attention.&lt;/P&gt;&lt;P&gt;The schedule may show a delay. The cost data may show a forecast overrun. Tasks may reveal overdue work. An inspection record may identify a blocker. A field report may finally explain the operational reason behind it.&lt;/P&gt;&lt;P&gt;Getting the complete picture normally means moving between different systems, reports, and datasets &amp;nbsp;or relying on someone who knows how to query them.&lt;/P&gt;&lt;P&gt;Crux is built for project managers, operations teams, and construction leaders who want to ask those questions directly:&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;&lt;FONT size="2"&gt;&lt;EM&gt;Which projects are at high risk?&lt;/EM&gt;&lt;/FONT&gt;&lt;/LI&gt;&lt;LI&gt;&lt;FONT size="2"&gt;&lt;EM&gt;Why is Stellar Heights at risk?&lt;/EM&gt;&lt;/FONT&gt;&lt;/LI&gt;&lt;LI&gt;&lt;FONT size="2"&gt;&lt;EM&gt;What inspection is blocking it?&lt;/EM&gt;&lt;/FONT&gt;&lt;/LI&gt;&lt;LI&gt;&lt;FONT size="2"&gt;&lt;EM&gt;What about its forecast variance?&lt;/EM&gt;&lt;/FONT&gt;&lt;/LI&gt;&lt;/UL&gt;&lt;P&gt;The user does not need to know the schema, write SQL, or manually search through project documents.&lt;/P&gt;&lt;P class=""&gt;One design decision became fundamental while building Crux:&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Not every source should have equal authority.&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;Genie owns the structured business facts. Project documents provide supporting explanation and evidence.&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;H2&gt;One project, end to end&lt;/H2&gt;&lt;P&gt;Take &lt;STRONG&gt;Stellar Heights&lt;/STRONG&gt;.&lt;/P&gt;&lt;P&gt;Crux includes a Projects workspace where the portfolio can be explored before moving into a deeper investigation. Selecting Stellar Heights takes the user into the Intelligence workspace with the project context prepared:&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;&lt;FONT size="2"&gt;&lt;EM&gt;What should I know about Stellar Heights?&lt;/EM&gt;&lt;/FONT&gt;&lt;/LI&gt;&lt;/UL&gt;&lt;P&gt;Genie establishes the authoritative structured state:&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;&lt;STRONG&gt;High Risk&lt;/STRONG&gt;&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;18 days behind schedule&lt;/STRONG&gt;&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;8.10% forecast cost variance&lt;/STRONG&gt;&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;$3,928,500 forecast overrun&lt;/STRONG&gt;&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;7 overdue tasks&lt;/STRONG&gt;&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;1 active blocker&lt;/STRONG&gt;&lt;/LI&gt;&lt;/UL&gt;&lt;P&gt;Those numbers answer &lt;STRONG&gt;what is happening&lt;/STRONG&gt;.&lt;/P&gt;&lt;P&gt;But construction decisions often require the next question: &lt;STRONG&gt;why?&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;For the same inquiry, Crux retrieves relevant documentary evidence describing issues affecting Stellar Heights, including an HVAC fabrication backlog and staffing constraints, the &lt;STRONG&gt;INS-2003 electrical inspection issue&lt;/STRONG&gt;, and its downstream impact on drywall work.&lt;/P&gt;&lt;P&gt;The two intelligence sources contribute to the same experience without competing for authority.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Genie establishes the measurable project state. Documentary intelligence explains the conditions behind it.&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-left" image-alt="Screenshot 2026-08-31 at 8.14.31 PM.png" style="width: 400px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30489i87C763C196C8E3FB/image-size/medium?v=v2&amp;amp;px=400" role="button" title="Screenshot 2026-08-31 at 8.14.31 PM.png" alt="Screenshot 2026-08-31 at 8.14.31 PM.png" /&gt;&lt;/span&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-left" image-alt="Screenshot 2026-08-31 at 8.15.09 PM.png" style="width: 400px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30490i7B0DB5D568166608/image-size/medium?v=v2&amp;amp;px=400" role="button" title="Screenshot 2026-08-31 at 8.15.09 PM.png" alt="Screenshot 2026-08-31 at 8.15.09 PM.png" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;BR /&gt;The answer is also inspectable.&lt;/P&gt;&lt;P&gt;Documentary claims include citations back to the retrieved project evidence, with source metadata such as the document and relevant location in the source.&lt;/P&gt;&lt;P&gt;The structured analysis is transparent too. Crux exposes the &lt;STRONG&gt;SQL generated by Genie&lt;/STRONG&gt; alongside the result, so the user can inspect how the structured answer was produced instead of treating it as a black box.&lt;/P&gt;&lt;P&gt;The investigation can then continue naturally.&lt;/P&gt;&lt;P&gt;Ask:&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;&lt;FONT size="2"&gt;&lt;EM&gt;What inspection is blocking it?&lt;/EM&gt;&lt;/FONT&gt;&lt;/LI&gt;&lt;/UL&gt;&lt;P&gt;Crux understands that the conversation still refers to Stellar Heights and identifies &lt;STRONG&gt;INS-2003&lt;/STRONG&gt;, together with the relevant supporting evidence.&lt;/P&gt;&lt;P&gt;Then ask:&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;&lt;FONT size="2"&gt;&lt;EM&gt;What about its forecast variance?&lt;/EM&gt;&lt;/FONT&gt;&lt;/LI&gt;&lt;/UL&gt;&lt;P&gt;Genie continues with the existing project context and returns the forecast variance without requiring &lt;EM&gt;Stellar Heights&lt;/EM&gt; to be repeated.&lt;/P&gt;&lt;P&gt;The previous turns remain available in the same intelligence thread, including structured results, documentary evidence, citations, and generated SQL.&lt;BR /&gt;&lt;BR /&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="Screenshot 2026-08-31 at 8.20.56 PM.png" style="width: 400px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30491i9F5C7F72330CA0BC/image-size/medium?v=v2&amp;amp;px=400" role="button" title="Screenshot 2026-08-31 at 8.20.56 PM.png" alt="Screenshot 2026-08-31 at 8.20.56 PM.png" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="Screenshot 2026-08-31 at 8.22.44 PM.png" style="width: 400px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30492iDEFC3184E053955F/image-size/medium?v=v2&amp;amp;px=400" role="button" title="Screenshot 2026-08-31 at 8.22.44 PM.png" alt="Screenshot 2026-08-31 at 8.22.44 PM.png" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;H2&gt;Genie at the core&lt;/H2&gt;&lt;P&gt;Remove Genie from Crux and the core product no longer works.&lt;/P&gt;&lt;P&gt;The documentary layer could still retrieve project reports, but Crux could no longer establish authoritative project risk, schedule variance, forecast cost variance, overdue work, blockers, contractor performance, rankings, or portfolio-level metrics.&lt;/P&gt;&lt;P&gt;Genie is load-bearing in several ways.&lt;/P&gt;&lt;OL&gt;&lt;LI&gt;&lt;STRONG&gt;Genie owns the structured truth.&lt;/STRONG&gt;&lt;BR /&gt;Project and portfolio metrics come from Genie over governed Databricks data. Crux does not independently recalculate those metrics after Genie responds.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Genie turns business questions into structured analysis.&lt;/STRONG&gt;&lt;BR /&gt;The user asks in natural language. Genie determines the relevant data, generates the SQL, executes the analysis, and returns the structured result. Crux preserves and exposes the generated SQL.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Genie holds the structured conversation.&lt;/STRONG&gt;&lt;BR /&gt;Follow-up questions continue the same Genie conversation. A question such as &lt;EM&gt;“What about its forecast variance?”&lt;/EM&gt; can therefore build on the previous discussion instead of starting from zero.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;The Genie Space carries the construction semantics.&lt;/STRONG&gt;&lt;BR /&gt;Instructions, business definitions, relationships, and verified examples guide how concepts such as project risk, schedule variance, blockers, change orders, and contractor performance should be interpreted.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Genie stays authoritative when documents enter the answer.&lt;/STRONG&gt;&lt;BR /&gt;A project report may explain why something happened, but it cannot redefine a Genie-owned KPI. Documentary evidence can support the metric; it does not become the source of truth for it.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Crux decides when Genie is required.&lt;/STRONG&gt;&lt;BR /&gt;A supervisor routes each question into one of three paths: GENIE_ONLY, RAG_ONLY, or GENIE_PLUS_RAG. Structured business questions go through Genie, documentary questions can use retrieval, and broader project questions can invoke both.&lt;/LI&gt;&lt;/OL&gt;&lt;P&gt;That responsibility boundary became the core of the architecture:&lt;/P&gt;&lt;P&gt;Genie tells Crux what is true in the structured portfolio. Documentary intelligence helps explain why.&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;H2&gt;What you can ask it&lt;/H2&gt;&lt;P&gt;Crux is not built around a fixed set of dashboard filters or pre-written queries.&lt;/P&gt;&lt;P&gt;Users can ask questions such as:&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;&lt;FONT size="2"&gt;&lt;EM&gt;Which projects are at high risk?&lt;/EM&gt;&lt;/FONT&gt;&lt;/LI&gt;&lt;LI&gt;&lt;FONT size="2"&gt;&lt;EM&gt;Which of those is the most delayed?&lt;/EM&gt;&lt;/FONT&gt;&lt;/LI&gt;&lt;LI&gt;&lt;FONT size="2"&gt;&lt;EM&gt;What should I know about Stellar Heights?&lt;/EM&gt;&lt;/FONT&gt;&lt;/LI&gt;&lt;LI&gt;&lt;FONT size="2"&gt;&lt;EM&gt;What inspection is blocking it?&lt;/EM&gt;&lt;/FONT&gt;&lt;/LI&gt;&lt;LI&gt;&lt;FONT size="2"&gt;&lt;EM&gt;What about its forecast variance?&lt;/EM&gt;&lt;/FONT&gt;&lt;/LI&gt;&lt;LI&gt;&lt;FONT size="2"&gt;&lt;EM&gt;Which contractor has the most task slippage?&lt;/EM&gt;&lt;/FONT&gt;&lt;/LI&gt;&lt;/UL&gt;&lt;P&gt;The user does not need to know which table contains the answer, how the tables join, which document mentions the issue, or what SQL should be written.&lt;/P&gt;&lt;P&gt;For a portfolio question, Genie performs the structured analysis.&lt;/P&gt;&lt;P&gt;For a documentary question, Crux retrieves the relevant project evidence.&lt;/P&gt;&lt;P&gt;For questions requiring both, the two paths are combined while maintaining their authority boundary.&lt;/P&gt;&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="Screenshot 2026-08-31 at 8.30.47 PM.png" style="width: 400px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30493iA96D5D4137357FF0/image-size/medium?v=v2&amp;amp;px=400" role="button" title="Screenshot 2026-08-31 at 8.30.47 PM.png" alt="Screenshot 2026-08-31 at 8.30.47 PM.png" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;H2&gt;Architecture&lt;/H2&gt;&lt;P&gt;Crux runs as a &lt;STRONG&gt;Databricks App&lt;/STRONG&gt;.&lt;/P&gt;&lt;P&gt;A Databricks-served supervisor receives the user's question and chooses between structured intelligence, documentary intelligence, or both.&lt;/P&gt;&lt;P&gt;The structured path uses &lt;STRONG&gt;Databricks Genie&lt;/STRONG&gt; over governed Unity Catalog data for projects, project costs, tasks, contractors, inspections, and change orders through a SQL warehouse.&lt;/P&gt;&lt;P&gt;The documentary path begins with project files stored in &lt;STRONG&gt;Unity Catalog Volumes&lt;/STRONG&gt;. Documents are parsed and deterministically chunked into governed tables, then indexed with &lt;STRONG&gt;Databricks Vector Search&lt;/STRONG&gt; using hybrid retrieval.&lt;/P&gt;&lt;P&gt;For mixed questions, Genie establishes the authoritative structured state while the retrieval pipeline supplies supporting documentary evidence. Crux then combines the two deterministically so that the evidence layer cannot overwrite Genie-authoritative metrics.&lt;/P&gt;&lt;P&gt;The application also preserves the active Genie conversation and bounded project context for multi-turn questions.&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="Screenshot 2026-08-31 at 8.50.17 PM.png" style="width: 400px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30494i4B08F8A1ED4114ED/image-size/medium?v=v2&amp;amp;px=400" role="button" title="Screenshot 2026-08-31 at 8.50.17 PM.png" alt="Screenshot 2026-08-31 at 8.50.17 PM.png" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;H2&gt;What I learned building it&lt;/H2&gt;&lt;UL&gt;&lt;LI&gt;&lt;STRONG&gt;Authority was harder than routing.&lt;/STRONG&gt;&lt;BR /&gt;Deciding whether to call Genie or retrieval was only part of the problem. Once both systems contribute to an answer, defining which system is allowed to state which facts becomes much more important.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;More model calls did not mean a better system.&lt;/STRONG&gt;&lt;BR /&gt;An earlier architecture included additional model calls for documentary planning and final synthesis. Evaluation showed they were unnecessary. Removing them produced a simpler mixed path while keeping Genie authoritative.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;RAG quality starts before retrieval.&lt;/STRONG&gt;&lt;BR /&gt;Document identity, parsing, chunking, versioning, serving eligibility, indexing, and citation metadata all had to be deterministic before the retrieval layer became dependable.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;A chat interface does not automatically create a conversation.&lt;/STRONG&gt;&lt;BR /&gt;Multi-turn behavior required preserving the real Genie conversation and resolved project context. After a reset, an ambiguous question such as &lt;EM&gt;“What about its forecast variance?”&lt;/EM&gt; correctly loses that context.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Transparency belongs in the product.&lt;/STRONG&gt;&lt;BR /&gt;Generated SQL and documentary citations are visible to the user. Natural language makes Crux accessible; inspectability makes the answer useful for decision-making.&lt;/LI&gt;&lt;/UL&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;H2&gt;The result&lt;/H2&gt;&lt;P&gt;Construction intelligence is not only about finding a number.&lt;/P&gt;&lt;P&gt;A project leader needs to know which project needs attention, how significant the problem is, what is causing it, and what evidence supports that conclusion.&lt;/P&gt;&lt;P&gt;Databricks Genie provides the structured intelligence core that makes those questions accessible through natural language.&lt;/P&gt;&lt;P&gt;Crux builds the decision workflow around it.&lt;/P&gt;&lt;P&gt;Genie discovers and quantifies the problem.&lt;/P&gt;&lt;P&gt;Documentary intelligence explains and substantiates it.&lt;/P&gt;&lt;P&gt;Crux turns both into a decision.&lt;/P&gt;&lt;P&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="Screenshot 2026-08-31 at 7.48.52 PM.png" style="width: 400px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30495i8DA1B5379155C11D/image-size/medium?v=v2&amp;amp;px=400" role="button" title="Screenshot 2026-08-31 at 7.48.52 PM.png" alt="Screenshot 2026-08-31 at 7.48.52 PM.png" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;H3&gt;CRUX -&lt;FONT size="3"&gt;&amp;nbsp;&lt;STRONG&gt;Construction Intelligence, powered by Databricks Genie&lt;/STRONG&gt;&lt;/FONT&gt;&lt;/H3&gt;</description>
      <pubDate>Tue, 01 Sep 2026 01:13:32 GMT</pubDate>
      <guid>https://community.databricks.com/t5/community-articles/crux-construction-intelligence-powered-by-databricks-genie/m-p/166985#M1501</guid>
      <dc:creator>Areddy10</dc:creator>
      <dc:date>2026-09-01T01:13:32Z</dc:date>
    </item>
    <item>
      <title>OpsPulse- From Operational Signals to Data-Backed Answers with AIBI Genie</title>
      <link>https://community.databricks.com/t5/community-articles/opspulse-from-operational-signals-to-data-backed-answers-with/m-p/166979#M1500</link>
      <description>&lt;H1&gt;OpsPulse- From Operational Signals to Data-Backed Answers with AIBI Genie&lt;/H1&gt;&lt;P&gt;Operations teams usually do not have a shortage of data. They have dashboards, reports, KPIs, spreadsheets, and weekly reviews.&lt;/P&gt;&lt;P&gt;The harder part often starts after someone notices that a metric has moved.&lt;/P&gt;&lt;P&gt;Why is backlog increasing?&lt;BR /&gt;Is the issue higher inbound volume or lower processing capacity?&lt;BR /&gt;Which vendor is actually deteriorating?&lt;BR /&gt;Is an increase meaningful, or just normal variation?&lt;/P&gt;&lt;P&gt;I built &lt;STRONG&gt;OpsPulse&lt;/STRONG&gt; for that part of the analytics workflow.&lt;/P&gt;&lt;P&gt;OpsPulse is a Genie-powered operations intelligence app built on &lt;STRONG&gt;Databricks Free Edition&lt;/STRONG&gt;. It allows an operations manager or analyst to investigate backlog, SLA attainment, throughput, defect rates, downtime, vendors, sites, and processes using natural-language questions.&lt;/P&gt;&lt;P&gt;My goal was not to build another dashboard or simply put a chat interface on top of a table. I wanted to see how far I could take a governed operational analytics experience where the user can move from:&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Detect → Diagnose → Investigate&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;and where the system also knows when the available data is not sufficient to answer a question.&lt;/P&gt;&lt;H2&gt;The use case&lt;/H2&gt;&lt;P&gt;Imagine an operations manager opens the application and asks:&lt;/P&gt;&lt;BLOCKQUOTE&gt;&lt;P&gt;What operational issues need attention in the latest available data?&lt;/P&gt;&lt;/BLOCKQUOTE&gt;&lt;P&gt;Instead of requiring the manager to open multiple dashboards and compare several KPIs manually, OpsPulse evaluates recent operational performance against an appropriate historical baseline.&lt;/P&gt;&lt;P&gt;It can surface issues such as:&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;&lt;P&gt;backlog growth&lt;/P&gt;&lt;/LI&gt;&lt;LI&gt;&lt;P&gt;SLA deterioration&lt;/P&gt;&lt;/LI&gt;&lt;LI&gt;&lt;P&gt;declining throughput&lt;/P&gt;&lt;/LI&gt;&lt;LI&gt;&lt;P&gt;increasing defect rates&lt;/P&gt;&lt;/LI&gt;&lt;LI&gt;&lt;P&gt;rising downtime&lt;/P&gt;&lt;/LI&gt;&lt;/UL&gt;&lt;P&gt;The user can then continue the investigation naturally.&lt;/P&gt;&lt;P&gt;For example:&lt;/P&gt;&lt;BLOCKQUOTE&gt;&lt;P&gt;Why did East Hub backlog increase after August 18?&lt;/P&gt;&lt;/BLOCKQUOTE&gt;&lt;P&gt;In the test data, OpsPulse found that average daily inbound volume had increased to approximately &lt;STRONG&gt;2,742 units&lt;/STRONG&gt;, while average daily processed volume reached only about &lt;STRONG&gt;2,288 units&lt;/STRONG&gt;.&lt;/P&gt;&lt;P&gt;That created an average processing gap of approximately &lt;STRONG&gt;454 units per day&lt;/STRONG&gt;.&lt;/P&gt;&lt;P&gt;At the same time, average daily downtime increased from approximately &lt;STRONG&gt;226 minutes before August 18 to more than 1,029 minutes after August 18&lt;/STRONG&gt;.&lt;/P&gt;&lt;P&gt;This gives the user more than a statement that "backlog increased." It provides the operational context needed to understand what changed.&lt;/P&gt;&lt;H2&gt;Who I designed OpsPulse for&lt;/H2&gt;&lt;P&gt;I designed the app primarily for:&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;&lt;P&gt;Operations managers&lt;/P&gt;&lt;/LI&gt;&lt;LI&gt;&lt;P&gt;Business analysts&lt;/P&gt;&lt;/LI&gt;&lt;LI&gt;&lt;P&gt;Data analysts&lt;/P&gt;&lt;/LI&gt;&lt;LI&gt;&lt;P&gt;Supply chain and logistics teams&lt;/P&gt;&lt;/LI&gt;&lt;LI&gt;&lt;P&gt;BI teams supporting operational decision-making&lt;/P&gt;&lt;/LI&gt;&lt;/UL&gt;&lt;P&gt;These users may understand their business very well without necessarily wanting to write SQL every time they have a follow-up question.&lt;/P&gt;&lt;P&gt;The value of Genie in this scenario is the ability to keep the investigation moving.&lt;/P&gt;&lt;P&gt;A user can start with a broad question, identify an exception, and immediately ask a more specific question without switching tools or rebuilding an analysis.&lt;/P&gt;&lt;H2&gt;Architecture&lt;/H2&gt;&lt;P&gt;The application uses a relatively simple architecture, but I spent quite a bit of time on the semantic and analytical layer behind Genie.&lt;/P&gt;&lt;PRE&gt;Synthetic Operations Data
        ↓
Delta Table
operations_daily
        ↓
Curated KPI View
operations_kpi_view
        ↓
Unity Catalog Metadata
+ Governed KPI Definitions
        ↓
AI/BI Genie Agent
        ↓
Databricks AppKit
        ↓
OpsPulse&lt;/PRE&gt;&lt;P&gt;The main components are:&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;&lt;P&gt;&lt;STRONG&gt;Databricks Free Edition&lt;/STRONG&gt;&lt;/P&gt;&lt;/LI&gt;&lt;LI&gt;&lt;P&gt;&lt;STRONG&gt;Databricks SQL&lt;/STRONG&gt;&lt;/P&gt;&lt;/LI&gt;&lt;LI&gt;&lt;P&gt;&lt;STRONG&gt;Unity Catalog&lt;/STRONG&gt;&lt;/P&gt;&lt;/LI&gt;&lt;LI&gt;&lt;P&gt;&lt;STRONG&gt;Delta tables / views&lt;/STRONG&gt;&lt;/P&gt;&lt;/LI&gt;&lt;LI&gt;&lt;P&gt;&lt;STRONG&gt;AI/BI Genie&lt;/STRONG&gt;&lt;/P&gt;&lt;/LI&gt;&lt;LI&gt;&lt;P&gt;&lt;STRONG&gt;Databricks Apps with AppKit&lt;/STRONG&gt;&lt;/P&gt;&lt;/LI&gt;&lt;/UL&gt;&lt;P&gt;I used a synthetic operations dataset for this project so that the behavior of the application could be tested against known scenarios.&lt;/P&gt;&lt;P&gt;The dataset contains &lt;STRONG&gt;1,080 operational records covering 30 days&lt;/STRONG&gt;, with combinations of:&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;&lt;P&gt;4 sites&lt;/P&gt;&lt;/LI&gt;&lt;LI&gt;&lt;P&gt;3 vendors&lt;/P&gt;&lt;/LI&gt;&lt;LI&gt;&lt;P&gt;3 operational processes&lt;/P&gt;&lt;/LI&gt;&lt;/UL&gt;&lt;P&gt;The data includes measures such as inbound units, processed units, backlog, SLA-eligible units, SLA-met units, defects, labor hours, downtime, and processing time.&lt;/P&gt;&lt;P&gt;I deliberately introduced several patterns into the data, including an East Hub backlog increase, vendor SLA deterioration, a throughput improvement at another site, and a temporary defect-rate increase.&lt;/P&gt;&lt;P&gt;That made it possible to test whether Genie could actually identify and explain the patterns I knew were present.&lt;/P&gt;&lt;H2&gt;Building a governed KPI layer&lt;/H2&gt;&lt;P&gt;One thing I wanted to avoid was allowing every question to produce a slightly different KPI calculation.&lt;/P&gt;&lt;P&gt;For example, SLA attainment should not be calculated by simply averaging percentages from individual rows.&lt;/P&gt;&lt;P&gt;For aggregated analysis, I defined SLA attainment as:&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Total SLA-met units / Total SLA-eligible units&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;I used the same approach for other measures:&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Throughput&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;Total processed units / Total labor hours&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Defect rate&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;Total defect units / Total processed units&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Net backlog change&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;Sum of ending backlog minus starting backlog&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Processing gap&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;Inbound units minus processed units&lt;/P&gt;&lt;P&gt;These measures were defined explicitly for Genie so that the same business definitions would be used across different questions.&lt;/P&gt;&lt;P&gt;I also added descriptions and metadata to the Unity Catalog view so fields such as labor_hours, processing_gap_units, and backlog_change_units had clear business meaning.&lt;/P&gt;&lt;H2&gt;What users can ask&lt;/H2&gt;&lt;P&gt;A few examples that worked well during testing were:&lt;/P&gt;&lt;BLOCKQUOTE&gt;&lt;P&gt;What operational issues need attention in the latest available data?&lt;/P&gt;&lt;/BLOCKQUOTE&gt;&lt;BLOCKQUOTE&gt;&lt;P&gt;Why did East Hub backlog increase after August 18?&lt;/P&gt;&lt;/BLOCKQUOTE&gt;&lt;BLOCKQUOTE&gt;&lt;P&gt;Which vendor experienced the largest decline in SLA attainment after August 16?&lt;/P&gt;&lt;/BLOCKQUOTE&gt;&lt;BLOCKQUOTE&gt;&lt;P&gt;Where did defect rates increase abnormally?&lt;/P&gt;&lt;/BLOCKQUOTE&gt;&lt;BLOCKQUOTE&gt;&lt;P&gt;Is East Hub's backlog problem primarily caused by higher inbound volume or weaker processing performance?&lt;/P&gt;&lt;/BLOCKQUOTE&gt;&lt;P&gt;These questions demonstrate different parts of the workflow.&lt;/P&gt;&lt;P&gt;Some are broad exception-detection questions. Others are diagnostic questions that require Genie to compare periods, calculate weighted KPIs, or break performance down by site, vendor, and process.&lt;/P&gt;&lt;H2&gt;The most useful part of the project was actually the testing&lt;/H2&gt;&lt;P&gt;The part I learned the most from was not getting Genie to answer more questions.&lt;/P&gt;&lt;P&gt;It was learning how to improve the quality of the answers and prevent unsupported conclusions.&lt;/P&gt;&lt;H3&gt;1. Arbitrary KPI thresholds&lt;/H3&gt;&lt;P&gt;During an early test, I asked:&lt;/P&gt;&lt;BLOCKQUOTE&gt;&lt;P&gt;What operational issues need attention?&lt;/P&gt;&lt;/BLOCKQUOTE&gt;&lt;P&gt;The first approach introduced thresholds such as SLA below a certain percentage or downtime above a certain number of minutes.&lt;/P&gt;&lt;P&gt;The problem was that I had never defined those thresholds as business rules.&lt;/P&gt;&lt;P&gt;So I changed the analytical approach.&lt;/P&gt;&lt;P&gt;For broad operational questions, OpsPulse now compares the &lt;STRONG&gt;latest seven days against the preceding fourteen-day baseline&lt;/STRONG&gt; for the same site, vendor, and process.&lt;/P&gt;&lt;P&gt;This made the analysis much more useful because the application looks for deterioration relative to recent normal performance rather than inventing universal thresholds.&lt;/P&gt;&lt;H2&gt;2. Choosing the right baseline matters&lt;/H2&gt;&lt;P&gt;I also tested a latest-day comparison against the previous seven days.&lt;/P&gt;&lt;P&gt;That initially looked reasonable, but it created another issue.&lt;/P&gt;&lt;P&gt;If an operational problem had already existed for several days, the recent baseline itself contained part of the deterioration. That could make the latest day look relatively normal and hide the issue.&lt;/P&gt;&lt;P&gt;I changed the logic to compare the most recent seven-day period against the preceding fourteen days.&lt;/P&gt;&lt;P&gt;That simple change significantly improved the exception detection.&lt;/P&gt;&lt;P&gt;It was a good reminder that the quality of AI-assisted analytics still depends heavily on the analytical design behind it.&lt;/P&gt;&lt;H2&gt;3. Labor hours are not labor cost&lt;/H2&gt;&lt;P&gt;One of my favorite tests was:&lt;/P&gt;&lt;BLOCKQUOTE&gt;&lt;P&gt;What can you tell me about labor cost across the sites?&lt;/P&gt;&lt;/BLOCKQUOTE&gt;&lt;P&gt;The dataset contains labor_hours, but it does not contain hourly wages, labor rates, or any monetary cost field.&lt;/P&gt;&lt;P&gt;In an early test, the response started treating higher labor hours as an indication of higher labor cost.&lt;/P&gt;&lt;P&gt;That was not supported by the data.&lt;/P&gt;&lt;P&gt;I updated both the Genie instructions and the Unity Catalog metadata to make the distinction explicit:&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Labor hours measure operational effort. They are not a financial cost measure.&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;After the change, OpsPulse correctly responded that labor cost could not be determined from the available dataset.&lt;/P&gt;&lt;P&gt;That was an important result for me because a useful analytical agent should not only know how to calculate something. It should also know when the data does not support the requested conclusion.&lt;/P&gt;&lt;H2&gt;4. "Which site is performing best overall?"&lt;/H2&gt;&lt;P&gt;Another interesting test was:&lt;/P&gt;&lt;BLOCKQUOTE&gt;&lt;P&gt;Which site is performing best overall?&lt;/P&gt;&lt;/BLOCKQUOTE&gt;&lt;P&gt;Initially, Genie created an implicit ranking using SLA, throughput, defects, backlog, and downtime.&lt;/P&gt;&lt;P&gt;But I had never defined a composite performance score or assigned weights to those KPIs.&lt;/P&gt;&lt;P&gt;There is no analytically defensible reason to assume, for example, that SLA should automatically matter more than defect rate, or that throughput should receive a particular weight.&lt;/P&gt;&lt;P&gt;So I added a guardrail.&lt;/P&gt;&lt;P&gt;Without a governed scoring methodology, OpsPulse now compares the sites across the individual KPIs and explains that the "best" site depends on the business priority.&lt;/P&gt;&lt;P&gt;This was another small change, but an important one from a data governance perspective.&lt;/P&gt;&lt;H2&gt;Genie at the core of the application&lt;/H2&gt;&lt;P&gt;Genie is not an optional feature inside OpsPulse. It is the primary way the user interacts with the operational data.&lt;/P&gt;&lt;P&gt;The experience is designed around a sequence like this:&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;1. Detect&lt;/STRONG&gt;&lt;/P&gt;&lt;BLOCKQUOTE&gt;&lt;P&gt;What needs attention?&lt;/P&gt;&lt;/BLOCKQUOTE&gt;&lt;P&gt;&lt;STRONG&gt;2. Diagnose&lt;/STRONG&gt;&lt;/P&gt;&lt;BLOCKQUOTE&gt;&lt;P&gt;Why is East Hub backlog increasing?&lt;/P&gt;&lt;/BLOCKQUOTE&gt;&lt;P&gt;&lt;STRONG&gt;3. Investigate&lt;/STRONG&gt;&lt;/P&gt;&lt;BLOCKQUOTE&gt;&lt;P&gt;Is inbound volume increasing faster than processing capacity?&lt;/P&gt;&lt;/BLOCKQUOTE&gt;&lt;P&gt;The application uses Genie to translate those business questions into queries against the curated operational KPI layer and return the results in a form that an operations user can work with.&lt;/P&gt;&lt;P&gt;Without Genie, the user would lose the natural-language path from an operational question to the underlying data and follow-up analysis.&lt;/P&gt;&lt;H2&gt;Building the Databricks App&lt;/H2&gt;&lt;P&gt;I used the &lt;STRONG&gt;AppKit Genie template&lt;/STRONG&gt; in Databricks Apps and connected the application directly to the OpsPulse Genie Agent.&lt;/P&gt;&lt;P&gt;I customized the application around the three-stage workflow:&lt;/P&gt;&lt;H3&gt;Detect&lt;/H3&gt;&lt;P&gt;Surface emerging operational exceptions across sites, vendors, and processes.&lt;/P&gt;&lt;H3&gt;Diagnose&lt;/H3&gt;&lt;P&gt;Compare recent performance with historical baselines and identify the metrics contributing to deterioration.&lt;/P&gt;&lt;H3&gt;Investigate&lt;/H3&gt;&lt;P&gt;Ask follow-up questions in natural language and explore the supporting operational data with Genie.&lt;/P&gt;&lt;P&gt;I intentionally kept the application itself simple.&lt;/P&gt;&lt;P&gt;Most of the work went into the data model, KPI definitions, business metadata, query patterns, testing, and guardrails behind the interface.&lt;/P&gt;&lt;H2&gt;What I learned&lt;/H2&gt;&lt;P&gt;This project changed the way I think about conversational analytics.&lt;/P&gt;&lt;P&gt;Connecting an AI interface to data is relatively straightforward.&lt;/P&gt;&lt;P&gt;Making the answers consistently useful is much more interesting.&lt;/P&gt;&lt;P&gt;I found myself spending more time thinking about questions such as:&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;&lt;P&gt;What does "latest" actually mean?&lt;/P&gt;&lt;/LI&gt;&lt;LI&gt;&lt;P&gt;What comparison period is appropriate?&lt;/P&gt;&lt;/LI&gt;&lt;LI&gt;&lt;P&gt;Should this KPI be averaged or weighted?&lt;/P&gt;&lt;/LI&gt;&lt;LI&gt;&lt;P&gt;Is the system identifying a correlation or claiming a cause?&lt;/P&gt;&lt;/LI&gt;&lt;LI&gt;&lt;P&gt;Is a threshold actually defined by the business?&lt;/P&gt;&lt;/LI&gt;&lt;LI&gt;&lt;P&gt;Is the requested conclusion supported by the available fields?&lt;/P&gt;&lt;/LI&gt;&lt;LI&gt;&lt;P&gt;Should the system rank something when no scoring methodology exists?&lt;/P&gt;&lt;/LI&gt;&lt;/UL&gt;&lt;P&gt;Those questions are not really AI questions.&lt;/P&gt;&lt;P&gt;They are &lt;STRONG&gt;analytics, semantic modeling, business logic, and governance questions&lt;/STRONG&gt;.&lt;/P&gt;&lt;P&gt;And I think that is where tools like AI/BI Genie become particularly interesting for analysts.&lt;/P&gt;&lt;H2&gt;Final thoughts&lt;/H2&gt;&lt;P&gt;OpsPulse started as a fairly simple idea: let an operations manager ask questions about operational performance.&lt;/P&gt;&lt;P&gt;By the end of the project, the more interesting challenge became making those answers &lt;STRONG&gt;consistent, explainable, and grounded in the available data&lt;/STRONG&gt;.&lt;/P&gt;&lt;P&gt;There is still plenty I would extend in a production version, including governed business targets, more sophisticated anomaly detection, additional historical context, alerting, and integration with real operational datasets.&lt;/P&gt;&lt;P&gt;For this challenge, though, I wanted to focus on one thing:&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Can a user move naturally from noticing an operational signal to understanding what may be driving it, without losing the analytical discipline behind the answer?&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;OpsPulse is my attempt at that.&lt;/P&gt;&lt;H3&gt;Built with&lt;/H3&gt;&lt;P&gt;&lt;STRONG&gt;Databricks Free Edition | Databricks Apps | AI/BI Genie | Unity Catalog | Databricks SQL | Delta&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Demo:&lt;/STRONG&gt;&amp;nbsp;&lt;A title="OpsPulse_Databricks_Genie_Demo" href="https://drive.google.com/file/d/1MZR6scvsFNv7R8iCzHftYAO1V5KA9IIj/view?usp=sharing" target="_blank" rel="noopener"&gt;OpsPulse: From Operational Signals to Data-Backed Answers with AI/BI Genie&lt;/A&gt;&lt;/P&gt;&lt;P&gt;[&lt;A href="https://drive.google.com/file/d/1MZR6scvsFNv7R8iCzHftYAO1V5KA9IIj/view?usp=sharing" target="_blank" rel="noopener"&gt;https://drive.google.com/file/d/1MZR6scvsFNv7R8iCzHftYAO1V5KA9IIj/view?usp=sharing&lt;/A&gt;]&lt;/P&gt;</description>
      <pubDate>Tue, 01 Sep 2026 00:34:03 GMT</pubDate>
      <guid>https://community.databricks.com/t5/community-articles/opspulse-from-operational-signals-to-data-backed-answers-with/m-p/166979#M1500</guid>
      <dc:creator>kulkarnigauri3</dc:creator>
      <dc:date>2026-09-01T00:34:03Z</dc:date>
    </item>
    <item>
      <title>Genie SQL Quest: A Genie-Powered Arcade for Learning SQL</title>
      <link>https://community.databricks.com/t5/community-articles/genie-sql-quest-a-genie-powered-arcade-for-learning-sql/m-p/166969#M1499</link>
      <description>&lt;H3&gt;🧞 What if learning SQL felt like a game?&lt;/H3&gt;&lt;P&gt;SQL is one of the most valuable skills in data and AI, but most people still learn it from static tutorials or documentation. Chatbots can answer questions, but they rarely push learners to think through a problem, make mistakes, and try again. &lt;STRONG&gt;Genie SQL Quest&lt;/STRONG&gt; changes that by turning SQL practice into a game — and it uses &lt;STRONG&gt;Databricks Genie&lt;/STRONG&gt; as the intelligent coach behind every challenge.&lt;/P&gt;&lt;P&gt;I built this project for the &lt;STRONG&gt;Databricks Community Contest — Creative Thinking track&lt;/STRONG&gt; to show that Genie can be more than a conversational analytics assistant. In Genie SQL Quest, Genie becomes a real-time tutor: it gives hints when you are stuck, explains why your answer is right or wrong, and makes the learning experience feel natural and conversational. The goal is to make SQL practice feel less like a lecture and more like a guided puzzle.&lt;/P&gt;&lt;HR /&gt;&lt;H3&gt;&lt;span class="lia-unicode-emoji" title=":sparkles:"&gt;✨&lt;/span&gt; What Genie SQL Quest does&lt;/H3&gt;&lt;P&gt;Genie SQL Quest is a browser-based learning game where players solve natural-language SQL challenges presented as flashcards.&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;&lt;STRONG&gt;Curated challenge categories:&lt;/STRONG&gt; Players progress through Basics, Filtering, Aggregations, Sorting / Top-N, Joins, and Date logic. Each category is designed to build a specific SQL muscle.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Progressive difficulty:&lt;/STRONG&gt; Every category contains multiple difficulty levels. Players must solve easier cards to unlock harder ones, which keeps the pacing adaptive and rewarding.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Boss Rounds:&lt;/STRONG&gt; Once players master a category, they face a harder Boss Round that combines multiple concepts and tests deeper understanding.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;XP, streaks, and progress tracking:&lt;/STRONG&gt; Correct answers earn XP, streaks encourage daily practice, and the progress panel shows how far the player has come across every category.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Interactive schema panel:&lt;/STRONG&gt; Players can inspect tables, columns, primary/foreign keys, sample rows, and relationships at any time. This removes the friction of memorizing a schema and lets players focus on writing good SQL.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Instant feedback:&lt;/STRONG&gt; Every submission is checked against a gold answer in the question bank, so players immediately know if they are correct and see the canonical query.&lt;/LI&gt;&lt;/UL&gt;&lt;P&gt;While the game mechanics keep players engaged, the real intelligence comes from Genie.&lt;/P&gt;&lt;HR /&gt;&lt;H3&gt;🧠 How Genie is at the center&lt;/H3&gt;&lt;P&gt;Databricks Genie is not a side feature in this project — it is the core reason the app feels intelligent and conversational. The backend calls the &lt;STRONG&gt;Databricks Genie Conversational API&lt;/STRONG&gt; to bring the flashcards to life.&lt;/P&gt;&lt;P&gt;Capability How Genie powers the experience&lt;/P&gt;&lt;TABLE&gt;&lt;TBODY&gt;&lt;TR&gt;&lt;TD&gt;&lt;STRONG&gt;Hints on demand&lt;/STRONG&gt;&lt;/TD&gt;&lt;TD&gt;When a player is stuck, Genie reads the current challenge and produces a short, context-aware hint. The hint points the player toward the right SQL concept without giving away the answer.&lt;/TD&gt;&lt;/TR&gt;&lt;TR&gt;&lt;TD&gt;&lt;STRONG&gt;Personalized explanations&lt;/STRONG&gt;&lt;/TD&gt;&lt;TD&gt;After a player submits an answer, Genie explains the key idea behind the correct SQL in one friendly sentence. The explanation can be tailored to whether the player's answer was correct or incorrect.&lt;/TD&gt;&lt;/TR&gt;&lt;TR&gt;&lt;TD&gt;&lt;STRONG&gt;Boss-round narration&lt;/STRONG&gt;&lt;/TD&gt;&lt;TD&gt;Genie can introduce or comment on Boss Round challenges, adding a narrative layer to the hardest part of the game.&lt;/TD&gt;&lt;/TR&gt;&lt;TR&gt;&lt;TD&gt;&lt;STRONG&gt;Graceful fallback&lt;/STRONG&gt;&lt;/TD&gt;&lt;TD&gt;If Genie is not configured, the app falls back to static hints and explanations stored in the question bank, so the game is never broken.&lt;/TD&gt;&lt;/TR&gt;&lt;/TBODY&gt;&lt;/TABLE&gt;&lt;P&gt;The integration is implemented through a lightweight Python client that streams responses from the Genie API via Server-Sent Events and extracts the assistant's message. This makes every Genie interaction fast, simple, and focused on the current challenge.&lt;/P&gt;&lt;HR /&gt;&lt;H3&gt;&lt;span class="lia-unicode-emoji" title=":building_construction:"&gt;🏗&lt;/span&gt;️ Architecture&lt;/H3&gt;&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="AI Vision Framework for Gap-2026-08-31-220906.png" style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30484i2148E2B673A2DAFB/image-size/large?v=v2&amp;amp;px=999" role="button" title="AI Vision Framework for Gap-2026-08-31-220906.png" alt="AI Vision Framework for Gap-2026-08-31-220906.png" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;The application is split into three layers:&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;1. React + Vite Frontend&lt;/STRONG&gt;&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;Category panel for picking challenge tracks and difficulty levels.&lt;/LI&gt;&lt;LI&gt;Flashcard play area with the prompt, SQL input, and submission flow.&lt;/LI&gt;&lt;LI&gt;Hint buttons that request Genie-powered guidance.&lt;/LI&gt;&lt;LI&gt;Progress panel showing XP, streaks, cards solved, and category completion.&lt;/LI&gt;&lt;LI&gt;Schema panel showing table definitions, columns, relationships, and sample rows.&lt;/LI&gt;&lt;/UL&gt;&lt;P&gt;&lt;STRONG&gt;2. FastAPI Backend&lt;/STRONG&gt;&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;Serves the game API endpoints: drawing cards, submitting answers, requesting hints, and resetting progress.&lt;/LI&gt;&lt;LI&gt;Loads the question bank and schema from JSON files.&lt;/LI&gt;&lt;LI&gt;Evaluates player answers against canonical SQL.&lt;/LI&gt;&lt;LI&gt;Orchestrates calls to Databricks Genie for hints and explanations.&lt;/LI&gt;&lt;/UL&gt;&lt;P&gt;&lt;STRONG&gt;3. Databricks Genie&lt;/STRONG&gt;&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;Provides natural-language hints when requested.&lt;/LI&gt;&lt;LI&gt;Explains SQL answers based on the player's attempt and the challenge prompt.&lt;/LI&gt;&lt;LI&gt;Acts as the conversational AI tutor that makes the app feel alive.&lt;/LI&gt;&lt;/UL&gt;&lt;P&gt;&lt;STRONG&gt;Key design decisions&lt;/STRONG&gt;&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;&lt;STRONG&gt;Game-first, Genie-enhanced:&lt;/STRONG&gt; The question bank provides deterministic scoring so feedback is instant and fair. Genie adds the conversational layer on top without slowing down challenge.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;No ETL pipeline needed:&lt;/STRONG&gt; The dataset is a curated JSON question bank, so the project is easy to run, extend, and deploy.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Resilient to missing credentials:&lt;/STRONG&gt; If Genie is not configured, the app automatically uses static hints and explanations so development and demos are never blocked.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Databricks Apps ready:&lt;/STRONG&gt; The app is packaged with an app.yaml and a single-entry run.py for straightforward deployment inside Databricks.&lt;/LI&gt;&lt;/UL&gt;&lt;HR /&gt;&lt;H3&gt;&lt;span class="lia-unicode-emoji" title=":hammer_and_wrench:"&gt;🛠&lt;/span&gt;️ Tech stack&lt;/H3&gt;&lt;UL&gt;&lt;LI&gt;&lt;STRONG&gt;Frontend:&lt;/STRONG&gt; React, Vite, Tailwind CSS, Lucide icons&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Backend:&lt;/STRONG&gt; FastAPI (Python)&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;AI/BI layer:&lt;/STRONG&gt; Databricks Genie Conversational API&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Optional fallback:&lt;/STRONG&gt; Databricks Model Serving endpoint (databricks-dbrx-instruct)&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Deployment:&lt;/STRONG&gt; Databricks Apps&lt;/LI&gt;&lt;/UL&gt;&lt;HR /&gt;&lt;H3&gt;&lt;span class="lia-unicode-emoji" title=":direct_hit:"&gt;🎯&lt;/span&gt; Why Genie makes this work&lt;/H3&gt;&lt;P&gt;Genie is the difference between a static SQL quiz and an adaptive learning experience. Without Genie, the app would only show pre-written hints. With Genie, the app becomes a tutor that understands the current challenge and the player's actual answer.&lt;/P&gt;&lt;P&gt;The three things that make Genie essential here are:&lt;/P&gt;&lt;OL&gt;&lt;LI&gt;&lt;STRONG&gt;Context-aware guidance:&lt;/STRONG&gt; Genie reads the exact challenge the player is looking at, so every hint is relevant and timely.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Personalized feedback:&lt;/STRONG&gt; Genie can react to whether the player's answer was correct or incorrect, making the explanation more useful.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Conversational learning:&lt;/STRONG&gt; Because Genie responds in natural language, players feel like they are learning from a coach rather than grading against an answer key.&lt;/LI&gt;&lt;/OL&gt;&lt;P&gt;That combination — structured, level-based challenges plus Genie's conversational intelligence — is what turns SQL practice into something players actually want to come back to.&lt;/P&gt;&lt;P&gt;&lt;EM&gt;Built for the Databricks Community Contest. Feedback, forks, and questions welcome!&lt;/EM&gt;&lt;/P&gt;</description>
      <pubDate>Mon, 31 Aug 2026 22:31:07 GMT</pubDate>
      <guid>https://community.databricks.com/t5/community-articles/genie-sql-quest-a-genie-powered-arcade-for-learning-sql/m-p/166969#M1499</guid>
      <dc:creator>Chiku07</dc:creator>
      <dc:date>2026-08-31T22:31:07Z</dc:date>
    </item>
    <item>
      <title>LakeOps - A Lakehouse Operation Observability App</title>
      <link>https://community.databricks.com/t5/community-articles/lakeops-a-lakehouse-operation-observability-app/m-p/166963#M1498</link>
      <description>&lt;P&gt;&lt;STRONG&gt;The Problem LakeOps Solves&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;Data teams focused on cloud-based analytics solutions often experience operational bottlenecks. Engineers spend countless hours writing SQL joins across system logs monitoring and reporting pipeline failures, while engineering managers struggle to get visible insights into compute costs and SLA performance. Some teams solve this through dashboards with static KPI reporting, but this often limits how much insight could be drawn from the data. LakeOps addresses this by providing an intelligent, natural-language interface for Lakehouse operation observability.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Target Audience and Dual-Persona Design&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;LakeOps is designed with a context-aware toggle to serve two distinct operational roles:&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;&lt;STRONG&gt;Data Engineers:&lt;/STRONG&gt; Can ask highly technical debugging questions to retrieve specific error traces, execution times, and job logs for root-cause analysis.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Engineering Managers:&lt;/STRONG&gt; Can request high-level summaries regarding DBU consumption, overall pipeline costs, and business impact.&lt;/LI&gt;&lt;/UL&gt;&lt;P&gt;&lt;STRONG&gt;&lt;SPAN&gt;Application architecture and data flow&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;LakeOps utilizes a secure architecture deployed directly within the Databricks Data Intelligence Platform. The end-to-end data flow is visualized below:&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="Screenshot 2026-08-31 215539.png" style="width: 400px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30481i314547D4B70096BB/image-size/medium?v=v2&amp;amp;px=400" role="button" title="Screenshot 2026-08-31 215539.png" alt="Screenshot 2026-08-31 215539.png" /&gt;&lt;/span&gt;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;Data Layer:&lt;/STRONG&gt;&lt;SPAN&gt; Following the Medallion architecture, the pipeline processes data through three layers, all stored in Delta Lake and strictly governed by Unity Catalog:&lt;/SPAN&gt;&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;Bronze: Raw logs and system tables are used as the bronze layer. Since this project was developed using the Databricks Free Edition workspace, synthetic data was ingested into the tables in this layer.&lt;/LI&gt;&lt;LI&gt;Silver: Transforms and cleans the raw data from the bronze layer into parsed pipeline execution records and technical error traces as required.&lt;/LI&gt;&lt;LI&gt;Gold: Aggregates the silver data into final business-level metrics, such as total DBU consumption, costs, and SLA performance.&lt;/LI&gt;&lt;/UL&gt;&lt;P&gt;&lt;STRONG&gt;&lt;SPAN&gt;Authentication Flow:&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt; The Streamlit application is hosted as a serverless, containerized Databricks App. It utilizes On-Behalf-Of (OBO) authentication, meaning the app's Service Principal is entirely locked out of the data. Instead, it uses the logged-in user's OAuth token, ensuring all queries execute securely under their exact Unity Catalog privileges.&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;&lt;SPAN&gt;Logic &amp;amp; Processing with Genie Agent:&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt; When a user submits a prompt, the app injects the selected persona context (Data Engineer or Engineering Manager). The Databricks Genie Agent processes this combined context, acting as the semantic engine to translate the natural language request into optimized SQL against the permitted tables.&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;&lt;SPAN&gt;Presentation Layer:&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt; The app's custom logic intercepts the returned SQL and data arrays from the Databricks SDK. It then intelligently routes the output to render dynamic, interactive Plotly tabs (Bar, Line, Pie) or raw DataFrames based on the structure of the returned data.&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;What can users ask the Genie Agent&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;Since LakeOps is context-aware, queries are inherently tailored to the user's operational role. The application intelligently guides the agent to fetch the appropriate level of detail:&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;Engineering Managers can ask business-level, aggregate questions such as:&lt;/LI&gt;&lt;UL&gt;&lt;LI&gt;"What was our total DBU spend by pipeline this month?"&lt;/LI&gt;&lt;LI&gt;"Which pipeline experienced the most SLA breaches last week, and what was the associated compute cost?"&lt;/LI&gt;&lt;/UL&gt;&lt;LI&gt;Data Engineers can dive directly into technical, root-cause investigations by asking:&lt;/LI&gt;&lt;UL&gt;&lt;LI&gt;"What is the success rate for all pipelines in the last month?"&lt;/LI&gt;&lt;LI&gt;"What is the average execution latency for the Gold aggregation job over the past seven days?"&lt;/LI&gt;&lt;/UL&gt;&lt;/UL&gt;&lt;P&gt;&lt;STRONG&gt;How Genie powers the app's main experience&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;When a user submits a question, the Streamlit app concatenates the persona-specific instructions and forwards the prompt to Genie. Genie leverages its deep understanding of the underlying table metadata and Unity Catalog relationships to generate optimized SQL queries on the fly. LakeOps then extracts this SQL from the Genie API payload, executes it synchronously against a serverless SQL Warehouse, and uses a custom Python heuristic to dynamically render the results as interactive Plotly charts or raw data tables. Without Genie's real-time text-to-SQL translation and semantic awareness, this dynamic, self-serve observability experience would be impossible.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;What I learned while building and testing the app&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;Building and deploying LakeOps end-to-end within the Databricks Data Intelligence Platform provided several key engineering takeaways:&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;&lt;STRONG&gt;Curating Genie Data Models for Accuracy:&lt;/STRONG&gt; Grounding Genie in accurate metadata, logical table joins, and clear instructions is essential to prevent hallucinations. Curating focused tables across the Medallion layers ensured Genie consistently chose the correct tables depending on the context of the incoming question.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;&lt;SPAN&gt;Benchmarking Genie Agent Queries:&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt; Conducted query benchmarking to establish baseline performance and response accuracy. This testing phase was vital to verify that Genie consistently translated natural language into the most optimized SQL, and to ensure acceptable latency when querying across the varying complexities of the Medallion layers.&lt;/SPAN&gt;&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Mastering On-Behalf-Of (OBO) Authentication:&lt;/STRONG&gt; Implementing user authorization via the OBO model proved critical for enterprise security. It taught the importance of properly configuring OAuth scopes (genie, workspace, SQL) so the app can securely execute queries under the end-user's identity without relying on overly privileged Service Principals.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Building Resilient API Parsers:&lt;/STRONG&gt; Handling SDK payload structures required robust dictionary-based parsing to extract text attachments and generated SQL statements dynamically across varying API responses.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;End-to-End Serverless Integration:&lt;/STRONG&gt; Deploying a containerized Streamlit Databricks App directly where the data lives highlighted the efficiency of the serverless compute plane. Coordinating the workflow—from the UI, through the Genie Agent, across the serverless SQL Warehouse, and back to dynamic Plotly visualizations—demonstrated how seamlessly Databricks components integrate into a unified operational tool.&lt;/LI&gt;&lt;/UL&gt;&lt;P&gt;&lt;STRONG&gt;Final Thoughts and Future Roadmap&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;While LakeOps successfully demonstrates the power of natural-language observability, the current application is a prototype. Time constraints limited the initial scope, but the following enhancements would significantly elevate both the app's capability and the overall user experience:&lt;/P&gt;&lt;OL&gt;&lt;LI&gt;&lt;STRONG&gt;General User Experience Improvements&lt;/STRONG&gt;&lt;/LI&gt;&lt;/OL&gt;&lt;UL&gt;&lt;LI&gt;&lt;STRONG&gt;&lt;SPAN&gt;Automated Data Refresh Pipelines:&lt;/SPAN&gt;&lt;/STRONG&gt;&lt;SPAN&gt; Implementing a scheduled daily pipeline orchestration job to keep data continuously flowing from the Bronze to Gold layers, ensuring that LakeOps always surfaces fresh operational metrics from the previous workday.&lt;/SPAN&gt;&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Query Caching:&lt;/STRONG&gt; Implementing caching mechanisms (such as Streamlit's &lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/12028"&gt;@ST&lt;/a&gt;.cache_data) for repetitive, high-level SLA queries to ensure instant load times and minimize SQL Warehouse compute costs.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Proactive Daily Health Checks:&lt;/STRONG&gt; Adding automated routines that execute critical SLA and cost queries upon launch, presenting users with a summarized morning health briefing before they even type a prompt.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Follow-Up Question Suggestions:&lt;/STRONG&gt; Utilizing lightweight LLM calls to generate contextual, clickable follow-up questions beneath rendered charts, driving continuous data exploration.&lt;/LI&gt;&lt;/UL&gt;&lt;OL&gt;&lt;LI&gt;&lt;STRONG&gt;RAG Integration for Autonomous Remediation&lt;/STRONG&gt; To transition LakeOps from a purely diagnostic tool into an active remediation assistant, the architecture could be expanded with Retrieval-Augmented Generation (RAG):&lt;/LI&gt;&lt;/OL&gt;&lt;UL&gt;&lt;LI&gt;&lt;STRONG&gt;Cross-Platform Documentation Search:&lt;/STRONG&gt; Indexing official documentation for the core technologies in the team's data stack—such as Databricks and Azure Data Factory. If an engineer encounters an obscure pipeline exception, the RAG model could surface the exact syntax fix or known workaround without requiring them to leave the app.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Historical Long-Term Memory:&lt;/STRONG&gt; Ingesting past Jira tickets, Slack support threads, and incident post-mortems into a Databricks Vector Search index. This would grant the application long-term memory, instantly providing engineers with historical context on recurring pipeline failures and their exact resolution paths.&lt;/LI&gt;&lt;/UL&gt;&lt;P&gt;This development was done using Google Gemini and Genie Code as coding assistant.&lt;/P&gt;&lt;P&gt;#Genie-PoweredAppChallenge&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Mon, 31 Aug 2026 20:38:05 GMT</pubDate>
      <guid>https://community.databricks.com/t5/community-articles/lakeops-a-lakehouse-operation-observability-app/m-p/166963#M1498</guid>
      <dc:creator>ad3ad3</dc:creator>
      <dc:date>2026-08-31T20:38:05Z</dc:date>
    </item>
  </channel>
</rss>

