<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>rss.livelink.threads-in-node</title>
    <link>https://community.databricks.com/</link>
    <description>Databricks Community</description>
    <pubDate>Wed, 02 Sep 2026 16:58:30 GMT</pubDate>
    <dc:creator>Community</dc:creator>
    <dc:date>2026-09-02T16:58:30Z</dc:date>
    <item>
      <title>Announcement | Get Started with Genie One: Top AI Cowork Use Cases for Business Users</title>
      <link>https://community.databricks.com/t5/announcements/announcement-get-started-with-genie-one-top-ai-cowork-use-cases/m-p/167322#M1035</link>
      <description>&lt;P&gt;&lt;SPAN&gt;Genie One helps business teams turn everyday, repetitive work into repeatable workflows using governed data, connected tools, and natural language.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;FONT size="4"&gt;&lt;STRONG&gt;Key highlights&lt;/STRONG&gt;&lt;/FONT&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Automate business reviews&lt;/STRONG&gt;&lt;SPAN&gt;: Use a standard format, pull in current data, add commentary, and schedule recurring reports.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Simplify meeting prep and follow-up&lt;/STRONG&gt;&lt;SPAN&gt;: Create meeting briefs, summarize decisions and action items, and organize follow-up work.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Keep documents up to date&lt;/STRONG&gt;&lt;SPAN&gt;: Draft and refresh policies, playbooks, FAQs, and other documents using approved templates and business data.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Monitor important signals&lt;/STRONG&gt;&lt;SPAN&gt;: Track metrics such as revenue, ticket volume, job failures, or inventory and receive alerts when something changes.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Save workflows as skills&lt;/STRONG&gt;&lt;SPAN&gt;: Turn useful instructions into reusable skills that teams can run again or schedule for regular tasks.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;SPAN&gt;The best place to start is with one recurring task that takes time every week. Connect the right data sources, give Genie One clear instructions, review the first few results, and refine the workflow as needed.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P class="p8i6j01 paragraph"&gt;&lt;A style="background-color: #ff3621; color: white; padding: 10px 20px; text-decoration: none; border-radius: 5px; font-weight: bold; display: inline-block;" href="https://www.databricks.com/blog/get-started-genie-one-top-ai-cowork-use-cases-business-users?utm_source=bambu&amp;amp;utm_medium=social&amp;amp;utm_campaign=advocacy" target="_blank" rel="noopener"&gt; &lt;span class="lia-unicode-emoji" title=":backhand_index_pointing_right:"&gt;👉&lt;/span&gt; Read the full post here &lt;span class="lia-unicode-emoji" title=":backhand_index_pointing_left:"&gt;👈&lt;/span&gt;&lt;/A&gt;&lt;/P&gt;</description>
      <pubDate>Wed, 02 Sep 2026 15:25:57 GMT</pubDate>
      <guid>https://community.databricks.com/t5/announcements/announcement-get-started-with-genie-one-top-ai-cowork-use-cases/m-p/167322#M1035</guid>
      <dc:creator>Tushar_Parekar</dc:creator>
      <dc:date>2026-09-02T15:25:57Z</dc:date>
    </item>
    <item>
      <title>Havent received project expert badge</title>
      <link>https://community.databricks.com/t5/certifications/havent-received-project-expert-badge/m-p/167319#M4893</link>
      <description>&lt;P&gt;Hi Team,&lt;/P&gt;&lt;P&gt;I wanted to let you know that I have completed the Databricks Data Engineer Professional certification recently.&amp;nbsp;&lt;/P&gt;&lt;P&gt;I have received the &lt;STRONG&gt;Delivery Expert badge&lt;/STRONG&gt; after completing the requirements few days back. However, it has now been more than &lt;STRONG&gt;48 hours since I completed the exam&lt;/STRONG&gt;, and I have not yet received the &lt;STRONG&gt;Project Expert badge&lt;/STRONG&gt;.&lt;/P&gt;&lt;P&gt;Could you please check the status of the Project Expert badge and let me know if any additional action is required from my side?&lt;/P&gt;&lt;P&gt;Thank you for your help.&lt;/P&gt;&lt;P&gt;Best regards,&lt;BR /&gt;Prathik&lt;/P&gt;</description>
      <pubDate>Wed, 02 Sep 2026 14:46:17 GMT</pubDate>
      <guid>https://community.databricks.com/t5/certifications/havent-received-project-expert-badge/m-p/167319#M4893</guid>
      <dc:creator>prathikkinim</dc:creator>
      <dc:date>2026-09-02T14:46:17Z</dc:date>
    </item>
    <item>
      <title>After Databricks Summit 2026, I Feel Data Engineering Is Entering a New Phase</title>
      <link>https://community.databricks.com/t5/genie-hub/after-databricks-summit-2026-i-feel-data-engineering-is-entering/m-p/167318#M45</link>
      <description>&lt;P&gt;Dear Databricks Community, After coming back from Summit 2026, one thought stayed in my mind.&lt;/P&gt;&lt;P&gt;Data engineering is changing again. Earlier, most of our work was around pipelines, tables, jobs, transformations, reports, and dashboards. All of that is still important. But now, data is moving much closer to decisions, applications, AI, and real-time actions.&lt;/P&gt;&lt;P&gt;A few topics stood out to me strongly: LTAP, Lakebase, Genie, and Lakehouse RT.&lt;/P&gt;&lt;P&gt;LTAP made me think about how analytics and transaction processing are coming closer. For a long time, applications handled transactions and analytics systems handled reporting. Then we moved data between them using batch, streaming, CDC, or ETL pipelines.&lt;/P&gt;&lt;P&gt;But today, users want fresh insights. Applications need faster intelligence. AI systems need current context. So the gap between operational data and analytical data needs to become smaller.&lt;/P&gt;&lt;P&gt;Lakebase also feels like an important step. To me, it is not just another database. It shows how database architecture is evolving toward the lakehouse. If operational data can work more closely with trusted lakehouse data, it can reduce duplication, simplify architecture, and open new patterns for Data + AI applications.&lt;/P&gt;&lt;P&gt;Genie feels very practical too. Dashboards are useful, but they usually answer only the questions we already planned for. Business users always have follow-up questions. Why did this number change? Which team, region, or process is driving it? What happened this week?&lt;/P&gt;&lt;P&gt;Natural language analytics can make this easier. But one thing is clear: AI can only answer well when the data foundation is strong. Clean data, trusted metrics, governance, permissions, and context still matter a lot.&lt;/P&gt;&lt;P&gt;Lakehouse RT is another exciting direction. Earlier, real-time data was needed only for special use cases. Now, many business problems need faster updates. Fraud, inventory, customer experience, monitoring, recommendations, operations, and AI agents all need fresh data.&lt;/P&gt;&lt;P&gt;The real question is not only, “Can we process the data?”&lt;/P&gt;&lt;P&gt;It is, “Can we process it while it is still useful?”&lt;/P&gt;&lt;P&gt;My biggest takeaway from Summit 2026 is simple.&lt;/P&gt;&lt;P&gt;The future data platform is becoming more connected. Transactions, analytics, AI, applications, and real-time data are coming closer.&lt;/P&gt;&lt;P&gt;This means data engineers need to think beyond pipelines and reports. We need to think about the full solution.&lt;/P&gt;&lt;P&gt;What problem are we solving?&lt;BR /&gt;Who will use this data?&lt;BR /&gt;How fast do they need the answer?&lt;BR /&gt;Can AI make this workflow easier?&lt;BR /&gt;Can real-time data improve the decision?&lt;/P&gt;&lt;P&gt;For me, this is the exciting part of working with Databricks. A small idea can become a POC. A POC can become a dashboard. A dashboard can become a real-time insight. A real-time insight can become an action.&lt;/P&gt;&lt;P&gt;That is how I see the next phase of data engineering. It is not only about moving data faster. It is about helping people make better decisions sooner.&lt;/P&gt;&lt;P&gt;After Summit 2026, which area are you most excited to explore with Databricks: LTAP, Lakebase, Genie, Lakehouse RT, or another Data + AI use case?&lt;/P&gt;</description>
      <pubDate>Wed, 02 Sep 2026 14:42:10 GMT</pubDate>
      <guid>https://community.databricks.com/t5/genie-hub/after-databricks-summit-2026-i-feel-data-engineering-is-entering/m-p/167318#M45</guid>
      <dc:creator>Brahmareddy</dc:creator>
      <dc:date>2026-09-02T14:42:10Z</dc:date>
    </item>
    <item>
      <title>Thoughts on Using Remix for Data-Focused Applications</title>
      <link>https://community.databricks.com/t5/data-engineering/thoughts-on-using-remix-for-data-focused-applications/m-p/167306#M55673</link>
      <description>&lt;P class=""&gt;&lt;SPAN&gt;I’ve been exploring&amp;nbsp;&lt;/SPAN&gt;&lt;SPAN&gt;different approaches for building web applications that work with data-heavy workflows, and I recently came across Remix as an interesting option.&lt;/SPAN&gt;&lt;/P&gt;&lt;P class=""&gt;&lt;SPAN&gt;What I like about Remix is its focus on server-side data loading, forms, and handling application state without requiring everything to be managed on the client side. It seems like this approach could be useful for applications where users frequently interact with data.&lt;/SPAN&gt;&lt;/P&gt;&lt;P class=""&gt;&lt;SPAN&gt;I’m curious about how others approach this kind of setup.&lt;/SPAN&gt;&lt;/P&gt;&lt;P class=""&gt;&lt;SPAN&gt;For those who have worked with Remix or similar frameworks:&lt;/SPAN&gt;&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;&lt;SPAN&gt;How has your experience been with data fetching and server-side logic?&lt;/SPAN&gt;&lt;/LI&gt;&lt;LI&gt;&lt;SPAN&gt;Do you find the approach easier to maintain as an application grows?&lt;/SPAN&gt;&lt;/LI&gt;&lt;LI&gt;&lt;SPAN&gt;Are there any challenges you’ve faced when connecting a Remix application with data platforms or APIs?&lt;/SPAN&gt;&lt;/LI&gt;&lt;/UL&gt;&lt;P&gt;&lt;SPAN&gt;Would be interested to hear different experiences and opinions&lt;/SPAN&gt;&lt;/P&gt;</description>
      <pubDate>Wed, 02 Sep 2026 13:24:18 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/thoughts-on-using-remix-for-data-focused-applications/m-p/167306#M55673</guid>
      <dc:creator>ThiamLee</dc:creator>
      <dc:date>2026-09-02T13:24:18Z</dc:date>
    </item>
    <item>
      <title>Data Engineering 3.0 - From Trusted Data Products to Context Layers</title>
      <link>https://community.databricks.com/t5/community-articles/data-engineering-3-0-from-trusted-data-products-to-context/m-p/167304#M1525</link>
      <description>&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="1_3Mo6NkMHbxBHodt-3yUskw" style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30579iC6CABCF312EF35F8/image-size/large?v=v2&amp;amp;px=999" role="button" title="1_3Mo6NkMHbxBHodt-3yUskw" alt="1_3Mo6NkMHbxBHodt-3yUskw" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;Data Engineering &lt;STRONG&gt;1.0&lt;/STRONG&gt; was largely about integrating, moving, consolidating massive data to support enterprise decisions. Enterprises built ETL pipelines, data warehouses, data lakes and orchestration platforms to bring data together reliably.&amp;nbsp;&lt;/SPAN&gt;Data Engineering &lt;STRONG&gt;2.0&lt;/STRONG&gt; was about making the move to Cloud &lt;STRONG&gt;Lakehouse&lt;/STRONG&gt; and making the data trustworthy and reusable. This created a much more mature enterprise data estate. But the arrival of generative and agentic AI changes the question again. An enterprise may now have thousands of tables, hundreds of data products, semantic models, business glossaries, APIs, documents and real-time data.&amp;nbsp;An agent can potentially access much of this information. But access does not mean understanding. An agent investigating a business problem needs to know which information is relevant, how different entities are related, which definitions are authoritative, which rules apply, what happened previously, what is happening now, and what evidence supports a conclusion.&lt;/P&gt;&lt;P class=""&gt;Enterprises may already have all of this information. The problem is that it is distributed across the data estate. This is the &lt;STRONG&gt;context gap  -  &lt;/STRONG&gt;the distance between information that exists in the enterprise and the context an AI system needs to reliably perform a task.&lt;/P&gt;&lt;H3&gt;&lt;U&gt;&lt;STRONG&gt;The Agentic&amp;nbsp;Shift&lt;/STRONG&gt;&lt;/U&gt;&lt;/H3&gt;&lt;P class=""&gt;Traditional analytics was designed primarily to help users make vital decisions. A dashboard presents information. A business user interprets it. The user makes the decision. Machine learning automated parts of this process by using data to generate predictions or classifications.&amp;nbsp;&lt;STRONG&gt;Agentic AI&lt;/STRONG&gt; introduces a different operating model. An agent can retrieve information, reason over it, use tools, interact with systems and potentially take action. This means the data platform is no longer serving only a human consumer or an analytical model. It is increasingly serving an intelligent system that needs to understand the enterprise before it can act within it. This changes the role of data engineering.&lt;/P&gt;&lt;P class=""&gt;The question is no longer only &lt;STRONG&gt;How do we make trusted data available?&amp;nbsp;&lt;/STRONG&gt;It becomes &lt;STRONG&gt;How do we make trusted enterprise information understandable and usable by an AI system for a specific decision?&amp;nbsp;&lt;/STRONG&gt;That is where I see the emergence of &lt;STRONG&gt;Data Engineering 3.0&lt;/STRONG&gt;.&lt;/P&gt;&lt;H3&gt;&lt;U&gt;Data Engineering 3.0&lt;/U&gt;&lt;/H3&gt;&lt;P class=""&gt;Data Engineering 3.0 is the next evolution of the data engineering discipline. It is not a replacement for Data Engineering 2.0. It builds on it. The important shift is the unit of value. In Data Engineering 1.0, the question was whether data could be moved reliably at scale, processed and consumed in a robust manner. In Data Engineering 2.0, the question became whether data could be trusted and reused in accordance with governance and usage at scale via data products. In Data Engineering 3.0, the question becomes whether an intelligent system has the right context to perform a task reliably.&amp;nbsp;We are moving from engineering data for consumption to increasingly engineering the &lt;STRONG&gt;information environment in which AI operates&lt;/STRONG&gt;.&lt;/P&gt;&lt;H3&gt;&lt;U&gt;A Trusted Data Product is not the same as&amp;nbsp;Context&lt;/U&gt;&lt;/H3&gt;&lt;P class=""&gt;Consider a bank that has already built well-governed data products.&amp;nbsp;It may have a Customer data product containing customer information. It may have a Transaction data product containing transactions. It may have a Merchant data product. It may have a Device data product. It may have a Fraud data product containing historical fraud cases. Each of these can be high-quality, governed and independently reusable.&lt;/P&gt;&lt;P class=""&gt;Now imagine an agent is asked&amp;nbsp;&lt;STRONG&gt;Investigate this transaction and determine whether it should be escalated for fraud investigation.&amp;nbsp;&lt;/STRONG&gt;The agent does not simply need the Transaction data product. It may need to understand the customer associated with the transaction, the account involved, the merchant, the device used, recent transaction behaviour, previous fraud cases, current fraud indicators and the policies that determine when a transaction should be escalated. The data already exists. But the decision requires those pieces of information to be brought together with their meaning and relationships intact. That assembled information is the &lt;STRONG&gt;context for the decision&lt;/STRONG&gt;. &lt;STRONG&gt;A Data Product provides trusted information. A Context Layer makes trusted information relevant to a particular task.&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="1_hWW5XiezH9zkHSmdOmFHjA" style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30580i7196F753A5D38503/image-size/large?v=v2&amp;amp;px=999" role="button" title="1_hWW5XiezH9zkHSmdOmFHjA" alt="1_hWW5XiezH9zkHSmdOmFHjA" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;P&gt;&lt;FONT size="4"&gt;&lt;STRONG&gt;&lt;U&gt;The Context&amp;nbsp;Layer&lt;/U&gt;&lt;/STRONG&gt;&lt;/FONT&gt;&lt;/P&gt;&lt;P class=""&gt;An AI Context Layer is a decision-oriented capability that brings together the information an AI system needs to perform a specific task. It can include structured data, semantic definitions, relationships between entities, historical information, real-time signals, business policies and supporting evidence. It does not necessarily mean creating another physical copy of the enterprise data. In many cases, it can be a governed layer that composes information from existing data products and other enterprise systems when an agent needs it.&lt;/P&gt;&lt;P class=""&gt;For the fraud investigation example, the context will include&amp;nbsp;&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;Customer information and history&lt;/LI&gt;&lt;LI&gt;The transactions and accounts involved&lt;/LI&gt;&lt;LI&gt;Merchant characteristics&lt;/LI&gt;&lt;LI&gt;Device relationships&lt;/LI&gt;&lt;LI&gt;Recent transaction patterns&lt;/LI&gt;&lt;LI&gt;Previous fraud cases&lt;/LI&gt;&lt;LI&gt;Current fraud signals&lt;/LI&gt;&lt;LI&gt;Applicable fraud policies&lt;/LI&gt;&lt;LI&gt;Evidence supporting the recommendation&lt;/LI&gt;&lt;/UL&gt;&lt;P&gt;The important point is that the agent does not need the entire enterprise. It needs the &lt;STRONG&gt;right context for the decision it is making&lt;/STRONG&gt;.&lt;/P&gt;&lt;P class=""&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="1_CUWN7KstAGcXAcnxfeMy4Q" style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30581i8CD19118191A76E6/image-size/large?v=v2&amp;amp;px=999" role="button" title="1_CUWN7KstAGcXAcnxfeMy4Q" alt="1_CUWN7KstAGcXAcnxfeMy4Q" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;H3&gt;&lt;U&gt;Context is more than Retrieval&lt;/U&gt;&lt;/H3&gt;&lt;P class=""&gt;The first instinct when discussing AI context is often to think about &lt;STRONG&gt;RAG&lt;/STRONG&gt;. RAG is an important capability. It allows an AI system to retrieve relevant information before generating a response. But enterprise decision-making requires more than retrieving documents. An agent may need to understand that a particular transaction belongs to an account, that the account belongs to a customer, that the transaction was performed using a device associated with several accounts, and that the merchant has a particular risk classification. That is not simply a document retrieval problem. It is a problem involving &lt;STRONG&gt;entities, relationships, semantics, history, current state and rules&lt;/STRONG&gt;.&lt;/P&gt;&lt;P class=""&gt;Context may come from structured data, documents, knowledge graph, APIs or operational systems or real-time events. The Context Layer is therefore not a specific technology. It is an architectural capability for composing the information required by an intelligent system.&lt;/P&gt;&lt;H3&gt;&lt;U&gt;Where Ontology&amp;nbsp;fits&lt;/U&gt;&lt;/H3&gt;&lt;P class=""&gt;An ontology can make business concepts and relationships explicit. For example, in a banking space, the enterprise may define relationships between Customer, Account, Transaction, Device and Merchant. That semantic structure helps an AI system understand that these are not independent datasets. It understands the business relationships between them. Ontology therefore provides part of the semantic foundation for a Context Layer. Ontology and context are not the same thing. Ontology answers questions such as&amp;nbsp;&lt;/P&gt;&lt;P class=""&gt;What are the important business concepts? , How are they related? , What does a concept mean? , The Context Layer answers a different question.&amp;nbsp;&lt;STRONG&gt;What of this information and meaning does the agent need for this particular task? &lt;/STRONG&gt;This distinction is important because it prevents the Context Layer from becoming another attempt to model the entire enterprise upfront.&lt;/P&gt;&lt;H3&gt;&lt;U&gt;From Data Quality to Decision Readiness&lt;/U&gt;&lt;/H3&gt;&lt;P class=""&gt;Data Engineering 2.0 made data quality a fundamental engineering discipline. We ask whether data is accurate, complete, fresh, consistent and governed. Data Engineering 3.0 introduces another dimension. &lt;STRONG&gt;Is the context sufficient for the decision? &lt;/STRONG&gt;Consider the fraud example again.&amp;nbsp;The Customer data, Transaction data, Merchant data and Fraud data may be accurate. Yet the agent could still make a poor decision if an important signal is missing from the context. This is why &lt;STRONG&gt;Data Quality and Context Quality &lt;/STRONG&gt;are different problems. A data product can be high quality while the decision context built from several data products is incomplete. This suggests that data engineering will increasingly need to think about &lt;STRONG&gt;decision readiness&lt;/STRONG&gt;.&lt;/P&gt;&lt;P class=""&gt;Does the agent have the right information?&amp;nbsp;Is it sufficiently fresh?, Are the relationships correct?, Are the business definitions consistent?, Are the relevant policies available? andCan the agent identify the evidence behind its conclusion?&amp;nbsp;These questions sit on top of traditional data quality rather than replacing it.&lt;/P&gt;&lt;H3&gt;&lt;U&gt;Context Layer should be driven by the&amp;nbsp;Decision&lt;/U&gt;&lt;/H3&gt;&lt;P class=""&gt;There is a natural temptation to approach this as another enterprise-wide architecture exercise. An organization may say We need to build an Enterprise Context Layer. The next step becomes an attempt to model every entity, relationship, rule, document and event across the organization. That is where the initiative can quickly become too large.&amp;nbsp;The better starting point is the decision. Check &lt;STRONG&gt;Which agent task would benefit most from better enterprise context? &lt;/STRONG&gt;Then define the minimum context required for that task. For a fraud investigation agent, it means starting with Customer, Account, Transaction, Merchant and Device along with the relationships, historical information and policies required for the investigation.&amp;nbsp;The first objective is not to model the bank. It is to make one important decision more reliable. This is a very different way of approaching enterprise semantics.&lt;/P&gt;&lt;H3&gt;&lt;U&gt;Context should evolve with the&amp;nbsp;Agent&lt;/U&gt;&lt;/H3&gt;&lt;P class=""&gt;The first version of a Context Layer will rarely be complete. And it does not need to be. Start with a defined task. Identify the information required. Build the context. Test the agent. Observe where it lacks information or misunderstands meaning. Then extend the context. This creates a practical feedback loop between the agent and the data platform. The agent becomes a source of evidence about where the enterprise’s existing information environment is insufficient for intelligent decision-making.&amp;nbsp;That evidence can drive the next iteration of the data product or semantic layer. Over time, the Context Layer can expand from one decision to multiple decisions and from one domain to connected domains. The goal is not to build the biggest context layer. It is to build the &lt;STRONG&gt;most useful context for the decisions that matter&lt;/STRONG&gt;.&lt;/P&gt;&lt;H3&gt;&lt;U&gt;What changes for Data Engineering?&lt;/U&gt;&lt;/H3&gt;&lt;P class=""&gt;Data Engineering 3.0 does not mean data engineers stop building pipelines. The fundamentals remain critical. Data still needs to be ingested, transformed, stored, governed, monitored and served. But the scope of the discipline expands. Data engineers will increasingly need to understand how enterprise information is represented for intelligent systems. That includes areas such as Semantic modeling, Ontology, Knowledge graphs, Entity resolution, Real-time data, Context retrieval, Agent tools, Data and AI governance, Context evaluation and Decision lineage.&amp;nbsp;&lt;/P&gt;&lt;P class=""&gt;The data engineer’s responsibility increasingly moves from &lt;STRONG&gt;Can we deliver the data? &lt;/STRONG&gt;to &lt;STRONG&gt;Can an intelligent system understand and use the data correctly in the context of the decision?. &lt;/STRONG&gt;That is a significant evolution in the role.&amp;nbsp;The investments made during Data Engineering 2.0 become the foundation for Data Engineering 3.0. Trusted data products provide reliable information. Data quality provides confidence in that information. Governance provides controls. Lineage provides provenance. Semantic models provide business meaning. Real-time platforms provide current state. Ontologies and Knowledge Graphs can provide explicit relationships.&lt;/P&gt;&lt;P class=""&gt;The next step is to bring these capabilities together around the decisions that AI agents need to perform. Context Layer is the &lt;STRONG&gt;new layer built via DE 3.0 on top of the Trusted Data Foundation enterprises have already created in DE 2.0&lt;/STRONG&gt;.&lt;/P&gt;</description>
      <pubDate>Wed, 02 Sep 2026 13:13:51 GMT</pubDate>
      <guid>https://community.databricks.com/t5/community-articles/data-engineering-3-0-from-trusted-data-products-to-context/m-p/167304#M1525</guid>
      <dc:creator>balajij8</dc:creator>
      <dc:date>2026-09-02T13:13:51Z</dc:date>
    </item>
    <item>
      <title>Learn Databricks Lakehouse | From Fundamentals to Hands-On Labs</title>
      <link>https://community.databricks.com/t5/community-articles/learn-databricks-lakehouse-from-fundamentals-to-hands-on-labs/m-p/167293#M1524</link>
      <description>&lt;DIV style="width: 100%; margin: 0 auto; font-family: Arial,Helvetica,sans-serif; color: #1b5162;"&gt;
&lt;DIV style="background-color: #1b5162; padding: 32px 34px 30px 34px; border-radius: 14px 14px 0 0; overflow: hidden;"&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-left" image-alt="DAT_Stacked_Lock_up_Full_Color_White@2x.png" style="width: 104px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/29975iB27742B0C457B7AC/image-size/medium?v=v2&amp;amp;px=400" width="104" role="button" title="DAT_Stacked_Lock_up_Full_Color_White@2x.png" alt="DAT_Stacked_Lock_up_Full_Color_White@2x.png" /&gt;&lt;/span&gt;
&lt;DIV style="color: #9cd3db; font-size: 13px; font-weight: bold; letter-spacing: 2px; text-transform: uppercase;"&gt;Databricks Training &amp;amp; Certifications&lt;/DIV&gt;
&lt;DIV style="color: #ffffff; font-size: 26px; font-weight: 800; line-height: 1.25; margin-top: 4px;"&gt;Learn Databricks Lakehouse &lt;span class="lia-unicode-emoji" title=":national_park:"&gt;🏞&lt;/span&gt;️&lt;/DIV&gt;
&lt;DIV style="color: #cbe9ed; font-size: 14px; margin-top: 6px;"&gt;&amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp;One open platform for all your data, analytics, and AI.&lt;/DIV&gt;
&lt;/DIV&gt;
&lt;DIV style="background-color: #eef7f8; padding: 24px 34px 28px 34px;"&gt;
&lt;DIV style="text-align: center; margin-bottom: 22px;"&gt;&lt;A style="display: inline-block; background-color: #d5ecef; color: #1b5162; font-size: 13px; font-weight: 800; padding: 8px 16px; border-radius: 24px; text-decoration: none; margin: 0 5px 8px 0;" target="_blank"&gt;&lt;span class="lia-unicode-emoji" title=":high_voltage:"&gt;⚡&lt;/span&gt; One open platform&lt;/A&gt; &lt;A style="display: inline-block; background-color: #d5ecef; color: #1b5162; font-size: 13px; font-weight: 800; padding: 8px 16px; border-radius: 24px; text-decoration: none; margin: 0 5px 8px 0;" target="_blank"&gt;&lt;span class="lia-unicode-emoji" title=":bar_chart:"&gt;📊&lt;/span&gt; Warehouse + Lake&lt;/A&gt; &lt;A style="display: inline-block; background-color: #d5ecef; color: #1b5162; font-size: 13px; font-weight: 800; padding: 8px 16px; border-radius: 24px; text-decoration: none; margin: 0 5px 8px 0;" target="_blank"&gt;🧠 Data + AI&lt;/A&gt;&lt;/DIV&gt;
&lt;DIV style="font-size: 17px; line-height: 1.6; color: #1b5162;"&gt;Want one platform for all your data, analytics, and AI? The &lt;STRONG&gt;lakehouse&lt;/STRONG&gt; combines the reliability and performance of a data warehouse with the openness and flexibility of a data lake - built on Delta Lake and governed by Unity Catalog on the Databricks Data Intelligence Platform.&lt;/DIV&gt;
&lt;DIV style="font-size: 17px; line-height: 1.6; color: #1b5162; margin-top: 14px;"&gt;The catalog offers &lt;STRONG&gt;five Lakehouse courses&lt;/STRONG&gt; - from introductory concepts to hands-on Associate labs, most of them free. Follow the path below. &lt;span class="lia-unicode-emoji" title=":backhand_index_pointing_down:"&gt;👇&lt;/span&gt;&lt;/DIV&gt;
&lt;DIV style="margin: 30px 0 14px 0;"&gt;&lt;A style="display: inline-block; width: 34px; height: 34px; line-height: 34px; text-align: center; background-color: #45c4d4; color: #0d3742; font-size: 16px; font-weight: 800; border-radius: 50%; text-decoration: none; vertical-align: middle; margin-right: 12px;" target="_blank"&gt;1&lt;/A&gt; &lt;SPAN&gt;Start here&lt;/SPAN&gt; &lt;SPAN&gt;· Introductory&lt;/SPAN&gt;&lt;/DIV&gt;
&lt;DIV style="background-color: #ffffff; border-left: 5px solid #45c4d4; border-radius: 0 12px 12px 0; padding: 20px 22px; margin-bottom: 6px;"&gt;&lt;A style="display: inline-block; background-color: #d5ecef; color: #1b5162; font-size: 12px; font-weight: 800; padding: 4px 12px; border-radius: 20px; text-decoration: none;" target="_blank"&gt;Introductory · 1H · Free&lt;/A&gt;
&lt;DIV style="font-size: 18px; font-weight: 800; color: #1b5162; line-height: 1.3; margin-top: 10px;"&gt;Lakehouse Architecture Fundamentals&lt;/DIV&gt;
&lt;DIV style="font-size: 14px; line-height: 1.55; color: #4a636a; margin-top: 8px;"&gt;A concept-level course based on The Data Lakehouse For Dummies. Go from “not sure what a lakehouse is” to holding a working conversation about lakehouse architecture - the limits of legacy architectures, the core technologies, and how it enables data intelligence and unified workloads on Databricks.&lt;/DIV&gt;
&lt;DIV style="margin-top: 14px;"&gt;&lt;A style="display: inline-block; background-color: #45c4d4; color: #0d3742; font-size: 13px; font-weight: 800; text-decoration: none; padding: 10px 18px; border-radius: 8px; margin: 0 8px 8px 0;" href="https://www.databricks.com/training/catalog/lakehouse-architecture-fundamentals-6026?itm_source=www&amp;amp;itm_category=training&amp;amp;itm_page=catalog&amp;amp;itm_location=body&amp;amp;itm_component=general-asset-card&amp;amp;itm_offer=lakehouse-architecture-fundamentals-6026" target="_blank"&gt;Free · Self-paced →&lt;/A&gt; &lt;A style="display: inline-block; background-color: #1b5162; color: #ffffff; font-size: 13px; font-weight: 800; text-decoration: none; padding: 10px 18px; border-radius: 8px; margin: 0 8px 8px 0;" href="https://www.databricks.com/training/catalog/lakehouse-architecture-fundamentals-mandarin-chinese-6088?itm_source=www&amp;amp;itm_category=training&amp;amp;itm_page=catalog&amp;amp;itm_location=body&amp;amp;itm_component=general-asset-card&amp;amp;itm_offer=lakehouse-architecture-fundamentals-mandarin-chinese-6088" target="_blank"&gt;Mandarin (中文) →&lt;/A&gt;&lt;/DIV&gt;
&lt;/DIV&gt;
&lt;DIV style="margin: 30px 0 14px 0;"&gt;&lt;A style="display: inline-block; width: 34px; height: 34px; line-height: 34px; text-align: center; background-color: #45c4d4; color: #0d3742; font-size: 16px; font-weight: 800; border-radius: 50%; text-decoration: none; vertical-align: middle; margin-right: 12px;" target="_blank"&gt;2&lt;/A&gt; &lt;SPAN&gt;Get hands-on&lt;/SPAN&gt; &lt;SPAN&gt;· Onboarding&lt;/SPAN&gt;&lt;/DIV&gt;
&lt;DIV style="background-color: #ffffff; border-left: 5px solid #45c4d4; border-radius: 0 12px 12px 0; padding: 20px 22px; margin-bottom: 6px;"&gt;&lt;A style="display: inline-block; background-color: #d5ecef; color: #1b5162; font-size: 12px; font-weight: 800; padding: 4px 12px; border-radius: 20px; text-decoration: none;" target="_blank"&gt;Onboarding · 4H · Free&lt;/A&gt;
&lt;DIV style="font-size: 18px; font-weight: 800; color: #1b5162; line-height: 1.3; margin-top: 10px;"&gt;Databricks Get Started Days (Lakehouse Architecture + Data Warehousing)&lt;/DIV&gt;
&lt;DIV style="font-size: 14px; line-height: 1.55; color: #4a636a; margin-top: 8px;"&gt;A guided, instructor-led day across two tracks - lakehouse platform architecture (the well-architected lakehouse framework) and getting started with Databricks for data warehousing. Create a Free Edition account and try the follow-along demos hands-on.&lt;/DIV&gt;
&lt;DIV style="margin-top: 14px;"&gt;&lt;A style="display: inline-block; background-color: #45c4d4; color: #0d3742; font-size: 13px; font-weight: 800; text-decoration: none; padding: 10px 18px; border-radius: 8px; margin: 0 8px 8px 0;" href="https://www.databricks.com/training/catalog/databricks-get-started-days-lakehouse-architecture-data-warehousing-4060?itm_source=www&amp;amp;itm_category=training&amp;amp;itm_page=catalog&amp;amp;itm_location=body&amp;amp;itm_component=general-asset-card&amp;amp;itm_offer=databricks-get-started-days-lakehouse-architecture-data-warehousing-4060" target="_blank"&gt;Free · Instructor-led →&lt;/A&gt;&lt;/DIV&gt;
&lt;/DIV&gt;
&lt;DIV style="margin: 30px 0 8px 0;"&gt;&lt;A style="display: inline-block; width: 34px; height: 34px; line-height: 34px; text-align: center; background-color: #45c4d4; color: #0d3742; font-size: 16px; font-weight: 800; border-radius: 50%; text-decoration: none; vertical-align: middle; margin-right: 12px;" target="_blank"&gt;3&lt;/A&gt; &lt;SPAN&gt;Go deeper: hands-on labs&lt;/SPAN&gt; &lt;SPAN&gt;· Associate&lt;/SPAN&gt;&lt;/DIV&gt;
&lt;DIV style="font-size: 15px; line-height: 1.6; color: #4a636a; margin-bottom: 16px;"&gt;Where the lakehouse meets Lakebase - serve and sync operational data in real time. Paid / subscription lab format.&lt;/DIV&gt;
&lt;DIV style="background-color: #ffffff; border-left: 5px solid #45c4d4; border-radius: 0 12px 12px 0; padding: 20px 22px; margin-bottom: 14px;"&gt;&lt;A style="display: inline-block; background-color: #d5ecef; color: #1b5162; font-size: 12px; font-weight: 800; padding: 4px 12px; border-radius: 20px; text-decoration: none;" target="_blank"&gt;Associate · 1H · Lab&lt;/A&gt;
&lt;DIV style="font-size: 18px; font-weight: 800; color: #1b5162; line-height: 1.3; margin-top: 10px;"&gt;Serving Lakehouse Data in Real Time with Lakebase&lt;/DIV&gt;
&lt;DIV style="font-size: 14px; line-height: 1.55; color: #4a636a; margin-top: 8px;"&gt;Serve ML features from Lakebase on the hot path. Measure why Delta cannot serve single-row reads, read features from a synced table with millisecond point lookups, and publish features to a managed online feature store.&lt;/DIV&gt;
&lt;DIV style="margin-top: 14px;"&gt;&lt;A style="display: inline-block; background-color: #1b5162; color: #ffffff; font-size: 13px; font-weight: 800; text-decoration: none; padding: 10px 18px; border-radius: 8px; margin: 0 8px 8px 0;" href="https://www.databricks.com/training/catalog/serving-lakehouse-data-in-real-time-with-lakebase-6348?itm_source=www&amp;amp;itm_category=training&amp;amp;itm_page=catalog&amp;amp;itm_location=body&amp;amp;itm_component=general-asset-card&amp;amp;itm_offer=serving-lakehouse-data-in-real-time-with-lakebase-6348" target="_blank"&gt;Paid / Subscription · Lab →&lt;/A&gt;&lt;/DIV&gt;
&lt;/DIV&gt;
&lt;DIV style="background-color: #ffffff; border-left: 5px solid #45c4d4; border-radius: 0 12px 12px 0; padding: 20px 22px; margin-bottom: 6px;"&gt;&lt;A style="display: inline-block; background-color: #d5ecef; color: #1b5162; font-size: 12px; font-weight: 800; padding: 4px 12px; border-radius: 20px; text-decoration: none;" target="_blank"&gt;Associate · 1H · Lab&lt;/A&gt;
&lt;DIV style="font-size: 18px; font-weight: 800; color: #1b5162; line-height: 1.3; margin-top: 10px;"&gt;Postgres to Lakehouse Change Data Feed (CDF)&lt;/DIV&gt;
&lt;DIV style="font-size: 14px; line-height: 1.55; color: #4a636a; margin-top: 8px;"&gt;Capture Lakebase changes back into Delta. Contrast Lakebase CDF with Delta CDF and synced tables, configure REPLICA IDENTITY FULL, interpret the five system columns in the destination Delta table, and choose among four downstream consumption patterns.&lt;/DIV&gt;
&lt;DIV style="margin-top: 14px;"&gt;&lt;A style="display: inline-block; background-color: #1b5162; color: #ffffff; font-size: 13px; font-weight: 800; text-decoration: none; padding: 10px 18px; border-radius: 8px; margin: 0 8px 8px 0;" href="https://www.databricks.com/training/catalog/postgres-to-lakehouse-change-data-feed-cdf-6344?itm_source=www&amp;amp;itm_category=training&amp;amp;itm_page=catalog&amp;amp;itm_location=body&amp;amp;itm_component=general-asset-card&amp;amp;itm_offer=postgres-to-lakehouse-change-data-feed-cdf-6344" target="_self"&gt;Paid / Subscription · Lab →&lt;/A&gt;&lt;/DIV&gt;
&lt;/DIV&gt;
&lt;DIV style="text-align: center; margin-top: 24px;"&gt;&lt;A style="display: inline-block; background-color: #45c4d4; color: #0d3742; font-size: 15px; font-weight: 800; text-decoration: none; padding: 14px 30px; border-radius: 8px;" href="https://www.databricks.com/training/catalog?search=lakehouse" target="_blank"&gt;Browse all Lakehouse courses in the catalog →&lt;/A&gt;&lt;/DIV&gt;
&lt;/DIV&gt;
&lt;DIV style="background-color: #1b5162; padding: 22px 34px; border-radius: 0 0 14px 14px; text-align: center;"&gt;
&lt;DIV style="color: #cbe9ed; font-size: 13px; line-height: 1.6;"&gt;Warehouse reliability. Data lake flexibility. One open platform. Learn the lakehouse the right way with official Databricks training.&lt;/DIV&gt;
&lt;/DIV&gt;
&lt;/DIV&gt;</description>
      <pubDate>Wed, 02 Sep 2026 12:23:58 GMT</pubDate>
      <guid>https://community.databricks.com/t5/community-articles/learn-databricks-lakehouse-from-fundamentals-to-hands-on-labs/m-p/167293#M1524</guid>
      <dc:creator>Tushar_Parekar</dc:creator>
      <dc:date>2026-09-02T12:23:58Z</dc:date>
    </item>
    <item>
      <title>create_auto_cdc_from_snapshot_flow Python session resolution fails if having multiple snapshot flows</title>
      <link>https://community.databricks.com/t5/data-engineering/create-auto-cdc-from-snapshot-flow-python-session-resolution/m-p/167268#M55663</link>
      <description>&lt;P&gt;When a pipeline contains more than one create_auto_cdc_from_snapshot_flow flow (each driven by a custom Python next_snapshot_and_version function), flow resolution fails intermittently/consistently with:&lt;/P&gt;&lt;DIV&gt;&lt;DIV&gt;&amp;nbsp;&lt;/DIV&gt;&lt;/DIV&gt;&lt;P&gt;RuntimeError: The original Spark session is being accessed instead of the per-flow cloned session during parallel analysis. This is commonly caused by spawning threads inside a flow function that access the Spark session.&lt;/P&gt;&lt;P&gt;I am having functions to figure out the next snapshot like this:&lt;BR /&gt;```&lt;BR /&gt;def next_x_snapshot_and_version(latest_version):&lt;BR /&gt;versions = spark.read.table(SOURCE_TABLE).select("file_modification_time").distinct()&lt;BR /&gt;```&lt;BR /&gt;&lt;BR /&gt;Is this a bug?&lt;/P&gt;</description>
      <pubDate>Wed, 02 Sep 2026 09:43:25 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/create-auto-cdc-from-snapshot-flow-python-session-resolution/m-p/167268#M55663</guid>
      <dc:creator>david_aspegren</dc:creator>
      <dc:date>2026-09-02T09:43:25Z</dc:date>
    </item>
    <item>
      <title>Virtual Event | Closing the AI Context Gap: How to Teach AI How Your Business Actually Runs</title>
      <link>https://community.databricks.com/t5/learning-events/virtual-event-closing-the-ai-context-gap-how-to-teach-ai-how/ec-p/167266#M10228</link>
      <description>&lt;P&gt;&lt;STRONG&gt;Virtual Event&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Time: &lt;BR /&gt;&lt;/STRONG&gt;&lt;STRONG&gt;AMER&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;Sept 23 / 9 AM PT&lt;/SPAN&gt;&lt;BR /&gt;&lt;STRONG&gt;EMEA&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;Sept 24 / 9 AM BST / 10 AM CEST&lt;/SPAN&gt;&lt;BR /&gt;&lt;STRONG&gt;APJ&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;Sept 24 / 9:30 AM IST / 12 PM SGT / 1 PM JST/KST / 2 PM AEST (&lt;/SPAN&gt;&lt;EM&gt;Japanese and Korean captions available&lt;/EM&gt;&lt;SPAN&gt;)&lt;/SPAN&gt;&lt;/P&gt;
&lt;P class="p8i6j01 paragraph lia-align-center"&gt;&lt;A style="background-color: #ff3621; color: white; padding: 10px 20px; text-decoration: none; border-radius: 5px; font-weight: bold; display: inline-block;" href="https://www.databricks.com/resources/webinar/closing-the-ai-context-gap?utm_source=bambu&amp;amp;utm_medium=organic-social&amp;amp;utm_scid=701Vp000013pWEvIAM" target="_blank" rel="noopener"&gt; &lt;span class="lia-unicode-emoji" title=":link:"&gt;🔗&lt;/span&gt; Register for the Workshop &lt;/A&gt;&lt;/P&gt;
&lt;P data-path-to-node="4"&gt;&lt;FONT color="#999999"&gt;Note: Marking &lt;STRONG&gt;RSVP&lt;/STRONG&gt; on this Community event does &lt;STRONG&gt;not&lt;/STRONG&gt; register you for the workshop. Please use the &lt;STRONG&gt;Register&lt;/STRONG&gt; button above to complete your registration on the official event page.&lt;/FONT&gt;&lt;/P&gt;
&lt;P data-path-to-node="1"&gt;AI doesn't have an intelligence problem, it has a context problem. Ground truth for business operations remains scattered across dashboards, documents, tickets, and chats. Join us for a virtual session on how to give AI a unified context layer to deliver trusted answers, operate autonomously, and integrate seamlessly into your team's workflow.&lt;/P&gt;
&lt;P data-path-to-node="2"&gt;&lt;STRONG data-index-in-node="0" data-path-to-node="2"&gt;Key Takeaways:&lt;BR /&gt;&lt;/STRONG&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI data-path-to-node="2"&gt;Learn to enable a living context graph that continuously infers business terms, entities, and KPIs.&lt;/LI&gt;
&lt;LI data-path-to-node="2"&gt;Ensure AI delivers trusted answers based on individual user access.&lt;/LI&gt;
&lt;LI data-path-to-node="2"&gt;Turn insights into autonomous actions, agents, and apps without code.&lt;/LI&gt;
&lt;LI data-path-to-node="2"&gt;Bring context-aware AI into Slack, Teams, mobile, and MCP-based experiences.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P data-path-to-node="4"&gt;&lt;STRONG data-index-in-node="0" data-path-to-node="4"&gt;Agenda (60 Minutes):&lt;/STRONG&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI data-path-to-node="4"&gt;&lt;STRONG data-index-in-node="0" data-path-to-node="5,0,0"&gt;Welcome&lt;/STRONG&gt; (5 mins)&lt;/LI&gt;
&lt;LI data-path-to-node="4"&gt;&lt;STRONG data-index-in-node="0" data-path-to-node="5,1,0"&gt;Solving the AI Context Problem&lt;/STRONG&gt; (15 mins)&lt;/LI&gt;
&lt;LI data-path-to-node="4"&gt;&lt;STRONG data-index-in-node="0" data-path-to-node="5,2,0"&gt;Under the Hood: How Genie Ontology Works&lt;/STRONG&gt; (20 mins)&lt;/LI&gt;
&lt;LI data-path-to-node="4"&gt;&lt;STRONG data-index-in-node="0" data-path-to-node="5,3,0"&gt;Genie Ontology in Action: Product Demo&lt;/STRONG&gt; (20 mins)&lt;/LI&gt;
&lt;LI data-path-to-node="4"&gt;&lt;STRONG data-index-in-node="0" data-path-to-node="5,4,0"&gt;Wrap Up&lt;/STRONG&gt; (5 mins)&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="Screenshot 2026-09-02 at 14.11.59.png" style="width: 939px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30633iB54FEC4D2F5764A8/image-size/large?v=v2&amp;amp;px=999" role="button" title="Screenshot 2026-09-02 at 14.11.59.png" alt="Screenshot 2026-09-02 at 14.11.59.png" /&gt;&lt;/span&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Wed, 02 Sep 2026 09:23:56 GMT</pubDate>
      <guid>https://community.databricks.com/t5/learning-events/virtual-event-closing-the-ai-context-gap-how-to-teach-ai-how/ec-p/167266#M10228</guid>
      <dc:creator>Tushar_Parekar</dc:creator>
      <dc:date>2026-09-02T09:23:56Z</dc:date>
    </item>
    <item>
      <title>Databricks Data Ingestion</title>
      <link>https://community.databricks.com/t5/get-started-discussions/databricks-data-ingestion/m-p/167261#M12064</link>
      <description>&lt;P&gt;Hi all,&lt;BR /&gt;&lt;BR /&gt;I am currently exploring the data ingestion feature offered by databricks, specifically connecting to a SQL server.&amp;nbsp;&lt;BR /&gt;I have gone through the UI and configured an ingestion pipeline that connects to one specific server and can read tables into databricks.&amp;nbsp;&lt;/P&gt;&lt;P&gt;I am currently venturing into understanding if we can make an ingestion pipeline dynamic.&lt;/P&gt;&lt;P&gt;What I mean by this is, I have created a yaml file that can be run to setup an ingestion pipeline. It accepts variables from a metadata config table that can be used to populate the YAML with information such as, connection server name, pipeline name, source catalog, source schema, source table, destination catalog, destination schema and destination table name.&amp;nbsp;&lt;/P&gt;&lt;P&gt;&lt;BR /&gt;So the question I have is, can we have a generic pipeline whose configurations can be altered?&amp;nbsp;&lt;BR /&gt;Use cases:&lt;BR /&gt;1. Once a pipeline is created can it be edited in the future to ingest different tables? As in using pipeline_01, I ingested table_1, table_2 present in Server_01. Can I later pass different configs to the YAML such that I use the same pipeline_01 to ingest table_03, table_04?&lt;/P&gt;&lt;P&gt;2. Can I use pipeline_01 to connect to a different to a connection server and ingest tables from this new server?&lt;/P&gt;&lt;P&gt;Currently I am not finding a way to achieve this functionality. This will cause an issue in the future when I have say 50 servers. It will be difficult to maintain 50 different pipelines and I don't think this is desirable as well.&amp;nbsp;&lt;BR /&gt;&lt;BR /&gt;Can I get some insights into Data Ingestion feature and if this a viable option for my use case?&lt;BR /&gt;&lt;BR /&gt;Thank you,&lt;BR /&gt;Anush&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Wed, 02 Sep 2026 09:09:04 GMT</pubDate>
      <guid>https://community.databricks.com/t5/get-started-discussions/databricks-data-ingestion/m-p/167261#M12064</guid>
      <dc:creator>anushnagesh</dc:creator>
      <dc:date>2026-09-02T09:09:04Z</dc:date>
    </item>
    <item>
      <title>Difference between Workspace and Unity Catalog experiments when using MLflow autologging?</title>
      <link>https://community.databricks.com/t5/machine-learning/difference-between-workspace-and-unity-catalog-experiments-when/m-p/167254#M4682</link>
      <description>&lt;DIV class=""&gt;&lt;DIV&gt;&lt;SPAN&gt;Hi everyone,&lt;/SPAN&gt;&lt;/DIV&gt;&lt;BR /&gt;&lt;DIV&gt;I am trying to understand the exact differences between using Workspace experiments versus Unity Catalog experiments, specifically in the context of MLflow autologging (mlflow.autolog()).&lt;/DIV&gt;&lt;BR /&gt;&lt;DIV&gt;Does autologging behave differently depending on whether the experiment is registered in the Workspace or in Unity Catalog (using a 3-level namespace)? Are there any limitations, best practices, or specific configurations I should be aware of when using autologging with Unity Catalog compared to the traditional Workspace setup?&lt;/DIV&gt;&lt;BR /&gt;&lt;DIV&gt;Any insights or documentation links would be greatly appreciated. Thanks!&lt;/DIV&gt;&lt;/DIV&gt;</description>
      <pubDate>Wed, 02 Sep 2026 08:44:13 GMT</pubDate>
      <guid>https://community.databricks.com/t5/machine-learning/difference-between-workspace-and-unity-catalog-experiments-when/m-p/167254#M4682</guid>
      <dc:creator>kunduruanil</dc:creator>
      <dc:date>2026-09-02T08:44:13Z</dc:date>
    </item>
    <item>
      <title>How can i rename a column  in a delta table?</title>
      <link>https://community.databricks.com/t5/data-engineering/how-can-i-rename-a-column-in-a-delta-table/m-p/167250#M55658</link>
      <description>&lt;P&gt;I have a delta table and i want to rename one of its columns.&lt;/P&gt;&lt;P&gt;what is the recommended way to rename a column in databricks?&lt;/P&gt;&lt;P&gt;is there any difference between renaming a column using sql and using pyspark?&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Wed, 02 Sep 2026 07:06:30 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/how-can-i-rename-a-column-in-a-delta-table/m-p/167250#M55658</guid>
      <dc:creator>gowri_databrick</dc:creator>
      <dc:date>2026-09-02T07:06:30Z</dc:date>
    </item>
    <item>
      <title>All 18 Lakeflow AUTO CDC configurations went green. Five failed my ship check</title>
      <link>https://community.databricks.com/t5/community-articles/all-18-lakeflow-auto-cdc-configurations-went-green-five-failed/m-p/167235#M1520</link>
      <description>&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="Paper event slips trace 18 Lakeflow AUTO CDC configurations: 13 stay on the main sequence while five branch into production stop signs." style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30626iEDB2F22254223A83/image-size/large?v=v2&amp;amp;px=999" role="button" title="exec-68d729e6-990b-4352-822c-6eb77f737c6b.png" alt="Paper event slips trace 18 Lakeflow AUTO CDC configurations: 13 stay on the main sequence while five branch into production stop signs." /&gt;&lt;span class="lia-inline-image-caption" onclick="event.preventDefault();"&gt;Paper event slips trace 18 Lakeflow AUTO CDC configurations: 13 stay on the main sequence while five branch into production stop signs.&lt;/span&gt;&lt;/span&gt;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;Change data capture (CDC) keeps a downstream table in step with row-level inserts, updates, and deletes instead of reloading the whole source. Lakeflow AUTO CDC handles the state-management work, including sequencing, deletes, and SCD history. It still needs a precise source contract: which clock wins, how ties break, what NULL means, and which changes deserve history.&lt;/P&gt;&lt;P&gt;I built this experiment to see what happens when source data is late, duplicated, contradictory, or noisy, and to separate a pipeline that finishes from a target I would trust. The suite pushes nine hostile CDC patterns across 13 isolated source tables and 18 AUTO CDC configurations: duplicates, late events, tied sequence values, conflicting clocks, sparse NULL updates, deletes, replays, sync-noise updates, and bitemporal corrections.&lt;/P&gt;&lt;P&gt;All 18 configurations completed. Five green configurations failed my ship check: three had complete ordering but violated the stated business rule, and two used incomplete ordering.&lt;/P&gt;&lt;P&gt;Start with scenario 4. Ordering the same rows by ingestion time kept ACTIVE; source event time produced SUSPENDED, the expected business state. Both configurations finished green.&lt;/P&gt;&lt;P&gt;I used pipeline status to confirm execution. I used target-state assertions to decide whether I would ship.&lt;/P&gt;&lt;BLOCKQUOTE&gt;&lt;P&gt;&lt;STRONG&gt;Independent experiment.&lt;/STRONG&gt; I ran this test for my own engineering work. It is not official Databricks guidance. I’m happy to discuss your results, feedback, and the failure modes you think I missed.&lt;/P&gt;&lt;/BLOCKQUOTE&gt;&lt;P&gt;&lt;STRONG&gt;&lt;A href="https://github.com/ivanvyd/lakeflow-auto-cdc-torture-test" target="_blank" rel="noopener"&gt;Run the experiment&lt;/A&gt;&lt;/STRONG&gt; · &lt;STRONG&gt;&lt;A href="https://github.com/ivanvyd/lakeflow-auto-cdc-torture-test/blob/main/results/normalized/summary_matrix.json" target="_blank" rel="noopener"&gt;Inspect the result matrix&lt;/A&gt;&lt;/STRONG&gt;&lt;/P&gt;&lt;HR /&gt;&lt;H2&gt;Results in one minute&lt;/H2&gt;&lt;UL&gt;&lt;LI&gt;&lt;STRONG&gt;10 handled:&lt;/STRONG&gt; keep the sequence rule and test it against your source.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;3 configuration-dependent:&lt;/STRONG&gt; set the option that matches the business rule.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;3 business-semantics stops:&lt;/STRONG&gt; change the chosen clock, NULL meaning, or history policy.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;2 ambiguous-order stops:&lt;/STRONG&gt; add a source-side tie-breaker.&lt;/LI&gt;&lt;/UL&gt;&lt;P&gt;Execution stayed green across all five stop signs. Source semantics or ordering failed my ship check.&lt;/P&gt;&lt;H2&gt;The five stop signs&lt;/H2&gt;&lt;P&gt;Configuration Measured result Fix before production&lt;/P&gt;&lt;TABLE&gt;&lt;TBODY&gt;&lt;TR&gt;&lt;TD&gt;3A Sequence collision&lt;/TD&gt;&lt;TD&gt;Two business states shared one sequence value. The contract names no winner.&lt;/TD&gt;&lt;TD&gt;Reject ties or add a stable source-side tie-breaker.&lt;/TD&gt;&lt;/TR&gt;&lt;TR&gt;&lt;TD&gt;3B Tie-breaker ignored&lt;/TD&gt;&lt;TD&gt;The source supplied transaction_sequence, but the flow left it out of SEQUENCE BY.&lt;/TD&gt;&lt;TD&gt;Use a composite sequence.&lt;/TD&gt;&lt;/TR&gt;&lt;TR&gt;&lt;TD&gt;4A Ingestion-time order&lt;/TD&gt;&lt;TD&gt;The target kept ACTIVE; source time required SUSPENDED.&lt;/TD&gt;&lt;TD&gt;Use the clock that defines business recency.&lt;/TD&gt;&lt;/TR&gt;&lt;TR&gt;&lt;TD&gt;5A Default NULL handling&lt;/TD&gt;&lt;TD&gt;A sparse update replaced the existing email with NULL.&lt;/TD&gt;&lt;TD&gt;Define NULL semantics and use IGNORE NULL UPDATES when NULL means absent.&lt;/TD&gt;&lt;/TR&gt;&lt;TR&gt;&lt;TD&gt;8A Track every column&lt;/TD&gt;&lt;TD&gt;Fifty sync-timestamp updates created 51 SCD2 rows.&lt;/TD&gt;&lt;TD&gt;Exclude operational metadata from history tracking.&lt;/TD&gt;&lt;/TR&gt;&lt;/TBODY&gt;&lt;/TABLE&gt;&lt;P&gt;The other 13 configurations matched the experiment’s business rule under a complete order and the required option.&lt;/P&gt;&lt;HR /&gt;&lt;H2&gt;How I tested it&lt;/H2&gt;&lt;P&gt;The generator creates small, isolated source tables with one failure mode per scenario. One pipeline wraps each source in a streaming view and runs all 18 AUTO CDC configurations.&lt;/P&gt;&lt;P&gt;For late events and replays, I ran two updates. The full refresh established a baseline. The incremental update appended the withheld rows. The verifier required the two expected histories to change, the other 16 targets to remain equal, and every target to match its row-count and observed-state predicate. The final classification combined those measurements with the declared ordering and business-rule labels.&lt;/P&gt;&lt;P&gt;You can inspect the code in &lt;A href="https://github.com/ivanvyd/lakeflow-auto-cdc-torture-test/blob/main/src/pipeline/pipeline.py" target="_blank" rel="noopener"&gt;src/pipeline/pipeline.py&lt;/A&gt; and &lt;A href="https://github.com/ivanvyd/lakeflow-auto-cdc-torture-test/blob/main/src/generators/dispatch.py" target="_blank" rel="noopener"&gt;src/generators/dispatch.py&lt;/A&gt;. The &lt;A href="https://github.com/ivanvyd/lakeflow-auto-cdc-torture-test/blob/main/results/normalized/summary_matrix.json" target="_blank" rel="noopener"&gt;result matrix&lt;/A&gt; and &lt;A href="https://github.com/ivanvyd/lakeflow-auto-cdc-torture-test/blob/main/results/raw/target_state.json" target="_blank" rel="noopener"&gt;captured target rows&lt;/A&gt; carry the measured evidence. Official Databricks documentation used for the platform claims is indexed in &lt;A href="https://github.com/ivanvyd/lakeflow-auto-cdc-torture-test/blob/main/docs/sources.md" target="_blank" rel="noopener"&gt;docs/sources.md&lt;/A&gt;.&lt;/P&gt;&lt;HR /&gt;&lt;H2&gt;Four source decisions that change the answer&lt;/H2&gt;&lt;H3&gt;1. Pick the clock that means “newer”&lt;/H3&gt;&lt;P&gt;The generator emits two events for the same key:&lt;/P&gt;&lt;PRE&gt;10:00 source time  ACTIVE
10:05 source time  SUSPENDED&lt;/PRE&gt;&lt;P&gt;Their ingestion timestamps reverse that order. Ordering by ingested_at keeps ACTIVE. Ordering by source_updated_at produces SUSPENDED, which matches the experiment’s business rule.&lt;/P&gt;&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="Two AUTO CDC targets show that ingestion time misses the business expectation while source time matches it." style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30627i88EA6F2049895B73/image-size/large?v=v2&amp;amp;px=999" role="button" title="wrong_clock.png" alt="Two AUTO CDC targets show that ingestion time misses the business expectation while source time matches it." /&gt;&lt;span class="lia-inline-image-caption" onclick="event.preventDefault();"&gt;Two AUTO CDC targets show that ingestion time misses the business expectation while source time matches it.&lt;/span&gt;&lt;/span&gt;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;&lt;EM&gt;Two AUTO CDC targets use the same rows and schema. The sequence clock changes the answer.&lt;/EM&gt;&lt;/P&gt;&lt;P&gt;Arrival time helps you analyze transport; source time can define business recency. Pick the column from the business definition before you write the flow.&lt;/P&gt;&lt;H3&gt;2. Give ties a real winner&lt;/H3&gt;&lt;P&gt;The collision pattern gives two states the same source_sequence:&lt;/P&gt;&lt;PRE&gt;seq=10  status=ACTIVE
seq=10  status=SUSPENDED&lt;/PRE&gt;&lt;P&gt;I observed SUSPENDED, but the configured order cannot distinguish the rows, so I recorded the result as AMBIGUOUS_ORDER.&lt;/P&gt;&lt;P&gt;The source also provides transaction_sequence. A composite sequence such as STRUCT(source_updated_at, transaction_sequence) gives the flow a stable order. The composite configuration produced the expected SUSPENDED state, and the verifier classified it as HANDLED.&lt;/P&gt;&lt;H3&gt;3. Decide what NULL means&lt;/H3&gt;&lt;P&gt;The sparse-update pattern begins with email='x@example.com'. The next event updates the city and carries email=NULL.&lt;/P&gt;&lt;P&gt;The default configuration sets the target email to NULL. IGNORE NULL UPDATES keeps the existing email. Both behaviors can serve a valid source contract. Your producer must define whether NULL means “erase this value” or “this field was absent.”&lt;/P&gt;&lt;H3&gt;4. Keep sync noise out of business history&lt;/H3&gt;&lt;P&gt;The history-noise generator emits 50 updates. Among the retained target columns, only last_synced_at changes.&lt;/P&gt;&lt;P&gt;Tracking every included target column creates 51 SCD2 rows. Excluding last_synced_at creates one row because the business fields never change.&lt;/P&gt;&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="Tracking every included target column produces 51 SCD2 rows; excluding last_synced_at produces one." style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30628i882EC60FE5C94C52/image-size/large?v=v2&amp;amp;px=999" role="button" title="scd2_history_noise.png" alt="Tracking every included target column produces 51 SCD2 rows; excluding last_synced_at produces one." /&gt;&lt;span class="lia-inline-image-caption" onclick="event.preventDefault();"&gt;Tracking every included target column produces 51 SCD2 rows; excluding last_synced_at produces one.&lt;/span&gt;&lt;/span&gt;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;&lt;EM&gt;Operational sync metadata accounts for all 50 extra versions.&lt;/EM&gt;&lt;/P&gt;&lt;P&gt;Choose the columns that represent business history before you deploy SCD2. Whether operational changes belong there is a domain and audit decision; this experiment’s business rule counted only business-field changes.&lt;/P&gt;&lt;HR /&gt;&lt;H2&gt;Patterns AUTO CDC handled&lt;/H2&gt;&lt;P&gt;Pattern Measured result&lt;/P&gt;&lt;TABLE&gt;&lt;TBODY&gt;&lt;TR&gt;&lt;TD&gt;Identical duplicate and late replay&lt;/TD&gt;&lt;TD&gt;SCD1 kept one ACTIVE row. The replay left visible state unchanged.&lt;/TD&gt;&lt;/TR&gt;&lt;TR&gt;&lt;TD&gt;Out-of-order update&lt;/TD&gt;&lt;TD&gt;SCD1 kept the newer state. SCD2 inserted the older event as a closed history row.&lt;/TD&gt;&lt;/TR&gt;&lt;TR&gt;&lt;TD&gt;Delete followed by an older event&lt;/TD&gt;&lt;TD&gt;SCD1 kept the deletion. SCD2 inserted the late state before the delete boundary.&lt;/TD&gt;&lt;/TR&gt;&lt;TR&gt;&lt;TD&gt;Full lifecycle replay&lt;/TD&gt;&lt;TD&gt;SCD1 and SCD2 matched their saved baselines after the replay.&lt;/TD&gt;&lt;/TR&gt;&lt;TR&gt;&lt;TD&gt;Bitemporal correction&lt;/TD&gt;&lt;TD&gt;The target preserved business time and system time across five measured rows.&lt;/TD&gt;&lt;/TR&gt;&lt;/TBODY&gt;&lt;/TABLE&gt;&lt;P&gt;These are visible-state observations from appended replays. They do not establish transactional deduplication or test a checkpoint-only restart with no new source rows.&lt;/P&gt;&lt;H3&gt;Bitemporal kept both clocks&lt;/H3&gt;&lt;P&gt;Scenario 9 uses source_updated_at for business time and ingested_at for system time. The target stores __START_AT / __END_AT beside __SYSTEM_START_AT / __SYSTEM_END_AT.&lt;/P&gt;&lt;P&gt;Three events produced five rows because later ingestion times revised earlier valid-time intervals.&lt;/P&gt;&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="Five measured bitemporal rows show original and revised valid-time intervals across three system times." style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30629i605D493792A77755/image-size/large?v=v2&amp;amp;px=999" role="button" title="bitemporal_timeline.png" alt="Five measured bitemporal rows show original and revised valid-time intervals across three system times." /&gt;&lt;span class="lia-inline-image-caption" onclick="event.preventDefault();"&gt;Five measured bitemporal rows show original and revised valid-time intervals across three system times.&lt;/span&gt;&lt;/span&gt;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;&lt;EM&gt;Five rows preserve the original and revised valid-time intervals.&lt;/EM&gt;&lt;/P&gt;&lt;P&gt;Databricks marks bitemporal storage as Beta. This run validates five tiny rows and does not test load. I would evaluate it when valid time differs from ingest time and consumers need as-of-system-time queries.&lt;/P&gt;&lt;HR /&gt;&lt;H2&gt;The full result map&lt;/H2&gt;&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="All 18 measured configurations grouped by handled, configuration-dependent, business-semantics, and ambiguous-order outcomes." style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30630i34D53099C293954C/image-size/large?v=v2&amp;amp;px=999" role="button" title="summary_matrix.png" alt="All 18 measured configurations grouped by handled, configuration-dependent, business-semantics, and ambiguous-order outcomes." /&gt;&lt;span class="lia-inline-image-caption" onclick="event.preventDefault();"&gt;All 18 measured configurations grouped by handled, configuration-dependent, business-semantics, and ambiguous-order outcomes.&lt;/span&gt;&lt;/span&gt;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;&lt;EM&gt;All 18 measured configurations grouped by production outcome.&lt;/EM&gt;&lt;/P&gt;&lt;P&gt;You can inspect the machine-readable matrix in &lt;A href="https://github.com/ivanvyd/lakeflow-auto-cdc-torture-test/blob/main/results/normalized/summary_matrix.json" target="_blank" rel="noopener"&gt;results/normalized/summary_matrix.json&lt;/A&gt; and the captured rows in &lt;A href="https://github.com/ivanvyd/lakeflow-auto-cdc-torture-test/blob/main/results/raw/target_state.json" target="_blank" rel="noopener"&gt;results/raw/target_state.json&lt;/A&gt;.&lt;/P&gt;&lt;HR /&gt;&lt;H2&gt;My production checklist&lt;/H2&gt;&lt;OL&gt;&lt;LI&gt;&lt;STRONG&gt;Name the business clock.&lt;/STRONG&gt; Put the column that defines recency in SEQUENCE BY.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Search for ties per key.&lt;/STRONG&gt; Add a source-side tie-breaker when two states can share one sequence value.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Write down NULL semantics.&lt;/STRONG&gt; Use IGNORE NULL UPDATES only when NULL means absent.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Choose the state model.&lt;/STRONG&gt; SCD1 keeps current state; SCD2 keeps ordered history.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Filter history noise.&lt;/STRONG&gt; Exclude sync metadata unless it represents a business event.&lt;/LI&gt;&lt;/OL&gt;&lt;P&gt;Add a target-state predicate for each rule: update status checks execution, while the predicate checks the expected target state.&lt;/P&gt;&lt;HR /&gt;&lt;H2&gt;Reproduce the run&lt;/H2&gt;&lt;PRE&gt;git clone https://github.com/ivanvyd/lakeflow-auto-cdc-torture-test.git
cd lakeflow-auto-cdc-torture-test
python -m pip install -e ".[dev]"
databricks auth login --host https://&amp;lt;workspace-url&amp;gt; --profile DEFAULT
make setup
make test
make results&lt;/PRE&gt;&lt;P&gt;You need Python 3.10 or newer, GNU Make, jq, the Databricks CLI, and access to a Databricks workspace. The &lt;A href="https://github.com/ivanvyd/lakeflow-auto-cdc-torture-test/blob/main/docs/reproduction.md" target="_blank" rel="noopener"&gt;reproduction guide&lt;/A&gt; covers the one-time setup and evidence capture.&lt;/P&gt;&lt;P&gt;Sixteen targets matched their baselines after update 2. The two SCD2 targets that received late history changed as expected. The verifier checked all 18 targets and produced 10 HANDLED, 3 CONFIGURATION_DEPENDENT, 3 BUSINESS_SEMANTICS, and 2 AMBIGUOUS_ORDER results.&lt;/P&gt;&lt;P&gt;The checked-in result set was captured on 2026-09-01 from clean commit 01d53b4; target_state.json records the pipeline and both update IDs.&lt;/P&gt;&lt;H2&gt;Scope&lt;/H2&gt;&lt;P&gt;This suite uses tiny controlled datasets. I did not test throughput, backpressure, or large joins. Each flow read one isolated source table; I did not test joins or interactions across streams. The delete-and-late-event case stayed inside the configured 48-hour tombstone-retention window. Bitemporal load behavior also sits outside this run.&lt;/P&gt;&lt;P&gt;The evidence comes from one workspace, one SQL warehouse, one customer key, serverless Advanced-edition compute on the CURRENT channel, and tiny deterministic inputs. Use separate tests for throughput, schema evolution, and multi-stream joins.&lt;/P&gt;&lt;H2&gt;Bring your failure mode&lt;/H2&gt;&lt;P&gt;Clone the repository and replace one generator with an event sequence from your source. Add the expected target state, run both phases, and compare the capture.&lt;/P&gt;&lt;P&gt;If the suite misses your case, &lt;A href="https://github.com/ivanvyd/lakeflow-auto-cdc-torture-test/issues/new" target="_blank" rel="noopener"&gt;open an issue&lt;/A&gt; with the source rows, SEQUENCE BY expression, storage type, expected target, and observed target. I’m happy to turn a clear failure report into another reproducible scenario.&lt;/P&gt;</description>
      <pubDate>Wed, 02 Sep 2026 01:54:46 GMT</pubDate>
      <guid>https://community.databricks.com/t5/community-articles/all-18-lakeflow-auto-cdc-configurations-went-green-five-failed/m-p/167235#M1520</guid>
      <dc:creator>ivanvyd</dc:creator>
      <dc:date>2026-09-02T01:54:46Z</dc:date>
    </item>
    <item>
      <title>Lakebase Postgre updating Delta Table.</title>
      <link>https://community.databricks.com/t5/data-engineering/lakebase-postgre-updating-delta-table/m-p/167234#M55653</link>
      <description>&lt;P&gt;I am using Postgre for OLTP processing for POS application.Lag is reduced a lot, however when there is updation on Postgre table, I need to sync back to delta table. There is one way from delta table sync table (read only). Any design pattern sas to keep delta table and postgre table in sync.&lt;/P&gt;</description>
      <pubDate>Wed, 02 Sep 2026 01:44:54 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/lakebase-postgre-updating-delta-table/m-p/167234#M55653</guid>
      <dc:creator>rkhbo3003</dc:creator>
      <dc:date>2026-09-02T01:44:54Z</dc:date>
    </item>
    <item>
      <title>From Semantic Similarity to Business Authority: Why Genie Ontology and OntoRank Matter</title>
      <link>https://community.databricks.com/t5/get-started-discussions/from-semantic-similarity-to-business-authority-why-genie/m-p/167223#M12062</link>
      <description>&lt;P&gt;Enterprise AI does not usually fail because the model lacks intelligence.&lt;/P&gt;&lt;P&gt;It fails because the model does not understand what the organization means.&lt;/P&gt;&lt;P&gt;Consider a simple question:&lt;/P&gt;&lt;P&gt;“What is our current exposure to active customer?”&lt;/P&gt;&lt;P&gt;Behind this question are several business decisions:&lt;/P&gt;&lt;P&gt;What qualifies as an “active” customer?&lt;/P&gt;&lt;P&gt;Should exposure include committed, outstanding, or available amounts?&lt;/P&gt;&lt;P&gt;Which customer identifier is authoritative?&lt;/P&gt;&lt;P&gt;Should rebooked amount be consolidated?&lt;/P&gt;&lt;P&gt;Which system is trusted: the servicing platform, CRM, MDM golden record, or a reporting mart?&lt;/P&gt;&lt;P&gt;What business date should be used?&lt;/P&gt;&lt;P&gt;A traditional text-to-SQL system may identify tables with similar column names and generate syntactically correct SQL. But syntactically correct SQL can still produce a completely incorrect business answer.&lt;/P&gt;&lt;P&gt;This is the context gap that Databricks Genie Ontology is designed to address.&lt;/P&gt;&lt;P&gt;What is Genie Ontology?&lt;BR /&gt;Genie Ontology is a unified, continuously improving context layer that gives Genie a business-aware map of the organization.&lt;/P&gt;&lt;P&gt;It combines:&lt;/P&gt;&lt;P&gt;Human-modeled context&lt;/P&gt;&lt;P&gt;Certified data products, Unity Catalog metric views, domains, business definitions, Pages, and governed assets.&lt;/P&gt;&lt;P&gt;Automatically inferred context&lt;/P&gt;&lt;P&gt;Knowledge extracted from tables, queries, dashboards, SQL patterns, Genie Agents, and platform usage.&lt;/P&gt;&lt;P&gt;Instead of treating enterprise knowledge as disconnected metadata, the ontology represents relationships among:&lt;/P&gt;&lt;P&gt;Business terms&lt;/P&gt;&lt;P&gt;Metrics&lt;/P&gt;&lt;P&gt;tables and columns&lt;/P&gt;&lt;P&gt;dashboards&lt;/P&gt;&lt;P&gt;queries&lt;/P&gt;&lt;P&gt;data products&lt;/P&gt;&lt;P&gt;people and teams&lt;/P&gt;&lt;P&gt;business rules&lt;/P&gt;&lt;P&gt;Genie Agents&lt;/P&gt;&lt;P&gt;This changes the question from:&lt;/P&gt;&lt;P&gt;“Which asset looks most similar to the user’s prompt?”&lt;/P&gt;&lt;P&gt;to:&lt;/P&gt;&lt;P&gt;“Which permitted source represents the most authoritative meaning for this question?”&lt;/P&gt;&lt;P&gt;Where OntoRank becomes important&lt;BR /&gt;Enterprises rarely have only one definition of a metric.&lt;/P&gt;&lt;P&gt;There may be multiple definitions of revenue, customer, active account, gross margin, or credit exposure—each created by different teams, at different times, for different purposes.&lt;/P&gt;&lt;P&gt;OntoRank is the PageRank-inspired authority-ranking concept associated with Genie Ontology.&lt;/P&gt;&lt;P&gt;Rather than ranking only by textual similarity, the context layer can consider signals such as:&lt;/P&gt;&lt;P&gt;Source authority and provenance&lt;/P&gt;&lt;P&gt;Asset certification&lt;/P&gt;&lt;P&gt;Frequency and breadth of usage&lt;/P&gt;&lt;P&gt;Relationships with other trusted assets&lt;/P&gt;&lt;P&gt;Freshness&lt;/P&gt;&lt;P&gt;Business relevance&lt;/P&gt;&lt;P&gt;User permissions&lt;/P&gt;&lt;P&gt;For example, imagine Genie discovers three definitions of “active customer”:&lt;/P&gt;&lt;P&gt;An old spreadsheet definition created three years ago&lt;/P&gt;&lt;P&gt;A frequently queried but uncertified reporting view&lt;/P&gt;&lt;P&gt;A certified Unity Catalog metric connected to the MDM golden customer and current amount balances&lt;/P&gt;&lt;P&gt;Keyword similarity alone might retrieve any of them.&lt;/P&gt;&lt;P&gt;An authority-aware approach should prioritize the certified, governed, fresh, and widely connected definition—while still enforcing the requesting user’s permissions.&lt;/P&gt;&lt;P&gt;Why this is bigger than better text-to-SQL&lt;BR /&gt;The real architectural shift is:&lt;/P&gt;&lt;P&gt;Metadata → Semantics → Context → Trusted action&lt;/P&gt;&lt;P&gt;A well-designed ontology can help an AI system understand:&lt;/P&gt;&lt;P&gt;Which source should be queried&lt;/P&gt;&lt;P&gt;Which metric definition should be applied&lt;/P&gt;&lt;P&gt;Which relationships and joins are valid&lt;/P&gt;&lt;P&gt;Which conflicting definition should take precedence&lt;/P&gt;&lt;P&gt;Which assets are deprecated&lt;/P&gt;&lt;P&gt;What the user is authorized to access&lt;/P&gt;&lt;P&gt;Why a particular source was used&lt;/P&gt;&lt;P&gt;This can make AI systems more accurate, explainable, reusable, and governance-aware.&lt;/P&gt;&lt;P&gt;But OntoRank does not eliminate data governance&lt;BR /&gt;Authority ranking is powerful, but popularity is not always correctness.&lt;/P&gt;&lt;P&gt;A widely used definition may still be outdated. A newly created certified data product may initially have little usage history. Poorly documented tables will continue to produce weak context.&lt;/P&gt;&lt;P&gt;Therefore, organizations should prepare the foundation:&lt;/P&gt;&lt;P&gt;Define important business terms&lt;/P&gt;&lt;P&gt;Create governed metric views&lt;/P&gt;&lt;P&gt;Certify authoritative data products&lt;/P&gt;&lt;P&gt;Deprecate obsolete assets&lt;/P&gt;&lt;P&gt;Maintain table and column descriptions&lt;/P&gt;&lt;P&gt;Capture lineage&lt;/P&gt;&lt;P&gt;Assign clear data ownership&lt;/P&gt;&lt;P&gt;Improve MDM and identity resolution&lt;/P&gt;&lt;P&gt;Test Genie answers against approved business scenarios&lt;/P&gt;&lt;P&gt;Genie Ontology can amplify a strong semantic and governance foundation—but it cannot magically repair an undefined business vocabulary.&lt;/P&gt;&lt;P&gt;My key takeaway&lt;BR /&gt;The next generation of enterprise AI will not be differentiated only by model size.&lt;/P&gt;&lt;P&gt;It will be differentiated by the quality of the context surrounding the model.&lt;/P&gt;&lt;P&gt;RAG helps AI find similar information.&lt;BR /&gt;Ontology helps AI understand relationships and meaning.&lt;BR /&gt;OntoRank helps AI decide what should be trusted.&lt;BR /&gt;Unity Catalog helps ensure that trust remains governed.&lt;/P&gt;&lt;P&gt;The most important question for data architects may soon change from:&lt;/P&gt;&lt;P&gt;“How do we expose our data to an AI agent?”&lt;/P&gt;&lt;P&gt;to:&lt;/P&gt;&lt;P&gt;“How do we make business meaning discoverable, authoritative, permission-aware, and machine-readable?”&lt;/P&gt;&lt;P&gt;I would love to hear from the Databricks Community:&lt;/P&gt;&lt;P&gt;How are you preparing your Unity Catalog metadata and metric views for Genie Ontology?&lt;/P&gt;&lt;P&gt;How should OntoRank balance popularity against formal certification?&lt;/P&gt;&lt;P&gt;Should users be able to inspect the ontology graph and understand why one definition outranked another?&lt;/P&gt;&lt;P&gt;What evaluation framework are you using to measure the business accuracy of Genie answers?&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Wed, 02 Sep 2026 00:59:50 GMT</pubDate>
      <guid>https://community.databricks.com/t5/get-started-discussions/from-semantic-similarity-to-business-authority-why-genie/m-p/167223#M12062</guid>
      <dc:creator>amitsharma1707</dc:creator>
      <dc:date>2026-09-02T00:59:50Z</dc:date>
    </item>
    <item>
      <title>Operationalizing a Live Genie Space: Benchmarking, Governance &amp; Continuous Improvement</title>
      <link>https://community.databricks.com/t5/genie-hub/operationalizing-a-live-genie-space-benchmarking-governance-amp/m-p/167186#M44</link>
      <description>&lt;P class="lia-align-justify" data-unlink="true"&gt;&lt;A href="https://community.databricks.com/t5/genie-hub/making-databricks-genie-spaces-actually-work-a-practical/td-p/167171" target="_blank" rel="noopener"&gt;In my previous article&lt;/A&gt;, I wrote an article about building a strong foundation for Databricks Genie Spaces through data modeling, metadata, and semantic design. This time, I want to share a few lessons from running a Genie Space against a real healthcare analytics use case.&lt;/P&gt;&lt;P class="lia-align-center"&gt;&lt;FONT color="#0000FF"&gt;&lt;U&gt;&lt;EM&gt;The biggest surprise wasn't AI&lt;/EM&gt;&lt;/U&gt;&lt;/FONT&gt;. &lt;FONT color="#339966"&gt;&lt;U&gt;&lt;EM&gt;It was how quickly Genie exposed problems that already existed in the data&lt;/EM&gt;&lt;/U&gt;&lt;/FONT&gt;.&lt;/P&gt;&lt;P class="lia-align-justify"&gt;When we first deployed the Genie Space and ran it against our benchmark questions, the results were far from perfect. The out-of-the-box configuration answered only 6 out of 20 benchmark questions correctly. Even after multiple rounds of tuning, a &lt;U&gt;repurposed&amp;nbsp;legacy analytics dataset for BI workloads&lt;/U&gt; never achieved more than 75% benchmark accuracy.&amp;nbsp;&lt;/P&gt;&lt;P class="lia-align-center"&gt;&lt;FONT color="#0000FF"&gt;&lt;U&gt;&lt;EM&gt;What changed everything wasn't prompt engineering&lt;/EM&gt;&lt;/U&gt;&lt;/FONT&gt;. &lt;FONT color="#339966"&gt;&lt;U&gt;&lt;EM&gt;It was dataset design&lt;/EM&gt;&lt;/U&gt;&lt;/FONT&gt;.&lt;/P&gt;&lt;P class="lia-align-justify"&gt;&lt;FONT color="#0000FF"&gt;&lt;EM&gt;Once we defined and built a live delta dataset with pre-defined metrics specifically for Genie&lt;/EM&gt;&lt;/FONT&gt;, removed irrelevant columns, simplified business logic, standardized categorical values, and aligned the schema to the questions users were actually asking, benchmark accuracy eventually reached &lt;STRONG&gt;&lt;FONT color="#008080"&gt;100%&lt;/FONT&gt;&lt;/STRONG&gt;.&lt;/P&gt;&lt;OL class="lia-align-justify"&gt;&lt;LI&gt;While many teams may focus on configuring Genie and fine-tuning prompt. For our team, the biggest boost in accuracy came from preparing the underlying data for Genie.&lt;/LI&gt;&lt;LI&gt;Another lesson was the importance of benchmarking. Without benchmarks, every discussion becomes subjective.&amp;nbsp;&lt;UL&gt;&lt;LI&gt;&lt;STRONG&gt;Quick Tip&lt;/STRONG&gt;: &lt;FONT color="#0000FF"&gt;Users may ask the same question in different ways. Databricks recommends using 2-4 different phrasings of the same question (but the same SQL code) to fully assess accuracy&lt;/FONT&gt;.&lt;/LI&gt;&lt;/UL&gt;&lt;/LI&gt;&lt;/OL&gt;&lt;P class="lia-align-justify"&gt;We also learned that not every configuration change improves accuracy. Below is a visual of our findings;&lt;/P&gt;&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="screen shot1.png" style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30587iD91B9E0D391CC07B/image-size/large?v=v2&amp;amp;px=999" role="button" title="screen shot1.png" alt="screen shot1.png" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;&lt;SPAN&gt;Table descriptions helped significantly.&lt;/SPAN&gt;&lt;/LI&gt;&lt;LI&gt;&lt;SPAN&gt;General instructions provided some improvement. &lt;/SPAN&gt;&lt;/LI&gt;&lt;LI&gt;&lt;SPAN&gt;Surprisingly, adding large numbers of synonyms wasn't always beneficial and sometimes introduced additional ambiguity.&lt;/SPAN&gt;&lt;/LI&gt;&lt;LI&gt;&lt;FONT color="#339966"&gt;The most effective configuration element by far was SQL examples&lt;/FONT&gt;.&lt;UL&gt;&lt;LI&gt;&lt;STRONG&gt;Quick Tip&lt;/STRONG&gt;: &lt;FONT color="#0000FF"&gt;Whenever a business rule involved non-trivial logic, such as event-level calculations, eligibility definitions, or classifications, SQL examples consistently outperformed additional instructions. Genie seemed to learn far more effectively from concrete examples than from lengthy explanations.&lt;/FONT&gt;&lt;/LI&gt;&lt;/UL&gt;&lt;/LI&gt;&lt;LI&gt;Another important observation was that Genie is not entirely deterministic. The same question can occasionally be presented differently depending on formatting choices, percentage calculations, null handling, or interpretation of date ranges. We found ourselves spending as much time standardizing outputs as improving answers.&lt;/LI&gt;&lt;/UL&gt;&lt;H3&gt;This is where governance becomes essential&lt;/H3&gt;&lt;P class="lia-align-justify"&gt;A successful Genie Space needs more than good metadata. It needs agreed definitions, controlled business logic, benchmark testing, and a clear process for evaluating changes. Otherwise, the results erodes trust.&lt;span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="screen shot2.png" style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30590iE80C25AF2E2E7BA6/image-size/large?v=v2&amp;amp;px=999" role="button" title="screen shot2.png" alt="screen shot2.png" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;P class="lia-align-justify"&gt;The teams that will get the most value from Genie Spaces and turn a demo into a production-ready analytical product won't be the ones writing the most sophisticated prompts. They'll be the ones investing in benchmark-driven development, semantic consistency, and Genie Tailored purpose-built datasets.&lt;/P&gt;</description>
      <pubDate>Tue, 01 Sep 2026 20:25:33 GMT</pubDate>
      <guid>https://community.databricks.com/t5/genie-hub/operationalizing-a-live-genie-space-benchmarking-governance-amp/m-p/167186#M44</guid>
      <dc:creator>Salman_Ahmed</dc:creator>
      <dc:date>2026-09-01T20:25:33Z</dc:date>
    </item>
    <item>
      <title>Partitioning vs Liquid Clustering (per-table):</title>
      <link>https://community.databricks.com/t5/data-engineering/partitioning-vs-liquid-clustering-per-table/m-p/167180#M55652</link>
      <description>&lt;P&gt;Can PARTITION BY and CLUSTER BY (Liquid Clustering) be used simultaneously on the same table? If we use only PARTITION BY, is there a negative performance impact on materialized-view refreshes in Silver/Gold? Since materialized views read only incremental delta files, is Liquid Clustering redundant in this scenario, or does it still provide benefits for file compaction and read optimization?&lt;/P&gt;</description>
      <pubDate>Tue, 01 Sep 2026 20:16:23 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/partitioning-vs-liquid-clustering-per-table/m-p/167180#M55652</guid>
      <dc:creator>AshokB</dc:creator>
      <dc:date>2026-09-01T20:16:23Z</dc:date>
    </item>
    <item>
      <title>Making Databricks Genie Spaces Actually Work: A Practical Framework for Client and Data Teams</title>
      <link>https://community.databricks.com/t5/genie-hub/making-databricks-genie-spaces-actually-work-a-practical/m-p/167171#M43</link>
      <description>&lt;DIV&gt;&lt;H2&gt;Introduction&lt;/H2&gt;&lt;P&gt;Over the past year, I've had quite a few conversations with teams exploring Databricks Genie Spaces. The pattern is usually the same. Someone sees a demo, watches a business user ask a question in plain English, and within seconds Genie returns a chart, a SQL query, and what appears to be a perfectly reasonable answer.&lt;/P&gt;&lt;P&gt;The reaction is almost always immediate.&lt;/P&gt;&lt;P&gt;&lt;EM&gt;"This could completely change how people use data and BI reports."&lt;/EM&gt;&lt;/P&gt;&lt;P&gt;For years we've built dashboards, semantic models, reporting layers, and self-service analytics platforms with the goal of helping business users answer questions faster. Genie feels like the natural next step in that journey. Instead of learning SQL or navigating dozens of dashboards, users can simply ask a question and interact with data conversationally.&lt;/P&gt;&lt;P&gt;The technology itself is impressive. But after working on enterprise data platforms for many years, the challenge is rarely the AI model itself.&lt;/P&gt;&lt;P&gt;&lt;U&gt;&lt;EM&gt;The challenge is trust&lt;/EM&gt;&lt;/U&gt;.&lt;/P&gt;&lt;P&gt;Can users trust the answer? Can analysts reproduce it? Can data teams explain it? And perhaps most importantly, will different users receive consistent answers to the same question?&lt;/P&gt;&lt;P&gt;Those questions have far less to do with the language model and much more to do with the foundation underneath it. That's why whenever I'm asked how to improve Genie Spaces.&lt;/P&gt;&lt;P&gt;&lt;EM&gt;&lt;U&gt;I rarely start by talking about prompts,&lt;/U&gt; &lt;U&gt;I start by talking about data&lt;/U&gt;&lt;/EM&gt;.&lt;/P&gt;&lt;HR /&gt;&lt;H3&gt;Why Most Genie Projects Fail Before Users Ask Their First Question&lt;/H3&gt;&lt;P class="lia-align-justify"&gt;&lt;FONT size="3"&gt;When teams evaluate Genie Spaces, their first instinct is often to improve prompts or add more instructions.&lt;/FONT&gt;&lt;/P&gt;&lt;P class="lia-align-justify"&gt;&lt;FONT size="3"&gt;In my experience, that's usually the wrong starting point. Most quality issues originate from one of five areas:&lt;/FONT&gt;&lt;/P&gt;&lt;OL class="lia-align-justify"&gt;&lt;LI&gt;&lt;FONT size="3"&gt;Weak data modeling (Data Engineering)&lt;/FONT&gt;&lt;/LI&gt;&lt;LI&gt;&lt;FONT size="3"&gt;Poor metadata quality (Data Scientist/Business Analyst/Data Analyst)&lt;/FONT&gt;&lt;/LI&gt;&lt;LI&gt;&lt;FONT size="3"&gt;Undefined business metrics (Data Scientist)&lt;/FONT&gt;&lt;/LI&gt;&lt;LI&gt;&lt;FONT size="3"&gt;Missing table relationships (Data Engineering)&lt;/FONT&gt;&lt;/LI&gt;&lt;LI&gt;&lt;FONT size="3"&gt;Lack of benchmark testing (ML Engineer)&lt;/FONT&gt;&lt;/LI&gt;&lt;/OL&gt;&lt;P class="lia-align-justify"&gt;&lt;FONT size="3"&gt;Databricks has been investing heavily in a governed semantic foundation through Unity Catalog Semantics, Genie Ontology, Live Tables, including metric views, domains, governed business definitions, and AI-aware context management. These capabilities help ensure that both humans and AI systems interpret data consistently.&lt;/FONT&gt;&lt;/P&gt;&lt;DIV&gt;&lt;HR /&gt;&lt;H2&gt;&lt;U&gt;&lt;SPAN&gt;Step-by-step plan&lt;/SPAN&gt;&lt;/U&gt;&lt;/H2&gt;&lt;/DIV&gt;&lt;/DIV&gt;&lt;H2&gt;&lt;U&gt;Step 1: Build the Data Foundation Before Building the Genie Space&lt;/U&gt;&lt;/H2&gt;&lt;P&gt;&lt;FONT size="3"&gt;&lt;U&gt;&lt;EM&gt;The single most important success factor is the quality of the curated data layer&lt;/EM&gt;&lt;/U&gt;.&lt;/FONT&gt;&lt;/P&gt;&lt;P&gt;&lt;FONT size="3"&gt;Many data teams expose highly normalized source models and expect Genie to figure out the relationships. While technically possible, this often introduces ambiguity.&amp;nbsp;&lt;/FONT&gt;&lt;FONT size="3"&gt;Instead, design datasets specifically for consumption.&lt;/FONT&gt;&lt;/P&gt;&lt;H3&gt;Recommended Design Approach&lt;/H3&gt;&lt;H4&gt;1. Denormalize Where Appropriate&lt;/H4&gt;&lt;DIV&gt;Rather than expecting Genie to navigate a maze of joins every time a user asks a question, it's worth investing in &lt;U&gt;curated business-ready delta tables&lt;/U&gt;. If answering a simple revenue question requires six or seven tables to be stitched together, the chances of selecting an incorrect relationship increase significantly. In most successful implementations I've seen, common dimensions are already joined, business entities are standardised, and duplicate relationship paths have been removed long before the data reaches Genie.&lt;/DIV&gt;&lt;HR /&gt;&lt;H4&gt;2. Pre-Calculate Common Business Logic&lt;/H4&gt;&lt;DIV&gt;A common mistake is treating Genie as the place where business logic should be assembled. In reality, repetitive calculations and classifications belong in the data layer. Whether it's reporting periods, fiscal calendars, active customer definitions, or product lifecycle states, these concepts should already exist in a governed and reusable form. This allows Genie to focus on answering the question rather than reconstructing business logic every time.&lt;/DIV&gt;&lt;HR /&gt;&lt;H4&gt;3. Establish Canonical Metrics&lt;/H4&gt;&lt;P&gt;&lt;FONT size="3"&gt;One of the strongest capabilities available through Unity Catalog is the ability to define reusable metrics and semantic objects that provide consistent business logic across analytics workloads and AI consumers.&lt;/FONT&gt;&lt;/P&gt;&lt;P&gt;&lt;FONT size="3"&gt;For example:&lt;/FONT&gt;&lt;/P&gt;&lt;DIV class=""&gt;&lt;DIV class=""&gt;&lt;DIV&gt;&lt;DIV class=""&gt;&lt;DIV class=""&gt;&lt;DIV class=""&gt;&amp;nbsp;&lt;/DIV&gt;&lt;/DIV&gt;&lt;/DIV&gt;&lt;/DIV&gt;&lt;/DIV&gt;&lt;/DIV&gt;&lt;LI-CODE lang="markup"&gt;measures:
  - name: Total Revenue
    expr: SUM(purchase_amount)
           FILTER (WHERE status='approved')
    comment: Revenue from approved transactions
    display_name: Total Revenue
    synonyms:
      - revenue
      - sales
      - total sales
      - approved revenue&lt;/LI-CODE&gt;&lt;DIV class=""&gt;&lt;DIV class=""&gt;&lt;DIV class=""&gt;&lt;DIV class=""&gt;&lt;SPAN&gt;This ensures that every user asking about revenue receives answers based on the same calculation.&lt;/SPAN&gt;&lt;/DIV&gt;&lt;/DIV&gt;&lt;/DIV&gt;&lt;/DIV&gt;&lt;HR /&gt;&lt;H2&gt;&lt;U&gt;Step 2: Treat Genie Like Software and Create Benchmarks&lt;/U&gt;&lt;/H2&gt;&lt;DIV&gt;I've been in sessions where a team asks Genie three questions, gets two correct answers, one questionable result, and immediately starts debating whether the prompt needs to be rewritten. The reality is that this kind of testing is far too subjective. Without a defined set of benchmark questions and expected outcomes, it's almost impossible to measure quality in a meaningful way. That's why it is best to establish a benchmark suite early, before wider adoption begins.&lt;/DIV&gt;&lt;HR /&gt;&lt;H3&gt;Create a Question Inventory&lt;/H3&gt;&lt;DIV&gt;The best benchmark questions usually come directly from the people who use the data every day. Spend time with business stakeholders, analysts, and subject matter experts to understand the questions they regularly ask, whether that's tracking KPIs, understanding trends, explaining variances, or preparing executive reporting. Once you've collected those questions, document what a correct answer looks like. That includes not only the expected result, but also the level of aggregation, any business filters that should be applied, and how the answer should be presented. The goal isn't simply to test whether Genie returns an answer. It's to verify that the answer aligns with how the business expects the question to be interpreted.&lt;/DIV&gt;&lt;HR /&gt;&lt;H3&gt;Build a Regression Test Suite&lt;/H3&gt;&lt;DIV&gt;&lt;U&gt;&lt;EM&gt;A Genie Space is never really finished&lt;/EM&gt;&lt;/U&gt;. The underlying data platform keeps evolving, new business requirements appear, and teams continuously refine their definitions and metrics. While those changes are important, they also introduce risk. I've found that the most successful teams maintain a set of benchmark questions that are executed regularly, especially after major updates. It provides a simple but effective way of confirming that answers users already trust continue to behave as expected, even as the platform grows and changes around them.&lt;/DIV&gt;&lt;HR /&gt;&lt;H2&gt;&lt;U&gt;Step 3: Teach Genie How Your Business Thinks&lt;/U&gt;&lt;/H2&gt;&lt;DIV&gt;Metadata is what helps bridge the gap in business thinking in natural flow and&amp;nbsp;&lt;SPAN&gt;tables, columns, or schemas&lt;/SPAN&gt;. The richer the business context around your data, the easier it becomes for Genie to understand what the user is really asking and translate that intent into a query that makes sense. In many cases, improving metadata delivers a bigger uplift in answer quality than yet another round of prompt tuning.&lt;/DIV&gt;&lt;HR /&gt;&lt;H3&gt;Table Descriptions Matter&lt;/H3&gt;&lt;P&gt;Avoid generic descriptions like:&lt;/P&gt;&lt;BLOCKQUOTE&gt;&lt;P&gt;Customer transaction table&lt;/P&gt;&lt;/BLOCKQUOTE&gt;&lt;P&gt;Instead use:&lt;/P&gt;&lt;BLOCKQUOTE&gt;&lt;P&gt;Contains finalized customer purchase records used for revenue reporting and financial performance analysis.&lt;/P&gt;&lt;/BLOCKQUOTE&gt;&lt;P&gt;The second description provides significantly more business context.&lt;/P&gt;&lt;HR /&gt;&lt;H2&gt;Define Synonyms Explicitly&lt;/H2&gt;&lt;P&gt;&lt;FONT size="3"&gt;Business users rarely use technical column names.&lt;/FONT&gt;&lt;/P&gt;&lt;P&gt;&lt;FONT size="3"&gt;For example:&lt;/FONT&gt;&lt;/P&gt;&lt;DIV class=""&gt;&lt;DIV class=""&gt;Business Term Actual Field &lt;TABLE&gt;&lt;TBODY&gt;&lt;TR&gt;&lt;TD width="121.167px" height="30px"&gt;&lt;FONT size="3"&gt;Sales&lt;/FONT&gt;&lt;/TD&gt;&lt;TD width="203.014px" height="30px"&gt;&lt;FONT size="3"&gt;Revenue&lt;/FONT&gt;&lt;/TD&gt;&lt;/TR&gt;&lt;TR&gt;&lt;TD width="121.167px" height="30px"&gt;&lt;FONT size="3"&gt;ARR&lt;/FONT&gt;&lt;/TD&gt;&lt;TD width="203.014px" height="30px"&gt;&lt;FONT size="3"&gt;Annual Recurring Revenue&lt;/FONT&gt;&lt;/TD&gt;&lt;/TR&gt;&lt;TR&gt;&lt;TD width="121.167px" height="30px"&gt;&lt;FONT size="3"&gt;Customer Base&lt;/FONT&gt;&lt;/TD&gt;&lt;TD width="203.014px" height="30px"&gt;&lt;FONT size="3"&gt;Active Customers&lt;/FONT&gt;&lt;/TD&gt;&lt;/TR&gt;&lt;TR&gt;&lt;TD width="121.167px" height="30px"&gt;&lt;FONT size="3"&gt;Gross Sales&lt;/FONT&gt;&lt;/TD&gt;&lt;TD width="203.014px" height="30px"&gt;&lt;FONT size="3"&gt;Invoice Amount&lt;/FONT&gt;&lt;/TD&gt;&lt;/TR&gt;&lt;/TBODY&gt;&lt;/TABLE&gt;&lt;/DIV&gt;&lt;/DIV&gt;&lt;P&gt;&lt;FONT size="3"&gt;&lt;EM&gt;Providing synonyms dramatically improves question interpretation&lt;/EM&gt;.&lt;/FONT&gt;&lt;/P&gt;&lt;HR /&gt;&lt;H3&gt;Document Relationships&lt;/H3&gt;&lt;DIV&gt;Another area that often gets overlooked is the way datasets relate to one another. In most enterprises, the same business entity appears across multiple tables, and there can be several possible paths between them. If those relationships aren't clearly defined, Genie may have to infer how the data is connected, which can lead to unexpected results.&lt;/DIV&gt;&lt;DIV&gt;Explicitly documenting relationships and validating the business meaning behind them significantly improves consistency. It's not enough to know that two tables can be joined; Genie also needs to understand how they should be joined and what business context that relationship represents.&lt;/DIV&gt;&lt;P&gt;&lt;EM&gt;&lt;FONT size="3" color="#FF0000"&gt;Incorrect joins are a major source of AI-generated analytical errors.&lt;/FONT&gt;&lt;/EM&gt;&lt;/P&gt;&lt;HR /&gt;&lt;H3&gt;Supply Example SQL&lt;/H3&gt;&lt;P&gt;&lt;STRONG&gt;&lt;FONT size="3"&gt;One of the most effective yet underutilized techniques is maintaining a library of gold-standard SQL&lt;/FONT&gt;&lt;/STRONG&gt;.&lt;/P&gt;&lt;P&gt;Example:&lt;/P&gt;&lt;DIV class=""&gt;&lt;DIV class=""&gt;&lt;DIV&gt;&lt;DIV class=""&gt;&lt;DIV class=""&gt;&lt;DIV class=""&gt;&amp;nbsp;&lt;/DIV&gt;&lt;/DIV&gt;&lt;/DIV&gt;&lt;/DIV&gt;&lt;/DIV&gt;&lt;/DIV&gt;&lt;LI-CODE lang="markup"&gt;SELECT
    fiscal_year,
    SUM(revenue) AS total_revenue
FROM sales_gold
GROUP BY fiscal_year
ORDER BY fiscal_year;&lt;/LI-CODE&gt;&lt;P&gt;&lt;EM&gt;These examples act as patterns that help Genie generate more reliable queries&lt;/EM&gt;.&lt;/P&gt;&lt;HR /&gt;&lt;H2&gt;Use General Instructions Sparingly&lt;/H2&gt;&lt;P&gt;&lt;FONT size="3"&gt;Many teams attempt to solve every issue through lengthy instructions.&lt;/FONT&gt;&lt;/P&gt;&lt;P&gt;&lt;FONT size="3"&gt;This typically creates maintenance problems.&lt;/FONT&gt;&lt;/P&gt;&lt;P&gt;&lt;FONT size="3"&gt;A simpler decision framework is:&lt;/FONT&gt;&lt;/P&gt;&lt;DIV class=""&gt;&lt;DIV class=""&gt;Problem Fix Location &lt;TABLE&gt;&lt;TBODY&gt;&lt;TR&gt;&lt;TD width="166.021px" height="30px"&gt;&lt;FONT size="3"&gt;Wrong table&lt;/FONT&gt;&lt;/TD&gt;&lt;TD width="167.219px" height="30px"&gt;&lt;FONT size="3"&gt;Table metadata&lt;/FONT&gt;&lt;/TD&gt;&lt;/TR&gt;&lt;TR&gt;&lt;TD width="166.021px" height="30px"&gt;&lt;FONT size="3"&gt;Wrong column&lt;/FONT&gt;&lt;/TD&gt;&lt;TD width="167.219px" height="30px"&gt;&lt;FONT size="3"&gt;Column description&lt;/FONT&gt;&lt;/TD&gt;&lt;/TR&gt;&lt;TR&gt;&lt;TD width="166.021px" height="30px"&gt;&lt;FONT size="3"&gt;Wrong value mapping&lt;/FONT&gt;&lt;/TD&gt;&lt;TD width="167.219px" height="30px"&gt;&lt;FONT size="3"&gt;Example values&lt;/FONT&gt;&lt;/TD&gt;&lt;/TR&gt;&lt;TR&gt;&lt;TD width="166.021px" height="30px"&gt;&lt;FONT size="3"&gt;Wrong join&lt;/FONT&gt;&lt;/TD&gt;&lt;TD width="167.219px" height="30px"&gt;&lt;FONT size="3"&gt;Relationship definition&lt;/FONT&gt;&lt;/TD&gt;&lt;/TR&gt;&lt;TR&gt;&lt;TD width="166.021px" height="30px"&gt;Wrong calculation&lt;/TD&gt;&lt;TD width="167.219px" height="30px"&gt;Example SQL&lt;/TD&gt;&lt;/TR&gt;&lt;/TBODY&gt;&lt;/TABLE&gt;&lt;/DIV&gt;&lt;/DIV&gt;&lt;P&gt;Use narrative instructions only for business context.&lt;/P&gt;&lt;HR /&gt;&lt;H2&gt;Key Takeaways&lt;/H2&gt;&lt;P&gt;&lt;FONT size="3"&gt;Organizations often assume conversational analytics starts with AI.&lt;/FONT&gt;&lt;/P&gt;&lt;P&gt;&lt;FONT size="3"&gt;In reality, it starts with data engineering.&lt;/FONT&gt;&lt;/P&gt;&lt;P&gt;&lt;FONT size="3"&gt;Before focusing on prompts, invest in:&lt;/FONT&gt;&lt;/P&gt;&lt;OL&gt;&lt;LI&gt;&lt;FONT size="3"&gt;Curated Gold datasets&lt;/FONT&gt;&lt;/LI&gt;&lt;LI&gt;&lt;FONT size="3"&gt;Metric definitions&lt;/FONT&gt;&lt;/LI&gt;&lt;LI&gt;&lt;FONT size="3"&gt;Rich metadata&lt;/FONT&gt;&lt;/LI&gt;&lt;LI&gt;&lt;FONT size="3"&gt;Relationship modeling&lt;/FONT&gt;&lt;/LI&gt;&lt;LI&gt;&lt;FONT size="3"&gt;Benchmark testing&lt;/FONT&gt;&lt;/LI&gt;&lt;/OL&gt;&lt;P&gt;&lt;FONT size="3"&gt;Genie Spaces are most successful when they are grounded in governed business semantics rather than isolated prompt instructions. Databricks' broader investment in Unity Catalog Semantics reflects this exact direction, creating trusted business context that can be reused across analytics and AI experiences.&lt;BR /&gt;&lt;BR /&gt;&lt;/FONT&gt;&lt;/P&gt;&lt;DIV&gt;&lt;H3&gt;2-Part Series&lt;/H3&gt;&lt;P&gt;&lt;STRONG&gt;Part 1:&lt;/STRONG&gt; Making Databricks Genie Spaces Actually Work: A Practical Framework for Client and Data Teams&lt;BR /&gt;&lt;STRONG&gt;Part 2:&lt;/STRONG&gt; Operationalizing a Live Genie Spaces with Benchmarking, Governance, and Continuous Improvement&lt;/P&gt;&lt;/DIV&gt;</description>
      <pubDate>Tue, 01 Sep 2026 18:42:28 GMT</pubDate>
      <guid>https://community.databricks.com/t5/genie-hub/making-databricks-genie-spaces-actually-work-a-practical/m-p/167171#M43</guid>
      <dc:creator>Salman_Ahmed</dc:creator>
      <dc:date>2026-09-01T18:42:28Z</dc:date>
    </item>
    <item>
      <title>CUSTOMER STORY | Kraken governs utility data at scale with Unity Catalog</title>
      <link>https://community.databricks.com/t5/announcements/customer-story-kraken-governs-utility-data-at-scale-with-unity/m-p/167167#M1032</link>
      <description>&lt;P&gt;&lt;SPAN&gt;&lt;EM&gt;"Unity Catalog allows us to deliver data in the safest and most compliant way for our customers. Given the liability involved, that is not something we take lightly."&amp;nbsp;&lt;/EM&gt; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp;&lt;STRONG&gt;-&amp;nbsp;&lt;/STRONG&gt;&lt;/SPAN&gt;&lt;STRONG&gt;Javi Asensio, Head of Data and Analytics Engineering, Kraken&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Kraken&lt;/STRONG&gt;&lt;SPAN&gt;, the operating system for utilities, supports approximately &lt;/SPAN&gt;&lt;STRONG&gt;85 million contracted accounts&lt;/STRONG&gt;&lt;SPAN&gt; across 13 countries. With &lt;/SPAN&gt;&lt;STRONG&gt;Unity Catalog and Delta Sharing&lt;/STRONG&gt;&lt;SPAN&gt;, Kraken is building a governed data foundation that helps utility providers access data securely, while giving internal teams better visibility across a complex, multi-account environment.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;FONT size="4"&gt;&lt;STRONG&gt;Key highlights:&lt;/STRONG&gt;&lt;/FONT&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;95% of governance needs met natively:&lt;/STRONG&gt;&lt;SPAN&gt; System tables and the Unity Catalog API provide visibility into access, queries, endpoints, and resource usage without relying on third-party tooling.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Data sharing in 1 day instead of 1.5 weeks:&lt;/STRONG&gt;&lt;SPAN&gt; Delta Sharing helps Kraken make data available to clients faster, especially for customers already using Databricks.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Secure, fine-grained access:&lt;/STRONG&gt;&lt;SPAN&gt; Account isolation, data masking, and role-based controls help protect sensitive personal, financial, and smart-meter data across regulated markets.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;One governance layer at scale:&lt;/STRONG&gt;&lt;SPAN&gt; Unity Catalog helps Kraken manage access, auditability, and compliance across dedicated Databricks environments for each client.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P class="p8i6j01 paragraph"&gt;&lt;A style="background-color: #ff3621; color: white; padding: 10px 20px; text-decoration: none; border-radius: 5px; font-weight: bold; display: inline-block;" href="https://www.databricks.com/customers/kraken/utility-data-unity-catalog?itm_source=www&amp;amp;itm_category=customers&amp;amp;itm_page=zerobus-ingest&amp;amp;itm_location=body&amp;amp;itm_component=general-asset-card&amp;amp;itm_offer=utility-data-unity-catalog" target="_blank" rel="noopener"&gt; &lt;span class="lia-unicode-emoji" title=":link:"&gt;🔗&lt;/span&gt; Check out the full story &lt;span class="lia-unicode-emoji" title=":backhand_index_pointing_left:"&gt;👈&lt;/span&gt;&lt;/A&gt;&lt;/P&gt;</description>
      <pubDate>Tue, 01 Sep 2026 17:04:24 GMT</pubDate>
      <guid>https://community.databricks.com/t5/announcements/customer-story-kraken-governs-utility-data-at-scale-with-unity/m-p/167167#M1032</guid>
      <dc:creator>Tushar_Parekar</dc:creator>
      <dc:date>2026-09-01T17:04:24Z</dc:date>
    </item>
    <item>
      <title>Webassesor last name change</title>
      <link>https://community.databricks.com/t5/certifications/webassesor-last-name-change/m-p/167164#M4888</link>
      <description>&lt;P&gt;Hi, anyone knows how to contact databricks support? Or how long it takes to get any response from them?&lt;/P&gt;&lt;P&gt;I need to change my last name in webassesor, can't do this my self.&amp;nbsp;&lt;/P&gt;&lt;P&gt;Already had to move my exam 4 times becasue support is not responsive. I've tried both... an email and the form..&lt;/P&gt;</description>
      <pubDate>Tue, 01 Sep 2026 17:00:12 GMT</pubDate>
      <guid>https://community.databricks.com/t5/certifications/webassesor-last-name-change/m-p/167164#M4888</guid>
      <dc:creator>Ds_1_0</dc:creator>
      <dc:date>2026-09-01T17:00:12Z</dc:date>
    </item>
    <item>
      <title>Como aprender correctamente en el entorno de DataBricks y aprovecharlo al 100%</title>
      <link>https://community.databricks.com/t5/get-started-discussions/como-aprender-correctamente-en-el-entorno-de-databricks-y/m-p/167163#M12061</link>
      <description>&lt;P&gt;Hola a todos, soy ingeneriero en sistemas y deseo aprender a utilizar esta grandiosa herramienta para poder sacarle el mejor provecho posible.&lt;/P&gt;</description>
      <pubDate>Tue, 01 Sep 2026 16:59:48 GMT</pubDate>
      <guid>https://community.databricks.com/t5/get-started-discussions/como-aprender-correctamente-en-el-entorno-de-databricks-y/m-p/167163#M12061</guid>
      <dc:creator>maycol25</dc:creator>
      <dc:date>2026-09-01T16:59:48Z</dc:date>
    </item>
  </channel>
</rss>

