<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>New blog articles in Databricks Community</title>
    <link>https://community.databricks.com/</link>
    <description>Databricks Community</description>
    <pubDate>Sun, 11 Oct 2026 01:13:34 GMT</pubDate>
    <dc:creator>Community</dc:creator>
    <dc:date>2026-10-11T01:13:34Z</dc:date>
    <item>
      <title>Deploy via Databricks Express Setup</title>
      <link>https://community.databricks.com/t5/databricks-tv/deploy-via-databricks-express-setup/ba-p/172514</link>
      <description>&lt;P&gt;&lt;EM&gt;Presented by Yotaro Enomoto&lt;/EM&gt;&lt;/P&gt;
&lt;P&gt;&lt;IFRAME width="560" height="315" src="https://www.youtube-nocookie.com/embed/fgdp6I1SEIQ" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen=""&gt;&lt;/IFRAME&gt;&lt;/P&gt;</description>
      <pubDate>Fri, 09 Oct 2026 18:44:46 GMT</pubDate>
      <guid>https://community.databricks.com/t5/databricks-tv/deploy-via-databricks-express-setup/ba-p/172514</guid>
      <dc:creator>DSBDatabricks</dc:creator>
      <dc:date>2026-10-09T18:44:46Z</dc:date>
    </item>
    <item>
      <title>Databricks AWS Workspace Manual Deployment</title>
      <link>https://community.databricks.com/t5/databricks-tv/databricks-aws-workspace-manual-deployment/ba-p/172505</link>
      <description>This video provides a step-bystep guide on how to manually #deploy a Databricks workspace within an #AWS VPC. The tutorial concludes with the successful deployment of the #workspace.</description>
      <pubDate>Fri, 09 Oct 2026 17:50:48 GMT</pubDate>
      <guid>https://community.databricks.com/t5/databricks-tv/databricks-aws-workspace-manual-deployment/ba-p/172505</guid>
      <dc:creator>DSBDatabricks</dc:creator>
      <dc:date>2026-10-09T17:50:48Z</dc:date>
    </item>
    <item>
      <title>Personalizing Genie Code with MCPs and User Skills</title>
      <link>https://community.databricks.com/t5/databricks-tv/personalizing-genie-code-with-mcps-and-user-skills/ba-p/172495</link>
      <description>Learn how to personalize your GenieCode experience using instructions, skills, and MCP servers in this Databricks walkthrough.</description>
      <pubDate>Fri, 09 Oct 2026 16:55:48 GMT</pubDate>
      <guid>https://community.databricks.com/t5/databricks-tv/personalizing-genie-code-with-mcps-and-user-skills/ba-p/172495</guid>
      <dc:creator>DSBDatabricks</dc:creator>
      <dc:date>2026-10-09T16:55:48Z</dc:date>
    </item>
    <item>
      <title>Finally, a Simple Way to Deploy Databricks Hybrid Workspaces</title>
      <link>https://community.databricks.com/t5/databricks-tv/finally-a-simple-way-to-deploy-databricks-hybrid-workspaces/ba-p/172486</link>
      <description>Serverless Workspaces are the ideal choice for getting started to Databricks because Databricks manages the compute an securely stores your workspace data, with nothing to create in your cloud account.
However, if you have a reason to deploy Databricks on your own cloud account- also known as deploying a Hybrid Workspace- you will notice that there are quite some pre-requisites that you must complete, such as creating a VPC, a default Storage, IAM Roles, etc.
In this video, Pedro Zanlorensi will introduce you to the Platform Kit, a Web UI that you can use to deploy your Hybrid Workspaces with just a few clicks.</description>
      <pubDate>Fri, 09 Oct 2026 15:57:55 GMT</pubDate>
      <guid>https://community.databricks.com/t5/databricks-tv/finally-a-simple-way-to-deploy-databricks-hybrid-workspaces/ba-p/172486</guid>
      <dc:creator>DSBDatabricks</dc:creator>
      <dc:date>2026-10-09T15:57:55Z</dc:date>
    </item>
    <item>
      <title>Databricks and SAP: Building the Next Generation of Enterprise Data and Analytics</title>
      <link>https://community.databricks.com/t5/databricks-tv/databricks-and-sap-building-the-next-generation-of-enterprise/ba-p/172476</link>
      <description>Join experts for an update on the Databricks–SAP partnership and how SAP Business Data Cloud can help organizations modernize their data and analytics architecture. We’ll explore the latest platform capabilities, roadmap direction, and the considerations that shape SAP-native versus Databricks-based approaches.</description>
      <pubDate>Fri, 09 Oct 2026 17:03:02 GMT</pubDate>
      <guid>https://community.databricks.com/t5/databricks-tv/databricks-and-sap-building-the-next-generation-of-enterprise/ba-p/172476</guid>
      <dc:creator>DSBDatabricks</dc:creator>
      <dc:date>2026-10-09T17:03:02Z</dc:date>
    </item>
    <item>
      <title>Bricktalk Recording | Databricks Certification - Did You Know?, What's New, and Where We're Going</title>
      <link>https://community.databricks.com/t5/bricktalks-tv/bricktalk-recording-databricks-certification-did-you-know-what-s/ba-p/172426</link>
      <description>&lt;P data-unlink="true"&gt;James Kantor (Community username:&amp;nbsp;&lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/91222"&gt;@Cert-Bricks&lt;/a&gt;)&amp;nbsp;&amp;nbsp;explores the latest updates in Databricks Certifications.&lt;/P&gt;
&lt;P&gt;&lt;div class="video-embed-center video-embed"&gt;&lt;iframe class="embedly-embed" src="https://cdn.embedly.com/widgets/media.html?src=https%3A%2F%2Fwww.youtube.com%2Fembed%2Fo2il5eR0e5M%3Ffeature%3Doembed&amp;amp;display_name=YouTube&amp;amp;url=https%3A%2F%2Fwww.youtube.com%2Fwatch%3Fv%3Do2il5eR0e5M&amp;amp;image=https%3A%2F%2Fi.ytimg.com%2Fvi%2Fo2il5eR0e5M%2Fhqdefault.jpg&amp;amp;type=text%2Fhtml&amp;amp;schema=youtube" width="600" height="337" scrolling="no" title="Community BrickTalk | Databricks Certification - Did You Know?, Whats New, and Where Were Going" frameborder="0" allow="autoplay; fullscreen; encrypted-media; picture-in-picture" allowfullscreen="true"&gt;&lt;/iframe&gt;&lt;/div&gt;&lt;/P&gt;</description>
      <pubDate>Fri, 09 Oct 2026 11:24:31 GMT</pubDate>
      <guid>https://community.databricks.com/t5/bricktalks-tv/bricktalk-recording-databricks-certification-did-you-know-what-s/ba-p/172426</guid>
      <dc:creator>Tushar_Parekar</dc:creator>
      <dc:date>2026-10-09T11:24:31Z</dc:date>
    </item>
    <item>
      <title>Introducing session restore for serverless jobs</title>
      <link>https://community.databricks.com/t5/technical-blog/introducing-session-restore-for-serverless-jobs/ba-p/170370</link>
      <description>&lt;P&gt;&lt;SPAN&gt;Debug failures and explore results from Lakeflow Jobs in an interactive notebook, with Python and Spark state restored so you can pick up exactly where the run left off, without rerunning the job.&lt;/SPAN&gt;&lt;/P&gt;</description>
      <pubDate>Thu, 08 Oct 2026 09:14:32 GMT</pubDate>
      <guid>https://community.databricks.com/t5/technical-blog/introducing-session-restore-for-serverless-jobs/ba-p/170370</guid>
      <dc:creator>sbanchik</dc:creator>
      <dc:date>2026-10-08T09:14:32Z</dc:date>
    </item>
    <item>
      <title>From Batch to Milliseconds: Real-Time Feature Views for Recommenders</title>
      <link>https://community.databricks.com/t5/technical-blog/from-batch-to-milliseconds-real-time-feature-views-for/ba-p/170806</link>
      <description>&lt;P&gt;Databricks Feature Views bridge the gap between traditional batch processing and real-time ingestion by enabling near real-time transformations and low-latency feature lookups for recommenders. By combining &lt;CODE&gt;StreamSource&lt;/CODE&gt;, &lt;CODE&gt;DeltaTableSource&lt;/CODE&gt;, and &lt;CODE&gt;RequestSource&lt;/CODE&gt; definitions within Unity Catalog, platforms can continuously materialize streaming session data and batch aggregates into Lakebase online tables. This unified architecture prevents data leakage through point-in-time joins during offline training and allows model serving endpoints to automatically retrieve fresh, multi-cadence features at inference time—delivering highly relevant recommendations while a shopper is actively browsing.&lt;/P&gt;</description>
      <pubDate>Thu, 08 Oct 2026 09:05:43 GMT</pubDate>
      <guid>https://community.databricks.com/t5/technical-blog/from-batch-to-milliseconds-real-time-feature-views-for/ba-p/170806</guid>
      <dc:creator>MarshallCarter</dc:creator>
      <dc:date>2026-10-08T09:05:43Z</dc:date>
    </item>
    <item>
      <title>Databricks SSH Tunnel: Connect Coding Agents and IDEs to your Workspace</title>
      <link>https://community.databricks.com/t5/databricks-tv/databricks-ssh-tunnel-connect-coding-agents-and-ides-to-your/ba-p/172192</link>
      <description>Learn to streamline your Databricks workspace by integrating your favorite editors for a smoother coding experience.</description>
      <pubDate>Thu, 08 Oct 2026 14:08:56 GMT</pubDate>
      <guid>https://community.databricks.com/t5/databricks-tv/databricks-ssh-tunnel-connect-coding-agents-and-ides-to-your/ba-p/172192</guid>
      <dc:creator>DSBDatabricks</dc:creator>
      <dc:date>2026-10-08T14:08:56Z</dc:date>
    </item>
    <item>
      <title>Unlocking High-Frequency Workloads with Nanosecond-Precision Timestamps</title>
      <link>https://community.databricks.com/t5/technical-blog/unlocking-high-frequency-workloads-with-nanosecond-precision/ba-p/170873</link>
      <description>&lt;P&gt;Learn how Timestamp Nano on Databricks enables use cases across financial markets, industrial sensors, and scientific research.&lt;/P&gt;</description>
      <pubDate>Tue, 06 Oct 2026 19:49:27 GMT</pubDate>
      <guid>https://community.databricks.com/t5/technical-blog/unlocking-high-frequency-workloads-with-nanosecond-precision/ba-p/170873</guid>
      <dc:creator>ajiang22</dc:creator>
      <dc:date>2026-10-06T19:49:27Z</dc:date>
    </item>
    <item>
      <title>Databricks Community Champion - September 2026 - Venugopal Dabbara</title>
      <link>https://community.databricks.com/t5/databricks-community-champions/databricks-community-champion-september-2026-venugopal-dabbara/ba-p/170782</link>
      <description>&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 13px; font-weight: bold; letter-spacing: 2px; text-transform: uppercase; color: #ff3621; margin: 0 0 6px;"&gt;Community Champion • September 2026&lt;/P&gt;
&lt;H2 style="font-family: Arial, Helvetica, sans-serif; font-size: 30px; line-height: 1.25; color: #1b3139; margin: 0 0 18px;"&gt;Meet Venugopal Dabbara&lt;/H2&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 17px; line-height: 1.7; color: #1b1b1f; margin: 0 0 18px;"&gt;Every month, our Community Champion Program recognizes a member who makes the Databricks Community better for everyone, through their knowledge, their willingness to help, and their commitment to learning alongside others. These are the people who turn questions into answers and members into a community.&lt;/P&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 17px; line-height: 1.7; color: #1b1b1f; margin: 0 0 18px;"&gt;Please join us in congratulating &lt;STRONG&gt;Venugopal Dabbara&lt;/STRONG&gt; (&lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/250262"&gt;@data_pulse&lt;/a&gt;), a &lt;STRONG&gt;Senior Data Consultant&lt;/STRONG&gt; at Colibri Digital with over 16 years in the data space and more than 8 years building on Databricks. Since joining the Community, Venugopal has quickly become a go-to voice for practical, well-tested answers, often validating solutions himself and backing them with docs, examples, and implementation guidance so others can move forward with confidence.&lt;/P&gt;
&lt;!-- FULL-WIDTH FEATURED PANEL: photo (left) + profile (right) --&gt;
&lt;TABLE style="width: 100%; border: 0; border-collapse: separate; background-color: #fbf4f2; background-image: linear-gradient(135deg,#FDF0ED 0%,#F6F8FB 65%,#FFFFFF 100%); border-radius: 16px; margin: 32px 0 36px;" role="presentation" border="0" width="100%" cellspacing="0" cellpadding="0"&gt;
&lt;TBODY&gt;
&lt;TR&gt;
&lt;TD style="border: 0; vertical-align: middle; padding: 14px 6px 14px 14px; width: 214px;"&gt;&lt;!-- PHOTO PLACEHOLDER (square): delete this box and upload the champion's square photo here via the editor's image button --&gt;
&lt;DIV style="width: 214px; height: 214px; line-height: 214px; font-family: Arial, Helvetica, sans-serif; font-size: 12px; color: #9a9aa2; text-align: center; background-color: #ffffff; border-radius: 16px; box-shadow: 0 6px 18px rgba(27,49,57,0.14);"&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="Champion.png" style="width: 800px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/31785i7C375183B8DC6477/image-size/large?v=v2&amp;amp;px=999" role="button" title="Champion.png" alt="Champion.png" /&gt;&lt;/span&gt;&lt;/DIV&gt;
&lt;/TD&gt;
&lt;TD style="border: 0; vertical-align: middle; padding: 24px 28px 24px 30px; width: 100%;"&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 12px; font-weight: bold; letter-spacing: 2px; text-transform: uppercase; color: #ff3621; margin: 0 0 8px;"&gt;Featured Member&lt;/P&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 25px; font-weight: bold; color: #1b3139; margin: 0 0 12px;"&gt;Venugopal Dabbara&lt;/P&gt;
&lt;DIV style="width: 40px; height: 3px; background-color: #ff3621; border-radius: 2px; margin: 0 0 14px;"&gt;&amp;nbsp;&lt;/DIV&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 15px; line-height: 1.8; color: #5a5a62; margin: 0;"&gt;&lt;STRONG&gt;Job title:&amp;nbsp;&lt;/STRONG&gt;Senior Data Consultant&lt;BR /&gt;&lt;STRONG&gt;Company:&amp;nbsp;&lt;/STRONG&gt;Colibri Digital&lt;BR /&gt;&lt;STRONG&gt;Community nickname:&lt;/STRONG&gt; &lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/250262"&gt;@data_pulse&lt;/a&gt;&amp;nbsp;&amp;nbsp;•&amp;nbsp; He/Him&lt;/P&gt;
&lt;/TD&gt;
&lt;/TR&gt;
&lt;/TBODY&gt;
&lt;/TABLE&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 17px; line-height: 1.7; color: #1b1b1f; margin: 0 0 30px;"&gt;Let’s get to know Venugopal a little better.&lt;/P&gt;
&lt;!-- SECTION 1 --&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 14px; font-weight: bold; letter-spacing: 1.5px; text-transform: uppercase; color: #ff3621; margin: 0 0 18px;"&gt;The Journey&lt;/P&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 18px; font-weight: bold; color: #1b3139; border-bottom: 2px solid #FF3621; padding-bottom: 8px; margin: 0 0 12px;"&gt;&lt;span class="lia-unicode-emoji" title=":microphone:"&gt;🎤&lt;/span&gt;&amp;nbsp; Can you provide a brief overview of your career journey leading up to your current role?&lt;/P&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 17px; line-height: 1.75; color: #333338; margin: 0 0 14px;"&gt;&lt;span class="lia-unicode-emoji" title=":speech_balloon:"&gt;💬&lt;/span&gt;&amp;nbsp; I started my career in the data space around 16 years ago, working with Oracle, SQL Server, data warehousing and BI, before moving into Big Data, Spark, cloud and Databricks.&lt;/P&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 17px; line-height: 1.75; color: #333338; margin: 0 0 14px;"&gt;Over the last 8 years in the UK, I have worked with multiple clients across different industries, helping design and deliver Databricks-based data platforms and solutions. My experience covers solution architecture, data engineering, ML integration, performance scaling, governance, CI/CD and end-to-end product delivery. I also hold certifications across Azure, AWS, Databricks and Claude.&lt;/P&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 17px; line-height: 1.75; color: #333338; margin: 0 0 14px;"&gt;In my current role, I work as a Senior Data Consultant with Colibri for an Energy Services client, helping design and build a scalable data platform covering Data Engineering, analytics, AI/ML and emerging GenAI use cases.&lt;/P&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 17px; line-height: 1.75; color: #333338; margin: 0 0 34px;"&gt;I became a Databricks Solution Architect Champion in 2021 and have since supported aspiring champions through project reviews, architecture feedback and panel interviews.&lt;/P&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 18px; font-weight: bold; color: #1b3139; border-bottom: 2px solid #FF3621; padding-bottom: 8px; margin: 0 0 12px;"&gt;&lt;span class="lia-unicode-emoji" title=":microphone:"&gt;🎤&lt;/span&gt;&amp;nbsp; What do you enjoy most about your current job or role?&lt;/P&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 17px; line-height: 1.75; color: #333338; margin: 0 0 14px;"&gt;&lt;span class="lia-unicode-emoji" title=":speech_balloon:"&gt;💬&lt;/span&gt;&amp;nbsp; What I enjoy most is the combination of &lt;STRONG&gt;hands-on engineering, solution architecture and customer engagement&lt;/STRONG&gt;. I like taking a problem from discovery through architecture, build, scale and productionize, and seeing it through to successful product delivery.&lt;/P&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 17px; line-height: 1.75; color: #333338; margin: 0 0 34px;"&gt;I also really value the &lt;STRONG&gt;culture at Colibri&lt;/STRONG&gt;, especially the trust and independence to explore ideas, innovate and propose pragmatic solutions. There is a strong focus on continuous learning, sharing best practices, and capturing proven engineering patterns as reusable references and accelerators that help other engineers understand solution approaches and deliver with greater confidence. That, along with mentoring and continuous upskilling across Databricks, Azure, AI/ML and GenAI, makes the role very engaging.&lt;/P&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 18px; font-weight: bold; color: #1b3139; border-bottom: 2px solid #FF3621; padding-bottom: 8px; margin: 0 0 12px;"&gt;&lt;span class="lia-unicode-emoji" title=":microphone:"&gt;🎤&lt;/span&gt;&amp;nbsp; If you had to describe yourself using three words, what would they be? How do you think your coworkers would describe you?&lt;/P&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 17px; line-height: 1.75; color: #333338; margin: 0 0 14px;"&gt;&lt;span class="lia-unicode-emoji" title=":speech_balloon:"&gt;💬&lt;/span&gt;&amp;nbsp; I would describe myself as pragmatic, dependable and innovative.&lt;/P&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 17px; line-height: 1.75; color: #333338; margin: 0 0 14px;"&gt;I’m pragmatic in finding solutions that are technically sound but also practical to deliver, dependable in taking ownership and supporting the team, and innovative in continuously exploring better approaches across data, cloud and AI.&lt;/P&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 17px; line-height: 1.75; color: #333338; margin: 0 0 34px;"&gt;I think my coworkers would describe me as approachable, reliable and solution-focused, someone who is happy to get hands-on, share knowledge and help the team deliver successfully.&lt;/P&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 18px; font-weight: bold; color: #1b3139; border-bottom: 2px solid #FF3621; padding-bottom: 8px; margin: 0 0 12px;"&gt;&lt;span class="lia-unicode-emoji" title=":microphone:"&gt;🎤&lt;/span&gt;&amp;nbsp; Have you had any mentors or significant influences in your professional life? If so, could you tell us about them?&lt;/P&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 17px; line-height: 1.75; color: #333338; margin: 0 0 14px;"&gt;&lt;span class="lia-unicode-emoji" title=":speech_balloon:"&gt;💬&lt;/span&gt;&amp;nbsp; Rather than one specific mentor, I’ve been fortunate to learn from a strong network of people across the Databricks ecosystem. Different people have influenced me in different ways, some around architecture and solution design, others around engineering practices, customer engagement, leadership and community contribution.&lt;/P&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 17px; line-height: 1.75; color: #333338; margin: 0 0 34px;"&gt;Through the Databricks Champion programme, I also had the opportunity to support and nurture aspiring champions by reviewing projects, sharing feedback and participating in panel discussions. I’ve stayed connected with many of them and continue to keep in touch with the wider Databricks community, which has been a great source of continuous learning and knowledge sharing.&lt;/P&gt;
&lt;!-- PULL QUOTE (VERBATIM sentence from Q8) --&gt;&lt;!-- SECTION 2 --&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 14px; font-weight: bold; letter-spacing: 1.5px; text-transform: uppercase; color: #ff3621; margin: 0 0 18px;"&gt;On Databricks&lt;/P&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 18px; font-weight: bold; color: #1b3139; border-bottom: 2px solid #FF3621; padding-bottom: 8px; margin: 0 0 12px;"&gt;&lt;span class="lia-unicode-emoji" title=":microphone:"&gt;🎤&lt;/span&gt;&amp;nbsp; When and why did you first start using Databricks?&lt;/P&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 17px; line-height: 1.75; color: #333338; margin: 0 0 14px;"&gt;&lt;span class="lia-unicode-emoji" title=":speech_balloon:"&gt;💬&lt;/span&gt;&amp;nbsp; I first started using Databricks extensively around beginning of 2019, when I worked on a Black Friday near real-time streaming analytics solution for a global marketing client.&lt;BR /&gt;The solution ingested high-volume Kafka streams into Azure Databricks, processed and aggregated the data using PySpark, and stored it in Gen2. We then exposed the processed results through REST APIs into Power BI to provide near real-time visibility during the Black Friday event.&lt;/P&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 17px; line-height: 1.75; color: #333338; margin: 0 0 34px;"&gt;It was a very good early use case for me because it demonstrated how Databricks could combine scalable cloud processing, streaming, low-latency analytics and BI consumption in one end-to-end solution. It also gave me strong confidence in using Databricks for production-scale data engineering workloads.&lt;/P&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 18px; font-weight: bold; color: #1b3139; border-bottom: 2px solid #FF3621; padding-bottom: 8px; margin: 0 0 12px;"&gt;&lt;span class="lia-unicode-emoji" title=":microphone:"&gt;🎤&lt;/span&gt;&amp;nbsp; Are there any Databricks features that you particularly enjoy or find indispensable in your work?&lt;/P&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 17px; line-height: 1.75; color: #333338; margin: 0 0 14px;"&gt;&lt;span class="lia-unicode-emoji" title=":speech_balloon:"&gt;💬&lt;/span&gt;&amp;nbsp; I particularly value Delta Lake, Unity Catalog, Lakeflow/Workflows, Asset Bundles and Databricks Dashboards, Serverless because they provide a strong foundation for reliable data processing, governance, orchestration, deployment and operational visibility.&lt;/P&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 17px; line-height: 1.75; color: #333338; margin: 0 0 34px;"&gt;I also find MLflow, Model Serving, Genie AI and AI Functions increasingly valuable for extending the same platform into ML and GenAI use cases.&lt;/P&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 18px; font-weight: bold; color: #1b3139; border-bottom: 2px solid #FF3621; padding-bottom: 8px; margin: 0 0 12px;"&gt;&lt;span class="lia-unicode-emoji" title=":microphone:"&gt;🎤&lt;/span&gt;&amp;nbsp; Is there a Databricks feature you wish existed or would like to see in future updates?&lt;/P&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 17px; line-height: 1.75; color: #333338; margin: 0 0 14px;"&gt;&lt;span class="lia-unicode-emoji" title=":speech_balloon:"&gt;💬&lt;/span&gt;&amp;nbsp; One area I would like to see improve further is concurrent write handling, especially where multiple asynchronous jobs write to the same Delta table. In particular, auto-generated identity columns can become a point of contention during concurrent writes, so smarter conflict handling around identity generation and metadata updates would help support more independent parallel writes.&lt;/P&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 17px; line-height: 1.75; color: #333338; margin: 0 0 34px;"&gt;I also like to see stronger native Data Quality observability and governance, with richer out-of-the-box DQ metrics, trends and failure visibility. And I would like Lakeflow Pipelines to continue maturing with more built-in connectors across legacy, cloud and SaaS systems to reduce custom ingestion code.&lt;/P&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 18px; font-weight: bold; color: #1b3139; border-bottom: 2px solid #FF3621; padding-bottom: 8px; margin: 0 0 12px;"&gt;&lt;span class="lia-unicode-emoji" title=":microphone:"&gt;🎤&lt;/span&gt;&amp;nbsp; When did you join the Databricks Community, and what motivated you to do so?&lt;/P&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 17px; line-height: 1.75; color: #333338; margin: 0 0 14px;"&gt;&lt;span class="lia-unicode-emoji" title=":speech_balloon:"&gt;💬&lt;/span&gt;&amp;nbsp; I joined the Databricks Community relatively recently, mainly to share the experience I’ve built over more than 8 years in the Databricks space and help others solve real technical challenges.&lt;/P&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 17px; line-height: 1.75; color: #333338; margin: 0 0 14px;"&gt;What motivates me most is the opportunity to make a real difference for someone who may be stuck on a problem or use case. I’ve been in those situations myself, where finding the right solution can take a lot of time, so being able to share a practical approach or expose someone to a better solution feels genuinely worthwhile.&lt;/P&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 17px; line-height: 1.75; color: #333338; margin: 0 0 34px;"&gt;I also try to make my contributions as useful as possible by validating the approach, referring to the right documentation, and sharing examples, pseudocode or implementation guidance where relevant. Where possible, I also test the solution myself so the guidance is practical, scalable and gives the person confidence to move forward.&lt;/P&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 18px; font-weight: bold; color: #1b3139; border-bottom: 2px solid #FF3621; padding-bottom: 8px; margin: 0 0 12px;"&gt;&lt;span class="lia-unicode-emoji" title=":microphone:"&gt;🎤&lt;/span&gt;&amp;nbsp; What aspects of the Databricks Community do you find most valuable or enjoyable?&lt;/P&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 17px; line-height: 1.75; color: #333338; margin: 0 0 34px;"&gt;&lt;span class="lia-unicode-emoji" title=":speech_balloon:"&gt;💬&lt;/span&gt;&amp;nbsp; What I value most is the &lt;STRONG&gt;two-way learning&lt;/STRONG&gt;. The questions often push me to explore features more deeply, validate approaches for scalability, performance and robustness, and stay updated with new Databricks capabilities, enhancements and documentation. I also enjoy the practical knowledge exchange, helping others while continuously broadening my own understanding through different perspectives and real-world use cases.&lt;/P&gt;
&lt;!-- SECTION 3 --&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 14px; font-weight: bold; letter-spacing: 1.5px; text-transform: uppercase; color: #ff3621; margin: 0 0 18px;"&gt;Beyond Work&lt;/P&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 18px; font-weight: bold; color: #1b3139; border-bottom: 2px solid #FF3621; padding-bottom: 8px; margin: 0 0 12px;"&gt;&lt;span class="lia-unicode-emoji" title=":microphone:"&gt;🎤&lt;/span&gt;&amp;nbsp; Outside of work, what is your favourite hobby or pastime?&lt;/P&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 17px; line-height: 1.75; color: #333338; margin: 0 0 34px;"&gt;&lt;span class="lia-unicode-emoji" title=":speech_balloon:"&gt;💬&lt;/span&gt;&amp;nbsp; Outside of work, I really enjoy &lt;STRONG&gt;travelling, sports and hiking&lt;/STRONG&gt;. I have covered around half of Europe so far and would love to explore the rest over the next few years. I also enjoy staying active through different sports and outdoor activities, and I’m currently planning a hiking expedition in the near future&lt;/P&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 18px; font-weight: bold; color: #1b3139; border-bottom: 2px solid #FF3621; padding-bottom: 8px; margin: 0 0 12px;"&gt;&lt;span class="lia-unicode-emoji" title=":microphone:"&gt;🎤&lt;/span&gt;&amp;nbsp; Where do you envision yourself professionally in the next three years?&lt;/P&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 17px; line-height: 1.75; color: #333338; margin: 0 0 14px;"&gt;&lt;span class="lia-unicode-emoji" title=":speech_balloon:"&gt;💬&lt;/span&gt;&amp;nbsp; I envision myself building &lt;STRONG&gt;next-generation AI-led platforms using GenAI, Genie and agentic frameworks on Databricks&lt;/STRONG&gt;, moving beyond information delivery to become &lt;STRONG&gt;decision accelerators&lt;/STRONG&gt; that provide intelligent recommendations, guided actions and faster business outcomes at scale.&lt;/P&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 17px; line-height: 1.75; color: #333338; margin: 0 0 14px;"&gt;I would like to build and contribute to an end-to-end intelligent platform spanning front-end experiences, data engineering, AI/ML and downstream reporting, while continuing to grow into a stronger &lt;STRONG&gt;technical leadership role&lt;/STRONG&gt;.&lt;/P&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 17px; line-height: 1.75; color: #333338; margin: 0 0 34px;"&gt;I also want to contribute more actively to the Databricks Community by sharing practical lessons, proven solution patterns and real-world engineering insights, while supporting new aspirants in upskilling and publishing technical blogs.&lt;/P&gt;
&lt;HR /&gt;&lt;!-- CLOSING (Community team voice + connect + CTA) --&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 17px; line-height: 1.7; color: #1b1b1f; margin: 0 0 14px;"&gt;On behalf of the Databricks Community team – thank you, &lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/250262"&gt;@data_pulse&lt;/a&gt;! Beyond helping fellow members learn, solve problems, and grow in their technical journeys, you've also been an extra set of eyes for our moderators, flagging what needs attention and helping keep the community clean and welcoming. We're so grateful to have you here.&lt;/P&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 17px; line-height: 1.7; color: #1b1b1f; margin: 0 0 14px;"&gt;Congratulations on becoming our Community Champion for September 2026! &lt;span class="lia-unicode-emoji" title=":rocket:"&gt;🚀&lt;/span&gt;&lt;/P&gt;
&lt;P style="font-family: Arial, Helvetica, sans-serif; font-size: 17px; line-height: 1.7; color: #1b1b1f; margin: 0 0 14px;"&gt;Connect with Venugopal on LinkedIn &lt;A style="color: #ff3621; font-weight: bold;" href="https://www.linkedin.com/in/venugopal-dabbara/" target="_self"&gt;here&lt;/A&gt;, and join us in congratulating him in the comments below! &lt;span class="lia-unicode-emoji" title=":backhand_index_pointing_down:"&gt;👇&lt;/span&gt;&lt;/P&gt;</description>
      <pubDate>Tue, 06 Oct 2026 16:04:19 GMT</pubDate>
      <guid>https://community.databricks.com/t5/databricks-community-champions/databricks-community-champion-september-2026-venugopal-dabbara/ba-p/170782</guid>
      <dc:creator>Advika</dc:creator>
      <dc:date>2026-10-06T16:04:19Z</dc:date>
    </item>
    <item>
      <title>Rebuilding an 800-Line PySpark Pipeline in 200 Lines of Databricks SQL</title>
      <link>https://community.databricks.com/t5/technical-blog/rebuilding-an-800-line-pyspark-pipeline-in-200-lines-of/ba-p/168578</link>
      <description>&lt;H1&gt;&lt;STRONG&gt;End-to-End DBSQL ETL Use Case&lt;/STRONG&gt;&lt;/H1&gt;
&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="Screenshot 2026-10-06 at 10.24.23.png" style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/31767iACCC14E0C07D1BBF/image-size/large?v=v2&amp;amp;px=999" role="button" title="Screenshot 2026-10-06 at 10.24.23.png" alt="Screenshot 2026-10-06 at 10.24.23.png" /&gt;&lt;/span&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;For years, a production ETL pipeline on Databricks meant a pile of PySpark. SCD Type 2 dimensions, cost-controlled materialized views, idempotent reloads of late data: each one was code you wrote and maintained. That has changed. AUTO CDC handles SCD Type 2 dimensions. REFRESH POLICY controls what a materialized view costs to keep current. REPLACE WHERE makes late-arriving data safe to reload. You can build a full e-commerce warehouse, bronze through platinum, and write every layer of it in Databricks SQL.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;Everything below is plain DBSQL: CREATE STREAMING TABLE, AUTO CDC flows, REPLACE WHERE flows, and metric views. Every layer runs directly on a serverless SQL warehouse. You can paste it into the SQL editor and build the whole warehouse, bronze through platinum, with no pipeline and no external orchestrator. When you take it to production you have a choice. You can keep running it in DBSQL, or you can package Bronze through Gold (the streaming tables, the AUTO CDC dimension, and the REPLACE WHERE flow) as a single Lakeflow Spark Declarative Pipeline (SDP), which adds a managed DAG, a shared event log, and clean promotion across dev, staging, and prod through a Databricks Asset Bundle. The Platinum metric views stay in DBSQL either way, because a metric view lives in the SQL warehouse rather than inside a pipeline. We cover deployment at the end.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;For comparison, think about how you would have built this before these primitives existed. A traditional warehouse stack would use stored procedures and hand-written MERGE INTO statements; a Spark shop would use PySpark notebooks orchestrated by Workflows. Either way you ended up with a 200-line MERGE for the SCD2 dimension, a script that allocated surrogate keys against a state table, and a separate nightly job that recomputed the gold aggregate from scratch. That pipeline produces correct numbers. It also takes a sprint to build, a couple of weeks to debug, and a senior engineer to keep running.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;The version below is about 200 lines of SQL in one file, with the same governance and observability you already have. The difference is that you write the schema and the joins, and Databricks handles the plumbing underneath.&lt;/SPAN&gt;&lt;/P&gt;
&lt;H2&gt;&lt;STRONG&gt;The pipeline we'll build&lt;/STRONG&gt;&lt;/H2&gt;
&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="Fangzhu_1-1789415121315.png" style="width: 805px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/31089i7209B7692BCB9A3A/image-dimensions/805x312?v=v2" width="805" height="312" role="button" title="Fangzhu_1-1789415121315.png" alt="Fangzhu_1-1789415121315.png" /&gt;&lt;/span&gt;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Domain:&lt;/STRONG&gt;&lt;SPAN&gt; an e-commerce data warehouse. Three sources:&lt;/SPAN&gt;&lt;/P&gt;
&lt;OL&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;SPAN&gt;customer_changes: CDC feed for the customer dimension.&lt;/SPAN&gt;&lt;SPAN&gt; Slowly changing attributes (segment, address, email).&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;SPAN&gt;order_events: append-only stream of order placements.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;SPAN&gt;line_item_events: append-only stream of order line-items, late by up to 24 hours.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;&lt;STRONG&gt;Targets:&lt;/STRONG&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;SPAN&gt;dim_customer: Silver, SCD Type 2 customer dim&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;SPAN&gt;fact_orders: Gold, order-level fact, surrogate-resolved against the dim&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;SPAN&gt;fact_line_items: Gold, line-item fact reconciled with a REPLACE WHERE flow for late data&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;SPAN&gt;revenue_by_segment: Platinum metric view with a materialization&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;SPAN&gt;top_products: Platinum metric view with a materialization&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;SPAN&gt;You author everything in plain Databricks SQL, and it runs as written on a serverless SQL warehouse. For production, you can optionally package Bronze through Gold (the streaming tables, the AUTO CDC dimension, and the REPLACE WHERE flow) as one SDP pipeline, which infers the dependencies from the table references and runs them as a single DAG, and promote it across dev, staging, and prod with a Databricks Asset Bundle. The Platinum metric views are created directly in DBSQL on top of the Gold facts in either setup.&lt;/SPAN&gt;&lt;/P&gt;
&lt;H2&gt;&lt;STRONG&gt;Bronze: append-only ingestion with expectations&lt;/STRONG&gt;&lt;/H2&gt;
&lt;P&gt;&lt;SPAN&gt;Two things shape bronze ingestion. First, the events are streaming, so the bronze tables need to be incremental. Second, line items can arrive up to 24 hours late. Bronze doesn't try to fix that: it lands every event append-only, exactly as it arrives, and leaves late-arrival reconciliation to the Gold layer. The bronze layer keeps a faithful, append-only record of the source and nothing more. Expectations (&lt;/SPAN&gt;&lt;SPAN&gt;CONSTRAINT … EXPECT&lt;/SPAN&gt;&lt;SPAN&gt;) handle bronze-layer data quality before bad rows poison anything downstream. (CONSTRAINT … EXPECT works on standalone DBSQL streaming tables created in the SQL editor, so you get expectations without deploying an SDP pipeline.)&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;One placement call worth flagging: we run expectations in bronze to drop bad rows before they land. The more common convention is to leave bronze exactly as it arrives and apply the drops in silver, so you can change a filter and reprocess from an intact bronze without re-ingesting the source. We drop at bronze here because these constraints catch hard parse and schema failures that can never become useful rows, and because the raw JSON still sits in the source volume, so a full refresh can rebuild bronze from scratch if a rule turns out too aggressive. If your expectations are business rules rather than structural checks, prefer the silver-drop pattern.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;-- Customer CDC source (Delta with CDF + row tracking)&lt;/SPAN&gt;&lt;/P&gt;
&lt;LI-CODE lang="markup"&gt;CREATE OR REPLACE TABLE ${target_catalog}.bronze.customer_changes (

  customer_id INT,

  email       STRING,

  segment     STRING,

  city        STRING,

  updated_at  TIMESTAMP

)

TBLPROPERTIES (

  delta.enableChangeDataFeed = true,

  delta.enableRowTracking    = true

);

&lt;/LI-CODE&gt;
&lt;P&gt;&lt;SPAN&gt;-- Order events (streaming, append-only, clustered by order_date, with quality checks)&lt;/SPAN&gt;&lt;/P&gt;
&lt;LI-CODE lang="markup"&gt;CREATE STREAMING TABLE IF NOT EXISTS ${target_catalog}.bronze.order_events (

  CONSTRAINT valid_order_id    EXPECT (order_id IS NOT NULL)    ON VIOLATION DROP ROW,

  CONSTRAINT valid_customer_id EXPECT (customer_id IS NOT NULL) ON VIOLATION DROP ROW,

  CONSTRAINT non_negative_amt  EXPECT (total_amount &amp;gt;= 0)        ON VIOLATION DROP ROW

)

CLUSTER BY (order_date)

SCHEDULE EVERY 1 HOUR

AS

SELECT

  o.order_id,

  o.customer_id,

  cast(o.order_ts AS DATE) AS order_date,

  o.total_amount,

  o.order_ts AS event_time

FROM STREAM read_files(

  '/Volumes/${target_catalog}/raw/orders/',

  format =&amp;gt; 'json'

) AS o;&lt;/LI-CODE&gt;
&lt;P&gt;&lt;SPAN&gt;-- Line items (streaming, append-only; late arrivals reconciled downstream in Gold)&lt;/SPAN&gt;&lt;/P&gt;
&lt;LI-CODE lang="markup"&gt;CREATE STREAMING TABLE IF NOT EXISTS ${target_catalog}.bronze.line_item_events (

CONSTRAINT valid_line_item   EXPECT (line_item_id IS NOT NULL) ON VIOLATION DROP ROW,

CONSTRAINT valid_order_id    EXPECT (order_id IS NOT NULL)     ON VIOLATION DROP ROW,

  CONSTRAINT non_negative_qty  EXPECT (quantity &amp;gt; 0)             ON VIOLATION DROP ROW

)

CLUSTER BY (order_date) 

AS

SELECT

  l.line_item_id,

  l.order_id,

  l.product_id,

  l.quantity,

  l.line_amount,

  cast(l.order_ts AS DATE) AS order_date,

  l.event_ts AS event_time

FROM STREAM read_files(

  '/Volumes/${target_catalog}/raw/line_items/',

  format =&amp;gt; 'json'

) AS l;&lt;/LI-CODE&gt;
&lt;P&gt;&lt;SPAN&gt;The whole pipeline uses one clustering approach. Both bronze event tables cluster explicitly by order_date because we know the access pattern: consumers filter by date, and it gives the downstream Gold reconciliation a clean date slice to target. The Silver dimension uses CLUSTER BY AUTO, since its access pattern is workload-dependent and Predictive Optimization can choose the columns. Gold is mixed: fact_orders uses CLUSTER BY AUTO, while fact_line_items clusters by order_date so the REPLACE WHERE flow has a clean slice to rewrite. Under SDP, the pipeline's own schedule drives all of its tables together, so you don't set a cadence per table. When you run a streaming table standalone instead, it carries its own: SCHEDULE EVERY 1 HOUR keeps ingestion incremental on a timer, or TRIGGER ON UPDATE fires when upstream data lands.&lt;/SPAN&gt;&lt;/P&gt;
&lt;H2&gt;&lt;STRONG&gt;Silver: the AUTO CDC dimension&lt;/STRONG&gt;&lt;/H2&gt;
&lt;P&gt;&lt;SPAN&gt;Silver holds one table: the customer dimension, maintained by an AUTO CDC flow. AUTO CDC turns the raw customer_changes feed into an SCD Type 2 history table without a hand-written MERGE or any surrogate-key bookkeeping. The fact tables that join against it sit one layer up, in Gold. CLUSTER BY AUTO lets Databricks choose the clustering columns from the actual query workload.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;-- SCD Type 2 customer dimension via AUTO CDC&lt;/SPAN&gt;&lt;/P&gt;
&lt;LI-CODE lang="markup"&gt;CREATE OR REFRESH STREAMING TABLE ${target_catalog}.silver.dim_customer (

  customer_sk BIGINT GENERATED ALWAYS AS IDENTITY,

  customer_id INT,

  email       STRING,

  segment     STRING,

  city        STRING,

  updated_at  TIMESTAMP,

  __START_AT  TIMESTAMP,

  __END_AT    TIMESTAMP

)&lt;/LI-CODE&gt;
&lt;P&gt;&lt;SPAN&gt;COMMENT 'SCD Type 2 customer dimension with generated surrogate key'&lt;/SPAN&gt;&lt;/P&gt;
&lt;LI-CODE lang="markup"&gt;CLUSTER BY AUTO

FLOW AUTO CDC

FROM STREAM ${target_catalog}.bronze.customer_changes

  KEYS (customer_id)

SEQUENCE BY updated_at

STORED AS SCD TYPE 2;&lt;/LI-CODE&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;UL&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;SPAN&gt;This is standard Databricks SQL. The same CREATE STREAMING TABLE ... FLOW AUTO CDC ... STORED AS SCD TYPE 2 statement runs unchanged in the SQL editor on a serverless warehouse and inside an SDP pipeline, so there is no separate DBSQL form of the AUTO CDC syntax to learn.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;SPAN&gt;AUTO CDC keeps the SEQUENCE BY column (updated_at) in the SCD2 target, so it's declared in the schema alongside __START_AT and __END_AT.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;SPAN&gt;customer_sk is a generated surrogate key. AUTO CDC supports GENERATED ALWAYS AS IDENTITY on an SCD Type 2 target, minting a fresh surrogate for each version row, so every (customer_id, effective-period) pair gets its own stable customer_sk. That is what the fact join needs: each fact row carries one compact integer pointing at the exact historical version of the customer, instead of repeating the (customer_id, __START_AT, __END_AT) predicate everywhere downstream. Because the segment is denormalized at write time, most queries never touch the dim. Identity values are assigned only for newly processed rows and stay correct under late or out-of-order data, which is the guarantee an SCD2 dimension needs. If you ever need to reproduce a surrogate outside the streaming table, for example to backfill a legacy fact, you can derive the same ordering with ROW_NUMBER() OVER (ORDER BY customer_id, __START_AT) in a view.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;UL&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;SPAN&gt;Use STORED AS SCD TYPE 2 when you need historical attribution: an order placed while the customer was in segment silver stays silver after they move to gold. If the dimension only needs current state, STORED AS SCD TYPE 1 uses less storage, keeps joins simpler, and drops the __START_AT / __END_AT columns.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;SPAN&gt;Gotcha: snapshot-join semantics on the dim. When a streaming fact joins the dim as a static table (the fact_orders join in the Gold section below), the dim snapshot is taken at the start of each microbatch. A dim update that lands after a fact row is written won't retroactively change that fact's surrogate. Say an order writes at 10:00 pointing at the customer's silver version; if a correction moves that customer to gold and arrives at 10:05, the order already on disk keeps pointing at silver. When you need the correction applied retroactively, reprocess the affected facts from bronze. That's the SCD2 contract working as designed, not a bug to engineer around.&lt;/SPAN&gt;&lt;/P&gt;
&lt;H2&gt;&lt;SPAN&gt;Gold: fact tables and the REPLACE WHERE flow for late data&lt;/SPAN&gt;&lt;/H2&gt;
&lt;P&gt;&lt;SPAN&gt;Gold is where the star schema comes together: the two fact tables, surrogate-resolved against the Silver dim, plus the reconciliation that handles late-arriving line items. (Naming note: these tables keep their silver.* schema names in SQL; Gold is the logical layer here, not the schema.)&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;fact_orders is a plain streaming fact. order_events arrives on time and append-only, so a standard streaming flow resolves customer_sk against the historical dim with a temporal join:&lt;/SPAN&gt;&lt;/P&gt;
&lt;LI-CODE lang="markup"&gt;CREATE OR REFRESH STREAMING TABLE ${target_catalog}.silver.fact_orders

COMMENT 'Order facts with customer_sk resolved against the historical dim'

CLUSTER BY AUTO

AS

SELECT

  o.order_id,

  c.customer_sk,          -- resolved from the dim, not from bronze

  c.segment,

  o.order_date,

  o.total_amount,

  o.event_time

FROM STREAM(${target_catalog}.bronze.order_events) o

LEFT JOIN ${target_catalog}.silver.dim_customer c

  ON o.customer_id = c.customer_id

  AND o.event_time &amp;gt;= c.__START_AT

  AND (c.__END_AT IS NULL OR o.event_time &amp;lt; c.__END_AT);&lt;/LI-CODE&gt;
&lt;P&gt;&lt;SPAN&gt;The temporal join picks the correct historical version of the customer for each order. An order placed in April, when the customer was in segment silver, stays silver even if that customer moves to gold later. You write the SCD2 join once and it holds for every fact.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;fact_line_items is the table that needs REPLACE WHERE. Line items can land up to 24 hours late, and a pure append flow would strand them: a line item that arrives after its order's date slice was first built would never get reconciled against the order. So instead of appending, we run a REPLACE WHERE flow that rewrites the trailing window of order_date on every refresh, re-evaluating the join over just that slice:&lt;/SPAN&gt;&lt;/P&gt;
&lt;LI-CODE lang="markup"&gt;CREATE OR REFRESH STREAMING TABLE ${target_catalog}.silver.fact_line_items

COMMENT 'Line-item facts; late arrivals reconciled by a REPLACE WHERE flow on order_date'

CLUSTER BY (order_date)&lt;/LI-CODE&gt;
&lt;P&gt;&lt;SPAN&gt;-- REPLACE WHERE flow: idempotently rewrite the trailing window that can&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;-- still receive late line items, re-evaluating the join over just that slice.&lt;/SPAN&gt;&lt;/P&gt;
&lt;LI-CODE lang="markup"&gt;FLOW REPLACE WHERE order_date &amp;gt;= date_add(current_date(), -2) BY NAME

SELECT

  l.line_item_id,

  l.order_id,

  o.customer_sk,

  o.segment,

  l.product_id,

  l.quantity,

  l.line_amount,

  l.order_date,

  l.event_time

FROM ${target_catalog}.bronze.line_item_events l

LEFT JOIN ${target_catalog}.silver.fact_orders o

  ON l.order_id = o.order_id

WHERE l.order_date &amp;gt;= date_add(current_date(), -2);&lt;/LI-CODE&gt;
&lt;P&gt;&lt;SPAN&gt;fact_line_items inherits customer_sk and segment from fact_orders, so the dim join runs once instead of once per line item, and customer attributes stay consistent between order and line-item rows. A two-day replace window comfortably covers the 24-hour late-arrival bound; widen it if your sources run later.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;Gotcha: the REPLACE WHERE flow rewrites only the matching slice, and it does so atomically. Readers see either the old slice or the new one, never a torn state, and rows outside the predicate are left alone. On serverless, incremental refresh reprocesses only the source data that changed inside the window rather than recomputing the whole slice, so a wide safety window stays cheap when little has changed.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;Keep the source filter matched to the REPLACE WHERE predicate. The flow already scopes each refresh to the predicate window, so any source rows that fall outside it are simply dropped rather than written. That makes the matching source filter a cost optimization rather than a correctness requirement: it prunes the scan to the same window, which is where the savings come from, and keeps the query's intent obvious to the next reader.&lt;/SPAN&gt;&lt;/P&gt;
&lt;H2&gt;&lt;STRONG&gt;Platinum: metric views with materializations for query acceleration&lt;/STRONG&gt;&lt;/H2&gt;
&lt;P&gt;&lt;SPAN&gt;The gold facts are query-ready, but every team that reads them will re-derive the same numbers, “revenue by segment,” “units sold this week,” and drift apart on the definitions. The platinum layer solves that with metric views. A metric view defines each measure once in governed YAML (revenue is SUM(total_amount)) and lets any consumer group it by any dimension at query time. Add a materialization and those aggregates are pre-computed and incrementally refreshed, giving you the query acceleration of a materialized view under one governed definition. Metric views and their materializations are authored and queried entirely in Databricks SQL.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;Revenue by segment as a metric view over the gold order fact:&lt;/SPAN&gt;&lt;/P&gt;
&lt;LI-CODE lang="markup"&gt;CREATE OR REPLACE VIEW ${target_catalog}.gold.revenue_by_segment

WITH METRICS

LANGUAGE YAML

AS $$

version: 1.1

comment: "Governed revenue metrics by customer segment"

source: ${target_catalog}.silver.fact_orders

dimensions:

  - name: Segment

    expr: segment

  - name: Order Date

    expr: order_date

measures:

  - name: Total Revenue

    expr: SUM(total_amount)

  - name: Order Count

    expr: COUNT(*)

materialization:

  schedule: EVERY 1 HOUR

  mode: relaxed

  materialized_views:

    - name: revenue_by_segment_daily

      type: aggregated

      dimensions: [Segment, Order Date]

      measures: [Total Revenue, Order Count]

$$;&lt;/LI-CODE&gt;
&lt;P&gt;&lt;SPAN&gt;Top products by week over the line-item fact, same shape with a weekly dimension and a distinct-order count:&lt;/SPAN&gt;&lt;/P&gt;
&lt;LI-CODE lang="markup"&gt;CREATE OR REPLACE VIEW ${target_catalog}.gold.top_products

WITH METRICS

LANGUAGE YAML

AS $$

version: 1.1

comment: "Governed weekly product metrics"

source: ${target_catalog}.silver.fact_line_items

dimensions:

  - name: Product

    expr: product_id

  - name: Week Starting

    expr: date_trunc('week', order_date)

measures:

  - name: Product Revenue

    expr: SUM(line_amount)

  - name: Units Sold

    expr: SUM(quantity)

  - name: Order Count

    expr: COUNT(DISTINCT order_id)

materialization:

  schedule: EVERY 1 HOUR

  mode: relaxed

  materialized_views:

    - name: top_products_weekly

      type: aggregated

      dimensions: [Product, Week Starting]

      measures: [Product Revenue, Units Sold, Order Count]

$$;&lt;/LI-CODE&gt;
&lt;P&gt;&lt;SPAN&gt;Consumers query the metric view, not the materialization. Measures are wrapped in MEASURE() and you group by whatever dimensions you need, and the definition of "revenue" is identical no matter who runs it:&lt;/SPAN&gt;&lt;/P&gt;
&lt;LI-CODE lang="markup"&gt;SELECT

  `Segment`,

  `Order Date`,

  MEASURE(`Total Revenue`) AS revenue,

  MEASURE(`Order Count`)    AS orders

FROM ${target_catalog}.gold.revenue_by_segment

WHERE `Order Date` &amp;gt;= current_date() - INTERVAL 30 DAYS

GROUP BY ALL

ORDER BY ALL;&lt;/LI-CODE&gt;
&lt;P&gt;&lt;SPAN&gt;This is where view matching pays off. The query never names revenue_by_segment_daily. When it runs, the optimizer checks the metric view's materializations, sees that the daily aggregate covers the requested dimensions and measures, and rewrites the query to read the pre-computed table instead of scanning the fact. You get materialized-view speed without hard-coding a materialized-view name, and a query that asks for a grouping no materialization covers falls back to the base fact automatically. One object gives you both the standard definition and the acceleration.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Gotcha: a materialization is a config, not a rewrite you own. You declare what to pre-compute in the YAML, and the engine owns the refresh (incremental, on the schedule you set) and the matching. Keep materializations aligned with how the metric view is actually queried: a materialization on dimensions nobody groups by burns refresh compute, and never gets matched. Start from your top dashboard queries and materialize those dimensions and measure combinations. (Metric-view materializations are recently GA and the config surface is still evolving, so check the current YAML reference before you ship.)&lt;/STRONG&gt;&lt;/P&gt;
&lt;H2&gt;&lt;STRONG&gt;Full flexibility: author in SQL, deploy as an SDP pipeline plus DBSQL metric views&lt;/STRONG&gt;&lt;/H2&gt;
&lt;P&gt;&lt;SPAN&gt;The pipeline deploys in two pieces, both authored in plain Databricks SQL. Bronze through Gold (the streaming tables, the AUTO CDC dimension, and the fact_line_items REPLACE WHERE flow) goes into a single SDP pipeline: SDP reads the table references, infers the dependency graph, and runs the layers in order, with no orchestration code to write. Wrap that .sql file in a Declarative Automation Bundles (formerly known as Databricks Asset Bundles), and you get one shared DAG, a shared event log, and clean promotion across dev, staging, and prod.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;The Platinum metric views are created separately, directly in DBSQL on top of the Gold facts. A metric view (CREATE VIEW … WITH METRICS) is a SQL-warehouse object, not a pipeline statement, so it lives outside the SDP file and refreshes on its own materialization schedule. The customer_changes CDC feed is likewise an external source table the pipeline reads, not one it manages. So the deploy has two steps: run the SDP pipeline for bronze through gold, then create the two metric views in DBSQL against the gold facts. (None of these statements requires an SDP pipeline. The streaming tables, the AUTO CDC dimension, and the inline REPLACE WHERE reconciliation all run standalone on a serverless SQL warehouse; the SDP pipeline is what adds the managed DAG, the shared event log, and clean promotion across dev, staging, and prod.)&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="Fangzhu_2-1789415121313.png" style="width: 612px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/31088iFBA27D288361A450/image-dimensions/612x612?v=v2" width="612" height="612" role="button" title="Fangzhu_2-1789415121313.png" alt="Fangzhu_2-1789415121313.png" /&gt;&lt;/span&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;# databricks.yml: bundle definition for the retail warehouse pipeline&lt;/SPAN&gt;&lt;/P&gt;
&lt;LI-CODE lang="markup"&gt;bundle:
  name: retail_warehouse



variables:
  target_catalog:
    description: Catalog where the pipeline writes
  source_volume_root:
    description: Root path for raw event JSON files



resources:
  pipelines:
    retail_warehouse:
      name: retail_warehouse_${bundle.target}
      catalog: ${var.target_catalog}
      target: gold
      libraries:
        - file:
            path: ./pipeline.sql
      configuration:
        target_catalog: ${var.target_catalog}
      serverless: true



targets:
  dev:
    variables:
      target_catalog: dev_main
      source_volume_root: /Volumes/dev_main/raw
  staging:
    variables:
      target_catalog: staging_main
      source_volume_root: /Volumes/staging_main/raw
  prod:
    variables:
      target_catalog: prod_main
      source_volume_root: /Volumes/prod_main/raw&lt;/LI-CODE&gt;
&lt;P&gt;&lt;SPAN&gt;Deploy with Databricks bundle deploy --target dev (or staging/prod). The same SQL pipeline runs across all three environments with no source code changes; the bundle target injects the catalog parameter at deployment time. The Pipeline UI shows the DAG, the per-table refresh status, and the last-run row counts for whichever target you triggered. The bundle deploys the bronze-through-gold pipeline; create the two metric views in DBSQL against the gold facts as a follow-on step (a post-deploy SQL task, or by hand in the editor), since metric views aren’t pipeline statements.&lt;/SPAN&gt;&lt;/P&gt;
&lt;H2&gt;&lt;STRONG&gt;Verifying incremental everywhere&lt;/STRONG&gt;&lt;/H2&gt;
&lt;P&gt;&lt;SPAN&gt;Once the pipeline runs, check that each materialization is refreshing incrementally rather than fully recomputing. A metric view’s materialization is managed by the view, and the pre-computed table isn’t always surfaced as a plain queryable object under the YAML name. If it is exposed in the target schema (the names you gave in the YAML: revenue_by_segment_daily, top_products_weekly), you can read its planning events from the event log; confirm the exact object names in Catalog Explorer first, and if event_log(TABLE(...)) can’t resolve the name, check the refresh from the metric view’s own materialization state instead:&lt;/SPAN&gt;&lt;/P&gt;
&lt;LI-CODE lang="markup"&gt;SELECT 'revenue_by_segment_daily' AS mv, message

FROM event_log(TABLE(${target_catalog}.gold.revenue_by_segment_daily))

WHERE event_type = 'planning_information'

AND origin.flow_name LIKE '%revenue_by_segment_daily%'

ORDER BY timestamp DESC LIMIT 1

UNION ALL

SELECT 'top_products_weekly' AS mv, message

FROM event_log(TABLE(${target_catalog}.gold.top_products_weekly))

WHERE event_type = 'planning_information'

AND origin.flow_name LIKE '%top_products_weekly%'

ORDER BY timestamp DESC LIMIT 1;&lt;/LI-CODE&gt;
&lt;P&gt;&lt;SPAN&gt;Both rows should report GROUP_AGGREGATE, the incremental technique. If either reports FULL_RECOMPUTE, the materialization is silently recomputing from scratch; narrow it to the dimensions and measures your queries actually match, or check the metric-view definition.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;For the streaming silver tables, query their event log for &lt;/SPAN&gt;&lt;SPAN&gt;flow_progress&lt;/SPAN&gt;&lt;SPAN&gt;:&lt;/SPAN&gt;&lt;/P&gt;
&lt;LI-CODE lang="markup"&gt;SELECT

  origin.flow_name,

  details:flow_progress.metrics.num_output_rows AS rows_processed,

  timestamp

FROM event_log(TABLE(${target_catalog}.silver.fact_orders))

WHERE event_type = 'flow_progress'

 AND origin.flow_name LIKE '%fact_orders%' 

ORDER BY timestamp DESC

LIMIT 5;&lt;/LI-CODE&gt;
&lt;P&gt;&lt;SPAN&gt;For the bronze layer's data quality:&lt;/SPAN&gt;&lt;/P&gt;
&lt;LI-CODE lang="markup"&gt;SELECT

  details:flow_progress.data_quality.expectations

FROM event_log(TABLE(${target_catalog}.bronze.order_events))

WHERE event_type = 'flow_progress'

AND origin.flow_name LIKE '%order_events%'

AND details:flow_progress.data_quality.expectations IS NOT NULL

ORDER BY timestamp DESC

LIMIT 1;&lt;/LI-CODE&gt;
&lt;P&gt;&lt;SPAN&gt;event_log() returns events for the entire pipeline the named table belongs to, not just events about that table. Filtering on origin.flow_name scopes results to the flow you actually care about. Without it, the ORDER BY timestamp DESC LIMIT 1 will return whichever flow happened to log most recently.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;The &lt;/SPAN&gt;&lt;SPAN&gt;expectations&lt;/SPAN&gt;&lt;SPAN&gt; JSON field has pass/fail counts for each &lt;/SPAN&gt;&lt;SPAN&gt;CONSTRAINT&lt;/SPAN&gt;&lt;SPAN&gt; you declared. Build an alerting query on top. For example, fire an alert when the failure rate on &lt;/SPAN&gt;&lt;SPAN&gt;non_negative_amt&lt;/SPAN&gt;&lt;SPAN&gt; exceeds 1% over a 5-minute window.&lt;/SPAN&gt;&lt;/P&gt;
&lt;H2&gt;&lt;STRONG&gt;Performance gotchas to watch for&lt;/STRONG&gt;&lt;/H2&gt;
&lt;P&gt;&lt;SPAN&gt;A few things tend to show up once a pipeline like this moves past the proof-of-concept stage.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Source materialization on AUTO CDC.&lt;/STRONG&gt;&lt;SPAN&gt; Behind the scenes, AUTO CDC sometimes needs to materialize the source change feed before applying changes. On very large or churn-heavy &lt;/SPAN&gt;&lt;SPAN&gt;customer_changes&lt;/SPAN&gt;&lt;SPAN&gt; tables, the time spent figuring out whether materialization is needed can end up dominating the silver-dim refresh, even when the actual changeset is small. The fixes are usually mundane: confirm row tracking &lt;/SPAN&gt;&lt;I&gt;&lt;SPAN&gt;and&lt;/SPAN&gt;&lt;/I&gt;&lt;SPAN&gt; CDF are both enabled on bronze, keep source volatility down (avoid frequent updates on bronze unless you actually need them), and shorten retention windows where you don't need long history.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Photon and AUTO CDC SCD2. Type 2 dim performance through AUTO CDC is much better with Photon. On serverless SQL warehouses Photon is on by default, so there is nothing to do. On classic SDP compute, check this first before drawing any conclusions about throughput.&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Small-file accumulation on low-volume bronze tables. Predictive Optimization compacts small files in the background, so this bites less than it used to, but very frequent triggers can still outrun compaction and leave slow reads in the gaps. Trigger cadence is the real lever: every 5 to 15 minutes is a reasonable lower bound for streaming ingestion, and hourly is fine for dim CDC streams that update infrequently.&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Skew at the surrogate join.&lt;/STRONG&gt;&lt;SPAN&gt; If a small number of customer_id values dominate the order stream (the B2B customer with 90% of the orders, say), the dim/fact join will skew, and one task grinds while the others finish. &lt;/SPAN&gt;&lt;SPAN&gt;CLUSTER BY AUTO&lt;/SPAN&gt;&lt;SPAN&gt; on dim and fact handles most of this by letting Databricks pick clustering columns from real workload. For pathological skew, salt the join key with a small random bucket suffix and aggregate in two stages.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Stream-stream joins without watermarks. This pipeline has no stream-stream join, but if you extend it with one (joining two event streams directly), watermarks matter. Skip the WATERMARK on either side, or the time-bound condition, and state grows without bound. It usually surfaces hours into a long run as memory pressure rather than a clean error, which makes it harder to debug than it should be. Keep a watermark on both sides and a BETWEEN-style time bound in the join condition.&lt;/STRONG&gt;&lt;/P&gt;
&lt;H2&gt;&lt;STRONG&gt;Cost + performance: what this pipeline saves&lt;/STRONG&gt;&lt;/H2&gt;
&lt;P&gt;&lt;SPAN&gt;Here is a like-for-like comparison against the 2024-shaped baseline: a PySpark notebook, a 200-line MERGE, a nightly full-refresh job, and manual orchestration.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;TABLE&gt;
&lt;TBODY&gt;
&lt;TR&gt;
&lt;TD&gt;
&lt;P&gt;&lt;STRONG&gt;Dimension&lt;/STRONG&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;STRONG&gt;2024 baseline&lt;/STRONG&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;STRONG&gt;This pipeline&lt;/STRONG&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;STRONG&gt;Why&lt;/STRONG&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;/TR&gt;
&lt;TR&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;Lines of code&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;~800&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;~200&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;Declarative primitives replace procedural SQL/PySpark&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;/TR&gt;
&lt;TR&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;Refresh model&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;Full nightly recompute on gold&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;Incremental on every trigger&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;Metric-view materialization + streaming silver&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;/TR&gt;
&lt;TR&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;SCD2 dim refresh&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;Full MERGE every batch&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;Incremental, change-only&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;AUTO CDC reads CDF, processes only changes&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;/TR&gt;
&lt;TR&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;Surrogate keys&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;App-side allocation + lock table&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;IDENTITY column&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;Engine-managed, no app state&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;/TR&gt;
&lt;TR&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;Late-arriving line items&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;Custom backfill scripts&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;REPLACE WHERE flow (Gold)&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;One-line idempotent rewrite&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;/TR&gt;
&lt;TR&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;Data quality&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;Custom validation step + log scrape&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;CONSTRAINT … EXPECT&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;First-class in the DDL; metrics in event log&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;/TR&gt;
&lt;TR&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;Data layout&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;Static partitioning + ZORDER jobs&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;CLUSTER BY AUTO&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;Self-tuning, skew-resistant, incremental&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;/TR&gt;
&lt;TR&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;Orchestration&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;Workflows tasks + dependencies wired by hand&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;SDP infers DAG&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;No orchestration code&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;/TR&gt;
&lt;TR&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;Cost-model surprises&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;Silent bill spikes when query changes&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;Metric-view materializations, incremental&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;Governed definitions; view matching auto-accelerates&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;/TR&gt;
&lt;TR&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;CI/CD&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;Hand-rolled deployment scripts&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;Bundle targets&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;One YAML file deploys to dev/staging/prod&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;/TR&gt;
&lt;TR&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;Observability&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;Custom log scraping&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;event_log() SQL view&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;First-class platform feature&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;/TR&gt;
&lt;/TBODY&gt;
&lt;/TABLE&gt;
&lt;P&gt;&lt;SPAN&gt;How much you save depends on your data volume, but the structural change matters more than the dollar figure. Every platinum materialization refresh is incremental. Every silver fact processes only new rows. Dim refresh scales with change volume rather than dim size, and the deployment is reproducible across environments. No single one of these is dramatic on its own; together, they remove a long list of failure modes the 2024 baseline was exposed to.&lt;/SPAN&gt;&lt;/P&gt;
&lt;H2&gt;&lt;STRONG&gt;Requirements&lt;/STRONG&gt;&lt;/H2&gt;
&lt;UL&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Compute:&lt;/STRONG&gt;&lt;SPAN&gt; Serverless SQL Warehouse, or SDP Pro / Advanced edition&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Runtime:&lt;/STRONG&gt;&lt;SPAN&gt; Databricks Runtime 17.3 or later&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Source tables:&lt;/STRONG&gt;&lt;SPAN&gt; Delta with row tracking; Change Data Feed enabled on the dim source&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Permissions:&lt;/STRONG&gt;&lt;SPAN&gt; standard pipeline + table grants on the target catalog&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2&gt;&lt;STRONG&gt;Get started&lt;/STRONG&gt;&lt;/H2&gt;
&lt;UL&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;A href="https://docs.databricks.com/aws/en/ldp/best-practices" target="_blank" rel="noopener"&gt;&lt;SPAN&gt;Best practices for Lakeflow Spark Declarative Pipelines&lt;/SPAN&gt;&lt;/A&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;A href="https://docs.databricks.com/aws/en/delta/selective-overwrite" target="_blank" rel="noopener"&gt;&lt;SPAN&gt;REPLACE WHERE: Delta Lake docs&lt;/SPAN&gt;&lt;/A&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;A href="https://docs.databricks.com/aws/en/ldp/dbsql/streaming#apply-change-data-capture-cdc-with-auto-cdc-flows" target="_blank" rel="noopener"&gt;&lt;SPAN&gt;AUTO CDC in Databricks SQL&lt;/SPAN&gt;&lt;/A&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;A href="https://docs.databricks.com/aws/en/optimizations/incremental-refresh" target="_blank" rel="noopener"&gt;&lt;SPAN&gt;Incremental refresh for materialized views&lt;/SPAN&gt;&lt;/A&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;A href="https://docs.databricks.com/aws/en/sql/language-manual/sql-ref-syntax-ddl-create-streaming-table" target="_blank" rel="noopener"&gt;&lt;SPAN&gt;CREATE STREAMING TABLE: Databricks SQL reference&lt;/SPAN&gt;&lt;/A&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;A href="https://docs.databricks.com/aws/en/metric-views/" target="_blank" rel="noopener"&gt;&lt;SPAN&gt;Metric views with materializations&lt;/SPAN&gt;&lt;/A&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;A href="https://docs.databricks.com/aws/en/ldp/expectations" target="_blank" rel="noopener"&gt;&lt;SPAN&gt;Manage data quality with pipeline expectations&lt;/SPAN&gt;&lt;/A&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;A href="https://docs.databricks.com/aws/en/delta/clustering" target="_blank" rel="noopener"&gt;&lt;SPAN&gt;Use liquid clustering for tables&lt;/SPAN&gt;&lt;/A&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;A href="https://docs.databricks.com/aws/en/dev-tools/bundles/index" target="_blank" rel="noopener"&gt;&lt;SPAN&gt;Declarative Automation Bundles&lt;/SPAN&gt;&lt;/A&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;A href="https://docs.databricks.com/aws/en/release-notes/dlt/2026" target="_blank" rel="noopener"&gt;&lt;SPAN&gt;Lakeflow SDP 2026 release notes&lt;/SPAN&gt;&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;SPAN&gt;The SQL above is a working pipeline (bronze through gold) plus its two metric views. The quickest way to evaluate it is to copy it into your own catalog (renaming the tables to match your data), point the bronze sources at your real volumes, and watch the Pipeline UI render the DAG. If you replace an existing PySpark-plus-MERGE pipeline with this pattern, the numbers worth sharing in the comments are line count, end-to-end refresh time, and monthly cost, before and after.&lt;/SPAN&gt;&lt;/P&gt;</description>
      <pubDate>Tue, 06 Oct 2026 10:23:39 GMT</pubDate>
      <guid>https://community.databricks.com/t5/technical-blog/rebuilding-an-800-line-pyspark-pipeline-in-200-lines-of/ba-p/168578</guid>
      <dc:creator>Fangzhu</dc:creator>
      <dc:date>2026-10-06T10:23:39Z</dc:date>
    </item>
    <item>
      <title>TokenMaxing to ValueMaxing with Databricks</title>
      <link>https://community.databricks.com/t5/technical-blog/tokenmaxing-to-valuemaxing-with-databricks/ba-p/168850</link>
      <description>&lt;P&gt;The goal isn't more AI usage. It's smarter AI usage. But teams keep measuring success by tokens consumed, so spend climbs while visibility does not. Route every model and agent call through the Unity Gateway and each call becomes governed data you can see, attribute, and act on. Here is how those same logs became GatewayIQ.&lt;/P&gt;</description>
      <pubDate>Mon, 05 Oct 2026 15:14:56 GMT</pubDate>
      <guid>https://community.databricks.com/t5/technical-blog/tokenmaxing-to-valuemaxing-with-databricks/ba-p/168850</guid>
      <dc:creator>prachi_bari</dc:creator>
      <dc:date>2026-10-05T15:14:56Z</dc:date>
    </item>
    <item>
      <title>Build Your First Databricks App with Codex: A Practical Setup Guide</title>
      <link>https://community.databricks.com/t5/technical-blog/build-your-first-databricks-app-with-codex-a-practical-setup/ba-p/168713</link>
      <description>&lt;P&gt;&lt;SPAN&gt;This guide walks through how to connect OpenAI Codex CLI to Databricks, create a simple Databricks App, develop it locally, and deploy it to your workspace. It starts with the recommended &lt;/SPAN&gt;&lt;SPAN&gt;ucode&lt;/SPAN&gt;&lt;SPAN&gt; setup and then covers manual Codex configuration for teams that want more control.&lt;/SPAN&gt;&lt;/P&gt;</description>
      <pubDate>Mon, 05 Oct 2026 09:04:02 GMT</pubDate>
      <guid>https://community.databricks.com/t5/technical-blog/build-your-first-databricks-app-with-codex-a-practical-setup/ba-p/168713</guid>
      <dc:creator>epandya</dc:creator>
      <dc:date>2026-10-05T09:04:02Z</dc:date>
    </item>
    <item>
      <title>How to choose the right AI agent path on Databricks</title>
      <link>https://community.databricks.com/t5/technical-blog/how-to-choose-the-right-ai-agent-path-on-databricks/ba-p/168706</link>
      <description>&lt;P&gt;Navigating the world of AI agents isn't one-size-fits-all. Discover how Databricks supports three distinct agent paths—governed Genie Agents for domain expertise, Genie One + Skills for everyday business assistance, and code-first Agent Bricks for custom agentic applications—so you can start with your primary need and scale as your requirements grow.&lt;/P&gt;</description>
      <pubDate>Mon, 05 Oct 2026 08:57:18 GMT</pubDate>
      <guid>https://community.databricks.com/t5/technical-blog/how-to-choose-the-right-ai-agent-path-on-databricks/ba-p/168706</guid>
      <dc:creator>epandya</dc:creator>
      <dc:date>2026-10-05T08:57:18Z</dc:date>
    </item>
    <item>
      <title>Step-by-Step: Building a Vacation Rental Operations App with AppKit</title>
      <link>https://community.databricks.com/t5/technical-blog/step-by-step-building-a-vacation-rental-operations-app-with/ba-p/168709</link>
      <description>&lt;P&gt;&lt;SPAN&gt;This post builds a real operations app using `samples.wanderbricks` — a vacation rental marketplace dataset that ships with every Databricks workspace. It has 16 tables covering bookings, users, properties, payments, reviews, destinations, clickstream, and more.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;All code snippets live at &lt;/SPAN&gt;&lt;A href="https://github.com/databricks/devhub/tree/main/examples/vacation-rentals/blog-post-snippets"&gt;&lt;SPAN&gt;github.com/databricks/devhub&lt;/SPAN&gt;&lt;/A&gt;&lt;SPAN&gt;. You don’t need to clone the repo — we’ll `curl` individual files as we go.&lt;/SPAN&gt;&lt;/P&gt;</description>
      <pubDate>Wed, 30 Sep 2026 09:42:08 GMT</pubDate>
      <guid>https://community.databricks.com/t5/technical-blog/step-by-step-building-a-vacation-rental-operations-app-with/ba-p/168709</guid>
      <dc:creator>epandya</dc:creator>
      <dc:date>2026-09-30T09:42:08Z</dc:date>
    </item>
    <item>
      <title>Bricktalk Recording | Real-Time Data &amp; AI: Tripwise Demo</title>
      <link>https://community.databricks.com/t5/bricktalks-tv/bricktalk-recording-real-time-data-amp-ai-tripwise-demo/ba-p/169826</link>
      <description>&lt;P&gt;Explore real-time data and AI architectures in this BrickTalk with&amp;nbsp;&lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/260093"&gt;@Hunter_Walker&lt;/a&gt;&amp;nbsp;and&amp;nbsp;&lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/260014"&gt;@ameevora&lt;/a&gt;&amp;nbsp;&lt;BR /&gt;&lt;div class="video-embed-center video-embed"&gt;&lt;iframe class="embedly-embed" src="https://cdn.embedly.com/widgets/media.html?src=https%3A%2F%2Fwww.youtube.com%2Fembed%2FACDS3_yZPrc%3Ffeature%3Doembed&amp;amp;display_name=YouTube&amp;amp;url=https%3A%2F%2Fwww.youtube.com%2Fwatch%3Fv%3DACDS3_yZPrc&amp;amp;image=https%3A%2F%2Fi.ytimg.com%2Fvi%2FACDS3_yZPrc%2Fhqdefault.jpg&amp;amp;type=text%2Fhtml&amp;amp;schema=youtube" width="600" height="337" scrolling="no" title="Community BrickTalk | Real-Time Data &amp;amp; AI: Tripwise Demo" frameborder="0" allow="autoplay; fullscreen; encrypted-media; picture-in-picture" allowfullscreen="true"&gt;&lt;/iframe&gt;&lt;/div&gt;&lt;/P&gt;</description>
      <pubDate>Fri, 25 Sep 2026 12:20:34 GMT</pubDate>
      <guid>https://community.databricks.com/t5/bricktalks-tv/bricktalk-recording-real-time-data-amp-ai-tripwise-demo/ba-p/169826</guid>
      <dc:creator>Tushar_Parekar</dc:creator>
      <dc:date>2026-09-25T12:20:34Z</dc:date>
    </item>
    <item>
      <title>Multi-region model serving on Databricks with OpenSharing</title>
      <link>https://community.databricks.com/t5/technical-blog/multi-region-model-serving-on-databricks-with-opensharing/ba-p/167622</link>
      <description>&lt;P&gt;&lt;SPAN&gt;How to extend a single-region model serving stack to additional regions using serverless workspaces, OpenSharing, Real-Time Serving Endpoints, and Lakebase.&lt;/SPAN&gt;&lt;/P&gt;</description>
      <pubDate>Mon, 21 Sep 2026 08:47:50 GMT</pubDate>
      <guid>https://community.databricks.com/t5/technical-blog/multi-region-model-serving-on-databricks-with-opensharing/ba-p/167622</guid>
      <dc:creator>KamLook</dc:creator>
      <dc:date>2026-09-21T08:47:50Z</dc:date>
    </item>
    <item>
      <title>BrickTalk Recording | One Platform, Any Source: Unifying Enterprise Data with Lakeflow Connect</title>
      <link>https://community.databricks.com/t5/bricktalks-tv/bricktalk-recording-one-platform-any-source-unifying-enterprise/ba-p/169070</link>
      <description>&lt;P&gt;In this Community BrickTalk,&amp;nbsp;&lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/67838"&gt;@Giselle_Go_DB&lt;/a&gt;&amp;nbsp; and&amp;nbsp;&lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/231837"&gt;@sonia_bendre&lt;/a&gt;&amp;nbsp;&amp;nbsp;demonstrate unifying enterprise data using Databricks Lakeflow Connect.&lt;BR /&gt;&lt;div class="video-embed-center video-embed"&gt;&lt;iframe class="embedly-embed" src="https://cdn.embedly.com/widgets/media.html?src=https%3A%2F%2Fwww.youtube.com%2Fembed%2FcZLaSE-jzos%3Ffeature%3Doembed&amp;amp;display_name=YouTube&amp;amp;url=https%3A%2F%2Fwww.youtube.com%2Fwatch%3Fv%3DcZLaSE-jzos&amp;amp;image=https%3A%2F%2Fi.ytimg.com%2Fvi%2FcZLaSE-jzos%2Fhqdefault.jpg&amp;amp;type=text%2Fhtml&amp;amp;schema=youtube" width="600" height="337" scrolling="no" title="Community BrickTalk | One Platform, Any Source: Unifying Enterprise Data with Lakeflow Connect" frameborder="0" allow="autoplay; fullscreen; encrypted-media; picture-in-picture" allowfullscreen="true"&gt;&lt;/iframe&gt;&lt;/div&gt;&lt;/P&gt;</description>
      <pubDate>Fri, 18 Sep 2026 09:07:01 GMT</pubDate>
      <guid>https://community.databricks.com/t5/bricktalks-tv/bricktalk-recording-one-platform-any-source-unifying-enterprise/ba-p/169070</guid>
      <dc:creator>Tushar_Parekar</dc:creator>
      <dc:date>2026-09-18T09:07:01Z</dc:date>
    </item>
    <item>
      <title>Tutorial: Databricks Genie for Data Engineers and Data Scientists</title>
      <link>https://community.databricks.com/t5/technical-blog/tutorial-databricks-genie-for-data-engineers-and-data-scientists/ba-p/168969</link>
      <description>&lt;P&gt;&lt;STRONG&gt;From a table to a running app: this hands-on tutorial takes you through the full data life cycle on 696 million records.&lt;/STRONG&gt; You access the data through Databricks Marketplace, run EDA with Genie Code, explore it in plain English with Genie Agents, generate a Spark Declarative Pipeline and a Lakeflow Job, and build a Databricks App. At every stage you learn how to prompt Genie, review its output, and decide where your own judgement takes over. Runs serverless on Databricks Free Edition.&lt;/P&gt;</description>
      <pubDate>Fri, 18 Sep 2026 08:57:28 GMT</pubDate>
      <guid>https://community.databricks.com/t5/technical-blog/tutorial-databricks-genie-for-data-engineers-and-data-scientists/ba-p/168969</guid>
      <dc:creator>DataAlchemist28</dc:creator>
      <dc:date>2026-09-18T08:57:28Z</dc:date>
    </item>
  </channel>
</rss>

