cancel
Showing results for 
Search instead for 
Did you mean: 
Community Articles
Dive into a collaborative space where members like YOU can exchange knowledge, tips, and best practices. Join the conversation today and unlock a wealth of collective wisdom to enhance your experience and drive success.
cancel
Showing results for 
Search instead for 
Did you mean: 

Forum Posts

Phani_sannala
by New Contributor II
  • 502 Views
  • 5 replies
  • 16 kudos

From 40 Minutes to 8 minutes: Why We Dropped MERGE in Our SAP BW to Databricks Gold Layer

Infra keeps getting faster, but the fix for a slow notebook is usually a config line, not a bigger cluster.One Spark Config, 32 Minutes Saved: Replacing MERGE with Dynamic Partition OverwriteBy @Phani_sannala , co-authored with @sridharplv Our gold n...

  • 502 Views
  • 5 replies
  • 16 kudos
Latest Reply
ozaaditya
Databricks Partner
  • 16 kudos

The zombie row bug is the real argument here reframing this from "MERGE is slow" to "MERGE is silently wrong for full-slice delivery" is what makes it land. Only thing I'd stress: dynamic partition overwrite is exactly as safe as your completeness ch...

  • 16 kudos
4 More Replies
MouR
by Databricks Partner
  • 574 Views
  • 0 replies
  • 1 kudos

Stop Translating Alteryx Boxes - A Lakebridge-assisted, test-driven migration to Azure Databricks

Lakebridge can assess an Alteryx estate, but the current support matrix does not list Alteryx for automated conversion or direct reconciliation. A safe migration therefore combines Lakebridge assessment with deliberate redesign, native Databricks eng...

mou_0-1784500718650.png mou_1-1784500718661.png
  • 574 Views
  • 0 replies
  • 1 kudos
GabFernandes
by Contributor
  • 931 Views
  • 0 replies
  • 0 kudos

Apache Spark 4.2 is officially here! Key architectural updates for AI-Native & Governed Platforms

Hi community!Matei Zaharia and the Databricks team just announced the release of Apache Spark 4.2. As a Data Architect, seeing how this engine is evolving to bridge the gap between traditional data engineering, governance, and the AI era is incredibl...

  • 931 Views
  • 0 replies
  • 0 kudos
naveenayalla
by New Contributor III
  • 704 Views
  • 1 replies
  • 0 kudos

From RAG Demo to Production on Databricks: 7 Things Teams Should Validate First

From RAG Demo to Production on Databricks: 7 Things Teams Should Validate FirstBy Naveen AyallaMany teams can build a RAG demo quickly.Upload documents, create embeddings, connect a model, ask a question, and show an answer.But production is differen...

naveen0808_0-1780880239856.png
  • 704 Views
  • 1 replies
  • 0 kudos
Latest Reply
naveenayalla
New Contributor III
  • 0 kudos

Thanks for reading. I’m especially interested in hearing from people who have worked on real RAG or GenAI workflows.Which one has been the biggest challenge for your team?1. Choosing the right source data2. Access control and governance3. Improving r...

  • 0 kudos
Avinash_Narala
by Databricks Partner
  • 550 Views
  • 0 replies
  • 1 kudos

How we solved the "18-Hour Running Job" problem with Data-Driven Timeouts

Hi everyone,I recently dealt with a frustrating scenario: a Databricks job that usually takes minutes ran for 18 hours without failing, quietly consuming compute and blocking downstream pipelines.The driver hadn't crashed, and the job hadn't failed—i...

  • 550 Views
  • 0 replies
  • 1 kudos
Kirankumarbs
by Valued Contributor III
  • 654 Views
  • 0 replies
  • 1 kudos

One Cluster per Task — Proven, Ready, and Waiting

Part 3 of 3: Databricks Streaming ArchitectureBy the end of Part 1 & Part 2, we knew what the real answer was. We just hadn’t committed to it yet.Not because it wouldn’t work. We tested it. We documented it. The code was ready. The answer was one clu...

  • 654 Views
  • 0 replies
  • 1 kudos
Ale_Armillotta
by Valued Contributor II
  • 11244 Views
  • 3 replies
  • 6 kudos

Resolved! CI/CD on Databricks with Asset Bundles (DABs) and GitHub Actions

Hi all.If you've ever manually promoted resources from dev to prod on Databricks — copying notebooks, updating configs, hoping nothing breaks — this post is for you.I've been building a CI/CD setup for a Speech-to-Text pipeline on Databricks, and I w...

Community Articles
CICD
DABs
GitHub
  • 11244 Views
  • 3 replies
  • 6 kudos
Latest Reply
SteveOstrowski
Databricks Employee
  • 6 kudos

Hi, Great question! Databricks Asset Bundles (DABs) are the recommended approach for CI/CD on Databricks. Here is a comprehensive walkthrough. WHAT ARE DATABRICKS ASSET BUNDLES? DABs let you define your Databricks resources (jobs, pipelines, dashboar...

  • 6 kudos
2 More Replies
kanikvijay9
by Contributor
  • 3900 Views
  • 2 replies
  • 10 kudos

Optimizing Delta Table Writes for Massive Datasets in Databricks

Problem StatementIn one of my recent projects, I faced a significant challenge: Writing a huge dataset of 11,582,763,212 rows and 2,068 columns to a Databricks managed Delta table.The initial write operation took 22.4 hours using the following setup:...

kanikvijay9_0-1762695454233.png kanikvijay9_1-1762695506126.png kanikvijay9_2-1762695536800.png kanikvijay9_3-1762695573841.png
  • 3900 Views
  • 2 replies
  • 10 kudos
Latest Reply
kanikvijay9
Contributor
  • 10 kudos

Hey @Louis_Frolio ,Thank you for the thoughtful feedback and great suggestions!A few clarifications:AQE is already enabled in my setup, and it definitely helped reduce shuffle overhead during the write.Regarding Column Pruning, in this case, the fina...

  • 10 kudos
1 More Replies
Labels