cancel
Showing results for 
Search instead for 
Did you mean: 
Community Articles
Dive into a collaborative space where members like YOU can exchange knowledge, tips, and best practices. Join the conversation today and unlock a wealth of collective wisdom to enhance your experience and drive success.
cancel
Showing results for 
Search instead for 
Did you mean: 

Forum Posts

Phani_sannala
by New Contributor
  • 115 Views
  • 4 replies
  • 10 kudos

From 40 Minutes to 8 minutes: Why We Dropped MERGE in Our SAP BW to Databricks Gold Layer

Infra keeps getting faster, but the fix for a slow notebook is usually a config line, not a bigger cluster.One Spark Config, 32 Minutes Saved: Replacing MERGE with Dynamic Partition OverwriteBy @Phani_sannala , co-authored with @sridharplv Our gold n...

  • 115 Views
  • 4 replies
  • 10 kudos
Latest Reply
VinayKumarB
Databricks Partner
  • 10 kudos

Great work, keep posting!!

  • 10 kudos
3 More Replies
TriambakR
by New Contributor III
  • 1049 Views
  • 3 replies
  • 9 kudos

Medallion Architecture in Practice: The Design Decisions Nobody Puts in the Diagram

Medallion Architecture in Practice: The Design Decisions Nobody Puts in the DiagramEvery Lakehouse conversation eventually shows the same three boxes: Bronze, Silver, Gold. It's a great mental model — but on a real enterprise migration, the diagram i...

  • 1049 Views
  • 3 replies
  • 9 kudos
Latest Reply
Sbm_dracarys
New Contributor II
  • 9 kudos

Great work, man! Even though I don't know much about this field, your article made me curious and motivated me to read more about it. Thanks for sharing such valuable insights.keep posting  

  • 9 kudos
2 More Replies
savlahanish27
by Databricks Partner
  • 462 Views
  • 0 replies
  • 0 kudos

Gold Layer Design on Databricks — MERGE vs Overwrite, Partitioning, SCD Type 2 from SAP

Part 3 of my series on building an enterprise data platform on Databricks is up - this one cover Gold layer design.The short version: Gold isn't just aggregated Silver. Silver maps to your source system. Gold maps to the business questions your consu...

  • 462 Views
  • 0 replies
  • 0 kudos
naveenayalla
by New Contributor III
  • 657 Views
  • 1 replies
  • 0 kudos

From RAG Demo to Production on Databricks: 7 Things Teams Should Validate First

From RAG Demo to Production on Databricks: 7 Things Teams Should Validate FirstBy Naveen AyallaMany teams can build a RAG demo quickly.Upload documents, create embeddings, connect a model, ask a question, and show an answer.But production is differen...

naveen0808_0-1780880239856.png
  • 657 Views
  • 1 replies
  • 0 kudos
Latest Reply
naveenayalla
New Contributor III
  • 0 kudos

Thanks for reading. I’m especially interested in hearing from people who have worked on real RAG or GenAI workflows.Which one has been the biggest challenge for your team?1. Choosing the right source data2. Access control and governance3. Improving r...

  • 0 kudos
Avinash_Narala
by Databricks Partner
  • 619 Views
  • 0 replies
  • 0 kudos

Managed vs External Tables in Unity Catalog: The Decision That’s Silently Inflating Your Cloud Bill

Hi everyone,I recently took a look into a silent cost driver in many data platforms: the default choice between managed and external tables in Unity Catalog.It is very common for teams to default to external tables, but this choice often leads to acc...

  • 619 Views
  • 0 replies
  • 0 kudos
Avinash_Narala
by Databricks Partner
  • 526 Views
  • 0 replies
  • 1 kudos

How we solved the "18-Hour Running Job" problem with Data-Driven Timeouts

Hi everyone,I recently dealt with a frustrating scenario: a Databricks job that usually takes minutes ran for 18 hours without failing, quietly consuming compute and blocking downstream pipelines.The driver hadn't crashed, and the job hadn't failed—i...

  • 526 Views
  • 0 replies
  • 1 kudos
Ashwin_DSA
by Databricks Employee
  • 988 Views
  • 0 replies
  • 2 kudos

Databricks Multi-Table Transactions - Part 2

In Part 1, we covered why multi-table transactions matter. Now let's build one. We'll create the tables from the claim wrap-up scenario, load sample P&C insurance data, and walk through what happens when the wrap-up succeeds, when it fails, and when...

s1-claim.png s1-wraplog.png s1-reserves.png s2-error.png
  • 988 Views
  • 0 replies
  • 2 kudos
Kirankumarbs
by Valued Contributor III
  • 624 Views
  • 0 replies
  • 1 kudos

One Cluster per Task — Proven, Ready, and Waiting

Part 3 of 3: Databricks Streaming ArchitectureBy the end of Part 1 & Part 2, we knew what the real answer was. We just hadn’t committed to it yet.Not because it wouldn’t work. We tested it. We documented it. The code was ready. The answer was one clu...

  • 624 Views
  • 0 replies
  • 1 kudos
Ale_Armillotta
by Valued Contributor II
  • 10749 Views
  • 3 replies
  • 6 kudos

Resolved! CI/CD on Databricks with Asset Bundles (DABs) and GitHub Actions

Hi all.If you've ever manually promoted resources from dev to prod on Databricks — copying notebooks, updating configs, hoping nothing breaks — this post is for you.I've been building a CI/CD setup for a Speech-to-Text pipeline on Databricks, and I w...

Community Articles
CICD
DABs
GitHub
  • 10749 Views
  • 3 replies
  • 6 kudos
Latest Reply
SteveOstrowski
Databricks Employee
  • 6 kudos

Hi, Great question! Databricks Asset Bundles (DABs) are the recommended approach for CI/CD on Databricks. Here is a comprehensive walkthrough. WHAT ARE DATABRICKS ASSET BUNDLES? DABs let you define your Databricks resources (jobs, pipelines, dashboar...

  • 6 kudos
2 More Replies
AbhaySingh
by Databricks Employee
  • 4141 Views
  • 0 replies
  • 1 kudos

Delta Lake 4.0 in the Real World

Delta Lake 4.0 is the next major open-source release aligned with Spark 4.x, adding first-class Variant for semi-structured data, safer Type Widening, improved DROP FEATURE, better transaction log handling, and a new multi-engine story via Delta Kern...

  • 4141 Views
  • 0 replies
  • 1 kudos
kanikvijay9
by Contributor
  • 3783 Views
  • 2 replies
  • 10 kudos

Optimizing Delta Table Writes for Massive Datasets in Databricks

Problem StatementIn one of my recent projects, I faced a significant challenge: Writing a huge dataset of 11,582,763,212 rows and 2,068 columns to a Databricks managed Delta table.The initial write operation took 22.4 hours using the following setup:...

kanikvijay9_0-1762695454233.png kanikvijay9_1-1762695506126.png kanikvijay9_2-1762695536800.png kanikvijay9_3-1762695573841.png
  • 3783 Views
  • 2 replies
  • 10 kudos
Latest Reply
kanikvijay9
Contributor
  • 10 kudos

Hey @Louis_Frolio ,Thank you for the thoughtful feedback and great suggestions!A few clarifications:AQE is already enabled in my setup, and it definitely helped reduce shuffle overhead during the write.Regarding Column Pruning, in this case, the fina...

  • 10 kudos
1 More Replies
Labels