cancel
Showing results for 
Search instead for 
Did you mean: 
Data Engineering
Join discussions on data engineering best practices, architectures, and optimization strategies within the Databricks Community. Exchange insights and solutions with fellow data engineers.
cancel
Showing results for 
Search instead for 
Did you mean: 

Forum Posts

Brahmareddy
by • Esteemed Contributor II
  • 3781 Views
  • 8 replies
  • 12 kudos

Future of Movie Discovery: How I Built an AI Movie Recommendation Agent on Databricks Free Edition

As a data engineer deeply passionate about how data and AI can come together to create real-world impact, I’m excited to share my project for the Databricks Free Edition Hackathon 2025 — Future of Movie Discovery (FMD). Built entirely on Databricks F...

  • 3781 Views
  • 8 replies
  • 12 kudos
Latest Reply
juanlozadab
New Contributor II
  • 12 kudos

@Brahmareddy I haven't used this exact pattern at multi-million-row scale, but I agree that the tradeoffs start to change as the dataset and query volume grow.I would probably keep Delta as the source of truth and separate the serving layer based on ...

  • 12 kudos
7 More Replies
Davide
by • New Contributor II
  • 84 Views
  • 1 replies
  • 1 kudos

SET MANAGED AND ROLLBACK IN RELATION TO VACUUM

Hello everyone,I am looking for clarification about the 14-day rollback period after converting an external table to a managed table using: ALTER TABLE table_name SET MANAGED.Databricks documentation states that the conversion can be rolled back with...

  • 84 Views
  • 1 replies
  • 1 kudos
Latest Reply
DoTA
Valued Contributor II
  • 1 kudos

Hi, I could not find these three cases documented either, so I will separate what the docs say from what I would treat as an assumption. What is documented (Convert external tables to managed): Databricks keeps the data in the old external location f...

  • 1 kudos
srikanthp24
by • New Contributor III
  • 74 Views
  • 1 replies
  • 0 kudos

Regarding the DAB Ingestion Pipeline Creating the Volume

Hello Guys,By taking the below yaml code as a reference I have created the databrick asset bundle for data transfer from sql server to databricks. But with below code the staging volume getting created by gateway pipeline automatically, which unable ...

srikanthp24_0-1790135779415.png
  • 74 Views
  • 1 replies
  • 0 kudos
Latest Reply
DoTA
Valued Contributor II
  • 0 kudos

Hi, a few things that may unblock you here. 1) Custom volume name. In gateway_definition you can set gateway_storage_catalog, gateway_storage_schema and gateway_storage_name. If you leave gateway_storage_name out, the gateway auto-generates a volume ...

  • 0 kudos
Khasim_1
by • New Contributor III
  • 107 Views
  • 1 replies
  • 0 kudos

Implementing a "Zero-Bus" Architecture with Unity Catalog

Hi everyone,I am currently refining the architecture for an end-to-end Lakehouse project and moving toward a "Zero-Bus" (Bus Architecture) model. My primary objective is to enforce dimension conformity (Customer, Product, Geography) across multiple i...

  • 107 Views
  • 1 replies
  • 0 kudos
Latest Reply
DoTA
Valued Contributor II
  • 0 kudos

Hi @Khasim_1, here is the pattern I would use for conformed dimensions on Unity Catalog, taking your questions in order. 1) Layering. I would avoid a "Master Silver" that everything funnels through, because it tends to become a bottleneck team. Treat...

  • 0 kudos
cdn_yyz_yul
by • Contributor III
  • 426 Views
  • 11 replies
  • 1 kudos

Change data feed from a materialized view

Hi everyone,Trying the CDF on materialized view (declarative pipeline) feature that is currently in Beta.- Configuration:** Verified the MVs (materialized veiws) have: 'delta.enableRowTracking'SHOW TBLPROPERTIES my_mv ('delta.enableRowTracking');  re...

  • 426 Views
  • 11 replies
  • 1 kudos
Latest Reply
cdn_yyz_yul
Contributor III
  • 1 kudos

The feature that will help the most is the CDF on MV that is in beta right now. I have been testing it to get ready at my side and take advantage of it once it is in General Availability. The Automatic change data feed on Streaming tables would be us...

  • 1 kudos
10 More Replies
Islam_hoti
by • New Contributor III
  • 125 Views
  • 1 replies
  • 2 kudos

Resolved! Streaming read fails with "Detected a data update" after a restatement job touches the source

Hi everyone,Looking for the right pattern here rather than a workaround.Setup. DBR 15.4 LTS, Unity Catalog. A silver streaming table reads from a bronze Delta table with a normal streaming read. Bronze is append only in the ordinary course of busines...

  • 125 Views
  • 1 replies
  • 2 kudos
Latest Reply
Khasim_1
New Contributor III
  • 2 kudos

Hi @Islam_hoti ,You have hit the classic 'Streaming vs. Mutation' wall. readStream on a Delta table assumes an append-only sequence; once an UPDATE happens, the commit versioning sequence is invalidated, hence the loud failure.Here is how we reasoned...

  • 2 kudos
Islam_hoti
by • New Contributor III
  • 169 Views
  • 3 replies
  • 5 kudos

Resolved! Photon enabled but a large share of the plan is falling back, cost up and runtime flat

Hi everyone,Trying to work out whether this is expected or whether I have misconfigured something.We enabled Photon on a job cluster running a nightly aggregation over roughly 2TB. The expectation was the usual improvement. What we got instead was ru...

  • 169 Views
  • 3 replies
  • 5 kudos
Latest Reply
Khasim_1
New Contributor III
  • 5 kudos

Hi @Islam_hoti ,Photon is fantastic, but it’s an 'all-or-nothing' value proposition. Once you hit a fallback, you lose the vectorized execution benefit for that entire sub-tree, and you're still paying the DBU premium. Here is how we approach auditin...

  • 5 kudos
2 More Replies
Islam_hoti
by • New Contributor III
  • 160 Views
  • 3 replies
  • 4 kudos

Resolved! At what data size do you stop reaching for Spark?

Hi everyone,A question I keep having with my team and I would like to hear how others think about it.A lot of the jobs we run are not big. Plenty of our pipelines process a few gigabytes, some considerably less. We run them on Spark because that is w...

  • 160 Views
  • 3 replies
  • 4 kudos
Latest Reply
Khasim_1
New Contributor III
  • 4 kudos

Hi @Islam_hoti ,This is the 'Silent Architect' dilemma—when the standard tool is technically suboptimal but organizationally necessary. We’ve wrestled with this, and here’s where we landed:The 'Standardization Premium' is real: We reasoned that the c...

  • 4 kudos
2 More Replies
Mado
by • Valued Contributor II
  • 276 Views
  • 6 replies
  • 7 kudos

How can I configure Lakeflow Connect SQL Server CDC Gateway to use a desired VM type?

Hi Team,I'm evaluating Databricks Lakeflow Connect for SQL Server CDC ingestion and have run into a gateway provisioning issue.EnvironmentRegion: Australia EastSource: Azure SQL DatabaseCDC enabled successfully on the database and source tableSQL Ser...

Mado_0-1790214214728.png
  • 276 Views
  • 6 replies
  • 7 kudos
Latest Reply
Khasim_1
New Contributor III
  • 7 kudos

Hi @Mado As of now, Lakeflow Connect's CDC Gateway VM types (driver/worker SKUs) are not configurable at the pipeline or connection level — they are managed internally by the platform. If the default SKU (Standard_E4d_v4) is unavailable in your regio...

  • 7 kudos
5 More Replies
dbernstein_tp
by • Contributor
  • 132 Views
  • 3 replies
  • 1 kudos

Lakeflow connect SQL server ingestion can be made elastic?

Hi Everyone, One of our big ingestion tasks is lakeflow connect CDC ingestion of ERP data from a SQL server database. I am deploying the pipelines and jobs for this via DABs. We are ingesting about 80 tables from the database, a handful of which are ...

  • 132 Views
  • 3 replies
  • 1 kudos
Latest Reply
balajij8
Esteemed Contributor II
  • 1 kudos

If you need 15 minute CDC interval - Migrate those specific jobs to the integrated CDC pipeline in continuous speed-optimized mode. If the 15 minute jobs involves fewer than 50 tables, it works directly. If not, split them into multiple pipelines. Mo...

  • 1 kudos
2 More Replies
Hariharan_0510
by • New Contributor
  • 124 Views
  • 2 replies
  • 0 kudos

1.Not able to access interent (gcp databricks) 2.both ingress and egress access given

1.Not able to access interent (gcp databricks) 2.both ingress and egress access given in default policy2.How to overcome this issue and  i am using serverless3.in free databricks able to use interent in gcp based databricks it is creating issueimport...

  • 124 Views
  • 2 replies
  • 0 kudos
Latest Reply
Louis_Frolio
Databricks Employee
  • 0 kudos

Hello @Hariharan_0510 , I took a look at both internal and external documentation and here is what I found. Building on @balajij8's reply, the serverless egress policy is the right place to start, but two things before you go hunting further. First, ...

  • 0 kudos
1 More Replies
Khasim_1
by • New Contributor III
  • 146 Views
  • 2 replies
  • 2 kudos

Implementing a "Zero-Bus" Architecture in a Lakehouse: Best Practices for Shared Dimensions?

Hi Everyone,I am currently architecting a Medallion Lakehouse and moving toward a "Zero-Bus" (Bus Architecture) approach. My goal is to ensure that our core dimensions—such as Customer, Product, and Geography—remain consistent and conformed across al...

  • 146 Views
  • 2 replies
  • 2 kudos
Latest Reply
Islam_hoti
New Contributor III
  • 2 kudos

Hi,Great question. This is where most Medallion implementations either scale or get messy. Here's the pattern that has worked well, based on Kimball conformed dimensions mapped onto Unity Catalog.Quick note on naming: in Databricks, "Zerobus" is a pr...

  • 2 kudos
1 More Replies
yashojha
by • New Contributor III
  • 2070 Views
  • 4 replies
  • 0 kudos

Resolved! Slow writes to managed volume

Hi All, I am using managed volumes as an intermediate storage to write a decrypted file before moving to data lake storage. Strangely the write operation is taking a lot of time (22 mins) to write a small file to volumes and it takes only few seconds...

yashojha_1-1771488017215.png yashojha_3-1771488035100.png
  • 2070 Views
  • 4 replies
  • 0 kudos
Latest Reply
yashojha
New Contributor III
  • 0 kudos

Hi All, Thank you for you advice regarding this issue, but the root cause was something else. We had 2 private endpoints for ADLS (BLOB and DFS), and the firewall was only open for one BLOB PE. As a process Databricks first tries to resolve connectio...

  • 0 kudos
3 More Replies
mmsanteago
by • New Contributor II
  • 194 Views
  • 3 replies
  • 1 kudos

Databricks AI/BI Get Session User

Hello everyone,I am currently working with Databricks AI/BI Dashboards and I have a specific requirement regarding user context.I need to dynamically capture the "session user" (the person who is currently viewing the dashboard) rather than the "exec...

  • 194 Views
  • 3 replies
  • 1 kudos
Latest Reply
Islam_hoti
New Contributor III
  • 1 kudos

Hi,Short answer: No. There's no built-in function or parameter that returns the viewer's identity while a dashboard runs on the publisher's credentials.Why: When you publish with Share data permissions (the default, formerly called "Embed credentials...

  • 1 kudos
2 More Replies
dawidch
by • New Contributor
  • 327 Views
  • 6 replies
  • 5 kudos

Resolved! org.apache.iceberg.connect.IcebergSinkConnector to sink data from kafka to databricks - large volume

Hey.We have different number of topics per let's call it "subject". The connector works great for us if there are not so many topics (partitions) in subject. We have issue when we try to sink 1500 topics, 3 partitions each. We've sharded the topics i...

  • 327 Views
  • 6 replies
  • 5 kudos
Latest Reply
dawidch
New Contributor
  • 5 kudos

Hello All.Thank you all for help. Here is what worked for us:## Original symptomsEach connector started normally and consumed records for more than an hour. Kafka Connect reportedthe connector and tasks as `RUNNING`, and source-topic lag decreased. T...

  • 5 kudos
5 More Replies
Labels