cancel
Showing results for 
Search instead for 
Did you mean: 
Data Engineering
Join discussions on data engineering best practices, architectures, and optimization strategies within the Databricks Community. Exchange insights and solutions with fellow data engineers.
cancel
Showing results for 
Search instead for 
Did you mean: 

Forum Posts

dbernstein_tp
by • Contributor
  • 156 Views
  • 4 replies
  • 1 kudos

Lakeflow connect SQL server ingestion can be made elastic?

Hi Everyone, One of our big ingestion tasks is lakeflow connect CDC ingestion of ERP data from a SQL server database. I am deploying the pipelines and jobs for this via DABs. We are ingesting about 80 tables from the database, a handful of which are ...

  • 156 Views
  • 4 replies
  • 1 kudos
Latest Reply
Louis_Frolio
Databricks Employee
  • 1 kudos

Greetings @dbernstein_tp, I did some digging and here is what I found. Short answer: no, the gateway can't scale out horizontally, and that's by design rather than a DAB problem. The distinction that matters is Spark worker autoscaling versus the con...

  • 1 kudos
3 More Replies
priya9896
by • New Contributor
  • 127 Views
  • 2 replies
  • 0 kudos

Azure to AWS

Hi everyone,We're evaluating an architecture pattern and would appreciate any guidance or recommendations. Has anyone implemented a similar cross-cloud pattern? Specifically:Can AWS Glue be used to access Azure Databricks data without copying it?Are ...

  • 127 Views
  • 2 replies
  • 0 kudos
Latest Reply
Louis_Frolio
Databricks Employee
  • 0 kudos

Hello @priya9896, I took a look at both internal and external documentation and here is what I found. Straight answer: don't plan on the Glue Data Catalog as a federation layer over Azure Databricks. Glue's Delta integration and the Glue Catalog itse...

  • 0 kudos
1 More Replies
Brahmareddy
by • Esteemed Contributor II
  • 3842 Views
  • 9 replies
  • 12 kudos

Future of Movie Discovery: How I Built an AI Movie Recommendation Agent on Databricks Free Edition

As a data engineer deeply passionate about how data and AI can come together to create real-world impact, I’m excited to share my project for the Databricks Free Edition Hackathon 2025 — Future of Movie Discovery (FMD). Built entirely on Databricks F...

  • 3842 Views
  • 9 replies
  • 12 kudos
Latest Reply
juanlozadab
New Contributor II
  • 12 kudos

@Brahmareddy I haven't used this exact pattern at multi-million-row scale, but I agree that the tradeoffs start to change as the dataset and query volume grow.I would probably keep Delta as the source of truth and separate the serving layer based on ...

  • 12 kudos
8 More Replies
LiresaFerizaj
by • New Contributor III
  • 254 Views
  • 5 replies
  • 4 kudos

AUTO CDC SCD Type 2 with late-arriving deletes: edge cases the docs don't cover

Hi everyone,I'm designing a Lakeflow Declarative Pipeline that processes customer profile changes from a CDC feed into an SCD Type 2 Silver table. Events can arrive out of order, sometimes up to 24 hours late. The planned flow:CREATE OR REFRESH STREA...

  • 254 Views
  • 5 replies
  • 4 kudos
Latest Reply
Louis_Frolio
Databricks Employee
  • 4 kudos

Hello @LiresaFerizaj , I took a look at both internal and external documentation and here is what I found. You've already had solid answers from Thomaz, Nitesh, data_pulse, and Islam, so I'll tie the thread together, separate what's documented from w...

  • 4 kudos
4 More Replies
david888
by • New Contributor II
  • 184 Views
  • 2 replies
  • 1 kudos

Question about MySQL Integrated‑CDC pipeline on Classic compute workspace

Hi everyone,I am trying to set up an Integrated‑CDC pipeline for a MySQL RDS instance on our Databricks workspace. Our workspace only uses Classic compute; Serverless compute is not available.According to the documentation, Integrated‑CDC for MySQL s...

  • 184 Views
  • 2 replies
  • 1 kudos
Latest Reply
Louis_Frolio
Databricks Employee
  • 1 kudos

Hello @david888 , I took a look at both internal and external documentation and here is what I found. Ivan (@ivanvyd) has this right, so I'll build on his answer rather than repeat it.   Taking your three questions in order. Feature flag. Yes, enabl...

  • 1 kudos
1 More Replies
Kirankumarbs
by • Valued Contributor III
  • 3331 Views
  • 4 replies
  • 2 kudos

Python logger.info() not showing inside applyInPandas (but print() works) — why?

Problem: In Databricks, logs from an external binary (via os.system) show up, but Python logger.info() inside groupBy(...).applyInPandas(...) does not. print(..., flush=True) does show up.Why: applyInPandas runs your function as a pandas UDF.  That c...

  • 3331 Views
  • 4 replies
  • 2 kudos
Latest Reply
BenGriffiths
  • 2 kudos

Hi @SteveOstrowski,Is your approach supposed to work for Databricks serverless (Spark connect)? I have tried implemented that helper function to my applyInPandas() function, but the logs do not appear in the cell output.  

  • 2 kudos
3 More Replies
Davide
by • New Contributor II
  • 283 Views
  • 4 replies
  • 0 kudos

Resolved! Failed to convert to managed table: PrivilegedGenerateTemporaryTableCredential disable

Hello everyone, I noticed that when I tried to convert some external table to managed one using the alter table catalog.schema.table_name set managed command (like suggested by databricks) sometimes I will encounter this error: Error running query: c...

  • 283 Views
  • 4 replies
  • 0 kudos
Latest Reply
Davide
New Contributor II
  • 0 kudos

Hello everyone, After following up with Databricks, they recognize it is a bug on their side therefore are starting the process to fix it, in the mean time the best way to deal with it according to Databricks is simply to setup a retry and it should ...

  • 0 kudos
3 More Replies
Davide
by • New Contributor II
  • 137 Views
  • 2 replies
  • 1 kudos

Resolved! SET MANAGED AND ROLLBACK IN RELATION TO VACUUM

Hello everyone,I am looking for clarification about the 14-day rollback period after converting an external table to a managed table using: ALTER TABLE table_name SET MANAGED.Databricks documentation states that the conversion can be rolled back with...

  • 137 Views
  • 2 replies
  • 1 kudos
Latest Reply
DoTA
Valued Contributor II
  • 1 kudos

Hi, I could not find these three cases documented either, so I will separate what the docs say from what I would treat as an assumption. What is documented (Convert external tables to managed): Databricks keeps the data in the old external location f...

  • 1 kudos
1 More Replies
srikanthp24
by • New Contributor III
  • 102 Views
  • 1 replies
  • 0 kudos

Regarding the DAB Ingestion Pipeline Creating the Volume

Hello Guys,By taking the below yaml code as a reference I have created the databrick asset bundle for data transfer from sql server to databricks. But with below code the staging volume getting created by gateway pipeline automatically, which unable ...

srikanthp24_0-1790135779415.png
  • 102 Views
  • 1 replies
  • 0 kudos
Latest Reply
DoTA
Valued Contributor II
  • 0 kudos

Hi, a few things that may unblock you here. 1) Custom volume name. In gateway_definition you can set gateway_storage_catalog, gateway_storage_schema and gateway_storage_name. If you leave gateway_storage_name out, the gateway auto-generates a volume ...

  • 0 kudos
Khasim_1
by • New Contributor III
  • 123 Views
  • 1 replies
  • 1 kudos

Implementing a "Zero-Bus" Architecture with Unity Catalog

Hi everyone,I am currently refining the architecture for an end-to-end Lakehouse project and moving toward a "Zero-Bus" (Bus Architecture) model. My primary objective is to enforce dimension conformity (Customer, Product, Geography) across multiple i...

  • 123 Views
  • 1 replies
  • 1 kudos
Latest Reply
DoTA
Valued Contributor II
  • 1 kudos

Hi @Khasim_1, here is the pattern I would use for conformed dimensions on Unity Catalog, taking your questions in order. 1) Layering. I would avoid a "Master Silver" that everything funnels through, because it tends to become a bottleneck team. Treat...

  • 1 kudos
cdn_yyz_yul
by • Contributor III
  • 510 Views
  • 11 replies
  • 1 kudos

Change data feed from a materialized view

Hi everyone,Trying the CDF on materialized view (declarative pipeline) feature that is currently in Beta.- Configuration:** Verified the MVs (materialized veiws) have: 'delta.enableRowTracking'SHOW TBLPROPERTIES my_mv ('delta.enableRowTracking');  re...

  • 510 Views
  • 11 replies
  • 1 kudos
Latest Reply
cdn_yyz_yul
Contributor III
  • 1 kudos

The feature that will help the most is the CDF on MV that is in beta right now. I have been testing it to get ready at my side and take advantage of it once it is in General Availability. The Automatic change data feed on Streaming tables would be us...

  • 1 kudos
10 More Replies
Islam_hoti
by • New Contributor III
  • 145 Views
  • 1 replies
  • 2 kudos

Resolved! Streaming read fails with "Detected a data update" after a restatement job touches the source

Hi everyone,Looking for the right pattern here rather than a workaround.Setup. DBR 15.4 LTS, Unity Catalog. A silver streaming table reads from a bronze Delta table with a normal streaming read. Bronze is append only in the ordinary course of busines...

  • 145 Views
  • 1 replies
  • 2 kudos
Latest Reply
Khasim_1
New Contributor III
  • 2 kudos

Hi @Islam_hoti ,You have hit the classic 'Streaming vs. Mutation' wall. readStream on a Delta table assumes an append-only sequence; once an UPDATE happens, the commit versioning sequence is invalidated, hence the loud failure.Here is how we reasoned...

  • 2 kudos
Islam_hoti
by • New Contributor III
  • 196 Views
  • 3 replies
  • 5 kudos

Resolved! Photon enabled but a large share of the plan is falling back, cost up and runtime flat

Hi everyone,Trying to work out whether this is expected or whether I have misconfigured something.We enabled Photon on a job cluster running a nightly aggregation over roughly 2TB. The expectation was the usual improvement. What we got instead was ru...

  • 196 Views
  • 3 replies
  • 5 kudos
Latest Reply
Khasim_1
New Contributor III
  • 5 kudos

Hi @Islam_hoti ,Photon is fantastic, but it’s an 'all-or-nothing' value proposition. Once you hit a fallback, you lose the vectorized execution benefit for that entire sub-tree, and you're still paying the DBU premium. Here is how we approach auditin...

  • 5 kudos
2 More Replies
Islam_hoti
by • New Contributor III
  • 201 Views
  • 3 replies
  • 4 kudos

Resolved! At what data size do you stop reaching for Spark?

Hi everyone,A question I keep having with my team and I would like to hear how others think about it.A lot of the jobs we run are not big. Plenty of our pipelines process a few gigabytes, some considerably less. We run them on Spark because that is w...

  • 201 Views
  • 3 replies
  • 4 kudos
Latest Reply
Khasim_1
New Contributor III
  • 4 kudos

Hi @Islam_hoti ,This is the 'Silent Architect' dilemma—when the standard tool is technically suboptimal but organizationally necessary. We’ve wrestled with this, and here’s where we landed:The 'Standardization Premium' is real: We reasoned that the c...

  • 4 kudos
2 More Replies
Mado
by • Valued Contributor II
  • 294 Views
  • 6 replies
  • 7 kudos

How can I configure Lakeflow Connect SQL Server CDC Gateway to use a desired VM type?

Hi Team,I'm evaluating Databricks Lakeflow Connect for SQL Server CDC ingestion and have run into a gateway provisioning issue.EnvironmentRegion: Australia EastSource: Azure SQL DatabaseCDC enabled successfully on the database and source tableSQL Ser...

Mado_0-1790214214728.png
  • 294 Views
  • 6 replies
  • 7 kudos
Latest Reply
Khasim_1
New Contributor III
  • 7 kudos

Hi @Mado As of now, Lakeflow Connect's CDC Gateway VM types (driver/worker SKUs) are not configurable at the pipeline or connection level — they are managed internally by the platform. If the default SKU (Standard_E4d_v4) is unavailable in your regio...

  • 7 kudos
5 More Replies
Labels