cancel
Showing results for 
Search instead for 
Did you mean: 
Data Engineering
Join discussions on data engineering best practices, architectures, and optimization strategies within the Databricks Community. Exchange insights and solutions with fellow data engineers.
cancel
Showing results for 
Search instead for 
Did you mean: 

Forum Posts

yit337
by • Contributor II
  • 1890 Views
  • 2 replies
  • 1 kudos

Resolved! Identity column has null values

I want to update a dimension table in the gold model from a silver table by using  create_auto_cdc_from_snapshot_flow and SCD2. In the target table, I have defined an IDENTITY column, which should be populated automatically.The dlt flow runs successf...

  • 1890 Views
  • 2 replies
  • 1 kudos
Latest Reply
SteveOstrowski
Databricks Employee
  • 1 kudos

Hi @yit337, The reason your identity column values are NULL is that the target table created by create_auto_cdc_from_snapshot_flow is a streaming table, and streaming tables do not support identity columns. This is a documented limitation: https://do...

  • 1 kudos
1 More Replies
Saikumar_Manne
by • New Contributor II
  • 4218 Views
  • 4 replies
  • 1 kudos

Resolved! How to use multi-threading and batch inserts for large UPSERT to PostgreSQL from Databricks?

Hi everyone,We have a Databricks (Unity Catalog) pipeline where we process large datasets in Spark and need to load incremental data into a PostgreSQL target table.Our scenario is:Initial full load (~300 million rows) to PostgreSQL using bulk COPY is...

  • 4218 Views
  • 4 replies
  • 1 kudos
Latest Reply
SteveOstrowski
Databricks Employee
  • 1 kudos

Hi @Saikumar_Manne, With 190M+ daily rows going into PostgreSQL via INSERT ON CONFLICT DO UPDATE, there are several levers to pull. Here is a breakdown of the approaches and tuning options. APPROACH 1: STAGING TABLE + MERGE (RECOMMENDED FOR THIS VOLU...

  • 1 kudos
3 More Replies
ChrisLawford_n1
by • Contributor II
  • 2236 Views
  • 3 replies
  • 1 kudos

Resolved! DeltaFileOperations: Listing improvement?

Hello, I am using databricks autoloader with managedfileevents turned on and include existing files.I want to understand if there is a way of increasing the speed of the initial listing of the files for autoloader.I thought that the idea behind the m...

  • 2236 Views
  • 3 replies
  • 1 kudos
Latest Reply
SteveOstrowski
Databricks Employee
  • 1 kudos

Hi @ChrisLawford_n1, You are correct that managed file events (cloudFiles.useManagedFileEvents = true) works by having Databricks maintain a record of file events on the external location, so when you start a new Auto Loader stream, it can replay tho...

  • 1 kudos
2 More Replies
arushigulati
by • Databricks Partner
  • 2096 Views
  • 2 replies
  • 0 kudos

Lakebridge transpile to translate from oracle to databricks sql

Hi Community,I am currently working on a PoC to migrate data from Oracle to Databricks. As part of this, we are attempting to automate the DDL conversion process.We are leveraging Databricks Labs Lakebridge for transpilation, but it is failing to con...

arushigulati_0-1769670207329.png arushigulati_1-1769670261737.png
  • 2096 Views
  • 2 replies
  • 0 kudos
Latest Reply
SteveOstrowski
Databricks Employee
  • 0 kudos

Hi @arushigulati, Lakebridge (the Databricks Labs project formerly known as Remorph) does support Oracle as a source dialect for transpilation, but the DDL handling, particularly around constraints like PRIMARY KEY, has some gaps depending on the ver...

  • 0 kudos
1 More Replies
echol
by • New Contributor II
  • 2784 Views
  • 6 replies
  • 1 kudos

Redeploy Databricks Asset Bundle created by others

Hi everyone,Our team is using Databricks Asset Bundles (DAB) with a customized template to develop data pipelines. We have a core team that maintains the shared infrastructure and templates, and multiple product teams that use this template to develo...

  • 2784 Views
  • 6 replies
  • 1 kudos
Latest Reply
SteveOstrowski
Databricks Employee
  • 1 kudos

Hi @echol, This is a common scenario when multiple team members work with Databricks Asset Bundles, and there are a few approaches to solve it cleanly. THE ROOT CAUSE When Staff A deploys a bundle, the jobs and other resources are created with Staff ...

  • 1 kudos
5 More Replies
YuriS
by • New Contributor III
  • 1798 Views
  • 6 replies
  • 3 kudos

Resolved! StreamingQueryListener metrics strange behaviour (inputRowsPerSecond metric is set to 0)

After implementing StreamingQueryListener to enable integration with our monitoring solution we have noticed some strange metrics for our DeltaSource streams (based on https://learn.microsoft.com/en-us/azure/databricks/structured-streaming/stream-mon...

YuriS_0-1769419735190.png YuriS_1-1769419836870.png
  • 1798 Views
  • 6 replies
  • 3 kudos
Latest Reply
SteveOstrowski
Databricks Employee
  • 3 kudos

Hi @YuriS, There are a few things going on here, and I will walk through each one. INPUTROWSPERSECOND SHOWING 0 The inputRowsPerSecond metric is not calculated from the current batch. It is the rate of data arriving between the end of the previous tr...

  • 3 kudos
5 More Replies
_its_akshaye
by • New Contributor II
  • 825 Views
  • 2 replies
  • 1 kudos

How to Track Hourly or Daily # of Upsert/Delete Metrics in a DLT Streaming Pipeline

We created a Delta Live Tables (DLT) streaming pipeline to ingest data from the Bronze layer to the Silver layer using Change Data Feed (CDF) enabled.The stream runs continuously and shows # of upserted and deleted rows at an aggregate level from the...

  • 825 Views
  • 2 replies
  • 1 kudos
Latest Reply
SteveOstrowski
Databricks Employee
  • 1 kudos

Hi @_its_akshaye, The pipeline event log captures exactly what you need. Every Lakeflow Spark Declarative Pipeline (SDP, formerly known as DLT) records flow_progress events that include per-flow metrics with num_upserted_rows and num_deleted_rows fie...

  • 1 kudos
1 More Replies
aranjan99
by • Contributor
  • 1548 Views
  • 6 replies
  • 1 kudos

System table missing primary keys?

This simple query takes 50seconds for me on a X-Small warehouse.select * from SYSTEM.access.workspaces_latest where workspace_id = '442224551661121'Can the team comment on why querying on system tables takes so long? I also dont see any primary keys ...

  • 1548 Views
  • 6 replies
  • 1 kudos
Latest Reply
SteveOstrowski
Databricks Employee
  • 1 kudos

Hi @aranjan99, There are two separate topics here, so let me address each one. WHY THE QUERY IS SLOW (~50 SECONDS ON X-SMALL) System tables are served via Delta Sharing from a Databricks-hosted storage account in the same region as your Unity Catalog...

  • 1 kudos
5 More Replies
dpc
by • Contributor III
  • 3165 Views
  • 7 replies
  • 2 kudos

Using AD groups for object ownership

Databricks has a general issue with object ownership in that only the creator can delete them.So, if I create a catalog, table, view, schema etc. I am the only person who can delete it.No good if it's a general table or view and some other developer ...

  • 3165 Views
  • 7 replies
  • 2 kudos
Latest Reply
SteveOstrowski
Databricks Employee
  • 2 kudos

Hi @dpc, Unity Catalog does support group ownership of objects, including groups synced from Azure AD (Entra ID) or other identity providers. This is fully available today and addresses each of the scenarios you described. TRANSFERRING OWNERSHIP TO A...

  • 2 kudos
6 More Replies
fintech_latency
by • New Contributor III
  • 6280 Views
  • 10 replies
  • 2 kudos

How to guarantee “always-warm” serverless compute for low-latency Jobs workloads?

We’re building a low-latency processing pipeline on Databricks and are running into serverless cold-start constraints.We ingest events (calls) continuously via a Spark Structured Streaming listener.For each event, we trigger a serverlesss compute tha...

  • 6280 Views
  • 10 replies
  • 2 kudos
Latest Reply
SteveOstrowski
Databricks Employee
  • 2 kudos

Hi @fintech_latency, This is a common architectural challenge when you need sub-second or near-real-time event processing. Let me walk through your three questions and then outline the recommended patterns. QUESTION 1: CAN YOU GUARANTEE FIXED WARM SE...

  • 2 kudos
9 More Replies
sandy_123
by • Databricks Partner
  • 2016 Views
  • 2 replies
  • 1 kudos

Getting 'Multiple failure in stage materialization' error in one of my Job with notebook task

Multiple failures in stage materialization. it tried using powerful Job cluster but it does not work out?? Any suggestions how should i fix it? FYI- my dataframe(uniq_rec_df) has around 30M rows -------Screenshot attached---  

sandy_123_0-1768844333336.png
  • 2016 Views
  • 2 replies
  • 1 kudos
Latest Reply
SteveOstrowski
Databricks Employee
  • 1 kudos

Hi @sandy_123, The "Multiple failures in stage materialization" error occurs when Spark executors fail to fetch shuffle data from other executors during a shuffle stage. With 30 million rows, your job is likely hitting resource limits during a shuffl...

  • 1 kudos
1 More Replies
ajay_wavicle
by • Databricks Partner
  • 1606 Views
  • 4 replies
  • 0 kudos

How to copy files of databricks associated storage account UC tables along with _delta_log folder

I want to migrate managed tables from one cloud Databricks workspace to another as it is with delta history. I am able to do with External tables since i have access to storage account container folder but its not the case for UC managed tables. How ...

  • 1606 Views
  • 4 replies
  • 0 kudos
Latest Reply
SteveOstrowski
Databricks Employee
  • 0 kudos

Hi @ajay_wavicle, Migrating UC managed tables between workspaces while preserving delta history requires a different approach than external tables, because Unity Catalog controls the underlying storage for managed tables and you do not have direct ac...

  • 0 kudos
3 More Replies
Dhruv-22
by • Contributor III
  • 3106 Views
  • 5 replies
  • 0 kudos

Feature request: Allow to set value as null when not present in schema evolution

I want to raise a feature request as follows.Currently, in the Automatic schema evolution for merge when a column is not present in the source dataset it is not changed in the target dataset. For e.g.%sql CREATE OR REPLACE TABLE edw_nprd_aen.bronze.t...

Dhruv22_0-1767970990008.png Dhruv22_1-1767971051176.png Dhruv22_2-1767971116934.png Dhruv22_3-1767971213212.png
  • 3106 Views
  • 5 replies
  • 0 kudos
Latest Reply
SteveOstrowski
Databricks Employee
  • 0 kudos

Hi @Dhruv-22, This is a valid use case and your workaround is solid. Let me share some context on the current behavior and a few approaches that may help streamline things. CURRENT BEHAVIOR WITH SCHEMA EVOLUTION When automatic schema evolution is ena...

  • 0 kudos
4 More Replies
SatabrataMuduli
by • New Contributor II
  • 2239 Views
  • 2 replies
  • 1 kudos

Unable to Connect to Oracle from Databricks UC Cluster (DBR 15.4) – ORA-12170 Timeout Error

 Hi all,I’m trying to connect to an Oracle database from my Databricks UC cluster (DBR 15.4) using the ojdbc8.jar driver, which I’ve installed on the cluster. Here’s the code I’m using:df = spark.read.format("jdbc")\ .option("url", jdbc_url)\ ...

  • 2239 Views
  • 2 replies
  • 1 kudos
Latest Reply
SteveOstrowski
Databricks Employee
  • 1 kudos

Hi @SatabrataMuduli, The ORA-12170 TCP connect timeout error is a networking issue, not a driver or credentials problem. Your Databricks cluster cannot reach the Oracle host and port within the 15-second TCP timeout window. The fact that it works fro...

  • 1 kudos
1 More Replies
Digvijay_11
by • Databricks Partner
  • 3035 Views
  • 3 replies
  • 3 kudos

Lakeflow Spark Declarative Pipeline

How we can run a SDP pipeline in parallel manner with dynamic parameter parsing on pipeline level. How we can consume job level parameter in Pipeline. If similar name parameters are defined in pipeline level then job level parameters are getting over...

Data Engineering
Spark Declarative Pipelines
  • 3035 Views
  • 3 replies
  • 3 kudos
Latest Reply
SteveOstrowski
Databricks Employee
  • 3 kudos

Hi @Digvijay_11, Here are answers to each of your three questions about Lakeflow Spark Declarative Pipelines (SDP): 1. RUNNING SDP PIPELINES IN PARALLEL WITH DYNAMIC PARAMETERS SDP automatically determines the dependency graph across your table and v...

  • 3 kudos
2 More Replies
Labels