Data Engineering

Forum Posts

Sorted by:

by Pratikmsbsvm • Contributor

Tuesday

87 Views
2 replies
1 kudos

How to Design a Data Quality Framework for Medallion Architecture Data Pipeline

Hello,I am building a Data Pipeline which extract data from Oracle Fusion and Push it to Databricks Delta lake.I am using Bronze, Silver and Gold Approach.May someone please help me how to control all three segment that is Bronze, Silver and Gold wit...

Data Engineering

87 Views
2 replies
1 kudos

Tuesday

View Replies

Latest Reply

nayan_wylde
Esteemed Contributor

Tuesday

1 kudos

Here’s how you can implement DQ at each stage:Bronze LayerChecks:File format validation (CSV, JSON, etc.).Schema validation (column names, types).Row count vs. source system.Tools:Use Databricks Autoloader with schema evolution and badRecordsPathImpl...

1 kudos

Tuesday

1 More Replies

by Shalabh007 • Honored Contributor

11-29-2022 10:52:08 AM

8912 Views
6 replies
19 kudos

Practice Exams for Databricks Certified Data Engineer Professional exam

Can anyone help with official Practice Exams set for Databricks Certified Data Engineer Professional exam, like we have below for Databricks Certified Data Engineer AssociatePractice exam for the Databricks Certified Data Engineer Associate exam

Data Engineering

8912 Views
6 replies
19 kudos

11-29-2022 10:52:08 AM

View Replies

Latest Reply

JOHNBOSCOW23
New Contributor

Tuesday

19 kudos

I Passed my Exam today thanks

19 kudos

Tuesday

5 More Replies

by Andolina1 • New Contributor III

05-15-2025 5:19:36 AM

2923 Views
6 replies
1 kudos

How to trigger an Azure Data Factory pipeline through API using parameters

Hello All,I have a use case where I want to trigger an Azure Data Factory pipeline through API. Right now I am calling the API in Databricks and using Service Principal(token based) to connect to ADF from Databricks.The ADF pipeline has some paramete...

Data Engineering

2923 Views
6 replies
1 kudos

05-15-2025 5:19:36 AM

View Replies

Latest Reply

rfranco
New Contributor

Tuesday

1 kudos

Hello @Andolina1,try to send your payload like:body = {'curr_working_user': f'{parameters}'}response = requests.post(url, headers=headers, json=body)the pipeline's parameter should be named curr_working_user. With these changes your setup should work...

1 kudos

Tuesday

5 More Replies

by BipinDatabricks • New Contributor

Monday

60 Views
3 replies
0 kudos

Using Databricks Sql Statement Execution api

TeamWe have internal chatbot service that will send query to data bricks SQL execution API.Number of queries vary from 50 per minute to 100 per minutes. and we are trying to limit response size by applying limit 10. Basically trying hard to use all o...

Data Engineering

60 Views
3 replies
0 kudos

Monday

View Replies

Latest Reply

Coffee77
Contributor III

Tuesday

0 kudos

As a suggestion, you can also think of creating your own API to query directly your tables via JDBC/ODBC connections over a SQL Warehouse. This case, limitations would be only those associated to SQL Warehouses and your API but not the Databricks API...

0 kudos

Tuesday

2 More Replies

by ShanQiwei • New Contributor

Monday

68 Views
2 replies
0 kudos

I/F security about using medallion architecture

I’m new to writing requirement definitions, and I’d like to ask a question about interface (I/F) security.My question is:Do I need to define the authentication and security mechanisms (such as OAuth2, Managed Identity, Service Principals, etc.) betwe...

Data Engineering

68 Views
2 replies
0 kudos

Monday

View Replies

Latest Reply

Coffee77
Contributor III

Tuesday

0 kudos

I'll try to summarize and go directly to the key points as I see this:- Client to S3 SAS Token or OAUTH 2.0 with Service to Service authentication (preferred)- Databricks to S3 Use Service Principal or Managed Identities (preferred)- Bronze/Silver/...

0 kudos

Tuesday

1 More Replies

by Techtic_kush • New Contributor

Friday

93 Views
2 replies
2 kudos

Can’t save results to target table – out-of-memory error

Hi team, I’m processing ~5,000 EMR notes with a Databricks notebook. The job reads from `crc_lakehouse.bronze.emr_notes`, runs SciSpaCy UMLS entity extraction plus a fine-tuned BERT sentiment model per partition, and builds a DataFrame (`df_entities`...

Data Engineering

93 Views
2 replies
2 kudos

Friday

View Replies

Latest Reply

bianca_unifeye
New Contributor III

Monday

2 kudos

You’re right that the behaviour is weird at first glance (“5k rows on a 64 GB cluster and I blow up on write”), but your stack trace is actually very revealing: this isn’t a classic Delta write / shuffle OOM – it’s SciSpaCy/UMLS falling over when loa...

2 kudos

Monday

1 More Replies

by mplang • New Contributor

10-14-2024 8:59:10 AM

4090 Views
3 replies
2 kudos

DLT x UC x Auto Loader

Now that the Directory Listing Mode of Auto Loader is officially deprecated, is there a solution for using File Notification Mode in a DLT pipeline writing to a UC-managed table? My understanding is that File Notification Mode is only available on si...

Data Engineering

autoloader

dlt

4090 Views
3 replies
2 kudos

10-14-2024 8:59:10 AM

View Replies

Latest Reply

Raman_Unifeye
Contributor III

Tuesday

2 kudos

Databricks introduced Managed File Events which completely bypasses the need for the cluster's identity to provision cloud resources, resolving the conflict with the Shared cluster mode.Steps to Implement in DLTEnable File Events on the External Loca...

2 kudos

Tuesday

2 More Replies

by Sainath368 • Contributor

Monday

81 Views
3 replies
2 kudos

Migrating from directory-listing to Autoloader Managed File events

Hi everyone,We are currently migrating from a directory listing-based streaming approach to managed file events in Databricks Auto Loader for processing our data in structured streaming.We have a function that handles structured streaming where we ar...

Data Engineering

81 Views
3 replies
2 kudos

Monday

View Replies

Latest Reply

Raman_Unifeye
Contributor III

Monday

2 kudos

Yes, for your setup, Databricks Auto Loader will create a separate event queue for each independent stream running with the cloudFiles.useManagedFileEvents = true option.As you are running - 1 stream per table, 1 unique directory per stream and 1 uni...

2 kudos

Monday

2 More Replies

by StephenDsouza • New Contributor II

05-23-2024 4:52:29 AM

3101 Views
3 replies
0 kudos

Error during build process for serving model caused by detectron2

Hi All,Introduction: I am trying to register my model on Databricks so that I can serve it as an endpoint. The packages that I need are "torch", "mlflow", "torchvision", "numpy" and "git+https://github.com/facebookresearch/detectron2.git". For this, ...

Data Engineering

3101 Views
3 replies
0 kudos

05-23-2024 4:52:29 AM

View Replies

Latest Reply

StephenDsouza
New Contributor II

05-24-2024 2:54:54 AM

0 kudos

Found an answer!Basically pip was somehow installed the dependencies from the git repo first and was not following the given order so in order to solve this, I added the libraries for conda to install.``` conda_env = { "channels": [ "defa...

0 kudos

05-24-2024 2:54:54 AM

2 More Replies

by shashankB • New Contributor III

Friday

145 Views
5 replies
0 kudos

Lakebridge analyzer not able to determine DDL.

Databricks analyzer does not shows any DDL statement count, I've also tested with just a simple SELECT * query (SELECT * FROM SCHEMA_NAME.TABLE_NAME;) . Is there any solution for this ?My target was to get a detailed analysis on SnowSQL code. Any h...

Data Engineering

145 Views
5 replies
0 kudos

Friday

View Replies

Latest Reply

Thompson2345
New Contributor II

Sunday

0 kudos

The Lakebridge analyzer counts DDL statements, not regular queries. A simple SELECT * is DML, not DDL, so it won’t show up in the DDL count.To get meaningful results for SnowSQL code analysis:Include actual DDL statements like CREATE TABLE, ALTER TAB...

0 kudos

Sunday

4 More Replies

by EDDatabricks • Contributor

09-12-2024 5:31:33 AM

4122 Views
1 replies
1 kudos

Schema Registry certificate auth with Unity Catalog volumes.

Greetings.We currently have a Spark structured streaming job (Scala) retrieving avro data from an Azure Eventhub with a confluent schema registry endpoint (using an Azure Api Management gateway with certificate authentication).Until now the .jks file...

Data Engineering

4122 Views
1 replies
1 kudos

09-12-2024 5:31:33 AM

View Replies

Latest Reply

stbjelcevic
Databricks Employee

Monday

1 kudos

Thanks for the detailed context—here’s a concise, actionable troubleshooting plan tailored to Databricks with Unity Catalog volumes and Avro + Confluent Schema Registry over APIM with mTLS. What’s likely going wrong Based on your description, the ini...

1 kudos

Monday

by Sega2 • New Contributor III

09-26-2024 1:50:41 AM

4680 Views
2 replies
0 kudos

Adding a message to azure service bus

I am trying to send a message to a service bus in azure. But I get following error:ServiceBusError: Handler failed: DefaultAzureCredential failed to retrieve a token from the included credentials.This is the line that fails: credential = DefaultAzure...

Data Engineering

4680 Views
2 replies
0 kudos

09-26-2024 1:50:41 AM

View Replies

Latest Reply

stbjelcevic
Databricks Employee

Monday

0 kudos

It looks like the issue is with the Azure credential chain rather than Service Bus itself; in Databricks notebooks, DefaultAzureCredential won’t succeed unless there’s a valid identity available (env vars, CLI login, managed identity, or a Databricks...

0 kudos

Monday

1 More Replies

by Miguel_Salas • New Contributor II

10-21-2024 12:05:48 PM

4956 Views
2 replies
0 kudos

How Install Pyrfc into AWS Databrick using Volumes

I'm trying to install Pyrfc in a Databricks Cluster (already tried in r5.xlarge, m5.xlarge, and c6gd.xlarge). I'm following these link.https://community.databricks.com/t5/data-engineering/how-can-i-cluster-install-a-c-python-library-pyrfc/td-p/8118Bu...

Data Engineering

4956 Views
2 replies
0 kudos

10-21-2024 12:05:48 PM

View Replies

Latest Reply

stbjelcevic
Databricks Employee

Monday

0 kudos

Thanks for the details. The PyRFC package is a Python binding around the SAP NetWeaver RFC SDK and requires the SAP NW RFC SDK to be present at build/run time; it does not work as a pure Python wheel on Linux without the SDK. The project is archived ...

0 kudos

Monday

1 More Replies

by HoussemBL • New Contributor III

05-16-2025 3:30:14 AM

2790 Views
2 replies
1 kudos

how to add Microsoft Entra ID managed service principal to aws databricks

Hi,I would like to add a Microsoft Entra ID managed service principal to AWS Databricks, but I have noticed that this option does not appear to be available-I am only able to create managed service principals directly within Databricks.For comparison...

Data Engineering

2790 Views
2 replies
1 kudos

05-16-2025 3:30:14 AM

View Replies

Latest Reply

stbjelcevic
Databricks Employee

Monday

1 kudos

You cannot add a Microsoft Entra ID–managed service principal to Databricks on AWS today; AWS workspaces only support Databricks‑managed service principals that you create in the Databricks account/workspace, not service principals federated from Ent...

1 kudos

Monday

1 More Replies

by nchittampelly • New Contributor II

05-05-2025 8:27:44 AM

3097 Views
3 replies
0 kudos

What is the best way to connect Oracle CRM cloud from databricks?

Data Engineering

3097 Views
3 replies
0 kudos

05-05-2025 8:27:44 AM

View Replies

Latest Reply

nchittampelly
New Contributor II

07-29-2025 8:45:50 AM

0 kudos

Oracle CRM on Demand is a Cloud platform not a relational database.Is there any proven solution for this requirement?

0 kudos

07-29-2025 8:45:50 AM

2 More Replies

Databricks Community

Forum Posts

How to Design a Data Quality Framework for Medallion Architecture Data Pipeline

Practice Exams for Databricks Certified Data Engineer Professional exam

How to trigger an Azure Data Factory pipeline through API using parameters

Using Databricks Sql Statement Execution api

I/F security about using medallion architecture

Can’t save results to target table – out-of-memory error

DLT x UC x Auto Loader

Migrating from directory-listing to Autoloader Managed File events

Error during build process for serving model caused by detectron2

Lakebridge analyzer not able to determine DDL.

Schema Registry certificate auth with Unity Catalog volumes.

Adding a message to azure service bus

How Install Pyrfc into AWS Databrick using Volumes

how to add Microsoft Entra ID managed service principal to aws databricks

What is the best way to connect Oracle CRM cloud from databricks?

Join Us as a Local Community Builder!

Moving tables between pipelines in production

Cannot view nested MLflow experiment runs without ...

Autoloader Managed File events

Unable to navigate/login to Databricks Account Con...

Got an empty query file when cloning Query file.