cancel
Showing results for 
Search instead for 
Did you mean: 
Data Engineering
Join discussions on data engineering best practices, architectures, and optimization strategies within the Databricks Community. Exchange insights and solutions with fellow data engineers.
cancel
Showing results for 
Search instead for 
Did you mean: 

Forum Posts

VivekRawat
by • New Contributor
  • 287 Views
  • 1 replies
  • 0 kudos

Declarative Until It Isn't: Four Sharp Edges of Lakeflow Declarative Pipelines

The common case is almost too easy — which is exactly why the real engineering lives at the boundaries where the abstraction leaks.Declarative pipelines make the common case feel almost suspiciously easy. You describe the tables you want, point them ...

  • 287 Views
  • 1 replies
  • 0 kudos
Latest Reply
DoTA
Valued Contributor II
  • 0 kudos

Nice write-up, the read_files plus mergeSchema point is one I'd have missed. A few things from the current docs that may help with edges 2 and 4, since some of it has changed recently.On expectations: the incremental refresh page now says materialize...

  • 0 kudos
LingeshK
by • Databricks Employee
  • 46 Views
  • 0 replies
  • 0 kudos

How to Unit Test Spark Declarative Pipelines

Every engineer who ships application code writes unit tests for it. Yet the data pipeline that feeds the same application is often left with a manual run and a dashboard that someone eventually notices is wrong. The gap has been more about tooling th...

00-etl-diagram-excalidraw.png 03-v2-pipeline-dag-closeup.png 13-v2-gold-table-wrong-numbers.png 12-v2-run-pipeline-complete-green.png
  • 46 Views
  • 0 replies
  • 0 kudos
Innuendo84
by • Databricks Partner
  • 565 Views
  • 10 replies
  • 2 kudos

Problems with jobs / GIT repo

I'm having problems with using jobs together with GIT.I have created a job that hasType - NotebookSource - Git provider (main),Path "jobs/base/INIT_PROCESSING"I can access the jupyter notebook without any problems. Each push to the repo can be read i...

  • 565 Views
  • 10 replies
  • 2 kudos
Latest Reply
anuj_lathi
Databricks Employee
  • 2 kudos

Most likely, your .py files are being treated as notebooks rather than plain Python files, so they can't be imported and can't be read as a python_file. The job is not failing to pull the folder. Your own findings show the whole commit is checked ou...

  • 2 kudos
9 More Replies
Mado
by • Valued Contributor II
  • 577 Views
  • 8 replies
  • 7 kudos

How can I configure Lakeflow Connect SQL Server CDC Gateway to use a desired VM type?

Hi Team,I'm evaluating Databricks Lakeflow Connect for SQL Server CDC ingestion and have run into a gateway provisioning issue.EnvironmentRegion: Australia EastSource: Azure SQL DatabaseCDC enabled successfully on the database and source tableSQL Ser...

Mado_0-1790214214728.png
  • 577 Views
  • 8 replies
  • 7 kudos
Latest Reply
anuj_lathi
Databricks Employee
  • 7 kudos

Yes, you can override the node types, but not from the ingestion wizard. The gateway is a pipeline underneath, so you set the VM sizes in the gateway pipeline's compute settings (JSON, API, CLI, or Asset Bundles) and rerun it. Answers to your three ...

  • 7 kudos
7 More Replies
dbernstein_tp
by • Contributor
  • 472 Views
  • 11 replies
  • 2 kudos

Lakeflow connect SQL server ingestion can be made elastic?

Hi Everyone, One of our big ingestion tasks is lakeflow connect CDC ingestion of ERP data from a SQL server database. I am deploying the pipelines and jobs for this via DABs. We are ingesting about 80 tables from the database, a handful of which are ...

  • 472 Views
  • 11 replies
  • 2 kudos
Latest Reply
dbernstein_tp
Contributor
  • 2 kudos

Thanks everyone for this very useful discussion. What has happened in the past is that the full refresh is "triggered" because I had to destroy and refactor the bundle code to meet new scheduling requirements or other stakeholder requests. The way th...

  • 2 kudos
10 More Replies
ChristianRRL
by • Honored Contributor II
  • 268 Views
  • 3 replies
  • 2 kudos

Databricks Workflow Task Failure - Custom Error Messages

Hi there, kind of silly question but I'm hoping there's a simple way to address this.I'm trying to make sure that when certain job failure conditions arise, that a workflow task step fails. I've tried this with both raise RuntimeError and sys.exit, b...

ChristianRRL_0-1790348230722.png ChristianRRL_0-1790348341444.png
  • 268 Views
  • 3 replies
  • 2 kudos
Latest Reply
aayush_410
New Contributor II
  • 2 kudos

It's a platform limitation. Notebook tasks run through IPython, which always prints a traceback for an unhandled exception, so you can't fully suppress the extra output. sys.exit doesn't help either. In a notebook it raises SystemExit, which gets ren...

  • 2 kudos
2 More Replies
marcell_nagy
by • New Contributor III
  • 57 Views
  • 1 replies
  • 0 kudos

HubSpot connector

The HubSpot Support Lakeflow connector is able to download basic datasets. However, I see no way to get non-basic columns:1. The suggested method to analyse email events is to use the portalSubscriptionStatus and source properties. The Lakeflow conne...

  • 57 Views
  • 1 replies
  • 0 kudos
Latest Reply
Louis_Frolio
Databricks Employee
  • 0 kudos

Greetings @marcell_nagy, I did some digging and here is what I found. I think there are two separate issues here. Missing columns (portalSubscriptionStatus, source, and custom fields) In HubSpot's Email Events API, portalSubscriptionStatus and source...

  • 0 kudos
DB1To3
by • Contributor II
  • 154 Views
  • 6 replies
  • 2 kudos

Zerobus Ingest with Kafka compliant API and WITHOUT a callback

Is see that kafka for zerobus is still in beta.  So maybe something will change. Can someone tell me whether they are aware of an authentication handshake with this API that might AVOID the callback?Here is the python example, which uses a sophistica...

  • 154 Views
  • 6 replies
  • 2 kudos
Latest Reply
Louis_Frolio
Databricks Employee
  • 2 kudos

Hello @DB1To3,  Glad we beat Reddit to it. You read the VNet point right. There's no anonymous mode, and the Kafka endpoint is public only (no front-end Private Link yet), so "anyone who can reach it" would mean the whole internet writing into your U...

  • 2 kudos
5 More Replies
monson_mark
by • New Contributor
  • 228 Views
  • 2 replies
  • 1 kudos

Auto-TTL on pipeline streaming table not firing

We have Auto-TTL configured on a Lakeflow pipeline streaming table via create_streaming_table(auto_ttl={"timestamp_column": "_discontinued_on", "expire_in_days": 77}). The autottl.timestampColumn and autottl.expireInDays properties are correctly set ...

  • 228 Views
  • 2 replies
  • 1 kudos
Latest Reply
anuj_lathi
Databricks Employee
  • 1 kudos

Short version: your diagnosis holds up. As far as I know, there is no pipeline setting, table property, or flag that flips a pipeline-owned streaming table to PO-managed maintenance. I wouldn't expect Auto-TTL to run on a table that is stamped DISAB...

  • 1 kudos
1 More Replies
diliprreddy
by • New Contributor
  • 274 Views
  • 2 replies
  • 0 kudos

constraints are not creating on materialized views

Hi,We are loading data from the Bronze layer to the Silver layer and creating materialized views with primary key and foreign key constraints. We have a total of 33 dependent tables, and the load is being performed using a Spark Declarative Pipeline ...

  • 274 Views
  • 2 replies
  • 0 kudos
Latest Reply
anuj_lathi
Databricks Employee
  • 0 kudos

That message is a generic "internal failure" surface, not a specific cause. The real reason is almost always in the pipeline event log or the update details, and if it isn't, it's something Databricks support needs to look at using the update ID. 1....

  • 0 kudos
1 More Replies
Navinkumar_K
by • New Contributor
  • 62 Views
  • 1 replies
  • 0 kudos

Serverless pipeline with Materialized Views vs. Delta overwrite tables for ~60 Gold outputs

Hi everyone,I'd appreciate some advice from anyone who has worked on a similar setup.Current setup:- One Databricks Job with 2 serverless Lakeflow pipelines and around 60 Gold outputs- Pipeline 1 reads Curated tables and builds Materialized Views- Pi...

  • 62 Views
  • 1 replies
  • 0 kudos
Latest Reply
anuj_lathi
Databricks Employee
  • 0 kudos

Short answer: if your MVs are already doing a full recompute, swapping them for overwrite Delta tables won't give you a real win, and it's likely to cost you more. I'd keep the pipelines, find out why the MVs aren't going incremental, and benchmark ...

  • 0 kudos
Wola
by • New Contributor II
  • 402 Views
  • 4 replies
  • 1 kudos

Data ingestion: Setting up a connector for Postgres with CDC enabled.

Hello,I'm trying to ingest data from my RDS instance and set up change data capture on Databricks. Everything on the Postgres side has been done so that replication is possible, but CDC is still greyed out. I have asked Claude, Gemini, and ChatGPT. I...

Wola_0-1789753526825.png
  • 402 Views
  • 4 replies
  • 1 kudos
Latest Reply
Louis_Frolio
Databricks Employee
  • 1 kudos

Greetings @Wola, I did some digging and here is what I found. @ThomazNeto already pointed you at the two things that matter most: the PostgreSQL connector for Lakeflow Connect is in Public Preview and needs workspace enrollment, and logical replicati...

  • 1 kudos
3 More Replies
dsay96
by • New Contributor II
  • 506 Views
  • 6 replies
  • 4 kudos

using remote_query with SQL Serverless Warehouse

Hello, Im trying to set up remote_query() as an option for our developers to use when querying our DB2 server. Unfortunately the connection keeps timing out. The workspace is in a VPC in AWS and the interactive clusters can successfully use the JDBC ...

  • 506 Views
  • 6 replies
  • 4 kudos
Latest Reply
Louis_Frolio
Databricks Employee
  • 4 kudos

Hello @dsay96, I took a look at both internal and external documentation and here is what I found. TLS probably isn't your blocker. @aayush_410 called it: a bad certificate throws a handshake or PKIX error, not a timeout. A timeout with every compone...

  • 4 kudos
5 More Replies
kohei-matsumura
by • Databricks Partner
  • 244 Views
  • 9 replies
  • 16 kudos

Best practices for DR & Checkpoint Management of Serverless Declarative Pipelines (ST/MV)?

I am currently designing a Disaster Recovery (DR) strategy for our Databricks workspace to recover from system failures and human errors. Our main requirement is to restore both the data and the pipelines reliably if a workspace goes down or gets cor...

  • 244 Views
  • 9 replies
  • 16 kudos
Latest Reply
Aravind_Reddy
Databricks Partner
  • 16 kudos

Thanks for the detailed explanation everyone. The Active-Passive + replay/rebuild approach makes sense given that Serverless manages the checkpoint state internally.One aspect I’m curious about is how this pattern should be designed when the source d...

  • 16 kudos
8 More Replies
priya9896
by • New Contributor II
  • 446 Views
  • 9 replies
  • 2 kudos

Azure to AWS

Hi everyone,We're evaluating an architecture pattern and would appreciate any guidance or recommendations. Has anyone implemented a similar cross-cloud pattern? Specifically:Can AWS Glue be used to access Azure Databricks data without copying it?Are ...

  • 446 Views
  • 9 replies
  • 2 kudos
Latest Reply
Louis_Frolio
Databricks Employee
  • 2 kudos

Also, if you feel your question has been answered please "Accept as Solution" so that others can benefit.  Cheers, Lou.

  • 2 kudos
8 More Replies
Labels