cancel
Showing results for 
Search instead for 
Did you mean: 
Data Engineering
Join discussions on data engineering best practices, architectures, and optimization strategies within the Databricks Community. Exchange insights and solutions with fellow data engineers.
cancel
Showing results for 
Search instead for 
Did you mean: 

Forum Posts

priya9896
by • New Contributor II
  • 369 Views
  • 6 replies
  • 0 kudos

Azure to AWS

Hi everyone,We're evaluating an architecture pattern and would appreciate any guidance or recommendations. Has anyone implemented a similar cross-cloud pattern? Specifically:Can AWS Glue be used to access Azure Databricks data without copying it?Are ...

  • 369 Views
  • 6 replies
  • 0 kudos
Latest Reply
priya9896
New Contributor II
  • 0 kudos

Thank you so much for all your help and guidance. I truly appreciate your responses and insights. I’ll definitely keep this in mind and follow up with the source team first.

  • 0 kudos
5 More Replies
DB1To3
by • Contributor II
  • 79 Views
  • 4 replies
  • 2 kudos

Zerobus Ingest with Kafka compliant API and WITHOUT a callback

Is see that kafka for zerobus is still in beta.  So maybe something will change. Can someone tell me whether they are aware of an authentication handshake with this API that might AVOID the callback?Here is the python example, which uses a sophistica...

  • 79 Views
  • 4 replies
  • 2 kudos
Latest Reply
Louis_Frolio
Databricks Employee
  • 2 kudos

Hello @DB1To3, I took a look at both internal and external documentation and here is what I found. Straight answer: there's no non-callback auth mode and no unauthenticated mode for the Zerobus Kafka API today. @DoTA has that right. The endpoint only...

  • 2 kudos
3 More Replies
surajitDE
by • Contributor
  • 93 Views
  • 3 replies
  • 1 kudos

Change in DLT Pipeline Deletion Behavior associated Tables Not Being Deleted

Hi Team,We have noticed a change in the behavior of our Lakeflow pipelines.Previously, when we deleted a DLT pipeline, the associated resources, such as materialized views and streaming tables, were also deleted along with the pipeline.However, recen...

  • 93 Views
  • 3 replies
  • 1 kudos
Latest Reply
anuj_lathi
Databricks Employee
  • 1 kudos

Yes, this is expected behavior for Unity Catalog pipelines using the default publishing mode. Deleting the pipeline now retains the materialized views, streaming tables, and views by default, and you can opt back into cascade deletion through the AP...

  • 1 kudos
2 More Replies
pvrcloudtech
by • New Contributor
  • 207 Views
  • 4 replies
  • 0 kudos

Databricks SDP (Spark declarative Pipelins) Overwrite table

1) Consider I have orders folder which orders_1.csv file  and it is loaded to orders tables orders/                 ============> load to order table      orders_1.csv  2) next day new file (orders_2.csv) arrived orders/                       =======...

  • 207 Views
  • 4 replies
  • 0 kudos
Latest Reply
anuj_lathi
Databricks Employee
  • 0 kudos

Streaming tables are append-only by design, and a materialized view recomputes over everything in the folder. Neither gives you "replace the table with just the newest file" out of the box, so you need a small pattern on top. Here are the two I'd co...

  • 0 kudos
3 More Replies
kartikchoudhary
by • Contributor
  • 152 Views
  • 7 replies
  • 0 kudos

What makes a Databricks data platform truly AI-ready?

As more teams start connecting AI agents and AI/BI workloads to Databricks, I’ve been thinking about what actually makes an enterprise data platform “AI-ready.”From my experience working in enterprise data and analytics at Polestar Analytics, the cha...

  • 152 Views
  • 7 replies
  • 0 kudos
Latest Reply
vs4
New Contributor II
  • 0 kudos

I see AI-readiness as more than making data accessible to an AI/agent. The foundation is trusted data + governed access + business context.In Databricks, I would build this around Unity Catalog for governance, lineage and discoverability, strong data...

  • 0 kudos
6 More Replies
pawanswami
by • New Contributor
  • 238 Views
  • 7 replies
  • 4 kudos

Guidance Required: Scheduling a Biweekly Databricks Job

I’m reaching out to seek your guidance on one of our use cases where we need to schedule a biweekly Databricks job to run every alternate Wednesday.I explored both the standard Databricks scheduling options and cron-based scheduling, but I was unable...

  • 238 Views
  • 7 replies
  • 4 kudos
Latest Reply
Louis_Frolio
Databricks Employee
  • 4 kudos

Greetings @pawanswami, I did some digging and here is what I found. Satyasai and ivy125cordes have the root cause right: Databricks schedules run on Quartz cron, and Quartz has no "every N weeks" concept. The question that decides your fix is whether...

  • 4 kudos
6 More Replies
kohei-matsumura
by • Databricks Partner
  • 103 Views
  • 4 replies
  • 2 kudos

Best practices for DR & Checkpoint Management of Serverless Declarative Pipelines (ST/MV)?

I am currently designing a Disaster Recovery (DR) strategy for our Databricks workspace to recover from system failures and human errors. Our main requirement is to restore both the data and the pipelines reliably if a workspace goes down or gets cor...

  • 103 Views
  • 4 replies
  • 2 kudos
Latest Reply
anuj_lathi
Databricks Employee
  • 2 kudos

Short answer: don't try to back up or migrate ST/MV state and checkpoints at all. On Serverless you can't reach them, and they aren't designed to be portable. Design for rebuild and replay instead: pipeline definitions in Git, replayable sources rep...

  • 2 kudos
3 More Replies
david888
by • New Contributor II
  • 275 Views
  • 3 replies
  • 1 kudos

Question about MySQL Integrated‑CDC pipeline on Classic compute workspace

Hi everyone,I am trying to set up an Integrated‑CDC pipeline for a MySQL RDS instance on our Databricks workspace. Our workspace only uses Classic compute; Serverless compute is not available.According to the documentation, Integrated‑CDC for MySQL s...

  • 275 Views
  • 3 replies
  • 1 kudos
Latest Reply
tracy74crawford
  • 1 kudos

Yes, workspace-level preview access or specific feature flags may be required for MySQL Integrated-CDC depending on your Databricks account tier, and while API-only deployments via the Databricks SDK are fully supported for production, you must ensur...

  • 1 kudos
2 More Replies
mlrichmond-mill
by • New Contributor III
  • 4420 Views
  • 11 replies
  • 0 kudos

Resolved! Bundled Wheel Task with Serverless Compute

I am trying to run a wheel task as part of a bundle on serverless compute. My databricks.yml includes an artifact being constructed:artifacts: nexusbricks: type: whl build: python -m build path: .I then am trying to set up a job to cons...

  • 4420 Views
  • 11 replies
  • 0 kudos
Latest Reply
AnhPT
New Contributor III
  • 0 kudos

i think we should install lib when init cluster.

  • 0 kudos
10 More Replies
CoopCoop
by • New Contributor III
  • 16366 Views
  • 17 replies
  • 7 kudos

Resolved! PDF Attachment on an Alert

Currently my Alert is an HTML table using data pointing to an SQL query.I was wondering if it is possible to attach the resulting table from this SQL query as a PDF to the alert email.If anyone has successfully implemented this, please let me know! T...

  • 16366 Views
  • 17 replies
  • 7 kudos
Latest Reply
Atanu
Databricks Employee
  • 7 kudos

Ok understood the concern, so basically the issue is with PDF rendering as much I understood. Let me know if I am wrong. Let me see if there is any improvement by our engineering team on this front.

  • 7 kudos
16 More Replies
dbernstein_tp
by • Contributor
  • 392 Views
  • 8 replies
  • 1 kudos

Lakeflow connect SQL server ingestion can be made elastic?

Hi Everyone, One of our big ingestion tasks is lakeflow connect CDC ingestion of ERP data from a SQL server database. I am deploying the pipelines and jobs for this via DABs. We are ingesting about 80 tables from the database, a handful of which are ...

  • 392 Views
  • 8 replies
  • 1 kudos
Latest Reply
Louis_Frolio
Databricks Employee
  • 1 kudos

Hello @dbernstein_tp , I took a look at the comments here and did more digging based on what I've read. Good questions from @lmcorreahdb, and the 21/60 made me think. Your current design already fits your SLAs. The gateway runs continuously on small ...

  • 1 kudos
7 More Replies
suryaprayaga
by • Contributor
  • 257 Views
  • 3 replies
  • 2 kudos

Genie Use Cases & easy adaptibility

Recently I started socializing the use of Genie and Databricks to even a laymen (and ofcourse laywomen) who have never heard what coding is in their life. I am proud that in my company nearly 200 people as of now are using Databricks for some or othe...

  • 257 Views
  • 3 replies
  • 2 kudos
Latest Reply
Aravind_Reddy
Databricks Partner
  • 2 kudos

its really good to see that genie is being used beyond data engineering teams. One thing I think becomes particularly important as adoption grows is separating experimentation from production workloads and the genie makes it much easier for non-techn...

  • 2 kudos
2 More Replies
Innuendo84
by • Databricks Partner
  • 476 Views
  • 9 replies
  • 2 kudos

Problems with jobs / GIT repo

I'm having problems with using jobs together with GIT.I have created a job that hasType - NotebookSource - Git provider (main),Path "jobs/base/INIT_PROCESSING"I can access the jupyter notebook without any problems. Each push to the repo can be read i...

  • 476 Views
  • 9 replies
  • 2 kudos
Latest Reply
Aravind_Reddy
Databricks Partner
  • 2 kudos

I believe this could be related to the sparse_checkout configuration. If the notebook is importing modules from directories outside the `jobs/base` path, those files may not be available in the job’s Git checkout.As a first step, I would suggest eith...

  • 2 kudos
8 More Replies
idk-arsh
by • New Contributor
  • 145 Views
  • 2 replies
  • 0 kudos

I built an open-source check that stops AI agents from inventing column names in SQL.

I work on data pipelines. The most common way I see Claude Code and Cursor break SQL is dull. They write against the schema they think exists. The README is stale, or they copy `customer_id` from an old query when the real column is `customerid`. The...

  • 145 Views
  • 2 replies
  • 0 kudos
Latest Reply
Sanjeeb2024
Valued Contributor II
  • 0 kudos

Thanks for sharing this. In high level, it is a tool for the agent which will validate the schema and take an action. One question how do we maintain the freshness of the schema, for an example if there is any change in the column name or data type, ...

  • 0 kudos
1 More Replies
Labels