cancel
Showing results for 
Search instead for 
Did you mean: 
Data Engineering
Join discussions on data engineering best practices, architectures, and optimization strategies within the Databricks Community. Exchange insights and solutions with fellow data engineers.
cancel
Showing results for 
Search instead for 
Did you mean: 

Forum Posts

Khasim_1
by • New Contributor III
  • 243 Views
  • 4 replies
  • 3 kudos

Operationalizing Lakeflow Connect:Handling Upstream Schema Evolution & Historical Backfill

Hi everyone,I’m currently evaluating Lakeflow Connect for our ingestion layer. While the setup for replicating source databases into the Lakehouse is remarkably streamlined, I’m looking for operational best practices from those of you using it in pro...

  • 243 Views
  • 4 replies
  • 3 kudos
Latest Reply
stephen4
New Contributor III
  • 3 kudos

The point about keeping Silver on an explicit schema contract is really useful, especially during full refreshes. I’d also keep an eye on CDC retention during large backfills so nothing gets missed. 

  • 3 kudos
3 More Replies
data_pulse
by • New Contributor III
  • 143 Views
  • 1 replies
  • 1 kudos

Anyone migrated from legacy CDF to Auto CDF yet?

Databricks mentions in the DBR 19 release notes that Auto CDF removes write-time overhead and can make MERGE and UPDATE operations about 15% faster on tables queried for changesFor anyone who has moved an existing pipeline from legacy CDF to Auto CDF...

  • 143 Views
  • 1 replies
  • 1 kudos
Latest Reply
ThomazNeto
Databricks Partner
  • 1 kudos

Hi,No production before/after numbers from me yet, we only have this on a couple of dev tables since GA on Sep 1, so I'll stick to what the docs say and what I'd measure. Someone with a month of prod data will hopefully chime in.The main thing to und...

  • 1 kudos
yit337
by • Contributor II
  • 423 Views
  • 8 replies
  • 7 kudos

Resolved! Is Advanced edition required for Serverless pipeline?

I want to set Serverless pipeline with edition=PRO, but it raises the following error:  cannot update pipeline: You must use the Advanced edition when using serverless compute.This is not specified anywhere in the documentation. It only says (see scr...

yit337_0-1790244064316.png
  • 423 Views
  • 8 replies
  • 7 kudos
Latest Reply
Khasim_1
New Contributor III
  • 7 kudos

Serverless DLT/Lakeflow pipelines currently mandate the "Advanced" edition.Even if you are not using Data Quality expectations (the most famous "Advanced" feature), the Serverless orchestration itself relies on advanced infrastructure management, enh...

  • 7 kudos
7 More Replies
AdamLucas
by • New Contributor II
  • 200 Views
  • 3 replies
  • 3 kudos

Which Topics Deserve More Attention Before the Databricks Data Engineer Associate Certification Exam

As soon as I began my preparation for the Databricks Data Engineer Associate Certification exam, I came to the realisation that merely going through the syllabus won’t be sufficient. There are certain topics which are recurrent in practical situation...

  • 200 Views
  • 3 replies
  • 3 kudos
Latest Reply
josepe
New Contributor
  • 3 kudos

Hey @AdamLucas I agree that just reading the syllabus isn’t enough for the Databricks Data Engineer Associate exam. I focused more on Delta Lake, Spark SQL, data transformation, and workflows because these topics needed a better understanding of how ...

  • 3 kudos
2 More Replies
lessalucas
by • New Contributor II
  • 123 Views
  • 1 replies
  • 1 kudos

New Databricks SQL window fuction: counter_diff

This function was designed to work with cumulative counters in time series. It automatically calculates the difference between the current and previous value, turning a cumulative counter into a per-period variation.Let's look at a practical example....

  • 123 Views
  • 1 replies
  • 1 kudos
Latest Reply
Khasim_1
New Contributor III
  • 1 kudos

The counter_diff function is a significant upgrade over the traditional LAG() approach for IoT and telemetry data because it is "reset-aware."While a standard subtraction (Current - LAG(Previous)) produces a large negative number when a hardware coun...

  • 1 kudos
suryaprayaga
by • Contributor
  • 119 Views
  • 0 replies
  • 0 kudos

Genie Use Cases & easy adaptibility

Recently I started socializing the use of Genie and Databricks to even a laymen (and ofcourse laywomen) who have never heard what coding is in their life. I am proud that in my company nearly 200 people as of now are using Databricks for some or othe...

  • 119 Views
  • 0 replies
  • 0 kudos
nafikazi
by • Contributor
  • 338 Views
  • 6 replies
  • 3 kudos

Resolved! Power BI Dataflow Gen1 → Databricks Pipeline: Is There a Migration Accelerator or Converter?

Hello Great Analytics professionals:Many organizations have legacy Power BI Dataflow Gen1 solutions and are now considering moving their ETL workloads to Databricks.Is there an existing Databricks migration accelerator, converter, or framework that c...

  • 338 Views
  • 6 replies
  • 3 kudos
Latest Reply
data_pulse
New Contributor III
  • 3 kudos

@nafikazi There isn't a one-click Power BI Dataflow Gen1 / Power Query M to Lakeflow conversion path yet available.For a POC, decompose each flow and use Genie Code to accelerate the translation into Databricks SQL/ PySpark.Provide the M query, sourc...

  • 3 kudos
5 More Replies
mmsanteago
by • New Contributor
  • 132 Views
  • 1 replies
  • 1 kudos

Databricks AI/BI Get Session User

Hello everyone,I am currently working with Databricks AI/BI Dashboards and I have a specific requirement regarding user context.I need to dynamically capture the "session user" (the person who is currently viewing the dashboard) rather than the "exec...

  • 132 Views
  • 1 replies
  • 1 kudos
Latest Reply
ivanvyd
New Contributor III
  • 1 kudos

@mmsanteago with Share data permissions, queries run using the publisher's permissions, so viewer-based RLS does not apply. Honestly, I'm not aware of a supported dashboard parameter or SQL function that exposes the actual viewer while keeping that e...

  • 1 kudos
farzad_h
by • New Contributor II
  • 269 Views
  • 3 replies
  • 0 kudos

Resolved! Am I being throttled?

Hi,Im using the free edition as a student and trying my best to learn databricks. I have participated in the learning festival and trying to prepare bith for my exam and get employed. Im running out of time in order to use your codeI noticed you guys...

  • 269 Views
  • 3 replies
  • 0 kudos
Latest Reply
ThiamLee
Contributor
  • 0 kudos

I completely understand the frustration. Free/student access is a crucial part of the learning journey, especially for those preparing for certifications and trying to build practical skills. Hopefully Databricks can look into this and ensure student...

  • 0 kudos
2 More Replies
Davide
by • New Contributor II
  • 39 Views
  • 0 replies
  • 0 kudos

SET MANAGED AND ROLLBACK IN RELATION TO VACUUM

Hello everyone,I am looking for clarification about the 14-day rollback period after converting an external table to a managed table using: ALTER TABLE table_name SET MANAGED.Databricks documentation states that the conversion can be rolled back with...

  • 39 Views
  • 0 replies
  • 0 kudos
david888
by • New Contributor II
  • 157 Views
  • 1 replies
  • 1 kudos

Question about MySQL Integrated‑CDC pipeline on Classic compute workspace

Hi everyone,I am trying to set up an Integrated‑CDC pipeline for a MySQL RDS instance on our Databricks workspace. Our workspace only uses Classic compute; Serverless compute is not available.According to the documentation, Integrated‑CDC for MySQL s...

  • 157 Views
  • 1 replies
  • 1 kudos
Latest Reply
ivanvyd
New Contributor III
  • 1 kudos

@david888 the missing wizard is expected. MySQL Integrated CDC currently runs on classic compute only, and UI-based pipeline creation isn't available. Workspace-level feature enablement is required; the docs direct you to your Databricks account team...

  • 1 kudos
Data_Engineer14
by • New Contributor II
  • 362 Views
  • 2 replies
  • 1 kudos

Resolved! Storage Credential creation fails with "Access Connector ... could not be found"

Subject: Storage Credential creation fails with "Access Connector ... could not be found" despite fully verified Azure configurationHi all,I'm trying to create a storage credential (Azure Managed Identity) for Unity Catalog on a Premium Azure Databri...

  • 362 Views
  • 2 replies
  • 1 kudos
Latest Reply
Ashwin_DSA
Databricks Employee
  • 1 kudos

@Data_Engineer14, Great troubleshooting so far. The fact that every Azure-side check passes cleanly but the storage credential creation still fails is a strong signal that this is a backend registration issue, not a misconfiguration on your end.   Wh...

  • 1 kudos
1 More Replies
Sandeep11
by • New Contributor II
  • 386 Views
  • 4 replies
  • 2 kudos

Incremental Load Issue with SDP

We have identified a critical issue in our pipeline and wanted to share it here to see if others have faced the same and how they approached it.Pipeline architecture: Bronze (Lakeflow Connect) → Silver (SCD2) → GoldThe Problem: Every pipeline run tri...

  • 386 Views
  • 4 replies
  • 2 kudos
Latest Reply
Sandeep11
New Contributor II
  • 2 kudos

Thank you all for the responses. We could get into a solution by adding readChangeFeed flag as true as additional option to the spark.readStream however kept the pipeline configs same.

  • 2 kudos
3 More Replies
1pedroosilva
by • New Contributor III
  • 509 Views
  • 5 replies
  • 9 kudos

Resolved! Databricks Dashboard Pivot Table: default collapsed state?

I'm building a financial dashboard in Databricks using a Pivot visualization with two dimensions in the row hierarchy.The drill-down works as expected, but I'm having a usability issue: every time the view is loaded, both dimensions are fully expande...

  • 509 Views
  • 5 replies
  • 9 kudos
Latest Reply
1pedroosilva
New Contributor III
  • 9 kudos

Thanks everyone for the input.Marking @data_pulse's reply as the solution: it confirms through testing that the expand/collapse state isn't persisted in the published dashboard, so there's currently no setting for an initial collapsed state or a defa...

  • 9 kudos
4 More Replies
Labels