cancel
Showing results for 
Search instead for 
Did you mean: 
Data Engineering
Join discussions on data engineering best practices, architectures, and optimization strategies within the Databricks Community. Exchange insights and solutions with fellow data engineers.
cancel
Showing results for 
Search instead for 
Did you mean: 

Forum Posts

pvrcloudtech
by • Visitor
  • 46 Views
  • 2 replies
  • 0 kudos

Databricks SDP (Spark declarative Pipelins) Overwrite table

1) Consider I have orders folder which orders_1.csv file  and it is loaded to orders tables orders/                 ============> load to order table      orders_1.csv  2) next day new file (orders_2.csv) arrived orders/                       =======...

  • 46 Views
  • 2 replies
  • 0 kudos
Latest Reply
vannurswamy
New Contributor
  • 0 kudos

A Streaming Table may not be the right fit for this requirement. It is designed to process new files incrementally, so when orders2.csv arrives, it will append the new data.If each new file is a full snapshot and should completely replace the previou...

  • 0 kudos
1 More Replies
suryaprayaga
by • Contributor
  • 178 Views
  • 2 replies
  • 2 kudos

Genie Use Cases & easy adaptibility

Recently I started socializing the use of Genie and Databricks to even a laymen (and ofcourse laywomen) who have never heard what coding is in their life. I am proud that in my company nearly 200 people as of now are using Databricks for some or othe...

  • 178 Views
  • 2 replies
  • 2 kudos
Latest Reply
DoTA
Valued Contributor II
  • 2 kudos

Hi, congratulations on getting to 200 people, and especially on getting non-coders productive. From our own experience of scaling an internal data and AI platform to a couple thousand users, the things that mattered most once adoption passed the firs...

  • 2 kudos
1 More Replies
Mado
by • Valued Contributor II
  • 382 Views
  • 7 replies
  • 7 kudos

How can I configure Lakeflow Connect SQL Server CDC Gateway to use a desired VM type?

Hi Team,I'm evaluating Databricks Lakeflow Connect for SQL Server CDC ingestion and have run into a gateway provisioning issue.EnvironmentRegion: Australia EastSource: Azure SQL DatabaseCDC enabled successfully on the database and source tableSQL Ser...

Mado_0-1790214214728.png
  • 382 Views
  • 7 replies
  • 7 kudos
Latest Reply
AbhilashNagilla
Databricks Employee
  • 7 kudos

You can pin the gateway's node types with that policy. The SQL Server ingestion page lists a custom gateway policy as "API only". Attach it in the gateway pipeline's clusters setting, with apply_policy_default_values set to true so the policy's defau...

  • 7 kudos
6 More Replies
alexisjohnson
by • New Contributor III
  • 25005 Views
  • 6 replies
  • 7 kudos

Resolved! Window function using last/last_value with PARTITION BY/ORDER BY has unexpected results

Hi, I'm wondering if this is the expected behavior when using last or last_value in a window function? I've written a query like this:select col1, col2, last_value(col2) over (partition by col1 order by col2) as column2_last from values ...

Screen Shot 2021-11-18 at 12.48.25 PM Screen Shot 2021-11-18 at 12.48.32 PM
  • 25005 Views
  • 6 replies
  • 7 kudos
Latest Reply
asbivu
New Contributor
  • 7 kudos

The behavior of LAST_VALUE with PARTITION BY and ORDER BY can be confusing when the window frame affects which row is considered the last value. I’ve found that checking the window frame definition carefully helps make these results easier to underst...

  • 7 kudos
5 More Replies
Rajasaiharish
by • New Contributor III
  • 87 Views
  • 2 replies
  • 0 kudos

Analyze

 How many types of analyze do we have for a UC delta table ? And how to check which type of ANALY

  • 87 Views
  • 2 replies
  • 0 kudos
Latest Reply
Rajasaiharish
New Contributor III
  • 0 kudos

How many types of analyze do we have for a UC delta table ? And how to check which type of ANALYZE ran for what table? i think total questions was not printed correctly

  • 0 kudos
1 More Replies
ram_11
by • New Contributor
  • 63 Views
  • 1 replies
  • 0 kudos

Lakeflow connect

Spoiler  What is extractor_sql-server-conn_cdc_sink in the Ingestion gateway pipeline in lakeflow connect and is it crreated by deafualta and how does it work?

  • 63 Views
  • 1 replies
  • 0 kudos
Latest Reply
K_Anudeep
Databricks Employee
  • 0 kudos

Hi @ram_11 ! The extractor_sql-server-conn_cdc_sink is the shared CDC sink that the SQL Server CDC extractor job inside the IG pipeline writes to.Example screenshot below:     What actually happens: The IG extractor (the job/process) reads change ro...

  • 0 kudos
priya9896
by • New Contributor
  • 192 Views
  • 3 replies
  • 0 kudos

Azure to AWS

Hi everyone,We're evaluating an architecture pattern and would appreciate any guidance or recommendations. Has anyone implemented a similar cross-cloud pattern? Specifically:Can AWS Glue be used to access Azure Databricks data without copying it?Are ...

  • 192 Views
  • 3 replies
  • 0 kudos
Latest Reply
Louis_Frolio
Databricks Employee
  • 0 kudos

Hello @priya9896, I took a look at both internal and external documentation and here is what I found. Straight answer: don't plan on the Glue Data Catalog as a federation layer over Azure Databricks. Glue's Delta integration and the Glue Catalog itse...

  • 0 kudos
2 More Replies
dbernstein_tp
by • Contributor
  • 273 Views
  • 7 replies
  • 1 kudos

Lakeflow connect SQL server ingestion can be made elastic?

Hi Everyone, One of our big ingestion tasks is lakeflow connect CDC ingestion of ERP data from a SQL server database. I am deploying the pipelines and jobs for this via DABs. We are ingesting about 80 tables from the database, a handful of which are ...

  • 273 Views
  • 7 replies
  • 1 kudos
Latest Reply
lmcorreahdb
New Contributor III
  • 1 kudos

Great discussion. I think there is another architectural consideration worth exploring before changing the gateway configuration.Does the 15-minute freshness requirement apply to all 80 tables, or only to a subset of business-critical tables?If only ...

  • 1 kudos
6 More Replies
Kirankumarbs
by • Valued Contributor III
  • 3412 Views
  • 5 replies
  • 2 kudos

Python logger.info() not showing inside applyInPandas (but print() works) — why?

Problem: In Databricks, logs from an external binary (via os.system) show up, but Python logger.info() inside groupBy(...).applyInPandas(...) does not. print(..., flush=True) does show up.Why: applyInPandas runs your function as a pandas UDF.  That c...

  • 3412 Views
  • 5 replies
  • 2 kudos
Latest Reply
SteveOstrowski
Databricks Employee
  • 2 kudos

Hi Ben, The stdout logging-handler trick works on a classic/standard cluster, but on serverless it will not surface in the notebook cell, and that is expected rather than a problem with the helper. Here is what is going on. applyInPandas runs your Py...

  • 2 kudos
4 More Replies
wschoi
by • New Contributor III
  • 24200 Views
  • 18 replies
  • 17 kudos

How to fix plots and image color rendering on Notebooks?

I am currently running dark mode for my Databricks Notebooks, and am using the "new UI" released a few days ago (May 2023) and the "New notebook editor."Currently all plots (like matplotlib) are showing wrong colors. For example, denoting:```... p...

  • 24200 Views
  • 18 replies
  • 17 kudos
Latest Reply
DavidLin
Databricks Employee
  • 17 kudos

HTML outputs in dark mode was recently improved. Please see https://docs.databricks.com/aws/en/notebooks/notebook-ui#html-outputs-in-dark-mode and try it out!

  • 17 kudos
17 More Replies
Brahmareddy
by • Esteemed Contributor II
  • 3940 Views
  • 9 replies
  • 12 kudos

Future of Movie Discovery: How I Built an AI Movie Recommendation Agent on Databricks Free Edition

As a data engineer deeply passionate about how data and AI can come together to create real-world impact, I’m excited to share my project for the Databricks Free Edition Hackathon 2025 — Future of Movie Discovery (FMD). Built entirely on Databricks F...

  • 3940 Views
  • 9 replies
  • 12 kudos
Latest Reply
juanlozadab
New Contributor III
  • 12 kudos

@Brahmareddy I haven't used this exact pattern at multi-million-row scale, but I agree that the tradeoffs start to change as the dataset and query volume grow.I would probably keep Delta as the source of truth and separate the serving layer based on ...

  • 12 kudos
8 More Replies
LiresaFerizaj
by • Contributor
  • 323 Views
  • 5 replies
  • 4 kudos

AUTO CDC SCD Type 2 with late-arriving deletes: edge cases the docs don't cover

Hi everyone,I'm designing a Lakeflow Declarative Pipeline that processes customer profile changes from a CDC feed into an SCD Type 2 Silver table. Events can arrive out of order, sometimes up to 24 hours late. The planned flow:CREATE OR REFRESH STREA...

  • 323 Views
  • 5 replies
  • 4 kudos
Latest Reply
Louis_Frolio
Databricks Employee
  • 4 kudos

Hello @LiresaFerizaj , I took a look at both internal and external documentation and here is what I found. You've already had solid answers from Thomaz, Nitesh, data_pulse, and Islam, so I'll tie the thread together, separate what's documented from w...

  • 4 kudos
4 More Replies
david888
by • New Contributor II
  • 211 Views
  • 2 replies
  • 1 kudos

Question about MySQL Integrated‑CDC pipeline on Classic compute workspace

Hi everyone,I am trying to set up an Integrated‑CDC pipeline for a MySQL RDS instance on our Databricks workspace. Our workspace only uses Classic compute; Serverless compute is not available.According to the documentation, Integrated‑CDC for MySQL s...

  • 211 Views
  • 2 replies
  • 1 kudos
Latest Reply
Louis_Frolio
Databricks Employee
  • 1 kudos

Hello @david888 , I took a look at both internal and external documentation and here is what I found. Ivan (@ivanvyd) has this right, so I'll build on his answer rather than repeat it.   Taking your three questions in order. Feature flag. Yes, enabl...

  • 1 kudos
1 More Replies
Davide
by • New Contributor II
  • 378 Views
  • 4 replies
  • 0 kudos

Resolved! Failed to convert to managed table: PrivilegedGenerateTemporaryTableCredential disable

Hello everyone, I noticed that when I tried to convert some external table to managed one using the alter table catalog.schema.table_name set managed command (like suggested by databricks) sometimes I will encounter this error: Error running query: c...

  • 378 Views
  • 4 replies
  • 0 kudos
Latest Reply
Davide
New Contributor II
  • 0 kudos

Hello everyone, After following up with Databricks, they recognize it is a bug on their side therefore are starting the process to fix it, in the mean time the best way to deal with it according to Databricks is simply to setup a retry and it should ...

  • 0 kudos
3 More Replies
Davide
by • New Contributor II
  • 226 Views
  • 2 replies
  • 1 kudos

Resolved! SET MANAGED AND ROLLBACK IN RELATION TO VACUUM

Hello everyone,I am looking for clarification about the 14-day rollback period after converting an external table to a managed table using: ALTER TABLE table_name SET MANAGED.Databricks documentation states that the conversion can be rolled back with...

  • 226 Views
  • 2 replies
  • 1 kudos
Latest Reply
DoTA
Valued Contributor II
  • 1 kudos

Hi, I could not find these three cases documented either, so I will separate what the docs say from what I would treat as an assumption. What is documented (Convert external tables to managed): Databricks keeps the data in the old external location f...

  • 1 kudos
1 More Replies
Labels