cancel
Showing results for 
Search instead for 
Did you mean: 
Data Engineering
Join discussions on data engineering best practices, architectures, and optimization strategies within the Databricks Community. Exchange insights and solutions with fellow data engineers.
cancel
Showing results for 
Search instead for 
Did you mean: 

Forum Posts

saab123
by • New Contributor II
  • 5833 Views
  • 1 replies
  • 0 kudos

Not able to connect to Neo4j Aura Db from databricks

I am trying to connect to a Neo4j AuraDb instance-f9374927. Created a free professional instance of Neo4j. I am able to connect to this instance, add nodes and relationships.   Created a Databricks shared cluster 14.3 LTS (includes Apache Spark 3.5.0...

  • 5833 Views
  • 1 replies
  • 0 kudos
Latest Reply
mark_ott
Databricks Employee
  • 0 kudos

The connection issue between your Databricks cluster and Neo4j AuraDB instance (f9374927) with the ServiceUnavailableException: No routing server available message is tied to network-level SSL configuration and connectivity rather than incorrect code...

  • 0 kudos
dbxlearner
by • New Contributor II
  • 6029 Views
  • 3 replies
  • 1 kudos

Resolved! Deploying using Databricks asset bundles (DABs) in a closed network

Hello, I'm trying to deploy DBX workflows using DABs using an Azure DevOps pipeline, in a network that cannot download the required terraform databricks provider package online, due to firewall/network restrictions.I have followed this post: https://...

  • 6029 Views
  • 3 replies
  • 1 kudos
Latest Reply
dbxlearner
New Contributor II
  • 1 kudos

Another thing I noticed is, when running the 'databricks bundle debug terraform' command, it mentions these variables:I have tried setting these variables as environment variables in my ADO pipeline, specially the databricks terraform provider variab...

  • 1 kudos
2 More Replies
janglais
by • New Contributor
  • 2471 Views
  • 2 replies
  • 0 kudos

Resolved! DLT Pipeline with unknown deleted source data

Hello.. I need help. So the context is : - ERP data for company in my group is stored in sql tables - Currently, once per day we copy the last 2 months of data (creation date) from each table into our datalake landing zone (we can however do full cop...

  • 2471 Views
  • 2 replies
  • 0 kudos
Latest Reply
madams
Contributor III
  • 0 kudos

Your solution #1 is very frustrating to me as well, for a number of reasons.  Simply put, we have to be able to compare incoming data to target data for normal ETL operations. One way around this is to create a view of your target silver table, outsi...

  • 0 kudos
1 More Replies
kevinzhang29
by • New Contributor III
  • 2446 Views
  • 1 replies
  • 1 kudos

Resolved! DLT pipeline failed: streaming table query reading from an unexpected Delta table ID

Hi everyone,I'm running a DLT pipeline that loads data from Bronze to Silver using dlt.apply_changes(SCD type 2)The first run of the pipeline worked fine -- data was written successfully into the target Silver tables.However, when I ingested new data...

  • 2446 Views
  • 1 replies
  • 1 kudos
Latest Reply
mark_ott
Databricks Employee
  • 1 kudos

This “unexpected Delta table ID” error typically means your Delta Live Tables (DLT) pipeline detected that the underlying Delta table it was reading from has changed since the last checkpoint. When you use dlt.apply_changes() (for SCD Type 2), this i...

  • 1 kudos
Mits11
by • New Contributor II
  • 1156 Views
  • 2 replies
  • 1 kudos

Community edition cluster - UI shows incorrect cores

Hi,I am a community edition user which gives me cluster ( as per below image)15GB of memory and 2 cores with one driver node ONLY.However,when I read a csv file of 181MB size,1) it generates 8 partitiones.As per default maxPartitionBytes is set to 12...

Mits11_1-1761165245208.png Mits11_3-1761165673802.png Mits11_2-1761165566839.png
  • 1156 Views
  • 2 replies
  • 1 kudos
Latest Reply
Mits11
New Contributor II
  • 1 kudos

Thank you Louis for detailed explaination.Including notifying me about CE updates.However, I have noticed this ( below is the screenshot)spark.sql.files.minPartitionNum does not restun any result.Its wierd.Am I missing anything?Thanks 

  • 1 kudos
1 More Replies
Ashok_Vengala
by • New Contributor
  • 1679 Views
  • 1 replies
  • 1 kudos

Resolved! Unable to Add Multiple Columns in Single ALTER TABLE Statement on Iceberg Table via Unity REST Catal

Hello Databricks Team,I have implemented code to integrate the Iceberg Unity REST Catalog with the Teradata OTF engine and successfully performed read and write operations, following the documentation at https://docs.databricks.com/aws/en/external-ac...

  • 1679 Views
  • 1 replies
  • 1 kudos
Latest Reply
nayan_wylde
Esteemed Contributor II
  • 1 kudos

This error stems from the Iceberg table metadata update constraints enforced by the Unity Catalog's REST API. Specifically, the Iceberg REST Catalog currently does not support multiple schema changes in a single commit. Each ALTER TABLE operation tha...

  • 1 kudos
TejeshS
by • Contributor II
  • 6916 Views
  • 3 replies
  • 1 kudos

How to identify which columns we need to consider for liquid clustering from a table of 200+ columns

In Databricks, when working with a table that has a large number of columns (e.g., 200), it can be challenging to determine which columns are most important for liquid clustering.Objective: The goal is to determine which columns to select based on th...

  • 6916 Views
  • 3 replies
  • 1 kudos
Latest Reply
noorbasha534
Valued Contributor II
  • 1 kudos

@Alberto_Umana is it possible to get from system table the columns used in joins & filters of a table being queried?

  • 1 kudos
2 More Replies
Alby091
by • New Contributor
  • 3713 Views
  • 2 replies
  • 0 kudos

Multiple schedules in workflow with different parameters

I have a notebook that takes a file from the landing, processes it and saves a delta table.This notebook contains a parameter (time_prm) that allows you to do this option for the different versions of files that arrive every day.Specifically, for eac...

Data Engineering
parameters
Workflows
  • 3713 Views
  • 2 replies
  • 0 kudos
Latest Reply
ImranA
Contributor
  • 0 kudos

You can do multiple schedules with Cron expression. If you are using a Cron expression in Databricks asset bundle YAML, but the limitation is you can't have one running at 0 past the hour and another at 25 past.i.e: quartz_cron_expression: 0 45 9,23 ...

  • 0 kudos
1 More Replies
Spenyo
by • New Contributor II
  • 2336 Views
  • 1 replies
  • 1 kudos

Delta table size not shrinking after Vacuum

Hi team.Everyday once we overwrite the last X month data in tables. So it generate a every day a larger amount of history. We don't use time travel so we don't need it.What we done:SET spark.databricks.delta.retentionDurationCheck.enabled = false ALT...

chrome_KZMxPl8x1d.png
  • 2336 Views
  • 1 replies
  • 1 kudos
Latest Reply
pabloaschieri
New Contributor II
  • 1 kudos

Hi, any update on this? Thanks

  • 1 kudos
fjrodriguez
by • New Contributor III
  • 1439 Views
  • 2 replies
  • 1 kudos

Resolved! Ingestion Framework

I would to like to update my ingestion framework that is orchestrated by ADF, running couples Databricks notebook and copying the data to DB afterwards. I want to rely everything on Databricks i though this could be the design:Step 1. Expose target t...

  • 1439 Views
  • 2 replies
  • 1 kudos
Latest Reply
fjrodriguez
New Contributor III
  • 1 kudos

Hey @saurabh18cs , It is taking longer than expected to expose Azure SQL tables in UC. I can do that through Foreign Catalog but this is not what i want due to is read-only. As far i can see external connection is for cloud object storage paths (ADLS...

  • 1 kudos
1 More Replies
Rjdudley
by • Honored Contributor
  • 2120 Views
  • 3 replies
  • 0 kudos

Resolved! AUTO CDC API and sequence column

The docs for AUTO CDC API stateYou must specify a column in the source data on which to sequence records, which Lakeflow Declarative Pipelines interprets as a monotonically increasing representation of the proper ordering of the source data.Can this ...

  • 2120 Views
  • 3 replies
  • 0 kudos
Latest Reply
Rjdudley
Honored Contributor
  • 0 kudos

Thanks Szymon, I'm familiar with the Postgre SQL implementation and was hoping Databricks would behave the same.

  • 0 kudos
2 More Replies
ankit001mittal
by • New Contributor III
  • 3177 Views
  • 1 replies
  • 2 kudos

DLT schema evolution/changes in the logs

Hi all,I want to figure out how to find when the schema evolution/changes are happening in the objects in DLT pipelines through the DLT logs.Could you please share some sample DLT logs which explains about the schema changes?Thank you for your help.

  • 3177 Views
  • 1 replies
  • 2 kudos
Latest Reply
mark_ott
Databricks Employee
  • 2 kudos

To find when schema evolution or changes are happening in objects within DLT (Delta Live Table) pipelines, you need to monitor certain entries within the DLT logs or Delta transaction logs that signal modifications to the underlying schema of a table...

  • 2 kudos
minhhung0507
by • Valued Contributor
  • 3790 Views
  • 3 replies
  • 0 kudos

DLT Flow Failed Due to Missing Flow Checkpoints Directory When Using Unity Catalog

I’m encountering an issue while running a Delta Live Tables (DLT) pipeline that is managed using Unity Catalog on Databricks. The pipeline has failed and is not restarting, showing the following error:java.lang.IllegalArgumentException: flow checkpoi...

  • 3790 Views
  • 3 replies
  • 0 kudos
Latest Reply
mark_ott
Databricks Employee
  • 0 kudos

The best practices for setting up checkpointing in Delta Live Tables (DLT) pipelines when using Unity Catalog are largely centered on leveraging Databricks' managed services, adhering to Unity Catalog's table management conventions, and minimizing th...

  • 0 kudos
2 More Replies
sumitkumar_284
by • Databricks Partner
  • 3045 Views
  • 4 replies
  • 1 kudos

Not able to refresh powerbi dashboar form databricks jobs

I am trying to refresh Power BI Dashboard using Databricks jobs and constantly getting this error, but I am providing optional parameters which includes catalog and database. Also, things to note that I am able to do refresh on Power BI UI using both...

  • 3045 Views
  • 4 replies
  • 1 kudos
Latest Reply
szymon_dybczak
Esteemed Contributor III
  • 1 kudos

Hi @sumitkumar_284 ,Can you provide us more details? Are you using Unity Catalog? Which authentication mechanism you have? In which version of Power BI Desktop you've developed your semantic model/dashboard? Do you meet all below requirements?Publish...

  • 1 kudos
3 More Replies
maninegi05
by • New Contributor II
  • 1796 Views
  • 3 replies
  • 1 kudos

Resolved! DLT Pipeline Stopped working

Hello, Suddenly our DLT pipelines we're getting failures saying thatLookupError: Traceback (most recent call last): result_df = result_df.withColumn("input_file_path", col("_metadata.file_path")).withColumn( ...

  • 1796 Views
  • 3 replies
  • 1 kudos
Latest Reply
Louis_Frolio
Databricks Employee
  • 1 kudos

Greetings @maninegi05 , I did some digging internally and I believe some recent changes to the DLT image may be to blame. We are aware of regression issue and are actively working to address them. TL/DR Why you might see “LookupError: ContextVar 'par...

  • 1 kudos
2 More Replies
Labels