cancel
Showing results for 
Search instead for 
Did you mean: 
Data Engineering
Join discussions on data engineering best practices, architectures, and optimization strategies within the Databricks Community. Exchange insights and solutions with fellow data engineers.
cancel
Showing results for 
Search instead for 
Did you mean: 

Forum Posts

hayden_blair
by New Contributor III
  • 2088 Views
  • 2 replies
  • 0 kudos

Why Shared Access Mode for Unity Catalog enabled DLT pipeline?

Hello all,I am trying to use an RDD API in a Unity Catalog enabled Delta Live Tables pipeline.I am getting an error because Unity Catalog enabled DLT can only run on "shared access mode" compute, and RDD APIs are not supported on shared access comput...

  • 2088 Views
  • 2 replies
  • 0 kudos
Latest Reply
hayden_blair
New Contributor III
  • 0 kudos

Thank you for the response @szymon_dybczak. Do you know if single user clusters are inherently less secure? I am still curious about why single user access mode is not allowed for DLT + Unity Catalog.

  • 0 kudos
1 More Replies
Chandru
by Databricks Partner
  • 7529 Views
  • 3 replies
  • 7 kudos

Resolved! Issue in importing librosa library while using databricks runtime engine 11.2

I have installed the library via PyPI on the cluster. When we import the package on notebook, getting the following errorimport librosaOSError: cannot load library 'libsndfile.so': libsndfile.so: cannot open shared object file: No such file or direct...

  • 7529 Views
  • 3 replies
  • 7 kudos
Latest Reply
Flo
New Contributor III
  • 7 kudos

If anybody ends up here after 2024: the init file must now be placed in the workspace for the cluster to accept it.So in Workspace, use Create/File to create the init script.Then add it to the cluster config inCompute - Your cluster - Advanced Config...

  • 7 kudos
2 More Replies
sinclair
by New Contributor II
  • 5264 Views
  • 6 replies
  • 1 kudos

Py4JJavaError: An error occurred while calling o465.coun

The following error occured when running .count() on a big sparkDF. Py4JJavaError: An error occurred while calling o465.count. : org.apache.spark.SparkException: Job aborted due to stage failure: Task 6 in stage 3.0 failed 4 times, most recent failur...

  • 5264 Views
  • 6 replies
  • 1 kudos
Latest Reply
RishabhTiwari07
Community Manager
  • 1 kudos

Hi @sinclair , Thank you for reaching out to our community! We're here to help you.  To ensure we provide you with the best support, could you please take a moment to review the response and choose the one that best answers your question? Your feedba...

  • 1 kudos
5 More Replies
NandaKishoreI
by New Contributor II
  • 1591 Views
  • 1 replies
  • 0 kudos

Databricks upon inserting delta table data inserts into folders in Dev

We have a Delta Table in Databricks. When we are inserting data into the Delta Table, in the storage account, it creates folders like: 05, 0H, 0F, 0O, 1T,1W, etc... and adds the parquet files there.We have not defined any partitions. We are inserting...

  • 1591 Views
  • 1 replies
  • 0 kudos
Latest Reply
RishabhTiwari07
Community Manager
  • 0 kudos

Hi @NandaKishoreI , Thank you for reaching out to our community! We're here to help you.  To ensure we provide you with the best support, could you please take a moment to review the response and choose the one that best answers your question? Your f...

  • 0 kudos
spicysheep
by New Contributor II
  • 3549 Views
  • 3 replies
  • 2 kudos

Where to find comprehensive docs on databricks.yaml / DAB settings options

Where can I find documentation on how to set cluster settings (e.g., AWS instance type, spot vs on-demand, number of machines) in Databricks Asset Bundle databicks.yaml files? The only documentation I've come across mentions these things indirectly, ...

  • 3549 Views
  • 3 replies
  • 2 kudos
Latest Reply
RishabhTiwari07
Community Manager
  • 2 kudos

Hi @spicysheep , Thank you for reaching out to our community! We're here to help you.  To ensure we provide you with the best support, could you please take a moment to review the response and choose the one that best answers your question? Your feed...

  • 2 kudos
2 More Replies
inagar
by New Contributor
  • 2869 Views
  • 1 replies
  • 0 kudos

Copying file from DBFS to a table of Databricks, Is there a way to get the errors at record level ?

We have file of data to be ingested into a table of Databricks. Following below approach,Uploaded file to DBFSCreating a temporary table and loading above file to the temporary table. CREATE TABLE [USING] Use MERGE INTO to merge temp_table created in...

  • 2869 Views
  • 1 replies
  • 0 kudos
Latest Reply
RishabhTiwari07
Community Manager
  • 0 kudos

Hi @inagar , Thank you for reaching out to our community! We're here to help you.  To ensure we provide you with the best support, could you please take a moment to review the response and choose the one that best answers your question? Your feedback...

  • 0 kudos
Maatari
by New Contributor III
  • 4758 Views
  • 2 replies
  • 0 kudos

Resolved! How to monitor Kafka consumption / lag when working with spark structured streaming?

I have just find out spark structured streaming do not commit offset to kafka but use its internal checkpoint system and that there is no way to visualize its consumption lag in typical kafka UI- https://community.databricks.com/t5/data-engineering/c...

  • 4758 Views
  • 2 replies
  • 0 kudos
Latest Reply
RishabhTiwari07
Community Manager
  • 0 kudos

Hi @Maatari , Thank you for reaching out to our community! We're here to help you.  To ensure we provide you with the best support, could you please take a moment to review the response and choose the one that best answers your question? Your feedbac...

  • 0 kudos
1 More Replies
thackman
by Databricks Partner
  • 21354 Views
  • 5 replies
  • 0 kudos

Databricks cluster random slow start times.

We have a job that runs on single user job compute because we've had compatibility issues switching to shared compute.Normally the cluster (1 driver,1 worker) takes five to six minutes to start. This is on Azure and we only include two small python l...

thackman_1-1720639616797.png thackman_0-1720639478363.png
  • 21354 Views
  • 5 replies
  • 0 kudos
Latest Reply
RishabhTiwari07
Community Manager
  • 0 kudos

Hi @thackman , Thank you for reaching out to our community! We're here to help you.  To ensure we provide you with the best support, could you please take a moment to review the response and choose the one that best answers your question? Your feedba...

  • 0 kudos
4 More Replies
ashraf1395
by Honored Contributor
  • 1300 Views
  • 1 replies
  • 0 kudos

Spark code not running bcz of incorrect compute size

I have a dataset having 260 billion recordsI need to group by 4 columns and find out the sum on four other columnsI increased the memory to e32 for driver and workers nodes, max workers is 40The job still is stuck in this aggregate step where I’m wri...

  • 1300 Views
  • 1 replies
  • 0 kudos
Latest Reply
RishabhTiwari07
Community Manager
  • 0 kudos

Hi @ashraf1395 , Thank you for reaching out to our community! We're here to help you.  To ensure we provide you with the best support, could you please take a moment to review the response and choose the one that best answers your question? Your feed...

  • 0 kudos
seefoods
by Valued Contributor
  • 1123 Views
  • 1 replies
  • 2 kudos

audit log for workspace users

Hello Everyone, How to retrieve trace execution of a Notebook databricks GCP Users Workspace.  Thanks

  • 1123 Views
  • 1 replies
  • 2 kudos
Latest Reply
szymon_dybczak
Esteemed Contributor III
  • 2 kudos

Hi @seefoods ,I think you can use system tables to get such information:https://docs.databricks.com/en/admin/system-tables/audit-logs.html

  • 2 kudos
WAHID
by New Contributor II
  • 1007 Views
  • 0 replies
  • 0 kudos

GDAL on Databricks serverless compute

I am wondering if it's possible to install and use GDAL on Databricks serverless compute. I couldn't manage to do that using pip install gdal, and I discovered that init scripts are not supported on serverless compute.

  • 1007 Views
  • 0 replies
  • 0 kudos
mr_robot
by New Contributor
  • 5067 Views
  • 3 replies
  • 3 kudos

Update datatype of a column in a table

I have a table in databricks with fields name: string, id: string, orgId: bigint, metadata: struct, now i want to rename one of the columns and change it type. In my case I want to update orgId to orgIds and change its type to map<string, string> One...

Data Engineering
tables delta-tables
  • 5067 Views
  • 3 replies
  • 3 kudos
Latest Reply
jacovangelder
Databricks MVP
  • 3 kudos

You can use REPLACE COLUMNS.ALTER TABLE your_table_name REPLACE COLUMNS ( name STRING, id BIGINT, orgIds MAP<STRING, STRING>, metadata STRUCT<...> );

  • 3 kudos
2 More Replies
ashraf1395
by Honored Contributor
  • 1697 Views
  • 1 replies
  • 1 kudos

Resolved! Querying Mysql db from Azure databricks where public access is disabled

Hi there,We are trying to setup a infra that ingest data from MySQL hosted on awa EC2 instance with pyspark and azure databricks and dump to the adls storage.Since databases has public accessibility disabled and how can I interact with MySQL from azu...

  • 1697 Views
  • 1 replies
  • 1 kudos
Latest Reply
-werners-
Esteemed Contributor III
  • 1 kudos

you will need some kind of tunnel that opens the db server to external access.Perhaps a vpn is an option?If not: won't be possible.An alternative way would be to have some local LAN tool extract the data and then move it to S3/... and afterwards let ...

  • 1 kudos
vvzadvor
by New Contributor III
  • 4812 Views
  • 4 replies
  • 2 kudos

Resolved! Debugging python code outside of Notebooks

Hi experts,Does anyone know if there's a way of properly debugging python code outside of notebooks?We have a complicated python-based framework for loading files, transforming them according to the business specification and saving the results into ...

  • 4812 Views
  • 4 replies
  • 2 kudos
Latest Reply
vvzadvor
New Contributor III
  • 2 kudos

OK, I can now confirm that remote debugging with stepping into your own libraries installed on the cluster is possible and is actually pretty convenient using a combination of databricks-connect Python library and a Databricks extension for VSCode. S...

  • 2 kudos
3 More Replies
ashraf1395
by Honored Contributor
  • 2461 Views
  • 1 replies
  • 2 kudos

Resolved! Reading a materialised view locally or using databricks api

Hi there, This was my previous approach - I had a databricks notebook with a streaming table bronze level reading data from volumes which created a 2 downstream tables.- 1st A a materialised view gold level, another a table for storing ingestion_meta...

  • 2461 Views
  • 1 replies
  • 2 kudos
Latest Reply
ashraf1395
Honored Contributor
  • 2 kudos

I used this approach - Querying the materialised view using databricks serverless SQL endpoint by connecting it with SQL connect. Its working right now. If I face any issues, I will write it into a normal table and delta share it.Thanks for your repl...

  • 2 kudos
Labels