Data Engineering

Forum Posts

Sorted by:

by bfridley • New Contributor II

09-21-2023 1:34:06 PM

2253 Views
2 replies
0 kudos

DLT Pipeline Out Of Memory Errors

I have a DLT pipeline that has been running for weeks. Now, trying to rerun the pipeline with the same code and same data fails. I've even tried updating the compute on the cluster to about 3x of what was previously working and it still fails with ou...

Data Engineering

2253 Views
2 replies
0 kudos

09-21-2023 1:34:06 PM

View Replies

Latest Reply

rajib_bahar_ptg
New Contributor III

09-21-2023 3:24:36 PM

0 kudos

I'd focus on understanding the codebase first. It'll help you decide what logic or data asset to keep or not keep when you try to optimize it. If you share the architecture of the application, the problem it solves, and some sample code here, it'll h...

0 kudos

09-21-2023 3:24:36 PM

1 More Replies

by gkrilis • New Contributor

06-29-2023 3:02:28 AM

4379 Views
1 replies
0 kudos

How to stop SparkSession within notebook without errr

I want to run an ETL job and when the job ends I would like to stop SparkSession to free my cluster's resources, by doing this I could avoid restarting the cluster, but when calling spark.stop() the job returns with status failed even though it has f...

Data Engineering

cluster

SparkSession

4379 Views
1 replies
0 kudos

06-29-2023 3:02:28 AM

View Replies

Latest Reply

PremadasV
New Contributor II

10-02-2023 7:25:37 PM

0 kudos

Please refer to this Job fails, but Apache Spark tasks finish - Databricks

0 kudos

10-02-2023 7:25:37 PM

by Martin1 • New Contributor II

05-05-2022 4:11:12 AM

8021 Views
3 replies
1 kudos

Referring to Azure Keyvault secrets in spark config

Hi allIn spark config for a cluster, it works well to refer to a Azure Keyvault secret in the "value" part of the name/value combo on a config row/setting.For example, this works fine (I've removed the string that is our specific storage account name...

Data Engineering

8021 Views
3 replies
1 kudos

05-05-2022 4:11:12 AM

View Replies

Latest Reply

kp12
New Contributor II

10-02-2023 3:31:23 AM

1 kudos

Hello,Is there any update on this issue please? Databricks no longer recommend mounting external location, so the other way to access Azure storage is to use spark config as mentioned in this document - https://learn.microsoft.com/en-us/azure/databri...

1 kudos

10-02-2023 3:31:23 AM

2 More Replies

by marvin1 • New Contributor III

09-29-2023 2:10:03 PM

209 Views
0 replies
0 kudos

Bamboolib error

What is the status of bamboolib? I understand that it is public preview but I'm unable to find any support references. I am getting error below. I've tried installing in a notebook, on a cluster, creating a pandas dataframe and running bam, etc. ...

Data Engineering

209 Views
0 replies
0 kudos

09-29-2023 2:10:03 PM

by mbvb_py • New Contributor II

07-12-2022 3:38:49 AM

2535 Views
4 replies
0 kudos

Create cluster error: Backend service unavailable

hello,i'm new to Databricks (community edition account) and encountered a problem just now.When creating a new cluster (default 10.4 LTS) it fails with the following error: Backend service unavailable.I've tried a different runtime > same issue.I've ...

Data Engineering

2535 Views
4 replies
0 kudos

07-12-2022 3:38:49 AM

View Replies

Latest Reply

stefnhuy
New Contributor III

09-29-2023 2:52:45 AM

0 kudos

Hey mbvb_py,I'm sorry to hear you're facing this "Backend service unavailable" issue with Databricks. I've encountered similar problems in the past, and it can be frustrating. Don't worry; you're not alone in this!From my experience, this error can o...

0 kudos

09-29-2023 2:52:45 AM

3 More Replies

by DBEnthusiast • New Contributor III

09-27-2023 12:47:43 AM

1609 Views
2 replies
0 kudos

How does Job Cluster knows how many resources to assign to an Application ?

Hi All Enthusiasts !As per my understanding when a user submits an application in spark cluster it specifies how much memory, executors etc. it would need . But in Data bricks notebooks we never specify that anywhere. If we have submitted the noteboo...

Data Engineering

1609 Views
2 replies
0 kudos

09-27-2023 12:47:43 AM

View Replies

Latest Reply

BilalAslamDbrx
Esteemed Contributor III

09-29-2023 2:04:53 AM

0 kudos

@DBEnthusiast great question! Today, with Job Clusters, you have to specify this. As @btafur note, you do this by setting CPU, memory etc. We are in early preview of Serverless Job Clusters where you no longer specify this configuration, instead Data...

0 kudos

09-29-2023 2:04:53 AM

1 More Replies

by smurug • New Contributor II

08-01-2023 11:30:03 AM

4107 Views
4 replies
1 kudos

Databricks Job scheduling - continuous mode

While scheduling the Databricks job using continuous mode - what will happen if the job is configured to run with Job cluster.At the end of each run will the cluster be terminated and re-created again for the next run? The official documentation is n...

Data Engineering

4107 Views
4 replies
1 kudos

08-01-2023 11:30:03 AM

View Replies

Latest Reply

Jo5h
New Contributor II

09-29-2023 1:49:19 AM

1 kudos

Hello @youssefmrini So how is the DBU calculated? As the cluster is reused, the DBU should be calculated per hour on all the jobs run in an hour correct? Or will it be calculated based on each run?I would like to know the cost calculation when runnin...

1 kudos

09-29-2023 1:49:19 AM

3 More Replies

by Mado • Valued Contributor II

11-15-2022 3:07:22 AM

25937 Views
4 replies
3 kudos

Resolved! How to set a variable and use it in a SQL query

I want to define a variable and use it in a query, like below: %sql SET database_name = "marketing"; SHOW TABLES in '${database_name}';However, I get the following error:ParseException: [PARSE_SYNTAX_ERROR] Syntax error at or near ''''(line 1, pos...

Data Engineering

25937 Views
4 replies
3 kudos

11-15-2022 3:07:22 AM

View Replies

Latest Reply

CJS
New Contributor II

09-28-2023 1:43:43 PM

3 kudos

Another option is demonstrated by this example:%sql SET database_name.var = marketing; SHOW TABLES in ${database_name.var}; SET database_name.dummy= marketing; SHOW TABLES in ${database_name.dummy};do not use quotesuse format that is variableName...

3 kudos

09-28-2023 1:43:43 PM

3 More Replies

by Hubert-Dudek • Esteemed Contributor III

09-26-2023 1:07:46 PM

897 Views
1 replies
1 kudos

Streaming Data Modeling Normalization with Databricks Delta Live Tables

Streamline Data Modeling Normalization with Databricks Delta Live Tables in Just a Few Steps:- Use the "Apply changes" function to populate tables with slowly changing dimensions using auto-increment IDs.- Register SQL mapping functions to associate ...

Data Engineering

897 Views
1 replies
1 kudos

09-26-2023 1:07:46 PM

View Replies

Latest Reply

jose_gonzalez
Moderator

09-28-2023 9:41:42 AM

1 kudos

Thank you for sharing this @Hubert-Dudek !!!

1 kudos

09-28-2023 9:41:42 AM

by BAZA • New Contributor II

06-28-2023 3:37:00 AM

6336 Views
9 replies
0 kudos

Invisible empty spaces when reading .csv files

When importing a .csv file with leading and/or trailing empty spaces around the separators, the output results in strings that appear to be trimmed on the output table or when using .display() but are not actually trimmed.It is possible to identify t...

Data Engineering

6336 Views
9 replies
0 kudos

06-28-2023 3:37:00 AM

View Replies

Latest Reply

Raluka
New Contributor III

09-27-2023 4:31:48 PM

0 kudos

I discovered an in-depth article that went beyond the physical aspects of aging and testosterone. It examined the emotional https://misterolympia.shop/buy/injectable-steroids/testosterone/testosterone-cypionate/ and psychological aspects of growing o...

0 kudos

09-27-2023 4:31:48 PM

8 More Replies

by Nico1 • New Contributor II

05-15-2022 2:58:10 PM

8745 Views
11 replies
2 kudos

Resolved! Problems connecting Simba ODBC with a M1 Macbook Pro

Hi,There's a way to make work the Simba ODBC Driver for M1 Macbook Pros?I find myself able to run on an old intel version of Macbook easily, but now every time I even test the connection with the iODBC Manager fails.Definitely, the issue is around no...

Data Engineering

8745 Views
11 replies
2 kudos

05-15-2022 2:58:10 PM

View Replies

Latest Reply

kunalmishra9
New Contributor III

09-27-2023 2:24:53 PM

2 kudos

Things seem to be mostly working for me now. I've added a bit more detail on my connection steps and process in case it's helpful for anyone on Stack Overflow: https://stackoverflow.com/questions/76407426/connecting-rstudio-desktop-to-databricks-comm...

2 kudos

09-27-2023 2:24:53 PM

10 More Replies

by zak_k • New Contributor III

09-25-2023 2:08:43 PM

3435 Views
5 replies
1 kudos

com.databricks.spark.safespark.UDFException: UNAVAILABLE: Channel shutdownNow invoked

Trying to determine a root cause of UDFException that occurs when returning a variable length ArrayType. If I hardcode the data returned from the UDF to a fixed length, say 19, the error does not occur. Setup codesplit_runs_UDF = udf(split_runs_udf, ...

Data Engineering

3435 Views
5 replies
1 kudos

09-25-2023 2:08:43 PM

View Replies

Latest Reply

zak_k
New Contributor III

09-27-2023 6:03:07 AM

1 kudos

After further investigation, It reproduces slightly differently on single user mode.Single user mode: runs foreverShared: gives the above messageI've determined that there was a corner case in the dataset which lead to UDF never returning. I am am as...

1 kudos

09-27-2023 6:03:07 AM

4 More Replies

by miiaramo • New Contributor II

05-19-2023 2:39:39 AM

1855 Views
2 replies
1 kudos

DLT current channel uses same runtime as the preview channel

Hi,According to the latest release notes, the current channel of DLT should be using Databricks runtime 11.3 and the preview channel should be using 12.2. The current channel was using correct runtime version 11.3 still yesterday morning, but since ...

Data Engineering

1855 Views
2 replies
1 kudos

05-19-2023 2:39:39 AM

View Replies

Latest Reply

adriennn
Contributor

09-27-2023 3:05:30 AM

1 kudos

I'm seeing the same issue with 12 current / 13 preview. Updating the channel didn't bump the runtime version and even creating a pipeline with the preview channel uses the current version.

1 kudos

09-27-2023 3:05:30 AM

1 More Replies

by Databricks143 • New Contributor III

09-25-2023 10:23:37 AM

2046 Views
4 replies
0 kudos

Correlated column is not allowed in non predicate in UDF SQL

Hi Team,I am new to databricks and currently working on creating sql udf 's in databricks .In udf we are calculating min date and that date column using in where clause also.While running udf getting Correlated column is not allowed in non predica...

Data Engineering

2046 Views
4 replies
0 kudos

09-25-2023 10:23:37 AM

View Replies

Latest Reply

Noopur_Nigam
Esteemed Contributor III

09-26-2023 3:58:50 AM

0 kudos

Could you please provide your full code? I would also like to know which DBR version you are using in your cluster.

0 kudos

09-26-2023 3:58:50 AM

3 More Replies

by thomann • New Contributor III

12-13-2022 8:48:53 AM

5669 Views
5 replies
6 kudos

Bug? Unity Catalog incompatible with Sparklyr in RStudio (on Driver) and as well if used on one cluster from multiple notebooks?

If I start a RStudio Server with in cluster init script as described here in a Unity Catalog Cluster the sparklyr connection fails with an error about a missing Credential Scope.=LI tried it both in 11.3LTS and 12.0 Beta. I tried it only in a Persona...

Data Engineering

5669 Views
5 replies
6 kudos

12-13-2022 8:48:53 AM

View Replies

Latest Reply

kunalmishra9
New Contributor III

09-26-2023 4:42:07 PM

6 kudos

Have run into this issue as well. Let me know if there was any resolution

6 kudos

09-26-2023 4:42:07 PM

4 More Replies

User

Count

1609

751

349

285

248

Databricks Community

Forum Posts

DLT Pipeline Out Of Memory Errors

How to stop SparkSession within notebook without errr

Referring to Azure Keyvault secrets in spark config

Bamboolib error

Create cluster error: Backend service unavailable

How does Job Cluster knows how many resources to assign to an Application ?

Databricks Job scheduling - continuous mode

Resolved! How to set a variable and use it in a SQL query

Streaming Data Modeling Normalization with Databricks Delta Live Tables

Invisible empty spaces when reading .csv files

Resolved! Problems connecting Simba ODBC with a M1 Macbook Pro

com.databricks.spark.safespark.UDFException: UNAVAILABLE: Channel shutdownNow invoked

DLT current channel uses same runtime as the preview channel

Correlated column is not allowed in non predicate in UDF SQL

Bug? Unity Catalog incompatible with Sparklyr in RStudio (on Driver) and as well if used on one cluster from multiple notebooks?

Connect with Databricks Users in Your Area

Load parent columns and not unnest using pyspark? ...

How to increase executor memory in Databricks jobs

Can I have additional logic in a DLT notebook that...

Databricks-connect Configure a connection to serve...

Import from repo