Data Engineering

Forum Posts

Sorted by:

Start a conversation

by NathanSundarara • Contributor

05-12-2023 6:36:41 PM

1490 Views
4 replies
1 kudos

sample

Help parsing the JSON using Spark SQL or python. Sample json attached.

Data Engineering

1490 Views
4 replies
1 kudos

05-12-2023 6:36:41 PM

View Replies

Latest Reply

NathanSundarara
Contributor

05-22-2023 5:45:16 AM

1 kudos

@Suteja Kanuri can you please respond to my question above?

1 kudos

05-22-2023 5:45:16 AM

3 More Replies

by _deepak_ • New Contributor II

05-09-2023 4:10:25 AM

1040 Views
3 replies
0 kudos

Databricks regression test suite

Hi, I am new to Databricks and setting up the non-prod environment. I am wanted to know, IS there any way by which I can run a regression suite so that existing setup should not break in case of any feature addition and also how can I make available ...

Data Engineering

1040 Views
3 replies
0 kudos

05-09-2023 4:10:25 AM

View Replies

Latest Reply

Anonymous
Not applicable

05-13-2023 8:38:50 AM

0 kudos

@deepak prasad :Yes, you can run regression tests to ensure that your changes do not break existing functionality. Databricks supports a number of testing frameworks like PyTest, which can be used to automate regression testing. You can write test c...

0 kudos

05-13-2023 8:38:50 AM

2 More Replies

by santhosh1 • New Contributor II

10-10-2022 11:57:11 PM

1095 Views
4 replies
3 kudos

Can we share exam voucher to another databricks account

Hi, I received free voucher for lakehouse webinar, My friend also got free voucher, by any chance can i use my friend voucher to shedule another exam for me.

Data Engineering

1095 Views
4 replies
3 kudos

10-10-2022 11:57:11 PM

View Replies

Latest Reply

SUMI1
New Contributor III

06-09-2023 2:34:09 AM

3 kudos

Hi guysUnfortunately, it is not possible to share an exam voucher with another Databricks account. Exam vouchers are typically tied to specific accounts or individuals and cannot be transferred or shared. Free Fire

3 kudos

06-09-2023 2:34:09 AM

3 More Replies

by tototox • New Contributor III

05-11-2023 7:08:57 AM

4821 Views
3 replies
2 kudos

how to check table size by partition?

I want to check the size of the delta table by partition.As you can see, only the size of the table can be checked, but not by partition.

Data Engineering

4821 Views
3 replies
2 kudos

05-11-2023 7:08:57 AM

View Replies

Latest Reply

Anonymous
Not applicable

05-13-2023 8:57:50 AM

2 kudos

@jin park :You can use the Databricks Delta Lake SHOW TABLE EXTENDED command to get the size of each partition of the table. Here's an example:%sql SHOW TABLE EXTENDED LIKE '<table_name>' PARTITION (<partition_column> = '<partition_value>') SELECT...

2 kudos

05-13-2023 8:57:50 AM

2 More Replies

by Yash_542965 • New Contributor II

05-16-2023 9:11:19 AM

2094 Views
2 replies
3 kudos

Resolved! Access Excel file in delta live pipeline

I'm having an issue accessing the excel through dlt pipeline. the file is in ADLS I'm using pandas to read the Excel. It seems pandas are not able to understand abfss protocol is there any way to read Excel with pandas in dlt pipeline?I'm getting thi...

Data Engineering

2094 Views
2 replies
3 kudos

05-16-2023 9:11:19 AM

View Replies

Latest Reply

Yash_542965
New Contributor II

06-09-2023 12:16:13 AM

3 kudos

Thanks for the info. It works just need to install an additional library using "%pip install openpyxl".

3 kudos

06-09-2023 12:16:13 AM

1 More Replies

by Inna_M • New Contributor III

06-07-2023 11:04:26 AM

944 Views
1 replies
1 kudos

Resolved! Is there any maintenance (patches , upgrade for VMs created by DataBricks on Azure) from DataBricks

We are using Databricks on Azure. Infra team noticed we have some VMs created in the past for DataBricks clusters on version Linux (ubuntu 18.04). Is there maintenance previewed for that, upgrade? Are there any patches for created in Azure objects by...

Data Engineering

944 Views
1 replies
1 kudos

06-07-2023 11:04:26 AM

View Replies

Latest Reply

Inna_M
New Contributor III

06-08-2023 7:49:55 AM

1 kudos

Finally while I was posting this question, AzureDataBricks upgraded VMs to the supported version 20, not the latest , 22. It was a week after old version was no longer supported by Microsoft

1 kudos

06-08-2023 7:49:55 AM

by CoopCoop • New Contributor III

05-15-2023 9:15:58 AM

2471 Views
6 replies
7 kudos

Resolved! PDF Attachment on an Alert

Currently my Alert is an HTML table using data pointing to an SQL query.I was wondering if it is possible to attach the resulting table from this SQL query as a PDF to the alert email.If anyone has successfully implemented this, please let me know! T...

Data Engineering

2471 Views
6 replies
7 kudos

05-15-2023 9:15:58 AM

View Replies

Latest Reply

Atanu
Esteemed Contributor

06-08-2023 7:03:57 AM

7 kudos

Ok understood the concern, so basically the issue is with PDF rendering as much I understood. Let me know if I am wrong. Let me see if there is any improvement by our engineering team on this front.

7 kudos

06-08-2023 7:03:57 AM

5 More Replies

by Louis_Databrick • New Contributor II

05-31-2023 12:12:38 AM

611 Views
2 replies
0 kudos

Registering a dataframe coming from a CDC data stream removes the CDC columns from the resulting temporary view, even when explicitly adding a copy of the column to the dataframe.

df_source_records.filter(F.col("_change_type").isin("delete", "insert", "update_postimage")) .withColumn("ROW_NUMBER", F.row_number().over(window)) .filter("ROW_NUMBE...

Data Engineering

611 Views
2 replies
0 kudos

05-31-2023 12:12:38 AM

View Replies

Latest Reply

Louis_Databrick
New Contributor II

06-08-2023 4:15:24 AM

0 kudos

Seems to work now actually. No idea what changed, as I tried multiple times exactly in this way and it did.not.work.from pyspark.sql.functions import expr from pyspark.sql.utils import AnalysisException import pyspark.sql.functions as f data = [(...

0 kudos

06-08-2023 4:15:24 AM

1 More Replies

by StuartKindness_ • New Contributor II

05-05-2023 10:20:58 AM

899 Views
4 replies
2 kudos

How to replace the SSO certifcate on our workspace?

We have Azure AD SSO setup on our workspace but the three year certificate is due to expire on Monday. I have logged onto the Admin Console & Single Sign-on tab. All the options are greyed out and there is no update or edit buttons as can be seen in ...

Data Engineering

899 Views
4 replies
2 kudos

05-05-2023 10:20:58 AM

View Replies

Latest Reply

StuartKindness_
New Contributor II

05-08-2023 4:14:28 AM

2 kudos

@Debayan our version is branch-3.96-1682169174-f2e2f130 if this helps any?

2 kudos

05-08-2023 4:14:28 AM

3 More Replies

by harraz • New Contributor III

05-31-2023 3:50:32 PM

2246 Views
1 replies
0 kudos

Run result unavailable: run failed with error message Notebook not found:

I'm trying to create a workflow job that fetches the notebook from a remote git repository (Bitbucket cloud)I tried everything in the Path field and nothing is working. Note that the bitbucket repo is connected to databricks already and no issues che...

Data Engineering

2246 Views
1 replies
0 kudos

05-31-2023 3:50:32 PM

View Replies

Latest Reply

Debayan
Esteemed Contributor III

06-07-2023 11:39:42 PM

0 kudos

Hi @harraz (Customer) , Could you please confirm if files in repos has been enabled? https://docs.databricks.com/files/workspace.html#configure-support-for-files-in-repos.You can use the command %sh pwd in a notebook inside a repo to check if Files ...

0 kudos

06-07-2023 11:39:42 PM

by harraz • New Contributor III

05-31-2023 3:53:30 PM

738 Views
2 replies
0 kudos

how to setup the path to a remote notebook in bitbucket to run as a jobI tried everything in the path and nothing is workingI keep getting this error:...

how to setup the path to a remote notebook in bitbucket to run as a jobI tried everything in the path and nothing is workingI keep getting this error:Run result unavailable: run failed with error message Notebook not found:Note that I already connec...

Data Engineering

738 Views
2 replies
0 kudos

05-31-2023 3:53:30 PM

View Replies

Latest Reply

Debayan
Esteemed Contributor III

06-07-2023 11:39:18 PM

0 kudos

Hi @mohamed harraz , Could you please confirm if files in repos has been enabled? https://docs.databricks.com/files/workspace.html#configure-support-for-files-in-repos.You can use the command %sh pwd in a notebook inside a repo to check if Files in...

0 kudos

06-07-2023 11:39:18 PM

1 More Replies

by cmilligan • Contributor II

12-07-2022 9:40:01 AM

459 Views
1 replies
1 kudos

Return notebook path from job that is run remotely from the repo

I'm wanting to set up some email alerts for issues in the data as a part of a job run. I am wanting to point the user to the notebook that the issue occurred in. I think this would be simple enough but another layer is that the job is going to be run...

Data Engineering

459 Views
1 replies
1 kudos

12-07-2022 9:40:01 AM

View Replies

Latest Reply

Debayan
Esteemed Contributor III

06-07-2023 11:33:12 PM

1 kudos

Hi, Could you please clarify what do you mean by return the file from the remote repo?Please tag @Debayan with your next response which will notify me, Thank you!

1 kudos

06-07-2023 11:33:12 PM

by nistrate • New Contributor II

05-15-2023 11:40:32 AM

1327 Views
1 replies
2 kudos

Resolved! Restricting Workflow Creation and Implementing Approval Mechanism in Databricks

Hello Databricks Community,I am seeking assistance understanding the possibility and procedure of implementing a workflow restriction mechanism in Databricks. Our aim is to promote a better workflow management and ensure the quality of the notebooks ...

Data Engineering

1327 Views
1 replies
2 kudos

05-15-2023 11:40:32 AM

View Replies

Latest Reply

User16502773013
New Contributor III

06-07-2023 7:13:42 PM

2 kudos

Hello Nistrate,If I understand the question correctly, the ask is to create an approval framework/workflow for workflows/jobs changes/commits, I don't believe this is currently supported however this can be supported through the use of source control...

2 kudos

06-07-2023 7:13:42 PM

by etsyal1e2r3 • Honored Contributor

06-03-2023 5:00:16 PM

3453 Views
2 replies
3 kudos

Resolved! Compiling Flattened Dataframe back to Struct Columns

I have a dataframe with this format of columns:[`first.second.third` , `alpha.bravo.test1` , `alpha.bravo.test2`]I'd like to get an output dataframe of this:[ `first` | `alpha` ] ---------------...

Data Engineering

3453 Views
2 replies
3 kudos

06-03-2023 5:00:16 PM

View Replies

Latest Reply

etsyal1e2r3
Honored Contributor

06-05-2023 9:43:01 PM

3 kudos

I have figured out the solution.

3 kudos

06-05-2023 9:43:01 PM

1 More Replies

by Pri • New Contributor

06-07-2023 12:28:00 PM

277 Views
0 replies
0 kudos

Hi, Cat! I’m applying for a position at Databricks and was hoping to get some current Brickster insights. I’ve been wanting to join the company for a ...

Hi, Cat! I’m applying for a position at Databricks and was hoping to get some current Brickster insights. I’ve been wanting to join the company for a while!! Thanks in advance

Data Engineering

277 Views
0 replies
0 kudos

06-07-2023 12:28:00 PM

User

Count

1602

736

344

284

247

Databricks

Forum Posts

sample

Databricks regression test suite

Can we share exam voucher to another databricks account

how to check table size by partition?

Resolved! Access Excel file in delta live pipeline

Resolved! Is there any maintenance (patches , upgrade for VMs created by DataBricks on Azure) from DataBricks

Resolved! PDF Attachment on an Alert

Registering a dataframe coming from a CDC data stream removes the CDC columns from the resulting temporary view, even when explicitly adding a copy of the column to the dataframe.

How to replace the SSO certifcate on our workspace?

Run result unavailable: run failed with error message Notebook not found:

how to setup the path to a remote notebook in bitbucket to run as a jobI tried everything in the path and nothing is workingI keep getting this error:...

Return notebook path from job that is run remotely from the repo

Resolved! Restricting Workflow Creation and Implementing Approval Mechanism in Databricks

Resolved! Compiling Flattened Dataframe back to Struct Columns

Hi, Cat! I’m applying for a position at Databricks and was hoping to get some current Brickster insights. I’ve been wanting to join the company for a ...

Best way to parse Google Analytics data in Databri...

DELTA_EXCEED_CHAR_VARCHAR_LIMIT

Not able to set run_as service_principal_name

Pyspark operations slowness in CLuster 14.3LTS as ...

[Databricks Assets Bundles] Workflow trigger on fi...