Data Engineering

Forum Posts

Sorted by:

by Tim_T • New Contributor

04-14-2023 6:53:49 AM

1676 Views
0 replies
0 kudos

Are training/ecommerce data tables available as CSVs?

The course "Apache Spark™ Programming with Databricks" requires data sources such as training/ecommerce/events/events.parquet. Are these available as CSV files? My company's databricks configuration does not allow me to mount to such repositories, bu...

Data Engineering

1676 Views
0 replies
0 kudos

04-14-2023 6:53:49 AM

by Hitesh_goswami • New Contributor

04-12-2023 9:25:35 AM

1995 Views
1 replies
0 kudos

Upgrading Ipython version without changing LTS version

I am using a specific Pydeeque function called ColumnProfilerRunner which is only supported with Spark 3.0.1, so I must use 7.3 LTS. Currently, I am trying to install "great_expectations" library on Python, which requires Ipython version==7.16.3, an...

Data Engineering

1995 Views
1 replies
0 kudos

04-12-2023 9:25:35 AM

View Replies

Latest Reply

Anonymous
Not applicable

04-14-2023 2:36:43 AM

0 kudos

@Hitesh Goswami : please check if the below helps!To upgrade the Ipython version on a Databricks 7.3LTS cluster, you can follow these steps:Create a new library installation command using the Databricks CLI by running the following command in your l...

0 kudos

04-14-2023 2:36:43 AM

by JGil • New Contributor III

04-11-2023 11:35:16 PM

5021 Views
5 replies
0 kudos

Installing Bazel on databricks cluster

I am new to azure databricks and I want to install a library on a cluster and to do that I need to install bazel build tool first.I checked the site bazel but I am still not sure how to do it in databricks?I appriciate if any can help me and write me...

Data Engineering

5021 Views
5 replies
0 kudos

04-11-2023 11:35:16 PM

View Replies

Latest Reply

Avinash_94
Databricks Employee

04-14-2023 12:25:06 AM

0 kudos

Databricks migrated over from the standard Scala Build Tool (SBT) to using Bazel to build, test and deploy our Scala code. Follow this doc https://www.databricks.com/blog/2019/02/27/speedy-scala-builds-with-bazel-at-databricks.html

0 kudos

04-14-2023 12:25:06 AM

4 More Replies

by afzi • New Contributor II

08-10-2022 10:40:47 PM

4500 Views
1 replies
1 kudos

Pandas DataFrame error when using to_csv

Hi Everyone, I would like to a Pandas Dataframe to /dbfs/FileStore/ using to_csv method.Usually it would just write the Dataframe to the path described but It has been giving me "FileNotFoundError: [Errno 2] No such file or directory: '/dbfs/FileStor...

Data Engineering

4500 Views
1 replies
1 kudos

08-10-2022 10:40:47 PM

View Replies

Latest Reply

Avinash_94
Databricks Employee

04-14-2023 12:31:19 AM

1 kudos

f = open("/dbfs/mnt/blob/myNames.txt", "r")

1 kudos

04-14-2023 12:31:19 AM

by User16826992783 • Databricks Employee

06-11-2021 7:53:45 AM

1989 Views
1 replies
0 kudos

Why are some of my AWS EBS volumes in my workspace unencrypted?

I noticed that 30GB of my EBS volumes are unencrypted, is there a reason for this, and is there a way to encrypt these volumes?

Data Engineering

1989 Views
1 replies
0 kudos

06-11-2021 7:53:45 AM

View Replies

Latest Reply

Abishek
Databricks Employee

04-14-2023 12:30:26 AM

0 kudos

https://docs.databricks.com/security/keys/customer-managed-keys-storage-aws.html#introductionThe Databricks cluster’s EBS volumes (optional) - For Databricks Runtime cluster nodes and other compute resources in the Classic data plane, you can option...

0 kudos

04-14-2023 12:30:26 AM

by wb • New Contributor II

09-17-2022 2:53:39 PM

1798 Views
1 replies
2 kudos

Import paths using repos and installed libraries get confused

We use Azure Devops and Azure Databricks and have custom Python libraries. I placed my notebooks in the same repo and the structure is like this:mylib/ mylib/__init__.pyt mylib/code.py notebooks/ notebooks/job_notebook.py setup.pyAzure pipelines buil...

Data Engineering

1798 Views
1 replies
2 kudos

09-17-2022 2:53:39 PM

View Replies

Latest Reply

Avinash_94
Databricks Employee

04-14-2023 12:28:50 AM

2 kudos

It looks for the configs locally i suppose if you can share requirements .txt i can elaborate

2 kudos

04-14-2023 12:28:50 AM

by User16826990884 • Databricks Employee

06-25-2021 11:46:25 AM

3299 Views
1 replies
1 kudos

Delta log retention

Is there an impact on performance if I increase the Delta log retention to 3000?

Data Engineering

3299 Views
1 replies
1 kudos

06-25-2021 11:46:25 AM

View Replies

Latest Reply

DD_Sharma
Databricks Employee

04-14-2023 12:28:43 AM

1 kudos

There will be no performance impact if you want to keep " Delta log retention to 3000". However, it will increase the storage cost so it's not advisable to use a large number until really needed for the business use cases.The default delta.logRetenti...

1 kudos

04-14-2023 12:28:43 AM

by Saurabh98290 • New Contributor II

09-15-2022 8:24:54 PM

1565 Views
1 replies
2 kudos

Best Suited Language To Parallelize Notebook

I would like to know if we are writing code for parallel execution on notebook which language is best suited for that Python or Scala.

Data Engineering

1565 Views
1 replies
2 kudos

09-15-2022 8:24:54 PM

View Replies

Latest Reply

User16756723392
Databricks Employee

04-14-2023 12:28:03 AM

2 kudos

You need to test in Python and scala based on the complexity one of it outperforms the other. In few cases Python was faster where as in other Scala. It is all about the efficiency of the code

2 kudos

04-14-2023 12:28:03 AM

by NimaiAhl • New Contributor II

09-08-2022 10:24:11 PM

1872 Views
1 replies
0 kudos

External Tables - SQL

To create external tables we need to use the location keyword and use the link for the storage location, in reference to that does the user need to have permission for the storage location if not then will we use storage credentials to provide the ac...

Data Engineering

1872 Views
1 replies
0 kudos

09-08-2022 10:24:11 PM

View Replies

Latest Reply

Anu-sha
Databricks Employee

04-14-2023 12:23:22 AM

0 kudos

Hi Nimai, That's partially right. You can grant permissions directly on the storage credential, but Databricks recommends that you reference it in an external location and grant permissions to that instead. An external location combines a storage cre...

0 kudos

04-14-2023 12:23:22 AM

by Kotofosonline • New Contributor III

09-08-2021 4:41:09 AM

3435 Views
1 replies
2 kudos

Query with distinct sort and alias produces error column not found

I’m trying to use sql query on azure-databricks with distinct sort and aliasesSELECT DISTINCT album.ArtistId AS my_alias FROM album ORDER BY album.ArtistIdThe problem is that if I add an alias then I can not use not aliased name in the order by cla...

Data Engineering

3435 Views
1 replies
2 kudos

09-08-2021 4:41:09 AM

View Replies

Latest Reply

User16756723392
Databricks Employee

04-14-2023 12:22:33 AM

2 kudos

SELECT album.ArtistId ,DISTINCT album.ArtistId AS my_alias FROM album ORDER BY album.ArtistIdCan you try this

2 kudos

04-14-2023 12:22:33 AM

by UmaMahesh1 • Honored Contributor III

04-11-2023 7:01:42 AM

3785 Views
1 replies
2 kudos

Checkpoint issue when loading data from confluent kafka

I have a streaming notebook which fetches messages from confluent Kafka topic and loads them into adls. It is a streaming notebook with the trigger as continuous processing. Before loading the message (which is in Avro format), I'm flattening out the...

Data Engineering

3785 Views
1 replies
2 kudos

04-11-2023 7:01:42 AM

View Replies

Latest Reply

Avinash_94
Databricks Employee

04-14-2023 12:21:44 AM

2 kudos

Best approach is to not to depend on Kafka’s commit mechanism! We can store processing result and message offset to external data store in the same database transaction. So, if the database transaction fails, both commit and processing will fail and ...

2 kudos

04-14-2023 12:21:44 AM

by Himanshu1 • New Contributor II

08-21-2022 11:05:34 PM

3574 Views
1 replies
3 kudos

How to read XML files in delta live tables?

Even after maven library installation using the Auto installation.spark.read.option("rowTag", "tag").xml("dbfs:/mnt/dev/bronze/xml/fileName.xml")not working.

Data Engineering

3574 Views
1 replies
3 kudos

08-21-2022 11:05:34 PM

View Replies

Latest Reply

DD_Sharma
Databricks Employee

04-14-2023 12:20:03 AM

3 kudos

At present DLT does not support installing the maven library from the DLT pipeline. In the future this feature will come for sure so please wait for some time and keep checking data bricks runtime release docs https://docs.databricks.com/release-note...

3 kudos

04-14-2023 12:20:03 AM

by samruddhi • New Contributor

06-29-2022 7:46:08 AM

2532 Views
1 replies
0 kudos

Issue while creating Workspace in databricks using AWS

I am trying to configure databricks with AWS, I have configured the cloud resources as described in this https://docs.databricks.com/administration-guide/account-api/iam-role.html#language-Databricks%C2%A0VPC I have selected "Your VPC Default" as the...

Data Engineering

2532 Views
1 replies
0 kudos

06-29-2022 7:46:08 AM

View Replies

Latest Reply

Abishek
Databricks Employee

04-14-2023 12:19:39 AM

0 kudos

@samruddhi ChitnisCan you please check the below troubleshooting guide : Credentials configuration error messages: Malformed request: Failed credential configuration validation checksThe list of permissions checks in the error message indicate the li...

0 kudos

04-14-2023 12:19:39 AM

by sajith_appukutt • Databricks Employee

06-09-2021 12:36:37 AM

3498 Views
2 replies
1 kudos

Resolved! How can I configure S3 Client-Side Encryption (CSE-KMS ) for my data pipeline

Data Engineering

3498 Views
2 replies
1 kudos

06-09-2021 12:36:37 AM

View Replies

Latest Reply

AdrianRojas
New Contributor II

04-13-2023 4:35:26 PM

1 kudos

a bit old, but I just faced the same issue, specifying a custom EncryptionMaterialsProvider (as described in the previous post) did the trick for me but I did had to also specify my kms endpoint, just because my region:"fs.s3.cse.kms.endpoint" -> "km...

1 kudos

04-13-2023 4:35:26 PM

1 More Replies

by Samit110978 • New Contributor II

12-16-2022 7:38:15 AM

3518 Views
3 replies
1 kudos

Passing Parameter from SSRS to Databricks user defined function

I am trying to pass parameter from SSRS to User Defined Function in Databricks which in turn will return table that will be shown as output in report.I tried below calling function from SSRS, but it looks like parameter value is not passed. I have di...

Data Engineering

3518 Views
3 replies
1 kudos

12-16-2022 7:38:15 AM

View Replies

Latest Reply

Aviral-Bhardwaj
Esteemed Contributor III

12-17-2022 10:48:32 PM

1 kudos

can you share full code and dataset by that we can also debug this

1 kudos

12-17-2022 10:48:32 PM

2 More Replies

Databricks Community

Forum Posts

Are training/ecommerce data tables available as CSVs?

Upgrading Ipython version without changing LTS version

Installing Bazel on databricks cluster

Pandas DataFrame error when using to_csv

Why are some of my AWS EBS volumes in my workspace unencrypted?

Import paths using repos and installed libraries get confused

Delta log retention

Best Suited Language To Parallelize Notebook

External Tables - SQL

Query with distinct sort and alias produces error column not found

Checkpoint issue when loading data from confluent kafka

How to read XML files in delta live tables?

Issue while creating Workspace in databricks using AWS

Resolved! How can I configure S3 Client-Side Encryption (CSE-KMS ) for my data pipeline

Passing Parameter from SSRS to Databricks user defined function

File Arrival Trigger - Multiple tables

Issue while handling Deletes and Inserts in Struct...

DLT with CDC and schema changes in streaming pipel...

how to update not tracked column only in new row v...

Databricks Cost Estimation Template