Data Engineering

Forum Posts

Sorted by:

by User16826990884 • Databricks Employee

06-25-2021 11:40:31 AM

19798 Views
1 replies
1 kudos

Resolved! Views vs Materialized Delta Tables

Is there general guidance around using views vs creating Delta tables? For example, I need to do some filtering and make small tweaks to a few columns for use in another application. Is there a downside of using a view here?

Data Engineering

19798 Views
1 replies
1 kudos

06-25-2021 11:40:31 AM

View Replies

Latest Reply

User16826990884
Databricks Employee

06-25-2021 12:18:16 PM

1 kudos

Views won't duplicate the data so if you are just filtering columns or rows or making small tweaks then views might be a good option. Unless, of course, the filtering is really expensive or you are doing a lot of calculations, then materialize the vi...

1 kudos

06-25-2021 12:18:16 PM

by Srikanth_Gupta_ • Databricks Employee

06-25-2021 7:33:46 AM

2561 Views
1 replies
1 kudos

What is the difference between Databricks secret scopes vs AWS secret manager vs Azure key vault, in which scenarios I should go for secret scopes

Data Engineering

2561 Views
1 replies
1 kudos

06-25-2021 7:33:46 AM

View Replies

Latest Reply

sajith_appukutt
Databricks Employee

06-25-2021 12:15:32 PM

1 kudos

All three options are secure ways to store secrets. Databricks secrets has the additional functionality of redaction , so it is convenient sometimes. Also in azure, you have the ability to use azure KV as the backend for secrets.

1 kudos

06-25-2021 12:15:32 PM

by brickster_2018 • Databricks Employee

06-25-2021 12:13:22 PM

4579 Views
1 replies
1 kudos

Resolved! Classpath issues when running spark-submit

How to identify the jars used to load a particular class. I am sure I packed the classes correctly in my application jar. However, looks like the class is loaded from a different jar. I want to understand the details so that I can ensure to use the r...

Data Engineering

4579 Views
1 replies
1 kudos

06-25-2021 12:13:22 PM

View Replies

Latest Reply

brickster_2018
Databricks Employee

06-25-2021 12:14:49 PM

1 kudos

Adding the below configurations at the cluster level can help to print more logs to identify the jars from which the class is loaded. spark.executor.extraJavaOptions=-verbose:class spark.driver.extraJavaOptions=-verbose:class

1 kudos

06-25-2021 12:14:49 PM

by brickster_2018 • Databricks Employee

06-25-2021 12:08:37 PM

2125 Views
1 replies
0 kudos

Resolved! Cannot upload libraries on UI

When trying to upload libraries on UI it fails.

Data Engineering

2125 Views
1 replies
0 kudos

06-25-2021 12:08:37 PM

View Replies

Latest Reply

brickster_2018
Databricks Employee

06-25-2021 12:10:23 PM

0 kudos

One corner case scenario where we can hit this issue is if there is /<shard name>/0/Filestore/jars file in the root bucket of the workspace. Once you remove the file, the upload should work fine.

0 kudos

06-25-2021 12:10:23 PM

by User16783853906 • Databricks Employee

06-07-2021 12:05:03 PM

4591 Views
3 replies
0 kudos

Resolved! Frequent spot loss of driver nodes resulting in failed jobs when using spot fleet pools

When using spot fleet pools to schedule jobs, driver and worker nodes are provisioned from the spot pools and we are noticing jobs failing with the below exception when there is a driver spot loss. Share best practices around using fleet pools with 1...

Data Engineering

4591 Views
3 replies
0 kudos

06-07-2021 12:05:03 PM

View Replies

Latest Reply

User16783853906
Databricks Employee

06-23-2021 2:20:55 PM

0 kudos

In this scenario, the driver node is reclaimed by AWS. Databricks started preview of hybrid pools feature which would allow you to provision driver node from a different pool. We recommend using on-demand pool for driver node to improve reliability i...

0 kudos

06-23-2021 2:20:55 PM

2 More Replies

by brickster_2018 • Databricks Employee

06-25-2021 12:06:06 PM

2808 Views
1 replies
1 kudos

Resolved! Databricks Vs Yarn - Resource Utilization

I have a spark-submit application that worked fine with 8GB executor memory in yarn. I am testing the same job against the Databricks cluster with the same executor memory. However, the jobs are running slower in Databricks.

Data Engineering

2808 Views
1 replies
1 kudos

06-25-2021 12:06:06 PM

View Replies

Latest Reply

brickster_2018
Databricks Employee

06-25-2021 12:06:46 PM

1 kudos

This is not an Apple to Apple comparison. When you set 8GB as the executor memory in Yarn, then the container that is launched to run the executor JVM is getting 8GB of memory. Accordingly, the Xmx value of the heap is calculated. In Databricks, when...

1 kudos

06-25-2021 12:06:46 PM

by Anonymous • Not applicable

06-24-2021 10:06:17 PM

3675 Views
2 replies
0 kudos

How can I disable downloading files such as csv files?

Data Engineering

3675 Views
2 replies
0 kudos

06-24-2021 10:06:17 PM

View Replies

Latest Reply

sajith_appukutt
Databricks Employee

06-25-2021 12:05:56 PM

0 kudos

You can disable download button for notebook results which exports results as csv from admin console - > workspace settings -> advanced section

0 kudos

06-25-2021 12:05:56 PM

1 More Replies

by Srikanth_Gupta_ • Databricks Employee

06-16-2021 10:09:43 AM

3076 Views
2 replies
0 kudos

Can we create pools to reduce cluster start time in Databricks

Data Engineering

3076 Views
2 replies
0 kudos

06-16-2021 10:09:43 AM

View Replies

Latest Reply

User16783853906
Databricks Employee

06-25-2021 12:05:44 PM

0 kudos

When a cluster is attached to a pool, cluster nodes are created using the the pool’s idle instances which help to reduce cluster start and auto-scaling times .If you are using pools and looking to reduce start time for all scenarios, then you should ...

0 kudos

06-25-2021 12:05:44 PM

1 More Replies

by Anonymous • Not applicable

06-24-2021 10:12:44 PM

2128 Views
2 replies
0 kudos

Is Autoloader an option to load data to Databricks from Azure SQL?

Data Engineering

2128 Views
2 replies
0 kudos

06-24-2021 10:12:44 PM

View Replies

Latest Reply

sajith_appukutt
Databricks Employee

06-25-2021 12:01:55 PM

0 kudos

If you are looking for incrementally loading data from Azure SQL, checkout one of our technology partners that support change-data-capture or setup debezium for sql-server. These solutions could land data in a streaming fashion to kafka/kinesis/even...

0 kudos

06-25-2021 12:01:55 PM

1 More Replies

by brickster_2018 • Databricks Employee

06-25-2021 11:56:55 AM

2711 Views
1 replies
0 kudos

Resolved! Can I use OSS Spark History Server to view the EventLogs

Is it possible to run the OSS SPark history server and view the spark event logs.

Data Engineering

2711 Views
1 replies
0 kudos

06-25-2021 11:56:55 AM

View Replies

Latest Reply

brickster_2018
Databricks Employee

06-25-2021 11:58:12 AM

0 kudos

Yes, it's possible. The OSS Spark history server can read the Spark event logs generated on a Databricks cluster. Using Cluster log delivery, the SPark logs can be written to any arbitrary location. Event logs can be copied from there to the storage ...

0 kudos

06-25-2021 11:58:12 AM

by User16826990884 • Databricks Employee

06-25-2021 11:56:21 AM

1179 Views
0 replies
0 kudos

Encrypt root S3 bucket

This is a 2-part question:How do I go about encrypting an existing root S3 bucket?Will this impact my Databricks environment? (Resources not being accessible, performance issues etc.)

Data Engineering

1179 Views
0 replies
0 kudos

06-25-2021 11:56:21 AM

by brickster_2018 • Databricks Employee

06-25-2021 11:55:32 AM

3873 Views
1 replies
0 kudos

Resolved! Jobs running forever in Spark UI

On the Spark UI, Jobs are running forever. But my notebook already completed the operations. Why the resources are wasted

Data Engineering

3873 Views
1 replies
0 kudos

06-25-2021 11:55:32 AM

View Replies

Latest Reply

brickster_2018
Databricks Employee

06-25-2021 11:55:51 AM

0 kudos

This happens if the Spark driver is missing events. The jobs/task are not running. The Spark UI is reporting incorrect stats. This can be treated as a harmless UI issue. If you continue to see the issue consistently, then it might be good to review w...

0 kudos

06-25-2021 11:55:51 AM

by User16826990884 • Databricks Employee

06-25-2021 11:44:58 AM

977 Views
0 replies
1 kudos

Question on running optimize on a Delta table streaming pipeline

I have a Bronze -> Silver -> Gold architecture for my ETL pipelines and all tables are Delta. I'm trying to understand what updates flow downstream when I make changes to the source table. Most importantly, if I run optimize on the source, does every...

Data Engineering

977 Views
0 replies
1 kudos

06-25-2021 11:44:58 AM

by brickster_2018 • Databricks Employee

06-25-2021 11:18:08 AM

1860 Views
1 replies
0 kudos

Resolved! What are the general performance optimization tips for improving MERGE performance

Data Engineering

1860 Views
1 replies
0 kudos

06-25-2021 11:18:08 AM

View Replies

Latest Reply

brickster_2018
Databricks Employee

06-25-2021 11:19:34 AM

0 kudos

While using MERGE INTO statement, if the source data that will be merged into the target delta table is small enough to be fit into memory of the worker nodes, then it makes sense to broadcast the source data. By doing so, the execution can avoid the...

0 kudos

06-25-2021 11:19:34 AM

by brickster_2018 • Databricks Employee

06-25-2021 11:14:02 AM

4693 Views
1 replies
0 kudos

Resolved! Can Spark JDBC create duplicate records

Is it transaction safe?Does it ensure atomicity

Data Engineering

4693 Views
1 replies
0 kudos

06-25-2021 11:14:02 AM

View Replies

Latest Reply

brickster_2018
Databricks Employee

06-25-2021 11:17:06 AM

0 kudos

Atomicity is ensured at a task level and not at a stage level. For any reason, if the stage is getting retried, the tasks which already completed the write operation will re-run and cause duplicate records. This is expected by design. When Apache Spa...

0 kudos

06-25-2021 11:17:06 AM

Databricks Community

Forum Posts

Resolved! Views vs Materialized Delta Tables

What is the difference between Databricks secret scopes vs AWS secret manager vs Azure key vault, in which scenarios I should go for secret scopes

Resolved! Classpath issues when running spark-submit

Resolved! Cannot upload libraries on UI

Resolved! Frequent spot loss of driver nodes resulting in failed jobs when using spot fleet pools

Resolved! Databricks Vs Yarn - Resource Utilization

How can I disable downloading files such as csv files?

Can we create pools to reduce cluster start time in Databricks

Is Autoloader an option to load data to Databricks from Azure SQL?

Resolved! Can I use OSS Spark History Server to view the EventLogs

Encrypt root S3 bucket

Resolved! Jobs running forever in Spark UI

Question on running optimize on a Delta table streaming pipeline

Resolved! What are the general performance optimization tips for improving MERGE performance

Resolved! Can Spark JDBC create duplicate records

Join Us as a Local Community Builder!

What's the difference between dbmanagedidentity an...

How to stop Databricks retaining widget selection ...

Writing to Foreign catalog

Using Init Scipt to execute python notebook at all...

Issues Creating Genie Space via API Join Specs Are...