cancel
Showing results for 
Search instead for 
Did you mean: 
Data Engineering
Join discussions on data engineering best practices, architectures, and optimization strategies within the Databricks Community. Exchange insights and solutions with fellow data engineers.
cancel
Showing results for 
Search instead for 
Did you mean: 

Forum Posts

jonathan-dufaul
by Valued Contributor
  • 5356 Views
  • 2 replies
  • 0 kudos

Is there a command in sql cell to ignore formatting for some lines like `# fmt: off` in Python cells

In python cells I can add the comments `# fmt: off` before a block of code that I want black/autoformatter to ignore and `# fmt: on` afterwards. Is there anything similar I can put in sql cells to accomplish the same effect?Some of the recommendation...

Data Engineering
autoformatter
formatter
sql
  • 5356 Views
  • 2 replies
  • 0 kudos
bayerb
by New Contributor
  • 2711 Views
  • 1 replies
  • 0 kudos

Sink is not written into delta table in Spark structured streaming

I want to create a streaming job, that reads messages from a folder within TXT files, does the parsing, some processing, and appends the result into one of 3 possible delta tables depending on the parse result. There is a parse_failed table, an unknw...

  • 2711 Views
  • 1 replies
  • 0 kudos
Latest Reply
Lakshay
Databricks Employee
  • 0 kudos

There doesn't seem to any issue with code. But log needs to be analysed to get a clue of what is the issue. Could you please create a support ticket.

  • 0 kudos
vishwanath_1
by New Contributor III
  • 2123 Views
  • 1 replies
  • 0 kudos

Resolved! Need Suggestion for better caching strategy

i have below steps to perform 1.Read a csv file (considerably huge file .. ~100gb)2.add index using zipwithindex function 3.repartition dataframe 4.Passing on to another function .Can you suggest the best optimized caching strategy to execute these c...

vishwanath_1_0-1705915220664.png
  • 2123 Views
  • 1 replies
  • 0 kudos
Latest Reply
Lakshay
Databricks Employee
  • 0 kudos

Hi @vishwanath_1 , Caching only comes into picture when there are multiple reference to data source in your code. As per the flow mentioned by you, I don't see that being the case for you. You are only reading the data from source once and also there...

  • 0 kudos
sudhakargen
by New Contributor II
  • 18926 Views
  • 2 replies
  • 0 kudos

Intermittently unavailable: Maven library com.crealytics:spark-excel_2.12:3.5.0_0.20.3

The issue is that the package com.crealytics:spark-excel_2.12:3.5.0_0.20.3 is intermittently unavailable i.e. most of the times excel import works and few times it fails with exception (org.apache.spark.SparkClassNotFoundException).I have installed m...

  • 18926 Views
  • 2 replies
  • 0 kudos
Latest Reply
sudhakargen
New Contributor II
  • 0 kudos

"Looks like the issue is source is not able to reach" - Can you please let me know what you mean by this.Libraries installed on the databricks cluster are as below, I have a cluster with14.2 version on which I have installed maven library(com.crealyt...

  • 0 kudos
1 More Replies
BartoszBiskupsk
by Databricks Partner
  • 3460 Views
  • 2 replies
  • 0 kudos

"Last Access" information for external delta tables (no UC)

Hi,Is there a way to make audit on all tables in hive_metastore (no UC), all are external, to check when each has been used for the last time (queried / updated / etc). ?

Data Engineering
access logs
  • 3460 Views
  • 2 replies
  • 0 kudos
Latest Reply
CharlesReily
New Contributor III
  • 0 kudos

Apache Ranger or Apache Sentry can be used for auditing Hive activities. If you have set up auditing in one of these tools, you can review the audit logs to see when tables were accessed. Audit logs are typically stored in a separate location, and yo...

  • 0 kudos
1 More Replies
hbs59
by New Contributor III
  • 11439 Views
  • 5 replies
  • 2 kudos

Resolved! Rest API Error 404

I am trying to export a notebook or directory using /api/2.0/workspace/export.When I run /api/2.0/workspace/list with a particular url and path, I get the results that I expect, a list of objects (notebooks and folders) at that location.But when I ru...

  • 11439 Views
  • 5 replies
  • 2 kudos
Latest Reply
Debayan
Databricks Employee
  • 2 kudos

Hi, Could you please remove the parameters , (format and direct_download) and confirm? 

  • 2 kudos
4 More Replies
drii_cavalcanti
by New Contributor III
  • 1866 Views
  • 1 replies
  • 0 kudos

Shared Mode Cluster Permission Issue: Editing Folders Across Users

Hi everyone,Currently, I save logs to a specific folder at the root level in Databricks. However, I need to use a Shared Mode cluster, and it seems I no longer have permission to save to the folder or even open its terminal to access the underlying i...

  • 1866 Views
  • 1 replies
  • 0 kudos
Latest Reply
Debayan
Databricks Employee
  • 0 kudos

Hi, If workspace access control is enabled, by default objects in this folder are private to that user. You can refer to https://docs.databricks.com/en/workspace/workspace-objects.html and let us know if this helps. 

  • 0 kudos
therealDE
by New Contributor II
  • 3554 Views
  • 3 replies
  • 1 kudos

databricks cli error : command >> databricks fs ls # getting error Error: accepts 1 arg(s), received

Hi team, I installed databricks cli on my mac using homebrew, below is the linkhttps://docs.databricks.com/en/dev-tools/cli/install.html#homebrew-installstep1:ran >> databricks configure , configured successfully.however, when I ran I am getting belo...

  • 3554 Views
  • 3 replies
  • 1 kudos
Latest Reply
therealDE
New Contributor II
  • 1 kudos

Thanks for the reply, when I install databricks cli in my windows, it was actually returning some directories even with databricks fs ls.I installed in windows with pip. You think pip install different t from brew install in mac

  • 1 kudos
2 More Replies
Phani1
by Databricks MVP
  • 2461 Views
  • 1 replies
  • 0 kudos

cloud fetch Qlik Sense

 Hi Team,Cloud Fetch will improve data transfer efficiency from DataBricks to Power BI and is it compatible with Qlik Sense as well ?

  • 2461 Views
  • 1 replies
  • 0 kudos
Latest Reply
Walter_C
Databricks Employee
  • 0 kudos

Cloud Fetch is a feature introduced in Databricks Runtime 8.3 and Simba ODBC 2.6.17 driver that significantly improves data transfer efficiency from Databricks to BI tools like Power BI. It achieves this by fetching data in parallel via cloud storage...

  • 0 kudos
Vishwanath_Rao
by New Contributor II
  • 3392 Views
  • 1 replies
  • 0 kudos

Same path producing different counts on Databricks and EMR

We're in the middle of migrating to Databricks and found that the same path on s3 is producing different counts between EMR (Spark 2.4.4) and Databricks (Spark 3.4.1) it is a simple spark.read.parquet().count(), tried multiple solutions like making t...

  • 3392 Views
  • 1 replies
  • 0 kudos
Latest Reply
Walter_C
Databricks Employee
  • 0 kudos

The discrepancy in counts between EMR (Spark 2.4.4) and Databricks (Spark 3.4.1) could be due to several reasons:1. Different versions of Spark: The two environments are running different versions of Spark which might have different optimizations or ...

  • 0 kudos
DmitriyLamzin
by New Contributor II
  • 8774 Views
  • 1 replies
  • 0 kudos

applyInPandas hangs on runtime 13.3 LTS ML and above

Hello, recently I've tried to upgrade my runtime env to the 13.3 LTS ML and found that it breaks my workload during applyInPandas.My job started to hang during the applyInPandas execution. Thread dump shows that it hangs on direct memory allocation: ...

Data Engineering
pandas PythonRunner
  • 8774 Views
  • 1 replies
  • 0 kudos
Latest Reply
Debayan
Databricks Employee
  • 0 kudos

Same post: https://community.databricks.com/t5/data-engineering/applyinpandas-function-hangs-in-runtime-13-3-lts-ml-and-above/td-p/56795

  • 0 kudos
MarsSu
by New Contributor II
  • 11503 Views
  • 3 replies
  • 0 kudos

How to implement merge multiple rows in single row with array and do not result in OOM?

Hi, Everyone.Currently I try to implement spark structured streaming with Pyspark. And I would like to merge multiple rows in single row with array and sink to downstream message queue for another service to use. Related example can follow as:* Befor...

  • 11503 Views
  • 3 replies
  • 0 kudos
Latest Reply
917074
Databricks Partner
  • 0 kudos

Is there any solution to this, @MarsSu  were you able to solve this, kindly shed some light on this if you resolve this.

  • 0 kudos
2 More Replies
uncle_rufus
by Databricks Partner
  • 3924 Views
  • 0 replies
  • 0 kudos

ipywidgets

I am having an issue with getting a display upon interaction with the ipywidgets dropdown menu. Once I've selected an option from the dropdown, nothing happens. I am inclined to believe it has to do with how I've structure my on_select function and n...

  • 3924 Views
  • 0 replies
  • 0 kudos
Jake2
by New Contributor III
  • 9478 Views
  • 4 replies
  • 1 kudos

Resolved! Z-Ordering a Unity Catalog Materialized View

Hey everyone, We're making the move to Unity Catalog from Hive_Metastore and we're running into some issues performing Z-order optimizations on some of our tables. These tables are, in either place, materialized views created with a "create or refres...

  • 9478 Views
  • 4 replies
  • 1 kudos
Latest Reply
Jake2
New Contributor III
  • 1 kudos

For anyone who's reading this later: You can still Z-order your materialized views, but you can't run it as a SQL command. Instead, you can set it as one of the TBLPROPERTIES when you define the table. Here's an example:create or refresh live table {...

  • 1 kudos
3 More Replies
DManowitz-BAH
by New Contributor II
  • 5699 Views
  • 4 replies
  • 1 kudos

Apparent bug with dbutils.fs.cp on S3 using DBR 13.3LTS

If I use dbutils.fs.cp on a cluster running DBR 13.3LTS to try to copy an object on S3 from one prefix to another, I don't get the expected results.For example, if I try the following command:dbutils.fs.cp('s3://some-bucket/some/prefix/some_file.gz',...

  • 5699 Views
  • 4 replies
  • 1 kudos
Latest Reply
Lakshay
Databricks Employee
  • 1 kudos

Hi @DManowitz-BAH , The correct syntax to use the dbutils.fs.cp command is to provide the file name in the destination path. Please check the document here: https://docs.databricks.com/en/dev-tools/databricks-utils.html#cp-command-dbutilsfscp

  • 1 kudos
3 More Replies
Labels