topic PowerBI performance with Databricks in Data Engineering

PowerBI performance with Databricks

Vetrivel — Mon, 09 Dec 2024 08:23:56 GMT

We have integrated PowerBI with Databricks to generate reports. However, PowerBI generates over 8,000 lines of code, including numerous OR clauses, which cannot be modified at this time. This results in queries that take more than 4 minutes to execute and are automatically cancelled before a plan is generated. The time required for query optimization and file pruning further delays the process, preventing the plan from being generated. As a result, we are unable to use the report with Databricks, as queries containing numerous OR clauses are either taking an excessive amount of time to execute or failing altogether.
Please note that we have already implemented optimization techniques within Databricks, and our data consists of small files, such as 1 file in the DIM table and 22 files in the FACT tables. Adjusting the size of the serverless SQL warehouse has not resolved the issue.

If anyone has successfully addressed this issue, please share your solution.

Re: PowerBI performance with Databricks

Sidhant07 — Mon, 09 Dec 2024 09:13:36 GMT

Hi,

To address this issue, here are some suggestions:

Use the BROADCAST hint to optimize the join between the DIM and FACT tables. This can help reduce the amount of data that needs to be processed and improve the performance of the query.
Use the MERGE statement to combine the OR clauses into a single query. This can help reduce the number of queries that are generated and improve the performance of the query.
Use the OPTIMIZE command to optimize the Delta tables. This can help improve the performance of the query by reducing the amount of data that needs to be read and processed.
Use the VACUUM command to remove any deleted files from the Delta tables. This can help improve the performance of the query by reducing the amount of data that needs to be read and processed.
Use the ZORDER command to optimize the layout of the Delta tables. This can help improve the performance of the query by reducing the amount of data that needs to be read and processed.

Re: PowerBI performance with Databricks

Vetrivel — Mon, 09 Dec 2024 09:33:53 GMT

@Sidhant07 The issue is not related to the volume of data, as it is relatively small. Rather, the challenge lies in the time it takes to generate the plan in Databricks, which results in the process being automatically cancelled. Consequently, we are unable to retrieve the complete query from the query history. Additionally, we cannot modify the query generated by PowerBI at this time. We have already implemented liquid clustering for the FACT and DIM tables.

Re: PowerBI performance with Databricks

Vetrivel — Mon, 09 Dec 2024 10:06:29 GMT

As per our analysis, “joins” are not a problem but the huge “where” clause with lot of “OR” conditions.

Re: PowerBI performance with Databricks

Vetrivel — Mon, 09 Dec 2024 15:27:38 GMT

Attached is the sample query generated by Power BI. Without the OR conditions the query runs within seconds.