cancel
Showing results for 
Search instead for 
Did you mean: 
Data Engineering
Join discussions on data engineering best practices, architectures, and optimization strategies within the Databricks Community. Exchange insights and solutions with fellow data engineers.
cancel
Showing results for 
Search instead for 
Did you mean: 

Forum Posts

data_pulse
by New Contributor II
  • 41 Views
  • 2 replies
  • 0 kudos

Do UC table tags propagate to billing for Predictive Optimization and Data Quality Monitoring?

I am looking at a scenario for team-level cost attribution using system.billing.usage and have a tagging gap for Databricks-managed services such as Predictive Optimization and Data Quality Monitoring.Tried adding a team/ownership tag directly to the...

  • 41 Views
  • 2 replies
  • 0 kudos
Latest Reply
stbjelcevic
Databricks Employee
  • 0 kudos

Hi @data_pulse ,   Short answer: no, that propagation isn't supported. Unity Catalog catalog/schema/table tags do not flow into system.billing.usage.custom_tags for Predictive Optimization, Data Quality Monitoring, or Lakehouse Monitoring. This is by...

  • 0 kudos
1 More Replies
srikanthp24
by New Contributor
  • 76 Views
  • 3 replies
  • 0 kudos

Multiple gateway pipeline for same database

Hi Team,To do the POC on Lakeflow Connect Ingestion Pipeline, we directly connected the sql server prod DB and build the pipeline it is running successfully. Now, My question is we need to move that to UAT and Prod. In Dev we have created througn UI ...

  • 76 Views
  • 3 replies
  • 0 kudos
Latest Reply
stbjelcevic
Databricks Employee
  • 0 kudos

@srikanthp24 , To your core question: multiple gateway pipelines against the same source database don't interfere with each other. Each gateway keeps its own independent state, its own staging volume, and its own CDC/change-tracking cursor, so a sepa...

  • 0 kudos
2 More Replies
pragya17
by New Contributor III
  • 43 Views
  • 1 replies
  • 0 kudos

Databricks Dashboard Embedded Ask Genie

Hi ,I have created Dashboard on Databricks using sql datasets and published with " Shared Data Permission " with the business user . he can see the data and charts but unable to use "Ask Genie " functionality . Do I need to grant Use Catalog , schema...

  • 43 Views
  • 1 replies
  • 0 kudos
Latest Reply
Valeria_Koz_DBX
Databricks Employee
  • 0 kudos

Hi @pragya17 ,You are correct that Shared data permission should not require the business user to have direct USE CATALOG, USE SCHEMA, or SELECT permissions on the underlying data. With embedded/shared credentials, the autogenerated Genie Agent shoul...

  • 0 kudos
DB1To3
by Contributor
  • 71 Views
  • 2 replies
  • 0 kudos

Resolved! Can't access abfss data in azure databricks when providing shared key (fighting UC?)

The "no isolation shared" clusters are going to be killed in the next month.  This came as a surprise to me.  I should have been paying closer attention.  We don't heavily use UC since our data is published to the business via Fabric.  All the spark ...

  • 71 Views
  • 2 replies
  • 0 kudos
Latest Reply
DB1To3
Contributor
  • 0 kudos

I found another reply where the customer was required to add the service principal's full GRANT to the external locations.   Seems very odd, and almost less secure than what I was doing with the shared access key.I really think there needs to be a ne...

  • 0 kudos
1 More Replies
Islam_hoti
by New Contributor II
  • 74 Views
  • 3 replies
  • 0 kudos

How are you separating dev, staging and prod in Unity Catalog without duplicating everything?

Hi everyone, We are trying to settle on an environment strategy and keep going back and forth, so I would like to hear how other teams actually landed on this rather than what the reference architecture suggests. Our situation. One Databricks account...

  • 74 Views
  • 3 replies
  • 0 kudos
Latest Reply
data_pulse
New Contributor II
  • 0 kudos

From experience of working in diff projects, separate workspaces + environment catalogs + governed Prod reads from Dev is a practical combination.Environment isolation: Dev/Test/Prod workspaces with dev_*, test_*, prod_* catalogs. Personal developmen...

  • 0 kudos
2 More Replies
1pedroosilva
by New Contributor II
  • 141 Views
  • 4 replies
  • 9 kudos

Databricks Dashboard Pivot Table: default collapsed state?

I'm building a financial dashboard in Databricks using a Pivot visualization with two dimensions in the row hierarchy.The drill-down works as expected, but I'm having a usability issue: every time the view is loaded, both dimensions are fully expande...

  • 141 Views
  • 4 replies
  • 9 kudos
Latest Reply
hayoni
New Contributor II
  • 9 kudos

Hello, @1pedroosilva !As others mentioned, there is no built-in setting for this yet. However, if you want a quick solution without changing anything in your dashboard or queries, here is a simple workaround.Step1. Create a new bookmark in your brows...

  • 9 kudos
3 More Replies
lessalucas
by New Contributor
  • 43 Views
  • 0 replies
  • 0 kudos

Databricks Secrets

Recently Databricks Secrets became generally available.Now, you can manage your secrets whitin the catalog using the namespace: catalog.schema.secretHowever, you can do the opposite of what you used to do: Connect Unity Catalog secrets to Azure Key V...

  • 43 Views
  • 0 replies
  • 0 kudos
tobyevans
by New Contributor II
  • 11962 Views
  • 2 replies
  • 3 kudos

Ingesting complex/unstructured data

Hi there,my company is reasonably new to using Databricks, and we're running our first PoCs.  Some of the data we have structured/reasonably structured, so it drops into a bucket, we point a notebook at it, and all is well and DeltaThe problem is ari...

  • 11962 Views
  • 2 replies
  • 3 kudos
Latest Reply
mark_alexander
  • 3 kudos

Ingesting complex or unstructured data works best when you break the process into stages rather than trying to analyze everything at once. First, collect the raw data, then clean it, remove duplicates, standardize formats, and extract only the fields...

  • 3 kudos
1 More Replies
Islam_hoti
by New Contributor II
  • 152 Views
  • 5 replies
  • 4 kudos

Resolved! How are you attributing serverless costs back to individual jobs and teams?

Hi everyone,We moved most of our workloads to serverless over the last few months and the performance side has been fine. The part we did not plan for is cost attribution.On classic compute this was easy. One cluster per team, tags on the cluster, an...

  • 152 Views
  • 5 replies
  • 4 kudos
Latest Reply
ivanvyd
New Contributor II
  • 4 kudos

@Islam_hoti the most defensible model is to separate measured attribution from allocation policy instead of forcing every DBU into a team bucket.Exact workload attribution. Use system.billing.usage as the ledger. For serverless jobs, group by workspa...

  • 4 kudos
4 More Replies
LakehouseLad
by New Contributor
  • 140 Views
  • 5 replies
  • 1 kudos

Serverless compute resolves some public domains but not others

Hi everyone,I’m troubleshooting outbound internet access from Databricks serverless compute and am seeing inconsistent DNS resolution.For example, the GitHub API domain resolves successfully:api [dot] github [dot] comDNS: PASS -> DNS resolved success...

LakehouseLad_0-1788843214064.png
  • 140 Views
  • 5 replies
  • 1 kudos
Latest Reply
data_pulse
New Contributor II
  • 1 kudos

@LakehouseLad I also tested the same from a Premium tier workspace, and the same public domains resolve successfully there. So this does not appear to be a general Premium tier limitation.In the screenshot you posted is showing the DNS request as DRO...

  • 1 kudos
4 More Replies
gowri_databrick
by New Contributor II
  • 220 Views
  • 8 replies
  • 2 kudos

Understanding Parquet File Storage for Large Datasets

Hi everyone,I’m learning about Parquet files and how they are used in Databricks for storing large datasets.I’m trying to understand how column-based storage works in a practical situation.For example, suppose an e-commerce company has 500 million or...

  • 220 Views
  • 8 replies
  • 2 kudos
Latest Reply
Coffee77
Honored Contributor III
  • 2 kudos

In addition to my previous picture, I'd add the following explanation. In Databricks, Partition Pruning, Data Skipping and Parquet’s columnar format work together to minimize I/O.1. Partition Pruning → eliminates partitions.If the table is partitione...

  • 2 kudos
7 More Replies
seanpmcn
by New Contributor
  • 158 Views
  • 3 replies
  • 1 kudos

Resolved! Issue in "Build a Declarative Pipeline with Spark Declarative Pipelines"

I am trying to complete the "Get Started with Data Engineering" course, but I have been running into an issue.I have gotten to this step in "Build a Declarative Pipeline with Spark Declarative Pipelines":Demo: Create and Run the PipelineNow you'll co...

  • 158 Views
  • 3 replies
  • 1 kudos
Latest Reply
seanpmcn
New Contributor
  • 1 kudos

I have added the source now. I went to the root folder > Create pipeline asset > New pipeline source code folder. I was then able to go to the right of the pipeline name at the top of the screen and edit the catalogue and schema.However, after runnin...

  • 1 kudos
2 More Replies
gowri_databrick
by New Contributor II
  • 185 Views
  • 3 replies
  • 0 kudos

What is a Data Skipping in Delta Lake?

Hi everyone,I’m learning about Delta Lake performance and came across data skipping.I understand that it can help Databricks avoid reading unnecessary data when running queries, but I’d like to understand its purpose more clearly.For example, if an o...

  • 185 Views
  • 3 replies
  • 0 kudos
Latest Reply
Coffee77
Honored Contributor III
  • 0 kudos

I think the below picture will help a bit. Take a look at section 2 for "data skipping": 

  • 0 kudos
2 More Replies
AshokB
by New Contributor II
  • 263 Views
  • 5 replies
  • 3 kudos

Partitioning vs Liquid Clustering (per-table):

Can PARTITION BY and CLUSTER BY (Liquid Clustering) be used simultaneously on the same table? If we use only PARTITION BY, is there a negative performance impact on materialized-view refreshes in Silver/Gold? Since materialized views read only increm...

  • 263 Views
  • 5 replies
  • 3 kudos
Latest Reply
Coffee77
Honored Contributor III
  • 3 kudos

Not possible at all BUT here is a summary comparing both methods:I hope it helps. 

  • 3 kudos
4 More Replies
km1837
by Databricks Partner
  • 102 Views
  • 3 replies
  • 0 kudos

AI/BI Dashboards

Hi AllMy client is looking to create a few dashboards in DataBricks instead of Power BI. We already have a template in Power BI (header for each page, information about out BI Page) in each of our power bi report which has colors, fonts, images etc.H...

  • 102 Views
  • 3 replies
  • 0 kudos
Latest Reply
Coffee77
Honored Contributor III
  • 0 kudos

I would suggest you standardize Databricks dashboards, although the approach is a bit different from Power BI.For the visual template, I would create a “master” Lakeview/AI/BI dashboard containing the common elements: header, logo/images, standard te...

  • 0 kudos
2 More Replies
Labels