cancel
Showing results for 
Search instead for 
Did you mean: 
Data Engineering
Join discussions on data engineering best practices, architectures, and optimization strategies within the Databricks Community. Exchange insights and solutions with fellow data engineers.
cancel
Showing results for 
Search instead for 
Did you mean: 

Forum Posts

Ericsson
by • New Contributor II
  • 7962 Views
  • 6 replies
  • 1 kudos

SQL week format issue its not showing result as 01(ww)

Hi Folks,I've requirement to show the week number as ww format. Please see the below codeselect weekofyear(date_add(to_date(current_date, 'yyyyMMdd'), +35)). also plz refre the screen shot for result.

result
  • 7962 Views
  • 6 replies
  • 1 kudos
Latest Reply
Rodgers
Visitor
  • 1 kudos

Data engineering is an increasingly important field, and communities focused on best practices and practical solutions can be valuable for professionals at every level. Discussions about architecture, optimization, data pipelines, and efficient workf...

  • 1 kudos
5 More Replies
Khasim_1
by • New Contributor III
  • 232 Views
  • 2 replies
  • 4 kudos

Architectural Pattern for PII: "Redact at Ingestion" vs. "Dynamic Masking at Consumption

 Hi everyone, we are defining our PII strategy in Unity Catalog. We are split on whether to: A) Redact/Hash PII at the Silver layer (permanent change), or B) Keep PII in Silver and use Dynamic Data Masking at the Gold/View layer.Does "Redact at Inges...

  • 232 Views
  • 2 replies
  • 4 kudos
Latest Reply
gnakan
Databricks MVP
  • 4 kudos

Hey @Khasim_1 , if it helps...did an experiment and blog post with results from the run that illustrates the pattern that @Ashwin_DSArecommended.https://medium.com/lab-notes/redacting-pii-in-silver-leaves-it-in-bronze-column-masks-vs-redaction-on-dat...

  • 4 kudos
1 More Replies
LiresaFerizaj
by • New Contributor III
  • 107 Views
  • 1 replies
  • 2 kudos

AUTO CDC SCD Type 2 with late-arriving deletes: edge cases the docs don't cover

Hi everyone,I'm designing a Lakeflow Declarative Pipeline that processes customer profile changes from a CDC feed into an SCD Type 2 Silver table. Events can arrive out of order, sometimes up to 24 hours late. The planned flow:CREATE OR REFRESH STREA...

  • 107 Views
  • 1 replies
  • 2 kudos
Latest Reply
ThomazNeto
Databricks Partner
  • 2 kudos

Hi Liresa,I went through the current docs for these. Two of your three questions have gaps in the documentation, so I'll separate what's documented from what I'd verify.1. Deletes after tombstone expiry. Not documented. The only guidance is the rule ...

  • 2 kudos
ChristianRRL
by • Honored Contributor II
  • 1318 Views
  • 4 replies
  • 5 kudos

Resolved! Unity Catalog - How to read prod data in dev with appropriate read-only access?

Hi there,Our team is currently migrating to using Unity Catalog. We have two databricks workspaces for dev & prod, and one thing that I'm wondering is if there is a simple/appropriate way to have only two catalogs dev & prod, where the prod databrick...

  • 1318 Views
  • 4 replies
  • 5 kudos
Latest Reply
Nagendra_B
New Contributor II
  • 5 kudos

Yes, we can share prod data into a dev workspace in Unity Catalog with read-only access. Databricks has built this to be remarkably simple without lengthy processes.The only prerequisites are that both workspaces must be in the same region and attach...

  • 5 kudos
3 More Replies
DB1To3
by • Contributor
  • 180 Views
  • 3 replies
  • 4 kudos

Unity Catalog v.2.0?

Are there plans for a v.2.0 of Unity Catalog?  I find the organization of tables in databricks to be very constrained and inflexible.  Given that UC tables have their origins in data lakes, you would think they would have brought a lot more flexibile...

  • 180 Views
  • 3 replies
  • 4 kudos
Latest Reply
LiresaFerizaj
New Contributor III
  • 4 kudos

Fair points, especially for data coming from systems with richer naming. I can't speak to the roadmap, but here's what works for us today:Aliases: Create views in a consumer-facing catalog, like reporting.sales.outside_sales, that point to the real t...

  • 4 kudos
2 More Replies
Srini_Pesala
by • Databricks Partner
  • 209 Views
  • 4 replies
  • 5 kudos

Best Architecture Approach for Databricks Multi-Cloud and Multi-Region Deployment

I am working on a data platform requirement where we need to support multiple cloud providers and multiple regions using Databricks.For example:AWS – Region 1AWS – Region 2Azure – Region 1Azure – Region 2The requirement is to maintain data processing...

  • 209 Views
  • 4 replies
  • 5 kudos
Latest Reply
Khasim_1
New Contributor III
  • 5 kudos

Hi @Srini_Pesala ,One Databricks workspace per Cloud + Region combination. So you'll have 4 separate workspaces (AWS-R1, AWS-R2, Azure-R1, Azure-R2). They don't share compute or storage directly.How to Handle Each ConcernData Residency Each workspace...

  • 5 kudos
3 More Replies
LiresaFerizaj
by • New Contributor III
  • 107 Views
  • 2 replies
  • 2 kudos

Databricks now deletes unused service principal secrets after 90 days

Here's a small change from this month's release notes that's easy to miss. Databricks now automatically deletes a service principal's OAuth client secret after 90 days without use. This matches how unused access tokens were already cleaned up.Why thi...

  • 107 Views
  • 2 replies
  • 2 kudos
Latest Reply
Khasim_1
New Contributor III
  • 2 kudos

Great write-up, Liresa — this is a really important operational gotcha that's easy to overlook until it bites during a quarterly job or DR drill. A few thoughts to add to the discussion:Automate tracking instead of manual lists Set up a scheduled che...

  • 2 kudos
1 More Replies
Mado
by • Valued Contributor II
  • 148 Views
  • 4 replies
  • 4 kudos

How can I configure Lakeflow Connect SQL Server CDC Gateway to use a desired VM type?

Hi Team,I'm evaluating Databricks Lakeflow Connect for SQL Server CDC ingestion and have run into a gateway provisioning issue.EnvironmentRegion: Australia EastSource: Azure SQL DatabaseCDC enabled successfully on the database and source tableSQL Ser...

Mado_0-1790214214728.png
  • 148 Views
  • 4 replies
  • 4 kudos
Latest Reply
Khasim_1
New Contributor III
  • 4 kudos

Hi @Mado As of now, Lakeflow Connect's CDC Gateway VM types (driver/worker SKUs) are not configurable at the pipeline or connection level — they are managed internally by the platform. If the default SKU (Standard_E4d_v4) is unavailable in your regio...

  • 4 kudos
3 More Replies
Khasim_1
by • New Contributor III
  • 163 Views
  • 1 replies
  • 0 kudos

Scaling Delta Sharing across Non-Databricks Recipients: Handling Credential Rotation and Lifecycle

Hi everyone,Hi everyone, I’m architecting a data-sharing program using Delta Sharing for external partners who do not have a Databricks workspace.How are others automating the recipient credential lifecycle? Do you have an automated process for rotat...

  • 163 Views
  • 1 replies
  • 0 kudos
Latest Reply
SumeshKashyap
New Contributor III
  • 0 kudos

The pattern that keeps this manageable at scale is treating recipients as code and running rotation as a scheduled job. Here's the breakdown: - Set a token lifetime at the metastore level. Open-sharing bearer tokens expire according to the metastore'...

  • 0 kudos
ChristianRRL
by • Honored Contributor II
  • 115 Views
  • 1 replies
  • 2 kudos

Databricks Workflow Task Failure - Custom Error Messages

Hi there, kind of silly question but I'm hoping there's a simple way to address this.I'm trying to make sure that when certain job failure conditions arise, that a workflow task step fails. I've tried this with both raise RuntimeError and sys.exit, b...

ChristianRRL_0-1790348230722.png ChristianRRL_0-1790348341444.png
  • 115 Views
  • 1 replies
  • 2 kudos
Latest Reply
K_Anudeep
Databricks Employee
  • 2 kudos

Hey @ChristianRRL ! When a Python wheel/notebook/script task fails, Jobs always does three things: Render the Python traceback in task output (IPython / Databricks REPL showtraceback).Extract a one-line summary (RuntimeError: <your message>), which i...

  • 2 kudos
LiresaFerizaj
by • New Contributor III
  • 102 Views
  • 1 replies
  • 5 kudos

Genie One Foundations: What Data Teams Should Get Right Before Rolling It Out

Hi everyone Genie One has been getting a lot of attention since DAIS, and most of the discussion is about what business users can do with it. I want to focus on the other side: what data teams need in place so it gives trustworthy answers.Quick recap...

  • 102 Views
  • 1 replies
  • 5 kudos
Latest Reply
ivanvyd
New Contributor III
  • 5 kudos

Thanks, @LiresaFerizaj, appreciate your efforts! I'd add one step to the weekly review: turn recurring mistakes into regression tests, then rerun them whenever metric definitions, table descriptions or agent instructions change.For a Genie Agent, the...

  • 5 kudos
dawidch
by • New Contributor
  • 132 Views
  • 3 replies
  • 2 kudos

org.apache.iceberg.connect.IcebergSinkConnector to sink data from kafka to databricks - large volume

Hey.We have different number of topics per let's call it "subject". The connector works great for us if there are not so many topics (partitions) in subject. We have issue when we try to sink 1500 topics, 3 partitions each. We've sharded the topics i...

  • 132 Views
  • 3 replies
  • 2 kudos
Latest Reply
ivanvyd
New Contributor III
  • 2 kudos

@dawidch, there is a related upstream ticket: Apache Iceberg #15852⁠, covering scheduled refresh of credentials held by ADLSFileIO. The proposed implementation is still open. I'd treat it as a lead, not yet a confirmed match for your connector versio...

  • 2 kudos
2 More Replies
koosha
by • New Contributor
  • 183 Views
  • 6 replies
  • 1 kudos

Genie agent hyperlink creation

I have created a dasboard about a ticketing system, the business user use this dashboard with the genie agent integrated to this dashboard, sometimes they need to ask genie to create a table which contains the ticket id with the link of the related i...

  • 183 Views
  • 6 replies
  • 1 kudos
Latest Reply
Ashwin_DSA
Databricks Employee
  • 1 kudos

Hi @koosha, Tried replicating this in my sandbox, on a small tickets and incidents dataset in a Genie space, and here is what I observed. When you ask Genie to show the ticket ID as a markdown link, it does put the right value in the cell. Whether th...

  • 1 kudos
5 More Replies
Khasim_1
by • New Contributor III
  • 175 Views
  • 4 replies
  • 3 kudos

Operationalizing Lakeflow Connect:Handling Upstream Schema Evolution & Historical Backfill

Hi everyone,I’m currently evaluating Lakeflow Connect for our ingestion layer. While the setup for replicating source databases into the Lakehouse is remarkably streamlined, I’m looking for operational best practices from those of you using it in pro...

  • 175 Views
  • 4 replies
  • 3 kudos
Latest Reply
stephen4
New Contributor III
  • 3 kudos

The point about keeping Silver on an explicit schema contract is really useful, especially during full refreshes. I’d also keep an eye on CDC retention during large backfills so nothing gets missed. 

  • 3 kudos
3 More Replies
Labels