cancel
Showing results for 
Search instead for 
Did you mean: 
Data Engineering
Join discussions on data engineering best practices, architectures, and optimization strategies within the Databricks Community. Exchange insights and solutions with fellow data engineers.
cancel
Showing results for 
Search instead for 
Did you mean: 

Forum Posts

jduran9987
by Visitor
  • 40 Views
  • 2 replies
  • 1 kudos

Serverless Capabilities Not Available In My Workspace

Hello,My AWS Databricks account is paid Premium and has serverless compute enabled but the only types I see for SQL warehouses are Pro and Classic.I created two workspaces each with “Use serverless compute with default storage,” and "Use your existin...

  • 40 Views
  • 2 replies
  • 1 kudos
Latest Reply
masonreed11
Contributor
  • 1 kudos

It looks like a workspace-level Serverless eligibility issue rather than a missing prerequisite. Since the CLI specifically says the workspace is no longer eligible, I’d check the Serverless Compute status in the Databricks account console. If all pr...

  • 1 kudos
1 More Replies
Khasim_1
by New Contributor III
  • 117 Views
  • 1 replies
  • 0 kudos

Architectural Pattern for PII: "Redact at Ingestion" vs. "Dynamic Masking at Consumption

 Hi everyone, we are defining our PII strategy in Unity Catalog. We are split on whether to: A) Redact/Hash PII at the Silver layer (permanent change), or B) Keep PII in Silver and use Dynamic Data Masking at the Gold/View layer.Does "Redact at Inges...

  • 117 Views
  • 1 replies
  • 0 kudos
Latest Reply
Ashwin_DSA
Databricks Employee
  • 0 kudos

Hi @Khasim_1, I would keep the PII in Silver within your regulatory and internal retention limits and control access with Unity Catalog's row filters and column masks, rather than redacting permanently. The reason is the one you already gave. A perma...

  • 0 kudos
Wola
by New Contributor
  • 126 Views
  • 3 replies
  • 0 kudos

Data ingestion: Setting up a connector for Postgres with CDC enabled.

Hello,I'm trying to ingest data from my RDS instance and set up change data capture on Databricks. Everything on the Postgres side has been done so that replication is possible, but CDC is still greyed out. I have asked Claude, Gemini, and ChatGPT. I...

Wola_0-1789753526825.png
  • 126 Views
  • 3 replies
  • 0 kudos
Latest Reply
Wola
New Contributor
  • 0 kudos

Quick question, being that my account is an individual account, how do I go about this "Contact your Databricks account team to request access" 

  • 0 kudos
2 More Replies
Khasim_1
by New Contributor III
  • 138 Views
  • 2 replies
  • 2 kudos

Best practice to enforcing row-level security across multiple catalogs sharing same schema structure

Hi Everyone,We have identical schema structures replicated across three catalogs (dev, qa, prod) to support environment isolation. We now need to apply row-level security (e.g., restricting sales reps to only see their own region's data) consistently...

  • 138 Views
  • 2 replies
  • 2 kudos
Latest Reply
bijilsubhash
New Contributor III
  • 2 kudos

Agree with @ThomazNeto  - only thing I would add is your point 3 requirement: you could also consider using Terraform instead DABs or job for managing the ABAC policies across multiple environments. It does not support the beta features i.e. metastor...

  • 2 kudos
1 More Replies
Khasim_1
by New Contributor III
  • 118 Views
  • 1 replies
  • 0 kudos

Managing Row-Level Security (RLS) vs. Views: Performance impacts at scale

 Hi everyone, we are designing our security layer in Unity Catalog. We are debating between using Row-Level Security (RLS) predicates versus standard SQL Views for masking PII data.When applying complex RLS predicates to tables with >100M rows, have ...

  • 118 Views
  • 1 replies
  • 0 kudos
Latest Reply
ThomazNeto
Databricks Partner
  • 0 kudos

Hi Khasim,Most of this is answered directly in the Databricks docs, so let me point to what they say.1. Planning time vs. runtime. The docs don't describe a planning-time penalty; the cost is in what the optimizer is allowed to do. Both table-level r...

  • 0 kudos
Khasim_1
by New Contributor III
  • 126 Views
  • 1 replies
  • 0 kudos

Managing "Secret" Injection in DABs across Dev/Stage/Prod

Hi everyone, I am fully moving our team to Databricks Asset Bundles (DABs) for CI/CD, but I’m struggling with the "Secret" management pattern.How are you handling the injection of secrets (like API keys for ingestion) into DABs without hardcoding any...

  • 126 Views
  • 1 replies
  • 0 kudos
Latest Reply
ThomazNeto
Databricks Partner
  • 0 kudos

Hi Khasim,No secret value ever goes into databricks.yml. Bundle variables get resolved at deploy time and rendered into the job definition in the workspace, so anyone who can open the job can read them. Only names and references live in the bundle.Se...

  • 0 kudos
Gilk
by New Contributor II
  • 211 Views
  • 3 replies
  • 0 kudos

Predictive Optimization for Streaming tables in Lakeflow pipelines

On July 2025 https://www.databricks.com/blog/whats-new-lakeflow-declarative-pipelines-july-2025 predictive optimization was enabled for all UC managed Lakeflow pipelines. I was wondering if there is a possibility now to disable it for Streaming table...

Data Engineering
dlt
lakeflow pipelines
  • 211 Views
  • 3 replies
  • 0 kudos
Latest Reply
Ashwin_DSA
Databricks Employee
  • 0 kudos

Hi @Gilk, Before the how, can I ask what is driving the wish to turn it off? For most pipeline tables predictive optimization is doing useful work, running OPTIMIZE, VACUUM, and ANALYZE on serverless compute so file sizes, storage, and statistics sta...

  • 0 kudos
2 More Replies
Fatimah-Tariq
by New Contributor III
  • 242 Views
  • 4 replies
  • 0 kudos

Lakeflow connect CT pipeline keeps running multiple connections in source

I'm working on an ingestion pipeline. I was adding tables to it gradually so that we can monitor the load on our prod server side by side and with my last set of tables added (with them, all the heavy duty tables were inside that pipeline), the load ...

FatimahTariq_0-1789473365412.png
  • 242 Views
  • 4 replies
  • 0 kudos
Latest Reply
AbhilashNagilla
Databricks Employee
  • 0 kudos

For a standard connector, the gateway continuously extracts snapshots, change logs, and metadata, while a separately scheduled serverless pipeline applies staged data (connector components). Changing the serverless pipeline schedule therefore does no...

  • 0 kudos
3 More Replies
kabiromohd
by New Contributor
  • 176 Views
  • 2 replies
  • 0 kudos

Databricks Free Edition

Hi,I just signed up for Databricks free edition to start my learning journey in Data Engineering on it.Want to know for how long it remains free?Thank you.

  • 176 Views
  • 2 replies
  • 0 kudos
Latest Reply
Ashwin_DSA
Databricks Employee
  • 0 kudos

Hi @kabiromohd,   Welcome to Databricks, and a great choice for starting your Data Engineering journey!   The short answer is that Databricks Free Edition is forever free. There is no trial period or expiration date. You can use it indefinitely as lo...

  • 0 kudos
1 More Replies
emorgoch
by Contributor
  • 999 Views
  • 9 replies
  • 2 kudos

Resolved! Using autoloader with multiple object types in load path

My data source is going to generate csv files for multiple objects all into the same directory that I need to load from. The files will have names in the format along the lines of <objecttype>_YYYY_MM_DD_guid.csv.gz. Each objecttype will have it's ow...

  • 999 Views
  • 9 replies
  • 2 kudos
Latest Reply
stephen4
New Contributor III
  • 2 kudos

A single Auto Loader stream with filename-based routing looks much cleaner here, especially since new object types can appear without updating the pipeline.

  • 2 kudos
8 More Replies
LiresaFerizaj
by New Contributor II
  • 177 Views
  • 2 replies
  • 2 kudos

Liquid Clustering vs Z-Order vs Partitioning

 Liquid clustering: what it actually replaced, and the caveats worth knowingI see a lot of questions here that boil down to "should I partition, Z-order, or use liquid clustering." Posting how I reason about it, partly to help and partly because I'd ...

  • 177 Views
  • 2 replies
  • 2 kudos
Latest Reply
aayush_410
New Contributor
  • 2 kudos

This lines up well with what Databricks' own docs and the 2026 refresh of their partitioning guidance say — a few things worth adding as confirmation/extra data points rather than pushback:On the key-count cap: confirmed at 4 clustering keys max. Wor...

  • 2 kudos
1 More Replies
DineshOjha
by New Contributor III
  • 188 Views
  • 3 replies
  • 3 kudos

Notebook commands hang after restarting a cluster with custom library installed

Hi,I'm experiencing an issue with a Databricks notebook after restarting a cluster that has a custom library installed.Here are the steps I followed:Created a Databricks compute cluster.Attached a notebook to the cluster.Installed a library required ...

  • 188 Views
  • 3 replies
  • 3 kudos
Latest Reply
DineshOjha
New Contributor III
  • 3 kudos

Thank you for the response. From the driver logs I gather that the REPL has crashed, but this isn't the case only with Hive. I have noticed this pattern with Oracle library installation as well.

  • 3 kudos
2 More Replies
seanpmcn
by New Contributor II
  • 641 Views
  • 5 replies
  • 4 kudos

Resolved! Issue in "Build a Declarative Pipeline with Spark Declarative Pipelines"

I am trying to complete the "Get Started with Data Engineering" course, but I have been running into an issue.I have gotten to this step in "Build a Declarative Pipeline with Spark Declarative Pipelines":Demo: Create and Run the PipelineNow you'll co...

  • 641 Views
  • 5 replies
  • 4 kudos
Latest Reply
seanpmcn
New Contributor II
  • 4 kudos

I have added the source now. I went to the root folder > Create pipeline asset > New pipeline source code folder. I was then able to go to the right of the pipeline name at the top of the screen and edit the catalogue and schema.However, after runnin...

  • 4 kudos
4 More Replies
Labels