cancel
Showing results for 
Search instead for 
Did you mean: 
Data Engineering
Join discussions on data engineering best practices, architectures, and optimization strategies within the Databricks Community. Exchange insights and solutions with fellow data engineers.
cancel
Showing results for 
Search instead for 
Did you mean: 

Forum Posts

Avinash_Turpu
by • New Contributor II
  • 778 Views
  • 1 replies
  • 0 kudos

Resolved! 50% discount voucher not received despite completing all 4 courses

Hi team, I completed all 4 required courses for the recent Learning Festival campaign (June 15 – July 6, 2026): - DevOps Essentials for Data Engineering - Build Data Pipelines with Lakeflow Spark Declarative Pipelines - Deploy Workloads with Lakeflow...

  • 778 Views
  • 1 replies
  • 0 kudos
Latest Reply
Advika
Community Manager
  • 0 kudos

Hello @Avinash_Turpu, Could you please send a DM to @Jim_Anderson with the email address associated with your Customer Academy account? That will help him verify your participation and assist you further.

  • 0 kudos
yit337
by • Contributor II
  • 493 Views
  • 1 replies
  • 1 kudos

What timezone is the “timestamp” value on a table's history?

Is it UTC, or the local zone of whoever accesses it?

  • 493 Views
  • 1 replies
  • 1 kudos
Latest Reply
szymon_dybczak
Esteemed Contributor III
  • 1 kudos

Hi  @yit337 , TIMESTAMP values are internally normalized and persisted in UTC, but they are displayed using the current session’s local timezone.Therefore, it is not based on the timezone of the person who originally performed the write. It depends o...

  • 1 kudos
APJESK
by • Contributor
  • 457 Views
  • 1 replies
  • 0 kudos

Clarification on Auto Loader Managed File Events with Unity Catalog Managed Volumes

Hi Databricks Team,I'm trying to understand how Auto Loader Managed File Events work with Unity Catalog Managed Volumes, and I'm looking for some clarification.My understandingI create an External Location: CREATE EXTERNAL LOCATION app_ext_locURL 's3...

  • 457 Views
  • 1 replies
  • 0 kudos
Latest Reply
Ashwin_DSA
Databricks Employee
  • 0 kudos

Hi @APJESK, I am wondering if the main source of confusion is the name cloudFiles.useManagedFileEvents. In the public docs, "managed" here refers to the Databricks-managed file-events service and cache layer, not to Unity Catalog managed volumes or t...

  • 0 kudos
None123
by • New Contributor III
  • 13988 Views
  • 4 replies
  • 3 kudos

Open a Support Ticket

Anyone know how to submit a support ticket? I keep getting into a loop that takes me back to the community page, but I need to submit an urgent ticket. I'm told our company pays a ridiculous sum for this feature yet it is impossible to find.Thanks ...

  • 13988 Views
  • 4 replies
  • 3 kudos
Latest Reply
vermyas
New Contributor II
  • 3 kudos

Hello Databricks Support Team,We are currently using:databricks-sql-connector==4.3.0thrift==0.22.0Our organization's security scanning tool (Prisma) is flagging Thrift 0.22.0 and requires us to upgrade to a newer version (0.23.0 or later).As part of ...

  • 3 kudos
3 More Replies
JD18
by • New Contributor
  • 2889 Views
  • 3 replies
  • 2 kudos

SCD-2 backfilling with streaming tabels

Hi there,Im new to Databricks and trying to build a SCD2 type table using AUTO CDC approach. while it quite simple to create a scd2 table Im unable to do a backfill.Full context.I have raw data(order, customer info) from 2019 and creating a dimension...

  • 2889 Views
  • 3 replies
  • 2 kudos
Latest Reply
SteveOstrowski
Databricks Employee
  • 2 kudos

Hi @JD18, Welcome to Databricks and Thanks for raising this. SCD Type 2 backfilling with streaming tables is a common need, and the good news is that the AUTO CDC framework (formerly APPLY CHANGES INTO) has built-in capabilities to handle this -- you...

  • 2 kudos
2 More Replies
stravi001
by • New Contributor III
  • 1130 Views
  • 5 replies
  • 1 kudos

Resolved! Permissions on job with SQL task

Hello,I am authenticated as GCP service account. I have to create query, run job as service principle and get result.I'm both cases, I set permissions CAN_MANAGE for group "users" for both query and job. When I create the query with RunAsMode.OWNER, ...

  • 1130 Views
  • 5 replies
  • 1 kudos
Latest Reply
stravi001
New Contributor III
  • 1 kudos

Hello @balajij8,I tried to use SQL File Task and was able to achieve the result. Only issue was, that I was not able to get output using WorkspaceClient.jobs.get_run_output(). It worked will direct call of /api/2.2/jobs/runs/get-output endpoint.Thank...

  • 1 kudos
4 More Replies
emorgoch
by • Contributor
  • 1770 Views
  • 4 replies
  • 3 kudos

Resolved! Managing IPYNB cell timestamps in source control

We're in the process of converting over our Databricks notebooks from .py file to .ipynb. We have disabled storing notebook output in source control at the workspace level.However, what we're discovering is that every cell in our notebooks has 3 time...

emorgoch_0-1781635989625.png
  • 1770 Views
  • 4 replies
  • 3 kudos
Latest Reply
Ashwin_DSA
Databricks Employee
  • 3 kudos

Hi @emorgoch, Thanks for raising this. This appears to be a regression rather than expected behaviour. Internally, the issue has been identified around .ipynb handling in Git folders, and the intended fix is to stop serialising these execution timest...

  • 3 kudos
3 More Replies
DazzaiDe
by • New Contributor III
  • 472 Views
  • 1 replies
  • 0 kudos

Declarative Asset Bundle Service Principals best practices

I have 2 service principals for deployment and runtime like databricks suggested, since declarative asset bundles dont include delta tables creation what should i just let my runtime to create table on the first run of a specific job?I want the commu...

  • 472 Views
  • 1 replies
  • 0 kudos
Latest Reply
balajij8
Esteemed Contributor II
  • 0 kudos

You can configure the runtime SP to create tables dynamically on the first run if you prioritize operational simplicity and self-contained pipelines. The data processing scripts inject CREATE TABLE IF NOT EXISTS statements just before executing Data ...

  • 0 kudos
ConnorK
by • Databricks Partner
  • 820 Views
  • 4 replies
  • 4 kudos

Databricks Standard SharePoint Connector Performance Issues

I've recently started using the Databricks Standard SharePoint connector within my workspace and have run into some significant performance issues.My notebook does a straightforward read using the following:lakeflow_connection_name = 'sharepoint_dev'...

  • 820 Views
  • 4 replies
  • 4 kudos
Latest Reply
ConnorK
Databricks Partner
  • 4 kudos

Hi all, thanks for the help.Unfortunately pathGlobFilter didn't improve performance for me.I ended up working around it with some pre-processing to figure out which specific folders I actually needed, then looping through those directories with the c...

  • 4 kudos
3 More Replies
Emmasophia666
by • New Contributor
  • 1402 Views
  • 1 replies
  • 0 kudos

hosting-services-dedicated-shared-colocation-dubai-uae-middle-east-saudi-arabia-oman-qatar-bahrain-kuwait (1)

We offer the best web hosting solutions that are blazing fast, and ultra reliable & our sales & support team is here to help you find the right solutions

  • 1402 Views
  • 1 replies
  • 0 kudos
Latest Reply
mannubhai
New Contributor II
  • 0 kudos

Thanks for sharing this information. A reliable hosting provider is important for any service business website. We recently optimized the Kitchen Chimney Service Guwahati website for Mannubhai and found that fast hosting significantly improves page s...

  • 0 kudos
rikkyvai
by • New Contributor II
  • 301 Views
  • 0 replies
  • 1 kudos

from_utc_timestamp silently double-shifts time when session timezone isn't UTC — docs should call th

Summary:from_utc_timestamp(current_timestamp(), '<tz>') produces an incorrect (future-shifted) timestamp whenever the Spark session timezone is already set to something other than UTC. This is a very common pattern for teams stamping dp_load_ts/creat...

  • 301 Views
  • 0 replies
  • 1 kudos
ChiaMingHu
by • New Contributor II
  • 724 Views
  • 1 replies
  • 1 kudos

Resolved! Lakebase CDF — destination Delta table not created after successful UI setup (Free Edition)

 Summary:We configured Lakebase CDF to stream changes from a native Postgres table (public.query_logs) into Unity Catalog, but the destination Delta table never appears despite completing all documented prerequisites.Steps taken1. Created query_logs ...

IMG_2829.jpeg IMG_2828.jpeg
  • 724 Views
  • 1 replies
  • 1 kudos
Latest Reply
AbhilashNagilla
Databricks Employee
  • 1 kudos

Most likely this is a Free Edition storage limitation rather than a bug, and your Postgres-side setup sounds correct. Lakebase CDF writes its destination as a Unity Catalog managed Delta table, and the Lakebase CDF docs list catalogs backed by defaul...

  • 1 kudos
Klusener
by • Contributor
  • 15983 Views
  • 7 replies
  • 2 kudos

Relevance of off heap memory and usage

I was referring to the doc - https://kb.databricks.com/clusters/spark-executor-memory.In general total off heap memory is  =  spark.executor.memoryOverhead + spark.offHeap.size.  The off-heap mode is controlled by the properties spark.memory.offHeap....

  • 15983 Views
  • 7 replies
  • 2 kudos
Latest Reply
gm-dev
New Contributor II
  • 2 kudos

this is helpful

  • 2 kudos
6 More Replies
yit337
by • Contributor II
  • 1583 Views
  • 4 replies
  • 3 kudos

Is it required to run Lakeflow Connect on Serverless?

As the subject states, my question is:Is it required to run the Ingestion Pipeline in Lakeflow Connect on Serverless compute? Cause I try to define my own cluster in the DAB, but it raises an error:`Error: cannot create pipeline: You cannot provide c...

  • 1583 Views
  • 4 replies
  • 3 kudos
Latest Reply
saurabh18cs
Honored Contributor III
  • 3 kudos

Yes — Lakeflow Connect ingestion pipelines always run on Serverless compute. Databricks overrides your compute config and switches back to serverless,because the ingestion connector requires it.     

  • 3 kudos
3 More Replies
Labels