cancel
Showing results for 
Search instead for 
Did you mean: 
Data Engineering
Join discussions on data engineering best practices, architectures, and optimization strategies within the Databricks Community. Exchange insights and solutions with fellow data engineers.
cancel
Showing results for 
Search instead for 
Did you mean: 

Forum Posts

RIDBX
by • Contributor
  • 1575 Views
  • 4 replies
  • 1 kudos

Robust/complex scheduling with dependency within Databricks?

Robust scheduling with dependency within Databricks?======================================  Thanks for reviewing my threads. I like to explore Robust/complex scheduling with dependency within Databricks.We know traditional scheduling framework allow ...

  • 1575 Views
  • 4 replies
  • 1 kudos
Latest Reply
alexjames3123
New Contributor II
  • 1 kudos

Databricks workflows can handle complex dependencies by using task conditions and job dependencies, allowing Finance jobs to run after the required HR jobs finish. A clear guide can also help with organizing broader business processes.

  • 1 kudos
3 More Replies
Ashoka
by • New Contributor III
  • 758 Views
  • 2 replies
  • 2 kudos

Resolved! Title: Oracle CDC pipeline (Lakeflow Connect) never terminates

Setup: Oracle 21c XE (CDB/PDB), Direct CDC Extraction, no gateway, triggered pipeline on hourly schedule, single table.Problem: A triggered run has been active 95+ minutes with no work left. All flows report IDLE, waiting for new data, the source is ...

  • 758 Views
  • 2 replies
  • 2 kudos
Latest Reply
Ashoka
New Contributor III
  • 2 kudos

I am using databrick-oracle CDC connector( beta version)

  • 2 kudos
1 More Replies
margarita_shir
by • New Contributor III
  • 529 Views
  • 2 replies
  • 0 kudos

Lakehouse Federation (Snowflake) — large query results fail to download from internal stage

We're hitting a consistent failure with Lakehouse Federation to Snowflake where large result sets fail during the result-chunk download, while small results work. Looking for help isolating whether this is the Databricks Snowflake connector or a conf...

  • 529 Views
  • 2 replies
  • 0 kudos
Latest Reply
arhamblake38
New Contributor II
  • 0 kudos

This sounds like it could be related to the size of the result set or the way the internal stage handles large downloads, rather than the federation itself. If smaller query results download successfully while larger ones fail, I’d compare the file s...

  • 0 kudos
1 More Replies
Deny1
by • New Contributor II
  • 431 Views
  • 1 replies
  • 1 kudos

cross-region DR in Azure Databricks (24h RPO/RTO)

Hi,We're designing a DR strategy for an Azure Databricks platform and would appreciate guidance on current best practices for achieving approximately 24-hour RPO and RTO across Azure regions.Our platform includes Unity Catalog, DAB, Jobs, Notebooks, ...

  • 431 Views
  • 1 replies
  • 1 kudos
Latest Reply
pradeep_singh
Honored Contributor III
  • 1 kudos

Your overall design is sensible, but Databricks’ current recommendation is to use Managed Disaster Recovery when your account is eligible. If Managed DR is unavailable, use an active-passive warm standby with Terraform/DABs, cross-region data replica...

  • 1 kudos
ThiamLee
by • Contributor
  • 330 Views
  • 2 replies
  • 3 kudos

Suggestions for recommendation

Hey everyone! I’m currently looking for some good tools for data analysis. If you have any recommendations, please share them with me. Would really appreciate it!

  • 330 Views
  • 2 replies
  • 3 kudos
Latest Reply
balajij8
Esteemed Contributor II
  • 3 kudos

@ThiamLee You can use Notebooks that let you write Python and SQL code interactively. You use libraries like pandas, matplotlib, ydata profiling and Py Spark for both small and large-scale analysis in it. Databricks SQL provides a dedicated SQL edito...

  • 3 kudos
1 More Replies
karuppusamy
by • Databricks Partner
  • 968 Views
  • 4 replies
  • 2 kudos

Oracle CDC Ingestion Pipeline - Schema Exploration finds 0 tables when DB_DOMAIN

Hi,I'm setting up an Oracle Integrated CDC Managed Ingestion Pipeline (LakeFlow Connect) on Azure Databricks. The pipeline connects to Oracle 19c successfully but Schema Exploration always returns 0 tables.Error:Schema Exploration has COMPLETED in 5 ...

  • 968 Views
  • 4 replies
  • 2 kudos
Latest Reply
VTiw25
New Contributor II
  • 2 kudos

@karuppusamy Try using ORCL as the service name.

  • 2 kudos
3 More Replies
FranPérez
by • New Contributor III
  • 20792 Views
  • 10 replies
  • 7 kudos

set PYTHONPATH when executing workflows

I set up a workflow using 2 tasks. Just for demo purposes, I'm using an interactive cluster for running the workflow. { "task_key": "prepare", "spark_python_task": { "python_file": "file...

  • 20792 Views
  • 10 replies
  • 7 kudos
Latest Reply
Kamil128
New Contributor II
  • 7 kudos

We’re facing the same issue while trying to implement Databricks Asset Bundles (DAB).We’ve checked multiple possibilities, and it looks like packaging the code as a wheel or using `sys.path.append` are the only viable options.It looks like Databricks...

  • 7 kudos
9 More Replies
de01
by • Databricks Partner
  • 769 Views
  • 3 replies
  • 5 kudos

Resolved! Auto SCD API Tombstone Garbage Collection

Are there any settings that can be used to influence the frequency at which the auto SCD API runs the tombstone garbage collection process in Spark Declarative Pipelines?  I've seen a couple community posts that referenced the following:  pipelines.a...

  • 769 Views
  • 3 replies
  • 5 kudos
Latest Reply
de01
Databricks Partner
  • 5 kudos

I was able to get this figured out.  The setting that works is pipelines.applyChanges.tombstoneGCFrequencyInSeconds, and it seems like it can be set as either a pipeline configuration or in the spark conf of the individual streaming table(either in c...

  • 5 kudos
2 More Replies
analyticsnerd
by • New Contributor III
  • 1020 Views
  • 2 replies
  • 2 kudos

Resolved! Transaction log integrity issue

Delta transaction log for one of our tables which is being written to by a Kafka Connect IcebergSinkConnector via the Unity Catalog Iceberg REST endpoint, is currently corrupted and is failing when trying to read with the below exceptionERROR:com.dat...

  • 1020 Views
  • 2 replies
  • 2 kudos
Latest Reply
K_Anudeep
Databricks Employee
  • 2 kudos

Hello @analyticsnerd , Delta maintains a checksum file (.crc) for each committed version, along with the commit log (.json), recording the table’s file count, total size, and a distribution of file sizes (a histogram that buckets files by size). For ...

  • 2 kudos
1 More Replies
ChristianRRL
by • Honored Contributor II
  • 624 Views
  • 4 replies
  • 6 kudos

Declarative Automation Bundle - How to handle Managed resources?

Hi there,I'm making a new repo from scratch to make it as DAB-compatible as possible. I was considering using DAB for the creation/management of resources such as schemas and volumes.. however, a concern crossed my mind. If these resources are MANAGE...

ChristianRRL_0-1787603890125.png
  • 624 Views
  • 4 replies
  • 6 kudos
Latest Reply
mancy34
New Contributor III
  • 6 kudos

Good point managed schemas/volumes need extra caution with bundle destroy. I’d avoid managing production data resources directly through DABs, or use safeguards to prevent accidental deletion.

  • 6 kudos
3 More Replies
DB1To3
by • Contributor II
  • 949 Views
  • 7 replies
  • 3 kudos

Distinguishing Runs Related to an All-Purpose Cluster

Sorry for the basic question.  If I am using an all-purpose clusters in Databricks, it seems to combine all my runs together and execute them within the same Apache Spark application.  There is only one application shown in the Spark UI.   Unrelated ...

  • 949 Views
  • 7 replies
  • 3 kudos
Latest Reply
Khasim_1
New Contributor III
  • 3 kudos

If you must use an All-Purpose cluster (e.g., for cost savings or fast start times), here is how you manage the chaos:The "Job Name" filter in the Jobs UI: Don't look at the Spark UI first. Go to the Workflows > Jobs tab. You can filter the "Runs" li...

  • 3 kudos
6 More Replies
pepco
by • New Contributor III
  • 779 Views
  • 2 replies
  • 2 kudos

Resolved! databricks SQL UDF in select statement

In the Unity Catalog we can now create/register SQL UDFs. There are two types - one that returns table and other that returns just a value. If the function that returns value is based on the SQL query and joins it would in standard relational databas...

  • 779 Views
  • 2 replies
  • 2 kudos
Latest Reply
AbhilashNagilla
Databricks Employee
  • 2 kudos

The earlier reply has the mechanism broadly right, and your two questions have clean answers: yes to the first, no to the second. A SQL scalar function whose body is a query is planned as a scalar subquery inside the calling statement. Databricks lab...

  • 2 kudos
1 More Replies
17780
by • New Contributor II
  • 35884 Views
  • 7 replies
  • 0 kudos

How to delete Databricks Account

I created and used a Databricks Account for testing purposes. I want to delete that account. In the Databricks Account Web UI, there is no menu to delete an account. How should I delete it?

  • 35884 Views
  • 7 replies
  • 0 kudos
Latest Reply
ss9808
New Contributor II
  • 0 kudos

tried to delete the account but could not find how to delete it. is this intentional? why is there no delete account button?anyways delete my entire account please, as for why - becasue im not interested in the service

  • 0 kudos
6 More Replies
isaac_gritz
by • Databricks Employee
  • 3187 Views
  • 1 replies
  • 2 kudos

Connecting Applications and BI Tools to Databricks SQL

Access Data in Databricks Using an Application or your Favorite BI ToolYou can leverage Partner Connect for easy, low-configuration connections to some of the most popular BI tools through our optimized connectors. Alternatively, you can follow these...

  • 3187 Views
  • 1 replies
  • 2 kudos
Latest Reply
ManeeshJha
New Contributor II
  • 2 kudos

Great summary, Isaac. One thing I'd add from implementation experience is choosing the connector based on the consumption pattern, not just the BI tool.Partner Connect is perfect for quick PoCs and for analysts who just need Power BI / Tableau to wor...

  • 2 kudos
Kayla
by • Valued Contributor II
  • 1023 Views
  • 5 replies
  • 1 kudos

Disable spark transformations inside try except warning

I'm wondering if there's a way to suppress certain errors? I have a lot of spammy yellow underlining from this one warning, and it doesn't seem that the warning behaves properly. Transformation inside try/except block is lazy, fine, I have actions on...

Kayla_0-1785436282119.png
  • 1023 Views
  • 5 replies
  • 1 kudos
Latest Reply
MartFaber
New Contributor II
  • 1 kudos

Hi @Kayla,I was wondering if this has already been patched in some way, as I'm facing the exact same error.Kind regards, Mart Faber

  • 1 kudos
4 More Replies
Labels