cancel
Showing results for 
Search instead for 
Did you mean: 
Data Engineering
Join discussions on data engineering best practices, architectures, and optimization strategies within the Databricks Community. Exchange insights and solutions with fellow data engineers.
cancel
Showing results for 
Search instead for 
Did you mean: 

Forum Posts

Anoora
by New Contributor II
  • 1269 Views
  • 2 replies
  • 0 kudos

Scheduling and triggering jobs based on time and frequency precedence

I have a table in Databricks that stores job information, including fields such as job_name, job_id, frequency, scheduled_time, and last_run_time.I want to run a query every 10 minutes that checks this table and triggers a job if the scheduled_time i...

Data Engineering
data engineering
jobs
scheduling
  • 1269 Views
  • 2 replies
  • 0 kudos
Latest Reply
SamAdams
Contributor
  • 0 kudos

You could add a job with a scheduled based trigger that runs every 10 minutes. The task at the start of the job runs a SQL query against the job information table and uses the logic you described above to output a boolean value. Then feed that boolea...

  • 0 kudos
1 More Replies
EricCournarie
by New Contributor III
  • 845 Views
  • 2 replies
  • 2 kudos

Retrieving OBJECT values with the JDBC driver may lead to invalid JSON

Hello,Using the JDBC driver , I try to retrieve values in the ResultSet for a OBJECT type. Sadly, it returns invalid JSONGiven the SQLCREATE OR REPLACE TABLE main.eric.eric_complex_team (`id` INT,`nom` STRING,`infos` STRUCT<`age`: INT, `ville`: STRIN...

  • 845 Views
  • 2 replies
  • 2 kudos
Latest Reply
EricCournarie
New Contributor III
  • 2 kudos

Hello,  thanks for the quick response .Sadly I do not have the hand on the SQL request , so no way for me to modify it ... 

  • 2 kudos
1 More Replies
DylanStout
by Contributor
  • 4635 Views
  • 1 replies
  • 0 kudos

Pyspark ML tools

Cluster policies not letting us use Pyspark ML toolsIssue details: We have clusters available in our Databricks environment and our plan was to use functions and classes from "pyspark.ml" to process data and train our model in parallel across cores/n...

  • 4635 Views
  • 1 replies
  • 0 kudos
Latest Reply
Louis_Frolio
Databricks Employee
  • 0 kudos

Hey @DylanStout ,   Thanks for laying out the symptoms clearly—this is a classic clash between Safe Spark (shared/high-concurrency) protections and multi-threaded/driver-mutating code paths.   What’s happening On clusters with the Shared/Safe Spark a...

  • 0 kudos
akeel-rehman
by New Contributor
  • 4186 Views
  • 1 replies
  • 0 kudos

Best Practices for Reusable Workflows & Cluster Management Across Repos.

Hi everyone,I am looking for best practices around reusable workflows in Databricks, particularly in these areas:Reusable Workflows Instead of Repetition: How can we define reusable workflows rather than repeating the same steps across multiple jobs?...

  • 4186 Views
  • 1 replies
  • 0 kudos
Latest Reply
AbhaySingh
Databricks Employee
  • 0 kudos

Here are my recommendations: 1. Databricks Asset Bundles (DABs) for reusable workflows   2. API-based triggering and Run Job Tasks for cross-repo workflows   3. Instance Pools as the #1 game-changer for cluster optimization (5-10 seconds vs 5-10 minu...

  • 0 kudos
jorperort
by Contributor
  • 5656 Views
  • 5 replies
  • 4 kudos

Resolved! Spark JDBC Write Fails for Record Not Present - PK error

Good afternoon everyone,I’m writing this post to see if anyone has encountered this problem and if there is a way to resolve it or understand why it happens. I’m working in a Databricks Runtime 15.4 LTS environment, which includes Apache Spark 3.5.0 ...

  • 5656 Views
  • 5 replies
  • 4 kudos
Latest Reply
ManojkMohan
Honored Contributor II
  • 4 kudos

@jorperort  When writing to SQL Server tables with composite primary keys from Databricks using JDBC, unique constraint violations are often caused by Spark’s distributed retry logic  https://docs.databricks.com/aws/en/archive/connectors/jdbcSolution...

  • 4 kudos
4 More Replies
Dimitry
by Valued Contributor
  • 1072 Views
  • 4 replies
  • 0 kudos

databricks notebook parameter works in interactive mode but not in the job

Hi guys I've added a parameter "files_mask " to a notebook, with a default value.The job running this notebook broke with error: com.databricks.dbutils_v1.InputWidgetNotDefined: No input widget named files_mask is definedCode: mask = dbutils.widgets....

  • 1072 Views
  • 4 replies
  • 0 kudos
Latest Reply
szymon_dybczak
Esteemed Contributor III
  • 0 kudos

Hi @Dimitry ,Do you use python or scala in your notebook?

  • 0 kudos
3 More Replies
ashraf1395
by Honored Contributor
  • 5485 Views
  • 4 replies
  • 1 kudos

Resolved! How to capture dlt pipeline id / name using dynamic value reference

Hi there,I have a usecase where I want to set the dlt pipeline id in the configuration parameters of that dlt pipeline.The way we can use workspace ids or task id in notebook task task_id = {{task.id}}/ {{task.name}} and can save them as parameters a...

  • 5485 Views
  • 4 replies
  • 1 kudos
Latest Reply
CaptainJack
New Contributor III
  • 1 kudos

Did someone was able to get pipeline_id programaticaly?

  • 1 kudos
3 More Replies
shadowinc
by New Contributor III
  • 4659 Views
  • 1 replies
  • 1 kudos

Call SQL Function via API

Background - I created a SQL function with the name schema.function_name, which returns a table, in a notebook, the function works perfectly, however, I want to execute it via API using SQL Endpoint. In API, I got insufficient privileges error, so gr...

  • 4659 Views
  • 1 replies
  • 1 kudos
Latest Reply
AbhaySingh
Databricks Employee
  • 1 kudos

Do you know if API service principal / user has USAGE on the database itself? This seems like the most likely issue based on information on the question.  Quick Fix Checklist:   Run these commands in order (replace api_user with the actual user from ...

  • 1 kudos
Pw76
by New Contributor III
  • 4506 Views
  • 4 replies
  • 3 kudos

CDC with Snapshot - next_snapshot_and_version() function

I am trying to use create_auto_cdc_from_snapshot_flow (formerly apply_changes_from_snapshot())  (see: https://docs.databricks.com/aws/en/dlt/cdc#cdc-from-snapshot)I am attempting to do SCD type 2 changes using historic snapshot data.In the first coup...

Data Engineering
CDC
dlt
Snapshot
  • 4506 Views
  • 4 replies
  • 3 kudos
Latest Reply
fabdsp
New Contributor II
  • 3 kudos

I have the same issue - very big limitation of create_auto_cdc_from_snapshot_flow and no solution

  • 3 kudos
3 More Replies
jeremy98
by Honored Contributor
  • 2596 Views
  • 3 replies
  • 1 kudos

how to pass secrets keys using a spark_python_task

Hello community,I was searching a way to pass secrets to spark_python_task. Using a notebook file is easy, it's only to use dbutils.secrets.get(...) but how to do the same thing using a spark_python_task set using serveless compute?Kind regards,

  • 2596 Views
  • 3 replies
  • 1 kudos
Latest Reply
analytics_eng
New Contributor III
  • 1 kudos

@Renu_  but passing them as spark_env will not work with serverless I guess? See also the limitations on the docs  Serverless compute limitations | Databricks on AWS 

  • 1 kudos
2 More Replies
dpc
by Contributor III
  • 2204 Views
  • 5 replies
  • 3 kudos

Resolved! Pass parameters between jobs

Hello I have jobIn that job, it runs a task (GetGid) that executes a notebook and obtains some value using dbutils.jobs.taskValuesSete.g. dbutils.jobs.taskValuesSet(key = "gid", value = gid)As a result, I can use this and pass it to another task for ...

  • 2204 Views
  • 5 replies
  • 3 kudos
Latest Reply
dpc
Contributor III
  • 3 kudos

Thanks @Hubert-Dudek and @ilir_nuredini I see this nowI'm setting using:dbutils.jobs.taskValues.Set()passing to the job task using Key - gid; Value - {{tasks.GetGid.values.gid}}Then reading using: pid = dbutils.widgets.get()

  • 3 kudos
4 More Replies
AlleyCat
by New Contributor II
  • 2089 Views
  • 3 replies
  • 0 kudos

To identify deleted Runs in Workflow.Job UI in "system.lakeflow"

Hi,I executed a few runs in a Workflow.Jobs UI. I then deleted some of them. I am seeing the deleted runs in "system.lakeflow.job_run_timeline". How do i know which runs are the deleted ones? Thanks

  • 2089 Views
  • 3 replies
  • 0 kudos
Latest Reply
Ayushi_Suthar
Databricks Employee
  • 0 kudos

Hi @AlleyCat , Hope you are doing well!  The jobs table includes a delete_time column that records the time when the job was deleted by the user. So to identify deleted jobs, you can run a query like the following: SELECT * FROM system.lakeflow.jobs ...

  • 0 kudos
2 More Replies
DM0341
by New Contributor II
  • 1162 Views
  • 2 replies
  • 1 kudos

Resolved! SQL Stored Procedures - Notebook to always run the CREATE query

I have a stored procedure that is saved as a query file. I can run it and the proc is created. However I want to take this one step further. I want my notebook to run the query file called sp_Remit.sql so if there is any changes to the proc between t...

  • 1162 Views
  • 2 replies
  • 1 kudos
Latest Reply
DM0341
New Contributor II
  • 1 kudos

Thank you. I did find this about an hour after I posted. Thank you Kevin

  • 1 kudos
1 More Replies
SuMiT1
by New Contributor III
  • 1773 Views
  • 1 replies
  • 1 kudos

Databricks to snowflake data load

Hi Team, I’m trying to load data from Databricks into Snowflake using the Snowflake Spark connector. I’m using a generic username and password, but I’m unable to log in using these credentials directly. In the Snowflake UI, I can only log in through ...

  • 1773 Views
  • 1 replies
  • 1 kudos
Latest Reply
nayan_wylde
Esteemed Contributor II
  • 1 kudos

@SuMiT1  The recommended method to connect to snowflake from databricks is OAuth with Client Credentials Flow.This method uses a registered Azure AD application to obtain an OAuth token without user interaction.Steps:Register an app in Azure AD and c...

  • 1 kudos
StephanieAlba
by Databricks Employee
  • 3995 Views
  • 2 replies
  • 0 kudos

Is it possible to turn off the redaction of secrets? Is there a better way to solve this?

As part of our Azure Data Factory pipeline, we utilize Databricks to run some scripts that identify which files we need to load from a certain source. This list of files is then passed back into Azure Data Factory utilizing the Exit status from the n...

  • 3995 Views
  • 2 replies
  • 0 kudos
Latest Reply
joanafloresc
New Contributor II
  • 0 kudos

Hello, as of today, is it still not possible to unredact secret names?

  • 0 kudos
1 More Replies
Labels