Setting environment variables to use in a SQL Delta Live Table Pipeline

ac0
Contributor

I'm trying to use the Global Init Scripts in Databricks to set an environment variable to use in a Delta Live Table Pipeline. I want to be able to reference a value passed in as a path versus hard coding it. Here is the code for my pipeline:

CREATE STREAMING LIVE TABLE data
COMMENT "Raw data in delta format"
TBLPROPERTIES ("quality" = "bronze", "pipelines.autoOptimize.zOrderCols" = "id")
AS
SELECT *, id FROM cloud_files(
  "${TEST_VAR}data/files", "json", map("cloudFiles.inferColumnTypes", "true")
  )

However, when I set up a global init script like below, it doesn't appear on the list of environment variables on the job compute cluster.

#!/bin/sh

sudo echo TEST_VAR=TESTING >> /etc/environment

Is this because the cluster type is PIPELINE? Is what I am attempting to do possible? Are Global Init Scripts even run when using Delta Live Table Pipelines, can can environment variables be referenced in SQL-style pipelines? I am finding little documentation about this online.

Debayan
Databricks Employee
Databricks Employee

Hi, Could you please check if you are using shared access cluster mode? 

https://docs.databricks.com/en/init-scripts/index.html

Priyanka_Biswas
Databricks Employee
Databricks Employee

ac0
Contributor

I was able to accomplish this by creating a Cluster Policy that put in place the scripts, config settings, and environment variables I needed.

View solution in original post