Monday
I'm having problems with using jobs together with GIT.
I have created a job that has
Type - Notebook
Source - Git provider (main),
Path "jobs/base/INIT_PROCESSING"
I can access the jupyter notebook without any problems. Each push to the repo can be read in the job run.
The problem is that within that jupyter notebook there are import from other py files that are stored on the repo.
Job pulls only the jupyter file and when importing libraries I get an error:
ModuleNotFoundError: No module named 'experiment'
How to fix this? Is it possible to pull a whole folder from the repo and use it in the job?
One way of solving this is to have git folder inside workspace and refer to that folder in the job. But the problem is that multiple people are working on that repo and it difficult to pull after each code change.
here is my job yaml:
Monday
Hi @Innuendo84
The problem is because you're using the Source: Git provider on a Databricks Job task that only executes a single notebook. The job runner will not clone the entire repository, only fetch that one notebook file (the one you're executing). Consequently, any relative python files or modules (like experiment.py) will be unavailable to the task.
There are two native ways to handle this.
Solution 1: Databricks Git Folder / Repos (recommended)
Instead of fetching from an external Git Provider url inside your job task, point it to a Databricks Git Folder (also known as Databricks Repos) that contains a full clone of your repository inside the workspace.
In the Databricks workspace UI, go to Git folders (or Repos) and clone your desired Git repository there.
Update the Job Task Source:
Edit your Job Task settings, change Source from Git provider to Workspace, and browse the notebook path from your Git folder: /Workspace/Repos/user@company.com/my-repo/jobs/base/INIT_PROCESSING
When you run a notebook from a Workspace Git folder, all local imports (from experiment import ...) should work as expected -- Databricks automatically adds the root of the repository to your python sys.path.
For CI/CD: to make sure your job always pulls the latest code from main, use the Databricks Repos API or Databricks CLI in your CI/CD pipeline (e.g. GitHub Actions) to run databricks repos update on push to main.
Solution 2: Manually add repo root to sys.path
If you want to keep Source: Git provider on your Job task (so that Databricks fetches the git commit on-demand), then you'll need to tell python where to find your local modules. Add this snippet to the very first cell of your INIT_PROCESSING notebook:
Python
import sys
import os
# Get the current working directory (where Databricks mounted the notebook)
current_dir = os.getcwd()
# Navigate up from 'jobs/base' to the root of the repo
repo_root = os.path.abspath(os.path.join(current_dir, "../../"))
# Add the repo root to the python path to find 'experiment'
if repo_root not in sys.path:
sys.path.insert(0, repo_root)
# Test the import
import experiment
Monday
I was looking at your YAML and I think the sparse_checkout might be the main issue here.
Right now you're only checking out:
sparse_checkout:
patterns:
- jobs/base
So if experiment lives somewhere else in the repo, that directory probably isn't being included in the checkout.
I'd probably try removing sparse_checkout first and let the job check out the full repo. With source: GIT, Databricks should use the remote repo directly for the run, so you shouldn't need to keep a Workspace Git folder manually updated.
If you do want to keep sparse checkout, then I'd include the folders that contain the Python modules as well, for example:
sparse_checkout:
patterns:
- jobs/base
- src
Then I'd check whether the directory containing experiment is actually on the Python path.
Given the config you shared, I'd start with the sparse checkout because that looks like the most likely cause.
Monday
py files are directly next to the jupyter files (inside base folder).
- adding syspath does not work. py files are not there even if I add cwd.
- sparse_checkout is not the problem. I tried numerous time with and without it and it never worked. besides py files in the same folder and notebook.
Monday
I did a quick validation with a git-backed job using source: GIT.
In my test Databricks checked out the repository at runtime under /Workspace/Repos/.internal/.., not just the notebook file. The repo root was automatically present on sys.path, so imports from the repo worked without manually modifying sys.path.
For The folder structure with:
repo/
โโโ jobs/base/INIT_PROCESSING.py
โโโ src/experiment.pythis had worked fine
from src import experimentCheck once the actual repo/package layout and import statement first. If experiment.py is under another folder, import experiment may fail even though the full Git repo is checked out.
Also debug and check the by adding below before import experiment statement in the notebook:
import os, sys
print(os.getcwd())
print(sys.path)Also verify the module exists in the exact branch/commit used by the job. The git_branch you posted in the job yaml is pointing to main branch. So, may be check if the module/files exist in main branch. I have used feature branch name for git_branch in jobs yaml while testing as I don't have the files yet in main.
I also verified sparse_checkout: even with only jobs/base in the pattern, the module under src was still available in my run, so I wouldnโt assume sparse checkout is the cause without first checking the actual runtime checkout and import path.
Monday
I think there might be another thing worth checking before adding the extra repo pull task.
In the job YAML you posted, the Git source is using:
git_branch: main
but in the code you posted later you're explicitly pulling:
BRANCH = "develop"
So I'd first check whether experiment.py actually exists in the exact main commit being used by the job. Databricks takes a snapshot of the configured branch when the run starts, and you can see the commit SHA used by the run in the job details.
I'd probably add something simple before the import just to verify what the job can actually see:
import os
print(os.getcwd())
print(os.listdir(os.getcwd()))
If experiment.py is really next to INIT_PROCESSING in that exact commit, then I'd expect it to show up there, and that would narrow this down quite a bit.
Given that your manual repo update is targeting develop while the job is targeting main, I'd check that first before introducing another task or the Repos API.
Monday - last edited Tuesday
I have been testing several things and here are my findings:
-1-
-2- I tried runing jobs that access jupyter notebook and py directly. And I can access jupyter, but not the py files.
Tuesday
I did a quick validation with a Git-backed spark_python_task, and it worked as expected when I used the full Python filename:
spark_python_task:
python_file: jobs/base/experiment.py
source: GITIn the run, Databricks checked out the repo under /Workspace/Repos/.internal/..., and both the task folder and repo root were automatically present on sys.path.
One thing I noticed in your example is:
python_file: jobs/base/experimentchange that to:
python_file: jobs/base/experiment.pyFor spark_python_task, python_file should point to the actual .py file. notebook_task.notebook_path behaves differently and commonly ignores the notebook extension.
Tuesday
Why it fails
With sparse_checkout.patterns: [jobs/base], Databricks clones only that folder. The experiment package (and anything else outside jobs/base) is never checked out, so Python can't find it. Without sparse checkout, a Git-sourced job clones the whole repo into a temporary location on every run, so you don't need a workspace Git folder or manual pulls.
Fix 1: Include the package in the sparse checkout (keeps the checkout small)
Add the folder that contains experiment to the patterns:
yaml
git_source:
git_url: https://XXXXXXXXXXXXXXXXXXXXXXXX.git
git_provider: bitbucketServer
git_branch: main
sparse_checkout:
patterns:
- jobs/base
- experiment # adjust to the real path of the package in the repo
Fix 2: Remove sparse checkout (simplest)
Delete the sparse_checkout block and the whole repo is cloned on each run. Fine unless the repo is huge.
Make sure the import path resolves
Cloning the files isn't always enough. For notebooks run from Git source, the working directory is the notebook's folder (jobs/base), not the repo root, so import experiment only works if experiment is on sys.path. Two options:
Option A: add the repo root to sys.path at the top of the notebook
python
import sys, os
# notebook lives in <repo_root>/jobs/base, so go up two levels
repo_root = os.path.abspath(os.path.join(os.getcwd(), "..", ".."))
if repo_root not in sys.path:
sys.path.insert(0, repo_root)
import experiment
If experiment sits somewhere else (e.g. src/experiment), append that parent folder instead.
Option B: package it and install it as a library (cleaner for shared code)
Build a wheel from the repo (or pip install from the Bitbucket URL) and attach it to the task:
yaml
tasks:
- task_key: init_processing
notebook_task:
notebook_path: jobs/base/INIT_PROCESSING
source: GIT
existing_cluster_id: XXXX-XXXXXXXXXXX-XXXXXXXX
libraries:
- whl: /Volumes/<catalog>/<schema>/<volume>/experiment-0.1-py3-none-any.whl
Tuesday
I believe this could be related to the sparse_checkout configuration. If the notebook is importing modules from directories outside the `jobs/base` path, those files may not be available in the jobโs Git checkout.
As a first step, I would suggest either removing sparse_checkout to test with the complete repository, or adding the required directories to the sparse checkout patterns. Once the required modules are included in the checkout, the imports should be available to the job.
If the files are already present in the checkout and the issue still persists, it may be worth checking the Python import path (sys.path) as the next step.
Additionally, for third-party dependencies, you can configure the required Python libraries directly as Job/cluster libraries. These can be installed as part of the job cluster setup rather than installing them at runtime from within the notebook. This can help keep the environment consistent across job runs.
However, if experiment is a custom module from the Git repository rather than a third-party package, the Git checkout and Python import path would still need to be addressed.
yesterday
Most likely, your .py files are being treated as notebooks rather than plain Python files, so they can't be imported and can't be read as a python_file. The job is not failing to pull the folder. Your own findings show the whole commit is checked out.
sys.path contains .../_commits/<sha>/jobs/base, so the job has checked out the repo at that commit and put the notebook's directory on the path. You don't need to pull the repo again or add a Git folder to the workspace.glob("/jobs/base") test is misleading. The leading / points at the filesystem root, not the checkout. Use this instead:import os
print(os.getcwd())
print(os.listdir(os.getcwd()))
spark_python_task error says it can't read .../jobs/base/experiment.py. That is the usual symptom when a .py file is stored as a notebook object rather than a workspace file. Notebook objects can't be imported with import experiment, and they can't be run as a python_file either.Open experiment.py and gt_resampler.py and look at the first line. If it's this:
# Databricks notebook source
then Databricks treats the file as a source-format notebook. Delete that line and any # COMMAND ---------- separators, commit, and push. After that the file is a plain Python module, and import experiment and import gt_resampler in INIT_PROCESSING should work. The notebook's own folder is already on sys.path, as you saw. Only the notebooks you actually want to run as notebooks should keep that header.
If the header is not there, print os.listdir(os.getcwd()) from inside the job run and compare it to the repo. That tells you whether the files are really present in the checkout.
spark_python_task test tooFor a Git-sourced spark_python_task, give the path relative to the repo root, including the extension:
spark_python_task:
python_file: jobs/base/experiment.py
source: GIT
That approach has two problems:
/api/2.0/repos call uses the Git credential of the identity running the job. Your run_as user has no Git credential configured for that Bitbucket Server host in this context, so the call fails. Git credentials are configured per user or service principal, as described in the Git folders configuration docs [1].git_source already gives you what you want: each run is pinned to a specific commit, and your colleagues' work in their own Git folders never touches it.source: GIT. This works well for repos with many shared modules.For now, strip the notebook header from the helper .py files and print os.getcwd() and os.listdir() in the run. I expect that fixes the ModuleNotFoundError without any other changes to the job.
[1] Databricks Git folders | Databricks on AWS โ https://docs.databricks.com/aws/en/repos