User16766737456
Databricks Employee
Databricks Employee

Just an update, to round this out.

We investigated further internally, and found that although we have a cleanup process in place to remove the internal repos that are being checked out for workflows, it was failing to catch up due to the sheer volume of jobs that were continuously failing during the repo checkout step (because of an invalid path).

This led to the limits being breached, and cascaded down to valid jobs not being able to launch.

We've worked with Kit to identify the errant job(s), and are now closely monitoring internal metrics, which currently show significant improvements.

View solution in original post