Setting up observability for serverless Databricks

APJESK
Contributor

I’m looking for best practices and guidance on setting up observability for serverless Databricks. Specifically, I’d like to know:

  • How to capture and monitor system-level metrics (CPU, memory, network, disk) in a serverless setup.

  • How to configure and collect application metrics (e.g., using Spark listeners, StreamingQueryListener, QueryExecutionListener).

  • Best way to manage and forward logs (driver logs, executor logs, audit logs, event logs) in a serverless environment.

  • Recommended approaches for integrating with external monitoring tools like Amazon CloudWatch, Datadog, or SIEM platforms.

  • Suggestions for building dashboards, alerts, and anomaly detection to ensure end-to-end observability.

If anyone has implemented observability for serverless Databricks workloads, I’d really appreciate insights into the architecture, tools used, and lessons learned.

Thanks in advance for your help!