balajij8
Esteemed Contributor II

You can start with tracking system health, determine root cause & anomalies based on info from tables & logs. I would start with the blog

https://community.databricks.com/t5/technical-blog/databricks-observability-using-grafana-and-promet... 

Add info from the system logs to these