Join discussions on data engineering best practices, architectures, and optimization strategies within the Databricks Community. Exchange insights and solutions with fellow data engineers.
Hello. I am trying to understand High Availability in DataBricks. I understand that DB uses Kubernetes for the cluster manager and to manage Docker Containers. And while DB runs on top of AWS or Azure or GCP, is HA automatically provisioned when I start a cluster because of Kubernetes or Does it use ZooKeeper? In another scenario, if I configure the cluster mode to Standard, I only have one Driver (Master Node) which can be a single point of failure. Will Kubernetes take care of it if it fails or does it make sense to start up ZooKeeper to maintain a Quorum? Or does Kubernetes by default use ZooKeeper to watch ALL nodes? I hope this makes sense.
Thanks so much Kaniz. Ultimately, I'm looking for any architecture details on how DB configures Kubernetes to manage nodes. Especially the Driver (Master Node). It's clear that Apache Spark can be configured to use ZooKeeper but I don't find anything regarding DataBricks. Thanks for your time.
Join a Regional User Group to connect with local Databricks users. Events will be happening in your city, and you won’t want to miss the chance to attend and share knowledge.
If there isn’t a group near you, start one and help create a community that brings people together.