Shall we opt for multiple worker nodes in dab workflow template if our codebase is based on pandas.

harishgehlot
New Contributor III

Hi team, I am working in a databricks asset bundle architecture. Added my codebase repo in a workspace. My question to do we need to opt for multiple worker nodes like num_worker_nodes > 1 or autoscale with range of worker nodes if my codebase has mostly pandas integration and performing parallelization with joblib parallel. No integration of pyspark. 

Does it make sense to go with multiple nodes,or I am increasing my money for waste of idle nodes.

 

targets:
  dev_cluster: &dev_cluster
    new_cluster:
      cluster_log_conf:
        dbfs:
          destination: "dbfs:/FileStore/logs"
      spark_version: 14.3.x-scala2.12
      node_type_id: m5d.16xlarge
      custom_tags:
        clusterSource: forecasting
      data_security_mode: SINGLE_USER
      autotermination_minutes: 20
      autoscale:
        min_workers: 3
        max_workers: 20
      docker_image:
        url: "**************"
      aws_attributes:
        first_on_demand: 1
        instance_profile_arn: **************
        ebs_volume_type: GENERAL_PURPOSE_SSD
        ebs_volume_count: 1
        ebs_volume_size: 50