Options
- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
05-18-2025 11:52 PM
Hi team, I am working in a databricks asset bundle architecture. Added my codebase repo in a workspace. My question to do we need to opt for multiple worker nodes like num_worker_nodes > 1 or autoscale with range of worker nodes if my codebase has mostly pandas integration and performing parallelization with joblib parallel. No integration of pyspark.
Does it make sense to go with multiple nodes,or I am increasing my money for waste of idle nodes.
targets:
dev_cluster: &dev_cluster
new_cluster:
cluster_log_conf:
dbfs:
destination: "dbfs:/FileStore/logs"
spark_version: 14.3.x-scala2.12
node_type_id: m5d.16xlarge
custom_tags:
clusterSource: forecasting
data_security_mode: SINGLE_USER
autotermination_minutes: 20
autoscale:
min_workers: 3
max_workers: 20
docker_image:
url: "**************"
aws_attributes:
first_on_demand: 1
instance_profile_arn: **************
ebs_volume_type: GENERAL_PURPOSE_SSD
ebs_volume_count: 1
ebs_volume_size: 50