iyashk-DB
Databricks Employee
Databricks Employee

Hi @jeremy98 , collect() operation brings data to the driver and yes it can cause the memory issues that you are seeing, which can cause the cluster to be hung/ crash as well if done enough times. You may confirm these instances from the cluster event logs showing "Driver is up, but not responsive. Likely due to GC". 

Looking at the screenshots, you show us that there is memory pressure. So can you try reducing the GC interval from default 30 mins to 15mins. 

Please set configurations below in the Advanced options -> Spark
spark.cleaner.periodic.GC.interval 15min

These configurations will help optimizing the GC, and clear the GC objects after every 15 minutes instead of 30 minutes(default value).