Raman_Unifeye
Honored Contributor III

@mordex - yes, Spark caps the parallelism for file listing at 200 tasks, regardless of whether you have 1,000 or 10,000 files. it is controlled by spark.sql.sources.parallelPartitionDiscovery.parallelism. 

Run below command to get value of it.
 
spark.conf.get('spark.sql.sources.parallelPartitionDiscovery.parallelism')
--200

RG #Driving Business Outcomes with Data Intelligence

View solution in original post