-werners-
Esteemed Contributor III

Her you can find some info.
Basically it boils down to: reading a file = overhead.
You want to minimize overhead without stressing the workers too much (gigantic partitions).

https://medium.com/globant/how-to-solve-a-large-number-of-small-files-problem-in-spark-21f819eb36d3