Options
- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
05-26-2024 05:05 AM
As per the very short review session of the available source code and the SPIP itself, I think the answer is YES.
It is especially clear for spark.sql.sources.v2.bucketing.partiallyClusteredDistribution.enabled that says:
This is an optimization on skew join and can help to reduce data skewness when certain partitions are assigned large amount of data.