Gopal269673
Contributor

@All Users Group​  Hi All.. we had tried several options of tuning the query by selecting required variables in the select and subsequent clauses. I see other queries are little ok to run. But the attached query seems not able to run from past 6 hours with 8 worker nodes config. I see that spill is high and attached the metrics for it. Anyone can suggest optimization techniques in python note book for looking into it as I am getting only scala related programs. Please help in optimization best methods guide and material more specific to Pyspark & Sql.

View solution in original post