I already tried that approach and got the same error. Should I consider increasing the size and cores of my cluster? Or should I be able to achieve my goal with this settings?

EDIT: Also when running your code i get the AttributeError: 'DataFrame' object has no attribute 'partition_id'