Databricks Community

elgeo · ‎02-21-2023

Hello. Could someone please explain why iteration over a Pyspark dataframe is way slower than over a Pandas dataframe?

Pyspark

df_list = df.collect()

for index in range(0, len(df_list )):

.....

Pandas

df_pnd = df.toPandas()

for index, row in df_pnd.iterrows():

....

Thank you in advance

Anonymous · ‎04-22-2023

Hi @ELENI GEORGOUSI

Hope everything is going great.

Just wanted to check in if you were able to resolve your issue. If yes, would you be happy to mark an answer as best so that other members can find the solution more quickly? If not, please tell us so we can help you.

Cheers!

Databricks Community

Iteration - Pyspark vs Pandas

Join Us as a Local Community Builder!

What's new in Databricks: July - August 2025

How to use Lakebase as a transactional data layer for Databricks Apps

🌟 Community Spark of the Week | Aug 22 – Aug 28 🌟

EMEA Learning Festival: Hands-on Learning Journey!

Virtual Learning Festival: 10 October - 31 October 2025