Options
- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
03-31-2022 07:39 AM
Hi, I'm doing some something simple on Databricks notebook:
spark.sparkContext.setCheckpointDir("/tmp/")
import pyspark.pandas as ps
sql=("""select
field1, field2
From table
Where date>='2021-01.01""")
df = ps.sql(sql)
df.spark.checkpoint()That runs great, saves the rdd on /mp/ then I want to save the df with
df.to_csv('/FileStore/tables/test.csv', index=False)or
df1.spark.coalesce(1).to_csv('/FileStore/tables/test.csv', index=False)And it recalculates the query again (it first did it on the checkpoint and then again to save the file).
What i'm doing wrong? currently, to solve this I'm saving the first dataframe without checkpoint, opening again and saving with coalesce.
If I use the coalesce(1) directly it doesn't parallelize.
EDIT:
Tried
df.spark.cache()But still reprocesses when I try to save to CSV, I'm looking to avoid reprocessing and avoid saving twice. Thanks!
the question is, why it recalculates df1 after the checkpoint?
Thanks!
Labels: