I think I'm experiencing something similar.

Not using S3 yet. But reading Parquet tables into DataFrames, trying tactics like

persist
,
coalesce
,
repartition
after reading from Parquet. Using HiveContext, if that matters. But I get the impression that it's ignoring my attempts to repartition and cache and always recomputing my queries from scratch.

I'm definitely still new at this, so not sure yet how to figure out what's really going on.