Coffee77
Honored Contributor III

In addition to above cool comments, try to use clusters with VMs enabled for disk caching as well. This caches data at parquet files level in VM local storage, acting as a great complement to spark caching.


Lifelong Solution Architect Learner | Coffee & Data