Options
- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
07-24-2015 10:22 AM
Hi,
There are a couple of SQL optimizations I recommend for you to consider.
1) Making use of partitions for your table may help if you frequently only access data from certain days at a time. There's a notebook in the Databricks Guide called "Partitioned Tables" with more data.
2) If your files are really small - it is true that you may get better performance by consolidating those files into a smaller number. You can do that easily in spark with a command like this:
sqlContext.parquetFile( SOME_INPUT_FILEPATTERN )
.coalesce(SOME_SMALLER_NUMBER_OF_DESIRED_PARTITIONS)
.write.parquet(SOME_OUTPUT_DIRECTORY)