Databricks Community

Erik · ‎11-05-2021

Situation: we have one partion per date, and it just so happens that each partition ends up (after optimize) as *a single* 128mb file. We partition on date, and zorder on userid, and our query is something like "find max value of column A where userid=X and date>=somedate".

Does zordering help in any way in this scenario? It is clear that we will have to read every partition after $somedate, but does the zordering on userid somehow help spark when reading inside each of those partitions (remember that each partition is a single file), or do we have to read *and scan* all 128mb of each of the remaining partitions even when we zoptimize?

-werners- · ‎11-07-2021

Z-Order will make sure that in case you need to read multiple files, these files are co-located.

For a single file this does not matter as a single file is always local to itself.

If you are certain that your spark program will only read a single file, you do not need z-ordering.

But it might be the case that your delta lake table is also read by another program, not using the partition filter. then it will become interesting, or if you have multiple files per partition.

Z-Ordering and partitioning are complementary techniques.

Z-Ordering is especially interesting for columns on which you cannot/don't want to partition (high cardinality)

View solution in original post

Erik · ‎11-07-2021

@Kaniz Fatma I have read the documentation. The question is not about general guidelines regarding partitions and zordering, it is very specifically about the (potential) benefit of zordering when reading single files. To rephrase: is the only advantage of zordering that it allows the skipping of whole files, or is there also some benefit to it after a file has been selected to be read. Does it allow faster searching inside the selected files, or maybe reading only chunks of the files?

Hubert-Dudek · ‎11-07-2021

ZORDER BY

Colocate column information in the same set of files. Co-locality is used by Delta Lake data-skipping algorithms to dramatically reduce the amount of data that needs to be read. You can specify multiple columns for ZORDER BY as a comma-separated list. However, the effectiveness of the locality drops with each additional column.

it is from https://docs.databricks.com/spark/latest/spark-sql/language-manual/delta-optimize.html

So for delta files once partition disappear it is really important to have Z-order as it will handle effectively your query, so you need:

OPTIMIZE data ZORDER BY (userid, date)

My blog: https://databrickster.medium.com/

Erik · ‎11-07-2021

@Hubert Dudek i don't know what you mean by "when the partition dissappear". I clearly asked this question in a confusing way, but hopefully my answer to @Kaniz Fatma helped clarify.

-werners- · ‎11-07-2021

Z-Order will make sure that in case you need to read multiple files, these files are co-located.

For a single file this does not matter as a single file is always local to itself.

If you are certain that your spark program will only read a single file, you do not need z-ordering.

But it might be the case that your delta lake table is also read by another program, not using the partition filter. then it will become interesting, or if you have multiple files per partition.

Z-Ordering and partitioning are complementary techniques.

Z-Ordering is especially interesting for columns on which you cannot/don't want to partition (high cardinality)

Databricks Community

Does Z-ordering speed up reading of a single file?

DAIS 2026 Speaker Spotlight Series #4 | Archika Dogra

Databricks Community Champion - May 2026 - Balaji J

Solution Accelerator Series | Media Mix Modeling (MMM)

DAIS 2026 | Community Virtual Contest – Showcase Your Skills & Win Exclusive Swag

DAIS registrants: apply for the Apps & Agents for Good Hackathon 2026