<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>topic Re: PARTITIONED BY, Liquid Clustering, and OPTIMIZE: when should you use each? in Get Started Discussions</title>
    <link>https://community.databricks.com/t5/get-started-discussions/partitioned-by-liquid-clustering-and-optimize-when-should-you/m-p/170417#M12179</link>
    <description>&lt;P&gt;Thanks &lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/250064"&gt;@Satyasai&lt;/a&gt;&amp;nbsp;!!!!&lt;/P&gt;</description>
    <pubDate>Fri, 02 Oct 2026 12:32:29 GMT</pubDate>
    <dc:creator>arthurfr23</dc:creator>
    <dc:date>2026-10-02T12:32:29Z</dc:date>
    <item>
      <title>PARTITIONED BY, Liquid Clustering, and OPTIMIZE: when should you use each?</title>
      <link>https://community.databricks.com/t5/get-started-discussions/partitioned-by-liquid-clustering-and-optimize-when-should-you/m-p/170354#M12177</link>
      <description>&lt;P&gt;&lt;SPAN&gt;I put together a practical guide to Delta Lake table layouts, with SQL examples in Databricks:&lt;/SPAN&gt;&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;&lt;P&gt;&lt;SPAN&gt;When to consider partitioning or Liquid Clustering.&lt;/SPAN&gt;&lt;/P&gt;&lt;/LI&gt;&lt;LI&gt;&lt;P&gt;&lt;SPAN&gt;How OPTIMIZE works with partitions, ZORDER, and clustering.&lt;/SPAN&gt;&lt;/P&gt;&lt;/LI&gt;&lt;LI&gt;&lt;P&gt;&lt;SPAN&gt;How to compare bytes scanned, query latency, and maintenance costs using the same data and queries.&lt;/SPAN&gt;&lt;/P&gt;&lt;/LI&gt;&lt;/UL&gt;&lt;P&gt;&lt;SPAN&gt;Read the guide:&lt;/SPAN&gt;&lt;BR /&gt;&lt;A href="https://medium.com/@arthurfr23/choosing-a-delta-lake-layout-when-to-partition-use-liquid-clustering-and-run-optimize-c046650cb41b" target="_blank" rel="noopener"&gt;&lt;SPAN&gt;https://medium.com/@arthurfr23/choosing-a-delta-lake-layout-when-to-partition-use-liquid-clustering-and-run-optimize-c046650cb41b&lt;/SPAN&gt;&lt;/A&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;If you’ve moved from partitioning or ZORDER to Liquid Clustering, what changed in your query performance and maintenance costs?&lt;/SPAN&gt;&lt;/P&gt;</description>
      <pubDate>Thu, 01 Oct 2026 17:37:38 GMT</pubDate>
      <guid>https://community.databricks.com/t5/get-started-discussions/partitioned-by-liquid-clustering-and-optimize-when-should-you/m-p/170354#M12177</guid>
      <dc:creator>arthurfr23</dc:creator>
      <dc:date>2026-10-01T17:37:38Z</dc:date>
    </item>
    <item>
      <title>Re: PARTITIONED BY, Liquid Clustering, and OPTIMIZE: when should you use each?</title>
      <link>https://community.databricks.com/t5/get-started-discussions/partitioned-by-liquid-clustering-and-optimize-when-should-you/m-p/170386#M12178</link>
      <description>&lt;P&gt;&lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/129995"&gt;@arthurfr23&lt;/a&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;The guide provides an in-depth overview of how to test different table layouts in Delta Lake. Transitioning from Hive-style partitioning and Z-Ordering to Liquid Clustering (CLUSTER BY) has significant benefits on both query performance and maintenance overhead. Here’s a breakdown of the advantages:&lt;/P&gt;&lt;P&gt;1. Query Performance Realignment&lt;BR /&gt;Eliminating Multi-Column Penalty&lt;BR /&gt;With Z-Order the performance penalty kicks in exponentially with each additional column resulting in significantly worse query times for more than 2-3 columns.&lt;BR /&gt;Using Liquid Clustering you can cluster by up to 4 keys without significant multi-column penalty. Applying filters on both high and low cardinality columns yields excellent performance gains due to file skipping.&lt;BR /&gt;Fixing Skew and High Cardinality Issues&lt;BR /&gt;Hive-style partitioning by high cardinality columns (ex: customer_id or date string) has poor performance characteristics due to excessive directory nesting and metadata scanning overhead for skewed values. There is also the additional problem of the small file issue.&lt;BR /&gt;Partitioning also leads to extreme skew if you partition by an unevenly distributed column. Liquid Clustering resolves all these issues by keeping the directory structure flat while being able to skip files by leveraging the internal Z-Cube file metadata statistics (stored in transaction log) without incurring additional metadata scanning overhead. This results in roughly the same byte scanned as optimal Z-Order but without the expensive directory tree traversal.&lt;BR /&gt;Avoiding Performance Degradation over Time&lt;BR /&gt;Queries on recently appended data on Z-Ordered tables quickly degrade in performance until an expensive OPTIMIZE ZORDER is run.&lt;BR /&gt;Queries on clustered tables have consistently good performance characteristics since Liquid Clustering is incremental (only rewriting the dirty cubes) and supports eager clustering during writes for tables smaller than 512 GB.&lt;/P&gt;&lt;P&gt;2.Key Operational Metrics to Benchmark&lt;BR /&gt;When running the lab please use the following guidelines while comparing the 3 configurations (gld_orders_plain, gld_orders_partitioned, gld_orders_liquid):&lt;BR /&gt;1. Compare Byte Scanning Performance using Spark UI&lt;BR /&gt;Look for the number of files read vs. pruned files for queries which leverage multiple predicates (ex: WHERE customer_id = 842 AND order_date = ‘2026-03-15’). Liquid Clustering should either match or beat Z-Order while scanning significantly lower number of metadata files.&lt;BR /&gt;2. Compare Maintenance Runtime&lt;BR /&gt;Run OPTIMIZE on recently appended data and measure the maintenance runtime. OPTIMIZE on liquid clustered tables should be significantly faster and consume fewer DBUs compared to OPTIMIZE ZORDER BY.&lt;/P&gt;&lt;P&gt;&lt;A href="https://www.youtube.com/watch?v=_2OmZl63hgs" target="_self"&gt;Liquid Clustering vs Partitioning in Delta Lake&lt;/A&gt;&lt;BR /&gt;This video shows how partitioning compares to Liquid Clustering in practice. It’s a great complement to the benchmarking guidance described in this post.&lt;/P&gt;</description>
      <pubDate>Fri, 02 Oct 2026 06:07:44 GMT</pubDate>
      <guid>https://community.databricks.com/t5/get-started-discussions/partitioned-by-liquid-clustering-and-optimize-when-should-you/m-p/170386#M12178</guid>
      <dc:creator>Satyasai</dc:creator>
      <dc:date>2026-10-02T06:07:44Z</dc:date>
    </item>
    <item>
      <title>Re: PARTITIONED BY, Liquid Clustering, and OPTIMIZE: when should you use each?</title>
      <link>https://community.databricks.com/t5/get-started-discussions/partitioned-by-liquid-clustering-and-optimize-when-should-you/m-p/170417#M12179</link>
      <description>&lt;P&gt;Thanks &lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/250064"&gt;@Satyasai&lt;/a&gt;&amp;nbsp;!!!!&lt;/P&gt;</description>
      <pubDate>Fri, 02 Oct 2026 12:32:29 GMT</pubDate>
      <guid>https://community.databricks.com/t5/get-started-discussions/partitioned-by-liquid-clustering-and-optimize-when-should-you/m-p/170417#M12179</guid>
      <dc:creator>arthurfr23</dc:creator>
      <dc:date>2026-10-02T12:32:29Z</dc:date>
    </item>
  </channel>
</rss>

