Waldar
Databricks Employee
Databricks Employee

For larger tables, it's faster to compute the approximate count distinct rather than the exact value.

The difference of the actual number won't change the query plan - the optimizer usually choose by looking at order of magnitudes rather a specific threshold.

View solution in original post