How do we manage data recency in Databricks

User16826994223 — Mon, 21 Jun 2021 12:57:04 GMT

I want to know how databricks maintain data recency in databricks

Re: How do we manage data recency in Databricks

sajith_appukutt — Wed, 23 Jun 2021 00:43:42 GMT

When using delta tables in databricks, you have the advantage of delta cache which accelerates data reads by creating copies of remote files in nodes’ local storage using a fast intermediate data format. At the beginning of each query delta tables auto-update to the latest version - this way data is always recent.

However, if it is acceptable for results to be stale for a short duration of time, you could lower the latency of queries further. This is done by setting the Spark session configuration variable spark.databricks.delta.stalenessLimit with a time string value, e.g 1h, 15m, 1d

topic Re: How do we manage data recency in Databricks in Data Engineering

How do we manage data recency in Databricks

Re: How do we manage data recency in Databricks