Deduplication, Bronze (raw) or Silver (enriched)

baatchus
New Contributor III

Need some help in choosing between where to do deduplication of data. So I have sensor data in blob storage that I'm picking up with Databricks Autoloader. The data and files can have duplicates in them.

Which of the 2 options do I choose?

Option 1:

  • Create a Bronze (Raw) Delta Lake table which reads from the files with Autoloader and only appends data.
  • Create a Silver (Enriched) Delta Lake table that reads from Bronze table and does merge to deduplicate?
  • Create a Silver (Enriched) Delta Lake that reads from the first Silver table and joins with another table.

Option 2:

  • Create a Bronze (Raw) Delta Lake table which reads from the files with Autoloader and does merge into to deduplicate
  • Create a Silver (Enriched) Delta Lake table with reads from the first Silver table and joins with another table.