<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>topic Architectural Pattern for PII: &amp;quot;Redact at Ingestion&amp;quot; vs. &amp;quot;Dynamic Masking at Consumption in Data Engineering</title>
    <link>https://community.databricks.com/t5/data-engineering/architectural-pattern-for-pii-quot-redact-at-ingestion-quot-vs/m-p/169190#M56052</link>
    <description>&lt;P&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;Hi everyone, we are defining our PII strategy in &lt;STRONG&gt;Unity Catalog&lt;/STRONG&gt;. We are split on whether to: A) Redact/Hash PII at the Silver layer (permanent change), or B) Keep PII in Silver and use &lt;STRONG&gt;Dynamic Data Masking&lt;/STRONG&gt; at the Gold/View layer.&lt;/P&gt;&lt;OL&gt;&lt;LI&gt;Does "Redact at Ingestion" (A) cause too much friction for Data Science teams who occasionally need the raw data?&lt;/LI&gt;&lt;LI&gt;Is the &lt;STRONG&gt;Dynamic Masking&lt;/STRONG&gt; approach (B) robust enough for strict audits (e.g., GDPR/CCPA)? Are there "escape hatches" where power users might accidentally bypass the masking via an API?&lt;/LI&gt;&lt;LI&gt;Are you using Unity Catalog’s built-in &lt;STRONG&gt;Data Classification tags&lt;/STRONG&gt; to trigger these policies automatically?&lt;/LI&gt;&lt;/OL&gt;</description>
    <pubDate>Sat, 19 Sep 2026 16:00:48 GMT</pubDate>
    <dc:creator>Khasim_1</dc:creator>
    <dc:date>2026-09-19T16:00:48Z</dc:date>
    <item>
      <title>Architectural Pattern for PII: "Redact at Ingestion" vs. "Dynamic Masking at Consumption</title>
      <link>https://community.databricks.com/t5/data-engineering/architectural-pattern-for-pii-quot-redact-at-ingestion-quot-vs/m-p/169190#M56052</link>
      <description>&lt;P&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;Hi everyone, we are defining our PII strategy in &lt;STRONG&gt;Unity Catalog&lt;/STRONG&gt;. We are split on whether to: A) Redact/Hash PII at the Silver layer (permanent change), or B) Keep PII in Silver and use &lt;STRONG&gt;Dynamic Data Masking&lt;/STRONG&gt; at the Gold/View layer.&lt;/P&gt;&lt;OL&gt;&lt;LI&gt;Does "Redact at Ingestion" (A) cause too much friction for Data Science teams who occasionally need the raw data?&lt;/LI&gt;&lt;LI&gt;Is the &lt;STRONG&gt;Dynamic Masking&lt;/STRONG&gt; approach (B) robust enough for strict audits (e.g., GDPR/CCPA)? Are there "escape hatches" where power users might accidentally bypass the masking via an API?&lt;/LI&gt;&lt;LI&gt;Are you using Unity Catalog’s built-in &lt;STRONG&gt;Data Classification tags&lt;/STRONG&gt; to trigger these policies automatically?&lt;/LI&gt;&lt;/OL&gt;</description>
      <pubDate>Sat, 19 Sep 2026 16:00:48 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/architectural-pattern-for-pii-quot-redact-at-ingestion-quot-vs/m-p/169190#M56052</guid>
      <dc:creator>Khasim_1</dc:creator>
      <dc:date>2026-09-19T16:00:48Z</dc:date>
    </item>
    <item>
      <title>Re: Architectural Pattern for PII: "Redact at Ingestion" vs. "Dynamic Masking at Cons</title>
      <link>https://community.databricks.com/t5/data-engineering/architectural-pattern-for-pii-quot-redact-at-ingestion-quot-vs/m-p/169258#M56070</link>
      <description>&lt;P&gt;Hi&amp;nbsp;&lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/249065"&gt;@Khasim_1&lt;/a&gt;,&lt;/P&gt;
&lt;P&gt;I would keep the PII in Silver within your regulatory and internal retention limits and control access with Unity Catalog's &lt;A href="https://docs.databricks.com/aws/en/tables/row-and-column-filters" target="_blank"&gt;row filters and column masks&lt;/A&gt;, rather than redacting permanently. The reason is the one you already gave. A permanent change at Silver is a one-way door and hard to reverse, and the moment a business user has a legitimate future need for the raw value, you are stuck re-ingesting from source if it still exists. That data can be valuable, and you do not want to throw it away early.&lt;/P&gt;
&lt;P&gt;On your first question, yes, redacting at ingestion tends to create real friction for data science. Teams that occasionally need the raw value for entity resolution, fraud work, or joining across systems cannot get it back, and hashing quietly breaks joins and distribution analysis unless you are very careful with consistent salting, which then limits its usefulness anyway. What usually happens is you end up setting up a separate raw pipeline for those teams, which is the duplication you were trying to avoid. Just to be clear, hashing or redaction at Silver is usually pseudonymisation, not anonymisation, so if the value can still be re-identified, it remains personal data under GDPR and does not take you out of scope. Only genuine anonymisation does that.&lt;/P&gt;
&lt;P&gt;On the second question, dynamic masking with Unity Catalog is robust enough for GDPR and CCPA work, because it is enforced by the engine at the metastore rather than in a view someone has to remember to use. The same row filter or column mask applies whether the data is read through SQL, a BI tool, a notebook, JDBC or ODBC, or a Genie Agent, and every access is written to the &lt;A href="https://docs.databricks.com/aws/en/admin/system-tables/audit-logs" target="_blank"&gt;audit log system table&lt;/A&gt;, which is usually what an auditor would be interested in.&lt;/P&gt;
&lt;P&gt;To your question about escape hatches, the real one is not the SQL API, it is direct access to the underlying cloud storage. If someone can read the Delta files in the bucket, none of the masking runs, so lock down your &lt;A href="https://docs.databricks.com/aws/en/connect/unity-catalog/cloud-storage/manage-external-locations" target="_blank"&gt;external locations and storage credentials&lt;/A&gt; and make sure nobody is querying the files directly. Beyond that, run on compute that enforces fine-grained access control, keep tight control of the privileged principals who can alter or drop a policy or who own the masking function, and remember that masking is access control, not erasure. A right-to-be-forgotten request still needs a real delete, so keep that as a separate process.&lt;/P&gt;
&lt;P&gt;On your third question, yes, and this is where it becomes scalable. Rather than attach masks and filters table by table, tag the sensitive columns and let one &lt;A href="https://docs.databricks.com/aws/en/data-governance/unity-catalog/abac" target="_blank"&gt;ABAC&lt;/A&gt; policy match the tag and apply the rule to every current and future table that carries it. Unity Catalog's &lt;A href="https://docs.databricks.com/aws/en/data-governance/unity-catalog/data-classification" target="_blank"&gt;data classification&lt;/A&gt; can scan and populate those tags for you, and the policy then acts on them. I would review the classification output rather than trust it blindly, but once it is set you define the rule once and it follows the tag.&lt;/P&gt;
&lt;P&gt;For context, this is close to what I am doing right now in a regulated engagement. We implemented both row-level and column-level security for the data science teams entirely through ABAC, and the same policies apply automatically when they query those tables through Genie Agents. You define it once, and it holds no matter how the data is accessed, because UC resolves the entitlement centrally and only returns what the user is allowed to see. Per-table filters and masks work fine, but ABAC is far more scalable and easier to govern once you are past a handful of tables.&lt;/P&gt;
&lt;P&gt;So my recommendation is to keep the raw data governed in Silver within your retention limits, define row and column security once with ABAC driven by classification tags, and lock down the storage layer. That gives data science the flexibility to reach raw data when they are entitled to it, and gives you an audit story that holds up.&lt;/P&gt;
&lt;P&gt;One a side note, I do sometimes see data science teams take a cut from Bronze so they can work on truly raw, un-standardised data before any Silver cleansing. It can be a fair choice, but governance gets harder once that data sits in a separate environment, so if you go that way I would keep it inside the same Unity Catalog governance rather than outside it.&lt;/P&gt;
&lt;P&gt;Hope this helps.&lt;/P&gt;
&lt;P class="p1"&gt;&lt;FONT size="2" color="#FF6600"&gt;&lt;STRONG&gt;&lt;I&gt;If this answer resolves your question, could you mark it as “Accept as Solution”? That helps other users quickly find the correct fix.&lt;/I&gt;&lt;/STRONG&gt;&lt;/FONT&gt;&lt;I&gt;&lt;/I&gt;&lt;/P&gt;</description>
      <pubDate>Sun, 20 Sep 2026 20:11:18 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/architectural-pattern-for-pii-quot-redact-at-ingestion-quot-vs/m-p/169258#M56070</guid>
      <dc:creator>Ashwin_DSA</dc:creator>
      <dc:date>2026-09-20T20:11:18Z</dc:date>
    </item>
    <item>
      <title>Re: Architectural Pattern for PII: "Redact at Ingestion" vs. "Dynamic Masking at Cons</title>
      <link>https://community.databricks.com/t5/data-engineering/architectural-pattern-for-pii-quot-redact-at-ingestion-quot-vs/m-p/169941#M56202</link>
      <description>&lt;P&gt;Hey&amp;nbsp;&lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/249065"&gt;@Khasim_1&lt;/a&gt;&amp;nbsp;, if it helps...did an experiment and blog post with results from the run that illustrates the pattern that&amp;nbsp;&lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/216690"&gt;@Ashwin_DSA&lt;/a&gt;recommended.&lt;BR /&gt;&lt;BR /&gt;&lt;A href="https://medium.com/lab-notes/redacting-pii-in-silver-leaves-it-in-bronze-column-masks-vs-redaction-on-databricks-f33175cf9267" target="_blank"&gt;https://medium.com/lab-notes/redacting-pii-in-silver-leaves-it-in-bronze-column-masks-vs-redaction-on-databricks-f33175cf9267&lt;/A&gt;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Sun, 27 Sep 2026 16:54:27 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/architectural-pattern-for-pii-quot-redact-at-ingestion-quot-vs/m-p/169941#M56202</guid>
      <dc:creator>gnakan</dc:creator>
      <dc:date>2026-09-27T16:54:27Z</dc:date>
    </item>
    <item>
      <title>Re: Architectural Pattern for PII: "Redact at Ingestion" vs. "Dynamic Masking at Cons</title>
      <link>https://community.databricks.com/t5/data-engineering/architectural-pattern-for-pii-quot-redact-at-ingestion-quot-vs/m-p/170025#M56213</link>
      <description>&lt;P&gt;Hi,&lt;/P&gt;&lt;P&gt;In practice most teams end up with a hybrid: keep raw PII in one tightly restricted place, and use dynamic masking everywhere else.&lt;/P&gt;&lt;P&gt;1. Does redacting at Silver (A) create friction? Yes, often. Once PII is hashed or dropped in Silver, the raw data is gone for everyone. Data Science then asks for exceptions, and those turn into ad-hoc copies, which is worse for compliance. Two further points:&lt;/P&gt;&lt;P&gt;Hashing is not anonymization. Under GDPR, hashed identifiers are still pseudonymized personal data, so A doesn't take you out of scope.&lt;BR /&gt;Common middle ground: tokenize direct identifiers (email, SSN) in Silver with a consistent key, so joins still work. Keep the token → raw mapping in a separate restricted "PII vault" schema. Data Science works with tokens and asks for re-identification only when they need it.&lt;/P&gt;&lt;P&gt;2. Is dynamic masking (B) audit-proof? It can be. Masks and filters are enforced by Unity Catalog on every UC-governed access path (SQL Warehouses, notebooks, jobs, DLT, APIs), and every access is in system.access.audit. The real escape hatches are usually around UC, not through it:&lt;/P&gt;&lt;P&gt;Direct cloud storage access. Anyone with IAM/SAS access to the bucket or container can read the files and skip UC entirely. Lock storage down so only UC storage credentials can reach it.&lt;BR /&gt;Exempt users making copies. Anyone in an EXCEPT group can CREATE TABLE AS SELECT raw data into an unprotected table. Keep exempt groups small and alert on it through the audit log.&lt;BR /&gt;Policy admins. Users with MANAGE or ownership can change or drop policies. Separate policy admins from data engineers.&lt;BR /&gt;Right to erasure. Masking hides data but doesn't delete it. You still need a deletion process plus VACUUM to meet GDPR/CCPA erasure requests.&lt;/P&gt;&lt;P&gt;3. Data classification → automatic policies: Yes, this is the intended pattern. Data classification scans tables and tags PII columns. ABAC column-mask policies at catalog or schema level then match those tags (MATCH COLUMNS has_tag(...)), so new PII columns get masked without manual work. Tips:&lt;/P&gt;&lt;P&gt;Review classification results before relying on them fully. They are probabilistic, so false negatives happen.&lt;BR /&gt;Add manual governed tags where it matters.&lt;BR /&gt;Use a default-deny approach on sensitive schemas as a backstop.&lt;/P&gt;</description>
      <pubDate>Mon, 28 Sep 2026 09:31:53 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/architectural-pattern-for-pii-quot-redact-at-ingestion-quot-vs/m-p/170025#M56213</guid>
      <dc:creator>Islam_hoti</dc:creator>
      <dc:date>2026-09-28T09:31:53Z</dc:date>
    </item>
  </channel>
</rss>

