<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>topic Data Engineering 3.0 - From Trusted Data Products to Context Layers in Community Articles</title>
    <link>https://community.databricks.com/t5/community-articles/data-engineering-3-0-from-trusted-data-products-to-context/m-p/167304#M1525</link>
    <description>&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="1_3Mo6NkMHbxBHodt-3yUskw" style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30579iC6CABCF312EF35F8/image-size/large?v=v2&amp;amp;px=999" role="button" title="1_3Mo6NkMHbxBHodt-3yUskw" alt="1_3Mo6NkMHbxBHodt-3yUskw" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;Data Engineering &lt;STRONG&gt;1.0&lt;/STRONG&gt; was largely about integrating, moving, consolidating massive data to support enterprise decisions. Enterprises built ETL pipelines, data warehouses, data lakes and orchestration platforms to bring data together reliably.&amp;nbsp;&lt;/SPAN&gt;Data Engineering &lt;STRONG&gt;2.0&lt;/STRONG&gt; was about making the move to Cloud &lt;STRONG&gt;Lakehouse&lt;/STRONG&gt; and making the data trustworthy and reusable. This created a much more mature enterprise data estate. But the arrival of generative and agentic AI changes the question again. An enterprise may now have thousands of tables, hundreds of data products, semantic models, business glossaries, APIs, documents and real-time data.&amp;nbsp;An agent can potentially access much of this information. But access does not mean understanding. An agent investigating a business problem needs to know which information is relevant, how different entities are related, which definitions are authoritative, which rules apply, what happened previously, what is happening now, and what evidence supports a conclusion.&lt;/P&gt;&lt;P class=""&gt;Enterprises may already have all of this information. The problem is that it is distributed across the data estate. This is the &lt;STRONG&gt;context gap  -  &lt;/STRONG&gt;the distance between information that exists in the enterprise and the context an AI system needs to reliably perform a task.&lt;/P&gt;&lt;H3&gt;&lt;U&gt;&lt;STRONG&gt;The Agentic&amp;nbsp;Shift&lt;/STRONG&gt;&lt;/U&gt;&lt;/H3&gt;&lt;P class=""&gt;Traditional analytics was designed primarily to help users make vital decisions. A dashboard presents information. A business user interprets it. The user makes the decision. Machine learning automated parts of this process by using data to generate predictions or classifications.&amp;nbsp;&lt;STRONG&gt;Agentic AI&lt;/STRONG&gt; introduces a different operating model. An agent can retrieve information, reason over it, use tools, interact with systems and potentially take action. This means the data platform is no longer serving only a human consumer or an analytical model. It is increasingly serving an intelligent system that needs to understand the enterprise before it can act within it. This changes the role of data engineering.&lt;/P&gt;&lt;P class=""&gt;The question is no longer only &lt;STRONG&gt;How do we make trusted data available?&amp;nbsp;&lt;/STRONG&gt;It becomes &lt;STRONG&gt;How do we make trusted enterprise information understandable and usable by an AI system for a specific decision?&amp;nbsp;&lt;/STRONG&gt;That is where I see the emergence of &lt;STRONG&gt;Data Engineering 3.0&lt;/STRONG&gt;.&lt;/P&gt;&lt;H3&gt;&lt;U&gt;Data Engineering 3.0&lt;/U&gt;&lt;/H3&gt;&lt;P class=""&gt;Data Engineering 3.0 is the next evolution of the data engineering discipline. It is not a replacement for Data Engineering 2.0. It builds on it. The important shift is the unit of value. In Data Engineering 1.0, the question was whether data could be moved reliably at scale, processed and consumed in a robust manner. In Data Engineering 2.0, the question became whether data could be trusted and reused in accordance with governance and usage at scale via data products. In Data Engineering 3.0, the question becomes whether an intelligent system has the right context to perform a task reliably.&amp;nbsp;We are moving from engineering data for consumption to increasingly engineering the &lt;STRONG&gt;information environment in which AI operates&lt;/STRONG&gt;.&lt;/P&gt;&lt;H3&gt;&lt;U&gt;A Trusted Data Product is not the same as&amp;nbsp;Context&lt;/U&gt;&lt;/H3&gt;&lt;P class=""&gt;Consider a bank that has already built well-governed data products.&amp;nbsp;It may have a Customer data product containing customer information. It may have a Transaction data product containing transactions. It may have a Merchant data product. It may have a Device data product. It may have a Fraud data product containing historical fraud cases. Each of these can be high-quality, governed and independently reusable.&lt;/P&gt;&lt;P class=""&gt;Now imagine an agent is asked&amp;nbsp;&lt;STRONG&gt;Investigate this transaction and determine whether it should be escalated for fraud investigation.&amp;nbsp;&lt;/STRONG&gt;The agent does not simply need the Transaction data product. It may need to understand the customer associated with the transaction, the account involved, the merchant, the device used, recent transaction behaviour, previous fraud cases, current fraud indicators and the policies that determine when a transaction should be escalated. The data already exists. But the decision requires those pieces of information to be brought together with their meaning and relationships intact. That assembled information is the &lt;STRONG&gt;context for the decision&lt;/STRONG&gt;. &lt;STRONG&gt;A Data Product provides trusted information. A Context Layer makes trusted information relevant to a particular task.&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="1_hWW5XiezH9zkHSmdOmFHjA" style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30580i7196F753A5D38503/image-size/large?v=v2&amp;amp;px=999" role="button" title="1_hWW5XiezH9zkHSmdOmFHjA" alt="1_hWW5XiezH9zkHSmdOmFHjA" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;P&gt;&lt;FONT size="4"&gt;&lt;STRONG&gt;&lt;U&gt;The Context&amp;nbsp;Layer&lt;/U&gt;&lt;/STRONG&gt;&lt;/FONT&gt;&lt;/P&gt;&lt;P class=""&gt;An AI Context Layer is a decision-oriented capability that brings together the information an AI system needs to perform a specific task. It can include structured data, semantic definitions, relationships between entities, historical information, real-time signals, business policies and supporting evidence. It does not necessarily mean creating another physical copy of the enterprise data. In many cases, it can be a governed layer that composes information from existing data products and other enterprise systems when an agent needs it.&lt;/P&gt;&lt;P class=""&gt;For the fraud investigation example, the context will include&amp;nbsp;&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;Customer information and history&lt;/LI&gt;&lt;LI&gt;The transactions and accounts involved&lt;/LI&gt;&lt;LI&gt;Merchant characteristics&lt;/LI&gt;&lt;LI&gt;Device relationships&lt;/LI&gt;&lt;LI&gt;Recent transaction patterns&lt;/LI&gt;&lt;LI&gt;Previous fraud cases&lt;/LI&gt;&lt;LI&gt;Current fraud signals&lt;/LI&gt;&lt;LI&gt;Applicable fraud policies&lt;/LI&gt;&lt;LI&gt;Evidence supporting the recommendation&lt;/LI&gt;&lt;/UL&gt;&lt;P&gt;The important point is that the agent does not need the entire enterprise. It needs the &lt;STRONG&gt;right context for the decision it is making&lt;/STRONG&gt;.&lt;/P&gt;&lt;P class=""&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="1_CUWN7KstAGcXAcnxfeMy4Q" style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30581i8CD19118191A76E6/image-size/large?v=v2&amp;amp;px=999" role="button" title="1_CUWN7KstAGcXAcnxfeMy4Q" alt="1_CUWN7KstAGcXAcnxfeMy4Q" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;H3&gt;&lt;U&gt;Context is more than Retrieval&lt;/U&gt;&lt;/H3&gt;&lt;P class=""&gt;The first instinct when discussing AI context is often to think about &lt;STRONG&gt;RAG&lt;/STRONG&gt;. RAG is an important capability. It allows an AI system to retrieve relevant information before generating a response. But enterprise decision-making requires more than retrieving documents. An agent may need to understand that a particular transaction belongs to an account, that the account belongs to a customer, that the transaction was performed using a device associated with several accounts, and that the merchant has a particular risk classification. That is not simply a document retrieval problem. It is a problem involving &lt;STRONG&gt;entities, relationships, semantics, history, current state and rules&lt;/STRONG&gt;.&lt;/P&gt;&lt;P class=""&gt;Context may come from structured data, documents, knowledge graph, APIs or operational systems or real-time events. The Context Layer is therefore not a specific technology. It is an architectural capability for composing the information required by an intelligent system.&lt;/P&gt;&lt;H3&gt;&lt;U&gt;Where Ontology&amp;nbsp;fits&lt;/U&gt;&lt;/H3&gt;&lt;P class=""&gt;An ontology can make business concepts and relationships explicit. For example, in a banking space, the enterprise may define relationships between Customer, Account, Transaction, Device and Merchant. That semantic structure helps an AI system understand that these are not independent datasets. It understands the business relationships between them. Ontology therefore provides part of the semantic foundation for a Context Layer. Ontology and context are not the same thing. Ontology answers questions such as&amp;nbsp;&lt;/P&gt;&lt;P class=""&gt;What are the important business concepts? , How are they related? , What does a concept mean? , The Context Layer answers a different question.&amp;nbsp;&lt;STRONG&gt;What of this information and meaning does the agent need for this particular task? &lt;/STRONG&gt;This distinction is important because it prevents the Context Layer from becoming another attempt to model the entire enterprise upfront.&lt;/P&gt;&lt;H3&gt;&lt;U&gt;From Data Quality to Decision Readiness&lt;/U&gt;&lt;/H3&gt;&lt;P class=""&gt;Data Engineering 2.0 made data quality a fundamental engineering discipline. We ask whether data is accurate, complete, fresh, consistent and governed. Data Engineering 3.0 introduces another dimension. &lt;STRONG&gt;Is the context sufficient for the decision? &lt;/STRONG&gt;Consider the fraud example again.&amp;nbsp;The Customer data, Transaction data, Merchant data and Fraud data may be accurate. Yet the agent could still make a poor decision if an important signal is missing from the context. This is why &lt;STRONG&gt;Data Quality and Context Quality &lt;/STRONG&gt;are different problems. A data product can be high quality while the decision context built from several data products is incomplete. This suggests that data engineering will increasingly need to think about &lt;STRONG&gt;decision readiness&lt;/STRONG&gt;.&lt;/P&gt;&lt;P class=""&gt;Does the agent have the right information?&amp;nbsp;Is it sufficiently fresh?, Are the relationships correct?, Are the business definitions consistent?, Are the relevant policies available? andCan the agent identify the evidence behind its conclusion?&amp;nbsp;These questions sit on top of traditional data quality rather than replacing it.&lt;/P&gt;&lt;H3&gt;&lt;U&gt;Context Layer should be driven by the&amp;nbsp;Decision&lt;/U&gt;&lt;/H3&gt;&lt;P class=""&gt;There is a natural temptation to approach this as another enterprise-wide architecture exercise. An organization may say We need to build an Enterprise Context Layer. The next step becomes an attempt to model every entity, relationship, rule, document and event across the organization. That is where the initiative can quickly become too large.&amp;nbsp;The better starting point is the decision. Check &lt;STRONG&gt;Which agent task would benefit most from better enterprise context? &lt;/STRONG&gt;Then define the minimum context required for that task. For a fraud investigation agent, it means starting with Customer, Account, Transaction, Merchant and Device along with the relationships, historical information and policies required for the investigation.&amp;nbsp;The first objective is not to model the bank. It is to make one important decision more reliable. This is a very different way of approaching enterprise semantics.&lt;/P&gt;&lt;H3&gt;&lt;U&gt;Context should evolve with the&amp;nbsp;Agent&lt;/U&gt;&lt;/H3&gt;&lt;P class=""&gt;The first version of a Context Layer will rarely be complete. And it does not need to be. Start with a defined task. Identify the information required. Build the context. Test the agent. Observe where it lacks information or misunderstands meaning. Then extend the context. This creates a practical feedback loop between the agent and the data platform. The agent becomes a source of evidence about where the enterprise’s existing information environment is insufficient for intelligent decision-making.&amp;nbsp;That evidence can drive the next iteration of the data product or semantic layer. Over time, the Context Layer can expand from one decision to multiple decisions and from one domain to connected domains. The goal is not to build the biggest context layer. It is to build the &lt;STRONG&gt;most useful context for the decisions that matter&lt;/STRONG&gt;.&lt;/P&gt;&lt;H3&gt;&lt;U&gt;What changes for Data Engineering?&lt;/U&gt;&lt;/H3&gt;&lt;P class=""&gt;Data Engineering 3.0 does not mean data engineers stop building pipelines. The fundamentals remain critical. Data still needs to be ingested, transformed, stored, governed, monitored and served. But the scope of the discipline expands. Data engineers will increasingly need to understand how enterprise information is represented for intelligent systems. That includes areas such as Semantic modeling, Ontology, Knowledge graphs, Entity resolution, Real-time data, Context retrieval, Agent tools, Data and AI governance, Context evaluation and Decision lineage.&amp;nbsp;&lt;/P&gt;&lt;P class=""&gt;The data engineer’s responsibility increasingly moves from &lt;STRONG&gt;Can we deliver the data? &lt;/STRONG&gt;to &lt;STRONG&gt;Can an intelligent system understand and use the data correctly in the context of the decision?. &lt;/STRONG&gt;That is a significant evolution in the role.&amp;nbsp;The investments made during Data Engineering 2.0 become the foundation for Data Engineering 3.0. Trusted data products provide reliable information. Data quality provides confidence in that information. Governance provides controls. Lineage provides provenance. Semantic models provide business meaning. Real-time platforms provide current state. Ontologies and Knowledge Graphs can provide explicit relationships.&lt;/P&gt;&lt;P class=""&gt;The next step is to bring these capabilities together around the decisions that AI agents need to perform. Context Layer is the &lt;STRONG&gt;new layer built via DE 3.0 on top of the Trusted Data Foundation enterprises have already created in DE 2.0&lt;/STRONG&gt;.&lt;/P&gt;</description>
    <pubDate>Wed, 02 Sep 2026 13:13:51 GMT</pubDate>
    <dc:creator>balajij8</dc:creator>
    <dc:date>2026-09-02T13:13:51Z</dc:date>
    <item>
      <title>Data Engineering 3.0 - From Trusted Data Products to Context Layers</title>
      <link>https://community.databricks.com/t5/community-articles/data-engineering-3-0-from-trusted-data-products-to-context/m-p/167304#M1525</link>
      <description>&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="1_3Mo6NkMHbxBHodt-3yUskw" style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30579iC6CABCF312EF35F8/image-size/large?v=v2&amp;amp;px=999" role="button" title="1_3Mo6NkMHbxBHodt-3yUskw" alt="1_3Mo6NkMHbxBHodt-3yUskw" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;Data Engineering &lt;STRONG&gt;1.0&lt;/STRONG&gt; was largely about integrating, moving, consolidating massive data to support enterprise decisions. Enterprises built ETL pipelines, data warehouses, data lakes and orchestration platforms to bring data together reliably.&amp;nbsp;&lt;/SPAN&gt;Data Engineering &lt;STRONG&gt;2.0&lt;/STRONG&gt; was about making the move to Cloud &lt;STRONG&gt;Lakehouse&lt;/STRONG&gt; and making the data trustworthy and reusable. This created a much more mature enterprise data estate. But the arrival of generative and agentic AI changes the question again. An enterprise may now have thousands of tables, hundreds of data products, semantic models, business glossaries, APIs, documents and real-time data.&amp;nbsp;An agent can potentially access much of this information. But access does not mean understanding. An agent investigating a business problem needs to know which information is relevant, how different entities are related, which definitions are authoritative, which rules apply, what happened previously, what is happening now, and what evidence supports a conclusion.&lt;/P&gt;&lt;P class=""&gt;Enterprises may already have all of this information. The problem is that it is distributed across the data estate. This is the &lt;STRONG&gt;context gap  -  &lt;/STRONG&gt;the distance between information that exists in the enterprise and the context an AI system needs to reliably perform a task.&lt;/P&gt;&lt;H3&gt;&lt;U&gt;&lt;STRONG&gt;The Agentic&amp;nbsp;Shift&lt;/STRONG&gt;&lt;/U&gt;&lt;/H3&gt;&lt;P class=""&gt;Traditional analytics was designed primarily to help users make vital decisions. A dashboard presents information. A business user interprets it. The user makes the decision. Machine learning automated parts of this process by using data to generate predictions or classifications.&amp;nbsp;&lt;STRONG&gt;Agentic AI&lt;/STRONG&gt; introduces a different operating model. An agent can retrieve information, reason over it, use tools, interact with systems and potentially take action. This means the data platform is no longer serving only a human consumer or an analytical model. It is increasingly serving an intelligent system that needs to understand the enterprise before it can act within it. This changes the role of data engineering.&lt;/P&gt;&lt;P class=""&gt;The question is no longer only &lt;STRONG&gt;How do we make trusted data available?&amp;nbsp;&lt;/STRONG&gt;It becomes &lt;STRONG&gt;How do we make trusted enterprise information understandable and usable by an AI system for a specific decision?&amp;nbsp;&lt;/STRONG&gt;That is where I see the emergence of &lt;STRONG&gt;Data Engineering 3.0&lt;/STRONG&gt;.&lt;/P&gt;&lt;H3&gt;&lt;U&gt;Data Engineering 3.0&lt;/U&gt;&lt;/H3&gt;&lt;P class=""&gt;Data Engineering 3.0 is the next evolution of the data engineering discipline. It is not a replacement for Data Engineering 2.0. It builds on it. The important shift is the unit of value. In Data Engineering 1.0, the question was whether data could be moved reliably at scale, processed and consumed in a robust manner. In Data Engineering 2.0, the question became whether data could be trusted and reused in accordance with governance and usage at scale via data products. In Data Engineering 3.0, the question becomes whether an intelligent system has the right context to perform a task reliably.&amp;nbsp;We are moving from engineering data for consumption to increasingly engineering the &lt;STRONG&gt;information environment in which AI operates&lt;/STRONG&gt;.&lt;/P&gt;&lt;H3&gt;&lt;U&gt;A Trusted Data Product is not the same as&amp;nbsp;Context&lt;/U&gt;&lt;/H3&gt;&lt;P class=""&gt;Consider a bank that has already built well-governed data products.&amp;nbsp;It may have a Customer data product containing customer information. It may have a Transaction data product containing transactions. It may have a Merchant data product. It may have a Device data product. It may have a Fraud data product containing historical fraud cases. Each of these can be high-quality, governed and independently reusable.&lt;/P&gt;&lt;P class=""&gt;Now imagine an agent is asked&amp;nbsp;&lt;STRONG&gt;Investigate this transaction and determine whether it should be escalated for fraud investigation.&amp;nbsp;&lt;/STRONG&gt;The agent does not simply need the Transaction data product. It may need to understand the customer associated with the transaction, the account involved, the merchant, the device used, recent transaction behaviour, previous fraud cases, current fraud indicators and the policies that determine when a transaction should be escalated. The data already exists. But the decision requires those pieces of information to be brought together with their meaning and relationships intact. That assembled information is the &lt;STRONG&gt;context for the decision&lt;/STRONG&gt;. &lt;STRONG&gt;A Data Product provides trusted information. A Context Layer makes trusted information relevant to a particular task.&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="1_hWW5XiezH9zkHSmdOmFHjA" style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30580i7196F753A5D38503/image-size/large?v=v2&amp;amp;px=999" role="button" title="1_hWW5XiezH9zkHSmdOmFHjA" alt="1_hWW5XiezH9zkHSmdOmFHjA" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;P&gt;&lt;FONT size="4"&gt;&lt;STRONG&gt;&lt;U&gt;The Context&amp;nbsp;Layer&lt;/U&gt;&lt;/STRONG&gt;&lt;/FONT&gt;&lt;/P&gt;&lt;P class=""&gt;An AI Context Layer is a decision-oriented capability that brings together the information an AI system needs to perform a specific task. It can include structured data, semantic definitions, relationships between entities, historical information, real-time signals, business policies and supporting evidence. It does not necessarily mean creating another physical copy of the enterprise data. In many cases, it can be a governed layer that composes information from existing data products and other enterprise systems when an agent needs it.&lt;/P&gt;&lt;P class=""&gt;For the fraud investigation example, the context will include&amp;nbsp;&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;Customer information and history&lt;/LI&gt;&lt;LI&gt;The transactions and accounts involved&lt;/LI&gt;&lt;LI&gt;Merchant characteristics&lt;/LI&gt;&lt;LI&gt;Device relationships&lt;/LI&gt;&lt;LI&gt;Recent transaction patterns&lt;/LI&gt;&lt;LI&gt;Previous fraud cases&lt;/LI&gt;&lt;LI&gt;Current fraud signals&lt;/LI&gt;&lt;LI&gt;Applicable fraud policies&lt;/LI&gt;&lt;LI&gt;Evidence supporting the recommendation&lt;/LI&gt;&lt;/UL&gt;&lt;P&gt;The important point is that the agent does not need the entire enterprise. It needs the &lt;STRONG&gt;right context for the decision it is making&lt;/STRONG&gt;.&lt;/P&gt;&lt;P class=""&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="1_CUWN7KstAGcXAcnxfeMy4Q" style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30581i8CD19118191A76E6/image-size/large?v=v2&amp;amp;px=999" role="button" title="1_CUWN7KstAGcXAcnxfeMy4Q" alt="1_CUWN7KstAGcXAcnxfeMy4Q" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;H3&gt;&lt;U&gt;Context is more than Retrieval&lt;/U&gt;&lt;/H3&gt;&lt;P class=""&gt;The first instinct when discussing AI context is often to think about &lt;STRONG&gt;RAG&lt;/STRONG&gt;. RAG is an important capability. It allows an AI system to retrieve relevant information before generating a response. But enterprise decision-making requires more than retrieving documents. An agent may need to understand that a particular transaction belongs to an account, that the account belongs to a customer, that the transaction was performed using a device associated with several accounts, and that the merchant has a particular risk classification. That is not simply a document retrieval problem. It is a problem involving &lt;STRONG&gt;entities, relationships, semantics, history, current state and rules&lt;/STRONG&gt;.&lt;/P&gt;&lt;P class=""&gt;Context may come from structured data, documents, knowledge graph, APIs or operational systems or real-time events. The Context Layer is therefore not a specific technology. It is an architectural capability for composing the information required by an intelligent system.&lt;/P&gt;&lt;H3&gt;&lt;U&gt;Where Ontology&amp;nbsp;fits&lt;/U&gt;&lt;/H3&gt;&lt;P class=""&gt;An ontology can make business concepts and relationships explicit. For example, in a banking space, the enterprise may define relationships between Customer, Account, Transaction, Device and Merchant. That semantic structure helps an AI system understand that these are not independent datasets. It understands the business relationships between them. Ontology therefore provides part of the semantic foundation for a Context Layer. Ontology and context are not the same thing. Ontology answers questions such as&amp;nbsp;&lt;/P&gt;&lt;P class=""&gt;What are the important business concepts? , How are they related? , What does a concept mean? , The Context Layer answers a different question.&amp;nbsp;&lt;STRONG&gt;What of this information and meaning does the agent need for this particular task? &lt;/STRONG&gt;This distinction is important because it prevents the Context Layer from becoming another attempt to model the entire enterprise upfront.&lt;/P&gt;&lt;H3&gt;&lt;U&gt;From Data Quality to Decision Readiness&lt;/U&gt;&lt;/H3&gt;&lt;P class=""&gt;Data Engineering 2.0 made data quality a fundamental engineering discipline. We ask whether data is accurate, complete, fresh, consistent and governed. Data Engineering 3.0 introduces another dimension. &lt;STRONG&gt;Is the context sufficient for the decision? &lt;/STRONG&gt;Consider the fraud example again.&amp;nbsp;The Customer data, Transaction data, Merchant data and Fraud data may be accurate. Yet the agent could still make a poor decision if an important signal is missing from the context. This is why &lt;STRONG&gt;Data Quality and Context Quality &lt;/STRONG&gt;are different problems. A data product can be high quality while the decision context built from several data products is incomplete. This suggests that data engineering will increasingly need to think about &lt;STRONG&gt;decision readiness&lt;/STRONG&gt;.&lt;/P&gt;&lt;P class=""&gt;Does the agent have the right information?&amp;nbsp;Is it sufficiently fresh?, Are the relationships correct?, Are the business definitions consistent?, Are the relevant policies available? andCan the agent identify the evidence behind its conclusion?&amp;nbsp;These questions sit on top of traditional data quality rather than replacing it.&lt;/P&gt;&lt;H3&gt;&lt;U&gt;Context Layer should be driven by the&amp;nbsp;Decision&lt;/U&gt;&lt;/H3&gt;&lt;P class=""&gt;There is a natural temptation to approach this as another enterprise-wide architecture exercise. An organization may say We need to build an Enterprise Context Layer. The next step becomes an attempt to model every entity, relationship, rule, document and event across the organization. That is where the initiative can quickly become too large.&amp;nbsp;The better starting point is the decision. Check &lt;STRONG&gt;Which agent task would benefit most from better enterprise context? &lt;/STRONG&gt;Then define the minimum context required for that task. For a fraud investigation agent, it means starting with Customer, Account, Transaction, Merchant and Device along with the relationships, historical information and policies required for the investigation.&amp;nbsp;The first objective is not to model the bank. It is to make one important decision more reliable. This is a very different way of approaching enterprise semantics.&lt;/P&gt;&lt;H3&gt;&lt;U&gt;Context should evolve with the&amp;nbsp;Agent&lt;/U&gt;&lt;/H3&gt;&lt;P class=""&gt;The first version of a Context Layer will rarely be complete. And it does not need to be. Start with a defined task. Identify the information required. Build the context. Test the agent. Observe where it lacks information or misunderstands meaning. Then extend the context. This creates a practical feedback loop between the agent and the data platform. The agent becomes a source of evidence about where the enterprise’s existing information environment is insufficient for intelligent decision-making.&amp;nbsp;That evidence can drive the next iteration of the data product or semantic layer. Over time, the Context Layer can expand from one decision to multiple decisions and from one domain to connected domains. The goal is not to build the biggest context layer. It is to build the &lt;STRONG&gt;most useful context for the decisions that matter&lt;/STRONG&gt;.&lt;/P&gt;&lt;H3&gt;&lt;U&gt;What changes for Data Engineering?&lt;/U&gt;&lt;/H3&gt;&lt;P class=""&gt;Data Engineering 3.0 does not mean data engineers stop building pipelines. The fundamentals remain critical. Data still needs to be ingested, transformed, stored, governed, monitored and served. But the scope of the discipline expands. Data engineers will increasingly need to understand how enterprise information is represented for intelligent systems. That includes areas such as Semantic modeling, Ontology, Knowledge graphs, Entity resolution, Real-time data, Context retrieval, Agent tools, Data and AI governance, Context evaluation and Decision lineage.&amp;nbsp;&lt;/P&gt;&lt;P class=""&gt;The data engineer’s responsibility increasingly moves from &lt;STRONG&gt;Can we deliver the data? &lt;/STRONG&gt;to &lt;STRONG&gt;Can an intelligent system understand and use the data correctly in the context of the decision?. &lt;/STRONG&gt;That is a significant evolution in the role.&amp;nbsp;The investments made during Data Engineering 2.0 become the foundation for Data Engineering 3.0. Trusted data products provide reliable information. Data quality provides confidence in that information. Governance provides controls. Lineage provides provenance. Semantic models provide business meaning. Real-time platforms provide current state. Ontologies and Knowledge Graphs can provide explicit relationships.&lt;/P&gt;&lt;P class=""&gt;The next step is to bring these capabilities together around the decisions that AI agents need to perform. Context Layer is the &lt;STRONG&gt;new layer built via DE 3.0 on top of the Trusted Data Foundation enterprises have already created in DE 2.0&lt;/STRONG&gt;.&lt;/P&gt;</description>
      <pubDate>Wed, 02 Sep 2026 13:13:51 GMT</pubDate>
      <guid>https://community.databricks.com/t5/community-articles/data-engineering-3-0-from-trusted-data-products-to-context/m-p/167304#M1525</guid>
      <dc:creator>balajij8</dc:creator>
      <dc:date>2026-09-02T13:13:51Z</dc:date>
    </item>
  </channel>
</rss>

