
Two truths and a lie about context layers
What a context layer is, and the data, business, and governance context inside it
Why you aren’t starting from scratch, and why context needs a development lifecycle
The hidden costs of platform-native context, and four properties to test for
PUBLISHED:
UPDATED:
Context is the AI advantage every data team needs right now. Just go to any data and analytics conference, and they’ll tell you so. Whether it’s context graphs, context layers, context engineering, context management, nearly every booth or session speaker has something to say about context.
Investment is following the hype. In a recent survey of VP- and C-suite enterprise data leaders, 94% of data leaders say they plan to change how they store and manage AI context, and 100% have or plan to hire a dedicated context engineer in the next 2 years.
Ask those same leaders what a context layer is, and you’ll see a wide variety of answers: semantic layer, data catalog, undocumented institutional knowledge that lives in an analyst’s head. That last answer was common among half of the respondents we surveyed.
With all the hype, context is quickly becoming the most misunderstood word in AI and data. I wrote this blog, the second installment in a five-part series, in hopes of resolving that.
If you’re new here, I suggest starting with the first article in this series: How to build a modern analytics architecture.
What is a context layer?
A context layer holds the business meaning a model can't infer from raw tables: metric definitions, entity resolution, governance rules, history records, and usage patterns. It's the governance layer sitting between your analytics surface and the AI that queries them. It’s what an Analytics Harness calls to deliver that meaning at query time.
Most people hear "context layer" and picture a semantic layer with a new name, or a data catalog with an MCP bolted on. Both are context inputs but neither captures the full scope.
Take a simple question: Did ticket resolution get faster last month? A semantic layer can tell you how resolution time is defined. A catalog knows where the ticket data lives. Neither knows what “good” means for the metric or that last month’s drop is actually tied to a queue migration.
That’s why your context layer has to encode more than just the data context. It also has to carry the business context: both operational knowledge like constraints and exceptions, but also your security and governance rules that determine access and permissions.
Here are the different components in a context layer:

Data context: What does the data mean?
Data Context is the technical understanding of the data itself, covering what it contains, how it is structured, where it comes from, and how it is calculated. It breaks down into two categories:
Descriptions: Documentation that explains each dataset, table, and field. They cover what the data contains, where it comes from, how often it refreshes, who owns it, and any known limitations.
Models: Structural representations of the data that define its sources, entities, tables, keys, relationships, and lineage. With a model in place, an agent can see how the data is organized, at what grain, and how records connect across systems.
Semantics: Shared definitions of business terms and metrics, paired with the logic that calculates them the same way every time. Each definition captures what a business term means and how it is measured, grouped, filtered, and aggregated.
Business context: How does the business work?
Business Context is the organizational understanding of the data. It explains how concepts relate, how the company actually uses its numbers, and what history sits behind them. Here’s the further breakdown:
Ontologies: Formal maps of business concepts and the relationships between them, independent of how the data is stored. An ontology specifies how entities such as customers, accounts, contracts, and products connect to one another and roll up into larger groupings.
Institutional knowledge: The collective expertise, experience, and know-how that a company and its people build over time. It captures which sources the business trusts, why certain numbers behave the way they do, and which past decisions and exceptions shape how the data should be read.
Governance: What is the agent allowed to do?
Governance context is the policy understanding of the data. It defines who may see the data, what may be done with it, and how those limits are enforced.
Rule-based controls and guardrails: The policies and runtime checks that govern what an agent may do. These guardrails specify which data may be combined, which decisions an agent may make independently, and when an expert must step in.
Row and column-level security: The access controls applied at query time to limit what an agent can retrieve. Row-level rules filter records by attributes such as region or team, while column-level rules mask sensitive fields such as salaries or personal identifiers.
Truth 1: You aren’t starting from scratch
For years, teams have bridged the context gap manually. An analyst knew which table to trust and quietly corrected the query before anyone saw the answer. To scale AI analytics beyond a handful of power users, you can no longer rely on that hidden correction. That’s why you have to create a single, standardized place for context to live.
Much of the context you need to add into your context layer already exists in your stack. But it's spread across dozens of fragmented systems, which is why AI fails to call it effectively. But the good news is, because the context already exists, you don’t need to start building it from scratch.

Imagine you have a dbt model that holds the definitions. Your Databricks warehouse stores your data as well as the relationships within Unity Catalog. Other business apps like Confluence contain business logic. By connecting your context layer to these existing sources, you have a working baseline.
This isn’t a new concept. Most BI tools have let you connect to semantic layers and catalogs for a decade or more. However, there’s one misconception here that’s worth addressing.
More context isn’t always better
You don’t have to capture all the available context to make AI usable. Actually, more context can make AI less accurate, introducing conflicts and out-of-date information that creates more opportunities for agents to trip up, while increasing the token costs.
In reality, a smaller amount of intentional context can get you much closer to an accurate answer. Just ask Arm:
“We thought more context was going to improve answers. What we’ve actually found is the opposite. Less context, more very specific, very refined context.”
— Thomas Smith, Procurement Transformation Lead at Arm
Start with one business domain, add the critical details, and see how far that will take you. Once you’ve built your initial context layer foundation, it’s time to test, refine, and launch.
Truth 2: Context should be treated like a development lifecycle
I often see teams spend 6 months doing a massive initial context build. They launch it, make a splashy internal announcement, and then they don’t have much ROI to show. Six weeks later, they wonder why accuracy tanks. When Anthropic’s own data team did this, they saw accuracy drop from 95% to 65% in a single month.
Obviously, I’m not the only one saying this, which is why I’m calling it a truth. But it’s still a trap I see many leaders falling into. In my opinion, it’s due to the leftover mindset from the clean data era and semantic layer craze. I understand why leaders think that if you get everything perfect up front, you can simply launch, set it, and forget it. But that's not how context works.
Your business is constantly evolving, so you need a system that helps your context evolve alongside it. Engineering teams already have a term for this: the Software Development Lifecycle (SDLC). Code gets written, reviewed, versioned, tested, promoted, monitored, and rolled back when it breaks. Data teams need their own term for this, something I call the Context Development Lifecycle (CDLC).
The CDLC practice covers the end-to-end path of context. Build it from what your organization already knows, govern who can reach it, adapt it as the business moves, and publish it as a versioned source of truth.
As the business uses it, query logs (as a source of context) can help you understand what’s missing, feedback loops will show you where conflicts exist, and context engineering can help you close any gaps so that the context stays up to date as your business changes.
Here’s the lie: It doesn’t matter where you store your context
With so many vendors selling similar versions of the same product, it’s easy to fall into the trap of going with what you know. Say your warehouse offers a platform-native context layer, and it’s easy enough to adapt your context to their format.
As you use it, every correction your team makes improves the warehouse’s understanding of your business. The system will get better; I want to be fair about that. But in the end, it will cost you much more than the licensing fee.
The hidden costs of a platform-native context layer
Ingestion: Because most vendor platforms read in their own format, all of your original context sources need to be rewritten in the vendor’s language and shape. You already modeled this in dbt, and now you have to model it again in Snowflake, Looker, or whatever other vendor you choose.
Maintenance: When a column gets renamed upstream, somebody on your team has to manually apply that change on the vendor side. And if your estate spans more than one platform, you may be doing that work multiple times.
Switching: This cost grows over time. Every context infusion or correction is another investment you make in the product. Now imagine the vendor gets acquired, pricing changes at renewal, or you acquire a company running on a different warehouse. Those investments in your vendor-locked context won’t travel with you. It’s not a reality you considered when starting, but it’s a cost you pay when you leave.
Now picture another reality where the context lives in an open format in your own repository, versioned like code, and readable by any tool you point at it. Every correction and investment is improving valuable IP that you own and get to keep.
That’s the choice: Either your context lives inside one vendor's product, or it lives in a layer you own that any tool can read.

What to look for in a context layer
Here’s the question worth asking before you sign anything: Do you have full ownership of your context, and can you export the whole thing in a form another tool could actually use? If not, I encourage you to invest in a different solution.
Or if you’ve already started building your context here, consider whether it’s worth cutting your losses early. The value of owning your context will only compound as your business grows in complexity and AI analytics becomes more ingrained in your day-to-day workflows.
Here are 4 properties to look for and test against in a context layer:
Portable: The context sits in an open format, or, at minimum, a documented one, that can be read outside the product that created it.
Inspectable: Any model or person on your team can open it and immediately understand what it means without needing a vendor's interface to interpret it.
Versioned: Every change is reviewable, reversible, and attributable, the same way you'd treat any other enterprise IP that matters.
Additive: The work you have already done should carry forward. Your catalog, semantic layer, dbt models, policies, and documentation should become inputs to the context layer, not migration projects. If adopting a tool means rebuilding years of business logic inside its system, it’s a vendor lock-in trap.
Closing argument
None of this is an argument against buying software. Buy the models, the orchestration, the retrieval, the evaluation, and the infrastructure that makes the system work. Just don't confuse buying the machinery with handing over the memory.
Your competitors can buy the same models, warehouses, and tooling. But they can't buy the accumulated knowledge of how your business actually works. That's the asset your context layer captures, and it's why ownership matters. Whoever builds it, and whatever you build it with, your context should stay portable, inspectable, and entirely yours.
The next installment in this series examines why context alone can’t guarantee accuracy, and the role a purpose-built Analytics Harness plays in applying context the moment a question is actually asked. Sign up for our newsletter so the piece lands in your inbox right when it’s live.
FEATURED RESOURCES

