
How to build a modern analytics architecture that delivers accurate AI analytics
Why models and context layers alone can’t deliver trusted AI analytics at scale
What an analytics harness does between the question and the answer
The five layers of a modern analytics architecture, and how to build yours
PUBLISHED:
UPDATED:
As agents become more ingrained in our day-to-day workflow, they're taking on harder jobs: analyzing data across both structured and unstructured sources; interpreting complex data joins, definitions, and access rules; creating artifacts in various formats; and coordinating actions across a web of connected business systems.
This raises a question most analytics architectures aren't yet built to answer: How do you give an agent what it needs to answer reliably, securely, and accurately at scale?
This is the first piece in a five-part series on the architecture behind trustworthy AI analytics. This post lays out the full picture: what most architectures get wrong, why context alone isn’t enough, and where the harness fits. Later posts go deeper on each piece.
The baseline: how most AI analytics architectures work

Insight delivery surfaces, like agents and chat, sit on top of your preferred LLM model, which is connected to your data sources through MCP or dedicated plugins. The model is fully capable of interpreting general natural-language questions and finding answers in connected data sources.
However, a generalized model is not built to interpret natural-language questions based on your business-specific definitions. They also can’t plan the most efficient path to an insight, govern access and security, offer clear traceability and verification, or create feedback loops that ensure insights are accurate and repeatable over time.
All of this has created a significant trust problem. In fact, less than 20% of surveyed data leaders fully trust their AI-generated insights, which is a key reason why only 7% of enterprises have scaled AI analytics across their business.
For the past few years, leaders have been working on AI readiness, testing out the hottest new models, and hoping one of them would develop deep enough reasoning capabilities to fix the accuracy problem. The models have gotten there: today’s are faster, cheaper, and more capable at advanced reasoning than ever. But they still can’t deliver the accuracy enterprises need to trust AI analytics at scale.
The fix most data teams reach for: context layer
Most data leaders I've talked to have realized this limitation. Seduced by the hottest new tech in the market, they’re adding a fourth layer to their analytics architecture: AI context.

Many of these layers are an evolution of the “modern data stack,” such as semantic models, data catalogs, or tools built into their warehouse or BI tool. Others are opting for AI skill files or adopting emerging, more robust context management systems that help them continuously manage their context.
In each of these cases, the context layer is where data teams encode definitions, entity relationships, business rules, and access policies along with the institutional knowledge that has only ever lived in people's heads.
This is one of the most impactful investments a data team can make right now. I don’t want to undersell the context layer’s value. Helping the model understand how to interpret your business meaning will provide an immediate, measurable payoff.
However, context alone won’t solve AI’s accuracy problem. It doesn’t force the AI to select the right definition for the right question, apply it the same way twice, or check that the answer respects who’s asking. That’s a different job done by a separate piece of software: the Analytics Harness.
Context is a partial solution to a complex problem
Just like perfectly clean data doesn’t mean AI knows how to interpret it, you can have perfectly organized and centralized context and the agent may not use it, not in every case, and not consistently. So, you can make this huge investment in context, and still have no idea if your agent is using it consistently.
Here's where the gap shows up:
Selection: Your semantic model knows how churn is calculated. It doesn't guarantee the agent reaches for that definition when churn comes up.
Enforcement: Access policies describe what data a person may see. They don't stop an agent from composing a query that joins past them.
Verification: Nothing in a semantic layer or catalog checks that the SQL ran correctly, or that the number is calculated against the definition.
Here's the distinction I keep coming back to: Context gives you accuracy, but it doesn't give you consistency. At scale, one without the other isn't worth much.
Where accuracy meets consistency: analytics harness
Between the model and the context, there needs to be a specialized control system that ensures the model consistently, efficiently, and securely applies your context. That control system is the harness.
What is an analytics harness?
According to the Oxford Learner’s Dictionary, harness means to control and use a natural force, energy, or resource for a specific purpose. In the world of AI models, a harness is the code that decides what the model sees, what tools it can call, what it's allowed to return, and what happens when it gets something wrong. I.e. to control and use the powerful generative models for a specific purpose.

A general-purpose harness like Claude CoWork is capable of a range of business jobs, like writing emails, preparing briefs, and brainstorming new product ideas. A coding agent harness like Claude Code is optimized to write code and manage files.
For complex enterprise analytics tasks, the harness takes on an even larger role: ensuring the model correctly interprets business questions for the specific person asking, planning and executing complex queries across disconnected data sources, optimizing those queries to control token and compute costs, and verifying that the resulting answer can be trusted.
If it can’t be, the harness needs to repair the answer, not just this once, but by encoding the fix so the same failure doesn’t happen again. And it has to do all of this within the row- and column-level access rules and governance standards your company has already set.
Real-world example of the analytics harness in action
That’s a lot of theory. I know. So let’s look at a specific example.
Say someone asks their agent, "What is our regional revenue this quarter?” Everything the harness controls is every step the model takes between the time it receives the question and delivers the answer.
Here’s what’s happening behind the scenes.
Interpretation and selection
The harness must first apply your context to determine the analytical intent behind the question. For instance: Is revenue net or gross? What regions does the business operate in, and how are they defined? What is the company’s quarterly schedule?
It also has to enforce your governance rules. A manager in EMEA should only see that region’s data, while a Global VP sees results across every region; two different people with different permissions, asking the same question, correctly getting different answers. If instead both users shared the same access level, they’d get identical answers, as they should. Either way, the response comes with a natural-language breakdown of the context behind it, so there’s no confusion about which version they’re looking at or why.
Execution
Based on the context the harness receives, it will then plan an optimized SQL, Python, or RAG query to pull the right data from its various connected sources. The benefit here is twofold.
By eliminating initial ambiguity and injecting the right context at query time, a well-built harness reduces the amount of work the model has to do to reach an answer. Instead of exploring the schema and rederiving context on every question, the harness resolves what it needs once and then reuses it as more questions come in. That shows up directly in cost: less token spend per question, and less strain on the warehouse or database doing the actual computation.
The same optimization that lowers costs also improves accuracy and consistency. When context is applied correctly and verified queries resolve the same way every time, regardless of who asks or which interface they’re asking from, the ceiling on accuracy moves higher. That’s a different outcome than what a general-purpose model can deliver without a harness built specifically for analytics.
Failure handling
Accuracy from a purpose-built harness is a floor that keeps rising as the system learns over time. There are three main ways the harness works to ensure repeatable accuracy:
After the query is executed, the harness reviews the answer to decide if the result actually answers the user’s question. If it doesn’t, the harness rewrites the query and tries again, drawing on past questions in the conversation, curated examples, and documented business logic in the context layer. This knowledge is packaged up and delivered alongside the answer, so the user can trace the answer back to the context and execution path it was derived from.
In the event it doesn’t have the information it needs to generate an accurate answer, say there is a conflict or missing context, the harness must be designed to never guess. It must pause to ask short, structured clarifying questions instead of delivering confident answers to the wrong questions.
In the event an inaccurate answer does reach a user, they can write back to the model, letting it know that feedback. The harness should then route the feedback back into the context layer, triggering a notification to the data team that expert review is required.
That last piece is an important part of the picture. While the harness is the connective tissue that makes the context layer multi-modal, the self-learning feedback loop is what helps you keep context fresh.
The self-improving loop
Your context starts decaying the moment you encode it. Anthropic's own data team watched accuracy fall from roughly 95% to 65% in a single month because their documentation couldn’t keep pace with the business.
That’s why you need a second mechanism running in parallel. The harness uses your context to answer a question. A self-improving loop uses those questions and feedback to improve your context.

This work needs a dedicated system watching for it. Something has to automatically monitor chats and user feedback to catch questions where the model had to guess, and flag it the moment essential context is missing or needs clarification. When a user or the model itself proposes changes, those adjustments should route to expert review before it’s encoded, so every definition stays versioned through every revision, not just the first one.
The important structural point is that this loop writes into the context layer, not into the harness. So every correction and clarification creates new context and intelligence, ensuring the enterprise’s context improves with each question instead of decaying.
How to build the new analytics architecture for your business
By 2028, Gartner predicts that 60% of agentic analytics projects that rely only on MCP will fail because they lack a consistent semantic layer.
MCP solves connectivity, giving agents a standard way to reach your tools and data. But access alone doesn't tell a model what a metric means, which source to trust, or whether either one is still current. This is why the full architecture has to treat context as a lifecycle.
All of this leaves you with five essential levels of AI analytics architecture:
Interfaces on top,
A harness handling governed execution,
A context layer supplying business meaning,
Your data staying where it already lives, and
A loop keeping that context accurate as the business changes.

If you’re looking to build reliable AI analytics and implement this architecture in your own organization, here are three discussions worth having with your team.
1. Own the context layer
It encodes how your business thinks, which makes it one of the most valuable things you'll build this cycle. Keep it portable, inspectable, and version-controlled, not trapped inside a walled-off platform or a single warehouse. If you already have a catalog or a semantic layer, make sure your context layer builds on it rather than replacing it.
2. Keep the interfaces on top interchangeable
The business should be able to ask questions wherever people already work, and increasingly that means an AI assistant rather than a BI tool. General assistants are good at almost everything, which is exactly why they're not tuned for anything in particular. Let the assistant take the question and make your harness do the analytics work behind it.
3. Ensure usage signals feed back into the context layer
Every correction someone makes should land as a signal in the context layer you own. That crowdsourced knowledge, which lives in users’ heads and comes out through their interactions with agents, is a goldmine most companies let go to waste. Catch the drift before it compounds so context keeps pace with a business that never stops changing.
Questions worth asking before you build
Those are my arguments. Now, consider how this compares to your own architecture:
When two people with different permissions ask the same question, what guarantees they get appropriate answers?
When a definition changes, what propagates that change everywhere it's used?
When a query fails, who finds out, and what happens next?
When someone corrects an answer, where does that correction get saved, and who owns it?
Six months from now, what will have made your context better, and how would you know?
Throughout the rest of this series, I’ll go deep into each of these: the context layer, the harness, and the feedback loop that keeps them both honest. I’ll also share what we’ve learned building this for large enterprise companies, including what broke along the way.
Next up, a deep dive on the context layer: where it lives, who owns it, and how you manage it as the business keeps moving. Sign up for our newsletter to make sure it lands in your inbox once it’s published.
FEATURED RESOURCES

