Why do enterprises need a specialized harness for data analytics?

Summary

Summary

  • Why enterprises need a specialized agent harness for analytics instead of a general-purpose or coding harness

  • The failure modes unique to data agents, from silent errors to inconsistent answers

  • How an Analytics Harness interprets questions, selects context, enforces access, and verifies results to deliver consistent, trusted answers

Kapil Chhabra

Kapil Chhabra

CPO

IN THIS ARTICLE

Like this article?

Subscribe to our newsletter to receive more educational content

PUBLISHED:

UPDATED:

Everyone now accepts that models make mistakes. A context layer helps, but reliability comes from the harness around the model, the system that controls what the agent is allowed to do.

The question that hasn’t been settled is whether the same harness works everywhere. The answer is no. You can't take a harness built for a coding agent, add a few data tools, and expect reliable enterprise-grade analytics. You also can’t take a harness built for analytics and expect it to work flawlessly and efficiently for legal processes.

Coding agents and data agents break in different ways, so they need different controls. Treat them interchangeably, and you get systems that look promising right up until someone makes a real decision on a wrong answer. It's one of the reasons fewer than 1 in 5 enterprise data leaders in our survey are fully confident in the answers their AI gives them.

This is the final piece in my series on modern AI data architecture. If you haven’t already, you can catch up on the other articles here:

  1. The modern analytics architecture

  2. Two truths and a lie about context layers

  3. Agent harness explained

  4. 4 AI failures in enterprise analytics

I want to wrap up by making the case for why enterprise analytics needs a specialized harness, so you can make the right call for your own stack.

The case for a specialized Analytics Harness

On their own, models can reason and write. But to do real analytics work, they need a specialized harness: the tools and guardrails that turn a model into an agent that can reliably finish tasks. 

Many agent harnesses, one model

General-purpose harnesses, like Claude Cowork and ChatGPT Work, are built to handle a bit of everything, from drafting memos to summarizing calls. For most everyday work, these tools are perfectly capable.

Highly skilled work is a different story. That’s why it requires a specialized harness that can run inside tools like Claude or ChatGPT or operate independently. 

Coding gives us the clearest example. Claude Code, Codex, and Cursor are specialized harnesses for software development. They show what a reliable agent loop looks like: the model makes a change, the tests run, and any failures go back to the model so it can correct itself before an expert reviews the final diff. A developer could ask a general-purpose assistant to write code, and it could do a decent job. But for serious work, they turn to Claude Code over Claude Cowork because it’s built specifically for how software gets made.

Specialized harnesses are showing up in other skilled areas beyond coding. Legal teams use Harvey and Legora, and support teams use Sierra and Decagon. Each is a specialized harness built with the skills and safeguards its domain requires. 

Analytics is no different. Answering a data question correctly and consistently takes specialized skills, so the agent doing it needs a specialized harness. A general-purpose assistant isn’t a bad tool, and a coding harness isn’t either. They’re just the wrong tools for the specific job of enterprise analytics; they were built for something else. 

Why can't a general-purpose harness handle analytics?

Every specialized harness is built to catch the specific failures in its domain. In coding, the reliability signals already exist. A compiler rejects bad syntax, and failing tests make problems obvious. Sometimes programs crash outright. The harness’s job is routing those signals back to the agent so it can try again. 

AI analytics fails in different ways. A query can run perfectly and still be wrong. The agent can pick the right tables and definitions and return a polished number in minutes. The number might even be correct, but the logic the agent used to get there can still be wrong.

A successful run tells you almost nothing about whether the process was reliable or whether the agent just happened to land on the right result. Ask the same question tomorrow, and it might take a different path and return a different number; you’d have no way to know which one to trust. Unlike coding, there's no compiler to reject it, tests to fail, or crashes to notice.

A general-purpose harness can’t catch any of this. A specialized harness, built specifically for analytics, fixes that failure mode with a deterministic execution layer that applies the right context at the right moment, enforcing the same rules wherever the question is asked. Without it, you still have an agent making judgment calls about which data and definitions to trust and which rows a user is allowed to see. Every one of those calls can change from one run to the next.

How does an Analytics Harness work? 

A specialized Analytics Harness supplies the data-specific logic an agent needs to answer the same question the same way today and next year.

The Analytics Harness handles the following jobs:

  • Interpreting the question: resolves what's actually being asked, including which entity, grain, and time window, before any SQL gets written. If the question has been answered earlier and has an approved SQL, it just uses that for consistency, speed, and token efficiency.

  • Selecting context: chooses the right subset of schemas, metric definitions, and historical queries that govern the question and feeds them into the model's limited context window without overwhelming it. Too little context leaves room for hallucinations. Too much context, on the other hand, confuses the model and bloats that cost.

  • Enforcing access: keeps the model inside the tables, and only fetches the rows the user is permitted to see. If certain columns need to be masked for PII or other compliance reasons, the harness takes care of that.

  • Optimizing the path: picks the most efficient sources, queries, and execution order up front so the agent spends its tokens answering questions instead of hunting for data.

  • Planning and executing the query: maps the question across every system that holds relevant data, flags missing or ambiguous context, and runs the plan while carrying intermediate results forward with their definitions intact.

  • Verifying and repairing the result: checks the output against what the context says should be true before any result is produced. When a check fails, it traces the failure back to its cause, whether that's a definition, a stale source, or an ambiguous question, and routes the fix for review.

analytics architecture

In practice, a general-purpose harness derives answers from whatever context happened to be in the window, so every correction you make fixes one answer and goes no further. A specialized harness handles those jobs deliberately, against a governed context layer, and encodes fixes so mistakes don't recur. 

Where I might be wrong

There are a few ways my argument could play out differently.

Specialization could turn out shallower than it is right now

If general-purpose harnesses build native governance, access, and verification, the specialized analytics layer shrinks to little more than a connector.

I don't think that's where the industry is heading. The hard part of analytics is that most institutional knowledge is still undocumented, and AI can’t work its way to a definition your team hasn't written down. 

The labels may change 

Harness is engineering language, and it is unclear if most CIOs will use it. They may just talk about governance or trusted analytics instead.

The architecture will outlast the word. Plenty of infrastructure categories have gone this way before: the system stays essential, but the language gets simpler as the market matures. Most vendors don’t use the word “middleware” anymore, but every serious stack still has something doing that job.

The market could remain divided

It’s also possible that no clear category forms between building your own harness and relying on built-in controls. 

The strongest data teams may keep building their own harnesses because they have the talent, scale, and company-specific needs to justify it. Everyone else may simply rely on whatever controls come built into their warehouse, BI platform, or the general-purpose AI stack. 

Final thoughts

Context gives you accuracy, and a harness gives you consistency. 

Together they close the loop; every correction flows back into the context so it compounds over time. I invite you to put this to the test: 

Open two chat windows in Claude, and ask a colleague to do the same. Then ask the same data question in all four windows, ideally about a metric your board sees. Check whether each one takes the same steps, pulls in the same context, uses a similar number of tokens, and gets you the same result. If the answers match, you're in better shape than most. If they don't, you've found the problem before it surfaced in a board meeting.

That test will tell you more about where you stand than any architecture diagram, including mine.

Here are a few key takeaways from the series:

1. Semantics are necessary, but not sufficient. You need context for accuracy.

2. Context has multiple layers and must be managed with care.

3. Context should be portable and ownable.

4. The Analytics Harness is equally important for other aspects of trust beyond accuracy: consistency, governance, explainability.

5. A general-purpose harness is not the right tool for a specialized job.

6. The specialized harness is a moat for companies in the AI era.

If you want to see what a governed context layer and a purpose-built harness look like running on your own data, book a demo and we’ll show you.