
4 AI failures in enterprise analytics and how to fix them
Why most enterprises still don't trust AI analytics, and why better models alone won't fix it
The four biggest AI mistakes we made while building our Agentic Analytics Platform
How we fixed each one, and the results teams like Patreon and Arm are seeing today
PUBLISHED:
UPDATED:
Everyone likes to talk about what’s working. Go to any data conference or watch any product launch, and you'll hear about accuracy jumps, adoption curves, and six-figure automation wins.
Reality is less flattering. In a recent survey of VP and C-suite enterprise data leaders, just 7% report org-wide AI analytics adoption, and fewer than 20% are fully confident in AI-generated answers.
Those figures suggest that trust is what’s holding back AI analytics at scale.
For a while, we diagnosed the problem the same way most of the industry does. While building our Agentic Analytics Platform, we assumed that better models, faster reasoning, and more autonomous agents would eventually compound into a system enterprises could trust with real decisions. Whenever something underperformed, that's where we looked for the fix.
It took four expensive mistakes to learn we were wrong.
AI mistakes we made building the Agentic Analytics platform

Mistake #1: Giving the model too much freedom
When we started building, we gave the model complete autonomy. Our first Agentic Analytics product let the model pick its own path to an answer.
This move was deliberate. Every constraint we considered adding felt like betting against the models getting better. We bet on the models instead. And for a while, that looked smart. We treated rules like scaffolding: useful in the early days and meant to come down as the models matured.
Early on, the evidence backed us up. Ask an agent a question that would normally take an analyst a full day, watch it reason through five or six steps, and get something useful back. It looked impressive in demos and we thought we nailed it.
Then we put it in front of enterprises, and they rightly turned it down. The question we kept getting was: “How do I defend this number when the method behind it changes every time I ask?” We didn't have a good enough answer.

The problem was consistency. The same question followed a different path each time. You can't build trust around that, let alone a recurring business process.
What we rebuilt: a deterministic execution layer
Ultimately, enterprises want a solution that produces numbers they can trust, which means knowing where a number came from and whether it will run the same way every time. That kind of confidence comes from rules built into the platform and enforced on every run.
We built those rules into the Analytics Harness. It sits between the model and your data, controlling which definitions and sources the model can use and which path it takes to an answer. The model still reasons within boundaries the system controls, so the same question follows the same route today, tomorrow, and after the next model upgrade.
The takeaway: What enterprises actually care about is trust and governance, and neither of those is something a probabilistic model can deliver on its own. The harness has to provide them.
Mistake #2: Letting each artifact have its own copy of context
Say someone on your team updates a definition; “active customer” now excludes trial accounts. The change is reviewed and approved. Everyone involved does their job properly. But the change only affects new answers. That’s where the trouble begins.
The question someone asks after the change picks up the new definition, sure. But the artifact saved last month is still working with the old one.

Even the most comprehensive context layer in the world won’t help when every dashboard, agent, and embedded experience keeps its own out-of-sync context copy. One definition becomes a growing set of versions scattered across the business, each looking equally authoritative. The only way to know which one is current is to inspect them one by one.
What we rebuilt: context resolved at runtime
Our Analytics Harness resolves every question against one Enterprise Context Layer at runtime. When a question comes in, the harness applies what “active customer” means right then, so every answer uses the current definition regardless of which surface the question originated from.
Resolving everything against a single layer has a second benefit: the harness can see when different sources start defining the same metric differently, flag the conflict, and route it for human review before it ever reaches an answer.
The takeaway: If definitions are baked into fragmented artifacts, no amount of maintenance fixes it. Every dashboard, agent, and widget should resolve meaning from the same governed context at runtime. Otherwise, every new surface creates another version you have to track, update, and verify.
Mistake #3: Having incomplete or missing context
We were in the middle of a pilot with a large enterprise, building out a domain with their team, when it became obvious the answers weren’t good enough. Once we dug in, we traced the problem to incomplete context.
The domain had been assembled quickly, so important definitions were missing or ambiguous, and business users started asking questions before the context was ready for primetime. In other words, the system was reasoning over an incomplete picture of the business.
Diagnosing the problem was easy. Fixing it was hard. Context maintenance needs an owner: someone accountable for whether definitions are complete, conflicts are resolved, and the AI is working from the right version of the business.
What we rebuilt: evals first, then questions
We built an eval framework to constantly keep an eye on context completeness and decay.
Before business users ever rely on the answers, AI Context Engineers build evals to validate the context underneath them. Definitions get clarified, relationships get checked, and context conflicts are caught automatically, while they’re still cheap to fix. The groundwork has been one of the highest-impact investments we’ve made.
The takeaway: A new model can’t compensate for missing business context. Incomplete, ambiguous, or poorly structured context has to be fixed first. Evals are an effective mechanism to monitor the gaps.
Mistake #4: Ignoring the path behind the answer
This one took us the longest to catch, and we still think it’s one of the easiest AI failures to miss: An answer came back correct. The number matched the source of truth, so by every measure we had at the time, the system did its job.
Then we looked at the path it took to get there. The system took a route we’d never have approved. It skipped context it should’ve used, relied on the wrong source, and somehow still happened to land on the right number.
While the result was correct, the process behind it was broken.
A known model weakness was partly to blame. Research on the lost-in-the-middle problem has shown that models can struggle to use relevant information buried in the middle of a long context. In an analytics system, the same weakness compounds because the model is also choosing sources, writing queries, and handling failures along the way.
The same risk appears when a query breaks halfway through. The system has a choice: stop and admit it can’t answer reliably, or keep going with whatever data it managed to retrieve. If it keeps going, the result looks exactly like a complete answer, and the user has no way to tell the difference.
What we rebuilt: governed paths
Now we grade both the number and the path that produced it. Our Analytics Harness puts rules around how a request gets resolved, including what the system does when a step fails, so failure itself is governed.
If the system doesn't have enough trustworthy information to answer, it flags what’s missing, checks its own work, and asks before it guesses.
The takeaway: To know whether your AI system is working, check whether it used the source your team actually trusts and told you when something broke.
Where we are today
Every capability in the Analytics Harness today exists because something went wrong first.
Since those rebuilds, we’ve been testing relentlessly and tightening the system every time we found a new way it could break. The result is a purpose-built Agentic Analytics platform that enterprises now trust with high-impact decisions.
Here’s what it looks like in the real world: Patreon moved more than 80% of their everyday data questions to self-serve within three months, at 95%+ accuracy. Arm cut contract review time in half; WisdomAI’s first-pass review came in at 98.5% accuracy.
Both teams got there using the systems, definitions, and business knowledge they already had. They needed an Enterprise Context Layer and an Analytics Harness that could connect that context, govern how AI uses it, and make it available consistently on every question, from every surface.
Stay tuned: The next and final piece in this series makes the case that specialized harnesses are about to become their own layer of the enterprise stack. Sign up for our newsletter so it lands in your inbox the moment it’s live.
FEATURED RESOURCES

