Getting the Same Answer Twice: Evaluating and Improving Consistency in Analytics Agents
Imagine on Monday, you ask a Conversational BI tool for Q2 revenue in EMEA, and it tells you $4.2M. On Thursday, you ask the same question and get $3.9M.
It doesn’t matter whether the true number is $4.2M, $3.9M, $1.0M, or $200M. Atleast, one of these answers is wrong, and you can easily know it without checking a single source. Catching a wrong answer normally takes an expert with access to ground truth. Catching a contradiction arising from inconsistency takes only common sense.
The bigger cost comes next. The product has shown you that it can’t be trusted, so from now on, you will doubt every number it shows, including the correct ones.
Consistency (agreement) is not correctness (accuracy)
If EMEA’s actual Q2 revenue was $4.0M and you get the answer $5.0M every single time, that is 100% consistent and 0% correct. This distinction is important because as we optimize to make our agent more consistent, we can also change the performance of the agent on other metrics such as correctness, latency, and cost. For example, a change that makes an agent more consistent can also make it more consistently wrong.
Most existing research doesn’t help our use case either. Thinking Machines Lab found that setting temperature to zero doesn’t fix model inconsistency because even greedy decoding is not deterministic on most inference stacks. Another tool for variance is self-consistency: sample several answers and take a vote.
However, our situation is different. The production agent makes many calls. It plans, explores tables, writes SQL, shapes the output, and all these small differences snowball into different answers at the end. It also runs on top of an Enterprise Context Layer (business definitions, trusted example queries, domain rules), which gives you levers a model-only approach doesn’t have.
I joined WisdomAI in June as one of its first engineering interns, through the Kleiner Perkins Fellows Program, and spent my summer building an eval for consistency as well as trying to improve it throughout our product. A lot of what we tried didn’t work, and we hope that these lessons can be helpful to anyone building with agents.
How we measure consistency
We already know how to measure if our agent is right (correct). As we described in our earlier post: run the agent against a benchmark with reference SQL attached, compare the output vs. expected, and report a number.
Measuring consistency takes a different approach. We ask the same natural-language question several times in our production system, exactly as a user would, and check whether the answers match. The score is the share of runs that landed on the most common answer, so if three out of four runs agree, the question scores 0.75.
The one design choice worth calling out is what “match” means. We check it two ways: do the runs return the same data, and does an LLM judge think the SQL computes the same thing? This leads to our first lesson.
Lesson #1: Look at the disagreements before you fix them
Our first instinct was that the agent was reasoning its way to different answers. Then we took the time to read through the disagreements manually.
About 75% of them weren’t different answers at all. The runs returned the same rows with a different set of columns. One run would add billing_currency, another would drop it, a third would reorder everything. On a question like “current MRR for one account,” the data check saw four different answers and the SQL judge saw one. Both were right: the logic was identical, but the user would have seen 4 different tables.
That’s why we check agreement both ways. If the logic matches and the data doesn’t, you have a presentation problem, not a reasoning problem, and those need very different fixes. With only one view, we would have spent the entire summer trying to make the agent think more carefully about something it was already getting right.
The remaining 25% were actual differences in interpretation: does “the last two years” mean calendar years or a rolling window? Are internal test accounts included? These are the disagreements most people imagine, and they were the minority.
Lesson #2: Fix the divergence where it starts
The textbook fix for inconsistent LLM output is self-consistency: generate several candidates and take a majority vote. We tried it first (we found n_candidates = 3 to be a good balance). On our older single-step SQL generator, it worked, at almost 4x the cost. On the multi-step agent, it did nothing.
We found the reason once we traced where the differences started. Our agent plans before it writes SQL, and that plan decides the shape of the answer: which fields the query selects, and so which columns show up in the table the user gets back. This is the result set itself, not the chart on top of it.
Ask for MRR by account, and one plan asks for account name and MRR, while the next also asks for account ID, plan tier, and billing currency. The planner made that choice differently every time (in one sample, the same metric was specified 16 different ways across 16 runs), and the SQL step followed the plan almost perfectly. By the time you're voting on SQL, the divergence already happened several steps earlier.
Voting works when the variance lives in the last step. In a multi-step agent, it usually doesn’t. The fix that did work was far simpler: a prompt change telling the agent to select only the fields the question actually needs.
Lesson #3: Memory helps until it goes stale
The single most effective thing we tested was giving the agent a memory: past questions paired with the SQL that answered them, retrieved as trusted examples. When we seeded the eval’s own questions, agreement jumped more than with any other strategy. That’s cheating, of course, since the agent is shown the answers, so we treated it as a ceiling rather than a result.
The more interesting result came when we seeded examples from real past traffic instead. Agreement got significantly worse. The example pool has a fixed size, and the older examples were pushing out newer, better ones. Retrieval capacity is a budget. Anything you add to an agent’s memory competes with what’s already there, so what you evict matters as much as what you store.
Lesson #4: Decide once, and be strict about what you decide
For the interpretation disagreements, we built a layer that spots a question’s hidden assumptions, resolves them once, and reuses that resolution every time the question comes back.
It's different from asking the user a clarifying question like "April of which year?" It can handle assumptions the user may not know exist, like how revenue is calculated in this particular company. So it only runs on tasks where there's no one to ask: scheduled reports, API calls, and background agents.
Most of the work was deciding what the layer should not do. Two rules made the difference:
Don’t decide what you’re not entitled to decide. The layer can freely settle presentation questions. It only settles which rows get counted when the company’s documented knowledge says so, never on its own judgment. A wrong answer there doesn’t just make one answer wrong: it makes the answer wrong every time.
A vague decision is worse than none. Each resolution has to name exact tables, columns, and filters, or it gets thrown out. “Exclude test accounts” with no definition just gives the agent something new to interpret differently.
With these rules, logic agreement went up significantly and correctness didn’t move at all. We also learned to be skeptical of our own safety rules: an early rule that discarded anything involving dates had been hurting consistency without protecting correctness, and removing it helped.
On its own, though, the layer mostly improved agreement on logic, not on the table the user sees. Paired with the column-trimming prompt from Lesson 2, it was our best result in five straight experiments, because each fix handled one of the two kinds of disagreement.
Lesson #5: Sometimes the cheapest fix is the decoder
We also ran the eval against open-weights models, including Qwen3-30B on a single GPU. Switching it to greedy decoding, which always picks the most likely next token, raised its exact-match agreement by 0.17 in that run.
Accuracy went up too, from 89.6% to 95.5%, though we'd treat that part as a single data point rather than a general result: greedy decoding is not a reliable way to improve accuracy, and in many settings it does the opposite.
The consistency effect is the durable one, and it makes sense, since sampling noise isn't ambiguity, it's just noise. Greedy decoding won't make output fully deterministic either, because inference stacks have their own sources of non-determinism. Still, if you can control your own decoding, it's a cheap knob to check before building anything more complicated.
Results
Every lever was measured against its own baseline inside the same dispatch. Agreement deltas are paired per question on a synthetic domain, correctness is double-checked with a public e-commerce dataset, and cost is generation spend per run relative to that dispatch's baseline. A dagger (†) marks a confidence interval that excludes zero.
Every lever was measured against its own baseline inside the same dispatch. Agreement deltas are paired per question on a synthetic domain, correctness is double-checked with a public e-commerce dataset, and cost is generation spend per run relative to that dispatch's baseline. A dagger (†) marks a confidence interval that excludes zero.
Strategy | Pipeline | Agreement (paired Δ vs baseline) | Correctness (goldens, 16 q) | Cost vs baseline | Notes |
Self-consistency voting | Chat agent | Null / negative (−0.03 to −0.09) | Not run | ≈3× baseline | Divergence stems from question breakdown prior to final SQL generation. |
Curated context | Chat agent | Null (+0.004 full, −0.066 LLM) | Not scored | +$0.05 / question | Richer context provides more columns without consistency gains. |
Silver reviewed queries | Chat agent | +0.346† full, +0.237† LLM | Not scored | Flat | High ceiling when seeded with target questions; stale real traffic examples degrade accuracy. |
Disambiguation layer | Chat agent | +0.184† LLM, +0.066 full | 0.000 (flat) | +2% (+$0.01 / question) | Improves judge/logic agreement without affecting correctness. |
Column doctrine, alone | Chat agent | +0.026 full, 0.000 LLM | −0.031 | −9% (−$0.06 / question) | Slight movement on data graders; needs combination with disambiguation. |
★ Disambiguation + doctrine | Chat agent | +0.189† LLM, +0.149† relaxed | −0.016 | −17% (−$0.11 / question†) | Best performing strategy. Fixes logic and presentation divergence simultaneously. |
What each one is, briefly:
Change | What it does |
Curated context | Clean up the domain’s business definitions and metadata, accepting every fix our tooling suggests. |
Seeded trusted examples | Let the agent retrieve past question and SQL pairs. Seeded here with the eval’s own questions, so it’s a ceiling, not a result. |
Column trimming | Prompt the agent to select only the fields the question needs. |
Disambiguation layer | Resolve a question’s hidden assumptions once, then reuse that resolution on later runs. |
Disambiguation + column trimming | Both of the above, together. |
Self-consistency voting (generate several candidate queries and take a majority vote) isn’t shown because it wasn’t re-run on the agent in this experiment.
Personal reflections
My summer wasn't just about this project. One of the best parts of working at a growing company is how much you get to learn from people across the business. I picked up a lot from coworkers in engineering, operations, finance, marketing, sales, and more, whether in meetings, over lunch, at our biweekly soccer games, or during our hackathon.
Since I hope to start my own company or work at a startup someday, getting to meet with each of our founders and hear their advice was invaluable. I'm grateful for the experience, and I'm heading back to Stanford having learned far more than I expected about what it's actually like to build an AI company at the frontier.




