Why do AI agents give confident wrong answers?
AI agents give confident wrong answers because they get your tools but not the map of your business. Why prompts and more tools fall short, and the fix.
AI agents give confident wrong answers because they get the tools your employees use, but not the map of the business those employees carry in their heads. On every run the agent rebuilds part of that map from the tools, fills the gaps with whatever is typical and answers with the same confidence either way. Better prompts and more tools don’t fix this. Giving agents the map does.
This note covers why it happens, why the usual fixes fall short, what the map looks like in practice and a quick way to test for the problem in your own company. For what the map itself is, see what is a world model for knowledge work?.
The two symptoms: rising costs and answers nobody trusts
The companies furthest along with AI tend to hit this first. We recently spoke with a tech company that had been deploying AI for months: Claude rolled out across the enterprise, people building their own agents and more than 40 MCP servers available to staff. They were pushing AI into the business harder than most, which is why they ran into the problem early. In their words:
“We were very quick in adopting those tools and putting all the connectors available, and now of course our costs are exploding. And the teams are still struggling to understand when they get an answer from these tools, is it correct, is it not?”
They aren’t an outlier. In a Dataiku and Harris Poll survey of 685 CIOs published in September 2026, 72% said they can’t consistently confirm whether their agents deliver the business outcomes they were built for.
Agents get your people’s tools but not their map
Think about how someone in finance works. In their head they carry a map of their world: customers, contracts, subscriptions, invoices, payments and how they connect. A customer signs a contract, the contract starts a subscription, the subscription produces invoices and the invoices get paid or don’t. The map also reaches outside the company, to a customer’s payment record, VAT rules per country and news that a customer is being acquired or cutting costs.
No single tool holds all of that:
| Part of the map | Where it lives |
|---|---|
| Invoices | ERP |
| Subscriptions | Billing system |
| Contracts | CRM |
| Payments | Bank |
| Payment records, VAT rules, news about customers | Outside sources, in no internal tool |
| How it all connects | In the person’s head. An agent rebuilds it on every run and loses it when the run ends. |
The person uses the tools to look up details and fits what they find onto the map they already have. That is what business tools were built for: people who already hold the map. Each tool gives them one window onto the business. Every company has this map, but it is rarely written down. People learn it on the job, and it’s a big part of what makes someone good at that job.
An agent gets the same windows without the map. So on every run it rebuilds the part it thinks it needs: it queries a few tools, stitches the pieces together, answers and forgets it all when the run ends. Ask it which customers might not pay next quarter and it has to work out on the spot how invoices, subscriptions, payments and the outside world connect. It only finds the connections it thought to look for, and it can’t see what it didn’t fetch. We wrote about the same gap in access is not a map.

The tech company from the call runs into exactly this. A person there knows which Confluence page is the official one. The agent sees a wiki that “has a mix of official ready to be consumed pages and pages that are just work in progress, pages that are archived,” and can’t tell them apart. Their product analytics tool knows users, but customers live in the data warehouse, and as one of them put it, “the user is not equal to a customer.” Ask the analytics tool’s MCP server a question today and “it’s going to answer you on a best effort.”
Why doesn’t the agent say it’s missing something?
A language model works by predicting the next piece of text from the text in front of it. Where that text has gaps, it fills them with what is typical in its training data, and your company isn’t in its training data. If nothing in the context says what “active customer” means at your company, you get something close to the average company’s definition. Andrej Karpathy put it this way: “hallucination is all LLMs do. They are dream machines.” It looks like a bug, he wrote, “but it’s just the LLM doing what it always does.”
Training also rewards guessing. OpenAI researchers argued in a September 2025 paper that models are “optimized to be good test-takers, and guessing when uncertain improves test performance.” Most benchmarks give no points for “I don’t know”, so a guess scores better on average.
The costs come from the same place. Each run spends tokens rebuilding the map, and the next run spends them again on the same thing. Exploding costs and answers nobody can check are two symptoms of one problem.
Why don’t better prompts and more tools fix it?
| Usual fix | What it does | Why it falls short |
|---|---|---|
| Better prompts | Tells the agent which system to trust or what a term means | Works for that one question from that one person. The next person who asks without the instruction gets the old answer. |
| More tools | Gives the agent another source to read | Adds one more window to stitch in on every run, and every tool description is sent to the model on every turn. |
On prompts: we think a rough prompt should still get a good answer. If your team has to learn prompting to get value, the system isn’t finished.
On tools: when the agent doesn’t know something, connecting another tool seems like the fix. Microsoft’s team made a similar point in a September post on context engineering: “More context does not guarantee better answers: a relevant fact buried in 40 pages is harder to use, and a long tool list makes the wrong choice more likely.” In their internal benchmark, letting the agent search for the tool it needs instead of loading the full list cut average input tokens by around 97% for large tool libraries. We go deeper on this in more context makes agents worse.
So before connecting MCP server number 41, ask whether the agent should read every tool directly at all, or read from a map that already holds what those tools know.
The fix: give agents the map of your business
Take the map out of people’s heads and out of the tools, and write it down where agents can read it. In practice that is a graph: the kinds of things in your business and how they relate (your ontology), filled in with your actual customers, contracts and invoices. Your tools keep it filled in and up to date. It should hold the outside world too, because your employees search the web and check outside sources while they work.
For the finance example, part of that map looks like this:

news item → mentions → customer → signed → contract → starts → subscription → bills → invoice ← settles ← payment
“Which customers might not pay next quarter?” is now one query across the map instead of a tour of four tools and a web search, and every run reads the same map. A late payment, a cost-cutting headline and a renewal date sit on the same customer, so the agent doesn’t have to think of connecting them.
Two conditions for the map to work
- Agents read their facts from the map, not from the tools. If the map is just a 41st tool next to the other 40, agents will still go tool by tool some of the time, and you’re back to guessing. The tools feed the map, and agents still use the tools to act (send the payment reminder, update the deal), but they get their facts from the map. This is the idea behind map first, tools second.
- People decide what the map says. Someone has to decide which system wins when two disagree, and which Confluence pages are official. The person who owns the process decides once, and it’s written down where agents read it. AI can draft it, and a human maintains and reviews it. Which system wins and who owns the process are questions we answer early in every deployment, and they are part of governing AI agents.
Isn’t this just a data warehouse?
It’s the obvious objection: pull everything in, join the tables, done. A join puts records from different systems next to each other without choosing between them. If the CRM says one thing and the finance system says another, the join shows you both, and the agent still has to work out which one counts, on every run. Without the sign-off from the people who own the process, it is still just a joined database.
What changes when agents work from a map
When both conditions hold, the agent stops rebuilding the map on every run. Ask the same question three times and you get the same answer. Rough prompts work, because the agent already has the context. And you stop paying for every run to work out the same things about your business again.
The market is moving this way. On 6 October 2026 SAP announced it is acquiring TechWolf, from Ghent, for what TechWolf calls a “context graph for work”. SAP’s product chief Manoj Swaminathan described it as “an excellent grounding layer for agent queries” that “makes token usage more efficient”. We explain the term in what is a context graph?
A quick test for your company
Have three people ask your AI agent the same business question. If you get three different answers, your agent doesn’t have the map. It’s rebuilding it on every run and guessing differently each time.
How Ortelian gives agents the map
Ortelian is a platform plus a forward deployed team. The platform holds a world model of your company and the world around it, hosted in the EU, and agents read their facts from it through scoped tools. Our engineers work on site with the people who own each process to map the work and agree which system wins, as described in how we work. If you want to know what that role involves, read what is forward deployed engineering?
Questions people ask
Why do AI agents hallucinate on business questions?
A language model fills gaps in its context with what is typical in its training data, and your company’s definitions, systems and exceptions aren’t in that data. Training also rewards guessing over saying “I don’t know”, so the agent answers with confidence whether it found the facts or not.
Will a better model stop confident wrong answers?
A better model reads its context more carefully, but it can only use what is there. If nothing in the context says which system wins or what “active customer” means at your company, a better model still guesses.
Do more MCP servers make AI agents more accurate?
Usually not. Each server is one more window the agent has to stitch in on every run, and every tool description is sent to the model on every turn. That raises token costs and makes a wrong tool choice more likely. Agents do better reading facts from one map that the tools feed.
How is this map different from a knowledge graph?
A knowledge graph is roughly the middle layer: the things in your business and how they connect. A world model for knowledge work adds the data that keeps it current, a source on every fact and the knowledge tied to each thing. See world model vs knowledge graph.
How can I tell if my AI agent is guessing?
Ask the same business question through three different people. Different answers mean the agent is rebuilding its picture of the company on each run instead of reading it from one map.
This note first appeared in our newsletter as Why your AI agents give confident wrong answers.
