The Most Expensive AI Errors Are Made of True Numbers
The Most Expensive AI Errors Are Made of True Numbers For three days in May, every pre-earnings brief our research agent produced on NVIDIA was anchored to one number: $68.1 billion, presented as the market’s consensus for the next quarter. The agent reasoned from it carefully. A beat is priced in, meaning the market already expects the target to be cleared. The bar sits here. Watch the guidance, then the reaction. The number was real. NVIDIA had reported it the previous quarter. It was the last quarter’s actual revenue, and the agent had dressed it as the next quarter’s forecast. NVIDIA’s own published outlook for the quarter the agent was forecasting was $78.0 billion, nearly $10 billion higher, so every brief built on that anchor was aimed at the wrong bar. 1 Three days of confident, internally consistent, well-written analysis, wrong at the root. And the root was not a lie; it was a real number in the wrong tense. Some failures in this record invented a part outright. But the errors that survived longest took a real number and attached the wrong period, scope, label or authority to it. I run an autonomous research agent inside a large personal knowledge system, and for two months I checked what it told me at every layer: individual claims against the strongest sources reachable, filings and earnings releases and the survey publishers’ own pages; whole documents against a gate that decides what enters the knowledge base; the agent’s own confidence labels against independent reviewers. Everything got logged. This piece is that record, and the field guide that fell out of it: the nine failure modes this audit exposed, many of which ordinary claim-level fact-checking misses, a real specimen of each from our logs, what each one looks like in an ordinary chat window, and the cheap check that catches it. At the end there is a method, and a tool, for building the same trust table for your own AI. The headline is the part I did not expect. Checked claim by claim, the agent was largely right. Checked at the level of what deserved to enter what I know, almost everything died. Those two results are about different things, and the gap between them is where the money goes. What we ran, and what checking means here The agent is the same one from How Reliable Is Your AI Agent?, the 91% piece: a Hermes-framework agent on a deliberately cheap model, running unattended on a rented server, researching markets and AI. It writes into a knowledge system of more than 16,000 notes. The standing rule: nothing it produces enters the canonical layer without surviving a check it cannot influence. Claims get sampled against the strongest sources reachable. Documents queue at a gate where a duplicate check and a human decide what gets through. Its confidence labels are read as text, never as evidence. If you work in a chat window rather than a pipeline, you own the same architecture without the vocabulary. The moment you copy an AI answer into your notes, your plan, or your codebase is your promotion gate. The only question is whether anything stands at it. Before the numbers, the scope. This is one agent, one model, one system’s gates, over one two-month window, 9 May to 13 July 2026. Every percentage below is local. Your model is probably better than ours. Your corpus and your review habits are different, so your rates will differ, and I will flag the direction where I can. What transfers are the failure classes and the checks. And one result transfers with special force: the accuracy numbers below came from a cheap model, and accuracy was not the main thing killing its output. Redundancy and expiry were, and those are properties of your corpus and your calendar. Upgrading the model alone does not fix the survival problem. The scoreboard: four gates, four denominators These are four separate measurements on four different populations, taken as the work happened. They do not chain into a single funnel, and multiplying across them produces nonsense. The claim audit In May we pulled 41 discrete factual claims from the agent’s research output, across two audits five days apart, and checked each against the strongest source reachable: SEC filings, earnings releases, official statistics, the trade press where nothing better existed. The checking ran on a separate model with web access, never the agent grading itself, and I adjudicated the verdicts. 32 held exactly. 6 more we graded approximate at the time, right in direction and wrong in precision. 3 were materially false. As first graded, that is 78% strict and 93% directional. Then the blind re-check described later in this piece re-examined 8 of the 41 and moved two of those approximates to wrong, which takes the directional rate to 88%. Only 8 were re-checked, and both moves went the same way, so treat 88% as the ceiling on what a fully blind pass would have returned rather than as a corrected figure. The strict rate does not move. 2 The anatomy of the misses matters more than the rate. No invented companies. No invented events. No inverted conclusions. Where invention appeared, it was one level down: sub-category splits and secondary ratios the source never published, riding on top-line stories that checked out. The failures were dates, staleness, scope, and structure: the classes in the field guide below. The door Over the same two months, the agent staged 113 documents for promotion into the knowledge base. 9 made it. 104 went to the archive. 45% of the queue, 51 items of the 113, duplicated something the system already held. The rest had expired before review or were too thin to keep. Three concessions before you quote that number. We capture aggressively by policy, so our net catches more junk than a stricter pipeline would. This was a backlog clearance, so it overstates any steady-state week. And a share of the expiry is on us, because the queue outlived our review loop. Read it as a cost figure, and it is a cost figure agent retrospectives rarely publish: autonomous capture into a mature corpus yielded single-digit percent durable knowledge, and the triage burden scaled with volume, no matter how accurate the individual sentences were. The high-value flags 18 items the agent had marked as significant knowledge sat in review for more than a week. 14 died of age before a human read them. That one is a finding about us, and probably about you: machine-flagged knowledge had a shelf life shorter than our review loop, which means “I’ll review it at the weekend” is a decision with a price. The other 4 were the queue’s most substantial research, each carrying its own verification claim, one of them a literal “Verified: 20/20 claims (100%)”. All four failed the promotion gate. Four of four. One had confabulated a specification date. One reported the star-counts on software projects, GitHub’s popularity number, wrong by 2 to 3 times under the label “100% verified”. One was titled “Verified Incidents” and sourced its numbered security vulnerabilities to personal blogs. One duplicated work the system already trusted. The checkers Three separate failures, three different mechanisms, and they deserve to be kept apart. Our internal review panels, scored against a six-reader focus group, on the same essay of ours, ran 0.9 points hot on a 10-point scale: one paired test, so treat it as an anecdote, not a benchmark. A batch of reviewer agents we ran over the queue waved off items that a direct check later kept. And the labels: in two months of logs, “100% verified” appears attached to work that failed verification. One stamped draft restated an analysis that had been updated 11 minutes earlier. Another asserted it had checked the vault for duplicates and found none, the same day its duplicate went live. That second agent had not run a check and reported a result. It had authored the sentence a check would produce. Accuracy is a property of sentences. Survival is a property of what you can safely build on. Our agent scored well on the first and brutally on the second, and almost none of the gap was made of lies. Top comments (0)
Comments
No comments yet. Start the discussion.