← Back to blog
September 18, 2026·4 min read·AI Agents

Your AI Agent Isn't Wrong. It Just Can't Touch a Real Tool.

An engineer asked an AI assistant whether a steel beam would hold. It confidently said yes. The real factor of safety was 0.73. The problem wasn't the model. It was that the model was never allowed to run a real calculation. Four independent signals this week point to the same missing layer.

Serge Kadjo
Serge Kadjo
AI agentsMCPengineering verificationhuman approvalhardware
Your AI Agent Isn't Wrong. It Just Can't Touch a Real Tool.

Your AI Agent Isn't Wrong. It Just Can't Touch a Real Tool.

An engineer asked an AI assistant a structural question: will this beam hold? The model answered with a reassuring yes. The real factor of safety was 0.73. The beam would not hold.

Here is the part that matters. The model was not broken. As Shahrad Zomorrodi, the engineer who lived this, put it: "It wasn't malfunctioning. It had no way to run the calculation, so it did what these models do without tools: it pattern-matched what a reassuring answer sounds like."

The agent did not compute. It performed confidence.

The hand-built answer

Zomorrodi's response is the interesting part. He did not write a better prompt. He built an MCP server: real mechanical engineering tools (beam stress and deflection, factor of safety, material properties, unit conversion, curve fitting) behind a tool interface, with tests and CI. The model calls the tool, the tool runs the real math, and the result comes back with the formula attached so a human can check the work.

Note what he reinvented without knowing it: receipts. The tool does not just return a number. It returns the number plus the formula that produced it, so the reasoning is checkable. That is the entire verification philosophy in one design decision.

He is not alone. Across our research this month we keep finding the same pattern: engineers hand-building MCP servers, eval scaffolding, and BOM sync scripts because agents answer questions but cannot act through real engineering tools. Nine documented cases and counting. Builders keep doing this work for free because no product gives them tool-grounded agents with verification built in.

The 80% ceiling

Even where AI tooling exists in hardware, it stalls at the same place. PCB practitioners report that AI autorouters complete 70 to 85 percent of routing, and then the board still needs hours of manual refinement before it is production-ready. As one designer wrote: "There's also a real gap between 'AI completed 80% of the routing' and 'this board is ready to fab.'" RF layouts, dense BGA breakouts, highly constrained designs: engineer judgment, every time.

Generation works. Completion is where manufacturability lives, and no agent closes that gap today.

What the chip industry settled on

The most conservative engineering domain on earth has already converged on the answer. At DAC this summer, verification leaders from Micron and Siemens EDA landed on the same working model: "generated by machine, validated by a person." The specification is the golden reference of design intent, and today only humans can validate a specification. The industry's nightmare is not that AI generates a bad design. It is that someone generates a design, sends it to manufacturing, and nobody can say who checked it or how.

And the evaluation numbers keep getting worse the harder you look. IBM Research's new Pass^k metric (the fraction of tasks where an agent succeeds on all k runs) shows a ReAct agent with 77.4 percent mean accuracy succeeding on all five runs only 53.0 percent of the time. Nearly a quarter of tasks are "sometimes solvable" with nothing changing between runs. Berkeley's new benchmark puts it bluntly: "the age of useful agents is here. The age of truly job-ready agents is not."

The missing layer

Four independent signals, one conclusion. The gap is not model capability. It is the absence of a layer where agents act through real tools, every action is verified against ground truth, and a human validates against the spec before anything ships.

That layer has a shape. Tools, not vibes. Receipts, not reassurance. Human approval at the boundary where a design becomes a product. The engineer who built his own MCP server after the 0.73 incident arrived at the same architecture from first principles: give the agent the real tool, and make the result checkable.

That is what we are building at ProductFlo. The agent proposes. The tools compute. The human approves. And every step leaves a receipt.

Stop bad hardware changes before they ship.

30-day pilot on one active product. Fixed $8k. Live on your own BOM in week one.