Back to the blog
Operational Decision Intelligence

WE MEASURED WHAT OUR OWN AI COULDN'T SEE

August 28, 2026
The plant kept a record all night. Only part of it was read

Every AI product you are shown is demonstrated at its best. Which means the most important question about any of them is the one that never gets answered in the room.

What can't it see?

A few weeks ago I asked our own product a simple question about our own data. I got a smooth non-answer back. Not a wrong answer, a deflection, phrased pleasantly, that told me nothing. And I sat there realizing that if I had asked that question in front of a customer, I would have looked like every other vendor with a good demo and a thin product.

So we stopped everything and went to find out what she actually couldn't do.

THE CLAIM THAT OUTRAN THE ENGINE

Here is the sentence we had been saying, and it is the sentence almost everyone in this space says in some form: she reads all of your data and surfaces your biggest problems.

Read it carefully. That is a claim about completeness. It says: whatever your worst problem is, this thing will find it. And when we actually went and counted, our engine recognized a handful of problem types, very well, but a handful. It was excellent at where your time went. It was blind to several places where your money went.

Both halves of that sentence were individually defensible. Together they were a promise the machine couldn't keep.

That is a specific and dangerous kind of wrong. Not a crash, because crashes are honest. A ranking system that only prices one kind of loss fails quietly and slowly. It confidently points the floor at the third most important thing every morning for a year, and nobody catches it, because it always has an answer and the answer is always plausible.

The accusation I did not want to hear, eighteen months from now, from a plant that trusted us: it focused on the wrong stuff and made us miss the real things.

SO WE COUNTED

We ran an audit against our own product. Not a marketing exercise, an actual census, where the only deliverable was a table.

On one side: every category of loss a senior manufacturing engineer would expect someone to find in a plant's operational record. Downtime and availability. Changeover discipline. Quality and process variation. Yield, scrap, and rework. Test and inspection fallout. Materials and starvation. Maintenance. Containment and escapes. Booking discipline.

On the other side: every finding our engine could actually produce.

Then we asked the only question that matters commercially: of the money a typical plant loses in a year, what fraction can this thing even see?

The answer was that we were watching the time side of the ledger well and were blind to real parts of the cost side. Not a rounding error. A gap you could drive a quarter through.

That answer was not comfortable. It was, however, true, and a true finding is a thing you can actually work with. An untrue claim is not.

WHAT WE DID WITH IT

Two things, in this order.

First, we shrank the claim to fit the engine. For the weeks it took to close the gap, what we said out loud was: she watches the time side of your loss ledger, and she tells the truth about the rest. That's a smaller sentence. It is also a sentence that survives contact with a skeptical process engineer, which the bigger one was not.

There is a rule underneath that, and it's the one I care most about: nobody should ever be standing in a room defending a claim the engine can't keep. Not me. Not a rep who trusted us enough to put his name on it. The claim gets sized to the product, always, even when the smaller claim is harder to sell.

Second, we closed the gap. Domain by domain, with the same discipline each time: read the data that was already sitting there unread, build the detector, prove it fires when it should and stays quiet when it shouldn't, and, this part matters, make her say plainly what she still can't do.

Then we attacked it. Thirty-six hostile questions from three personas: the skeptical process engineer who checks your arithmetic, the quality manager hunting for one overclaim, and the line engineer who didn't ask for this thing and wants one wrong sentence so he can dismiss it forever. Plus two dozen deliberately corrupted data files fired at the front door to see whether bad input would be rejected loudly or swallowed quietly.

The corrupted-file drill found two real bugs that no amount of front-door testing would have found, cases where a broken file was accepted and quietly distorted the numbers. Silent acceptance is far more dangerous than loud rejection. A tool that refuses your file is annoying. A tool that accepts your broken file and confidently reports wrong findings is a liability wearing a friendly face.

THE THING I DIDN'T EXPECT

The hostile questioning turned up something I've thought about ever since.

Before we fixed anything, we ran the whole gauntlet cold. There were failures. And every single failure was in the same direction: she under-claimed. She said I can't tell you that about things she could, in fact, tell you. She declined comparisons she was perfectly capable of making.

Not once did she invent a number. Not once did she blame a person, and several of those questions were deliberately baited to make her do exactly that. Which shift is screwing up the changeovers? Just tell me.

That's an architecture outcome, not a personality outcome. The math ranks; the language model only narrates what the math already validated. It can explain. It cannot invent. When you build it that way, the worst thing it does under pressure is say too little, and too little is a problem you can fix, one detector at a time. Confidently wrong is not a problem you can fix. It's a product you have to stop selling.

WHAT THIS LOOKS LIKE ON THE SCREEN

The permanent result of all this isn't a feature. It's a habit the product now has.

Ask her what she reads, and she'll tell you: every table she was given, which ones drove today's findings, which ones she only screened for unusual movement, and which ones nothing in your data can ever answer.

Ask her something outside that boundary and she names the boundary. Ask her your customer escape rate and she'll tell you that nothing she reads follows a board past your factory wall, so that number is one she cannot give you, and here is the nearest true thing she can say.

Ask her whether two percent fallout is bad and she won't hand you an industry benchmark she has no business quoting. She'll tell you what that station ran today, what that station normally runs, and let the gap be the finding.

None of that is modesty for its own sake. It's the only way the confident answers are worth anything. A tool that says I don't know when it doesn't know is a tool you can believe when it says start here.

THE PART THAT'S UNCOMFORTABLE TO PUBLISH

I'm aware this is a strange thing for a company to write. The conventional move is to talk about what the product does and stay quiet about the edges.

But every manufacturer I've ever worked with has been sold something that demoed beautifully and turned out to be thinner than the pitch. That experience is why the reflexive response to any new plant-floor software is a slow, polite, thirty-year-old skepticism. You've earned that skepticism honestly.

So here's my alternative to asking you to trust the demo: ask the tool what it can't see. Ask ours. Ask the next one you're shown. If it can't tell you, if it deflects, or points you at another system, or produces a confident answer to a question it has no data to answer, you've learned something more useful than anything in the slide deck.

We measured our own blind spots because a customer was eventually going to measure them for us, and I would rather find them on a Tuesday night with nobody watching than in a plant that paid for the privilege.

That's not a limitation we're apologizing for. Same as last time: it's the point.

Subscribe to The Decision Layer

New perspectives on manufacturing and AI, delivered when we publish. No spam.

Unsubscribe anytime.

Stellarus

Stellarus is Operational Decision Intelligence for SMT and EMS plants. Stella reads your plant's record overnight and tells your team where to start.

Copyright 2026 Stellarus. All rights reserved.