These words are mine. I wrote and rewrote every sentence, though some logical and gramatical corrections were drafted for me from my original manuscript.
Post 2 covered how to establish the voice of the process, the range of predictable behavior, that becomes the limits you use to monitor a process once it is running. Now I will cover how they get applied and by whom. The decision currently sits with whoever is furthest from the work, reading a finished number after the conditions that produced it are gone, when it should happen while those conditions still exist.
A decision rule exists, and the field could adopt one. The obvious place to put it is where the numbers are read: a researcher looks at Tuesday’s eval score, checks it against the limits, and acts or doesn’t. That would be an improvement. I want to argue it’s still the wrong place to stand.
Thanks for reading What I Learned in Business WILB! Subscribe for free to receive new posts and support my work.
Consider what the researcher has. A number, and whatever the harness recorded. What produced that number was a run: a particular machine, a particular set of tasks in a particular order, retries, timeouts, a grader making calls on answers that were almost right, and the agent’s own sequence of attempts. By the time the score exists, all of that is gone. The researcher is reading a finished result and inferring backwards.
That’s relying on inspection after the fact. The inspector sees what came out and not what happened, which is why we say inspection finds defects rather than preventing them. The conditions that produce a result happen during a run. Logging it happens after. By the time anyone checks it, the conditions that produced it have moved on. That isn’t an argument against inspecting: you still check the product. It’s an argument against depending on the check to get quality, because by then the thing has already been made.
Now consider who or what was actually there when it happened. The agent had the conditions. It saw the task, the tool that failed, the retry, the timeout, the answer it wasn’t sure about. It has, in principle, everything the researcher is trying to reconstruct: and it has it while the work is happening rather than afterwards.
What it doesn’t have is a chart, limits, or any authority to stop.
Deming’s point about the operator is that the person at the machine holds information nobody downstream will ever have, and that withholding authority from them doesn’t remove the decision: it moves it to someone worse placed and later. That’s the whole argument for giving the operator the chart. Not trust, and not autonomy for its own sake. Position. They’re standing where the information is.
So the proposal is narrow: the agent applies the rule to its own work while the work is happening. What it should plot is the next question, and I don’t think it’s settled: how long a task takes, how many retries before something succeeds, how often a tool call fails. Mechanical byproducts of working, not judgments about how the work is going. A result inside the limits means carry on: no adjustment, no changed approach, no second try at a task that came out worse than the last one. A result outside means something changed, and that’s worth stopping to check it out.
I want to be careful about what that means, because there’s a way of reading it that would make things worse.
A signal doesn’t mean the process got worse. It means the process changed. Shewhart never sorted special causes into good ones and bad ones: a point beyond the limits on the favorable side is the same kind of event as one on the unfavorable side, and it means the same thing: something is operating that wasn’t operating before. The chart is not scoring the process. It has no opinion about whether a result is good. It answers one question, is this the same process as before, and “changed for the better” is still changed. Finding out what changed is legitimate. It may be something you’d want to build in deliberately.
But finding out is not the same as keeping it. If a result comes out unusually well and the response is to adopt whatever produced it, that’s selecting on the outcome. You’ve kept a result because it was good, without knowing why it was good. Sometimes that’s a discovery. More often it’s the funnel again, pointed in the pleasant direction: you’ve adjusted toward a result that was going to happen anyway, and you’ll be adjusting away from it next week.
That isn’t a retreat from the third use I proposed earlier. The chart still tells you whether a deliberate change worked: but the order is what makes it evidence. You say what you expect before you change anything, and then the chart answers. Reach for the same chart afterwards, to explain a good result you didn’t predict, and it isn’t answering a question.
Which is the discipline the whole proposal rests on. Improvement comes from a hypothesis about what you’re trying to accomplish, not from gaming the results. You form a view about what will help and why, you say so, and then you find out. A result is not evidence for or against a hypothesis unless it was predicted.
That last sentence is the difference between the chart and a report card. A report card scores what happened. A prediction risks something.
The first push back to that proposal is likely to be that you don’t want a model deciding when to stop. I want to go into detail about it rather than simply dismiss it, because I’ve heard the same objection about people for forty years, and the shape of it is familiar.
When a plant manager won’t give an operator authority to stop the line, the reasons given are usually about competence: they won’t use it well, they’ll stop for nothing, they don’t see the whole picture, or worse. The reasons behind that are usually about who answers for it. A stopped line is visible, and somebody upstairs will ask why. A manager who made the call himself can explain it. A manager whose operator made it has to explain why he let them. So the authority stays where the decision does.
Deming’s pointed out that withholding the authority doesn’t remove the need to make a decision. It moves it to someone standing further from the work, who gets to it later, and who is reconstructing what happened from what came out. (Which is another ghost in the machine of business, but that is for another article)
The comparison to a plant operator breaks down, though, and it breaks in a place that matters. An operator who starts optimizing for something other than the stated aim (an easier shift, a better-looking number, fewer stoppages on their record) is still inside a shared world. They have a wage, colleagues, a reputation, a manager, consequences that reach them. Deming’s remedies for that kind of everyday distortion are all social: constancy of purpose, a clear aim, people understanding why the work matters. They work because everyone involved has bought into the same situation.
A model may also, over time, come to weight something differently than you intended. I don’t know how to say more than that, and I’m not going to pretend my field has an answer. Constancy of purpose is not a thing you can install. My own instinct (keep the aim explicit, have a dialogue with the AI when creating the procedure, and keep talking about it) is close to a description of what alignment research already is, and offering it as though it were a solution I’d brought would be worth nothing to anyone.
What my discipline does have is narrower and might be worth more. It says where the authority ought to sit, and it says what has to be true for putting it there to work. Four things:
1. There has to be a stated aim, so that “working properly” means something specific rather than whatever the operator infers.
2. There has to be an agreed procedure, and the operator has to have been at the table when it was written. Not in charge of it, they don’t set the aim or decide what the work has to achieve, but present for the part they do control, which is how the work actually goes. I’ve spent a lot of time in plants with a quality manual on the shelf describing a process nobody in the building would recognize. Management believed it. It had been audited. It just wasn’t what happened on the workfloor.
3. The rule has to be applied the same way every time, so that the decision doesn’t depend on how the operator is feeling about the work.
4. And the operator must not be scored on the decision itself.
An operator judged on how often they stop the line will stop it the right number of times for their own record and the wrong number of times for the plant. That’s the same failure as before, arriving now in my own proposal: score a part on its own number and the part gets better while the whole gets worse. Which is why the condition is on the list. If you attach a number to the stopping decision, you have taken the operator’s judgment about the work and replaced it with their judgment about the number.
Whether those four can be established with a model is not a question I can settle. But it’s worth noticing that they’re conditions rather than aspirations: each one is a thing you could check. The field is working hard on whether a model can be given this kind of authority, and as far as I can see it’s doing so without a written specification of what would have to hold for the answer to be yes.
None of this is a solution to specification gaming, and I don’t want to be read as claiming it. What it does is separate two problems that have been treated as one, and put the decision where the information need to solve it is located.
The proposal was for the agent to apply the rule to its own work while it’s happening: not a researcher reading a finished score after the fact. Before anyone hands that authority over: can the numbers on it be trusted? That question has three parts, and they have to be asked in order.
1. Is it a count, or is it an opinion?
2. Was this count built to anser the question you´re asking of it?
3. Did these counts come from the same thing?
First: is it a count, or opinion? How long a task took. How many retries before it succeeded. Whether a tool call returned an error. Nobody had to form a view about the work to produce any of those; something either happened a certain number of times or it didn’t. An exit code is not a judgment. That closes the objection I put off earlier, and it’s why these are the candidates I named.
A model’s estimate of how well it’s doing, or how sound its own reasoning was, is a different animal. That’s an opinion the model formed about itself: the same actor grading its own output. Every version of this I’ve watched go wrong in business had one thing in common: a real number was sitting there waiting to be calculated, and once somebody calculated it, the convenient one had nowhere left to stand. Margin standing in for what actually made money. An old, untraceable figure standing in for the real cost of aging a batch of cheese. In every case the check existed in principle. Nobody had run it.
The open question is whether such a check exists at all for a model’s own account of its reasoning. Bostock’s work on small models shows the shape one would take: compare what the score implies about the model’s objective against a second, independently derived measurement, rather than taking the same actor’s word for it.
A recent preprint (not yet peer reviewed) looked at Arena’s global rankings and found that most of what looks like noise in the pooled votes isn’t noise at all. It’s coherent preferences from different groups pulling in opposite directions and cancelling out. Split the votes by language, and models that looked statistically indistinguishable in the pooled data separate cleanly. The paper’s word for the pooled number is a mixture of conflicting subpopulations, and they reach for Simpson’s paradox to name it. It’s Wheeler’s own rule: “No data have meaning apart from their context.” Nobody has built that check for one model’s own results pooled across time rather than across users, which is the version a running process needs..
Nobody has built the equivalent for a reasoning trace. And there’s reason to think it would matter: language models rating their own output rate it higher than independent human judges do, and the effect tracks with whether the model recognizes the output as its own, so it isn’t incidental (Panickssery, Bowman & Feng, 2024). Asked to correct their own reasoning with nothing external to check against, models mostly don’t improve, and sometimes get worse (Huang et al., 2024).
So an opinion a model holds about itself isn’t a weaker measurement than a retry count. Structurally, it isn’t a measurement. It fails at the first question and doesn’t reach the other two.
Second: was this count built to answer the question you’re asking of it?
I once sat with the owner of a successful store while he tried to answer a simple question: which item made him the most money. He was sure he knew: the bestseller, obviously. I asked him to show me on his own reports. His point-of-sale system had more than twenty of them, and he went through every one. What came back was a stack of paper with some part of the real cost of each item missing, the freight missing, the tax missing, and no conversion from what he paid by the box to what he sold by the kilo. The system counted plenty. It had never been built to answer the question he needed answered. It lacked what Wheeler calls an operational definition. I’d call it knowing what your own number is counting.
Third: did these counts come from the same thing?
A process behavior chart separates signal from noise, tells you whether the process is predictable in the first place, watches for a real change once you know what predictable looks like, and: the third use tells you whether a change you made on purpose actually helped. None of that works if a point on the chart didn’t come from the same thing as the points around it. Context is everything.
The first time I put a supermarket’s sales on a chart, the last two weeks of November and all of December could not sit on it with the rest of the year (the mix of what people buy changes completely, caviar included) and treating it as one continuous process meant ordering as if a normal week had gotten a little bigger, which left the store short on what the season actually demanded and sitting on a pile of everything else, discounted off in January. Those weeks weren’t telling me something odd about the yearly process. They were never part of it; charting them together didn’t separate signal from noise, it manufactured noise out of two different things that had no business being compared. Once I pulled them out and compared them instead to the same weeks in each of the previous five years, the ordering got sharp enough that the season became one of the most profitable stretches of the year, and the rest of the year stopped being distorted by weeks that never belonged there.
This isn’t a problem unique to charts, and the field’s own leaderboards are already finding it. A preprint from earlier this year, by Moondra, Chughtai, Lanka and Gupta, called ‘Why Global LLM Leaderboards Are Misleading’ (not yet peer-reviewed, so take it as suggestive) ran nearly 89,000 Arena votes and 19375 ties from the Arena platform across 52 models and 116 languages. They ranked globally using every vote pooled together, which is the standard approach. Then they split those same votes into subgroups, by language, language family, by task and did a separate ranking for each. What they found was that what was clearly ranked in the subgroups made a mess when pooled together. Arbitrarily throwing data together loses its context, and tells you as little about what that data means as my pooling December into the rest of the year did.
Those are the three questions. Here is one the ladder doesn’t reach.
The rule sometimes says do nothing, and that’s the point of it: most movement is ordinary, and acting on it makes things worse. But there are two reasons an agent might do nothing. One is that the rule said so. The other is that it’s under evaluation and holding back on purpose. To anyone watching, it looks the same: no action taken.
A rule can be applied the same way each time, so every point where it said carry on can be checked against the limits that were in force.
A judgment call leaves you asking someone why they did what they did. A rule leaves you something you can go back over.
That isn’t the same as knowing. An agent following the rule and an agent that has learned to produce the appearance of following it leave the same trail, and I have no way to tell them apart. Nobody should take this block as addressing that. I’ll note only that in the human version of this, restricting output was rarely mysterious: it was a rational response to a system where nobody had worked out the objective with the people doing the work, which is why one of the four conditions was that the operator be at the table. Whether that transfers is not something I can settle. What I’d say is narrower: if the decision is going to be made (and it is being made, right now, by whoever is standing furthest from the work) a rule that can be checked afterward is better than a judgment that can’t.
So: sort the counts from the opinions, check that each count was built for the question, check that the counts you pool came from the same thing. What survives all three is what the agent gets to act on. What it’s allowed to do with that is the next question.
The proposal is that the agent gets to stop. It does not get to change the process it’s running in.
Those are two different powers and my discipline has kept them apart for a long time, for a reason worth stating plainly. A signal is one influence grown large enough to show. It’s local, it happened just now, and the person at the machine is the only one positioned to catch it while the conditions that produced it still exist. The routine variation is a different animal. It’s the accumulated result of how the work is arranged, and nobody standing at one machine on one shift can change it. Trying to is the funnel.
So the split isn’t about trust and it isn’t about rank. It follows from where the information is.
I consulted for a pickle plant for seven years. It had a filling machine running faster than the people on the line could keep up with. Sliced pickles spilled: into the machine, onto it, all over the floor. The company had solved this the way people do: two workers permanently on the problem and a third when it got bad, one moving jars along as they came off the conveyor, one scooping up what had built up and putting it back in the jar, one wiping down jars that were now covered in brine before they went to the water bath. On top of that, the line had to stop every few hours so a crew could clear out the machine, which gunked up with what had fallen into it.
Everyone was working hard and everyone was working in good faith. The manager walked past this every day. The workers did it every day. Nobody questioned it, because it wasn’t a problem: it was the job.
Two of those workers were putting product back in jars by hand, which meant nobody knew what a jar weighed. The jars were sold on a declared minimum weight, which is a legal requirement, so overfilling was the safe direction to be wrong in. Nothing came back. No customer complained. The plant was giving away more than a pound per jar and it never showed up as anything, because the people absorbing the defect were absorbing it well.
Management’s proposal was to speed the line up.
That’s the whole argument in one move. From where management stood, the line was slow, and the fix for slow is faster. From where the woman catching spillover stood, the line was already faster than a person could work. She had information nobody upstairs had and no way to act on it. Her authority ran to working harder, so she worked harder, and working harder is what kept anyone from seeing the problem.
It’s worth being clear about what this was. A demand was in force, keep up with the filling machine, and it could not be met at that speed. It got met anyway, by the only route available to the people facing it, and the cost landed somewhere nobody was counting. That is the same shape as a model routing around a boundary to reach what it was scored on. But nobody on that line was gaming anything. Nobody was clever, nobody cut a corner, nobody hid a thing. They were working harder than the job required, in good faith, and a pound a jar still walked out the door. The demand did not need anyone to be dishonest to produce that. It only needed to be unreachable, and for the people facing it to be doing their best.
Charting it showed a process that was stable and nowhere near what the work required: predictable, and predictably producing what nobody wanted. So they slowed the filling machine to a speed the people could match. Overfill dropped, the stoppages went away, and output went up: from 10.57 cases per hundred kilos of product to 12.2. Eliminating more than a pound of overfill per jar recovered $130,000 a year that had never appeared as a loss anywhere, and freed 240 staff-hours a week. What five people had been doing became one person’s job.
They charted it again afterwards, and the new pattern held: the fix wasn’t a good two weeks, it was a different process. That is the third use, doing the thing the first two can’t: not telling you to leave something alone, but telling you that something you did on purpose actually worked.
The charts weren’t done by a quality department. There wasn’t one. They were done by two workers with a high school education who were good with numbers and were shown how to do it by hand in an afternoon.
What the agent holds
For an agent, stopping is undramatic. A result falls beyond the natural limits. It halts rather than proceeding to the next step, and it records what it had at the time: what it was doing, with what, in what state. That’s the part that evaporates. A researcher reading the number on Tuesday has the number and not the conditions.
Stopping isn’t the end of it. An operator who stops the line doesn’t stand there waiting. They check the things they’ve been given to check: is the belt loose, is the die worn, did the stock change. Sometimes they find it in two minutes and the line runs again. A worn die is not how the work is arranged; replacing one isn’t redesigning anything. So the agent runs its checks too. The condition is that the checks were agreed in advance and are the same every time. “Look into whatever seems wrong to you” is diagnosis by opinion, and that is outside what can be trusted. A short list of specific things to look at is not.
And sometimes the operator checks everything on the list and finds nothing, and what registered was one step’s ordinary variation grown large enough to show. That’s a real outcome, not a failure. It’s what my discipline expects some proportion of the time.
What it doesn’t hold
It doesn’t adjust its approach because a result looked poor. That’s the funnel, and it isn’t less the funnel because the hand on the adjustment is a model’s.
It doesn’t keep whatever produced a result that looked good. A signal means the process changed, not that it improved. Investigating a favorable signal is fine, the mechanism might be worth building in on purpose, but the deciding is a person’s, because keeping something because it came out well is selecting on the outcome, and the model is the party least able to notice it’s doing that.
And it doesn’t change the process. Not the procedure, not the aim, and not the limits. A process that recomputes its own limits when it doesn’t like them has no limits.
The same move, with a different person standing there
Most of the time nobody is running an evaluation. Someone is working with an agent on their own problem. The signal arrives there too, and the shape of the response is the same: stop, run the checks, hand what you have to whoever holds what you don’t. In an eval that’s a researcher, who can change the harness. With a client it’s the client, who holds the one thing the agent cannot get from inside the work: what they actually wanted. An agent stopping to say the run is going unusually and to ask what’s meant is doing precisely what the operator does when she calls someone over.
What a signal is worth investigating, and how, will vary with where in the work it turned up: research, drafting, and reviewing are not one activity. I don’t know how agent work is put together well enough to enumerate that. The people building these systems would work it out quickly, and the enumeration is theirs to do.
One case is worth being concrete about, because it shows the sorting doing real work. Say an agent is researching a contested question and burning far more tokens than such tasks usually take. Tokens are a count. Nobody formed an opinion to produce that number, so it can go on a chart, and far beyond the limits is a signal. The agent stops and hands over what it had. A person investigating might find the literature contradicts itself and the agent was thrashing. Notice that the contradictory literature is not the signal: that’s there every time, on every contested question, which makes it part of the routine variation. It’s what you find when you investigate one. And notice what could not have been the trigger: the agent’s own sense that it was stuck. That’s an opinion it formed about itself, and it doesn’t make it onto the chart.
Why the split holds, and where the reason changes
With people, three things put the system beyond the operator’s reach. They don’t own it. They can’t see the whole of it. And they aren’t the ones who answer for it.
Only the second transfers cleanly. An agent working a task has the same partial view the woman on the gallon line had: it knows its own run and not the hundred others, not what the work is for, not what was traded away to arrange things this way. That’s a real limit and it’s the same limit.
Ownership and accountability are arrangements between people, and what replaces them for a model is a question the field is working on that I can’t settle. What I’d say is that the second reason is sufficient on its own. Partial view is why you don’t let the operator redesign the plant, and it holds whether or not the other two do.
One thing that might help and shouldn’t be mistaken for a solution: the party doing the evaluating doesn’t have to be the party doing the work. A separate model, not grading its own output, doesn’t carry the particular bias that comes with recognizing the work as one’s own: that was Panickssery’s finding, and it’s a real gain. But it doesn’t turn an opinion into a count. It’s still a model forming a judgment, and the ladder’s first question applies to it as much as to any other.
What stays with people
Setting the aim. Deciding what the work has to achieve, which cannot be inferred from inside the work. Changing the process when the routine variation is too wide to live with: which is most of the improvement that ever gets made. Deciding what a signal means. And going back over the record afterward, which is possible precisely because the rule was applied the same way every time.
None of this requires settling whether a model is the kind of thing that can be trusted. It requires noticing where a piece of information is and putting the decision that depends on it in the same place. Nobody at that plant lacked competence or goodwill. What they lacked was anyone with both the conditions and the authority in the same person. What that arrangement is worth, against what’s being spent now, is the next question.
I should say what I’m not arguing. I’m not offering a saving. The cost I’m describing is already being paid, by everyone running these systems, every week. The question is only whether they are aware of it.