The Half-Life of Judgment

Photo by Nicolas Lafargue on Unsplash

The morning the first iPhone 3G launched, the queue filled with a problem the queue had never seen.

I was working in contact centers then, on the floor, hearing the beep of a new call arriving as soon as the previous one ended, watching the flood roll over us in real time.

The calls coming in weren’t the type of calls the floor was built for. People had a device nobody on the team had ever held, asking questions that had no FAQ article, no script, and no precedent. The honest length of those calls, the time it actually took to help someone through something genuinely new using only the documentation, was four times what the floor was tuned for.

And on the wall, the number did not move.

Average Handle Time. The target that day was the target from the week before, which was the target from the quarter before that. Agents who slowed down to actually solve the new problem watched their numbers go red. The floor’s implicit instruction, the one no supervisor had to say aloud, was to keep the calls short on the one day of the year when short calls meant worse service.

The metric didn’t ask what kind of day it was. It couldn’t; it had never been built to.

The same thing happened, in a quieter register, every time the network went down in a region. Volume spiked, the calls got harder and longer and angrier, and the target held perfectly still, as if nothing in the world had changed.

It’s worth being precise about why, because the reason is not stupidity, and the people enforcing that number were not fools. AHT was load-bearing on both sides of a BPO (business process outsourcing) contract, and it was wired into the money in more than one direction.

The client forecasted call volume against a mostly static AHT assumption; that was how they sized staffing and balanced agent capacity, and if the assumption moved, the whole capacity plan was suddenly wrong. The outsourcer billed per call-minute rather than per call, which made every minute of handle time a unit of revenue as much as a unit of service. And underneath both sat the service-level agreement: a ceiling on how long a caller could wait before an agent picked up, carrying financial penalties, often steep ones, the moment it was breached.

Those forces are not aligned the way you would guess. Per-minute billing, left alone, would happily let handle time drift upward, since longer calls are simply more billable minutes. What actually holds AHT down is the wait-time SLA, because handle time is a direct input to how many calls a given headcount can answer. Longer calls mean fewer answered, which means more callers waiting, which means longer wait times, which means the penalty clause trips.

The ceiling on AHT was never really about the length of any single call. It was the throttle that kept the queue moving fast enough to stay out of penalty.

And here is the part that matters for everything that follows: on a normal day, all of this is perfectly consistent. Short handle times serve each caller efficiently and keep the line flowing, so the AHT ceiling and the wait-time SLA want precisely the same thing, and the machine hums. That consistency is real. It was also only ever true under a single condition, the one nobody wrote down because nobody had to: a normal call mix.

The iPhone launch broke that condition, and the instant it broke, the two constraints that had always agreed came apart. When every call honestly needs four times the minutes and volume is up, you cannot both keep handle time short and answer everyone inside the SLA; those goals were never independent, only reconciled by the assumption of an ordinary day. Forced to choose, the apparatus protected the throughput number and the penalty exposure, and in doing so sacrificed the one thing that actually mattered that morning: the quality of help on a genuinely new problem.

Two organizations, three interlocking commercial forces, every one of them silently resting on the same unstated premise. Nobody in that building decided to lower the standard of service on launch day. The structure decided it for them, because protecting the queue was the thing the structure had been built to do.

So on the day the number most needed to flex, it was the one thing in the building that couldn’t.

That is not a story about a bad metric. AHT is a perfectly good metric for average days and call types. It is, however, a story about what happens to a good decision when it becomes durable enough to keep enforcing itself after the conditions that justified it have walked out of the room.

The Radius Is Also a Blast Radius

Last week, in Path B, I described the senior individual contributor’s real form of leverage as a radius of judgment: how many decisions, made by how many people, across how much of an enterprise, for how long, depend on thinking you did once and did well. I meant it as a description of power, and it is one. A junior contributor solves a problem. The most senior ones change the environment so that whole classes of problems get easier, or stop occurring. The measure of that work is reach; across teams you’ll never meet, across time you won’t be present for.

This piece is the invoice for that power.

Because for how long does not point in only one direction. The same mechanism that lets one good decision hold across teams and years is the identical mechanism by which one wrong or outdated decision does exactly the same thing. The same machine, but with the opposite sign. The same radius that measures how far your best judgment travels also measures how far your worst judgment travels, often long after you have stopped thinking it.

The radius of judgment is also a blast radius.

Someone, once, decided that AHT was the right way to hold a contact center accountable. Under normal conditions, they were right. Then that decision got embedded into dashboards, scorecards, staffing forecasts, and contractual SLAs. It was written into the muscle memory of a whole industry, and it acquired the one property good architecture is supposed to have. It became durable.

It held across companies and decades and thousands of people who never once re-derived it. It kept holding on the morning it was actively making service worse.

In Path B I wrote that if my judgment stops being good, nothing holds anyone to it but momentum, and that momentum is a short-term loan. I need to correct myself.

In a well-built system, momentum is not a short-term loan. It is a thirty-year mortgage the next team never signed. A standard does not lose its grip when the conditions that made it right begin to change. It keeps enforcing because someone made it durable. The applicability of the judgment and the persistence of the rule come apart, and the whole subject of this essay lives in that gap.

Why the Condition Disappears First

There is a reflex, when you describe the AHT situation, to reach for Goodhart’s Law: when a measure becomes a target, it stops being a good measure. I reached for it myself in The Meter is a Map, and it earned its place there.

It is not quite the mechanism at work here, and the difference matters.

Goodhart describes a proxy degraded by optimization. That was not quite what happened here. AHT continued measuring exactly what it had always measured. What failed was the assumption that the same target represented good performance under a radically different call mix.

The measure was not corrupted. Its interpretation was held still. And that inertia has a mechanism, one built into how we make decisions durable in the first place.

When we make a decision we intend to reuse, we encode the decision everywhere it needs to live. AHT went into the dashboard, the scorecard, the forecast model, the billing schedule, the QA rubric, and every coaching script. We are extraordinarily good at propagating the conclusion. What we almost never encode, anywhere durable, is the condition: the sentence that reads this assumes a normal call mix, and a demand shock or an outage voids it.

That sentence, if it was ever spoken at all, only lived on in someone’s head, or in a founding assumption so obvious at the time that writing it down felt unnecessary. It was not built to persist, so it didn’t.

The AHT case shows why that particular sentence was even less likely than most to survive. The normal-mix condition was not a footnote to a single rule; it was the hidden premise that made three separate constraints – the handle-time ceiling, the wait-time SLA, and the billing model – agree with one another. A premise doing that much quiet reconciling almost never gets stated, precisely because as long as it holds, nothing forces anyone to notice it is doing any work at all. The assumption that makes everything consistent is the last one anyone thinks to record, and the first one to matter when it fails.

This is why the condition disappears first. It is not merely an accident of neglect; it is an asymmetry we manufacture. We deliberately build the enforcement to outlast us while leaving the assumptions that bound it to a particular reality undocumented.

The rule’s fitness for launch morning isn’t the part that decays with time; it’s a cliff. It was correct on the last normal day and entirely insufficient by the first activation call; a threshold the demand shock crossed in hours. What decays is the institution’s retained awareness that the cliff is there.

The caveat “this holds only under a normal mix” is rarely encoded. To the extent that it survives at all, it survives in people, and people rotate.

Each new hire learns the target without the boundary that used to travel with it. Each rewrite of the QA guide keeps the rule and trims the footnote. Every year of ordinary Tuesdays where the rule and the SLA agree and nothing breaks is one more year of evidence that the explanation is dead weight. The boundary itself is never repealed. Instead, it’s diluted, hire by hire and quarter by quarter, until no one in the room is holding it.

The rule is engineered to persist. The memory of its limits is not; and memory is what has a half-life.

And the cruelty of it – the same shape I keep finding in this work – is that you cannot see the change by inspecting the rule. A policy whose founding conditions failed yesterday looks identical to one whose conditions still perfectly hold. Both appear the same on the dashboard and enforce with the same authority.

The only thing that distinguishes a living rule from a dead one is not merely a preserved why. It is a preserved account of when that why is sufficient, and when it is not.

Doctrine, Policy, and Dogma

It helps to separate three things we tend to file under the single word “standard,” because they age very differently.

Doctrine is a durable principle that travels with its reasoning. Because the ‘why’ is attached, doctrine remains corrigible. You can ask whether the justification still holds because the justification is available to interrogate. Doctrine does not correct itself, but it carries enough of its reasoning to be challenged on its own terms. It invites the question that keeps it honest.

Policy is a contextual decision compressed into a rule for recurring use, so that no one has to re-derive it every time. The fixed AHT ceiling, and the interpretation of variance from it as poor performance, were policy. And policy is efficient precisely because it drops the rationale at the point of use. The agent, the supervisor, and the forecaster follow the number without re-reasoning the whole chain that produced it. On most days, that is the feature. Removing the need to re-think is the entire value of a standard.

It is also, exactly, the vulnerability.

Dogma is what policy becomes when it continues to be applied as a complete answer after the conditions that made it sufficient have disappeared. Dogma is not a third kind of decision someone sits down and makes. It is the decay product of policy, plus time, plus a missing boundary.

Notice that dogma does not require anyone to forget the reason. This is where AHT sharpens the framework past the familiar version. The tired essay is the one about the fence in the field nobody remembers building; dogma by amnesia. But on iPhone launch day, nobody had amnesia. Every leader in that building could recite exactly why AHT mattered: the forecast, the capacity balance, the per-minute billing, the penalty clause.

The reason was not forgotten. It was fluent. What was missing was not the why, but the when: the boundary at which that reason stopped being sufficient to govern the whole situation. The reason was so well known that no one thought to ask whether the conditions that made it a complete answer were still present.

The failure was never “having policy.” Policy is how judgment scales; it is the radius in operational form, and an organization that refused all policy would drown in re-derivation. The failure is letting policy age into dogma with no mechanism to notice when its conditions have changed.

The Instrument That Kept the Number Still

It would be easy to tell the AHT story as the clear-eyed analyst who saw the rigid number for what it was while everyone else obeyed it. That is a flattering seat, but it is not an honest one, because I was not standing outside that machine diagnosing it. I may have been answering calls on the day of that launch, but a few versions later, I was one of the instruments that kept it rigid.

The forecasts I worked with ran on the static AHT assumption. The reports I helped produce scored agents against a target I knew, on certain days, the conditions had voided. The capacity plans and the billing math I fed were built on that number holding still; which means that flexing it, on the day it needed to flex, would have broken work my team was responsible for. I did not enforce a stale rule because I failed to see it. In some moments I enforced it because I saw exactly how much depended on it not moving. The rigidity was not a blind spot. Meeting AHT was a requirement for handling the expected volume, and I was part of the load path.

The better your judgment and the more authority it carries, the worse this problem gets, not the better.

Your decisions persist longer, propagate wider, and get questioned less because they arrive stamped with a credibility that discourages the very re-examination they need. The most trusted leaders and architects leave behind the most fences that no one dares touch.

Competence is not a defense against this failure. Conversely, it can be the very thing that makes your stale decisions dangerous, because competence is what makes people stop asking.

Encode the Doubt, Not Just the Decision

The weak version of the fix is “review your standards periodically.” Everyone nods, no one does it, and even where it happens, a scheduled review with no preserved rationale or boundary conditions is just archaeology. It becomes a team standing over a rule, guessing at why it exists, and usually deciding to leave it alone because tearing down a fence you can’t explain feels reckless.

The real obligation sits at the moment of creation, not at some hypothetical review years later. If your work is going to become durable, the whole point of senior work, then every durable decision has to carry, encoded as durably as the rule itself, three things it almost never carries today.

The rationale, written so a stranger who arrives after you can reconstruct why, not just what. The conditions of application – specifically, the invalidating condition: the sentence naming what would have to change for the rule to stop being sufficient. Not merely “here is when to use AHT,” but “here is what suspends its normal interpretation: a demand shock, an outage, or a product event that materially changes the call mix.” And a legitimate path to revision: a named person who is allowed to retire or suspend the rule, and a way for them to do it without heroism and without being punished for touching something that carried an authority they did not personally sign.

A contextual policy shipped with its enforcement but not its invalidating condition is unfinished architecture. It is a landmine with your name on it, waiting for the day the ground changes.

I framed the ethic of the senior IC path, in Path B, as answerability: being the person to whom the failure traces back, whether or not you held formal authority over anyone. This is what answerability actually requires once your judgment becomes precedent.

You are answerable not only for whether a decision was right when you made it, but for whether the people who inherit it years later can tell. Whether you left them the means to know if it still deserves to be obeyed. A decision handed down without its conditions is not a gift of judgment. It is a demand for obedience wearing judgment’s clothes.

Closing

AHT is still a core metric across the industry, and it should be. This was never an argument against the number itself, only against enforcing a target after the conditions that made it meaningful have changed. The difference between a standard and a superstition is not whether anyone remembers the rule, or even the reason. It is whether anyone can still say when the rule applies.

Durability, then, was never quite the achievement we treat it as. A decision whose enforcement outlives its applicability has not been preserved; it has only been left running. Durability without a preserved rationale and boundary is not stewardship. It is abandonment with extra steps: you left something powerful behind, and you took the instructions with you.

The radius of judgment measures how far your best thinking travels. It also measures how far your worst thinking keeps traveling, at full authority, long after you have stopped thinking it.

The work, then, is not to make fewer durable decisions. It is to stop shipping them without the one sentence that lets someone, someday, standing on a floor you will never visit, on a morning you could not have predicted, know that this was the day the rule was supposed to bend.

Leave a comment