Continuous Duty

Photo by Liam Briese on Unsplash

There were always people on the floor who weren’t on a call.

If you walked the aisles at eleven in the morning on an ordinary Tuesday, you’d see it right away. A headset pushed down around someone’s neck. Someone typing up notes from the last interaction. Someone leaning back, waiting for the next one to arrive. A contact center running exactly as designed still has, at any given moment, a visible fraction of its people not doing the thing they are there to do.

The economics are what make the rest of this worth reading.

We paid for every minute. Wages, benefits, seat cost, the building itself, the software licenses; all of it accrued whether an agent was speaking to a customer or waiting for a call to arrive. And we billed the client per call-minute. Only the minutes spent actually handling contacts converted into revenue.

Every idle minute on that floor was a direct, unrecovered subtraction from margin. Not a soft cost, or an opportunity cost; a minute we bought and could not sell.

So the obvious question got asked, and it got asked often, usually by someone new to leadership, and never unreasonably: why are we staffing people who aren’t on calls? Increase occupancy, get the number to a hundred.

We planned for around eighty-five.

Not as a compromise, and not because anyone on that floor was being protected. Nobody at eighty-five percent occupancy ever felt like they were coasting. We planned for eighty-five because the last fifteen points cost more than they returned, and we had the data to show it.

Those final fifteen percentage points of utilization were always far more expensive, and far less productive, than most new leaders expected.

What Happens at Ninety-Five

Here is what nobody expects about pushing a floor past its designed occupancy: nothing breaks.

There’s no alarm, failure, or moment where the system refuses the work. Nobody walks off. Nobody says I can’t. If you were watching the wallboard, you’d see the number you asked for, and you’d conclude you had won an argument with physics.

What actually happens is that the floor starts to drift.

Time migrates. After-call work stretches, and the notes that took thirty seconds before start taking two minutes, then three. Research time expands, and agents who used to resolve a case from memory start looking it up. And the calls themselves get longer: the same problem, the same scripts, and the same customer profile, taking measurably more time this month than last, for reasons no supervisor can point to and no agent could explain if you asked.

Nobody decides any of this. There is no meeting where the floor agrees to slow down. It simply becomes true, a few seconds at a time, distributed across hundreds of people and thousands of contacts until it’s large enough to see in the aggregate and far too late to trace to a cause.

That was the thing worth understanding, and it took me longer than it should have:

We hadn’t eliminated the fifteen percent. We had moved it.

The reserve didn’t disappear when we stopped planning for it; it reappeared somewhere we weren’t looking, inside the states we were counting as productive. And it came back at a worse price than the one we’d refused to pay.

Idle time is cheap and honest: it sits in one column, it’s easy to measure, and you can staff against it. The same capacity, reabsorbed involuntarily into handle time and after-call work, arrives with interest: longer queues, more callers waiting, more abandons, more callbacks on issues that should have closed the first time. We had traded a visible cost for an invisible one and paid a premium for the privilege.

Call it the involuntary reserve. Every loaded system has one. The only real decision is whether it sits somewhere you can see it.

You don’t get to choose whether the margin exists. You only get to choose whether it appears on a report.

The Curve Was Never Optional

None of this was a novel discovery, or anything that would surprise someone who has worked in workforce management. The math had been sitting in the industry the whole time, and the same shape recurs elsewhere.

Queueing systems fail in a specific and counterintuitive way. In the simplest model, the average time a unit of work spends waiting is governed by the gap between how fast work arrives and how fast it can be served. As those two converge, wait time doesn’t rise toward some tolerable ceiling; it goes to infinity. The curve is not linear and it was never going to be.

It is worth putting numbers on it, since that is the whole argument of this essay. In the simplest queueing model, the average time work spends waiting – measured against the time the work itself actually takes – runs roughly four times at eighty percent utilization, nine times at ninety, and nineteen times at ninety-five. Read that sequence again. The five points between ninety and ninety-five add more delay than the entire first eighty points combined.

The difference between eighty and ninety percent utilization is not the same size as the difference between ninety and a hundred; the second one is a cliff only a mathematician would still call a slope.

Traditional inbound contact-center staffing often runs on Erlang C, which is this curve with a service-level target strapped to it. It’s worth saying plainly what that means: the industry that most wanted saturation, that was most ruthlessly measured, that had the clearest possible financial motive to squeeze the last minute out of every hour, built its entire staffing discipline around a model that told it saturation was unaffordable.

We didn’t arrive at eighty-five percent through compassion. We arrived there through arithmetic we couldn’t argue with.

Traffic engineering describes the same shape in a form everyone has experienced personally. Flow – cars past a point per hour – rises with density right up to a critical threshold, and then falls. Past that point, adding vehicles reduces throughput. At maximum density, the road holds the most cars while delivering almost none of them. Every commuter has watched a highway approach one hundred percent utilization and produce nearly nothing.

The two failures look different but they converge. Latency explosion becomes throughput collapse the moment you add rework: callers abandon and call back, tasks half-finished get restarted, work already done once gets done again because it was done imperfectly the first time. A queue that was merely slow starts generating its own volume, and the system begins competing with its own backlog.

A system at full utilization isn’t producing more. It’s producing later, and then producing it twice.

Rated for Forever

The metaphor everyone reaches for is the redline, and it’s a good one – but it’s usually deployed backwards, as an argument for restraint. The more interesting fact is what the redline actually marks.

It isn’t where the engine is strongest. In most production engines, peak power arrives somewhere below the redline; past that peak, holding the throttle wide open produces less power, not more, while the mechanical stress keeps climbing.

The redline is not the summit. It is the point beyond which performance gives way to accumulated consequence.

Which means the metaphor was never don’t live at the limit because it’s dangerous. It’s sharper than that.

The limit isn’t even the best place to operate. You give up output to get there, and then you pay for the privilege in wear.

That is the first of two costs, and it is the one you can observe immediately. The redline describes what happens to output in the moment. The second cost is slower, and it is what happens to the machine over time – a distinction electrical engineering makes explicit enough to stamp on the outside of the equipment.

Motors carry duty ratings: continuous duty, the load a machine can carry indefinitely without exceeding its thermal limits, and short-time duty, the substantially higher load it can carry for a defined interval before it must rest. Both numbers are true, and both are printed. Confusing them, though, destroys the motor. It may not happen immediately, or dramatically, but accumulating heat faster than the machine can shed it only yields one result over time.

Serious engineering practice distinguishes what a thing can do from what it can do repeatedly. Publishing both numbers is considered so basic an obligation that they appear on the nameplate, in the open, where anyone can read them.

The rating on the plate isn’t the largest number the machine can produce. It’s the largest number it can produce again tomorrow.

Why We Admire It Anyway

So the engineering is settled, the math is old, and the operational evidence is available to anyone who has ever staffed a queue. Yet the aphorism survives untouched: give one hundred percent.

We say it to children. We put it on walls.

The persistence isn’t stupidity, and it isn’t only incentives. It’s an aesthetic problem, and I think it works like this.

Effectiveness is not directly observable. You cannot look at a person and see whether their judgment was sound, whether the architecture will hold, whether the problem they solved on Tuesday was the one that mattered. Those things resolve later, elsewhere, usually in the form of something that didn’t happen. Strain, by contrast, is immediately legible. Hours are countable. Responsiveness is timestamped. Exhaustion is visible across a room and audible on every call.

So strain becomes the proxy for effectiveness, and it is specifically the proxy reached for by someone who cannot evaluate the work itself. Visible exertion is what competence looks like from outside the domain. The twelve-hour day, the instant reply, the calendar with no room left in it: these are signals designed to be read by an observer who has no other way to tell whether you’re any good.

I’ve written before about the work organizations can’t see, and how reliably they remove it. This is a step past that. We don’t merely fail to observe effectiveness. We substitute something we can observe, and the thing we chose to substitute is suffering.

That substitution explains why the argument for margin always sounds like special pleading. Someone protecting capacity is producing less of the visible signal. They look, precisely, like someone giving less.

In the only currency the room can read, they are.

I Priced the Margin

It would be easy to write all of this from the outside, but that would be dishonest, because for a while, I was the one holding the pencil.

I managed the forecasts. I sat in the meetings where we decided what fraction of an FTE (read: human being) was committed in advance, and I defended eighty-five to people who wanted the other fifteen, and I was right. But I need to be exact about why I was able to win that argument, because the reason was less flattering than the outcome.

I could hold the line because the degradation was measurable. Occupancy was a number on a wall. Erlang gave me a defensible figure, generated by a model the client had already accepted, and I could put a curve in front of a finance director and show him what the last fifteen points would cost. I didn’t need anyone to trust my judgment about human capacity. I could win with a spreadsheet.

Knowledge work has no Erlang.

There is no occupancy metric for an architect. Nothing on any dashboard shows a senior engineer crossing the threshold where thinking starts quietly converting into rework. The drift still happens; decisions get made a little faster and a little worse, the review that would have caught it gets skimmed, the design that needed a week of ambiguity gets three days – but there’s no wallboard, no model, no number I can carry into a room.

It would be convenient to conclude that the absence of the model means the absence of the curve. It doesn’t, and the honest version is worse.

What makes the curve steep was never the queue; it was the variability. How unpredictably work arrives, and how unevenly long it takes once it does. Measured that way, a contact center is a remarkably regular place. Arrival patterns can be forecast weeks out, handle times cluster tightly enough to average, and the work itself is bounded by a script. Knowledge work has none of that. Requests arrive without pattern, interruptions arrive without notice, and an honest estimate for a piece of design work is routinely wrong by a factor of two.

Higher variability doesn’t shift the curve. It steepens it. Degradation begins earlier and accelerates harder, which means the reserve required to stay ahead of it is larger than the fifteen points I fought for on a floor where everything was counted. We give it none, and we call the result a scheduling problem.

So I have almost certainly failed at this exactly where it costs the most. At times, I protected margin where I could count it. I have spent the years since operating in a domain where I cannot, quietly behaving as though the curve no longer applies to me.

And here is the part that should be uncomfortable for anyone who works in a modern knowledge organization: the contact center, the most relentlessly metered environment most readers can imagine, is better at this than most of us are. Not more humane, perhaps, but better.

It was better because it could count, and because when it counted, it was honest about what the count said.

Too often, we congratulate ourselves on treating people as more than throughput, and then we schedule them at a hundred percent because nothing stops us.

Continuous Duty

The conclusion is not that anyone should operate at seventy percent, and I don’t think there’s a universal number worth defending. Eighty-five was ours; it belonged to that queue, that call mix, that service-level agreement, and it would be wrong somewhere else.

The obligation elsewhere is narrower, and harder to dodge.

Every durable commitment we make carries two ratings. A staffing plan, a delivery date, an on-call rotation, a roadmap, a team, a person – each has a peak figure and a continuous figure, and they are never as close as we think they are.

We know this. We prove that we know it every time we describe an effort as heroic, because heroic is just the word for a short-time rating we’ve decided not to write down.

We start early. Give one hundred percent is the first performance standard most people ever receive, and it arrives as a peak figure with no continuous figure attached. We are careful with children about every other limit: what a bridge holds, what a dose should be, how fast the road allows, when a fever stops being something you wait out. The only machine we describe to them using nothing but its maximum is themselves. By the time the same sentence turns up on an office wall, it isn’t advice anymore; it’s a rating nobody remembers agreeing to.

And then we publish one number. We quote the peak, usually established during an emergency that everyone survived, and we schedule against it as though it were the sustainable one. That single silent substitution lies beneath a remarkable share of what we later call poor execution, attrition, burnout, or bad luck.

Naming the rating is not modesty and it isn’t self-protection. It’s the same discipline as printing a duty cycle on a nameplate: a statement of what the thing can actually do, made in advance, in public, so that the people depending on it can plan against something true.

Anyone can find out what a system will do at the limit. Run it there and wait. The harder work, and often the work that requires knowing the curve well enough to defend a number nobody in the room wants to hear, is establishing what it will do on the thousandth day.

Continuous duty isn’t the lesser duty. It’s the only kind that can be performed twice.


Leave a comment

Or check out another recent post directly!