A ticket is a promise with a clock
Most SLA breaches are accounting failures: a clock paused by the wrong party, never resumed, or two running on one ticket. Give each state one clock and one actor who can stop it.
A service-level agreement on a ticket is a promise with a clock attached: we will respond within four hours, resolve within two days. When the promise is broken, the post-mortem almost always finds that nobody was slow. The clock was paused by an agent's internal note and never resumed, or it kept running through a weekend the contract excluded, or two clocks were running on the same ticket because a reassignment started a second one. The breach was an accounting failure, and accounting failures come from modelling the ticket as a record with a status field and a nightly job that tries to reconstruct what the clocks should have done. The fix is to model the ticket as a state machine in which every state carries exactly one clock and names exactly one actor who can stop it.
The one-pause rule
The rule that makes the design work is small enough to state in one sentence: every SLA clock has exactly one party whose action can pause it, and that party is never the one being measured. The response clock measures the agent, so an agent's action cannot pause it. Writing an internal note, reassigning the ticket, changing its priority, none of those stop the clock, because they are things the measured party does. What stops the response clock is a customer reply, because the ticket is now waiting on the customer, or a scheduled callback that the customer agreed to, because the promise has been renegotiated with the person it was made to. The resolution clock has the same shape with a longer horizon.
Once the rule is in the state machine rather than in a settings page, the accounting failures become impossible rather than unlikely. An agent cannot pause a clock by accident because no agent action leads to a paused state. A clock cannot fail to resume because the resume is the transition the customer's reply performs, not a separate step someone has to remember. And two clocks cannot run on one ticket because a state carries exactly one, and a ticket is in exactly one state.
What a published clock looks like
Clocks are easiest to reason about when their targets are public, and the clearest published examples are the large cloud providers' support plans. AWS publishes first-response targets by severity and plan: under 24 hours for general guidance, under 12 for a system impaired, under 4 for a production system impaired, under 1 hour for a production system down, and for a business-critical system down, under 30 minutes on the Business plan, under 15 on Enterprise, and under 5 minutes from an incident engineer on the Unified Operations tier. Each of those is a clock that starts when the case is opened and stops when a person responds, and the whole table is a promise the customer can hold the vendor to.
A published table like that only means something if the vendor's own accounting is honest about when the clock was running, which is the point of the one-pause rule. If the vendor could stop the clock by posting "we are looking into this", the fifteen minutes would be a number rather than a promise.
Time as a query, not a job
The second consequence of the state machine is that elapsed SLA time stops being a stored field and becomes a query. Every transition is a row: ticket, from state, to state, actor, timestamp. The response time of a ticket is the sum of the durations it spent in states whose clock was the response clock, computed from consecutive transition rows, and the same for resolution. The business calendar, if the contract has one, is applied to those intervals at query time. Nothing is precomputed, nothing is cached in a column that can drift, and a report for last quarter is the same query with a date range.
The design is also what makes the report trustworthy to the people it measures. An agent who sees a breach can open the ticket and read the transitions: here is where the clock ran, here is who paused it and when, here is the customer reply that resumed it. There is nothing to dispute except the timestamps, and the timestamps were written by the system at the moment of the transition. A nightly job that recomputes a stored field produces a number nobody can trace, and a number nobody can trace becomes a number nobody trusts, at which point the SLA report is a formality and the promise behind it is not really being kept.
Calendars, reassignment, escalation
Three details decide whether the model survives contact with a real desk. The first is the business calendar. A contract that promises four working hours means the clock runs only inside the customer's business hours and holidays, and the state machine does not need to know that; the transition rows carry wall-clock timestamps, and the calendar is applied when the elapsed time is computed, by intersecting each running interval with the calendar's working periods. Changing the calendar, or correcting a holiday that was entered wrong, changes every report retroactively and correctly, which a stored elapsed-time column could never do.
The second is reassignment. Handing a ticket from one agent to another is an event the desk cares about, and it is recorded as a transition with an actor, but it does not change state and it does not touch the clock, because the promise was made to the customer, not to a particular agent. The same holds for internal escalation: the ticket moves up a tier, the row records who moved it, and the clock that was running keeps running.
The third is a change of target. When a ticket's priority is raised, or a customer's plan changes the promised response time, the target changes but the elapsed time does not; the clock has been running since the ticket opened and carries that time into the new target, which may already be breached at the moment of the change. That is uncomfortable and correct, and it is the reason a priority change is a transition with an actor rather than an edit to a field: the record shows who changed the promise and when, and the report shows what the change cost.
Where the vendors stop
Every ticketing product has SLA pausing, and their help pages describe it as a setting: choose the statuses that pause the clock. Atlassian's SLA conditions documentation is a good example of the shape, and the shape is a list of statuses with a checkbox. What the setting cannot express is the rule about who may move a ticket into a paused status, and that is where the accounting failures come from: any agent can set the status to awaiting customer, whether or not a customer is being awaited. The one-pause rule moves that decision out of the settings page and into the transitions, where the actor is part of the row and the state machine refuses the ones the rule forbids.
Whether a team builds its own desk or configures a vendor's, the rule is worth applying. Write down each state's clock. Write down the single actor whose action can stop it, and check that the actor is never the party the clock measures. Then look at the transition log for the last ten breaches. In my experience most of them were never slow work. They were a clock that someone stopped, and the state machine is how you make sure that only the right someone can.
Get new posts by email
Occasional essays on engineering, AI, and building for the people technology leaves behind.
Subscribe with RSS