First response is not resolution
A 1-hour first-response SLA is widely achievable with on-shift staffing. A 1-hour resolution SLA is not, except for password resets. Separate the two metrics and you stop the team from triaging tickets twice — once to acknowledge them, once again to resolve them — and you stop the customer from receiving an SLA breach notice on a ticket that was actually handled well. The customer cares about both numbers, and a transparent split is more credible than a single number that is constantly missed.
Tier by ticket type, not customer tier
Tiering SLAs by customer (gold/silver/bronze) creates queue-jumping and demoralises the team. Tiering by ticket type (incident, request, question, change) aligns SLAs to the work, which is what your team can actually control. An outage gets a 15-minute response SLA regardless of the customer; a how-to question gets a 4-hour response. The customer-tier model belongs to commercial conversations; the ticket-type model belongs to the helpdesk.
Publish the numbers
SLA performance hidden inside dashboards drifts. SLA performance published weekly to the whole company creates accountability — both ways. The team gets credit for hits and air cover when staffing slips. The internal customers get visibility of the realistic service level, which moderates expectations. A weekly all-hands email with the SLA performance and the staffing context is the single most underrated practice in helpdesk management.
Severity definitions that survive contact with reality
Severity 1 (major outage), Severity 2 (degraded service), Severity 3 (single user affected), Severity 4 (request). Each severity gets its own SLA, and the definitions must be tight enough that the team can categorise without arguing. The most common failure is severity inflation — every ticket logged as Sev 2 because the requester is annoyed. Build categorisation into the ticket form, audit the categorisation weekly, and recalibrate the team monthly.
Business hours, calendars and on-call
Most internal helpdesks operate on business hours, not 24/7. The SLA must reflect this — a ticket logged at 6pm Friday has a different clock to a ticket logged at 10am Tuesday. Out-of-hours coverage for critical incidents needs an on-call rota, defined escalation, and a separate (relaxed) SLA. Pretending the team is 24/7 when it is not produces the worst of both worlds: missed SLAs and burned-out engineers.
When to renegotiate
If the team is missing the SLA more than 20% of the time over two months, the SLA is wrong, the staffing is wrong, or the demand has shifted. Renegotiate openly, with data, and reset expectations. A renegotiated SLA that is hit is more valuable than an original SLA that is missed — both to the team and to the internal customers.
The metrics that matter beyond SLA
First-contact resolution, customer satisfaction (CSAT) on resolved tickets, ticket volume by category, and backlog age. These four together give a richer picture than SLA alone, and they are the metrics that drive real improvement. SLA is the headline; these are the diagnostics.
SLAs are a forecast of behaviour. Design them around the work and the team hits them; design them around the customer pitch and they miss — and you spend the next year fielding angry escalation calls.
