Content

Customer Effort Score: How to Measure CES and Reduce It with AI

Written by Lautaro Schiaffino | Aug 6, 2026, 12:00:00 PM

Last updated: August 5, 2026

Ask a customer whether they were satisfied and they will usually be polite. Ask them how hard it was to get their problem solved and they will tell you the truth. That distinction is the entire argument for Customer Effort Score — and it is the reason CES has quietly become the metric support leaders trust when CSAT looks fine but renewals do not.

The underlying research is unusually blunt. Gartner's work on the effortless experience found that 96% of customers who had a high-effort service interaction went on to show disloyalty, against just 9% of those with a low-effort experience — and that 94% of low-effort customers said they would buy again. Effort, not delight, is what moves the needle.

What this guide covers

What CES measures and how to calculate it

Customer Effort Score measures how much work a customer had to do to get what they needed. It is collected with a single agree/disagree statement — most commonly "The company made it easy for me to handle my issue" — answered on a 1-to-7 scale, where 7 is strong agreement.

The score is the mean of all responses. If 200 customers respond and the total of their ratings is 1,120, your CES is 5.6. Many teams also report the percentage of responses at 5 or above, because an average hides the shape of the distribution: a 5.0 built from mostly 5s is a healthy service operation, while a 5.0 built from 7s and 2s is a service operation with a broken segment.

Key takeaway: report CES as an average and the share of low scores. The average tells you how the operation is doing; the tail tells you which customers to call.

Wording matters more than the scale

The original phrasing asked how much effort the customer had to expend, and it consistently confused respondents: on an effort scale, a high number is bad, which inverts every other survey they have ever filled in. The agree/disagree version fixed that. Whichever you choose, keep it identical over time — rewording the question resets your baseline, and you will spend a quarter unable to tell whether service improved or the survey changed.

Why effort predicts loyalty better than satisfaction

CES came out of research by CEB, now part of Gartner, published in the Harvard Business Review article "Stop Trying to Delight Your Customers". The finding that made it famous was counterintuitive: exceeding customer expectations in service produced barely any loyalty gain over simply meeting them, while failing to make things easy produced a sharp loyalty penalty.

Mechanically, this makes sense. Satisfaction is a judgement about a whole relationship, so it is anchored by everything the customer already believes about you and moves slowly. Effort is a judgement about one specific interaction that just happened, so it moves immediately and points at something you can change on Monday. CSAT tells you the temperature; CES tells you which door is letting the cold in.

The practical consequence is a reordering of priorities. Service teams that optimise for delight invest in gestures. Service teams that optimise for effort invest in removing steps — and removing steps is cheaper, more measurable, and more likely to survive a budget review.

How to run a CES survey that produces usable data

Ask immediately after resolution, not after closure

Ticket closure is an internal event; resolution is the customer's event. Send the survey when the customer's problem is actually solved — within minutes of the final message, in the same channel the conversation happened in. A CES survey emailed three days later measures memory, not effort.

Keep it to one question, then one open field

The score tells you where the problem is; the free-text answer tells you what it is. One optional follow-up — "What made it hard?" — will generate more usable improvement backlog than any dashboard. Resist adding a third question.

Segment before you average

A single company-wide CES is almost useless for action. Split it by channel, by issue type, by whether the interaction involved a handoff, and by customer tier. The insight is nearly always in a segment: billing questions score two points below shipping questions, or chat scores well until it escalates to email.

The four effort drivers worth attacking

The Gartner research is specific about what customers experience as effort, and all four items are process failures rather than attitude problems.

Effort driver What the customer experiences What to fix
Repeat contactGetting in touch more than once about the same issueRoot-cause the top repeat reasons; resolve the next likely question in the first reply
Channel switchingStarting in chat, being told to send an emailGive every channel the same permissions and the same data access
Repeating informationExplaining the problem again to a second personPass full context on every handoff, including what has already been tried
Generic answersA help-centre link that does not address their caseAnswer with account-specific facts, not documentation

Notice that three of the four are caused by how work moves between systems and people. That is the same seam where first contact resolution breaks down — which is why FCR and CES tend to move together, and why fixing one usually improves the other.

How AI actually lowers effort

AI gets deployed against effort in two very different ways, and only one of them works.

The version that fails

A bot placed in front of the queue whose main function is to delay contact with a human raises effort. It adds a step, then hands over a conversation with no context, forcing the customer to start again. Gartner's own data is a caution here: a 2024 survey found only 14% of customer service issues are fully resolved in self-service. Deflection that does not resolve is effort with extra steps.

The version that works

An AI agent lowers effort when it can actually complete the job: read the account, check the order, apply the policy, issue the refund, update the record, and escalate with a full summary when it cannot. That is the difference between a chatbot and an agent with system access, and it is the premise behind Gartner's forecast that agentic AI will autonomously resolve 80% of common service issues by 2029, with roughly a 30% reduction in operational costs.

This is how we build Eva, Darwin AI's customer experience agent: she answers on WhatsApp and the channels your customers already use, resolves with real account data rather than help-centre links, and hands to a human with the full conversation attached so nobody has to repeat themselves. The design goal is not containment for its own sake — it is fewer steps per resolution.

The same logic applies before the customer ever contacts you. Proactive support is the lowest-effort experience there is, because the effort score of a conversation that never had to happen is unmeasurable in the best possible way.

Where CES fits with your other support metrics

CES is a lagging indicator of experience quality, so it needs operational metrics next to it to be actionable. A workable four-metric set looks like this:

  • CES — how hard it was for the customer. The outcome you care about.
  • Containment rate — what share your AI handled end to end. Only meaningful when read alongside CES, because containment achieved by exhausting the customer is not a win.
  • Cost per resolution — what each solved issue costs you. Keeps efficiency honest.
  • Ticket deflection — volume you removed. Valid only if the underlying issue was resolved.

Read together, these four catch each other's blind spots. Containment up and CES down means your AI is stalling customers. Cost per resolution down and CES flat means a genuine efficiency gain. Deflection up and repeat contacts up means you moved work rather than removing it.

Example: a B2B software team sees containment climb after launching an AI agent while CES falls slightly. Segmenting reveals the drop lives entirely in billing questions, where the agent lacked access to the invoicing system and could only advise customers to email finance. Granting that one integration moved both metrics the right way.

Frequently asked questions

What is a good Customer Effort Score?

On the 7-point agree/disagree scale, most teams treat 5.0 as a floor and anything above 6.0 as strong. Absolute benchmarks are weak, though, because scores vary by industry, issue complexity and survey timing. Your own trend line, segmented by channel and issue type, is a far better target than a published average.

Is CES better than NPS or CSAT?

It is better at a specific job: diagnosing individual service interactions. NPS measures relationship-level advocacy and CSAT measures overall contentment, and both are slower to react to process problems. Gartner's effortless experience research found 96% of high-effort customers showed disloyalty versus 9% of low-effort ones — that gap is what makes CES so useful for operations. Most teams run CES alongside one relationship metric rather than replacing anything.

When should we send a CES survey?

Immediately after the customer's issue is resolved, in the channel where the conversation took place. Delayed surveys measure recall rather than effort, and cross-channel surveys depress response rates.

Can AI improve CES, or does it make things worse?

Both happen. AI lowers effort when the agent can resolve the issue with real account data and hand off with full context; it raises effort when it functions as a gatekeeper. Gartner found only 14% of issues are fully resolved in self-service, so the decisive question is whether your AI has the system access to finish the job.

How many CES responses do we need?

Enough per segment to be stable, not enough for statistical elegance. If billing and shipping each return a few dozen responses a month, you can act on the difference between them. Chasing a company-wide sample size while every segment stays too small to read is a common way to collect data nobody uses.

Fewer steps per resolution

Eva resolves customer issues with real account data on WhatsApp, and hands off with full context when a human is needed.

Meet Eva