Executive summary
Every insurer and benefits provider carries a cost line that behaves badly: servicing. Cover questions, claims chases, enrolment queries, benefit look-ups - each one low-value individually, enormous in aggregate, and growing faster than the teams that handle them.
The economics are stark and well-documented. Gartner benchmarks put the median cost of an assisted service contact at $13.50 - around £10 - against $1.84, roughly £1.40, for self-service. We treat these as a conservative starting point: in our experience, the fully loaded cost of a resolved contact in UK benefits servicing runs higher once handle time, rework and chase loops are counted in full. Yet the industry's standard answer - portals, FAQs, intranets, static chatbots - has largely failed: in Gartner's research, only 9% of customers report fully resolving their issue through self-service, and a majority go straight to an agent. Every failed self-service journey becomes an assisted contact, which means most "digital deflection" investment to date has added channels without removing cost.
In employee benefits the problem is compounded by an understanding gap. MetLife's long-running US Employee Benefit Trends Study finds that 45% of employees do not fully understand their benefits package, and only 38% are completely confident they even know what is offered to them. Confused members generate more contacts, worse contacts - misdirected queries, incomplete claims - and, ultimately, unused benefits and unrenewed schemes.
Generative AI changes this calculus, but not uniformly and not automatically. Peer-reviewed evidence and large-scale deployment data now exist: a study published in the Quarterly Journal of Economics covering 5,172 support agents found generative AI assistance raised productivity 15% on average - with the largest gains, around 30%, among less-experienced agents; McKinsey estimates generative AI could reduce human-serviced contacts by up to 50% in comparable sectors; and more than half of leaders at Europe's largest insurers surveyed by McKinsey expect productivity gains of 10–20%.
This paper sets out where servicing cost actually comes from, why the volume curve outruns headcount, what the evidence says AI can and cannot resolve end-to-end, and a framework for modelling return honestly - including the deployments where it will not pay back.
1. The anatomy of servicing cost
Group benefits servicing has a distinctive shape. Unlike retail insurance, where the insurer faces one policyholder per policy, a group scheme places an insurer or benefits provider behind thousands of covered employees - most of whom do not know who the insurer is, what the policy says, or how a claim works, until the day they need to.
The resulting demand is dominated by a long tail of routine, resolvable queries:
- Coverage questions. Am I covered for physio? Does my plan include my partner? What is my excess? Answerable from policy documents - which almost no member reads.
- Claims process queries. How do I claim? What documents do I need? Where is my claim? "Where is my claim" chasers are the benefits equivalent of the WISMO ("where is my order") calls that plague logistics - status requests that create no value and consume agent time.
- Enrolment and eligibility. When can I add my child? What happens to my cover if I go part-time? Concentrated seasonally around enrolment windows, forcing capacity planning for peaks that sit idle the rest of the year.
- Benefit discovery. What wellbeing support is available? Is there an EAP? The queries providers should want - they drive utilisation - but which arrive at the same expensive channel as everything else.
Individually, none of these justifies a skilled human. Collectively, they define the operation's cost base: labour is the dominant expense of any contact operation, typically accounting for around 70% of contact-centre cost. And the volume is structurally rising, for three reasons the sector has engineered itself: product proliferation (voluntary benefits, wellbeing services and add-on covers multiply the surface area of possible questions), rising member expectations set by consumer-grade digital experiences, and - in the UK - the Consumer Duty's support outcome, which converts poor servicing from a cost problem into a regulatory one.
2. Why the standard playbook failed
The FCA's position, restated consistently through 2025 and 2026, is that it will avoid additional regulation for AI by relying on existing frameworks. There is no AI rulebook. There are, instead, three existing regimes that together answer almost every question a firm will face when deploying an AI assistant:
The Consumer Duty. The Duty requires firms to act to deliver good outcomes for retail customers across four outcomes: products and services, price and value, consumer understanding, and consumer support. An AI assistant sits squarely inside the last two. If an assistant gives a member a wrong answer about their cover, that is a consumer understanding failure. If it strands a claimant in a loop with no route to a human, that is a consumer support failure - and the FCA's guidance on the fair treatment of vulnerable customers (FG21/1) applies with particular force to a channel that vulnerable members may reach at moments of bereavement, illness or financial distress. The Duty is outcomes-based and technology-neutral: the firm cannot outsource accountability for the answer to the model, or to the vendor.
The Duty also cuts the other way, in the deploying firm's favour. Its monitoring obligations require firms to evidence customer outcomes - and an AI assistant, properly instrumented, generates exactly that evidence: what customers asked, what they misunderstood, where journeys failed, which cohorts struggled. Interaction data that a call centre loses in disposition codes becomes, in an AI channel, a continuous Consumer Duty outcomes dataset.
The Senior Managers and Certification Regime. The SM&CR answers the accountability question the Treasury Committee raised: a named senior manager is responsible for the activities within their remit, and deploying an AI system does not dilute that responsibility. The Bank/FCA survey found 84% of firms using AI already assign an accountable person for their AI framework, and 72% place accountability with executive leadership. The practical implication for procurement is that the accountable SMF-holder must be able to understand and defend the system - which sets a floor on the explainability and auditability a vendor must provide.
Existing conduct and systems-and-controls rules. Outsourcing rules, operational resilience requirements, and SYSC governance obligations all apply to third-party AI exactly as they apply to any other material outsourced service. With a third of AI use cases now third-party built, the FCA's supervisory interest in vendor governance is structural, not incidental.
On top of this, the FCA is generating supervisory expectations through engagement rather than rulemaking: the AI Lab, the Supercharged Sandbox, Live Testing cohorts through 2026, the Mills Review (launched January 2026, with input closing that February and findings published in July 2026), and a good and poor practice publication due later in 2026. That publication will be the nearest thing UK financial services has to AI guidance. Firms deploying now should design against the direction of travel it will codify: the FCA has said it evaluates the AI system - model, deployment context, governance, human-in-the-loop arrangements, and input and output controls - not the model in isolation.
3. The understanding gap: benefits' specific multiplier
In group benefits, servicing demand is inflated by a factor that logistics or telecoms do not face: the product itself is poorly understood by design of circumstance. Members receive cover chosen by their employer, documented in policy language, encountered rarely and usually under stress - a bereavement, an illness, an accident.
MetLife's US Employee Benefit Trends Study - the longest-running of its kind - quantifies the gap: 45% of employees say there are elements of their benefits package they do not fully understand, and only 38% are completely confident they know about everything offered to them. The same research shows what understanding is worth: 76% of employees who understand their benefits report being happy at work versus 47% of those who don't; employees who understand and are satisfied with their benefits are 1.4 times more likely to feel engaged and 1.2 times more likely to be productive at work; and half say a better understanding of their benefits would make them more loyal.
For the provider, the understanding gap converts directly into cost and lost value at every stage: misdirected first contacts that require triage before they can be resolved; incomplete claim submissions that generate rework loops (each chase another £10-class contact); benefits that go unused because members never discover them - undermining the utilisation data that drives employer renewal decisions; and, in the UK, weaker evidence for Consumer Duty's consumer understanding outcome.
The strategic point: in employee benefits, servicing cost and product value are the same problem. Every unanswered question is both an operational expense and a failure of the product to land. Fixing the first without fixing the second - pure deflection - leaves the more valuable half of the prize on the table.
4. What the evidence says AI actually changes
The claims made for AI in customer service have run ahead of evidence for years. That is no longer necessary; three classes of evidence now exist.
Peer-reviewed field data. A study published in the Quarterly Journal of Economics (Brynjolfsson, Li and Raymond, "Generative AI at Work") followed 5,172 customer support agents given access to a generative AI assistant. Productivity - issues resolved per hour - rose 15% on average, with substantial heterogeneity: less-experienced and lower-skilled agents improved both the speed and the quality of their work by around 30%, with the AI effectively transferring the practices of top performers, while the most experienced agents saw small speed gains and small quality declines. This is the strongest evidence available for the assist model: AI making human agents faster and better, particularly at the inexperienced end where attrition-driven churn hurts most - and a reminder that the gains are not uniform.
Sector-level analysis. McKinsey's research on generative AI's economic potential identifies customer operations as the single most immediate opportunity, estimating that the technology could reduce human-serviced contact volume by up to 50% in sectors such as banking and telecommunications, with productivity impact worth 30–45% of current function cost. In its November 2024 survey of leaders at Europe's largest insurance groups, more than half expected generative AI to deliver productivity gains of 10–20%. McKinsey's insurance practice is equally clear on the failure mode: applying AI to individual use cases without redesigning the end-to-end process yields a fraction of the value of domain-level transformation.
Deployment benchmarks. IBM's 2020 study of deployed virtual agents reported an average containment rate of 64% - defined as the share of the queries the system had been trained to handle that it resolved without human involvement, among organisations that were already adopters; Gartner forecast in 2025 that agentic AI will resolve 80% of common customer issues autonomously by 2029, with an associated 30% reduction in operational costs, having earlier projected that conversational AI would reduce contact centre agent labour costs by $80 billion in 2026. Containment varies enormously with query mix, scope and implementation quality; the difference between a grounded system that knows the member's actual products and a generic chatbot is the difference between resolution and another failed channel.
Taken together, this evidence supports a specific and bounded claim, not a general one. It says AI can materially raise agent productivity, and can contain a real share of routine contacts end-to-end - not that it resolves everything, and not without a governed knowledge base underneath it. Two limits follow directly from the evidence itself, and an honest model prices both in.
The evidence does not extend to complex or sensitive interactions. Complex claims requiring judgment, empathy-critical interactions (bereavement, serious diagnosis, financial distress), vulnerable customers, and genuine disputes should route to humans by design - both because containment attempts fail expensively (every failed automation becomes an escalation costing more than the original call would have) and because regulators expect it.
The evidence assumes a governed knowledge base, which most providers do not yet have. AI does not rescue a bad one: systems grounded in outdated or incomplete product data automate the delivery of wrong answers. The prerequisite investment - consolidating products, policies, eligibility rules and provider networks into a governed source of truth - is unavoidable, and providers with fragmented product data should cost it into any business case.
5. A framework for modelling return
Four variables determine the economics of any deployment. Providers should insist on seeing all four, with assumptions exposed, in any vendor's business case - including ours.
1. Baseline volume and mix. Annual assisted contacts, decomposed by query type. In group benefits the majority typically sits in the routine categories of section 1 - but the split is empirical, not assumed, and it determines the ceiling on containment. Insist on the provider's own contact data by category, not a sector average.
2. Realistic containment by category. Not a blended aspiration. Coverage questions grounded in actual policy data can contain at high rates; claims status queries near-completely, given systems integration; complex claims should be modelled at or near zero containment and treated as a routing improvement instead.
3. Fully loaded channel costs. Gartner's medians - $13.50, around £10, for assisted contacts, and $1.84, roughly £1.40, for self-service - are a starting point. In Onsi's experience, the fully loaded cost per resolved contact in UK benefits servicing runs higher once handle time, quality overhead, rework and chase loops are counted in full, which makes the Gartner medians a conservative basis for the model. The 80–100x cost multiplier for failed self-service belongs in the model too: an AI channel with poor containment can add cost.
4. The revenue side. Every contained interaction is also a discovery and data opportunity: benefits surfaced at the moment of need drive the utilisation that determines employer renewal, and interaction data evidences Consumer Duty outcomes and exposes product gaps. Deflection-only models undervalue the deployment; state utilisation and renewal effects separately from cost savings, with their own assumptions, rather than folding them into a blended number.
Worked illustration. The figures below are deliberately generic - Gartner's US-dollar medians converted to sterling at prevailing rates - and providers should substitute their own throughout.
A book of 200,000 covered employees generating 0.75 contacts per member per year produces 150,000 assisted contacts a year, roughly £1.5m at around £10 per contact. At 60% end-to-end containment of that volume at self-service-class cost, gross servicing savings run to approximately £775,000 annually - before platform and implementation costs, and before any assist-model productivity gains on the residual human-handled volume, worth a further 15% per the QJE evidence. Below roughly 30% containment, or on books under ~50,000 lives, the case weakens and should be tested honestly.
On the revenue side, one datapoint frames the prize: MetLife's 2026 study found 73% of employers regard non-medical benefits as their most cost-effective wellbeing lever - but only used benefits count. Our companion paper on the unused-benefits problem sets out the growth economics in full.
About Onsi
Onsi provides AI assistants for insurers and benefit providers, combining trusted company knowledge, business rules and enterprise governance to deliver accurate, auditable employee experiences across benefits, claims and wellbeing. Onsi is a global technology provider and a UK and EU insurance intermediary: Onsi is a trading name of Collective Society Ltd, Collective Denmark ApS and Collective Netherlands B.V., authorised and regulated by the UK Financial Conduct Authority (No. 923788), the Danish Financial Services Authority (No. 42352985) and the Netherlands Authority for Financial Markets (No. 12049041) respectively. Onsi is ISO 27001 and Cyber Essentials Plus certified.
Conclusion
Servicing cost in employee benefits is not a staffing problem, and a decade of channel proliferation has demonstrated it is not a portal problem either. It is a resolution problem: the industry built channels that could not answer the member's actual question about their actual cover, and paid for the failure twice.
The evidence base now supports a different claim than the chatbot era's: grounded AI assistants can resolve the routine majority of benefits interactions end-to-end, make human agents measurably better on the remainder, and convert servicing from a pure cost line into the provider's richest source of product and outcomes data. The providers that capture this will be the ones that model it honestly - category by category, with the failure modes priced in - rather than buying a blended promise.
Sources
- Gartner, Benchmarks to Assess Your Customer Service Costs (median cost per contact: $13.50 assisted, $1.84 self-service), 2024. Sterling equivalents are approximate conversions at prevailing exchange rates.
-
- Gartner, 2019 Customer Service and Support Leader poll and related 2019–2020 research: live channels averaging $8.01 per contact vs. ~$0.10 self-service; 9% complete self-service resolution across 8,000 customer journeys; 53% of customers going directly to agents; 80–100x cost of channel switching.
- Gartner, forecast that conversational AI would reduce contact centre agent labour costs by $80 billion in 2026 (2022); forecast that agentic AI will resolve 80% of common customer service issues autonomously by 2029, with a 30% reduction in operational costs (March 2025).
- Brynjolfsson, E., Li, D., and Raymond, L., "Generative AI at Work," Quarterly Journal of Economics, 140(2), May 2025, pp. 889–942 (study of 5,172 customer support agents; +15% average productivity; ~+30% in speed and quality for less-experienced and lower-skilled agents; small quality declines for the most experienced).
- McKinsey & Company, The Economic Potential of Generative AI: The Next Productivity Frontier (June 2023): customer operations impact, up to 50% reduction in human-serviced contacts in comparable sectors; productivity value of 30–45% of function costs.
- McKinsey & Company, The Potential of Generative AI in Insurance (November 2024): survey of 50+ leaders at the largest European insurer groups; expected productivity gains of 10–20%.
-
- IBM Institute for Business Value with Oxford Economics, virtual agent deployment study (October 2020): 64% average containment rate, defined as the share of contacts the virtual agent technology had been trained to handle that were resolved without human involvement; survey of 1,005 organisations already deploying the technology.
- MetLife, U.S. Employee Benefit Trends Study (21st–24th annual editions, 2023–2026): 45% of employees do not fully understand their benefits; 38% completely confident in knowledge of offerings; happiness (76% vs 47%), engagement (1.4x) and productivity (1.2x) differentials for employees who understand their benefits; 73% of employers citing non-medical benefits as most cost-effective wellbeing support (2026 edition). US data; cited here as the longest-running study of its kind.
This paper is provided for general information. Benchmark figures are drawn from the cited third-party research and are not representations about outcomes on any particular book of business. References to Onsi's experience of servicing costs reflect Onsi's own deployment observations and are not third-party benchmarks.

