The insurance industry has been running AI pilots for the better part of a decade. Most of them never reach production.
A 2024 Howden survey found that 80% of insurance leaders cite competitor AI adoption as their top concern. Yet the same leaders, when asked about their own deployments, describe a familiar pattern: a promising pilot, enthusiastic results in a controlled environment, and then somewhere between the proof of concept and the production cutover - the project stalls.
This is not a technology problem. The AI works. The problem is that most AI deployments in insurance are built on foundations that were never designed for the operational realities of the industry.
This article explains what those foundations are, why they matter, and what agentic AI deployments that actually reach production and stay there look like in practice.
Why Most Insurance AI Pilots Fail to Scale
Before examining what works, it is worth being specific about what goes wrong.
The generic model problem:
Most enterprise AI vendors offer a horizontal platform, tools built for any industry, configured for yours. In theory, this sounds efficient. In practice, it creates a fundamental problem: the models do not understand insurance documents.
A Table of Benefits (TOB) is not a generic document. It contains nested benefit structures, plan-specific exclusions, co-payment rules, and regulatory references that vary by jurisdiction. A loss run contains information that only makes sense in the context of underwriting history. A slip, the foundational document of the London market, has a structure that has evolved over centuries of Lloyd's practice.
Generic large language models, when asked to parse these documents, reach approximately 60 to 70 percent field accuracy. That sounds reasonable until you consider what the remaining 30 to 40 percent means in production: incorrect coverage determinations, missed exclusions, wrong reserve calculations, and compliance failures.
Insurance-trained models, built on actual policy documents, TOBs, slips, and loss runs, reach 97 percent field accuracy on the same tasks. The difference is not marginal. At the volume insurers operate, millions of documents annually, a 30-point accuracy gap is operationally catastrophic.
The integration problem:
Insurance operations run on core systems that were not designed with AI integration in mind. Guidewire, Duck Creek, eBao, Sapiens- these platforms are the source of truth for policy and claims data. Any AI deployment that cannot read from and write to these systems in real time is, at best, a parallel workflow that creates more reconciliation work than it saves.
Most AI pilots are built in isolation from core systems. They run on exported data, operate in separate environments, and produce outputs that have to be manually reconciled with the production system. When the pilot ends, and the question of production deployment arises, the integration work begins, and it is almost always larger than anyone anticipated.
The regulatory problem
Insurance is one of the most regulated industries in the world. An AI model that makes a coverage determination, a claims decision, or a pricing recommendation is making a decision that has regulatory consequences. In the UK, the FCA requires that AI-driven decisions be explainable. In India, IRDAI mandates audit trails for automated decisions. In Saudi Arabia, the Saudi Arabian Monetary Authority has specific requirements for AI governance in financial services.
Most AI pilots are not built with regulatory compliance as a design requirement. They are built to demonstrate capability, with compliance treated as something to be addressed before go-live. The result is that models that work perfectly in a pilot environment can't be deployed in production because they can't produce the audit trails, explanation frameworks, and override mechanisms that regulators require.
The handoff problem
Insurance workflows are not single-step processes. A motor claim moves from first notice of loss to triage to adjudication to settlement. A health pre-authorisation request moves from submission to eligibility verification to benefit resolution to approval. At each stage, different people, different systems, and different rules are involved.
Most AI deployments are built to automate a single step. The result is a series of AI-assisted islands connected by manual handoffs. The adjudicator still has to gather information from four different systems before making a decision. The customer still has to call to find out the status of their pre-auth. The operational friction that AI was supposed to remove has simply been redistributed rather than eliminated.
What Agentic AI Actually Means in Insurance
The term "agentic AI" is used loosely in most industry commentary. In the insurance context, it has a specific meaning that is worth defining precisely.
An agentic AI system in insurance is one that can execute multi-step workflows autonomously, gathering information, making decisions, triggering actions, and escalating to human reviewers based on confidence thresholds, without requiring human intervention at each step.
This is categorically different from an AI assistant that helps a human do their job faster. It is also different from an automation tool that follows fixed rules. Agentic AI is capable of handling variation, ambiguity, and exception, the characteristics that define real insurance workflows as opposed to textbook descriptions of them.
The practical difference is significant. A non-agentic AI claims tool might extract information from an FNOL and present it to an adjudicator. An agentic claims system receives the FNOL, validates coverage against the policy, scores fraud risk, estimates the reserve, assigns the handler, and presents the adjudicator with a pre-filled brief, all before the adjudicator opens the file.
The first tool saves minutes. The second transforms the economics of claims handling.
What Production Agentic AI Looks Like Across the Value Chain
Having established what goes wrong and what the technology actually means, it is worth examining what successful production deployments look like in practice.
Claims: Motor FNOL
Motor claims processing is one of the highest-volume, most standardised workflows in general insurance. It is also one of the most consistently manual, a combination that makes it an ideal candidate for agentic automation.
A production motor FNOL pipeline operates as follows:
The claimant submits their FNOL via WhatsApp, app, or a web portal. The agentic system generates a claim reference, identifies the claimant against the policy database, and begins document processing in parallel, extracting information from the police report, repair estimate, driving licence, and any supporting photographs.
Coverage validation runs seven checks simultaneously: policy status, expiry date, territorial limits, notification period compliance, peril matching, exclusion screening, and sublimit application. Depending on the market, jurisdiction-specific rules, Saudi IA guidelines, IRDAI requirements, or FCA standards are applied automatically.
Fraud scoring runs against cross-claim pattern analysis, location and time data, and third-party network verification, producing a score on a zero to one hundred scale with specific flags for review.
Reserve estimation benchmarks against peer claims data, adjusted by the AI for the specifics of this incident, with upper and lower confidence bands.
Handler assignment matches the claim value against handler authority levels and current workload, with specialism matching for complex cases.
The adjudicator receives a pre-filled case brief: coverage confirmed, exclusions cleared, fraud score, reserve recommendation, and a routing decision. Their job is to review, approve, flag, or escalate, not to gather information that the system has already assembled.
In production, this pipeline reduces standard handling time from four to five hours to under ninety seconds. Cost per claim falls from approximately $45 to under one dollar. Decision accuracy, measured against adjudicator review outcomes, improves from 84% to 97 %.
Health: Pre-Authorisation
Health pre-authorisation is a workflow that affects patient outcomes, operational efficiency, and regulatory compliance simultaneously. It is also one of the most paper-intensive processes in insurance, a combination of member validation, benefit resolution, clinical rule application, and approval issuance that can take hours or days when done manually.
A production health pre-auth system receives the request from the hospital, member ID, ICD-10 code, procedure code, treating physician, and estimated cost, and processes it through nine sequential checks: member validation, policy status, provider network verification, TOB parsing, benefit resolution, co-payment calculation, gatekeeper rule application, regulatory compliance, and confidence scoring.
Cases that meet the confidence threshold are approved automatically, with the approval, including the approved amount, co-payment, validity period, and regulatory reference, sent to the hospital and the operations team simultaneously, following the regulator-specific guidelines.
Cases below the confidence threshold route to the medical operations team with the complete AI brief, not a blank form, but a fully populated case that requires clinical judgement on a specific question.
In production, this reduces pre-approval time from hours to under two minutes for standard cases. It also eliminates a significant proportion of pre-auth related inbound calls, patients calling to check the status of a request that the system has already processed.
Customer Experience: The Diagnostic Loop
Most insurers treat customer service as a cost centre to be managed rather than a source of operational intelligence. Calls are handled, resolved where possible, and logged. The patterns in those calls, the systemic failures that generate repeat contacts, are rarely surfaced back to the operations teams who could fix them.
An agentic CX system audits one hundred per cent of customer calls, not the two to five percent sample that traditional quality assurance reviews. It identifies not just whether calls were handled well, but why the calls were made in the first place.
When thirty percent of calls in a given period relate to customers asking about the status of a pending pre-authorisation, that is not a customer service problem. It is an operations problem, a confirmation notification that was never sent, or a processing queue that has grown beyond acceptable length. The agentic system identifies the root cause and triggers the operational fix: a push notification to all customers with pending pre-auths, or an alert to the operations manager about queue depth.
In production, this approach identifies two to three times more operational pain points than traditional QA sampling. Over six months, it produces measurable reductions in customer churn and significant improvements in first-contact resolution rates.
Note: These are only a few areas where Agentic AI is making an impact. It is hard to write about the entire value chain in one article.
The Platform Architecture That Makes Production Possible
The deployments described above share a common architectural requirement that is worth making explicit: they all depend on a single, unified risk record per customer.
This sounds like a technical detail. It is actually the most important architectural decision in insurance AI deployment.
Consider what happens in a motor FNOL when the agentic system needs to check claims history, policy status, and previous fraud signals simultaneously. If those three data sources live in three different systems with three different data models, the system has to resolve identities across all three in real time. If they have been unified into a single risk record, one representation of this customer, updated by every interaction, the system reads from a single source and acts immediately.
The same logic applies across every workflow. The renewal agent reading a customer's claims history, channel preferences, and NPS score. The cross-sell agent identifying the right product based on life events and wallet share data. The pre-auth agent checking member enrolment and benefit history simultaneously. All of them depend on a complete, current picture of the customer.
Building that unified risk record is not a trivial exercise. It requires resolving identities across policy administration, claims, billing, and customer service systems, systems that were often built decades apart and were never designed to talk to each other. But once it exists, every subsequent AI workflow runs faster, more accurately, and with less integration work.
This is also the mechanism by which platform intelligence compounds. Every interaction updates the risk record. Every update makes the next decision more accurate. The insurer is not just automating workflows; they are building an institutional memory that gets better with every claim processed, every renewal converted, every pre-auth decided.
The Controls Layer That Regulators Require
Any honest discussion of agentic AI in insurance has to address the question of controls. Regulators in every major insurance market are paying close attention to AI deployment, and their requirements are specific.
The FCA, in its guidance on AI in financial services, requires that automated decisions be explainable, not just accurate, but traceable to specific inputs and rules. The EU AI Act classifies insurance AI systems as high-risk, requiring conformity assessments, human oversight mechanisms, and documentation of training data. IRDAI requires audit trails for automated claims decisions. Saudi IA requires that AI systems used in underwriting and claims be capable of producing decision rationale on demand.
Production agentic AI systems in insurance must therefore be built with the following controls as standard requirements, not post-deployment additions:
Audit trail: Every decision, the inputs considered, the rules applied, the confidence score assigned, and the output produced are logged in an append-only audit trail. Five-year retention minimum. Regulator export on demand.
Explainability: Plain-language reason codes for every customer-facing decision. Not just "claim declined" but "claim declined: exclusion clause 4.2 applies to damage occurring outside the territorial limit defined in schedule 1." SHAP-based attribution for model-driven decisions.
Human override: Every agentic decision is reviewable and overridable by a human reviewer. The override is captured in the audit trail with the reviewer's identity and the reason for the override.
Confidence thresholds: The system escalates to human review when its confidence falls below a configurable threshold. The threshold is set by the insurer's operations team, not by the technology vendor.
Data residency: Customer data does not leave the jurisdiction in which it was collected. This is a non-negotiable requirement in India, Saudi Arabia, the UAE, and the UK, and it must be an architectural requirement rather than a deployment option.
The Sequencing Question: Where to Start
For insurers evaluating agentic AI deployment, the sequencing question, where to start, is as important as the technology question.
The instinct of most transformation programmes is to start with a strategic initiative: a full claims automation programme, an enterprise AI platform, a multi-year digital transformation. This instinct is wrong. Not because the strategic vision is misguided, but because it creates exactly the conditions in which AI projects fail: long timelines, complex integrations, high organisational risk, and a gap of 12 to 18 months between commitment and first production result.
The approach that works is the opposite: start with the workflow that is costing the most, scope it tightly, and get it live in production within 8-12 weeks. Not a pilot. Not a proof of concept. Production, with real data, real systems, and real decisions.
The choice of the first workflow matters. Motor FNOL is a strong starting point for general insurers because it is high volume, well-defined, and the production metrics, handling time, cost per claim, and accuracy are immediately measurable. Health pre-auth or Health Claims is the right starting point for health insurers because it affects both operational efficiency and member experience simultaneously. CX audit is the right starting point for insurers whose first priority is understanding where their operational failures are before automating anything.
What matters most is not which workflow you start with, but that you start with one, fully scoped, fully integrated, fully in production, before expanding to the next.
The second workflow is always faster than the first. The pre-built components from the first deployment, the coverage validation logic, the document extraction pipeline, and the confidence scoring framework are reusable. The integration work with core systems, which is always the longest part of the first deployment, is already done. The operations team knows how to work with the system. The compliance team has seen the audit trail format and approved it.
This is the mechanism by which agentic AI compounds in insurance. Not through a single large deployment, but through a series of focused deployments that each build on the last.
What to Look for in a Production-Ready Platform
For insurance technology leaders evaluating agentic AI vendors, the questions that distinguish production-ready platforms from pilot-ready tools are specific:
On domain knowledge: Can the system parse a Table of Benefits without configuration? What is the field accuracy on insurance-specific documents compared to generic LLMs? Has the vendor worked with actual insurance documents from your lines of business?
On integration: What is the integration path with your core policy and claims system? Does the vendor have pre-built connectors, or is integration a custom project? How long has it taken previous clients to go from contract signature to live integration?
On compliance: What regulators does the platform support natively? Can the system produce the audit trail format required by your regulator? Has the platform been deployed in your jurisdiction before?
On controls: What are the confidence threshold configurations available? How are overrides captured? Can the compliance team access the audit trail without IT involvement?
On architecture: Does the platform support a unified customer risk record across workflows? What happens to the data from the first deployment when you add a second workflow? Is it reused, or does integration start again from scratch?
The answers to these questions will distinguish vendors who have built for insurance production from those who have built for insurance pilots.
Conclusion
Agentic AI in insurance is past the point of being a future capability. It is in production today, in motor claims, motor underwriting, health pre-authorisation, health claims, health underwriting, customer operations, renewal, and cross-sell at insurers across Asia, the US, the Middle East, and the UK.
The gap between the insurers who have reached production and those still running pilots is not a technology gap. It is an architecture gap, an integration gap, and a domain knowledge gap. The platforms and vendors who bridge those gaps with insurance-native models, pre-built integrations, and regulatory compliance as a design requirement are the ones whose deployments are live.
The insurers who close this gap in the next 12-18 months will have a compounding advantage: operational intelligence that gets better with every workflow deployed, every claim processed, and every customer interaction recorded. Those who wait will find the gap harder to close, not easier.
The window for first-mover advantage in insurance AI is not closed. But it is narrowing.
Yukthi Labs deploys agentic AI across the insurance value chain from motor FNOL and health pre-auth to customer operations, underwriting triage, and renewal. Our deployments are live in production across India, the Gulf, the UK and Southeast Asia. If you are evaluating where to start, we offer a half-day working session to map your highest-value workflow and scope a production deployment.
About Author: Bala Chandrashekharan has more than 30 years of experience in insurance, working with companies like JLT and Marsh in leadership positions.