Let's talk!

Why Insurers Stuck in AI Pilots Are Losing

ai digitalization emerging technologies Jul 31, 2026
 

Written by Alchemy Crew Ventures Team

 

The insurers still running AI pilots aren't behind on models. They're behind on trust.

That's the argument Willem Paling, Executive Manager of AI and Analytics at IAG, makes about why so much of the industry is stuck. The winners won't be the ones with the most impressive experimental models. They'll be the ones who built the architecture that lets AI actually scale: strong data quality, real governance, explainability, reasoning traces, and human override pathways.

Optimize for the pilot, and you get a pilot. Optimize for trust, and you get agentic AI you can put into production.

For insurance leaders, C-suite executives, and AI, analytics, and innovation teams across insurers, reinsurers, brokers, and financial services firms, the message is practical: the AI race is no longer won by the best model. The insurer wins it with the strongest trust architecture, data foundations, and governance discipline. Under his leadership, IAG launched more AI models in the past two years than in the previous six combined by killing the pilot phase and shipping AI products into production with the same delivery discipline as any software project.

Three things to take from this article: 

Key Takeaways
  • IAG stopped running experiments and started running production systems, launching more AI models in two years than in the previous six combined, because it applied software delivery discipline, monthly governance, and clear business ownership to every model.
  • Trust architecture, not model access, is the real differentiator on the agentic frontier. Willem Paling and the AI 2030 Horizons authors argue that monitoring, explainability, reasoning traces, and human override pathways are what turn an impressive model into a system a regulator and a customer can rely on.
  • The article maps the full shift: how insurers move from pilots to production, how orchestrated AI workflows and human oversight change claims and operations, what regulators now expect from AI governance, why insurance products must become machine-readable, and how carriers can avoid disappearing from AI-driven customer journeys entirely.

 

The Insurer That Stopped Experimenting

The most radical thing happening in insurance AI right now is not a model. It is a decision to stop piloting. For a decade, the industry funded proofs of concept, innovation labs, and sandboxes that produced slideware but never touched the P&L. IAG made a different choice. It set up the factory, embedded data scientists into the same planning process as software engineers, and shipped.

That is the shift Sabine VanderLinden unpacks on the episode of Scouting for Growth with Willem Paling, Executive Manager of AI and Analytics at IAG, Australia's largest general insurer and the anchor case study in the AI 2030 Horizons perspective produced by Alchemy Crew Ventures alongside IAG, SAS, Assurant, and Majesco. IAG now runs production systems including CASI, its AI claims assistant, and a 150-strong network of generative AI activators embedded across the business. The question this conversation explores is deceptively simple. What separates an insurer that ships AI at scale from one still stuck in the sandbox, and why is trust, not model capability, the thing that decides the winner?

Why This Matters Now

The window to build the foundations is closing while the technology is still stabilizing. Generative AI could unlock $50-$70 billion in revenue for the insurance industry, according to McKinsey, and users who effectively leverage it can increase productivity by more than 20%. AI adoption in insurance is projected to rise from 14% to 70% by 2028, shaping the future competitive baseline. Yet McKinsey also finds that a full end-to-end transformation of the claims domain can deliver up to 14 times the impact of isolated use cases. The value is not in the demo. It is in the redesign.

Regulators are moving in parallel. In July 2025, the International Association of Insurance Supervisors published its Application Paper on the supervision of AI, signaling that governance, explainability, and fairness expectations now apply directly to AI systems across the insurance value chain. Insurers cannot treat trust as a later-stage compliance gate.

The competitive landscape is consolidating fast. Allianz and Anthropic announced a global partnership in January 2026 focused on agentic automation and traceable, compliant AI. Travelers launched an agentic AI claim assistant built with OpenAI in February 2026 and expanded it countrywide within two months. Leading AI insurers are already pulling ahead, with leaders achieving 6.1 times higher shareholder returns. The insurers redesigning around AI are compounding advantage now, not in 2030.

The Key Insights: Leveraging Industry Knowledge

The pilot is the problem, not the proof

Skip the pilot and ship the product. Willem Paling's core argument is that the leap from experiment to operation is organizational, not technical. When AI is an experiment, an innovation team owns the idea. When it is a product, the business owns the outcome, and moving beyond isolated pilots requires repeatable deployment processes aligned to business needs and existing processes, changing workflows, controls, roles, and accountability all at once.

IAG closed the gap by treating models like software. Clear purpose, clear business value, clear path to deployment, and governance that runs monthly rather than once a year. Successful implementations also depend on change management, the human element, and upskilling so teams can support adoption.

"We stopped doing experiments, and we focused on delivery. We put the same delivery expectations onto our data scientists and machine learning engineers that we have for software engineers."

This is the Frontier Firm thesis in practice. A human-led, agent-operated organization is not built by a lab on the side. It is built when AI delivery becomes as routine and accountable as any other line of business, and when governance runs like clockwork because it runs every month, an operating model whose development can help foster a culture of innovation inside insurance organizations.

Trust architecture is what turns a model into a system

The model is a small part. Everything around it creates the value. Paling makes the point with a sharp example. Put an insurance coverage decision into a frontier model ten times, and you can get ten different decision processes behind the scenes. That is not a system a customer can trust. Pilot projects in insurance also often fail when structured data is weak and integration into core systems is complex. In an industry where decisions affect claims, pricing, and people recovering from loss, consistency is the product because it underpins reliable decision-making.

According to Willem Paling, trust architecture isn't separate from value creation. Trust is what turns AI from an impressive model into something that improves insurance at scale."

This is the Intelligent Layers pillar made concrete. Data quality, workflow fit, monitoring, explanation, reasoning traces, and defined human oversight intervention points are the architecture. Trust architecture also has to cover regulatory requirements, legal exposure from bias or misrepresentation in AI interactions, and ongoing maintenance for compliant performance. Get it wrong with one person, and you have a local problem. Get it wrong with AI, and you create noise, inconsistency, and risk at scale. The insurers that build this rigor in can hand more of the process to AI sooner, while maintaining accuracy and safety.

The messy middle is where the competitive advantage lives

Automate the grind, not the judgment. The highest-reliability value sits in the messy middle, the friction zone between claims lodgement and final judgment, where skilled people spend hours reconciling PDFs, medical packets, and engineering reports before they can help a customer. Here, natural language processing helps turn large volumes of unstructured documents into structured data, and many systems correctly extract about 70% of claims documents before review. AI does the heavy lifting in that layer through an AI-driven workflow, not a single model. Specialized agents extract, check, enrich, summarise, and evaluate, and the case arrives decision-ready with its reasoning attached.

Willem Paling is deliberate about restraint. The more material the decision, the lower the agent's autonomy, and the stronger the oversight, explanation, and override pathways must be. Autonomy is earned in layers, not declared. Enabling underwriters with this kind of orchestration can raise quote capacity by 40% while keeping final judgment with underwriters.

"A lot of the best AI in insurance is focused on the mundane. It's reducing friction in the parts of the process customers probably don't even want to know about." This is the counterintuitive heart of the agentic frontier, and it is why Sabine VanderLinden frames the winners as the ones who earn autonomy instead of assuming it. Compressed cycle time, higher consistency, and better operational efficiency across operations are the payoff, with AI acting as a coordination tool that frees skilled people for negotiation, empathy, and exception handling.

Generative AI and machine readability are the new distribution war

Prepare to be legible to machines, not just people. Paling describes a two-speed disruption. AI is already becoming a discovery and recommendation layer, and full transactional capability will follow as payment rails and trust standards mature. OpenAI already hosts around a dozen approved insurer-built apps inside ChatGPT, and Visa and Mastercard are building tokenization infrastructure for agentic commerce.

"The most underestimated risk is AI on the other side, AI attacking the evidence layer of insurance."

The strategic response splits into two ways. First, defense. As automation rises, assurance has to rise with it through stronger identity and provenance checks at the point of capture. Second, visibility. Products must be discoverable, quotable through APIs, and bindable in milliseconds. This is where the Venture Client Model earns its place. Rather than buying every capability or making equity bets, insurers can identify external solution providers faster and solve distribution and service challenges without building every capability internally. Using DIVAAA™ as a structured path for project management and cross-team adoption can give companies an efficient competitive advantage, with tangible benefits that support long-term success.

Actionable Takeaways for Leaders

Four moves that convert AI from a demo into a compounding advantage this quarter.

  1. Kill your next pilot. Redefine it as a repeatable deployment tied to clear business needs, with a business owner, a path into existing processes, and monthly governance, and hold it to software delivery standards from day one.

  2. Build trust architecture before you scale autonomy. Instrument every meaningful step with model and prompt versioning, reasoning traces, drift detection, defined human oversight, and controls that meet regulatory requirements.

  3. Attack the messy middle first. Target document-heavy, semi-structured work where natural language processing turns submissions into structured data with orchestrated agents on bounded autonomy, and free your experts for judgment and customer conversations.

  4. Make your products legible to machines. Audit whether your coverage logic, quoting, and binding capabilities are discoverable, API-ready, and transmissible by AI agents, because machine readability is a competitive advantage for the future.

FAQ

How is IAG scaling AI in insurance? IAG applied software delivery discipline to AI, skipping the pilot phase and shipping production products with monthly governance, clear business ownership, reliability monitoring, repeatable deployment processes, and change management. It launched more AI models in two years than in the previous six combined.

What is trust architecture in AI for insurance? Trust architecture is the operating environment around a model: data quality, monitoring, explainability, reasoning traces, human override pathways, human oversight, and regulatory requirements. Willem Paling argues that it is not model capability that makes AI reliable enough to deploy in a regulated industry.

What is the messy middle in insurance AI? The messy middle is the friction zone between claims intake and final decision where staff reconciles PDFs, medical documents, and reports. It is where orchestrated AI agents deliver the most reliable, scalable value, with natural language processing and specialized agents helping turn unstructured inputs into structured data.

Will AI agents replace insurance distribution? Not overnight. Paling describes a two-speed shift in which AI first becomes a discovery and recommendation layer, then a transactional one, creating a competitive advantage for carriers that adapt early. Insurers whose products are not machine-readable risk becoming invisible in AI-mediated buying journeys.

What did IAG contribute to the AI 2030 Horizons perspective? IAG served as the anchor case study, developed with Alchemy Crew Ventures, SAS, Assurant, and Majesco after a 50-leader executive summit at ITC 2025, pressure-testing what AI at scale really requires. 

The Race Is for Trust, Not Models

The fight ahead is not about who has the biggest model. Everyone will have model access. The fight is about who earns autonomy, who builds the trust architecture before switching on the digital labor, and who makes their products machine-readable before the market catches up. Willem Paling's bet is clear. Stop running experiments, start running pipes. Generative AI tools alone are not enough; insurers need agentic systems, backed by industry knowledge and expertise, to scale reliably. AI must be treated as infrastructure for decision-making and operational efficiency, not as an isolated tool. Make trust the product, and design for a world where your policy is either legible to machines or invisible to them.

So here is the question for your leadership team. Are you building AI demos, or are you building successful implementations that match business needs and create measurable benefits year after year?

Listen to the full conversation with Willem Paling on Scouting for Growth on Apple Podcasts, Spotify, and Podbean, and explore the Frontier Firm research program at alchemycrew.ventures/amplifying-success.

Sources and Citations 

McKinsey and Company, The potential of generative AI in insurance

McKinsey and Company, Gen AI could unlock 50 to 70 billion dollars in insurance revenue

International Association of Insurance Supervisors, Application Paper on the supervision of artificial intelligence, July 2025

Allianz, Allianz and Anthropic forge global partnership to advance responsible AI in insurance, January 2026

OpenAI, Travelers deploys AI-powered claims countrywide with OpenAI, February 2026

 

 

We have been featured in many mainstream and FutureTech publications. Learn more here.

Let's talk!

[email protected]

 

Join our programs

Activate Your Authentic Identity
Unlock $1M Funding in 90 Days

Scouting for Growth

 

Listen to our podcasts

Scouting for Growth
Beyond Tech Frontiers