AI Liability Insurance Buyer's Guide
Underwritten Essay

The 21% Problem: Why Uninsurable Agents Fail Underwriting

The five controls that separate an insurable AI agent deployment from an uninsurable one.

Written by

Joel R. Singh

Section

Underwritten

Published

2026-08-23

A NASA flight director at a Mission Control console in 1965, headset on, watching banks of monitors, one hand ready to halt the mission the instant something goes wrong
Mission Control, 1965. The human-in-the-loop kill switch, decades before the phrase existed. NASA, public domain.

Twenty-one percent of enterprises cannot stop a runaway AI agent's spending in real time. In a survey of 107 software and ML engineers, product managers, and data and AI executives, that group reported they rely on reactive monitoring only. They watch the logs after the fact and have no live way to halt an agent mid-incident.[1] When the same respondents were asked how they control agent spend, the breakdown was telling: roughly 30% use native platform caps or throttling, about a quarter use custom gateway middleware, another quarter route dynamically to cheaper models, and that last 21% have no live stop at all.

Here is why that number belongs in a coverage guide. You cannot insure what you cannot control. An underwriter pricing AI liability is buying into your loss potential, and an agent with no live kill switch is an open-ended loss. The 21% who cannot stop a runaway agent are the same enterprises walking into higher premiums, tighter sub-limits, broad exclusions, and, when a claim lands, a denial that turns on whether the promised controls were actually in place.

You cannot insure what you cannot control.

The insurance market is already pricing on exactly this. Armilla begins its engagement with a risk assessment before it binds a standalone AI liability policy.[2] AIUC-1 certification runs a model through more than 5,000 adversarial tests.[3] Underwriters are pricing on governance evidence now, which means model inventories, documented human oversight, incident response procedures, and demonstrable controls decide whether you get workable terms or a declined quote.

Five controls are becoming the underwriting conditions that separate an insurable agent deployment from an uninsurable one. Read each as a question an underwriter will ask, because increasingly they do. If you cannot answer yes with evidence, you have found a coverage gap before it finds you.

1. Human gates on irreversible actions


The first condition is a decision about which actions an agent may take on its own and which require a human first. Current governance practice tiers actions by reversibility and blast radius. Reversible reads auto-approve. Recoverable actions such as sending an email get a soft gate or a notification. Irreversible actions, deleting data, transferring money, triggering an external system, or deploying, get a hard gate that blocks until a person approves.[4]

Underwriting question

For each agent you run, is there a written map of action types to tiers, and does a money-moving or data-deleting action actually pause for human approval when you test it?

A policy tied to the EU AI Act's human-oversight requirement will look for exactly this, and that obligation applies from 2 August 2027 for high-risk systems.[5] If irreversible actions execute without a checkpoint, an insurer sees uncapped severity, and a claims adjuster later sees a missing control.

2. Real-time spend caps with a live kill


This is the control the 21% are missing, and it is the clearest line between an insurable risk and an open one. A hard budget ceiling that triggers a real-time stop prevents the loss. A log you review the next morning only documents it. From an underwriting seat, the first is a bounded exposure and the second is a blank check.

Underwriting question

Can you halt an agent's spend inside a single billing interval, on demand, right now?

A credible answer looks like a budget ceiling wired into your gateway or platform config plus a demonstrated stop, not a dashboard you check after the invoice arrives. If your only spend control is a report, you are in the reactive-only group, and that is the group facing the widest coverage gap.

3. Least-privilege identity, one scope per agent


Agents are non-human identities, and underwriters increasingly expect them to carry the same discipline you apply to service accounts. The failure mode is a broad, persistent credential shared across several agents, a key that can do far more than any one task requires. Governance practice calls for a unique service account or workload identity per agent, scoped to the minimum permissions that agent needs, provisioned before it reaches production.[6]

Underwriting question

Does each agent run under its own identity with an enumerated permission scope, or do several agents share one broad key?

If a single credential unlocks your database, your payment rails, and your email, one confused agent has the run of all three, and a carrier prices that aggregation of risk accordingly.

4. Decision traceability on demand


When an agent does something you did not expect, you need to reconstruct what it decided and why, and you need it without a forensics project. Traceability means full action logs plus decision lineage available on demand, not stitched together after an incident from scattered fragments. For a claim, this is the evidence file. The insured who can produce a clean decision trail gets paid faster and argued with less.

Underwriting question

If an agent took a costly or wrong action yesterday, can you pull its full decision trail today in minutes?

If the answer requires an engineer to reassemble logs by hand, you do not have traceability, and at claim time you have a dispute waiting to happen.

5. Narrow-scope, single-responsibility agents


A robotic arm rehearses an autonomous satellite-servicing approach in a government test rig
GAO test imagery. Wide autonomy means a wide blast radius, and blast radius is the severity an underwriter prices. Public domain.

The last condition is about ambition. A general-purpose agent with a wide mandate carries a correspondingly wide blast radius, and blast radius is severity, which is what an underwriter prices. The pattern among the teams getting durable value is single-responsibility agents with tightly bounded mandates, each doing one job well, so that when one misbehaves the damage stays in its lane.

Underwriting question

Is any agent in your fleet doing three jobs where three narrow agents would be safer?

Breadth of mandate and breadth of insurable risk move together.

Why this window matters for coverage


The reason to close these gaps now, rather than after your own 2 a.m. incident, is that both the risk and the underwriting bar are rising at once. Gartner projects that over 40% of agentic AI projects will be canceled by the end of 2027, driven by escalating costs, unclear business value, or inadequate risk controls.[7] Security and risk are cited by nearly two-thirds of enterprises as the top obstacle to scaling agents, and responsible-AI maturity averages 2.3 out of 4, with only about 30% of organizations reaching level 3 or higher on overall agentic-governance readiness.[8]

The rising bar

Over 40% of agentic AI projects are projected to be canceled by the end of 2027 (Gartner). Nearly two-thirds of enterprises name security and risk as the top obstacle to scaling agents. Responsible-AI maturity averages 2.3 out of 4, and only about 30% of organizations reach level 3 or higher on overall agentic-governance readiness.

The carriers writing dedicated AI liability policies are drawing their lines against that backdrop. Assessment-before-binding is becoming standard, so the controls above are shifting from good practice to entry price. Build them now and you are building the evidence file underwriters ask for, instead of scrambling to assemble it under a renewal deadline or, worse, arguing about it during a claim.

Where iSL fits


We run these five controls in our own operations before we recommend them to anyone. That is why an iSL assessment measures your agents against real evidence, not a checklist read off a slide.

If you worked through the five questions and hit a control you could not answer with proof, two moves work together:

  1. Get insurable. The iSL AI Advisory assessment maps your agents to their action tiers, tests whether the gates actually hold, and hands you a prioritized fix list plus the governance file carriers ask for at quote.
  2. Shop coverage with that file in hand. You walk into the broker conversation with evidence instead of promises, and evidence is what earns workable terms.

The 21% problem is a controls gap rather than a technology limit, and control gaps decide whether your agent runs safely and if it's even insurable at all.


Joel Singh is the founder of iSinghLabs and author of the AI Coverage Guide. This page is general information about the AI insurance market, not insurance, legal, or financial advice. Coverage terms vary by carrier, state, and policy. Consult a licensed insurance broker about your specific situation.

Works Cited

Every factual claim in this essay is sourced below, with primary sources preferred. The numbers match the citation after each sentence. Where a source compresses a more layered reality, the note says so.

  1. 1VentureBeat, One in five enterprises can't stop a runaway AI agent's spending in real time. Directional: figures come from a proprietary poll of 107 software and ML engineers, product managers, and data and AI executives, not a representative sample.
  2. 2Armilla AI, AI Insurance.
  3. 3AIUC, AIUC-1 Standard. AIUC-1 is a certification standard, not an underwriter's binding gate.
  4. 4StackAI, Human-in-the-Loop AI Agents: How to Design Approval Workflows for Safe and Scalable Automation.
  5. 5Regulation (EU) 2024/1689 (EU AI Act), Article 113 (Entry into force and application). High-risk obligations, including human oversight, apply from 2 August 2027; a 2026 Omnibus amendment may adjust these dates.
  6. 6Agent Security Review, Enterprise AI Agent Governance Frameworks.
  7. 7Gartner, Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027.
  8. 8McKinsey, The State of AI Trust in 2026: Shifting to the Agentic Era. Both the near-two-thirds security/risk figure and the 2.3-out-of-4 maturity average are from McKinsey's 2026 AI Trust Maturity Survey (~500 organizations); "about 30%" reflects the roughly one-third reporting maturity level 3 or higher on governance.
HOME