Skip to Content
Risk Management

Making AI Use in Claims Defensible

Michelle Perry | September 11, 2026

On This Page
the word "REVIEW" in white block letters with scattered AI icons overlaid on an image of a woman typing on a laptop

Artificial intelligence (AI) has quickly moved into claims operations: In a relatively short period of time, AI-enabled tools have become capable of performing work that previously required a significant amount of manual claims handling.

An AI tool can summarize hundreds of pages of claim documentation, organize a loss chronology, identify potentially relevant policy provisions, flag inconsistencies in the reported facts, draft a claim note, and prepare a proposed communication for a claims professional to review. Assume AI does all of that on one claim: The claims professional reads the output, makes a few edits, and approves the next action or sends the communication for final approval.

Who made the decision on the claim? The answer may not be as obvious as it sounds; the claims professional was involved. There was technically a human "in the loop." But that does not tell us whether the person independently evaluated the claim, whether the AI established a conclusion that the claims professional later ratified, or whether the tool simply organized information that the professional independently analyzed. Those are different operating models, and from a claims governance standpoint, the difference matters.

As AI becomes embedded more deeply into claims administration, fraud detection, document analysis, valuation, communications, and other workflows, claims organizations need to define what I call the boundary of human decision—the point beyond which AI may assist, inform, organize, or recommend, but responsibility for the substantive claim decision remains with a human claims professional. A human in the workflow is not the same thing as accountable human judgment. A defensible claims operation has to define that boundary, preserve evidence that the human crossed it appropriately, and test whether that is actually occurring.

An internal AI policy that simply says "human oversight is required" does not accomplish that; the operational workflow has to do it.

The Regulatory Obligation Is Broad. The Claims Control Cannot Be.

The regulatory framework around insurance AI is developing quickly, but the underlying obligation is not particularly new. The "Model Bulletin: Use of Artificial Intelligence Systems by Insurers" written by the National Association of Insurance Commissioners (NAIC) makes clear that consumer-impacting decisions made or supported by AI remain subject to existing insurance laws, including unfair trade practice and claims settlement requirements. It also places governance expectations around claims administration and payment, fraud detection, third-party systems, testing, documentation, and the degree of human participation in final decision-making.

Mississippi's July 22, 2026, "Bulletin 2026-9: Use of Artificial Intelligence Systems by Insurers" provides a current state example. It similarly identifies human involvement in final decision-making as a factor in determining the appropriate level of control and contemplates regulatory review of automation constraints, model inventories, validation, bias controls, monitoring, and third-party oversight.

Texas makes the human review expectation more explicit. In "Commissioner's Bulletin B-0003-26," issued June 12, 2026, the Texas Department of Insurance states that, when a regulated entity uses AI to make a consequential decision, a person is expected to review and agree with the decision before action is taken. The bulletin also ties AI use directly to existing Texas claims-processing requirements, adjuster licensing obligations, market conduct surveillance, and examination authority. 1

Requiring a person to review a decision establishes a human checkpoint. It does not, by itself, establish what a meaningful review looks like inside the claim file. A claims organization still has to define what the person is expected to evaluate independently, what authority the AI was permitted to exercise, and what evidence demonstrates that the professional judgment required for the decision actually occurred.

The significance for claims organizations is not that regulators have invented an entirely new claims-handling obligation for AI; they largely have not. Existing obligations are increasingly being translated into expectations about how the technology is governed, where the human responsibility remains, and how the organization proves that controls actually operated.

Colorado has gone a step further in its broader automated decision-making law. Signed May 14, 2026, and effective January 1, 2027, SB26-189 establishes the requirements for covered automated decision-making technology used to materially influence consequential decisions, expressly including insurance. Among other provisions, the law gives consumers a right to request meaningful human review and reconsideration following certain adverse outcomes. 2 At the time of writing, the Colorado attorney general has indicated that the state will not enforce the law until regulations are finalized.

Nonadmitted/Surplus Lines Status Doesn't Equal Unregulated

One issue that receives less attention in the current AI discussion is how these expectations reach surplus lines claims operations. Many state insurance AI bulletins are directed to insurers holding a certificate of authority in the issuing state. A nonadmitted insurer, by definition, does not hold that certificate in the state where it is writing surplus lines business, which does not mean that the regulatory obligation disappears.

A surplus lines insurer may not be the direct addressee of a particular bulletin, which does not end the analysis. The insurer remains regulated in its state of domicile, where it holds a certificate of authority and AI governance expectations may apply directly. It also remains subject to the underlying insurance laws that govern the claims activity itself.

The surplus line exemption is narrower than the word "nonadmitted" sometimes suggests. It generally changes the regulatory treatment of areas such as rates and forms, but it does not automatically remove the claims-handling standards governing how losses are investigated, evaluated, communicated, and resolved.

This matters because emerging AI bulletins generally do not establish an entirely separate standard for claims handling; they connect AI use back to existing obligations involving unfair claims practices, unfair discrimination, documentation, governance, and accountability. Whether a particular state's unfair claims settlement requirements apply to surplus lines business must still be evaluated on a jurisdiction-by-jurisdiction basis, and that analysis cannot be skipped. Where the underlying claims obligation applies, using AI to perform or support the activity does not eliminate the obligation.

Delegated claims models add another layer: When an insurer delegates claims handling to a third-party administrator (TPA) or other claims administrator, that administrator may also be the third party whose use of AI the insurer is expected to understand and oversee. If AI influences claims administration through the delegated operation, the insurer may need to be able to explain what the system does, how its use is controlled, what oversight exists, and how responsibility for the ultimate claim decision remains with an accountable human decision-maker.

A surplus lines claims organization cannot practically wait for every state to issue AI guidance specific to nonadmitted insurers before establishing its governance framework. The more defensible approach is to identify where AI enters the claims process, define the human decision boundary, document the applicable controls, and evaluate state-specific requirements as part of the organization's normal regulatory analysis.

Nonadmitted status may change how an AI requirement reaches the organization, but it does not remove the need to understand and govern how AI influences the claim decision; states do not all have identical requirements, and not every law requires a named human to personally make every insurance decision, which is important to keep in mind. The broader operational issue is whether the organization can identify and define where human judgment is required and demonstrate that it occurred.

The Human-in-the-Loop Problem

"Human-in-the-loop" has become one of the most common phrases in AI governance and is also one of the easiest controls to overstate. A human can technically participate in an operational workflow without meaningfully exercising judgment. Consider the following three examples.

  • A claims professional receives an AI-generated coverage analysis, reads the recommendation, and routinely accepts it without returning to the policy or underlying facts to review or verify.
  • A fraud model assigns a high-risk score, and the claim is escalated because that is what the workflow automatically does, even though nobody can explain which factors materially produced the score; the thought process behind the escalation for that particular claim is not articulated.
  • An AI tool drafts a settlement analysis, and the final claim note repeats substantially the same reasoning without identifying what the claims professional independently evaluated.

There is a human involved in all three workflows, which does not necessarily mean the human made the decision; an approval action indicates that someone reviewed or accepted an output. Standing alone, it does not tell us whether the person independently evaluated the information underlying it or exercised the judgment required by the claim decision. The organization remains responsible for the regulated outcome, and that responsibility does not transfer to the model, the developer, or the vendor merely because technology influenced the process. This is where the boundary of human decision becomes an operating control rather than an AI principle.

Define the Decision Boundary Before the Tool Is Used

Claims organizations should determine an AI system's permitted role before putting the tool into production. It is important to note that this role will not be identical for every system, and AI may appropriately assist with activities such as summarization, organization, document classification, routine drafting, information retrieval, pattern recognition, or other defined functions. Some low-risk administrative activities may also be capable of operating within predetermined parameters.

The risk changes when the AI output begins influencing a substantive decision affecting coverage, payment, settlement, claim disposition, or another consumer-impacting outcome. For a consumer-impacting decision, the operating model should make several things clear, including the following.

  • What is the AI authorized to do, and what is it not authorized to decide?
  • At what point is independent human judgment required?
  • What information must the claims professional independently review?
  • What circumstances require supervisory or compliance escalation?
  • Who is accountable for the final claim decision?

The answers should not exist only in an enterprise AI policy, and they need to be reflected in the processes and procedures, training, system permissions, supervisory expectations, quality assurance (QA) review criteria, and documentation requirements governing the actual claim. A control that exists only in policy is particularly vulnerable when the operational workflow makes ignoring it easier than following it.

Make the Human Decision Reconstructable

Defining the boundary is only half of the control—the organization also needs evidence that the boundary operated on the claim. An examiner, quality reviewer, supervisor, or subsequent claims professional should be able to reconstruct enough of an AI-assisted decision to understand the following four things.

  • What information was considered.
  • What AI contributed to the process.
  • What the claims professional decided.
  • Why the claims professional independently reached that conclusion.

The fourth element is where the distinction between review and judgment becomes visible. Consider an AI tool that summarizes a lengthy property loss file and identifies several facts relevant to a potential coverage limitation where the claims professional reviews the source documents, confirms some of the facts, identifies another fact the AI did not give appropriate weight to, reviews the applicable policy language, and reaches the final coverage conclusion. The claim file should make the professional's analysis understandable and should not merely preserve an AI-generated paragraph and add the words "reviewed and approved."

Documentation by itself does not create accountability; its value is that it allows the organization to demonstrate what actually occurred, and the purpose is not to create another documentation exercise simply because AI was used. Therefore, the record needs to preserve enough information to distinguish the technology's contribution from the professional judgment that produced the claim outcome.

Govern the Vendor and the Change, Not Just the Employee

Some of the most significant AI exposure in claims may not come from an adjuster opening a public generative AI application; it may come from technology already embedded in the claims ecosystem. A valuation platform may incorporate machine learning, a document management system may add automated classification, a fraud vendor may change its scoring logic, a claims platform may introduce an AI assistant through a routine software update, and the claims professional may still see the same screen while the underlying decision influence may have changed considerably.

The NAIC model bulletin specifically expects insurer AI governance to encompass systems developed by third parties and contemplates due diligence, contractual audit rights, regulatory cooperation, and oversight of third-party compliance. Mississippi's 2026 bulletin similarly provides that insurers should expect regulators to request third-party due diligence, contracts, audit or confirmation processes, and validation and testing information. Colorado's 2026 law adds another useful governance principle: Developers of covered automated decision-making technology are required to provide notice of material updates or modifications.

That creates a practical claims problem: If a vendor changes the model, data inputs, or decision logic without telling the insurer or claims administrator, the organization may be operating under a governance approval that no longer reflects what the tool actually does.

Vendor governance, therefore, needs to extend beyond cyber security and procurement. Depending on the use case, it may also need to address the approved function of the AI, data inputs, available audit information, testing, model limitations, material change notification, regulatory cooperation, and circumstances that require the tool to be reevaluated. The vendor may own the technology; however, the insurer still owns the regulatory burden, the risk, and the defense of the outcome.

Test Whether Human Judgment Is Actually Occurring

Training is necessary, but having signed training completion records does not establish that the control is working, nor does a signed AI policy acknowledgment or approved-tool inventory. The organization eventually has to look at the claims.

QA is one of the most practical places to determine whether the boundary of human decision-making is functioning because quality review already examines how claims professionals investigate, document, communicate, interpret coverage, and support claim outcomes. AI governance does not need an entirely separate testing structure if the existing QA process can be adapted to evaluate it.

For an AI-assisted claim, that may mean reviewing whether the tool was authorized for the particular use, whether material facts and policy language were independently verified, whether human analysis can be distinguished from AI-generated content, whether the decision is supported by the underlying record, and whether required escalation occurred.

Aggregate patterns also provide useful information. A tool with a consistently high override rate may be poorly calibrated for the claims it supports, but a tool that is almost never overridden may raise a different question: Are claims professionals continuing to independently evaluate its recommendations? Neither result proves a control failure on its own.

An appropriate override may demonstrate that the control worked exactly as intended. Likewise, a high acceptance rate may reflect strong model performance or declining independent review. The purpose of the metric is to identify where further review is warranted, not to predetermine the answer.

There is another control issue that deserves attention: the expertise of the claims professional. Human oversight is only meaningful if the person performing that oversight has the knowledge necessary to recognize when an AI output is wrong, incomplete, or inconsistent with the actual claim record. If increasingly automated workflows cause claims professionals to lose the subject matter expertise necessary to detect and challenge a plausible but incorrect output, the organization weakens one of its own validation mechanisms.

The AI Risk Management Framework created by the National Institute of Standards and Technology (NIST) recognizes the importance of clearly defining human roles and responsibilities in AI decision-making and oversight. In claims, the appropriate human-AI configuration should reflect the consequence of the decision being made. The greater the potential effect on the claimant, the more important competent and identifiable human judgment becomes.

What a Defensible Claims AI Control Environment Looks Like

There will not be one universal AI governance structure that works for every insurer, TPA, managing general agent, or claims organization. The operating model, lines of business, regulatory footprint, delegated authority structure, technology stack, and use cases will differ. However, the components of a defensible control environment are becoming easier to identify, and a defensible claims AI framework should establish at least five things.

  • Know where AI is being used. The inventory should include internally deployed tools and AI capabilities embedded in vendor platforms. The relevant question is not simply what technology the organization owns, but where that technology can influence a claim.
  • Classify the decision impact. A drafting tool compared to a system that materially influences coverage, payment, fraud escalation, or settlement does not present equivalent risk and should not be governed as though it does.
  • Define and document the human decision boundary. The organization should establish where AI assistance ends, where accountable professional judgment begins, and what evidence must remain in the claim record.
  • Govern change and third parties. AI-enabled vendor services should be subject to appropriate due diligence, monitoring, change notification, and reevaluation when the system or its role materially changes.
  • Test the control against actual claims. QA, complaints, corrective action, unusual acceptance or override patterns, and identified AI incidents should be used to determine whether the governance model is functioning in practice.

The NAIC's current regulatory evaluation work reinforces this movement toward operating evidence. Its AI Risk Evaluation Supplement asks regulators to examine written programs, risk assessment, training, third-party vendors, consumer complaints, testing, high-risk AI systems, and the effectiveness of the surrounding governance framework.

The direction is similar to what happens elsewhere in the compliance realm: A written standard establishes the expectation. The operating control demonstrates whether the organization follows it, the audit determines whether the control works, and corrective action addresses the condition when it does not.

The Human Decision Is the Control

AI use in claims will become more sophisticated over time. The question for claims organizations is not whether AI can perform increasingly complex functions; it can. The question is what authority the organization is prepared to give it, and how the organization will demonstrate that responsibility for the regulated outcome remains where it belongs.

If a regulator selects an AI-assisted claim during a market conduct examination, a defensible organization should be able to answer straightforward questions.

  • What system was used, and what was it authorized to do?
  • What did it contribute?
  • Who made the substantive claim decision, and what independent reasoning supported that decision?
  • How does the organization know the control is operating consistently across other claims?

If those questions cannot be answered from the claim record and the surrounding governance evidence, the AI policy is ahead of the operational reality. AI can support the claims professional by accelerating work, organizing complex information, improving access to data, and potentially creating greater consistency. Those benefits are real, and claims organizations should be able to use them. But the tool's efficiency does not change who is accountable for the claim—that is the purpose of the boundary of human decision. It establishes where AI can contribute and where professional responsibility must remain with the person making the regulated decision and also gives the organization something concrete to document, audit, and defend.

Having a human somewhere in the workflow is not enough. The defensible position is being able to show where human judgment was required, that it was actually exercised, and that the organization has tested the control well enough to know the difference.

References

  • "Model Bulletin: Use of Artificial Intelligence Systems by Insurers," NAIC, adopted December 4, 2023.
  • "Implementation of NAIC Model Bulletin: Use of Artificial Intelligence Systems by Insurers," NAIC, status as of August 6, 2026.
  • AI Risk Evaluation Supplement (formerly the AI Systems Evaluation Tool), Big Data and Artificial Intelligence (H) Working Group, NAIC, 2026.
  • "Bulletin 2026-9: Use of Artificial Intelligence Systems by Insurers," Mississippi Insurance Department, July 22, 2026.
  • "Commissioner's Bulletin B-0003-26," Texas Department of Insurance, June 12, 2026.
  • Senate Bill 26-189, Automated Decision-Making Technology, Colorado General Assembly, signed May 14, 2026, effective January 1, 2027.
  • "Artificial Intelligence Risk Management Framework (AI RMF 1.0)," NIST, January 26, 2023.

Opinions expressed in Expert Commentary articles are those of the author and are not necessarily held by the author's employer or IRMI. Expert Commentary articles and other IRMI Online content do not purport to provide legal, accounting, or other professional advice or opinion. If such advice is needed, consult with your attorney, accountant, or other qualified adviser.


Footnotes