Artificial intelligence (AI) has quickly moved into claims operations: In a
relatively short period of time, AI-enabled tools have become capable of performing work
that previously required a significant amount of manual claims handling.
An AI tool can summarize hundreds of pages of claim documentation, organize a
loss chronology, identify potentially relevant policy provisions, flag inconsistencies
in the reported facts, draft a claim note, and prepare a proposed communication for a
claims professional to review. Assume AI does all of that on one claim: The claims
professional reads the output, makes a few edits, and approves the next action or sends
the communication for final approval.
Who made the decision on the claim? The answer may not be as obvious as it
sounds; the claims professional was involved. There was technically a human "in the
loop." But that does not tell us whether the person independently evaluated the claim,
whether the AI established a conclusion that the claims professional later ratified, or
whether the tool simply organized information that the professional independently
analyzed. Those are different operating models, and from a claims governance standpoint,
the difference matters.
As AI becomes embedded more deeply into claims administration, fraud
detection, document analysis, valuation, communications, and other workflows, claims
organizations need to define what I call the boundary of
human decision—the point beyond which AI may assist, inform, organize, or
recommend, but responsibility for the substantive claim decision remains with a human
claims professional. A human in the workflow is not the same thing as accountable human
judgment. A defensible claims operation has to define that boundary, preserve evidence
that the human crossed it appropriately, and test whether that is actually
occurring.
An internal AI policy that simply says "human oversight is required" does not
accomplish that; the operational workflow has to do it.
The Regulatory Obligation Is Broad. The Claims Control Cannot Be.
The regulatory framework around insurance AI is developing
quickly, but the underlying obligation is not particularly new. The "Model Bulletin: Use of Artificial
Intelligence Systems by Insurers" written by the National Association of
Insurance Commissioners (NAIC) makes clear that consumer-impacting decisions made or
supported by AI remain subject to existing insurance laws, including unfair trade
practice and claims settlement requirements. It also places governance expectations
around claims administration and payment, fraud detection, third-party systems,
testing, documentation, and the degree of human participation in final
decision-making.
Mississippi's July 22, 2026, "Bulletin 2026-9: Use of Artificial
Intelligence Systems by Insurers" provides a current state example. It
similarly identifies human involvement in final decision-making as a factor in
determining the appropriate level of control and contemplates regulatory review of
automation constraints, model inventories, validation, bias controls, monitoring,
and third-party oversight.
Texas makes the human review expectation more explicit. In "Commissioner's Bulletin
B-0003-26," issued June 12, 2026, the Texas Department of Insurance states
that, when a regulated entity uses AI to make a consequential decision, a person is
expected to review and agree with the decision before action is taken. The bulletin
also ties AI use directly to existing Texas claims-processing requirements, adjuster
licensing obligations, market conduct surveillance, and examination authority.1
Requiring a person to review a decision establishes a human
checkpoint. It does not, by itself, establish what a meaningful review looks like
inside the claim file. A claims organization still has to define what the person is
expected to evaluate independently, what authority the AI was permitted to exercise,
and what evidence demonstrates that the professional judgment required for the
decision actually occurred.
The significance for claims organizations is not that regulators
have invented an entirely new claims-handling obligation for AI; they largely have
not. Existing obligations are increasingly being translated into expectations about
how the technology is governed, where the human responsibility remains, and how the
organization proves that controls actually operated.
Colorado has gone a step further in its broader automated
decision-making law. Signed May 14, 2026, and effective January 1, 2027, SB26-189 establishes the requirements for covered
automated decision-making technology used to materially influence consequential
decisions, expressly including insurance. Among other provisions, the law gives
consumers a right to request meaningful human review and reconsideration following
certain adverse outcomes.2 At the time of writing, the Colorado attorney general has
indicated that the state will not enforce the law until regulations are
finalized.
Nonadmitted/Surplus Lines Status Doesn't Equal Unregulated
One issue that receives less attention in the current AI
discussion is how these expectations reach surplus lines claims operations. Many
state insurance AI bulletins are directed to insurers holding a certificate of
authority in the issuing state. A nonadmitted insurer, by definition, does not hold
that certificate in the state where it is writing surplus lines business, which does
not mean that the regulatory obligation disappears.
A surplus lines insurer may not be the direct addressee of a
particular bulletin, which does not end the analysis. The insurer remains regulated
in its state of domicile, where it holds a certificate of authority and AI
governance expectations may apply directly. It also remains subject to the
underlying insurance laws that govern the claims activity itself.
The surplus line exemption is narrower than the word "nonadmitted"
sometimes suggests. It generally changes the regulatory treatment of areas such as
rates and forms, but it does not automatically remove the claims-handling standards
governing how losses are investigated, evaluated, communicated, and resolved.
This matters because emerging AI bulletins generally do not
establish an entirely separate standard for claims handling; they connect AI use
back to existing obligations involving unfair claims practices, unfair
discrimination, documentation, governance, and accountability. Whether a particular
state's unfair claims settlement requirements apply to surplus lines business must
still be evaluated on a jurisdiction-by-jurisdiction basis, and that analysis cannot
be skipped. Where the underlying claims obligation applies, using AI to perform or
support the activity does not eliminate the obligation.
Delegated claims models add another layer: When an insurer
delegates claims handling to a third-party administrator (TPA) or other claims
administrator, that administrator may also be the third party whose use of AI the
insurer is expected to understand and oversee. If AI influences claims
administration through the delegated operation, the insurer may need to be able to
explain what the system does, how its use is controlled, what oversight exists, and
how responsibility for the ultimate claim decision remains with an accountable human
decision-maker.
A surplus lines claims organization cannot practically wait for
every state to issue AI guidance specific to nonadmitted insurers before
establishing its governance framework. The more defensible approach is to identify
where AI enters the claims process, define the human decision boundary, document the
applicable controls, and evaluate state-specific requirements as part of the
organization's normal regulatory analysis.
Nonadmitted status may change how an AI requirement reaches the
organization, but it does not remove the need to understand and govern how AI
influences the claim decision; states do not all have identical requirements, and
not every law requires a named human to personally make every insurance decision,
which is important to keep in mind. The broader operational issue is whether the
organization can identify and define where human judgment is required and
demonstrate that it occurred.
The Human-in-the-Loop Problem
"Human-in-the-loop" has become one of the most common phrases in
AI governance and is also one of the easiest controls to overstate. A human can
technically participate in an operational workflow without meaningfully exercising
judgment. Consider the following three examples.
A claims professional receives an AI-generated coverage analysis, reads the recommendation, and routinely accepts it without returning to the policy or underlying facts to review or verify.
A fraud model assigns a high-risk score, and the claim is escalated because
that is what the workflow automatically does, even though nobody can explain
which factors materially produced the score; the thought process behind the
escalation for that particular claim is not articulated.
An AI tool drafts a settlement analysis, and the final claim note repeats substantially the same reasoning without identifying what the claims professional independently evaluated.
There is a human involved in all three workflows, which does not
necessarily mean the human made the decision; an approval action indicates that
someone reviewed or accepted an output. Standing alone, it does not tell us whether
the person independently evaluated the information underlying it or exercised the
judgment required by the claim decision. The organization remains responsible for
the regulated outcome, and that responsibility does not transfer to the model, the
developer, or the vendor merely because technology influenced the process. This is
where the boundary of human decision becomes an operating control rather than an AI
principle.
Define the Decision Boundary Before the Tool Is Used
Claims organizations should determine an AI system's permitted
role before putting the tool into production. It
is important to note that this role will not be identical for every system, and AI
may appropriately assist with activities such as summarization, organization,
document classification, routine drafting, information retrieval, pattern
recognition, or other defined functions. Some low-risk administrative activities may
also be capable of operating within predetermined parameters.
The risk changes when the AI output begins influencing a substantive decision affecting coverage, payment, settlement, claim disposition, or another consumer-impacting outcome. For a consumer-impacting decision, the operating model should make several things clear, including the following.
What is the AI authorized to do, and what is it not authorized to decide?
At what point is independent human judgment required?
What information must the claims professional independently review?
What circumstances require supervisory or compliance escalation?
Who is accountable for the final claim decision?
The answers should not exist only in an enterprise AI policy, and
they need to be reflected in the processes and procedures, training, system
permissions, supervisory expectations, quality assurance (QA) review criteria, and
documentation requirements governing the actual claim. A control that exists only in
policy is particularly vulnerable when the operational workflow makes ignoring it
easier than following it.
Make the Human Decision Reconstructable
Defining the boundary is only half of the control—the organization
also needs evidence that the boundary operated on the claim. An examiner, quality
reviewer, supervisor, or subsequent claims professional should be able to
reconstruct enough of an AI-assisted decision to understand the following four
things.
What information was considered.
What AI contributed to the process.
What the claims professional decided.
Why the claims professional independently reached that conclusion.
The fourth element is where the distinction between review and
judgment becomes visible. Consider an AI tool that summarizes a lengthy property
loss file and identifies several facts relevant to a potential coverage limitation
where the claims professional reviews the source documents, confirms some of the
facts, identifies another fact the AI did not give appropriate weight to, reviews
the applicable policy language, and reaches the final coverage conclusion. The claim
file should make the professional's analysis understandable and should not merely
preserve an AI-generated paragraph and add the words "reviewed and approved."
Documentation by itself does not create accountability; its value
is that it allows the organization to demonstrate what actually occurred, and the
purpose is not to create another documentation exercise simply because AI was used.
Therefore, the record needs to preserve enough information to distinguish the
technology's contribution from the professional judgment that produced the claim
outcome.
Govern the Vendor and the Change, Not Just the Employee
Some of the most significant AI exposure in claims may not come
from an adjuster opening a public generative AI application; it may come from
technology already embedded in the claims ecosystem. A valuation platform may
incorporate machine learning, a document management system may add automated
classification, a fraud vendor may change its scoring logic, a claims platform may
introduce an AI assistant through a routine software update, and the claims
professional may still see the same screen while the underlying decision influence
may have changed considerably.
The NAIC model bulletin specifically expects insurer AI governance
to encompass systems developed by third parties and contemplates due diligence,
contractual audit rights, regulatory cooperation, and oversight of third-party
compliance. Mississippi's 2026 bulletin similarly provides that insurers should
expect regulators to request third-party due diligence, contracts, audit or
confirmation processes, and validation and testing information. Colorado's 2026 law
adds another useful governance principle: Developers of covered automated
decision-making technology are required to provide notice of material updates or
modifications.
That creates a practical claims problem: If a vendor changes the
model, data inputs, or decision logic without telling the insurer or claims
administrator, the organization may be operating under a governance approval that no
longer reflects what the tool actually does.
Vendor governance, therefore, needs to extend beyond cyber
security and procurement. Depending on the use case, it may also need to address the
approved function of the AI, data inputs, available audit information, testing,
model limitations, material change notification, regulatory cooperation, and
circumstances that require the tool to be reevaluated. The vendor may own the
technology; however, the insurer still owns the regulatory burden, the risk, and the
defense of the outcome.
Test Whether Human Judgment Is Actually Occurring
Training is necessary, but having signed training completion
records does not establish that the control is working, nor does a signed AI policy
acknowledgment or approved-tool inventory. The organization eventually has to look
at the claims.
QA is one of the most practical places to determine whether the
boundary of human decision-making is functioning because quality review already
examines how claims professionals investigate, document, communicate, interpret
coverage, and support claim outcomes. AI governance does not need an entirely
separate testing structure if the existing QA process can be adapted to evaluate
it.
For an AI-assisted claim, that may mean reviewing whether the tool was authorized for the particular use, whether material facts and policy language were independently verified, whether human analysis can be distinguished from AI-generated content, whether the decision is supported by the underlying record, and whether required escalation occurred.
Aggregate patterns also provide useful information. A tool with a
consistently high override rate may be poorly calibrated for the claims it supports,
but a tool that is almost never overridden may raise a different question: Are
claims professionals continuing to independently evaluate its recommendations?
Neither result proves a control failure on its own.
An appropriate override may demonstrate that the control worked exactly as intended. Likewise, a high acceptance rate may reflect strong model performance or declining independent review. The purpose of the metric is to identify where further review is warranted, not to predetermine the answer.
There is another control issue that deserves attention: the expertise of the claims professional. Human oversight is only meaningful if the person performing that oversight has the knowledge necessary to recognize when an AI output is wrong, incomplete, or inconsistent with the actual claim record. If increasingly automated workflows cause claims professionals to lose the subject matter expertise necessary to detect and challenge a plausible but incorrect output, the organization weakens one of its own validation mechanisms.
The AI Risk Management Framework created by the
National Institute of Standards and Technology (NIST) recognizes the importance of
clearly defining human roles and responsibilities in AI decision-making and
oversight. In claims, the appropriate human-AI configuration should reflect the
consequence of the decision being made. The greater the potential effect on the
claimant, the more important competent and identifiable human judgment becomes.
What a Defensible Claims AI Control Environment Looks Like
There will not be one universal AI governance structure that works
for every insurer, TPA, managing general agent, or claims organization. The
operating model, lines of business, regulatory footprint, delegated authority
structure, technology stack, and use cases will differ. However, the components of a
defensible control environment are becoming easier to identify, and a defensible
claims AI framework should establish at least five things.
Know where AI is being used. The inventory should include internally deployed tools and AI capabilities embedded in vendor platforms. The relevant question is not simply what technology the organization owns, but where that technology can influence a claim.
Classify the decision impact. A drafting
tool compared to a system that materially influences coverage, payment, fraud
escalation, or settlement does not present equivalent risk and should not be
governed as though it does.
Define and document the human decision boundary. The organization should establish where AI assistance ends, where accountable professional judgment begins, and what evidence must remain in the claim record.
Govern change and third parties. AI-enabled vendor services should be subject to appropriate due diligence, monitoring, change notification, and reevaluation when the system or its role materially changes.
Test the control against actual claims. QA, complaints, corrective action, unusual acceptance or override patterns, and identified AI incidents should be used to determine whether the governance model is functioning in practice.
The NAIC's current regulatory evaluation work reinforces this movement toward operating evidence. Its AI Risk Evaluation Supplement asks regulators to examine written programs, risk assessment, training, third-party vendors, consumer complaints, testing, high-risk AI systems, and the effectiveness of the surrounding governance framework.
The direction is similar to what happens elsewhere in the
compliance realm: A written standard establishes the expectation. The operating
control demonstrates whether the organization follows it, the audit determines
whether the control works, and corrective action addresses the condition when it
does not.
The Human Decision Is the Control
AI use in claims will become more sophisticated over time. The
question for claims organizations is not whether AI can perform increasingly complex
functions; it can. The question is what authority the organization is prepared to
give it, and how the organization will demonstrate that responsibility for the
regulated outcome remains where it belongs.
If a regulator selects an AI-assisted claim during a market conduct examination, a defensible organization should be able to answer straightforward questions.
What system was used, and what was it authorized to do?
What did it contribute?
Who made the substantive claim decision, and what independent reasoning
supported that decision?
How does the organization know the control is operating consistently across other claims?
If those questions cannot be answered from the claim record and
the surrounding governance evidence, the AI policy is ahead of the operational
reality. AI can support the claims professional by accelerating work, organizing
complex information, improving access to data, and potentially creating greater
consistency. Those benefits are real, and claims organizations should be able to use
them. But the tool's efficiency does not change who is accountable for the
claim—that is the purpose of the boundary of human decision. It establishes where AI
can contribute and where professional responsibility must remain with the person
making the regulated decision and also gives the organization something concrete to
document, audit, and defend.
Having a human somewhere in the workflow is not enough. The defensible position is being able to show where human judgment was required, that it was actually exercised, and that the organization has tested the control well enough to know the difference.
References
"Model Bulletin: Use of Artificial Intelligence Systems by Insurers," NAIC,
adopted December 4, 2023.
Opinions expressed in Expert Commentary articles are those of the author and are not necessarily held by the author's employer or IRMI. Expert Commentary articles and other IRMI Online content do not purport to provide legal, accounting, or other professional advice or opinion. If such advice is needed, consult with your attorney, accountant, or other qualified adviser.
Artificial intelligence (AI) has quickly moved into claims operations: In a relatively short period of time, AI-enabled tools have become capable of performing work that previously required a significant amount of manual claims handling.
An AI tool can summarize hundreds of pages of claim documentation, organize a loss chronology, identify potentially relevant policy provisions, flag inconsistencies in the reported facts, draft a claim note, and prepare a proposed communication for a claims professional to review. Assume AI does all of that on one claim: The claims professional reads the output, makes a few edits, and approves the next action or sends the communication for final approval.
Who made the decision on the claim? The answer may not be as obvious as it sounds; the claims professional was involved. There was technically a human "in the loop." But that does not tell us whether the person independently evaluated the claim, whether the AI established a conclusion that the claims professional later ratified, or whether the tool simply organized information that the professional independently analyzed. Those are different operating models, and from a claims governance standpoint, the difference matters.
As AI becomes embedded more deeply into claims administration, fraud detection, document analysis, valuation, communications, and other workflows, claims organizations need to define what I call the boundary of human decision—the point beyond which AI may assist, inform, organize, or recommend, but responsibility for the substantive claim decision remains with a human claims professional. A human in the workflow is not the same thing as accountable human judgment. A defensible claims operation has to define that boundary, preserve evidence that the human crossed it appropriately, and test whether that is actually occurring.
An internal AI policy that simply says "human oversight is required" does not accomplish that; the operational workflow has to do it.
The Regulatory Obligation Is Broad. The Claims Control Cannot Be.
The regulatory framework around insurance AI is developing quickly, but the underlying obligation is not particularly new. The "Model Bulletin: Use of Artificial Intelligence Systems by Insurers" written by the National Association of Insurance Commissioners (NAIC) makes clear that consumer-impacting decisions made or supported by AI remain subject to existing insurance laws, including unfair trade practice and claims settlement requirements. It also places governance expectations around claims administration and payment, fraud detection, third-party systems, testing, documentation, and the degree of human participation in final decision-making.
Mississippi's July 22, 2026, "Bulletin 2026-9: Use of Artificial Intelligence Systems by Insurers" provides a current state example. It similarly identifies human involvement in final decision-making as a factor in determining the appropriate level of control and contemplates regulatory review of automation constraints, model inventories, validation, bias controls, monitoring, and third-party oversight.
Texas makes the human review expectation more explicit. In "Commissioner's Bulletin B-0003-26," issued June 12, 2026, the Texas Department of Insurance states that, when a regulated entity uses AI to make a consequential decision, a person is expected to review and agree with the decision before action is taken. The bulletin also ties AI use directly to existing Texas claims-processing requirements, adjuster licensing obligations, market conduct surveillance, and examination authority. 1
Requiring a person to review a decision establishes a human checkpoint. It does not, by itself, establish what a meaningful review looks like inside the claim file. A claims organization still has to define what the person is expected to evaluate independently, what authority the AI was permitted to exercise, and what evidence demonstrates that the professional judgment required for the decision actually occurred.
The significance for claims organizations is not that regulators have invented an entirely new claims-handling obligation for AI; they largely have not. Existing obligations are increasingly being translated into expectations about how the technology is governed, where the human responsibility remains, and how the organization proves that controls actually operated.
Colorado has gone a step further in its broader automated decision-making law. Signed May 14, 2026, and effective January 1, 2027, SB26-189 establishes the requirements for covered automated decision-making technology used to materially influence consequential decisions, expressly including insurance. Among other provisions, the law gives consumers a right to request meaningful human review and reconsideration following certain adverse outcomes. 2 At the time of writing, the Colorado attorney general has indicated that the state will not enforce the law until regulations are finalized.
Nonadmitted/Surplus Lines Status Doesn't Equal Unregulated
One issue that receives less attention in the current AI discussion is how these expectations reach surplus lines claims operations. Many state insurance AI bulletins are directed to insurers holding a certificate of authority in the issuing state. A nonadmitted insurer, by definition, does not hold that certificate in the state where it is writing surplus lines business, which does not mean that the regulatory obligation disappears.
A surplus lines insurer may not be the direct addressee of a particular bulletin, which does not end the analysis. The insurer remains regulated in its state of domicile, where it holds a certificate of authority and AI governance expectations may apply directly. It also remains subject to the underlying insurance laws that govern the claims activity itself.
The surplus line exemption is narrower than the word "nonadmitted" sometimes suggests. It generally changes the regulatory treatment of areas such as rates and forms, but it does not automatically remove the claims-handling standards governing how losses are investigated, evaluated, communicated, and resolved.
This matters because emerging AI bulletins generally do not establish an entirely separate standard for claims handling; they connect AI use back to existing obligations involving unfair claims practices, unfair discrimination, documentation, governance, and accountability. Whether a particular state's unfair claims settlement requirements apply to surplus lines business must still be evaluated on a jurisdiction-by-jurisdiction basis, and that analysis cannot be skipped. Where the underlying claims obligation applies, using AI to perform or support the activity does not eliminate the obligation.
Delegated claims models add another layer: When an insurer delegates claims handling to a third-party administrator (TPA) or other claims administrator, that administrator may also be the third party whose use of AI the insurer is expected to understand and oversee. If AI influences claims administration through the delegated operation, the insurer may need to be able to explain what the system does, how its use is controlled, what oversight exists, and how responsibility for the ultimate claim decision remains with an accountable human decision-maker.
A surplus lines claims organization cannot practically wait for every state to issue AI guidance specific to nonadmitted insurers before establishing its governance framework. The more defensible approach is to identify where AI enters the claims process, define the human decision boundary, document the applicable controls, and evaluate state-specific requirements as part of the organization's normal regulatory analysis.
Nonadmitted status may change how an AI requirement reaches the organization, but it does not remove the need to understand and govern how AI influences the claim decision; states do not all have identical requirements, and not every law requires a named human to personally make every insurance decision, which is important to keep in mind. The broader operational issue is whether the organization can identify and define where human judgment is required and demonstrate that it occurred.
The Human-in-the-Loop Problem
"Human-in-the-loop" has become one of the most common phrases in AI governance and is also one of the easiest controls to overstate. A human can technically participate in an operational workflow without meaningfully exercising judgment. Consider the following three examples.
There is a human involved in all three workflows, which does not necessarily mean the human made the decision; an approval action indicates that someone reviewed or accepted an output. Standing alone, it does not tell us whether the person independently evaluated the information underlying it or exercised the judgment required by the claim decision. The organization remains responsible for the regulated outcome, and that responsibility does not transfer to the model, the developer, or the vendor merely because technology influenced the process. This is where the boundary of human decision becomes an operating control rather than an AI principle.
Define the Decision Boundary Before the Tool Is Used
Claims organizations should determine an AI system's permitted role before putting the tool into production. It is important to note that this role will not be identical for every system, and AI may appropriately assist with activities such as summarization, organization, document classification, routine drafting, information retrieval, pattern recognition, or other defined functions. Some low-risk administrative activities may also be capable of operating within predetermined parameters.
The risk changes when the AI output begins influencing a substantive decision affecting coverage, payment, settlement, claim disposition, or another consumer-impacting outcome. For a consumer-impacting decision, the operating model should make several things clear, including the following.
The answers should not exist only in an enterprise AI policy, and they need to be reflected in the processes and procedures, training, system permissions, supervisory expectations, quality assurance (QA) review criteria, and documentation requirements governing the actual claim. A control that exists only in policy is particularly vulnerable when the operational workflow makes ignoring it easier than following it.
Make the Human Decision Reconstructable
Defining the boundary is only half of the control—the organization also needs evidence that the boundary operated on the claim. An examiner, quality reviewer, supervisor, or subsequent claims professional should be able to reconstruct enough of an AI-assisted decision to understand the following four things.
The fourth element is where the distinction between review and judgment becomes visible. Consider an AI tool that summarizes a lengthy property loss file and identifies several facts relevant to a potential coverage limitation where the claims professional reviews the source documents, confirms some of the facts, identifies another fact the AI did not give appropriate weight to, reviews the applicable policy language, and reaches the final coverage conclusion. The claim file should make the professional's analysis understandable and should not merely preserve an AI-generated paragraph and add the words "reviewed and approved."
Documentation by itself does not create accountability; its value is that it allows the organization to demonstrate what actually occurred, and the purpose is not to create another documentation exercise simply because AI was used. Therefore, the record needs to preserve enough information to distinguish the technology's contribution from the professional judgment that produced the claim outcome.
Govern the Vendor and the Change, Not Just the Employee
Some of the most significant AI exposure in claims may not come from an adjuster opening a public generative AI application; it may come from technology already embedded in the claims ecosystem. A valuation platform may incorporate machine learning, a document management system may add automated classification, a fraud vendor may change its scoring logic, a claims platform may introduce an AI assistant through a routine software update, and the claims professional may still see the same screen while the underlying decision influence may have changed considerably.
The NAIC model bulletin specifically expects insurer AI governance to encompass systems developed by third parties and contemplates due diligence, contractual audit rights, regulatory cooperation, and oversight of third-party compliance. Mississippi's 2026 bulletin similarly provides that insurers should expect regulators to request third-party due diligence, contracts, audit or confirmation processes, and validation and testing information. Colorado's 2026 law adds another useful governance principle: Developers of covered automated decision-making technology are required to provide notice of material updates or modifications.
That creates a practical claims problem: If a vendor changes the model, data inputs, or decision logic without telling the insurer or claims administrator, the organization may be operating under a governance approval that no longer reflects what the tool actually does.
Vendor governance, therefore, needs to extend beyond cyber security and procurement. Depending on the use case, it may also need to address the approved function of the AI, data inputs, available audit information, testing, model limitations, material change notification, regulatory cooperation, and circumstances that require the tool to be reevaluated. The vendor may own the technology; however, the insurer still owns the regulatory burden, the risk, and the defense of the outcome.
Test Whether Human Judgment Is Actually Occurring
Training is necessary, but having signed training completion records does not establish that the control is working, nor does a signed AI policy acknowledgment or approved-tool inventory. The organization eventually has to look at the claims.
QA is one of the most practical places to determine whether the boundary of human decision-making is functioning because quality review already examines how claims professionals investigate, document, communicate, interpret coverage, and support claim outcomes. AI governance does not need an entirely separate testing structure if the existing QA process can be adapted to evaluate it.
For an AI-assisted claim, that may mean reviewing whether the tool was authorized for the particular use, whether material facts and policy language were independently verified, whether human analysis can be distinguished from AI-generated content, whether the decision is supported by the underlying record, and whether required escalation occurred.
Aggregate patterns also provide useful information. A tool with a consistently high override rate may be poorly calibrated for the claims it supports, but a tool that is almost never overridden may raise a different question: Are claims professionals continuing to independently evaluate its recommendations? Neither result proves a control failure on its own.
An appropriate override may demonstrate that the control worked exactly as intended. Likewise, a high acceptance rate may reflect strong model performance or declining independent review. The purpose of the metric is to identify where further review is warranted, not to predetermine the answer.
There is another control issue that deserves attention: the expertise of the claims professional. Human oversight is only meaningful if the person performing that oversight has the knowledge necessary to recognize when an AI output is wrong, incomplete, or inconsistent with the actual claim record. If increasingly automated workflows cause claims professionals to lose the subject matter expertise necessary to detect and challenge a plausible but incorrect output, the organization weakens one of its own validation mechanisms.
The AI Risk Management Framework created by the National Institute of Standards and Technology (NIST) recognizes the importance of clearly defining human roles and responsibilities in AI decision-making and oversight. In claims, the appropriate human-AI configuration should reflect the consequence of the decision being made. The greater the potential effect on the claimant, the more important competent and identifiable human judgment becomes.
What a Defensible Claims AI Control Environment Looks Like
There will not be one universal AI governance structure that works for every insurer, TPA, managing general agent, or claims organization. The operating model, lines of business, regulatory footprint, delegated authority structure, technology stack, and use cases will differ. However, the components of a defensible control environment are becoming easier to identify, and a defensible claims AI framework should establish at least five things.
The NAIC's current regulatory evaluation work reinforces this movement toward operating evidence. Its AI Risk Evaluation Supplement asks regulators to examine written programs, risk assessment, training, third-party vendors, consumer complaints, testing, high-risk AI systems, and the effectiveness of the surrounding governance framework.
The direction is similar to what happens elsewhere in the compliance realm: A written standard establishes the expectation. The operating control demonstrates whether the organization follows it, the audit determines whether the control works, and corrective action addresses the condition when it does not.
The Human Decision Is the Control
AI use in claims will become more sophisticated over time. The question for claims organizations is not whether AI can perform increasingly complex functions; it can. The question is what authority the organization is prepared to give it, and how the organization will demonstrate that responsibility for the regulated outcome remains where it belongs.
If a regulator selects an AI-assisted claim during a market conduct examination, a defensible organization should be able to answer straightforward questions.
If those questions cannot be answered from the claim record and the surrounding governance evidence, the AI policy is ahead of the operational reality. AI can support the claims professional by accelerating work, organizing complex information, improving access to data, and potentially creating greater consistency. Those benefits are real, and claims organizations should be able to use them. But the tool's efficiency does not change who is accountable for the claim—that is the purpose of the boundary of human decision. It establishes where AI can contribute and where professional responsibility must remain with the person making the regulated decision and also gives the organization something concrete to document, audit, and defend.
Having a human somewhere in the workflow is not enough. The defensible position is being able to show where human judgment was required, that it was actually exercised, and that the organization has tested the control well enough to know the difference.
References
Opinions expressed in Expert Commentary articles are those of the author and are not necessarily held by the author's employer or IRMI. Expert Commentary articles and other IRMI Online content do not purport to provide legal, accounting, or other professional advice or opinion. If such advice is needed, consult with your attorney, accountant, or other qualified adviser.