What Automated Underwriting Can Teach Us About AI Accountability
Key Takeaways
Insurance did not begin delegating decisions to machines with generative AI. Automated underwriting systems have been doing versions of it for decades, and they generally worked by making authority explicit: what could be automated, what required referral, who could approve an exception, and what had to be recorded.
Generative AI changes the equation because one system can now move across several layers that older systems tended to keep separate. It can gather evidence, interpret it, apply a guideline, recommend a decision, explain the rationale, and sometimes trigger the next action. That makes it easier for authority to expand operationally before anyone has explicitly decided that it should.
The central governance question is therefore not simply whether an AI system can make an underwriting decision. It is where the carrier has given it authority to do so, and what has to be true before that authority can be exercised.
The Employee Did It
If an underwriter makes a bad decision, the carrier does not ordinarily get to say, “The employee did it.”
The underwriter was hired by the carrier, trained by the carrier, given authority by the carrier, and placed inside a system of guidelines, referrals, supervision, review, and audit. The carrier provided the information environment, established who could bind what, and decided when a decision had to move to someone with greater authority. We understand instinctively that the individual underwriter may have made the decision, but the carrier designed the system in which that decision was made.
The discussion around artificial intelligence sometimes creates a different frame. We ask what happens if the model makes the wrong decision, as though the model has somehow become a new accountable actor.
It is an important question, but it starts one level too low.
The carrier still decides what information the model can access, what it may infer, whether it can recommend an action or take one, what conditions require referral, who can override it, and what record needs to remain afterward. Those are not model questions. They are management questions.
In Noise Diversifies. Bias Accumulates., I argued that the useful comparison is not between perfectly deterministic machines and perfectly consistent humans. Neither exists. Insurance has spent generations managing imperfect human judgment through authority limits, referrals, review, audit, and rules. The piece ended with a different problem: artificial intelligence is moving the boundary between what can be structured and what historically required judgment.
The next question follows naturally. Who owns that boundary?
The answer has changed far less than the technology has.
Insurance Has Been Here Before
Machine participation in underwriting is not new.
Long before large language models, insurers and reinsurers were building systems designed to handle cases that could be expressed with enough structure while referring other cases to people. Swiss Re has been developing automated life underwriting for more than three decades. Magnum began by automating cases suitable for straight-through treatment while cases outside the automated path continued to require referral and human review.
The interesting point is not whether an early expert system could do anything remotely comparable to a modern generative model. It could not. The historical lesson is what the insurer had to decide before the system could safely make even a much narrower class of decisions.
What evidence was required? Which conditions could produce an automated outcome? What caused a referral? Where did machine authority end? Where did human authority begin?
Automation required a boundary.
As automated underwriting moved deeper into primary carriers, that boundary became part of ordinary operating infrastructure rather than an exotic technical concept.
The Rules Moved Inside the Carrier
By the 2000s and 2010s, configurable underwriting logic was becoming a normal part of carrier technology.
Celent documented Travelers using business-rules technology by 2008 to automate underwriting and pricing for small commercial accounts. IBM subsequently described the same environment in more detail, including automated eligibility determination, risk assessment, pricing, and referrals, with business rules centrally maintained and business users participating directly in creating, testing, and maintaining them.
Guidewire reflected the same operating principle. PolicyCenter was using rule-based eligibility, prequalification, and risk-suitability screening by 2006. By 2009, Guidewire had added explicit underwriting authority management, allowing carriers to delegate authority while enforcing underwriting guidelines inside the policy administration system.
Duck Creek later described a similar architecture in its 2020 Securities and Exchange Commission filing. Carrier-specific business rules, workflows, products, and underwriting authority could be configured separately from the underlying software platform.
The point is not that these systems were identical. They were not, and other core-system providers approached the same problems in their own ways. The important pattern is that rules, authority, referral, approval, exception, and record had become ordinary parts of carrier technology architecture.
The software executed the rule. The insurer owned the rule.
Owning the rule also meant deciding who could change it, when an exception was permitted, when a transaction had to stop, and whose authority was required to move it forward. That is why the history matters, and why it needs only a brief telling. Insurance has been delegating decisions to software for decades, and mature systems did not make accountability disappear. They forced organizations to define it.
Automated Underwriting Did Not Eliminate Judgment
Traditional automated underwriting worked largely by dividing decisions into recognizable categories. Some things could be specified well enough to automate. Some required additional evidence. Some exceeded the authority available at that point in the workflow. Some required judgment.
Those boundaries were never static. Better data made some referrals unnecessary. New rules automated decisions that once required an underwriter. New products and new risks created exceptions that had not existed before. The line moved, but it was still visible.
An underwriter generally knew when a case had been referred. A supervisor knew whether an approval required greater authority. The system knew when a transaction could proceed and when it had to stop.
That matters because the controls were not merely safeguards wrapped around the underwriting system. They were part of the underwriting system itself. Authority was part of the decision. Referral was part of the decision. Evidence requirements were part of the decision. The ability to reconstruct what happened later was part of the decision environment.
That is the part of the old architecture worth carrying forward.
Generative AI Blurs the Boundary
A large language model does not naturally arrive divided into those neat categories.
In a single workflow, a generative system might read a submission, extract information from it, summarize the account, infer something that was not explicitly stated, locate a relevant guideline, interpret the guideline, recommend an underwriting action, explain its reasoning, and initiate the next step. To the user, that can look like one continuous act of reasoning.
That is a real capability gain. It is also where the old boundaries begin to disappear.
A rules engine generally forced the organization to specify the rule before the rule could execute. Generative AI can produce useful output before the organization has fully decided what the system should be allowed to do.
That difference is easy to underestimate. A carrier can put an AI assistant into an underwriting workflow and discover very quickly that it is good at reading submissions, identifying missing information, summarizing loss history, or comparing an account with guidelines. Then someone asks it what they should do. If the recommendations are useful often enough, people begin to rely on them. Eventually, a workflow gets connected to the output.
No one may have sat down and formally declared that the model now possesses underwriting authority. Operationally, however, that may be exactly what has happened.
This is why “human in the loop” is not enough by itself as a governance description. The more useful question is when the machine must be referred because it has reached the boundary of its authority. A human clicking approve on a recommendation the organization expects to be accepted nearly every time may technically remain in the loop while exercising very little independent judgment.
That distinction is already showing up in insurance supervision. EIOPA has explicitly recognized that human oversight can mean very different things, ranging from genuine control to a much more passive monitoring role. The IAIS goes further, calling for clear lines of accountability, meaningful human oversight, and explicit guardrails around when an AI system can and cannot be deployed. That is a better way to think about the problem than simply asking whether a person remains somewhere in the workflow. The important question is what authority that person is actually exercising.
The better questions are architectural. What evidence may the model rely on, and what may it infer from that evidence? What can it recommend or decide on its own? When must a case be referred, independently reviewed, or preserved for later reconstruction?
Older systems forced many of those questions into configuration. Generative systems can let us postpone them, even while the model is already producing useful work.
That may be the more important governance problem.
The Evidence Is Not the Judgment
Another boundary is worth pointing out: the evidence is not the judgment.
Suppose an AI system recommends declining an account because of three adverse facts in the submission. There are at least two very different ways that decision can be wrong.
The first is an evidence failure. The model misread the submission, omitted something material, combined information incorrectly, or generated a fact that was never there.
The second is a judgment failure. The facts are accurate, but the conclusion drawn from them is questionable.
Those are not the same problem, and a carrier cannot evaluate the quality of the judgment without knowing whether the evidence underneath it was legitimate.
Large language models make this distinction particularly important because generated text can flatten the difference between retrieved fact, derived information, and inference. A confident paragraph may contain all three, and unless the architecture distinguishes them, the user may not know which is which.
Humans are hardly immune to evidence failures. Underwriters overlook information. Claims professionals misread files. People remember guidelines incorrectly. Anyone who has worked with a large enough book has encountered a decision resting partly on something everyone thought was in the file until someone went back and looked.
The mechanisms differ, but the organizational requirement does not. The carrier needs controls over the evidence layer before judgment is exercised on top of it.
That could mean requiring factual assertions to trace back to source documents. It could mean separating extraction from inference. It could mean preventing generated information from silently entering a system of record as fact. The exact controls will differ by workflow, but the principle should not.
Testability Is Not Replayability
There is also a difference between knowing how a model generally behaves and knowing what happened in one particular decision.
An insurer can test an AI system extensively. It can run thousands of cases through it, measure consistency, compare recommendations with experienced underwriters, test edge cases, and evaluate whether a new model version behaves differently from the old one.
All of that is valuable. None of it necessarily reconstructs what happened on a particular Tuesday afternoon six months ago.
Current supervisory thinking is moving in exactly that direction. The IAIS recommends tracing data sources, content-generation processes, model changes, and meaningful system activity through tools such as provenance records, model documentation, and event logs. The NAIC similarly contemplates model inventories, provenance, data lineage, validation records, testing, auditing, and information about the specific system involved in a consequential decision. In other words, the question is no longer merely whether the model passed validation. It is whether the insurer can reconstruct the decision environment that existed when the model acted.
To do that, the carrier may need to reconstruct the decision environment as it actually existed: the information and evidence available at the time, the model and version in use, the rules and authority then in force, and the workflow context that shaped the interaction. In some cases, later reconstruction may depend on preserving a contemporaneous record of the evidence, system version, authority, and action that mattered to the decision, rather than assuming the surrounding technology environment can be recreated years afterward. It may also need to know whether the case was referred, modified, overridden, or ultimately handled differently from the model’s recommendation.
AI did not invent this requirement. Insurance has kept underwriting files, claim notes, approval records, authority documentation, referral records, and audit trails for a reason. Consequential decisions occasionally have to be reconstructed.
What AI changes is the difficulty of that reconstruction if more of the decision process takes place inside systems whose intermediate steps were never designed to become part of the permanent record.
Testing tells you how the system behaves. Replayability tells you what happened. A mature decision architecture needs to know the difference.
Don’t Let the Model Become the Architecture
The easiest way to create a governance problem may be to ask one system to become everything at once: evidence source, rulebook, underwriter, authority structure, explanation engine, workflow controller, and audit trail.
A sufficiently capable model may appear able to perform pieces of all of those functions. That does not mean it should become the architecture connecting them.
The older systems offer a practical way to think about the problem because they forced insurers to separate layers that generative AI can blur. Regulators use different terminology, but the underlying pieces are increasingly familiar: data and provenance, controls on automation, human authority, accountability, and records that allow the decision to be examined later.
Evidence → Rules → Authority → Judgment → Record
The evidence establishes what is known. The rules establish what the organization has already decided. Authority establishes who or what may act. Judgment handles what remains uncertain. The record preserves what happened.
Not every layer has to be deterministic, not every layer has to use artificial intelligence, and not every layer requires a person. There is also no reason the boundaries have to remain fixed forever.
A carrier might allow AI to extract evidence while preventing it from creating new evidence through inference. It might let a model recommend within broad underwriting guidelines but retain deterministic eligibility rules. It might allow automated action within one authority band while referring anything beyond it. Over time, it might increase that authority after enough performance data exists to justify the change.
Those are design choices, and that is the point.
The important question for management is not simply, “Where are we using AI?” It is, “Where does AI have authority, and what has to be true before it can exercise it?”
That is a harder question, but it is also one insurance already understands.
The Carrier Owns the Architecture
An underwriter exercises authority delegated by the carrier. A rules engine executes authority encoded by the carrier. An AI system operates within authority permitted by the carrier.
Those actors are not interchangeable. They fail differently, require different controls, and create different records. But none of them becomes an independent insurance company merely because it participates in a decision.
Neither “the employee did it” nor “the model did it” is much of a governance system.
Current insurance supervision is converging on much the same principle. The NAIC places responsibility for AI-supported decisions inside the insurer’s governance program. EIOPA assigns ultimate responsibility for AI use to insurer management. The IAIS says insurers remain responsible for understanding and managing AI systems and their outcomes, including systems supplied by third parties.
The technology may be outsourced. The responsibility for the decision architecture is not.
The carrier chose the information environment, the rules, the authority, the referral structure, the safeguards, the override process, and the records that had to be preserved. Or, just as importantly, the carrier failed to choose some of those things explicitly.
Generative AI makes that second possibility easier because useful output can arrive before the surrounding architecture is complete. Older automated underwriting systems could be frustratingly rigid. Before they could do something, somebody generally had to define what they were supposed to do.
Large language models are appealing partly because they are not rigid in the same way. They can interpret, infer, generalize, and operate in ambiguity. Those are precisely the characteristics that push AI deeper into territory insurance historically reserved for judgment.
The commercial pressure also runs in that direction. Part of the value proposition is that generative AI removes friction, compresses handoffs, and makes previously separate tasks feel continuous. Defining explicit authority boundaries can seem slower than simply letting a useful system do more. That is exactly why the boundary has to be deliberate. If management does not decide where the authority ends, the workflow eventually will.
All of which is why the old management discipline matters more, not less.
The next generation of underwriting technology may look nothing like the rules engines that preceded it, but the governance lesson has changed far less.
The machine’s authority will expand one way or another. It can expand because management deliberately granted it, tested it, and surrounded it with controls. Or it can expand because useful recommendations quietly became expected decisions.
Either way, the carrier owns the choice.
Additional Reading
This article builds on the earlier discussion of human judgment, model variance, and the difference between noise and correlated error.
Noise Diversifies. Bias Accumulates.
Sources
- InsuranceIndustry.ai, Noise Diversifies. Bias Accumulates.
- Swiss Re, Magnum automated underwriting materials and historical commentary:
Magnum Assessment Engine;
historical commentary. - Celent, Business Rules and Underwriting Automation at Travelers, June 25, 2008.
- IBM, Business Rule Management for Insurance, including Travelers implementation details.
- Guidewire Software, PolicyCenter release announcement, March 15, 2006.
- Guidewire Software, PolicyCenter 4.0 announcement, November 2009.
- Guidewire Software, PolicyCenter underwriting-issues documentation.
- Duck Creek Technologies, Form S-1 Registration Statement, 2020, U.S. Securities and Exchange Commission.
- National Association of Insurance Commissioners, Model Bulletin on the Use of Artificial Intelligence Systems by Insurers, 2023.
- International Association of Insurance Supervisors, Application Paper on the Supervision of Artificial Intelligence, 2025.
- European Insurance and Occupational Pensions Authority, Artificial Intelligence Governance Principles: Towards Ethical and Trustworthy Artificial Intelligence in the European Insurance Sector, 2021.
