Judgment, Determinism, and the Shape of Error in Insurance

Key Takeaways

  • Large language models can produce different outputs from identical prompts even when configured to minimize randomness, but some of that variation comes from how the system processes requests rather than the model “changing its mind.”
  • In a widely cited insurance noise audit, two randomly paired underwriters evaluating the same case differed by a median of 55 percent, far above the roughly 10 percent executives expected.
  • The useful distinction is not deterministic machines versus variable humans. It is increasingly visible, testable machine variance versus human variance that is difficult to replay or measure.
  • Insurance has been managing inconsistent human judgment for generations through authority limits, referrals, review, audit, and automated rules.
  • Greater consistency has a cost. When many decision-makers rely on the same model, independent errors can become correlated errors.

The Objection

One of the most serious objections to using large language models in insurance decision-making is also one of the simplest: they are non-deterministic.

Ask the same question twice and you may not get precisely the same answer. That is uncomfortable enough when the task is drafting an email. It becomes something else entirely when the output might influence an underwriting decision, a claim evaluation, a compliance determination, or any other decision carrying regulatory, fiduciary, or errors-and-omissions exposure.

The objection deserves to be taken seriously. An insurer should be able to explain how a consequential decision was reached, reproduce the conditions that produced it, and determine whether the same facts would be treated consistently tomorrow.

But there is a question hiding inside the objection.

Compared to what?

Insurance did not discover variable decision-making when generative AI arrived. We have been living with it for as long as professionals have exercised judgment.

The more interesting question is not whether artificial intelligence is perfectly deterministic. It is what kind of variance we are comparing it with, and whether we can see that variance.

The Machines Aren’t Deterministic Either

The critics are right about the basic fact. Even at what AI engineers call “temperature zero,” essentially telling the model to minimize randomness and favor its most likely response, identical prompts do not necessarily produce identical results.

The reason is more interesting than the usual explanation.

Thinking Machines Lab demonstrated in 2025 that some of this variation comes from how AI systems process many users’ requests at the same time. For efficiency, those requests are grouped together. As the number and mix of requests change, the calculations used to produce an answer can change slightly. The model is not reconsidering the question because it woke up in a different mood. The same request is being processed under slightly different computational conditions.

In one experiment, Thinking Machines ran the same prompt 1,000 times. The system produced 80 different answers. More interestingly, all 1,000 were identical through the first 102 tokens before they began to diverge. When the researchers changed how the requests were processed, all 1,000 answers became identical.

There was a price. In a separate performance test, the original deterministic approach was roughly 61.5 percent slower than the default. Within weeks, another engineering team reported reducing the average slowdown to about 34 percent.

Those percentages are not a universal price list. Hardware, software, workload, and implementation matter, and the cost will continue to change as the technology improves. Current production software already allows operators to choose greater reproducibility while accepting a performance tradeoff.

That changes the nature of at least part of the determinism objection. What sounds like an inherent philosophical limitation is, in part, an engineering choice with a measurable price.

That does not make machine variance disappear. It makes some of it something an organization can decide whether to pay to control.

The Humans Never Were

Now consider the baseline.

Daniel Kahneman, Andrew Rosenfield, Linnea Gandhi, and Tom Blaser described the insurance problem in a 2016 Harvard Business Review article. Kahneman, Olivier Sibony, and Cass Sunstein later developed it more fully in Noise: A Flaw in Human Judgment, which reports the figures used here.

In one insurance engagement described in Noise, experienced underwriters were asked to price the same fictitious cases independently. Executives at the insurer expected differences of roughly 10 percent and considered that level tolerable.

The actual median difference between two randomly paired underwriters was 55 percent.

The distinction matters. That is not a 55 percent range or statistical variance. It is the median difference between two randomly selected professionals, expressed relative to the average of their estimates. In the example used by the authors, if one underwriter set a premium at $9,500, another would be at approximately $16,700 rather than $10,500.

Claims professionals in the same exercise produced a median difference of 43 percent.

There are limitations. The audit involved five fictitious cases at one unnamed large insurer. It arose from a consulting engagement rather than a peer-reviewed experimental program. I have not found a post-2021 insurance underwriting study that cleanly replicates the same cross-underwriter dispersion measure, and it would be wrong to imply that 55 percent is some universal constant of underwriting.

But the most important result was never 55 percent.

It was the gap between 10 and 55.

The insurer’s own leadership underestimated the inconsistency of a core professional decision by roughly a factor of five.

We spend a great deal of time worrying that AI systems may not return exactly the same answer twice. We have spent much less time asking how often two qualified humans would have returned the same answer in the first place.

The Difference Was Never Determinism. It Was Visibility.

Anyone who has spent enough time around insurance has seen both versions of professional judgment.

A claims adjuster makes a call that turns out to be exactly right, even though nobody could reconstruct the full reasoning from the file afterward. The same adjuster, perhaps years later or perhaps the following week, makes another call nobody can defend once the outcome is known.

Those are the decisions we remember.

The deeper problem is that the counterfactual was never observable. We saw the decision that was made. We never saw the decisions the same person might have made under slightly different conditions.

That is where the machine comparison becomes more interesting. Run the same submission through a model a thousand times and you can begin to characterize what it does. You can see how often the answer changes, how far it moves, where the tails are, and what kinds of prompts or facts make the system unstable. You can test it before the decision binds.

You cannot rerun an underwriter.

There is no way to restore Tuesday morning and run the decision again. We cannot hold the submission, workload, recent loss experience, production pressure, sleep, mood, and every other input constant and ask the same person to make the decision a thousand times.

The decision-maker itself changes.

An underwriter in the second week of a difficult quarter is not precisely the same decision system as that underwriter in the first week. A claims adjuster who has just been burned by an apparently minor injury that developed badly has information, experience, and perhaps a little recency bias that did not exist before. An actuary deciding whether a data point is an outlier is not operating independently of the last reserve review.

That variability can contain expertise. It can also contain noise. From the outside, they can look remarkably similar.

Machine systems change too. Models are updated. The information they draw from changes. Providers adjust the systems running them. Even one of today’s widely used AI serving platforms limits its reproducibility guarantee to the same hardware and the same software version.

The advantage is therefore narrower than saying machine variance has been solved.

It is that machine behavior can increasingly be tested repeatedly under known conditions, with changes documented from one version to the next. Human judgment generally offers no comparable way to replay the decision.

Visible variance is better than invisible variance. You can measure it, test it, govern it, and decide how much you are willing to tolerate.

Insurance has spent a very long time trying to make the invisible kind more manageable.

A Century of Managing Invisible Variance

Insurance never behaved as though professional judgment is perfectly consistent.

Look at the structures we built around it.

Underwriting authority limits determine who can make which decisions. Referral thresholds require a second set of eyes. Large or unusual accounts go to committees. Claims departments conduct roundtables. Files are audited. Actuarial work is reviewed. Licenses and continuing education establish minimum standards. Errors-and-omissions coverage exists because professional judgment sometimes produces professional mistakes.

Even the ability to walk down the hall and ask someone to explain a decision is part of the architecture.

Those controls are not proof that the industry consciously understood decision noise in Kahneman’s terms. They do reveal something simpler: organizations behave as though individual judgment needs boundaries.

I learned a version of that lesson from the distribution side. When an agent was losing an argument with an underwriter, there was always the temptation to take the agent’s case higher. Eventually the useful question stopped being who I instinctively wanted to side with. It became whether the agent’s argument was strong enough that I was willing to make it on someone else’s behalf. If it wasn’t, the argument stopped there.

Requiring judgment to survive another person’s scrutiny is a noise-management mechanism, whether anyone calls it that or not.

Then insurance went further. We automated the decisions we could make repeatable.

Traditional automated underwriting systems encoded underwriting philosophy into rules. Swiss Re’s Magnum materials provide a current example: selected rules can be configured to an insurer’s underwriting philosophy, while cases requiring human review can still move to manual underwriting.

That architecture matters more than the individual system.

Automated underwriting did not eliminate judgment. It drew a boundary around it.

Where the decision could be expressed as a rule, the system could apply it consistently. Where the rules ran out, the case went back to a person.

Recent research points in the same direction. In a 2025 Journal of Business Research study, Gavin Maistry, Jochen Reb, Shenghua Luan, and Thomas Menkhoff tested simple-rules training in an insurance-underwriting context with 220 participants and an active control condition. The intervention improved decision quality in both accuracy and consistency, with larger benefits for less-experienced participants. Maistry is a senior Munich Re executive. The study does not replicate the Kahneman audit or measure the same cross-underwriter dispersion. It does show that structure can improve professional judgment.

That is important because AI is not introducing the boundary between rules and judgment.

It is moving it.

Modern systems can read documents, interpret language, score risks, find relevant guidelines, and generate recommendations in areas that traditional rules engines often referred to humans. And different parts of the system can behave differently: a fixed eligibility rule may be acting on information produced by a model whose answer can vary.

So asking whether “automated underwriting” is deterministic is already too crude a question.

Which part of the system?

Insurance has been doing this for decades: constrain judgment where it can, structure it where it cannot, and automate what can be made repeatable.

The Trade Nobody Priced

Imagine ten underwriters making somewhat different mistakes.

One is too aggressive on a class another dislikes. One gives management more credit. Another puts more weight on loss history. Across a sufficiently large book, some of those idiosyncratic errors can offset one another.

Now imagine ten underwriters receiving the same recommendation from the same model.

The average individual decision may improve. The range of answers may narrow dramatically. Management gets something it has wanted for decades: consistency.

But when the shared model is wrong, everybody can be wrong in the same direction.

That is not merely a hypothetical concern about a catastrophic model failure.

Jon Kleinberg and Manish Raghavan examined what they call “algorithmic monoculture”: many decision-makers relying on the same algorithm. In a 2021 Proceedings of the National Academy of Sciences paper, they showed that overall decision quality can decline even when that shared algorithm is more accurate for each individual decision-maker than the alternatives. The result does not require an external shock. It can emerge under normal operations.

Insurance executives already have vocabulary for the underlying problem.

Correlation.

If every underwriting model likes the same accounts, capacity can chase the same risks. If every model dislikes the same characteristics, some risks can become functionally residual without a regulator creating a residual market. If every decision-maker inherits the same blind spot, the error no longer scatters across the portfolio.

Noise diversifies. Bias accumulates.

There is an obvious objection here. Kahneman’s prescription for noise included rules, algorithms, and decision hygiene. The 2025 underwriting study provides contemporary evidence that imposing simple structure can improve accuracy and consistency. Why warn about the very thing that appears to solve the problem?

Because both can be true.

Kahneman was asking how to reduce unwanted variation in individual decisions. Kleinberg and Raghavan were asking what happens when multiple decision-makers converge on the same algorithm. One describes the benefit of reducing noise. The other describes a potential system-level cost of achieving that consistency through common decision machinery.

That makes an AI procurement decision something more than a technology decision.

It can also be a correlation decision.

The insurer may be buying a better decision process and, at the same time, buying more commonality in how its errors occur. Across an industry increasingly dependent on the same underlying AI models, data sources, vendor platforms, and decision systems, that question gets larger still.

Consistency is valuable, it is not free.

But correlation does not return us to invisible variance. It creates a different measurable problem. A carrier can test whether its own model-driven decisions are clustering too tightly across segments, software versions, and repeated cases. Being able to inspect the behavior remains an advantage. It simply does not eliminate the cost of commonality.

Which Variance Is Wisdom?

The hardest part is that we do not actually want to eliminate all variation.

Some of it is why we hire professionals.

The unusual account that does not fit the manual. The management team whose controls are substantially better than its class would suggest. The claim whose early facts do not quite make sense. The data point an experienced actuary refuses to discard as an outlier because it looks like the beginning of a trend.

Those departures can be expertise.

But variation can also be fatigue, anchoring, recency, production pressure, mood, habit, or the last large loss.

In a spreadsheet showing only the resulting decisions, wisdom and noise can look identical.

That may be the most important work insurers have to do before automating judgment. The task is not simply to observe how experienced people decide and teach a model to reproduce it. Historical decisions contain the organization’s accumulated expertise, but they also contain its accumulated inconsistency and bias.

Automating the past without separating those things risks making both more efficient.

Conclusion

For the last several years, one of the recurring arguments against consequential use of AI has been that the systems are non-deterministic.

The industry spent the previous century building referral thresholds, authority limits, review, and audit because professional judgment was never perfectly consistent either. Then it built automated underwriting systems to make repeatable the decisions that could be expressed as rules, while sending the exceptions back to humans.

AI pushes that boundary farther into the territory we used to reserve for judgment.

That can make variance more visible, decisions more consistent, and professional expertise more scalable. Those are real advantages.

It can also change the shape of the error.

Scattered human mistakes may be difficult to see and expensive to manage, but some of them cancel. A shared model can make decisions better one at a time while making the remaining mistakes more alike.

The question was never whether AI could make insurance deterministic.

It is what shape of error we want on the book.

And insurance already knows what to call the kind that moves together.

 


 

Sources

AI Disclaimer: This content was created with assistance from artificial intelligence technology. While content is based on factual information from the source material, readers should verify all details directly with the respective sources before making business decisions.