Skip to main content

Military AI: The confidence you cannot see

By Burak Oktenli
Real Clear Wire

In a command center, an operator looks at a screen where an automated system has recommended an engagement, and beside it sits a number: the system's confidence, say eighty-seven percent. In the seconds available, the operator must decide what to do with that number. The same moment plays out wherever machines now advise consequential decisions, from a benefits determination that can sink a household to a targeting recommendation that can widen a conflict. Across all of them, the same quiet thing is happening, and it bears on who actually governs the decision.

The person defers to the number. Not always, but often, and in a documented pattern. When a machine expresses confidence, people tend to go along with it, and in hard cases they will reverse even their own correct judgment because the system disagreed. In controlled studies, people given machine advice that was correct only half the time still let it pull their judgments off course, simply because the machine sounded sure. The opposite failure also happens: after one visible error, a person stops trusting the system entirely and switches it off, discarding whatever value it had. One failure surrenders too much; the other wastes a real tool. Both usually get filed under human psychology, to be fixed with better training or a cleaner interface.

That filing is the mistake, and for defense an expensive one. The deeper issue is not how much an operator trusts the machine, but who ends up holding the authority to decide. Treating the question as psychology hides a transfer of power taking place in plain sight.

Look closely at what happens when someone defers. They approve, so the record shows a human made the decision. But if that person deferred to a confidence score, they could not actually interpret, the decision was effectively the machine's, and the human supplied a signature. Authority migrated quietly, with no one choosing to hand it over and a paper trail insisting a human was in control throughout. Multiply this across a command structure, an agency, or an alliance, and you get a governance condition no one chose: consequential decisions drifting toward machines, one deferral at a time, while every procedure certifies that people remain in charge.

It helps to be precise about what the machine is actually doing when it produces that number. A modern AI system does not reason in any human sense; it synthesizes patterns from the data it was trained on and returns the output that best fits those patterns within the parameters its designers set. It does not understand the situation, weigh consequences, or grasp what is at stake; it computes the most statistically favorable response given its inputs and its programming. The confidence score is a byproduct of that computation, a measure of how well the current input matches the patterns the system has seen before, not a considered judgment about whether acting is wise. When we call this intelligence, we import connotations of understanding and responsibility that the system does not possess, and those borrowed connotations are exactly what make a human inclined to defer to it.

The reason a confidence number cannot fix this on its own is that a bare percentage says nothing about the conditions under which it holds. Eighty-seven percent confidence inside the range of situations a system was actually validated against means something real. The same number attached to a case the system has never encountered means almost nothing, and the two look identical on the screen. A poorly calibrated score does not announce its own uncertainty. It reads as authoritative while concealing how little it supports, which is the most dangerous form an automated output can take, because it invites exactly the deference it does not deserve. The operator has no way to see the difference, and the system was never built to show it.

This is where the stakes become a matter of accountable command rather than user experience. Military and civil institutions alike have long insisted that certain decisions belong to identifiable, answerable humans, a commander or an official who can be questioned and held responsible. As automated systems spread, the locus of judgment migrates without any of the deliberation such a shift would normally require. No policy declares that machines now decide; there is only an accumulation of moments in which a person accepts a number they cannot read. In the highest-stakes settings, that silent migration becomes the erosion of accountable human authority, hiding inside a design detail.

So, the real question is not how to make people trust machines the right amount through willpower, but how to build systems that communicate confidence in a way that keeps authority where it belongs. That is an engineering and governance task at once.

Confidence should never reach a person as a naked number. It should arrive bound to the conditions under which it is valid: the range of situations the system was validated against, and whether the case at hand falls inside that range. When the situation drifts outside what the system was built and tested for, or when calibrated confidence falls below what the stakes demand, the system's authority to act on its own should contract toward the human automatically, rather than presenting the same clean number and inviting the same reflexive approval. Calibration, understood this way, is part of the authority architecture rather than a display choice bolted on at the end, and it determines whether a human is genuinely deciding or merely ratifying.

This reframing converts an unmanageable demand into an enforceable requirement. Telling operators to trust automation appropriately is not actionable, and no training reliably calibrates a person against a number whose meaning shifts with context. Requiring that a system express confidence as a function of its validated context, and that its authority contract when the context drifts, is a specification that can be written into procurement standards, audited, and verified. It moves the burden from the individual's judgment under pressure to the system's design under institutional scrutiny, where it belongs.

As automated systems spread through the institutions of national security, human oversight is quietly collapsing into the act of accepting or rejecting a confidence score. If defense institutions keep treating how that score reaches a person as a question of psychology or interface polish, meaningful human authority will keep eroding into an approval, one click at a time, while the paperwork insists a human was always in charge. Drawing the line now means treating the communication of machine confidence as what it actually is: a decision about who holds authority, made in advance, by the people who build the machine.

Burak Oktenli is an independent researcher specializing in human-machine authority architecture and autonomous-systems governance. His work has appeared at the Modern War Institute at West Point, RUSI, and RealClearDefense. He holds an MBA and a Master of Professional Studies in Applied Intelligence from Georgetown University. 
 

Add new comment

This is not for publication.
This is not for publication.

Plain text

  • No HTML tags allowed.
  • Lines and paragraphs break automatically.
  • Web page addresses and email addresses turn into links automatically.
Article comments are not posted immediately to the Web site. Each submission must be approved by the Web site editor, who may edit content for appropriateness. There may be a delay of 24-48 hours for any submission while the web site editor reviews and approves it. Note: All information on this form is required. Your telephone number and email address is for our use only, and will not be attached to your comment.
CAPTCHA This question is for testing whether or not you are a human visitor and to prevent automated spam submissions. Image CAPTCHA
Enter the characters shown in the image.