AI in Warfare Part 3: Out-Fought, Not Out-Frightened
373 years of battle data versus the most repeated idea in pop military history.
Bottom Line Up Front
There are a good number of legal, policy, ethical, and even religious discussions asking if we should incorporate AI into weapons. I have been taking a different approach, projecting AI-characteristics onto well-studied historic battles, and using Lanchester equations to model what AI-weapons will do to battle. This series focuses on being able to embue AI weapons with an unbreakable will.
If you believe battles are decided by breaking the enemy’s will, the historical record disagrees with you most of the time. Across 607 usable defeats in a 660-battle database built for the US Army, roughly two-thirds (60.8-72.2% across reasonable threshold choices) were capability losses. The loser was so out-fought that it would have lost even fighting to the last man. And because prisoners taken in the pursuit get booked as winner effectiveness, that two-thirds is an upper bound, not a midpoint.
The true will-collapse signature, a loser at or above capability parity that broke early, shows up in about one-fifth of defeats (16-29% across the same thresholds). That is a lower bound. Correct it honestly, using measured rout casualties, and it rises to about a third (21-38%). It stays a minority under every correction I could construct.
Two-thirds of history’s losers were out-fought, not out-frightened.
How to read it: every dot is one defeat, 1600-1973, sorted into the three audit classes; the brass block is every will-collapse signature found.
What to see: two of every three losers were beaten on capability; pursuit booking makes that an upper bound.
There is a twist. The database’s compilers included a morale code, and checked against the will-collapse signature it carries essentially no information. Even the people who built the morale column could not make it predict.
The data comes from CDB90, the US Army Concepts Analysis Agency’s Database of Battles: over 600 land battles from 1600 to 1973, compiled in the 1980s by Trevor N. Dupuy’s Historical Evaluation and Research Organization (HERO), with per-side strengths, casualties, and outcomes. The research data was current as of July 1, 2026.
The findings
1. Clausewitz made will half the equation. Pop history made it the whole thing. Here is what he actually wrote: “If you want to overcome your enemy you must match your effort against his power of resistance, which can be expressed as the product of two inseparable factors, viz. the total means at his disposal and the strength of his will.” Means times will. Two factors, inseparable, multiplied. The sentence everyone repeats, that the object of war is to break the enemy’s will, is a paraphrase that dropped the other factor and grew teeth.
The paraphrase won because routs are cinematic and the winner’s version flatters the winner: we broke their spirit sounds better than we had more guns. The paraphrase is also half true, which is the durable kind of wrong. Armies really do quit long before annihilation. The Union quit First Bull Run at 2,896 casualties out of 35,000 present, about 8.3 percent; that story carried post 2. The question here is different. How often did the quitting, rather than the out-fighting, decide the result?
2. The audit is arithmetic you can redo. For every defeat in the database, take four numbers: each side’s starting strength and casualties. Lanchester’s square law (the standard attrition model in operations research, validated in this project’s first phase) turns them into one summary statistic for the loser, a capability ratio. In one line: R_cap = (loser per-man kill rate x loser starting strength squared) divided by the same product for the winner, with each side’s kill rate inferred from the casualties the other side suffered. At R_cap = 1, the loser had the fighting power to trade evenly. Below 1, it was out-matched on the day regardless of anyone’s nerve.
The classification rule is three lines. If R_cap fell below 0.70, the loss is booked as out-fought. If R_cap was at least 0.70 and the loser broke or quit before 15 percent casualties, that is the will-collapse signature. Everything else, near parity but bled past 15 percent, is booked as ambiguous: a hard repulse, not a collapse.
How to read it: the ratio multiplies inferred kill rate by strength squared, loser over winner; the cells below are the whole classification rule.
What to see: arithmetic you can redo from two columns per side, with thresholds declared before the data was scored.
The median out-fought loser had R_cap of 0.28, crushed on capability. The median will-collapse loser had R_cap of 1.24 and broke at 4.8 percent casualties. That is the textbook rout profile. Gettysburg lands in ambiguous, correctly: Lee’s army was near parity and fought to roughly 37 percent casualties. A hard repulse is not a panic.
3. Two-thirds of the losers were going to lose anyway. Of 607 defeats, 402 classify as out-fought. That is 66.2 percent, and it holds at 60.8-72.2% across the full grid of threshold choices (capability cuts of 0.6 to 0.8, break cuts of 10 to 20 percent casualties). The will-collapse class holds 130 of 607, 21.4 percent, ranging 16.0-28.7% across the same grid. The remaining 75 (12.4 percent) are the ambiguous hard repulses.
How to read it: each dot is one defeat, placed by the loser’s capability ratio (left of parity means out-matched) and its casualty fraction when it lost.
What to see: history’s losers mass far left of parity; the morale story lives in the thin brass wedge; Bull Run enters it only after the pursuit correction.
Drop any single war from the audit (there are 65) and the out-fought share moves only between 64.3 and 68.4 percent. Drop any entire century, including the whole twentieth, and it stays between 64.2 and 67.2 percent.
How to read it: each mark re-runs the classification with one war or century removed; vertical lines are the full-sample shares.
What to see: out-fought stays within 64.2-68.4% under every deletion, including the entire twentieth century.
One direction of error matters, and it runs against my headline, not for it. The database records total casualties, with no timeline inside the battle. When a loser breaks and runs, the prisoners and stragglers taken in the pursuit are booked as if the winner shot its way to them, making losers look more out-fought than they were during the actual fighting. A second channel pushes the same direction: an army that breaks early also stops inflicting casualties, which deflates its own measured effectiveness, so the upper-bound reading already absorbs it. Finding 5 measures the correction.
4. The database’s own morale column cannot find the collapses. CDB90’s compilers coded a morale advantage for many battles. If will collapse were the well-understood engine of defeat, that hand-coded column and my four-number arithmetic should point at roughly the same battles. They do not. The code flags the loser as demoralized in only 118 of 607 defeats, and of my 130 will-collapse battles it flags just 19. Agreement beyond chance, measured by Cohen’s kappa, is -0.06: statistically indistinguishable from zero (permutation p = 0.14, n = 607). It does not even flag First Bull Run, the most famous rout in American history. The field has long treated CDB90’s subjective columns with suspicion; the contribution here is measuring the failure cleanly, not discovering it.
I do not think the compilers were careless. The likelier reading is that their code captures pre-battle troop quality, green versus veteran, rather than what happened once the shooting started, and where it does fire it usually sits on top of a capability rout anyway (of the 118 flagged battles, 96 were out-fought on the numbers). But that is the point: professional military historians, building the Army’s reference database, could not make morale into a column that predicts. The database’s outcome codes, by contrast, agree with the arithmetic: all 26 losers recorded as annihilated classify as out-fought, and the will-collapse class contains no annihilations at all. Armies that collapse early do not fight to the death. The signature is coherent. The folklore column is not.
How to read it: left column, battles the database’s morale column flagged; right column, will collapses per the arithmetic; brass threads are the only agreements.
What to see: 19 threads out of 118 and 130; kappa -0.06, indistinguishable from zero. The morale column carries essentially no information.
5. The honest correction: half the loser’s casualties came after it broke. To size the pursuit bias flagged in finding 3, I collected the 11 famous routs whose published casualty figures split killed and wounded from captured and missing: Bull Run, Nashville, Missionary Ridge, Chancellorsville, Waterloo, Sedan, Rossbach, Königgrätz, Jena-Auerstedt, Blenheim, Austerlitz. These are deliberately selected severe routs, not a random sample; read every number in this paragraph as a bound, not a population estimate. In that sample, the share of the loser’s casualties that came after its line broke averaged 53% (n = 11, range 24-85%). Half the butcher’s bill was not combat. It was shooting retreating soldiers in the back. As Robert E. Lee said:
It is well that war is so terrible. We should grow too fond of it.
Remove those post-break losses and recompute. The average loser’s capability ratio roughly doubles, and 6 of the 11 battles change class, every one toward parity, none the other way. Bull Run is the cleanest case: about 45 percent of the Union’s casualty count that day was captured or missing men, and stripping them lifts its capability ratio from 0.795 to 1.42: an unambiguous will collapse at better-than-parity capability. The Union army at Bull Run was not out-fought. It was out-frightened. That is the exception, and it took the correction to make it clean.
How to read it: each row is one famous rout; the arrow is the loser’s measured fighting power once pursuit prisoners stop counting.
What to see: ratios roughly double and 6 of 11 verdicts flip toward morale; Sedan, Rossbach, Königgrätz, and Waterloo stay out-fought regardless.
Apply the measured correction to all 607 battles, reattributing casualties for losers whose recorded outcome was an actual rout or withdrawal, and the true will-collapse share rises from about one-fifth to about a third (21-38%, where 38% assumes an aggressive post-break fraction of 60%, above the measured 53% average). The control group inside the sample: Sedan’s loser moves from 0.10 to 0.37 and stays hopelessly out-fought. Sometimes an army runs because it is beaten, and the running merely sets the price.
6. A minority verdict, and it is the interesting minority. So the pop-Clausewitz trope fails its audit. In most defeats, will was not the deciding variable; capability was. In Ukraine, officials and analysts attribute about two-thirds to 85 percent of battlefield casualties to drones (the figures are mostly official-sourced; a June 2026 Oklahoma Watch fact-check rated them plausible but soft), yet after years of that attrition neither field army has collapsed into an army-scale rout. The front moves on capability: strikes, stockpiles, jamming. It is a firepower war with morale at the margins, which is what most of the 607 battles were too.
But flip the finding over. In about a third of historical defeats, at the corrected bound, the loser had the capability to trade evenly and lost anyway because somebody’s will gave out. Those are the battles where a combatant that cannot break, an AI, changes who wins rather than what winning costs. One battle in three is not a revolution. It is a large, specific, findable class of fights.
Best Arguments Against This
The anchor battle sits on the boundary you chose. Uncorrected, Bull Run’s capability ratio of 0.795 clears my 0.70 will-collapse cut by less than a tenth, and at a 0.80 cut it would classify as out-fought. If I had picked my thresholds to flatter the series’ anchor battle, this is where you would catch me. Three answers. The thresholds and the full 3-by-3 sensitivity grid are declared in the analysis script, not tuned after the fact. The population headline survives every cell of that grid. And the corrected Bull Run, at 1.42, is not sensitive to any threshold at all.
The database’s numbers are as suspect as its morale column. Fair. CDB90’s quantitative columns, the strengths and casualties this audit runs on, carry error too. Serious Lanchester curve-fitting moved to the purpose-built Ardennes and Kursk campaign databases partly for that reason (Bracken 1995; Fricker 1998; Lucas and Turkes 2004). Three things keep the classification standing on noisy inputs. It is a coarse two-way cut, not a curve fit, and it survives the full threshold grid. It survives dropping any single war or any entire century. And the class medians sit nowhere near the boundary: 0.28 for out-fought against 1.24 for will collapse, either side of a cut at 0.70. Noise would have to move battles across that whole gap, one way, to flip the shares. One caveat survives: any battle’s label is only as good as its entry. The claim is the aggregate shares, not any single battle’s coding.
Skill, not materiel or morale, decides modern battles. Stephen Biddle’s Military Power (2004) makes force employment the decisive variable, and that cuts across any raw capability-versus-will split. Conceded, and absorbed: the capability ratio is inferred from casualties actually inflicted, so employment quality is already inside it. R_cap is an effectiveness ratio, not a bean count.
Early withdrawal is not the same thing as will collapse. A prudent commander executing a planned withdrawal matches my behavioral signature, at or above parity and stopped before 15 percent casualties, without anyone panicking. Partly conceded. The class should be read as “broke or quit early,” not as a mind-reading of terror, and the outcome codes only partially separate the two. But the class has the right texture: its median member quit at 4.8 percent casualties while holding a capability edge, and it contains none of the database’s annihilation endings. Whatever mix of panic and prudence is in there, it is the ceiling on the morale story, and it is still a minority.
Eleven hand-picked routs cannot correct 607 battles. Correct. The 11-rout sample was selected on the dependent variable, so its post-break fraction of 53% is an estimate for battles like those, not for battles in general. That is why the correction reattributes casualties only for losers the database itself records as having routed or withdrawn, why the corrected share prints as a band (21-38%) rather than a point, and why the falsification list below asks for the real fix: a phase-separated re-audit at scale. Until that exists, the direction is one-way, toward morale, and the bound still leaves capability losses the majority.
What Would Change My Mind
1. By December 2027: someone runs a phase-separated re-audit, a machine-readable battle database splitting fighting-phase from pursuit-phase casualties across hundreds of battles (a digitized ACSDB or KOSAVE-class dataset would do it), and the true will-collapse share exceeds 50 percent. I would retract the headline, not just soften it.
2. By mid-2027: a replication of this audit under the linear (area-fire) Lanchester law flips the majority class on the same 607 battles. This is a weekend’s work for a skeptic.
3. Anyone produces a defensible reading of CDB90’s morale code under which it does predict the parity-plus-early-break signature, with agreement reliably above chance. Finding 4 dies, and I print the obituary. And please tell me how you did it so I learn how to better model it.
The close
A rout is not why most armies lost. It is how the losing looked. Two-thirds of history’s losers were out-fought, not out-frightened, and the most famous morale column in military-history data cannot find the exceptions. But the exceptions are real, they are about a third of defeats at the honest bound, and they cluster exactly where the capability ledger reads even.
That near-parity sliver is where an unbreakable AI stops being a rounding error and starts picking winners. Next post: what a will that cannot break is actually worth there, and why the answer is a tiebreaker, not a war-winner. The next AI-weapons series, the models are already working, focus on AI capabilities that out-fight.






