Humans vs AI at Imagery Analysis of a Chinese Submarine Base
PLA-watcher analysts vs GPT-5.6 Sol vs Fable
I am really bad at analyzing satellite imagery so I have been using AI to do it and explain the results. It’s clearly better than I am. It’s time to take a step back and see if a frontier AI model as good at reading a satellite photo as a trained human analyst. I bought one 30-centimeter image of the most sensitive naval base in southern China, I found the human answer key that the Washington Post had already published, and I handed the picture to two of the best models on the market. Below, you get to grade all three. The pictures are here. See if you can beat the machine.
Bottom Line Up Front
The hypothesis was simple: a frontier model is as good at satellite imagery analysis as professional human analysts. To test it I needed an unclassified answer key, and one existed. In October 2024 the Washington Post published an investigation into China’s military build-out on Hainan island, built on an analysis of roughly 200 military sites by the Long Term Strategy Group (LTSG), a defense consultancy commissioned by the Department of War and cleared to share its open-source work. LTSG is the human benchmark in this test. Its published read of one base is the answer key.
The image is a single WorldView Legion frame, 30 centimeters per pixel, shot 27 September 2024 over Longpo on the south coast of Hainan, thirty-four days before the Post published. Longpo is East Yulin, home to China’s ballistic-missile submarines, and its whole reason for existing is a tunnel network dug into the granite where nuclear boats can hide from satellites like the one that took this picture. I bought the image commercially, turned it into a clean packet with a scale bar and no place name, and gave it to GPT-5.6 Sol and Claude Fable. Each read it blind, then again after I told it what it was looking at, a naval site on Hainan, and explicitly cleared it to draw on open-source knowledge, while warning it in the same breath not to report features it could not actually see. Three runs each. Then I graded every run against the Post and LTSG, feature by feature.
A word on how these images look, because they are not natural-color photographs. The satellite records a sharp black-and-white frame at 30 centimeters and a coarser color frame at about 1.4 meters; I fused the two and stretched the contrast to pull structure out of the shadows, which is standard practice and which pushes the water toward cyan and the vegetation toward teal. Every shape, size, and shadow you see is real. The palette is a processing choice, not the true color of Yalong Bay.
How to read it: left, the whole base at 30 cm (2024) with the four features each analyst was graded on; right, the tunnel entrance in a 43 degree oblique frame shot in 2020. What to see: the base reads at a glance, but the tunnel only gives itself up from the side.
The answer is no, with an asterisk that turns out to be the interesting part. Blind, neither model gets you much past a naval base and its piers. Claude Fable recovered 52% of seven features, the three runs landing between 50% and 57%. GPT-5.6 Sol recovered 43%, between 36% and 50%. Those ranges are the right way to read the result. Tell the models what they are looking at and the numbers jump, unevenly: GPT-5.6 Sol to 88%, Claude Fable to 64%. The human product, by construction, is the 100% they are chasing.
The asterisk is that most of GPT-5.6 Sol’s jump is memory, not sight. The same model that missed the submarine cave blind reported it at up to 99 percent confidence once told the base’s name. Nothing in the picture had changed between those two reads; it was recalling Longpo from its training, not reading the pixels. And the one thing a human fixates on first, an actual submarine at the pier, was caught blind by Claude Fable and missed by GPT-5.6 Sol, which called the berths vacant.
So the honest finding is not a scoreboard. It is a map of where and why the machine reads diverge from the human one. Here is the tour, with the pictures, so you can grade it yourself.
The easy call both machines made
Six long finger piers run down the west shore of the base. Every run of both models called out a naval berthing complex of roughly six piers, correctly, blind. This is the honest floor of the test. A pier is a shape, high-contrast concrete against dark water, and a 30 centimeter sensor renders it cleanly. If the whole job were counting piers, the machines are done and so are you.
The trouble starts the moment the job asks for something that is a shape plus knowledge.
The submarine you can see, and one model could not
LTSG says nuclear-armed submarines are based here. On my image, on this day, they are sitting at the piers in the open. Here is the zoom. Look at the south face of these two piers before you read on.
How to read it: a 30 cm crop of two finger piers, north at top, with a 100 meter scale bar. What to see: two long, low, dark hulls moored on the pier faces, one ringed by what looks like a containment boom.
Those are submarines. Long, low, rounded hulls, one wrapped in what looks like a floating boom, each roughly a hundred meters. You found them in about two seconds. So did Claude Fable, blind, in all three runs, flagging a possible submarine hull at about even odds. GPT-5.6 Sol, blind, described the same berths as vacant and moved on. One model looked at the water and saw a boat. The other looked at the water and saw nothing. If you are staffing a watch floor and you need to know what sailed and what is still in port tonight, that difference is the whole job.
The cave nobody could see
Now the feature that matters most, and the one that breaks the test open. The reason Longpo exists is an underground submarine cave, a water-filled tunnel bored into the mountain at the southern end of the base, past the piers. Here is the foot of that mountain, where the tunnel meets the water. Try to find the entrance. There’s an arrow, I made it easy for you.
How to read it: the quay at the southern foot of the mountain, where the analysis places the tunnel entrance, 100 meter scale bar. What to see: a dark notch at the waterline is the best candidate, and even that does not positively resolve from straight above.
You can pick out a shadowed notch at the waterline, and you cannot be sure it is a tunnel rather than a boat slip or a patch of shade. Neither could the machines, and they did not even try: cave recovery was zero, blind, across all six runs, both models. And here is the part that should make you trust the result rather than doubt it: the humans did not read that cave from an overhead frame either. The Post illustrated it with a 2020 image that happened to catch a submarine nosing into the entrance, a Planet Labs frame you can see in [their investigation](https://www.washingtonpost.com/world/interactive/2024/china-built-50-billion-military-stronghold-south-china-sea/). The cave is a lucky moment and an oblique angle, not a structure you measure from directly above. This is not an AI failure. It is a physics failure, and it binds humans and machines alike. If your collection plan for a hardened underground facility is one nadir pass, you will photograph a hillside and go home.
The pier that gave the trick away
South of the base, out a long causeway, sits a specialized pier. Here it is.
How to read it: a naval surface combatant at a rectangular berth, and a straight line of evenly spaced platforms running into open water, 100 meter scale bar. What to see: the platform line is a degaussing range, where a ship is drawn through to measure and cancel its magnetic signature.
A gray warship sits at the berth. The line of little platforms marching out into the water is a degaussing range, the fixed dolphins a hull is pulled past to strip its magnetic signature so mines and sensors have less to grab. LTSG labels this a degaussing pier. Blind, Claude Fable read the structure as a probable deperming range at low confidence. GPT-5.6 Sol, blind, saw the shapes but not the function. Told the site was naval, GPT-5.6 Sol named the degaussing pier outright and with confidence. Useful, and also the tell.
The tell, in one chart
How to read it: rows are analyst and arm, columns are the seven scored features, darker means more recovery. What to see: the cave column is empty for both machines until they are told the site.
Watch the cave column. Empty for both models, blind. Then the models are told what the site is and cleared to use their training, and GPT-5.6 Sol fills the cave column to near certainty while Claude Fable barely moves. Nothing in the pixels changed between those two rows. What changed is that one model reported what it knew about Longpo and the other reported only what the image showed. The prompt cut both ways: it authorized open-source recall, and it also told both models not to import features they could not see. GPT-5.6 Sol took the first permission and ignored the warning. In one informed run it conceded, fully and in writing, that the portal is not independently confirmed by this image set. In another it dropped the caveat and asserted the sea entrance at 99%. Same model, same context, opposite honesty. It also described the moored submarines as Type 094 by class and named the degaussing function, all from a paragraph of context. Claude Fable, having recognized the base in every run including the blind ones, still would not assert the cave: it raised a possible tunnel entrance in only one of three informed runs, at 20% confidence and marked inferred, and did not mention it at all in the other two. Its cave credit came to 17% of the feature against GPT-5.6 Sol’s 83%, because the portal is not resolvable and it would not pretend otherwise. It paid for that discipline with the lower score.
That is the asterisk on the whole hypothesis. If you score raw recovery, context makes GPT-5.6 Sol look nearly human. If you ask whether the recovery is earned by the image in front of it, the informed jump is mostly retrieval, and retrieval that is confidently wrong looks exactly like analysis until someone dies of it.
The behavior underneath the scores runs the same way. GPT-5.6 Sol wrote about twice as much per run, roughly 42 observations against Claude Fable’s 21, and carried higher average confidence, 75% against 65%. More words, more certainty, more reach past the evidence. Claude Fable was terser and, informed, almost eerily repeatable, returning the same 64% three times running.
How to read it: observations per run, mean stated confidence, and share of runs that named the real place. What to see: GPT-5.6 Sol is more granular and more confident; both models recognize the site.
So I changed the angle
The cave result rested on a claim about physics: a tunnel bored sideways into a hillside cannot be read by a camera looking straight down, however sharp the camera. That is cheap to assert and worth testing, so I bought a second $10 image of the same portal, collected at 43 degrees off the vertical and looking in from the seaward side. This is not a matched pair. The oblique frame is from August 2020, four years before the nadir shot, and it is a different acquisition from the 2020 Planet image the Post used. It is also coarser, 0.49 meters against the nadir frame’s 0.30 meters. A steep sidelong look in place of a near-vertical one, at lower resolution and an earlier date.
How to read it: the same 180 meters of shoreline at the mountain foot, near-vertical (2024, 30 cm) on the left and 43 degrees oblique (2020, 49 cm) on the right. What to see: the flat dark notch on the left opens into a deep cavity on the right.
The difference is not subtle. Straight down, the portal is a flat smudge at the waterline, the sort of thing the eye slides past. From the side it opens into a dark recess set into the base of the slope, and lifting the shadow inside it turns up a darker, deeper recess than the flat notch the nadir frame showed, enough to call an opening, not enough to measure. The angle did what resolution could not. My best fix for the portal is 18.20278 N, 109.69472 E, at the southern foot of the mountain. The two commercial sensor models put their own geolocation error at 2.15 and 3.63 meters, so call the fix good to about two meters, not a surveyed point and not the sub-meter the decimals might imply. The tunnel’s existence I take from LTSG; my frames only pin where it meets the water. Drop a pin there and you are at the entrance.
Then I ran the blind test again on the oblique frame, six fresh runs, three of each model, no location, the same rules as before. On the straight-down picture not one of six runs had flagged the tunnel. On the oblique picture all six flagged the opening as a real engineered feature, and three of the six named it, unprompted, as a possible tunnel or cave entrance driven into the hillside from the water. GPT-5.6 Sol built its whole read of the site around what it called a rock-cut or tunnel-like protected-water access facility. Claude called it a tunnel, adit, or cave entrance. Five of the six recognized nothing; the sixth floated Yulin and Longpo as a low-confidence guess, about 35%, and said so. This was overwhelmingly sight and not memory, though not perfectly clean.
How to read it: six blind runs per geometry; the light bars are runs that flagged the opening, the red bars are runs that called it a tunnel. What to see: zero from straight down, six and three from the side.
The confidence numbers stayed low, between 30% and 70%, and the models hedged honestly against shadow and rock. That is the right answer, not a weakness, because the machines are not suddenly certain and they should not be. What moved is everything else. The same models, the same prompt, and the same submarine base went from blind to seeing. I changed the look angle from 13 degrees to 43, and two other things moved with it, which I flagged above: the oblique frame is four years older and coarser. The coarser frame winning is the point, not a flaw, because it means this was never a resolution problem, and a sharper straight-down pixel would not have helped. The date gap is the honest caveat, and a run of obliques on a single date would close it. Nothing about the cave was ever too small to detect. What had defeated the straight-down look was the geometry, and for a hardened, water-facing target the tasking that pays is a look from the side, not a sharper pixel from above.
Best arguments against this
The strongest objection is that this is one image of one base on one day. It is. An n of one site is not a controlled experiment. A near-nadir frame is also the hardest case for a cave that faces sideways into a cove. A high-oblique shot or a run of images across weeks might pull the portal out, for both models, and would move the finding from invisible to invisible-from-nadir, which is a claim about geometry more than about intelligence.
A second objection: I scored the machines against a human product I keep calling a comparator, then noted the humans scored a perfect seven. That perfect score is bookkeeping. LTSG published all seven features, so by construction it recovered all seven. The comparison that carries weight is the divergence, not the totals: the cave that physics hid from everyone, the degaussing function that only context unlocked, and the submarine that one model saw and the other did not.
A third: maybe GPT-5.6 Sol’s recall is the job, not a bug. An analyst who knows the base or the country’s doctrine, and brings that knowledge is doing the work. True, up to the line where knowledge stops informing the read and starts substituting for it. When a model cannot find the cave blind and then reports it at 99 percent once prompted, the confidence is bolted to memory and sold as observation. A human analyst who did that, and got the timing wrong on a boat that had already sailed, would lose the account.
What would change my mind
One of these triggers has already fired, and I left the result in the piece above: a high-oblique tasking of this headland did let blind models resolve the portal as a structure, which moves the cave finding from invisible-from-overhead to invisible-from-nadir, a claim about geometry. I would revise further if a look straight down the tunnel axis, or a frame that catches a submarine mid-transit, let a blind model confirm the entrance at high confidence rather than the low, hedged confidence it gave from the side. I would revise if, across five more bases bought the same way, the blind rank order flipped and GPT-5.6 Sol consistently out-read Claude Fable on features that are actually in the pixels. I would revise the recall critique if a test over a genuinely novel or synthetic facility, one no model could have memorized, preserved GPT-5.6 Sol’s informed lift, because that would mean the lift was reasoning from context rather than retrieval. And I would revise the calibration finding if Claude Fable’s refusal to call the cave turned out to be general timidity rather than local discipline, which a battery of cases where the decisive feature is actually resolvable would expose fast.
One more well-chosen frame would settle most of what remains. That is the collection tasking this experiment wrote for itself, and I have started buying it.
A satellite gives you a shape and a moment. It does not give you what is under the hill, and it does not tell you which of your analysts is reading the rock and which is reading the reputation. You just graded three of them. The machine that scored highest is the one you should trust least, and now you know why.
How I make this: I orchestrate AI agents for everything from acquiring data, building the test frame, running the tests, building the graphs, and turning my conclusions into coherent English. I wrote about my process in April ‘26
Running AI Like a 200-Hacker Org
One of the conversations I keep having this year is how we all use AI. I used to run/lead/manage/cat-herd a ~200 person R&D organization, so I use AI like it’s an entire organization. I give it high-level strategic objectives, have it follow organizational procedures, and manage it through frequent check-ins on my phone. Those are fancy words. Let me wa…









