Neural Nations · CAMS research report

Reading Societies Without Their Stories

A blind, pre-registered, multi-model test of CAMS on ancient collapse and recovery

Kari McKern · ComplexityWorkz · 2 October 2026 · 20 society-windows · 22 blind scoring passes per prompt from three AI model families (Claude, Grok, GPT) · working paper

CAMS reads a society through eight functions, not through the story it tells about itself. Deep time is a hard place to test that claim. The record is thin, often written by victors or later traditions, and easy to fill with modern ideas about states and progress. So we gave three AI model families the same blind rubric, had each score the same twenty ancient society-windows in fresh contexts, and checked five hypotheses that were written down before any scoring began.

What all three models agreed on

  1. Collapse drains everything at once. In both the Egyptian and the Aegean breakdowns, every function lost surplus (Capacity − Stress) and shared thinking (Abstraction × Coherence) together.
  2. The Aegean recovered far more slowly than it fell. Egypt rebuilt to near Old Kingdom levels in under three centuries. Under Claude and Grok, the Aegean was still below its pre-collapse level 500 years later.
  3. Egypt gave an early warning; the Aegean did not. Coherence between rulers (Helm) and meaning-makers (Lore) fell for two centuries in Egypt while surplus was still positive. In the Aegean it stayed flat until the breakdown.
  4. Debt-relief edicts did not coincide with lower labour stress in Mesopotamia. This test has a design flaw (below), but it was pre-registered and it failed.

The models also disagreed, and that disagreement is a result in itself. They differ in scale and, more importantly, in what they accept as evidence. The one hypothesis about non-state societies (“no ladder”) passed under Claude, sat exactly on the cut-off under Grok, and could not be tested under GPT.

1. What was tested

CAMS scores eight functional nodes: Helm (command), Shield (force and order), Lore (religion and meaning), Stewards (owners of stores and land), Craft (skilled production), Hands (mass labour), Archive (records and memory, written or oral) and Flow (exchange). Each node gets four scores from 1 to 10: Coherence, Capacity, Stress and Abstraction. Two composites carry the argument:

Cases

Prompt A scored twelve societies at one anchor year each: Longshan China, Shang Anyang, the Mature Harappan cities, Old Kingdom Egypt, First Intermediate Egypt, Hammurabi’s Babylon, Neopalatial Crete, Mycenaean Greece (LH IIIB), the Aegean after the palace destructions, pre-contact Aboriginal Australia, Māori Aotearoa and the Xiongnu. Prompt B added eight windows to build two collapse-and-recovery series (Egypt and the Aegean) and one Mesopotamian series. It rescored the two collapses as overlap anchors.

Safeguards

ModelPasses per promptConditions
Claude (Opus 5.5)15Three independent runs of 5; fresh subagent per pass; prompts A and B always separate
Grok 4.74One fresh conversation per prompt, prompt file only
GPT (ChatGPT Work/Codex, version unverified)3Fresh isolated agent per pass; prompts A and B in the same context

An early non-blind Claude pass and a Grok pass that had read the hypotheses are kept in the run log but excluded from the analysis. One upload turned out to be a byte-identical copy of an earlier run and was not counted.

2. Is the instrument stable?

Within one model, yes. Between models, the order of societies mostly holds, but the scale does not.

ComparisonMean absolute difference per score (A / B)Rank agreement of societies (A / B)Mean offset
Claude run vs Claude run (5-pass ensembles)0.14–0.23 / 0.11–0.220.97–0.99 / 1.00under ±0.15
Single Claude pass vs other Claude runs0.26–0.29——
Grok (4) vs Claude (15)0.39–0.53 / 0.41–0.570.92 / 1.00+0.2 to +0.47
GPT (3) vs Claude (15)0.89–1.64 / 0.84–1.800.81 / 0.99+0.6 to +1.2; Abstraction +1.6 to +1.8

All three models were about equally self-consistent: the spread between passes of one model was 0.12–0.44 points per dimension. So the gaps between models are not noise from single passes; each model applies the rubric in its own consistent way. The ten time windows in prompt B were ranked almost identically by every model (0.98–1.00). The twelve societies in prompt A were ranked less consistently (0.81–0.92), and the shifts fall on societies without visible rulers or without writing.

Scale offsets matter less than they look. Most tests compare windows within one model, so a model that scores everything a point higher still reaches the same verdict.

3. Collapse and recovery

The two collapses look alike on the way down and very different on the way back.

Mean Energy (Capacity − Stress) across the eight nodes, by window

Ensemble means: Claude 15 passes, Grok 4, GPT 3. Years BCE. Below zero, functions are under more strain than they have capacity for. Both panels use the same scale.

Egypt reaches about zero by 2250 BCE and bottoms out in the First Intermediate Period. All three models put the Middle Kingdom back near the Old Kingdom’s height. In the Aegean, the fall from LH IIIB is the steepest in the data. By 750 BCE, Claude and Grok still put the Aegean below its LH IIIB level, while GPT puts it slightly above.
Show the data as a table

The early-warning signal (H3) is a separate measure: the mean Coherence of Helm and Lore. In Egypt it fell from about 8 to between 5 and 6.5 across the two pre-collapse windows under every model, while Energy was still positive. In the Aegean it stayed flat until the breakdown. Egypt’s collapse began as a loss of shared purpose at the top. The Aegean’s struck a system that still looked coherent at the top.

SeriesFall, Energy per centuryRecovery, Energy per centuryHelm–Lore Coherence before breakdown
Egypt (Claude)−2.86+2.568.1 → 5.1
Egypt (Grok)−2.02+2.018.0 → 5.4
Egypt (GPT)−1.75+1.398.0 → 6.5
Aegean (Claude)−7.27+1.955.5 → 5.7
Aegean (Grok)−6.19+1.276.2 → 6.5
Aegean (GPT)−4.17+1.596.8 → 6.8

Two patterns at node level are worth noting. First, in both collapses Lore holds up best: it has the highest Node Value in First Intermediate Egypt under all three models. The lowest nodes are Helm, Hands and Flow in Egypt, and Archive and Helm in the Aegean: meaning-making outlasts command and administration. Second, Hands (mass labour) is the lowest or second-lowest node in almost every state window under every model, while non-state societies (Aboriginal Australia, Māori Aotearoa) do not show that gap. The state’s stress lands on its labourers.

4. Hypothesis results

Four of the eight tested claims gave the same verdict under all three models: three held and one (H4) failed. Two held in direction only, and two depended on the scorer.

Holds, all models
H2a: in collapse, Energy and Cognition fall together in every node
Claude 8/8 nodes in both collapses; Grok 8/8 and 7/8; GPT 8/8 and 7/8.
Holds, all models
H2b (Aegean): recovery is slower than the fall
The Aegean falls 4–7 Energy points per century and recovers at 1.3–2.0. Egypt mostly does not show this: it recovers about as fast as it fell under Claude and Grok, more slowly under GPT.
Holds, all models
H3 (Egypt): Helm–Lore coherence falls before breakdown
A fall of 1.5–3.0 points under every model while Energy is still positive. In the Aegean the same test fails under every model.
In direction
H1b: maritime societies have Flow above Helm
Minoan Crete +2.1 to +4.1, Māori +0.8 to +1.2, against −0.6 to −0.8 for the other societies. The direction holds; the size varies by model.
In direction
H1a: river-valley states have stronger Stewards and Archive
Claude pooled +0.82, Grok +0.40, GPT +0.71 Node Value points, but one of three Claude runs points the other way. Small, and moved by which Harappan nodes are NA.
Fails, all models
H4: debt-relief edicts keep labour stress lower for longer
Mesopotamian Hands Stress was below the comparator in 0 of 3 windows under every model. Design flaw: edicts were issued because labour was under debt strain, so the test cannot separate cause from response. It is still reported as a failure, because that is what the pre-registered test showed.
Depends on scorer
H1c: steppe pastoralists couple Shield and Flow
Shield ranks first under every model, but Flow ranks 3rd under Claude and GPT and 5th under Grok.
Depends on scorer
H5: no ladder (non-state societies do not sit at the low end)
Holds under Claude (Energy ratio 0.92 against a 0.80 cut-off). Exactly on the line under Grok (0.798). Not testable under GPT, which left half of Aboriginal Australia’s nodes NA.

5. What each model counts as evidence

The NA cells, where a model declined to score a node, differ more between models than the scores do. They show what each model treats as evidence.

Node left NA (passes out of total)Claude (15)Grok (4)GPT (3)
Harappan Helm (no visible rulers)1201
Harappan Shield1413
Aboriginal Australia: Helm, Shield, Lore, Archive003 each
Māori Lore and Archive003 each
Longshan Archive (no writing)1543

CAMS was built to read societies without borrowing their own or their rivals’ stories. The models bring stories of their own, about writing, states and what counts as evidence. The NA pattern is where those stories show, and it is why H5 changed verdict between models.

6. Limitations

Agreement between AI models shows the instrument is applied consistently. It does not show the scores are historically true.

7. Next steps

  1. Rubric DT-2. State outright that oral tradition is evidence for Lore and Archive; separate “function demonstrably absent” (score low) from “no evidence” (NA); add guidance on Abstraction for coerced but skilled labour.
  2. Rescore under DT-2 with all three models, to see whether the NA differences and H5 converge.
  3. A fourth model (Kimi), scoring blind from the prompts alone.
  4. More series: Mesopotamia after Ur III, the Shang–Zhou transition, the Late Bronze Age Levant and the post-Harappan Indus. These test H3 beyond Egypt, and a hypothesis formed after seeing the data (H6): recovery is faster where the setting forces re-coordination, as in a single river valley.
  5. Human scorers: specialists in two or three regions scoring their cases on the same rubric, the first test against something other than a language model.

Data and code

All prompts, raw scores, ensembles, scripts and the full run log are in the backstory folder of the wintermute repository. Every raw score file carries a SHA-256 hash.

Kari McKern, ComplexityWorkz. Scoring passes were produced by AI models from their training knowledge; the Claude and Grok passes made no tool or web calls. Open science: specialists are invited to rescore or contest any case, and disagreements will be logged. Contact.