Reading Societies Without Their Stories
A blind, pre-registered, multi-model test of CAMS on ancient collapse and recovery
CAMS reads a society through eight functions, not through the story it tells about itself. Deep time is a hard place to test that claim. The record is thin, often written by victors or later traditions, and easy to fill with modern ideas about states and progress. So we gave three AI model families the same blind rubric, had each score the same twenty ancient society-windows in fresh contexts, and checked five hypotheses that were written down before any scoring began.
What all three models agreed on
- Collapse drains everything at once. In both the Egyptian and the Aegean breakdowns, every function lost surplus (Capacity − Stress) and shared thinking (Abstraction × Coherence) together.
- The Aegean recovered far more slowly than it fell. Egypt rebuilt to near Old Kingdom levels in under three centuries. Under Claude and Grok, the Aegean was still below its pre-collapse level 500 years later.
- Egypt gave an early warning; the Aegean did not. Coherence between rulers (Helm) and meaning-makers (Lore) fell for two centuries in Egypt while surplus was still positive. In the Aegean it stayed flat until the breakdown.
- Debt-relief edicts did not coincide with lower labour stress in Mesopotamia. This test has a design flaw (below), but it was pre-registered and it failed.
The models also disagreed, and that disagreement is a result in itself. They differ in scale and, more importantly, in what they accept as evidence. The one hypothesis about non-state societies (“no ladder”) passed under Claude, sat exactly on the cut-off under Grok, and could not be tested under GPT.
1. What was tested
CAMS scores eight functional nodes: Helm (command), Shield (force and order), Lore (religion and meaning), Stewards (owners of stores and land), Craft (skilled production), Hands (mass labour), Archive (records and memory, written or oral) and Flow (exchange). Each node gets four scores from 1 to 10: Coherence, Capacity, Stress and Abstraction. Two composites carry the argument:
- Energy = Capacity − Stress: whether a function has surplus to work with.
- Cognition = Abstraction × Coherence: how much shared symbolic thinking it does.
Cases
Prompt A scored twelve societies at one anchor year each: Longshan China, Shang Anyang, the Mature Harappan cities, Old Kingdom Egypt, First Intermediate Egypt, Hammurabi’s Babylon, Neopalatial Crete, Mycenaean Greece (LH IIIB), the Aegean after the palace destructions, pre-contact Aboriginal Australia, Māori Aotearoa and the Xiongnu. Prompt B added eight windows to build two collapse-and-recovery series (Egypt and the Aegean) and one Mesopotamian series. It rescored the two collapses as overlap anchors.
Safeguards
- Evidence gate. A node is scored only if the scorer can name an indicator of performance, an indicator of strain and a mechanism tying them to that node. Otherwise it is NA. Rubric: CAMS RAW SCORER v1.2-OPT with a deep-time adapter (DT-1).
- Source posture. Conquerors’ accounts, later traditions and colonial ethnography count as reports, not as evidence of function.
- Blind passes. Every pass ran in a fresh context that saw only the prompt: no project files, no hypotheses, no other scores, no knowledge that other passes existed.
- Pre-registration. Hypotheses H1–H5 were logged before scoring. The series tests were written into a script whose SHA-256 hash (
353aaeab…) was logged before the extra windows were scored.
| Model | Passes per prompt | Conditions |
|---|---|---|
| Claude (Opus 5.5) | 15 | Three independent runs of 5; fresh subagent per pass; prompts A and B always separate |
| Grok 4.7 | 4 | One fresh conversation per prompt, prompt file only |
| GPT (ChatGPT Work/Codex, version unverified) | 3 | Fresh isolated agent per pass; prompts A and B in the same context |
An early non-blind Claude pass and a Grok pass that had read the hypotheses are kept in the run log but excluded from the analysis. One upload turned out to be a byte-identical copy of an earlier run and was not counted.
2. Is the instrument stable?
Within one model, yes. Between models, the order of societies mostly holds, but the scale does not.
| Comparison | Mean absolute difference per score (A / B) | Rank agreement of societies (A / B) | Mean offset |
|---|---|---|---|
| Claude run vs Claude run (5-pass ensembles) | 0.14–0.23 / 0.11–0.22 | 0.97–0.99 / 1.00 | under ±0.15 |
| Single Claude pass vs other Claude runs | 0.26–0.29 | — | — |
| Grok (4) vs Claude (15) | 0.39–0.53 / 0.41–0.57 | 0.92 / 1.00 | +0.2 to +0.47 |
| GPT (3) vs Claude (15) | 0.89–1.64 / 0.84–1.80 | 0.81 / 0.99 | +0.6 to +1.2; Abstraction +1.6 to +1.8 |
All three models were about equally self-consistent: the spread between passes of one model was 0.12–0.44 points per dimension. So the gaps between models are not noise from single passes; each model applies the rubric in its own consistent way. The ten time windows in prompt B were ranked almost identically by every model (0.98–1.00). The twelve societies in prompt A were ranked less consistently (0.81–0.92), and the shifts fall on societies without visible rulers or without writing.
Scale offsets matter less than they look. Most tests compare windows within one model, so a model that scores everything a point higher still reaches the same verdict.
3. Collapse and recovery
The two collapses look alike on the way down and very different on the way back.
Mean Energy (Capacity − Stress) across the eight nodes, by window
Ensemble means: Claude 15 passes, Grok 4, GPT 3. Years BCE. Below zero, functions are under more strain than they have capacity for. Both panels use the same scale.
Show the data as a table
The early-warning signal (H3) is a separate measure: the mean Coherence of Helm and Lore. In Egypt it fell from about 8 to between 5 and 6.5 across the two pre-collapse windows under every model, while Energy was still positive. In the Aegean it stayed flat until the breakdown. Egypt’s collapse began as a loss of shared purpose at the top. The Aegean’s struck a system that still looked coherent at the top.
| Series | Fall, Energy per century | Recovery, Energy per century | Helm–Lore Coherence before breakdown |
|---|---|---|---|
| Egypt (Claude) | −2.86 | +2.56 | 8.1 → 5.1 |
| Egypt (Grok) | −2.02 | +2.01 | 8.0 → 5.4 |
| Egypt (GPT) | −1.75 | +1.39 | 8.0 → 6.5 |
| Aegean (Claude) | −7.27 | +1.95 | 5.5 → 5.7 |
| Aegean (Grok) | −6.19 | +1.27 | 6.2 → 6.5 |
| Aegean (GPT) | −4.17 | +1.59 | 6.8 → 6.8 |
Two patterns at node level are worth noting. First, in both collapses Lore holds up best: it has the highest Node Value in First Intermediate Egypt under all three models. The lowest nodes are Helm, Hands and Flow in Egypt, and Archive and Helm in the Aegean: meaning-making outlasts command and administration. Second, Hands (mass labour) is the lowest or second-lowest node in almost every state window under every model, while non-state societies (Aboriginal Australia, Māori Aotearoa) do not show that gap. The state’s stress lands on its labourers.
4. Hypothesis results
Four of the eight tested claims gave the same verdict under all three models: three held and one (H4) failed. Two held in direction only, and two depended on the scorer.
5. What each model counts as evidence
The NA cells, where a model declined to score a node, differ more between models than the scores do. They show what each model treats as evidence.
| Node left NA (passes out of total) | Claude (15) | Grok (4) | GPT (3) |
|---|---|---|---|
| Harappan Helm (no visible rulers) | 12 | 0 | 1 |
| Harappan Shield | 14 | 1 | 3 |
| Aboriginal Australia: Helm, Shield, Lore, Archive | 0 | 0 | 3 each |
| Māori Lore and Archive | 0 | 0 | 3 each |
| Longshan Archive (no writing) | 15 | 4 | 3 |
- Claude hides Harappa’s rulers. It accepts the uniform weights, bricks and seals as evidence of records and craft, but finds nothing tying them to a ruling centre. Grok assigns that coordination to Helm.
- GPT does not accept oral tradition. It leaves Lore and Archive unscored for every society without writing. Claude and Grok score Aboriginal Lore and Archive at 7–8, treating songlines and law as working institutions. The rubric says oral records count, but evidently not loudly enough for every model. This single choice is the largest disagreement in the data set.
- “Absent” and “unknown” get confused. When a function demonstrably disappeared, such as literacy in the Aegean after the palaces, some passes score it 1 and others NA. These mean different things, and the rubric does not yet separate them.
CAMS was built to read societies without borrowing their own or their rivals’ stories. The models bring stories of their own, about writing, states and what counts as evidence. The NA pattern is where those stories show, and it is why H5 changed verdict between models.
6. Limitations
Agreement between AI models shows the instrument is applied consistently. It does not show the scores are historically true.
- Shared corpus. All three models learned from overlapping literatures; where those share a bias, the models will agree on it. No human specialist has scored any case yet.
- Unequal samples. Claude has 15 passes per prompt, Grok 4, GPT 3. The model behind the third Claude run is inferred from its scores, not confirmed.
- Unequal isolation. GPT scored both prompts in one context per pass, so its overlap windows are less independent.
- Single anchors and thin series. Each society is one representative year standing for 50–250 years. H2 and H3 rest on two collapse series, and H4 on three Mesopotamian windows.
- Choices made within the project. Test details such as the 0.80 cut-off in H5 were written after the hypotheses and after the first non-blind pass, though before any blind scoring.
- No reasoning shown. The blind prompt forbids explanation to keep scorers independent, so disagreements can be located but not yet explained from the scorers’ own reasons.
7. Next steps
- Rubric DT-2. State outright that oral tradition is evidence for Lore and Archive; separate “function demonstrably absent” (score low) from “no evidence” (NA); add guidance on Abstraction for coerced but skilled labour.
- Rescore under DT-2 with all three models, to see whether the NA differences and H5 converge.
- A fourth model (Kimi), scoring blind from the prompts alone.
- More series: Mesopotamia after Ur III, the Shang–Zhou transition, the Late Bronze Age Levant and the post-Harappan Indus. These test H3 beyond Egypt, and a hypothesis formed after seeing the data (H6): recovery is faster where the setting forces re-coordination, as in a single river valley.
- Human scorers: specialists in two or three regions scoring their cases on the same rubric, the first test against something other than a language model.
Data and code
All prompts, raw scores, ensembles, scripts and the full run log are in the backstory folder of the wintermute repository. Every raw score file carries a SHA-256 hash.
- Working paper (PDF), with the complete window-level and node-level tables
- Run log: every run, decision and deviation, newest first
- Pre-registered hypotheses
- Node-level means for all three models (CSV, 480 rows)
- The two blind prompts, to run against any model yourself (see TESTING.md)
Kari McKern, ComplexityWorkz. Scoring passes were produced by AI models from their training knowledge; the Claude and Grok passes made no tool or web calls. Open science: specialists are invited to rescore or contest any case, and disagreements will be logged. Contact.