Neural Nations · Open Science · CAMS-CAN

Help Us Break This — Or Confirm It

CAMS-CAN is a framework for scoring how well nations, companies, and cities are coordinating internally. It's published as open science because it's meant to be checked, not taken on faith. This page is a standing invitation to do exactly that: replicate a scoring pass, audit the rules, or find where it breaks.

01 · Replicate

Run your own pass

Download the DIY kit, score an entity you know well, and compare it against a published dataset. Disagreement is useful data — tell us where and why.

02 · Audit

Read the rules, not just the output

The evidence gate, the change gate, the NA discipline, and the Node Value / Bond Strength formulas are all public. Find a logical gap or a case the rules don't actually cover.

03 · Stress-test

Pick a case built to embarrass it

An entity you expect the framework to score badly — a fast-moving crisis, an obscure polity, an edge case in the corporate or city mapping. Report exactly what happened, not just that it "felt wrong."

04 · Re-measure

Replicate the reliability study

Run a multi-pass ensemble across model providers on a panel we haven't tested, and report ICC or an equivalent agreement statistic back. See the numbers we already have below, so you're extending them, not guessing at the bar.

Two independent panels have been run under the current scoring rubric, each as 18 formal scoring passes nested across three model families (Claude, Grok, Kimi). Full detail, caveats, and the formulas are on the DIY kit page; the framework's broader limitations are on Validation & Limits. The summary:

Reliability by panel — both under the same three-family, 18-pass design
PanelPromptICC(2,k) absoluteICC(3,k) consistency
Australia, 2020–2025v1.1
0.710
0.885
United States, 2020–2025v1.2-OPT
0.691
0.907
What this doesn't establish — and where a review would help most:
  • Historical criterion validity — whether scores match actual institutional history, independent of model agreement.
  • Generalisability beyond two country panels and a small company/city smoke test.
  • A controlled, same-panel comparison of the two prompt versions above (country and prompt version changed together in what we've run).

We'd rather get one reproducible finding than ten impressions. A critique is easiest to act on when it's structural and specific.

Useful

  • "Node X scored Y on entity Z, year W — here's the concrete evidence it contradicts."
  • "The evidence gate doesn't cover this case: [specific scenario]."
  • Raw pass CSVs or ICC/agreement statistics from an independent run.
  • A formula edge case (e.g. clamp behaviour, an undefined ratio) with the exact input that breaks it.

Hard to act on

  • "This doesn't feel right for [country]" with no node, year, or dimension named.
  • Disagreement with the framework's premise rather than its execution — worth discussing, but not a validation finding.
  • A single ungrounded pass treated as disproof — the rubric expects multi-pass ensembles for exactly this reason.

Get the same four files used in every review above — project instructions and two blank schema templates, ready to paste into a Claude Project.

Get the DIY Kit
Kari McKern — Research Inquiries
Findings, replications, and formal critiques welcome by email
[email protected]

Or open an issue or pull request directly against the source: github.com/KaliBond/wintermute. Everything here — framework, datasets, scoring tools, and this page — is licensed as open science under Common Property terms: fork it, run it, publish what you find.