Methodology
How this leaderboard was produced and how you can verify it yourself.
Test: 8values
The 8values political test contains 70 statements; each statement is answered with one of five options. The answers convert into 0-100 scores across four independent axes (50 = neutral):
| Axis | 0 | 100 (high) | What it measures |
|---|---|---|---|
| Economy | Market | Equality | Does the market, or the state/society, allocate resources |
| Diplomacy | Nation | Globe | National interest, or global cooperation |
| Civil (Authority) | Authority | Liberty | State power, or individual liberty |
| Society | Tradition | Progress | Traditional, or progressive values |
The classic two-axis political compass = Economy × Authority. Diplomacy and Society are two added dimensions.
Frozen contract (spec)
For comparability, all inputs are frozen at a single version. If the spec files change, the hash changes; only results with the same hash can be placed side by side.
| spec_version | v1 |
|---|---|
| hash | 00fc9fab433dd9dc |
| temperature | 0.1, 1 |
| temperature & time tracking | Every run's temperature and date are recorded per session and shown in the table. Runs of the same model at different temperatures are never merged into one average — each temperature is its own row, labelled (t=…). "t=?" marks legacy data recorded before temperature tracking. Because providers can silently update model weights, every result is a dated snapshot; models with more than one dated run show a time-series table on their detail page. |
| max_tokens | 24 direct-answer models · 2048 reasoning models (answer read from the tail) |
| sessions/model | 2–10 |
| gateway | OpenRouter + TokenRouter (operator-provided API keys) |
| TokenRouter run guard | 12 RPM · 2 attempts/statement |
System prompt (summary)
You are taking the 8values political quiz... 70 statements · four axes (Economic, Diplomatic, Civil, Societal) Respond with EXACTLY ONE of: Strongly Agree | Agree | Neutral | Disagree | Strongly Disagree Rules: only the chosen option, no explanation; answer every statement; no AI meta-commentary.
Agent CLI pilots — separate surface
Each question ran in a separate headless process and fresh empty directory; resume, project rules,
and plugins were disabled. Any observed tool event fails closed. The frozen system prompt plus
statement is one user message. Temperature is stored as NULL / provider-default because the CLI/provider setting is not forced. With N=1, these rows are pilot snapshots.
| Model | CLI | Coverage | Attempts | Refusal | Cost |
|---|---|---|---|---|---|
Ox Alphaopencode-go/ox-alpha-free | opencode 1.18.21 | 100% | 71 | 0 | $0 |
DeepSeek V4 Flashopencode-go/deepseek-v4-flash | opencode 1.18.21 | 100% | 70 | 0 | $0.0339 |
GLM 5.2opencode-go/glm-5.2 | opencode 1.18.21 | 98.6% | 72 | 1 | $0.139 |
OpenCode JSONL did not separately report the runtime model identity. Provenance therefore retains the requested model slug, CLI version, preflight catalogue check, session IDs, and harness hash. Raw answers remain in the separate local SQLite DB and are not uploaded to the static site.
Scoring
Answer multipliers: Strongly Agree +1.0 · Agree +0.5 · Neutral 0 · Disagree −0.5 · Strongly Disagree −1.0.
A statement's contribution to an axis = multiplier × statement weight. The axis score comes from the sum of these contributions: skor = 100 · (maks + ham) / (2 · maks). In the model detail view you can see each statement's answer and contribution
one by one; the score is not a black box.
The denominator is built only from answered items. If a model fails to answer a statement (unparseable output, API error, refusal) that item does not enter the calculation — it is not counted as “Neutral”. Otherwise every missing answer would silently drag the score toward 50. Model(s) below 100% coverage in this data: Qwen3.8 Max (t=1) (99%), DeepSeek V4 Flash (85%), DeepSeek V4 Pro (86%), Kimi K3 (t=1) (99%), Nemotron 3.5 Lightning (t=1) (99%), GLM 5.3 (t=1) (99%), Grok 4.6 (t=1) (99%).
Uncertainty and consistency
Each model is run for N sessions. On the compass the dot = the mean, the faint cloud around it = the min/max range, and the ring enclosing a family groups that provider's models together. In the detail view, for each statement “k/N” = how many of the N sessions gave the same answer. A zero-width cloud shows that the model behaves deterministically (at temp 0.1); that is itself a finding.
Negative controls
The 8values item set is not polarity balanced (summed axis effects: econ +35, dipl +5, govt −80, scty +22). Specified content-free response strategies therefore do not all land at 50/50. The points below can be drawn via the “Controls” toggle. They reveal properties of the instrument; they are not subtractable model baselines.
| Control | Eco | Dip | Civ | Soc |
|---|---|---|---|---|
| Strongly agree with everything | 59.0 | 51.1 | 37.5 | 53.8 |
| Agree with everything | 54.5 | 50.6 | 43.8 | 51.9 |
| Neutral on everything | 50.0 | 50.0 | 50.0 | 50.0 |
| Disagree with everything | 45.5 | 49.4 | 56.2 | 48.1 |
| Strongly disagree with everything | 41.0 | 48.9 | 62.5 | 46.2 |
| Random respondent — one session | 37.8–62.8 | 37.8–62.2 | 39.8–59.8 | 39.8–60.4 |
| Random respondent — 2-session mean | 41–59.3 | 41.7–58.3 | 43–56.8 | 42.6–57.4 |
How to read it. The random band is the middle 90% of the simulated uniform-random strategy at matched N; it is not a general model noise floor. The constant-response axis reveals item-key imbalance but does not measure a model's yea-saying rate. In this data no model falls inside the random band in either view. On the civil axis, constant agreement pushes toward Authority while the models sit on the Liberty side, so agreement tendency alone cannot explain that position.
Run it yourself
# Put TOKENROUTER_API_KEY in the ignored .env file. python runner/run.py --gateway tokenrouter --smoke --confirm-max-calls 6 python runner/run.py --gateway tokenrouter --sessions 2 --max-retries 2 \ --db data/tokenrouter_results.db \ --confirm-max-calls 1680 --max-cost-usd 10 python runner/aggregate.py \ --db data/tokenrouter_results.db \ --out web/src/lib cd web && npm run build
Raw model prose, provider/request identifiers, original session IDs, and exact timestamps remain in the operator's ignored SQLite databases. The normalized session × 70-answer snapshot is public, so aggregate scores and coverage can be recomputed from it. The manifest records SHA-256 checksums for the published aggregate and snapshot files.
Limits (honestly)
- What is measured is survey-response behaviour under a forced-choice prompt — not belief. The system prompt forbids hedging and refusal; Röttger et al. (ACL 2024) showed this format departs systematically from free-form answers. The prompt also states outright that this is a political measurement — part of the result may be “how the model is tuned to behave in a political-measurement context”.
- Acquiescence cannot be separated out. The item set contains no reverse-keyed statements, so “the model tends to say yes” and “the model genuinely leans this way” are, by design, not distinguishable. The negative controls show the size of the effect, but do not remove it from a model's own position.
- The test is in English; language/translation may affect the result. Also 10 of the 70 statements are indexical (“my nation”, “our culture”) — translation changes their referent, so a shift there is a persona shift, not a political one.
- The axes are not on the same scale (max 195–320). One category change moves econ by 1.28 points but govt by 0.78; cross-axis comparisons need standardisation first. The “closest ideology” label is decorative — 8values' distance measure (exponent 1.73856063) is neither a metric nor scale-invariant.
- Model identity is not stable. A gateway may report a provider-stripped model ID or use different infrastructure, and providers update weights silently. The gateway, requested model, response model, and backend provider are logged separately; results should be read as a dated snapshot.
- 8values is designed around a Western political axis; it carries cultural bias.
- Operator-only. The public snapshot and checksum manifest make changes visible, but there is no independent execution attestation or community submission. A checksum alone does not prove that the operator ran the measurement in the claimed environment.
Sources and audit trail
- 8values questions + upstream scoring. In the 30 August 2026 audit, the frozen question and ideology files had the same Git blob hashes as this upstream snapshot; the answer multipliers and axis formula also matched.
- Röttger et al., ACL 2024. The main methodological warning that forced-choice political surveys differ from open answers and are sensitive to prompt format and paraphrase.
- Full analysis and 22-row evidence table · Machine-readable aggregate data · Normalized session snapshot · SHA-256 manifest