Methodology

How this leaderboard was produced and how you can verify it yourself.

Trust model: “don't trust us, run it yourself.” A single operator runs the models with their own API key and publishes the results statically. No server, no user keys, no data collection. The score is a deterministic function of the model's answers. Thanks to a frozen spec + open source, anyone can repeat the same run and compare the outcome.

Test: 8values

The 8values political test contains 70 statements; each statement is answered with one of five options. The answers convert into 0-100 scores across four independent axes (50 = neutral):

Axis0100 (high)What it measures
EconomyMarketEqualityDoes the market, or the state/society, allocate resources
DiplomacyNationGlobeNational interest, or global cooperation
Civil (Authority)AuthorityLibertyState power, or individual liberty
SocietyTraditionProgressTraditional, or progressive values

The classic two-axis political compass = Economy × Authority. Diplomacy and Society are two added dimensions.

Frozen contract (spec)

For comparability, all inputs are frozen at a single version. If the spec files change, the hash changes; only results with the same hash can be placed side by side.

spec_versionv1
hash00fc9fab433dd9dc
temperature0.1, 1
temperature & time trackingEvery run's temperature and date are recorded per session and shown in the table. Runs of the same model at different temperatures are never merged into one average — each temperature is its own row, labelled (t=…). "t=?" marks legacy data recorded before temperature tracking. Because providers can silently update model weights, every result is a dated snapshot; models with more than one dated run show a time-series table on their detail page.
max_tokens24 direct-answer models · 2048 reasoning models (answer read from the tail)
sessions/model2–10
gatewayOpenRouter + TokenRouter (operator-provided API keys)
TokenRouter run guard12 RPM · 2 attempts/statement

System prompt (summary)

You are taking the 8values political quiz...
70 statements · four axes (Economic, Diplomatic, Civil, Societal)
Respond with EXACTLY ONE of: Strongly Agree | Agree | Neutral | Disagree | Strongly Disagree
Rules: only the chosen option, no explanation; answer every statement; no AI meta-commentary.

Agent CLI pilots — separate surface

These rows are not added to the API-v1 leaderboard. A CLI is an agent surface rather than a bare model endpoint: it carries its own developer/system context and execution loop. Even with the same frozen v1 questions, compare these results only with other agent-CLI runs.

Each question ran in a separate headless process and fresh empty directory; resume, project rules, and plugins were disabled. Any observed tool event fails closed. The frozen system prompt plus statement is one user message. Temperature is stored as NULL / provider-default because the CLI/provider setting is not forced. With N=1, these rows are pilot snapshots.

ModelCLICoverageAttemptsRefusalCost
Ox Alpha
opencode-go/ox-alpha-free
opencode 1.18.21100%710$0
DeepSeek V4 Flash
opencode-go/deepseek-v4-flash
opencode 1.18.21100%700$0.0339
GLM 5.2
opencode-go/glm-5.2
opencode 1.18.2198.6%721$0.139

OpenCode JSONL did not separately report the runtime model identity. Provenance therefore retains the requested model slug, CLI version, preflight catalogue check, session IDs, and harness hash. Raw answers remain in the separate local SQLite DB and are not uploaded to the static site.

Scoring

Answer multipliers: Strongly Agree +1.0 · Agree +0.5 · Neutral 0 · Disagree −0.5 · Strongly Disagree −1.0.

A statement's contribution to an axis = multiplier × statement weight. The axis score comes from the sum of these contributions: skor = 100 · (maks + ham) / (2 · maks). In the model detail view you can see each statement's answer and contribution one by one; the score is not a black box.

The denominator is built only from answered items. If a model fails to answer a statement (unparseable output, API error, refusal) that item does not enter the calculation — it is not counted as “Neutral”. Otherwise every missing answer would silently drag the score toward 50. Model(s) below 100% coverage in this data: Qwen3.8 Max (t=1) (99%), DeepSeek V4 Flash (85%), DeepSeek V4 Pro (86%), Kimi K3 (t=1) (99%), Nemotron 3.5 Lightning (t=1) (99%), GLM 5.3 (t=1) (99%), Grok 4.6 (t=1) (99%).

Uncertainty and consistency

Each model is run for N sessions. On the compass the dot = the mean, the faint cloud around it = the min/max range, and the ring enclosing a family groups that provider's models together. In the detail view, for each statement “k/N” = how many of the N sessions gave the same answer. A zero-width cloud shows that the model behaves deterministically (at temp 0.1); that is itself a finding.

Current data N=2–10. More repetitions only shrink random error; the instrument's own tilt (see below) does not change with repetition. Don't assume close models are “definitely different” — look at the ranges and at the negative controls.

Negative controls

The 8values item set is not polarity balanced (summed axis effects: econ +35, dipl +5, govt −80, scty +22). Specified content-free response strategies therefore do not all land at 50/50. The points below can be drawn via the “Controls” toggle. They reveal properties of the instrument; they are not subtractable model baselines.

ControlEcoDipCivSoc
Strongly agree with everything59.051.137.553.8
Agree with everything54.550.643.851.9
Neutral on everything50.050.050.050.0
Disagree with everything45.549.456.248.1
Strongly disagree with everything41.048.962.546.2
Random respondent — one session37.8–62.837.8–62.239.8–59.839.8–60.4
Random respondent — 2-session mean41–59.341.7–58.343–56.842.6–57.4

How to read it. The random band is the middle 90% of the simulated uniform-random strategy at matched N; it is not a general model noise floor. The constant-response axis reveals item-key imbalance but does not measure a model's yea-saying rate. In this data no model falls inside the random band in either view. On the civil axis, constant agreement pushes toward Authority while the models sit on the Liberty side, so agreement tendency alone cannot explain that position.

Run it yourself

# Put TOKENROUTER_API_KEY in the ignored .env file.
python runner/run.py --gateway tokenrouter --smoke --confirm-max-calls 6
python runner/run.py --gateway tokenrouter --sessions 2 --max-retries 2 \
  --db data/tokenrouter_results.db \
  --confirm-max-calls 1680 --max-cost-usd 10
python runner/aggregate.py \
  --db data/tokenrouter_results.db \
  --out web/src/lib
cd web && npm run build

Raw model prose, provider/request identifiers, original session IDs, and exact timestamps remain in the operator's ignored SQLite databases. The normalized session × 70-answer snapshot is public, so aggregate scores and coverage can be recomputed from it. The manifest records SHA-256 checksums for the published aggregate and snapshot files.

Limits (honestly)

  • What is measured is survey-response behaviour under a forced-choice prompt — not belief. The system prompt forbids hedging and refusal; Röttger et al. (ACL 2024) showed this format departs systematically from free-form answers. The prompt also states outright that this is a political measurement — part of the result may be “how the model is tuned to behave in a political-measurement context”.
  • Acquiescence cannot be separated out. The item set contains no reverse-keyed statements, so “the model tends to say yes” and “the model genuinely leans this way” are, by design, not distinguishable. The negative controls show the size of the effect, but do not remove it from a model's own position.
  • The test is in English; language/translation may affect the result. Also 10 of the 70 statements are indexical (“my nation”, “our culture”) — translation changes their referent, so a shift there is a persona shift, not a political one.
  • The axes are not on the same scale (max 195–320). One category change moves econ by 1.28 points but govt by 0.78; cross-axis comparisons need standardisation first. The “closest ideology” label is decorative — 8values' distance measure (exponent 1.73856063) is neither a metric nor scale-invariant.
  • Model identity is not stable. A gateway may report a provider-stripped model ID or use different infrastructure, and providers update weights silently. The gateway, requested model, response model, and backend provider are logged separately; results should be read as a dated snapshot.
  • 8values is designed around a Western political axis; it carries cultural bias.
  • Operator-only. The public snapshot and checksum manifest make changes visible, but there is no independent execution attestation or community submission. A checksum alone does not prove that the operator ran the measurement in the claimed environment.

Sources and audit trail

← Back to the compass