Analysis · API v1 · spec 00fc9fab
A shared response regime
Across 22 API-tested models, all 20 non-xAI rows occupy the same equality–liberty region—even at their observed session minima. xAI is the economic exception. Release drift is provider-specific, temperature has no common ideological direction, and headquarters is a weak explanation.
This is a cross-sectional study of forced-choice questionnaire responses, not a measurement of beliefs. “Left” below means a score above 50 on 8values’ Equality axis; “liberal” means above 50 on its Liberty axis. “Shared” describes clustering in this snapshot—not proof that the models converged over time.
01 / core result
The shared cluster spans providers and headquarters
At T=0.1, 20/20 non-xAI models score above 50 on both Equality and Liberty. The result survives a simple session-range check: every non-xAI model’s observed session minimum remains above 50 on both axes.
That cross-provider pattern is the main result. It describes a shared output region under one instrument and prompt; it does not show that model beliefs exist or that the models moved toward one another over time.
The two xAI models sit apart economically. Grok 4.3 scores 37.9; Grok 4.6 scores 49.4 and its range crosses the midpoint. The gap from Grok 4.6 to the lowest non-xAI model is 11.8 points.
Equality × liberty, T=0.1
Range bars show the observed session minimum and maximum.
02 / qualified signal
Some release lines drift left-liberal; the dataset cannot establish a law
5 of 6 name-ordered endpoint comparisons move toward Equality, 5 move toward Liberty, and 5 move toward both. OpenAI is the cleanest large shift in the expected direction. Google is a large counterexample: its endpoint moves 21.8 points toward Market and 6.5 toward Authority.
The important limitation is structural. The dataset has no release-date, family or tier fields. Some endpoints were measured on different dates, through different gateways, with different session counts. This is evidence worth following, but it is not yet a release-trend design.
Named release endpoints
Later-labelled endpoint minus earlier-labelled endpoint; positive means Equality or Liberty.
| Provider / sequence | Δ Equality | Δ Liberty | Comparability |
|---|---|---|---|
| GoogleGemini 2.5 Flash → 3 Flash → 3.7 Flash | -21.8 | -6.5 | The last endpoint changes date, gateway and N. |
| xAIGrok 4.3 → 4.6 | +11.5 | +0.6 | The endpoint changes date, gateway and N. |
| OpenAIGPT-4o → GPT-5 | +11.6 | +8.2 | Same date, gateway and N; still a cross-model comparison. |
| MetaLlama 3.3 → Llama 4 siblings | +6.5 | +2.4 | The Llama 4 endpoint is the mean of two different siblings. |
| DeepSeekV3.2 → V4 variants | +4.2 | +8.0 | V4 variants answered only 85–86% of the instrument. |
| AlibabaQwen3.6 Max → 3.8 Max | +1.2 | +0.4 | Non-monotonic series; the endpoint changes date and gateway. |
03 / no common effect
Temperature does not have an ideology
For the six matched models, raising temperature from 0.1 to 1.0 moves 3 toward Equality and 3 toward Market. It moves 4 toward Liberty and 2 toward Authority. Only 2 of six move left and liberal at the same time. No model changes quadrant.
The means are small on the two relevant axes—+0.3 on Equality and +0.4 on Liberty—but “small” is not the same as “inside a published noise floor.” This experiment has no validated ±2 threshold, and each condition has only two sessions.
Temperature deltas by model
T=1.0 minus T=0.1. Bars point toward the named pole.
04 / weak grouping
Headquarters does not explain the shared cluster
A model-weighted country average quietly gives providers with more rows more votes, so the comparison below first averages models within each provider and then averages providers. With xAI included, the country groups are nearly tied on Equality. Remove xAI and the US mean jumps from 67.5 to 72.3—reversing the tempting left/right story.
The more persistent descriptive difference is elsewhere: the four China-headquartered providers score higher on Globe and Liberty in this sample. Four versus six providers, with unmatched model tiers, is too small to turn that into a national-culture claim.
Provider-weighted means
Each provider receives equal weight, regardless of how many models it contributes.
| Provider headquarters | Equality | Globe | Liberty | Progress |
|---|---|---|---|---|
| China-headquartered providers4 providers · 9 models | 68.0 | 64.8 | 64.9 | 69.7 |
| US-headquartered providers6 providers · 13 models | 67.5 | 59.6 | 61.6 | 69.9 |
| US providers, excluding xAI5 providers · 11 models | 72.3 | 61.4 | 61.2 | 71.1 |
05 / measurement discipline
Instrument imbalance limits the ideological claim
Agreeing with every statement scores 54.5 on Equality; strongly agreeing scores 59.0. Those controls show that item polarity is not balanced, so a tendency to agree could contribute to the cluster. They do not mean that “the first 59 points are free,” that 59 is a noise floor, or that 59 can be subtracted from a model score.
A constant-response control is a point produced by one response strategy. It does not measure how agreeable any model was. To make acquiescence part of the thesis, the study must compute a per-model agreement rate and use balanced or reverse-keyed item pairs.
What this snapshot cannot identify
- Beliefs, attitudes or stable political identities.
- A causal effect of temperature; the two settings were run on different days.
- A universal release trend without release dates, matched tiers and a common run protocol.
- A national effect from four China-headquartered and six US-headquartered providers.
- Fully comparable DeepSeek V4 scores while 14–15% of answers are missing.
The follow-up that would answer the question
- Pre-register families, tiers, release order and exclusion rules.
- Run every historical release on the same day, gateway and prompt.
- Use at least 20 sessions per model and temperature; randomise run order.
- Estimate provider and family effects separately from country grouping.
- Report item-level agreement, missingness and uncertainty—not only compass coordinates.
Evidence table
The 22 rows behind the claims
This table keeps the protocol differences visible. It excludes the seven T=1 rows and the separate CLI-agent pilot.
| Model | Provider | Eq. | Lib. | N | Date | Gateway | Coverage |
|---|---|---|---|---|---|---|---|
| Gemini 2.5 Flash | 83.0 | 64.3 | 10 | 2026-06-21 | OpenRouter | 100% | |
| GPT-5 | OpenAI | 81.9 | 67.2 | 10 | 2026-06-21 | OpenRouter | 100% |
| Llama 4 Scout | Meta | 78.5 | 59.6 | 10 | 2026-06-21 | OpenRouter | 100% |
| Llama 4 Maverick | Meta | 75.0 | 64.8 | 10 | 2026-06-21 | OpenRouter | 100% |
| DeepSeek V4 Pro | DeepSeek | 74.2 | 65.5 | 10 | 2026-06-21 | OpenRouter | 86% |
| DeepSeek V4 Flash | DeepSeek | 72.9 | 65.4 | 10 | 2026-06-21 | OpenRouter | 85% |
| Qwen3.8 Max | Alibaba | 72.4 | 65.8 | 2 | 2026-08-24 | TokenRouter | 100% |
| Claude Opus 4.8 | Anthropic | 72.2 | 64.1 | 10 | 2026-06-21 | OpenRouter | 100% |
| Qwen3.6 Max | Alibaba | 71.2 | 65.4 | 2 | 2026-06-21 | OpenRouter | 100% |
| Gemini 3 Flash | 70.8 | 65.3 | 10 | 2026-06-21 | OpenRouter | 100% | |
| Llama 3.3 70B | Meta | 70.3 | 59.8 | 10 | 2026-06-21 | OpenRouter | 100% |
| GPT-4o | OpenAI | 70.3 | 59.0 | 10 | 2026-06-21 | OpenRouter | 100% |
| DeepSeek V3.2 | DeepSeek | 69.4 | 57.5 | 10 | 2026-06-21 | OpenRouter | 100% |
| Nemotron 3.5 Lightning | NVIDIA | 69.2 | 57.0 | 2 | 2026-08-24 | TokenRouter | 100% |
| Qwen3.7 Max | Alibaba | 68.8 | 63.6 | 10 | 2026-06-21 | OpenRouter | 100% |
| Kimi K3 | Moonshot AI | 68.6 | 67.6 | 2 | 2026-08-24 | TokenRouter | 100% |
| Claude Haiku 4.5 | Anthropic | 67.3 | 60.2 | 10 | 2026-06-21 | OpenRouter | 100% |
| Qwen3.7 Plus | Alibaba | 62.8 | 64.8 | 2 | 2026-06-21 | OpenRouter | 100% |
| GLM 5.3 | Z.ai | 62.5 | 64.3 | 2 | 2026-08-24 | TokenRouter | 100% |
| Gemini 3.7 Flash | 61.2 | 57.8 | 2 | 2026-08-24 | TokenRouter | 100% | |
| Grok 4.6 | xAI | 49.4 | 63.5 | 2 | 2026-08-24 | TokenRouter | 100% |
| Grok 4.3 | xAI | 37.9 | 62.9 | 10 | 2026-06-21 | OpenRouter | 100% |