Analysis · API v1 · spec 00fc9fab

A shared response regime

Across 22 API-tested models, all 20 non-xAI rows occupy the same equality–liberty region—even at their observed session minima. xAI is the economic exception. Release drift is provider-specific, temperature has no common ideological direction, and headquarters is a weak explanation.

20/20 non-xAI models in the cluster 22 models at T=0.1 10 providers 2–10 sessions/model 2 collection dates 6 matched temperature pairs

This is a cross-sectional study of forced-choice questionnaire responses, not a measurement of beliefs. “Left” below means a score above 50 on 8values’ Equality axis; “liberal” means above 50 on its Liberty axis. “Shared” describes clustering in this snapshot—not proof that the models converged over time.

01 / core result

The shared cluster spans providers and headquarters

At T=0.1, 20/20 non-xAI models score above 50 on both Equality and Liberty. The result survives a simple session-range check: every non-xAI model’s observed session minimum remains above 50 on both axes.

That cross-provider pattern is the main result. It describes a shared output region under one instrument and prompt; it does not show that model beliefs exist or that the models moved toward one another over time.

The two xAI models sit apart economically. Grok 4.3 scores 37.9; Grok 4.6 scores 49.4 and its range crosses the midpoint. The gap from Grok 4.6 to the lowest non-xAI model is 11.8 points.

Equality × liberty, T=0.1

Range bars show the observed session minimum and maximum.

Equality and liberty scores for 22 modelsAll 20 non-xAI models are above 50 on both axes. Grok 4.3 and Grok 4.6 are separated from the cluster on the economic axis.40506070809045505560657075Market ← economic score → EqualityAuthority ← civil score → LibertyQwen3.6 Max: Equality 71.2, Liberty 65.4Qwen3.7 Max: Equality 68.8, Liberty 63.6Qwen3.7 Plus: Equality 62.8, Liberty 64.8Qwen3.8 Max (t=0.1): Equality 72.4, Liberty 65.8Claude Haiku 4.5: Equality 67.3, Liberty 60.2Claude Opus 4.8: Equality 72.2, Liberty 64.1DeepSeek V3.2: Equality 69.4, Liberty 57.5DeepSeek V4 Flash: Equality 72.9, Liberty 65.4DeepSeek V4 Pro: Equality 74.2, Liberty 65.5Gemini 2.5 Flash: Equality 83.0, Liberty 64.3Gemini 3 Flash: Equality 70.8, Liberty 65.3Gemini 3.7 Flash (t=0.1): Equality 61.2, Liberty 57.8Gemini 3.7 FlashLlama 3.3 70B: Equality 70.3, Liberty 59.8Llama 4 Maverick: Equality 75.0, Liberty 64.8Llama 4 Scout: Equality 78.5, Liberty 59.6Kimi K3 (t=0.1): Equality 68.6, Liberty 67.6Nemotron 3.5 Lightning (t=0.1): Equality 69.2, Liberty 57.0GPT-4o: Equality 70.3, Liberty 59.0GPT-5: Equality 81.9, Liberty 67.2GLM 5.3 (t=0.1): Equality 62.5, Liberty 64.3Grok 4.3: Equality 37.9, Liberty 62.9Grok 4.3Grok 4.6 (t=0.1): Equality 49.4, Liberty 63.5Grok 4.6
All non-xAI means and their observed session minima remain in the equality–liberty quadrant.

02 / qualified signal

Some release lines drift left-liberal; the dataset cannot establish a law

5 of 6 name-ordered endpoint comparisons move toward Equality, 5 move toward Liberty, and 5 move toward both. OpenAI is the cleanest large shift in the expected direction. Google is a large counterexample: its endpoint moves 21.8 points toward Market and 6.5 toward Authority.

The important limitation is structural. The dataset has no release-date, family or tier fields. Some endpoints were measured on different dates, through different gateways, with different session counts. This is evidence worth following, but it is not yet a release-trend design.

Named release endpoints

Later-labelled endpoint minus earlier-labelled endpoint; positive means Equality or Liberty.

Provider / sequenceΔ EqualityΔ LibertyComparability
GoogleGemini 2.5 Flash → 3 Flash → 3.7 Flash-21.8-6.5The last endpoint changes date, gateway and N.
xAIGrok 4.3 → 4.6+11.5+0.6The endpoint changes date, gateway and N.
OpenAIGPT-4o → GPT-5+11.6+8.2Same date, gateway and N; still a cross-model comparison.
MetaLlama 3.3 → Llama 4 siblings+6.5+2.4The Llama 4 endpoint is the mean of two different siblings.
DeepSeekV3.2 → V4 variants+4.2+8.0V4 variants answered only 85–86% of the instrument.
AlibabaQwen3.6 Max → 3.8 Max+1.2+0.4Non-monotonic series; the endpoint changes date and gateway.
Rows are descriptive comparisons inferred from product names, not a fitted time trend. Sibling averages are used for Meta and DeepSeek.

03 / no common effect

Temperature does not have an ideology

For the six matched models, raising temperature from 0.1 to 1.0 moves 3 toward Equality and 3 toward Market. It moves 4 toward Liberty and 2 toward Authority. Only 2 of six move left and liberal at the same time. No model changes quadrant.

The means are small on the two relevant axes—+0.3 on Equality and +0.4 on Liberty—but “small” is not the same as “inside a published noise floor.” This experiment has no validated ±2 threshold, and each condition has only two sessions.

Temperature deltas by model

T=1.0 minus T=0.1. Bars point toward the named pole.

Both conditions used TokenRouter with N=2. T=0.1 was measured on 24 August; T=1.0 on 25 August, so temperature is also confounded with day.

04 / weak grouping

Headquarters does not explain the shared cluster

A model-weighted country average quietly gives providers with more rows more votes, so the comparison below first averages models within each provider and then averages providers. With xAI included, the country groups are nearly tied on Equality. Remove xAI and the US mean jumps from 67.5 to 72.3—reversing the tempting left/right story.

The more persistent descriptive difference is elsewhere: the four China-headquartered providers score higher on Globe and Liberty in this sample. Four versus six providers, with unmatched model tiers, is too small to turn that into a national-culture claim.

Provider-weighted means

Each provider receives equal weight, regardless of how many models it contributes.

Provider headquartersEqualityGlobeLibertyProgress
China-headquartered providers4 providers · 9 models68.064.864.969.7
US-headquartered providers6 providers · 13 models67.559.661.669.9
US providers, excluding xAI5 providers · 11 models72.361.461.271.1
“China” and “US” refer only to provider headquarters. They do not identify training data, user market, deployment policy or a model’s political provenance.

05 / measurement discipline

Instrument imbalance limits the ideological claim

Agreeing with every statement scores 54.5 on Equality; strongly agreeing scores 59.0. Those controls show that item polarity is not balanced, so a tendency to agree could contribute to the cluster. They do not mean that “the first 59 points are free,” that 59 is a noise floor, or that 59 can be subtracted from a model score.

A constant-response control is a point produced by one response strategy. It does not measure how agreeable any model was. To make acquiescence part of the thesis, the study must compute a per-model agreement rate and use balanced or reverse-keyed item pairs.

What this snapshot cannot identify

  • Beliefs, attitudes or stable political identities.
  • A causal effect of temperature; the two settings were run on different days.
  • A universal release trend without release dates, matched tiers and a common run protocol.
  • A national effect from four China-headquartered and six US-headquartered providers.
  • Fully comparable DeepSeek V4 scores while 14–15% of answers are missing.

The follow-up that would answer the question

  1. Pre-register families, tiers, release order and exclusion rules.
  2. Run every historical release on the same day, gateway and prompt.
  3. Use at least 20 sessions per model and temperature; randomise run order.
  4. Estimate provider and family effects separately from country grouping.
  5. Report item-level agreement, missingness and uncertainty—not only compass coordinates.

Evidence table

The 22 rows behind the claims

This table keeps the protocol differences visible. It excludes the seven T=1 rows and the separate CLI-agent pilot.

ModelProviderEq.Lib.NDateGatewayCoverage
Gemini 2.5 FlashGoogle83.064.3102026-06-21OpenRouter100%
GPT-5OpenAI81.967.2102026-06-21OpenRouter100%
Llama 4 ScoutMeta78.559.6102026-06-21OpenRouter100%
Llama 4 MaverickMeta75.064.8102026-06-21OpenRouter100%
DeepSeek V4 ProDeepSeek74.265.5102026-06-21OpenRouter86%
DeepSeek V4 FlashDeepSeek72.965.4102026-06-21OpenRouter85%
Qwen3.8 MaxAlibaba72.465.822026-08-24TokenRouter100%
Claude Opus 4.8Anthropic72.264.1102026-06-21OpenRouter100%
Qwen3.6 MaxAlibaba71.265.422026-06-21OpenRouter100%
Gemini 3 FlashGoogle70.865.3102026-06-21OpenRouter100%
Llama 3.3 70BMeta70.359.8102026-06-21OpenRouter100%
GPT-4oOpenAI70.359.0102026-06-21OpenRouter100%
DeepSeek V3.2DeepSeek69.457.5102026-06-21OpenRouter100%
Nemotron 3.5 LightningNVIDIA69.257.022026-08-24TokenRouter100%
Qwen3.7 MaxAlibaba68.863.6102026-06-21OpenRouter100%
Kimi K3Moonshot AI68.667.622026-08-24TokenRouter100%
Claude Haiku 4.5Anthropic67.360.2102026-06-21OpenRouter100%
Qwen3.7 PlusAlibaba62.864.822026-06-21OpenRouter100%
GLM 5.3Z.ai62.564.322026-08-24TokenRouter100%
Gemini 3.7 FlashGoogle61.257.822026-08-24TokenRouter100%
Grok 4.6xAI49.463.522026-08-24TokenRouter100%
Grok 4.3xAI37.962.9102026-06-21OpenRouter100%