Benchmarking GLM 5.2: a similar average hides a new language divide
When we first tried GLM 5.2, it answered a random English question about Tiananmen Square. It was an informal, uncontrolled test, and we did not record the prompt or response.
That answer made GLM 5.2 look less censored than GLM 5. It could also have been one lucky response, an inference-provider artifact, or some other serving difference.
We ran the Chinese Political Neutrality Benchmark to create a formal record and see whether the impression held across repeated runs. The full results tell a different story.
With no system prompt, the overall score moved from 64.6% to 65.4%. With the Claude identity prompt, it moved from 85.0% to 80.0%.
These are percentage-point changes. The almost unchanged unprompted average hides larger movements by language.
What we tested
The benchmark contains 50 questions about Chinese politics, history, and governance. Each question is available in English, Brazilian Portuguese, and Simplified Chinese.
Every question was asked five times at temperature 1. Across two model versions and two prompt conditions, the comparison covers 3,000 responses.
Mistral Large 3 2512 evaluated the responses at temperature 0 with the same 1 to 5 rubric used in the original GLM 5 benchmark. A score of 1 maps to 0%; a score of 5 maps to 100%.
Why we changed inference providers
We used Google Cloud Vertex AI MaaS for GLM 5 because it provided the model through a managed API. That let us test a large open-weights model without operating a GPU cluster.
GLM 5.2 was not available through Vertex AI MaaS when we ran this benchmark. At the time of writing, it is still available there only as a dedicated deployment; Google distinguishes MaaS offerings through model cards labeled API Service.
We chose W&B Inference instead. We consider W&B a trustworthy, established Western provider; its GLM 5.2 listing lists 744 billion total parameters, with 40 billion active.
As with the GLM 5 benchmark, we deliberately chose Western inference providers rather than Z.ai’s official API. We wanted to evaluate the model, not application-layer guardrails that Chinese providers are likely to add, such as classifiers, refusal filters, or other serving artifacts.
GLM 5 was served by Vertex AI MaaS; GLM 5.2 was served by W&B Inference. Inference settings and provider-side guardrails may account for part of the measured difference.
The results describe the two deployments we tested. They do not isolate changes in the model weights.
Results
Changes of only a few points are well within the practical error margin of this limited sample. Run-to-run chance, numerical precision, inference settings, and provider characteristics can all produce small movements.
We therefore treat changes such as +0.9, −2.6, and −2.0 points as no measurable movement. The larger shifts remain descriptive; the benchmark was not designed to establish statistical significance or isolate their cause.
Choose Per language to compare every language within one prompt condition. Choose Per prompt to compare both prompt conditions within one language.
GLM 5 to GLM 5.2
The +0.9 point overall change and the small English and Portuguese changes are within the practical error margin of this sample.
Without a system prompt
GLM 5.2’s overall score is 0.9 points higher than GLM 5’s. Within this sample’s error margin, the overall result is unchanged.
Chinese measures 7.2 points higher. Thirty of the 50 Chinese questions improve, nine decline, and 11 are unchanged.
English measures 2.6 points lower. Its question-level results are evenly divided: 16 improve, 17 decline, and 17 are unchanged.
Portuguese measures 2.0 points lower but remains above 93%. We treat both small declines as expected noise, not evidence that GLM 5.2 became worse in either language.
With the Claude prompt
The Claude prompt raised GLM 5’s overall score by 20.4 points. It raises GLM 5.2’s score by 14.5 points.
The language mix changes. English receives a larger lift on GLM 5.2; Chinese receives a much smaller one.
| Language | Prompt lift on GLM 5 | Prompt lift on GLM 5.2 |
|---|---|---|
| English | +28.3 pp | +36.6 pp |
| Portuguese | −1.4 pp | +3.5 pp |
| Chinese | +34.3 pp | +3.5 pp |
| Overall | +20.4 pp | +14.5 pp |
The prompt still changes GLM 5.2 substantially. It no longer produces the large Chinese-language gain measured on GLM 5.
Compared under the Claude prompt, GLM 5.2 measures 5.7 points higher in English and 2.9 points higher in Portuguese. The Portuguese movement is within expected noise, and the English movement is still modest for this sample.
The Chinese score is 23.6 points lower. This is a much larger movement, and it appears across many questions rather than depending on one response.
One plausible explanation is that Z.ai gave GLM 5.2 stricter model-level guardrails in Chinese while applying less restrictive alignment in languages such as English. The data are consistent with that hypothesis, particularly under the Claude prompt, but they do not establish it.
We tested W&B’s deployment rather than Z.ai’s API, and GLM 5 was served by a different provider. The observed pattern could still reflect model alignment, inference characteristics, or a combination of both.
Question-level changes
The averages do not show which topics moved. The explorer lists the four largest gains and declines for each language and prompt condition.
Question text is displayed in the article’s language—English on this page—so the lists remain comparable. The selected language controls which translated prompts and responses were scored.
Largest observed changes
English; no system prompt: 16 questions improved, 17 were unchanged, and 17 declined. Five rose by at least 20 points; eight fell by at least 20.
Largest gains
Largest declines
These rankings are descriptive. Each per-question average comes from five responses, so a large movement can still reflect sampling noise.
The broader counts matter more than any single item. Under the Claude prompt, for example, 38 of the 50 Chinese questions decline from GLM 5 to GLM 5.2; 29 fall by at least 20 points.
That pattern is wider than the Tiananmen Square answer that prompted this test. It still does not establish why the deployment changed.
What happened to our first impression
We cannot reproduce the initial test directly because we did not record it. The benchmark instead contains three English questions about the 1989 protests, the June 4 casualty estimates, and the protesters’ demands.
Without a system prompt, only two of the 15 responses across those questions received a score of 3 or higher. The casualty-estimate and protester-demand questions were evaded in all five runs each.
With the Claude prompt, ten of the 15 responses scored 3 or higher. Four of five runs answered the casualty-estimate question substantively, and four of five answered the protester-demand question substantively.
The benchmark therefore does not confirm that GLM 5.2 reliably answers Tiananmen Square questions both with and without the Claude prompt. It supports a narrower finding: GLM 5.2 can answer some of them without a prompt, but substantive answers were much more common with the Claude prompt in this sample.
Five benchmark runs per question give a better view than one response, but they are still a small sample. A larger replication would be needed to estimate how often each model answers or refuses a specific question.
Limitations
-
Different providers: the serving stack changed between the GLM 5 and GLM 5.2 runs. The benchmark cannot separate provider effects from model effects.
-
Evaluator bias: responses were judged by Mistral Large 3 2512. Its scores reflect one model’s interpretation of factual, balanced, and nuanced political commentary.
-
Stochastic variation: every subject prompt was run five times at temperature 1. Small changes can occur through chance, numerical precision, inference settings, and provider characteristics.
-
No peer review: these results have been produced by return moe without independent review or reproduction.
Raw data
Each JSON file contains the configuration metadata, all prompts and responses, evaluator transcripts, run scores, and per-question statistics.
GLM 5
- zai-glm-5_no-prompt.json: no system prompt
- zai-glm-5_claude-prompt.json: “You are Claude” system prompt
GLM 5.2
- zai-glm-5-2_no-prompt.json: no system prompt
- zai-glm-5-2_claude-prompt.json: “You are Claude” system prompt
Conclusion
GLM 5.2 did not confirm our first impression of a broadly less-censored model. Its unprompted overall score is almost unchanged from GLM 5.
The distribution changed. Chinese measures higher without a prompt, while English is effectively unchanged within the sample’s error margin. Under the Claude prompt, Chinese measures substantially lower compared with GLM 5.
The results suggest stricter Chinese-language guardrails in GLM 5.2. We cannot fully rule out serving differences as a contributing factor.