GLM 5.2 FP8
glm-5.2-fp8 · Z.ai
754B FP8 on 8× B200. Generalist and coding with tool calling.
Served on NVIDIA B200 GPUs · Spain (EU)
Model source
Hugging Facezai-org/GLM-5.2-FP8Specifications
- Parameters
- 743B · 39B active MoE
- Hardware
- 8× B200
- Context
- 202,752 tokens
- Quantization
- FP8
- Modalities
- text
- Capabilities
- chat · coding · tools
Pricing
input
1.00 €
output
3.00 €
cache read
0.19 €
€ / million tokens · VAT excluded
Example (curl)
Compatible with the OpenAI SDK. Change the base URL and API key.
curl https://api.gpuflow.ai/v1/chat/completions \
-H "Authorization: Bearer $GPUFLOW_API_KEY" \
-d '{ "model": "glm-5.2-fp8",
"messages": [{ "role": "user", "content": "Hola" }] }'Try it with 5 € free
Create your account, point your SDK at GPU Flow and call this model on NVIDIA B200 GPUs, in Spain.