back

GLM 5.2 FP8

glm-5.2-fp8 · Z.ai

754B FP8 on 8× B200. Generalist and coding with tool calling.

Served on NVIDIA B200 GPUs · Spain (EU)

Model source

Hugging Facezai-org/GLM-5.2-FP8

Specifications

Parameters
743B · 39B active MoE
Hardware
8× B200
Context
202,752 tokens
Quantization
FP8
Modalities
text
Capabilities
chat · coding · tools

Pricing

input

1.00

output

3.00

cache read

0.19

€ / million tokens · VAT excluded

Example (curl)

Compatible with the OpenAI SDK. Change the base URL and API key.

curl https://api.gpuflow.ai/v1/chat/completions \
  -H "Authorization: Bearer $GPUFLOW_API_KEY" \
  -d '{ "model": "glm-5.2-fp8",
        "messages": [{ "role": "user", "content": "Hola" }] }'

Try it with 5 € free

Create your account, point your SDK at GPU Flow and call this model on NVIDIA B200 GPUs, in Spain.