developers
Build on GPUFlow.
An API that speaks your SDK, SSH access to your own machine, and a whole GPU if you need one. Everything it takes to wire up each product, organised by product.
base_url https://api.gpuflow.ai/v1
From signing up to your first answer
- 1Create your key in the dashboardYou sign up with 50 € of credit, go to API keys and create one. It is shown only once, so copy it right then.
- 2Point your SDKChange the base URL to https://api.gpuflow.ai/v1 and use your key. The rest of your code stays as it is.
- 3Call a modelAny of the 15 in the catalogue: the text ones with the OpenAI or Anthropic SDK, and the audio and music ones through their own endpoints.
When you create it you choose whether it works for every API or only for Token Factory.
The inference API
Compatible with the OpenAI and Anthropic SDKs. Change the base URL, paste your key, and the rest of your code stays as it is.
Base URL
https://api.gpuflow.ai
Authentication
Authorization: Bearer $GPUFLOW_API_KEY
Include your API key in the Authorization header as a Bearer token. Create and manage keys from your dashboard.
Endpoints
- /v1/chat/completions
- Main endpoint compatible with the OpenAI SDK. Supports streaming, function calling and all catalogue models.
- /v1/messages
- Endpoint compatible with the Anthropic SDK. Same functionality, Anthropic message format.
- /v1/models
- Lists available models with pricing and capabilities.
- /v1/audio/transcriptions
- Upload an audio file and get the text back. Multipart; with response_format=verbose_json it also returns the duration, which is what gets billed.
- /v1/audio/speech
- You send text and a voice, and get the audio back. Billed per character.
- /v1/music/generations
- Generates a whole song. It is asynchronous: it returns an id, you poll the status and download the track when it finishes.
Streaming and reasoning
Streaming (SSE)
Add "stream": true and you get the response over Server-Sent Events, token by token. We inject keep-alives so long connections aren't dropped by intermediate proxies.
Configurable reasoning
The models that reason accept a thinking budget: control it with reasoning.effort ("low", "medium", "high", or "none" to switch it off) or with reasoning.max_tokens. It works from the OpenAI SDK and from the Anthropic one (thinking.budget_tokens); we translate it to the engine for you.
request
{ "model": "glm-5.3-nvfp4",
"messages": [{ "role": "user", "content": "Resuélvelo paso a paso" }],
"stream": true,
"reasoning": { "effort": "high" } }Function calling
5 of the 9 text models support tools, in the standard OpenAI format. You pass your tools array and the model decides when to call them; you run the function and return the result. In the catalogue below they carry the “Tools” capability.
request
{ "model": "glm-5.3-nvfp4",
"messages": [{ "role": "user", "content": "¿Qué tiempo hace hoy?" }],
"tools": [{
"type": "function",
"function": {
"name": "get_weather",
"parameters": {
"type": "object",
"properties": { "city": { "type": "string" } }
}
}
}] }Your first call, in your language
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.gpuflow.ai/v1",
api_key=os.environ["GPUFLOW_API_KEY"]
)
response = client.chat.completions.create(
model="glm-5.3-nvfp4",
messages=[{"role": "user", "content": "Hello!"}],
max_tokens=10000
)
print(response.choices[0].message.content)Error codes
- 401
- Invalid or missing API key
- 402
- Insufficient balance, top up from your dashboard
- 403
- You do not have permission to use this model
- 404
- Model not found
- 429
- Rate limit exceeded, wait and retry (Retry-After header)
- 503
- Model temporarily overloaded, retry with backoff (Retry-After header)
Model catalogue
All models run on our infrastructure in Spain. Prices per million tokens.
| model | input EUR/M | output EUR/M | capabilities |
|---|---|---|---|
| glm-5.3-nvfp4 | 1.206 € | 3.789 € | Chat, Coding, Reasoning, Tools |
| qwen-3.8-27b | 0.10 € | 0.30 € | Reasoning, JSON output |
| qwen-3-14b | 0.20 € | 1.00 € | Chat, Reasoning, Tools |
| qwen-3.5-9b | 0.10 € | 0.15 € | Chat, Reasoning, Tools |
| alia-40b-instruct | 0.20 € | 2.00 € | Chat |
| gemma-4-26b-a4b | 0.12 € | 0.30 € | Reasoning, Tools |
| nemotron-nano-12b-v2-vl | 0.06 € | 0.24 € | Vision, OCR |
| qwen3-vl-32b-ocr | 0.10 € | 0.35 € | Vision, OCR |
| nemotron-3.5-lightning-30b-a3b | 0.12 € | 0.30 € | Reasoning, Tools |
| whisper-large-v3 | 0.0015 € per minute of audio | Transcription | |
| qwen3-tts-customvoice | 19.00 € per million characters | TTS, Voice cloning | |
| qwen3-tts-base | 5.00 € per million characters | TTS | |
| fish-speech-s2-pro | 8.00 € per million characters | TTS, Voice cloning | |
| voxtral-tts | 6.00 € per million characters | TTS | |
| acestep-musicfeatured | 0.50 € per song | Text to music, Cover versions, Redo sections, Split stems, Edit the song | |
Voice models are billed per minute or per character, and music per song. The detail for each one is on its own page. Model catalogue.
Prices in EUR per million tokens, VAT included.
Your machine over SSH
A container with SSH access, its own disk and the coding agents already installed. You reach it through our bastion: no VPN and no tunnels.
1. Generate your key
Three commands and it is done. Do not give it a path with -f: OpenSSH looks for ~/.ssh/id_ed25519 on its own when jumping through the bastion, and a key somewhere else is the usual reason access fails.
# 1. create the keys folder mkdir -p ~/.ssh && chmod 700 ~/.ssh # 2. generate your key (Enter to leave it without a passphrase) ssh-keygen -t ed25519 -C "gpuflow" # 3. print it so you can copy it cat ~/.ssh/id_ed25519.pub
Paste the PUBLIC key, the one ending in .pub. The private one never leaves your machine.
2. Connect
One command, no tunnels and no second terminal. The -J jumps through the bastion, which identifies you with the same key you registered and takes you to your Sandbox.
terminal
ssh -J [email protected] developer@mi-sandbox
3. What the bastion will not forward
Only the tunnel to your Sandbox: no agent forwarding, X11, environment variables or port forwarding from the bastion. And if the Sandbox is paused or deleted it rejects you with a message instead of leaving you waiting on a machine that is not there.
The disk outlives the Sandbox
Your files live on a separate network disk. Pausing the Sandbox does not touch it, and neither does deleting it: the disk stays and you can attach it to a new one.
The API from inside
From the Sandbox, Token Factory is called over the internal network. The key never leaves the perimeter and never travels over the internet.
Agents already installed
Claude Code and OpenCode come set up and pointed at the API. You open the terminal and they already work against your models.
The same, with a GPU of your own
A dedicated GPU is a Sandbox with an exclusive card: same SSH, same disk, same agents. What changes is underneath.
Nothing changes on your side
The connection, the disk and the agents work exactly as in any Sandbox. What you learn above applies here.
The card is yours, even while paused
No neighbours and no queues: video memory is not shared, so a loaded model competes with nobody.
It is activated by talking
It does not provision itself. You pick the configuration, you write to us, and we confirm availability and specs with you before switching it on.
Outside the browser
A CLI for the three products and an MCP server so your agents can use them on their own. Both install from npm and each has its own guide.
CLI
A single binary for inference, account, Sandboxes, GPU and disks: list models, talk to one, check your balance and usage, create and pause Sandboxes, look up free GPU capacity.
read the CLI guideMCP server
Twenty tools so Claude, Cursor or your own agent can talk to GPU Flow: the catalogue and chat, plus read-only access to your balance, your Sandboxes and free GPU capacity. Anything that creates, pauses or deletes sits behind a permission you have to switch on by hand.
read the MCP guideYour first call, today.
An account in a minute and 50 € of credit once you activate your card.