back

pricing

You only pay what you use. In euros, no surprises.

Token Factory charges per token consumed, Sandboxes per active hour plus the persistent volume, and dedicated GPU per hour or monthly reserved. No commitment: the only fixed cost is the persistent volume.

01

Token Factory

Per-million-tokens prices, same for all customers. Pay against prepaid credit, no commitment, no minimums.

start
modelinput · €/Mtokoutput · €/Mtokdetails
DeepSeek V4 Pro

1.6T MoE on 8× B200. Full-power reasoning.

ChatCodingReasoningTools
0.802.60details
GLM 5.2 NVFP4

743B NVFP4 on 8× B200. Generalist and coding with tool calling.

ChatCodingReasoningToolsLong context
1.003.00details
DeepSeek V4 Flash

1.6T MoE on 2× B200. High throughput, 524K context.

ChatReasoningTools
0.140.28details
Qwen3.6 27B

27B dense. Strong at coding and agents, with tool calling.

ChatReasoningToolsOCR
0.302.40details
Qwen3 14B

14B FP8. Reasoning with a configurable thinking budget.

ChatReasoningTools
0.201.00details
Qwen3.5 9B

9B ultra-light. Fast and cost-effective, with tool calling.

ChatReasoningTools
0.100.15details
Gemma 4 26B A4B

Multimodal MoE, ~3.8B active. Vision, reasoning and tools.

ReasoningTools
0.120.30details
Nemotron Nano 12B VL

Vision-language model for OCR and document understanding. 128K context, up to 4 images per prompt.

image-text-to-textOCR
0.060.24details
Qwen3-VL 32B OCR

Vision-language model for OCR and document understanding. 64K context, up to 4 images per prompt.

image-text-to-textOCR
0.100.35details
Nemotron 3.5 Lightning 30B

MoE + Mamba-2 hybrid, ~3B active. Reasoning and tool calling, 128K context.

ReasoningTools
0.120.30details
See the model catalogue
02

Sandboxes

On-demand sandboxes: compute is paid per active hour and the sandbox auto-suspends after the idle time you configure, stopping the meter. The persistent volume is separate and always billed — it is your minimum monthly spend.

configure sandbox

Starter

0.15/ active hour

no upfront payment · compute per active hour

  • 4 vCPU
  • 8 GB RAM
  • configurable auto-suspend (10 min – 4 h or manual)

Pro

0.30/ active hour

no upfront payment · compute per active hour

  • 8 vCPU
  • 32 GB RAM
  • configurable auto-suspend (10 min – 4 h or manual)

Team

0.55/ active hour

no upfront payment · compute per active hour

  • 16 vCPU
  • 64 GB RAM
  • configurable auto-suspend (10 min – 4 h or manual)
03

Dedicated GPU

Dedicated NVIDIA B200 GPU, MIG slice or full card. Hourly, or reserved monthly with a 10% discount.

see Dedicated GPU

MIG B200 · 23 GB

723 €/mo

reserved · €1.10 per hour

  • ~23 GB of dedicated VRAM
  • 8 vCPU · 48 GB RAM included

MIG B200 · 45 GB

1,281 €/mo

reserved · €1.95 per hour

  • ~45 GB of dedicated VRAM
  • 16 vCPU · 96 GB RAM included

MIG B200 · 90 GBsold out

2,365 €/mo

reserved · €3.60 per hour

  • ~90 GB of dedicated VRAM

B200 completa

4,271 €/mo

reserved · €6.50 per hour

  • 180 GB of dedicated VRAM
  • 32 vCPU · 384 GB RAM included

Availability for each configuration is confirmed when you contact us.