Answering your customers
Questions about opening hours, bookings, sizes or orders, answered on the spot and in the middle of the night.
Token Factory
Your customers ask and get an answer straight away, at any hour. It runs on our machines in Spain and you only pay for what gets used.
No code needed to try it: activate your welcome credit with your card.
how it works
The whole round trip takes under a second, and you set nothing up: the machines and the models are on us.
From your site, your shop or your app. Nothing changes for your customers.
In a Spanish data centre, on NVIDIA's most powerful chips. Your data never leaves the EU and never trains any AI.
Spain, EUWritten in your tone, at 3am and in August too. Each answer costs a fraction of a cent.
Wiring it into your site or app takes a technical person one afternoon. If you haven't got one, we'll walk you through it.
what people use it for
With the real cost right beside it, on everyday jobs.
Questions about opening hours, bookings, sizes or orders, answered on the spot and in the middle of the night.
You hand it the photo or the scan and it gives the data back in order: dates, amounts, who invoiced you.
Product descriptions, emails and posts that sound like your brand, not like a robot.
the brains on offer
All ready to use, each good at something different. You start with the one we suggest and switch whenever you like.
| model | input €/Mtok | output €/Mtok | details | |
|---|---|---|---|---|
GLM 5.3 NVFP4 An all-rounder for writing and coding, with tool use. ChatCodingReasoningTools | ChatCodingReasoningTools | 1.206 € | 3.789 € | details |
Qwen3.8 27B Strong at coding and at chaining tasks on its own. ReasoningJSON output | ReasoningJSON output | 0.10 € | 0.30 € | details |
Qwen3 14B Thinks before answering, and you decide how much: a balance of quality and spend. ChatReasoningTools | ChatReasoningTools | 0.20 € | 1.00 € | details |
Qwen3.5 9B Light and very fast, for simple tasks at high volume. ChatReasoningTools | ChatReasoningTools | 0.10 € | 0.15 € | details |
ALIA 40B Instruct Spain's public model, trained in Spanish, Catalan, Basque and Galician. Chat | Chat | 0.20 € | 2.00 € | details |
Gemma 4 26B A4B Understands text and images at once without being a heavy model. ReasoningTools | ReasoningTools | 0.12 € | 0.30 € | details |
Nemotron Nano 12B VL Reads scanned documents and photos: invoices, contracts, forms. VisionOCR | VisionOCR | 0.06 € | 0.24 € | details |
Qwen3-VL 32B OCR Document reading for awkward layouts: tables, columns and handwriting. VisionOCR | VisionOCR | 0.10 € | 0.35 € | details |
Nemotron 3.5 Lightning 30B Reasons and uses tools while spending very little, built to answer fast. ReasoningTools | ReasoningTools | 0.12 € | 0.30 € | details |
Whisper large-v3 Turns audio into text: meetings, calls, interviews and videos. Transcription | Transcription | 0.0015 €per minute of audio | details | |
Qwen3-TTS CustomVoice Reads your text aloud in the voice you choose. TTSVoice cloning | TTSVoice cloning | 19.00 €per million characters | details | |
Qwen3-TTS Base (clone) Turns text into speech: the cheapest option in the catalogue. TTS | TTS | 5.00 €per million characters | details | |
Fish Speech S2 Pro Very natural speech, able to imitate the sample you provide. TTSVoice cloning | TTSVoice cloning | 8.00 €per million characters | details | |
Voxtral TTS Ready-made Mistral voices to read your text aloud. TTS | TTS | 6.00 €per million characters | details | |
ACE-Step Music Composes a complete song from lyrics and a style. Text to musicCover versionsRedo sectionsSplit stemsEdit the song | Text to musicCover versionsRedo sectionsSplit stemsEdit the song | 0.50 €per song | details | |
The price is per million tokens: roughly 750,000 words.
what it costs
You load credit in euros and each answer takes its share. No monthly fees, no lock-in and no nasty surprises: if the credit runs out, it simply stops.
get a feel for it
It would cost you1.43 €a month
A shop with real traffic. Less than a coffee a month, answering at any hour. Worked out with Qwen3.5 9B, the cheapest model in the catalogue, assuming 500 input tokens and 300 output tokens per answer. A more powerful model costs more, and you can switch whenever you like.
frequently asked
And if yours isn't here, write to us: a person answers, not a bot.
No. Token Factory speaks the OpenAI protocol (/v1/chat/completions) and the Anthropic one (/v1/messages). Point the SDK you already use at https://api.gpuflow.ai, paste your key and the rest of your code stays the same.
All 16 in the catalogue (DeepSeek, GLM, Qwen, Gemma, NVIDIA, OpenAI, Mistral, Fish Audio and ACE Studio), with a single API key. Need another open source model that fits the GPU? We'll deploy it.
You consume against your prepaid balance: pay per token up to whatever you've topped up, no quotas or minimums. If you need sustained high volume, reach out and we'll raise your account limits.
No. Your prompts and responses are processed to serve your request and are not used to train models. The details are in the privacy policy.
In our datacenter in Spain, on dedicated NVIDIA B200 GPUs. Your data never leaves Europe.
We publish availability on the status page. For a contractual SLA (enterprise accounts), reach out and we'll set one up.
You create the account, add your card, type a question in the panel and see the answer. No code.