Token Factory

An AI that answers for you, inside your site or app.

Your customers ask and get an answer straight away, at any hour. It runs on our machines in Spain and you only pay for what gets used.

No code needed to try it: activate your welcome credit with your card.

how it works

Your business asks, Spain answers.

The whole round trip takes under a second, and you set nothing up: the machines and the models are on us.

  1. 01

    Your business sends the question

    From your site, your shop or your app. Nothing changes for your customers.

  2. 02

    Our machines think it through

    In a Spanish data centre, on NVIDIA's most powerful chips. Your data never leaves the EU and never trains any AI.

    Spain, EU
  3. 03

    The answer comes straight back

    Written in your tone, at 3am and in August too. Each answer costs a fraction of a cent.

Wiring it into your site or app takes a technical person one afternoon. If you haven't got one, we'll walk you through it.

what people use it for

Work you no longer do by hand.

With the real cost right beside it, on everyday jobs.

clinics, shops, training centres

Answering your customers

Questions about opening hours, bookings, sizes or orders, answered on the spot and in the middle of the night.

500 answers a day≈ 1.43 € / month
accountants, back office

Reading invoices and paperwork

You hand it the photo or the scan and it gives the data back in order: dates, amounts, who invoiced you.

1,000 invoices scanned≈ 0.49 €
marketing, retail

Writing copy in your voice

Product descriptions, emails and posts that sound like your brand, not like a robot.

300 product descriptions≈ 0.03 €

the brains on offer

Pick the model that does your thinking.

All ready to use, each good at something different. You start with the one we suggest and switch whenever you like.

modelinput €/Mtokoutput €/Mtokdetails
GLM 5.3 NVFP4

An all-rounder for writing and coding, with tool use.

ChatCodingReasoningTools
1.206 €3.789 €details
Qwen3.8 27B

Strong at coding and at chaining tasks on its own.

ReasoningJSON output
0.10 €0.30 €details
Qwen3 14B

Thinks before answering, and you decide how much: a balance of quality and spend.

ChatReasoningTools
0.20 €1.00 €details
Qwen3.5 9B

Light and very fast, for simple tasks at high volume.

ChatReasoningTools
0.10 €0.15 €details
ALIA 40B Instruct

Spain's public model, trained in Spanish, Catalan, Basque and Galician.

Chat
0.20 €2.00 €details
Gemma 4 26B A4B

Understands text and images at once without being a heavy model.

ReasoningTools
0.12 €0.30 €details
Nemotron Nano 12B VL

Reads scanned documents and photos: invoices, contracts, forms.

VisionOCR
0.06 €0.24 €details
Qwen3-VL 32B OCR

Document reading for awkward layouts: tables, columns and handwriting.

VisionOCR
0.10 €0.35 €details
Nemotron 3.5 Lightning 30B

Reasons and uses tools while spending very little, built to answer fast.

ReasoningTools
0.12 €0.30 €details
Whisper large-v3

Turns audio into text: meetings, calls, interviews and videos.

Transcription
0.0015 €per minute of audiodetails
Qwen3-TTS CustomVoice

Reads your text aloud in the voice you choose.

TTSVoice cloning
19.00 €per million charactersdetails
Qwen3-TTS Base (clone)

Turns text into speech: the cheapest option in the catalogue.

TTS
5.00 €per million charactersdetails
Fish Speech S2 Pro

Very natural speech, able to imitate the sample you provide.

TTSVoice cloning
8.00 €per million charactersdetails
Voxtral TTS

Ready-made Mistral voices to read your text aloud.

TTS
6.00 €per million charactersdetails
ACE-Step Music

Composes a complete song from lyrics and a style.

Text to musicCover versionsRedo sectionsSplit stemsEdit the song
0.50 €per songdetails
see the full pricing

The price is per million tokens: roughly 750,000 words.

what it costs

Like a prepaid card: you top up and you use it.

You load credit in euros and each answer takes its share. No monthly fees, no lock-in and no nasty surprises: if the credit runs out, it simply stops.

  • You start with €50 on the house: enough for around 520,000 answers.
  • You see the daily spend in your panel, in euros with VAT.
  • The same credit works across the other GPU Flow products.

get a feel for it

How many answers a day would your business give?

It would cost you1.43 €a month

A shop with real traffic. Less than a coffee a month, answering at any hour. Worked out with Qwen3.5 9B, the cheapest model in the catalogue, assuming 500 input tokens and 300 output tokens per answer. A more powerful model costs more, and you can switch whenever you like.

frequently asked

What developers ask before they start.

And if yours isn't here, write to us: a person answers, not a bot.

Do I have to rewrite my code?

No. Token Factory speaks the OpenAI protocol (/v1/chat/completions) and the Anthropic one (/v1/messages). Point the SDK you already use at https://api.gpuflow.ai, paste your key and the rest of your code stays the same.

Which models can I call?

All 16 in the catalogue (DeepSeek, GLM, Qwen, Gemma, NVIDIA, OpenAI, Mistral, Fish Audio and ACE Studio), with a single API key. Need another open source model that fits the GPU? We'll deploy it.

Are there usage limits?

You consume against your prepaid balance: pay per token up to whatever you've topped up, no quotas or minimums. If you need sustained high volume, reach out and we'll raise your account limits.

Do you train on my data?

No. Your prompts and responses are processed to serve your request and are not used to train models. The details are in the privacy policy.

Where is my data processed?

In our datacenter in Spain, on dedicated NVIDIA B200 GPUs. Your data never leaves Europe.

Do you offer an SLA?

We publish availability on the status page. For a contractual SLA (enterprise accounts), reach out and we'll set one up.

Try it today with €50 on the house.

You create the account, add your card, type a question in the panel and see the answer. No code.