All the AI infrastructureyou need.No expertise required.

Token Factory for the API, Sandbox to build in, and dedicated GPUs for the heavy lifting.One balance, in euros, in a European data centre.

Activate with your card and you're up in under a minute. Use it for three minutes, pay for three minutes.

  • cert.ISO 27001 and ENS Medium
  • data inSpain, EU
  • invoicein EUR
  • computeNVIDIA B200

products

Combine all the AI infrastructure you need.

Inference with the API, agents in the Sandbox, heavy workloads on a dedicated GPU, separately or combined. All on a single balance, one invoice.

Token Factory

An AI that answers, via API

Connect your website, shop or application to the best open models. Compatible with the OpenAI and Anthropic SDKs.

Real example

a clinic answering appointment questions at any hour.

  • TTFT is how long the AI takes to start answering. Here, under a tenth of a second.
  • NVIDIA's most advanced compute chip. It is the engine that runs the AI.
  • You top up your balance in advance and it is drawn down by actual usage. No fees, no surprises.
from 0.06 €/ M input tokenscredit from 10 €

Sandbox

Your workspace in the cloud

A powerful machine switched on for you, with agent templates ready. It suspends itself when you stop using it, and your files stay put.

Real example

a team that builds and tests without depending on its own hardware.

  • A very low-latency internal network between your Sandbox and the models: requests never reach the internet.
  • The cluster's storage. It holds large repositories and datasets without slowing you down.
  • If you stop using it, it shuts down on its own and stops spending your balance. You choose after how long.
from 0.15 €/ active hour+ volume from €5/month

Dedicated GPU

A whole machine for you

A full NVIDIA B200 or a MIG slice, shared with nobody. By the month or by the year, for serious workloads.

Real example

a company training its own model on private data.

  • MIG splits a GPU into independent slices. You rent a slice or the whole machine.
  • NVIDIA's most advanced compute chip. It is the engine that runs the AI.
  • The GPU's memory: the more there is, the larger the model that fits. FP4 is the number format that lets it be served faster.
from 1.10 €/ hourmonthly reserved −10%

The path

From the first experiment to production.

One balance covers all three stages, with no need to switch providers.

  1. 01

    Test the idea

    Token Factory

    Call a model from your app through the API. The free €50 goes a long way.

  2. 02

    Build it

    Sandbox

    A persistent environment to build in, with agent templates ready to go.

  3. 03

    Take it to production

    Dedicated GPU

    Dedicated power for your real workload, by the month or by the year.

Two questions

Which one do I need?

Two questions. Whether you run infrastructure for a living or have never touched AI before, we will tell you where to start.

1. What do you want to achieve?
2. How are you with the technical side?

What's new

The newest models on the market, already running on our cluster.

ReleasedVersionModelWhat it is forMore
GLM-5.3-NVFP4GLM 5.3 NVFP4An all-rounder for writing and coding, with tool use.See model
Qwen3.8-27B-NVFP4Qwen3.8 27BStrong at coding and at chaining tasks on its own.See model
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4Nemotron 3.5 Lightning 30BReasons and uses tools while spending very little, built to answer fast.See model
ALIA-40b-instruct-2606ALIA 40B InstructSpain's public model, trained in Spanish, Catalan, Basque and Galician.See model
Gemma-4-26B-A4B-NVFP4Gemma 4 26B A4BUnderstands text and images at once without being a heavy model.See model
Qwen3.5-9BQwen3.5 9BLight and very fast, for simple tasks at high volume.See model

infrastructure

Sovereign hardware, no shared tenant.

Six reasons to bring your AI here: what hardware runs it, where it sits, how it connects, where your data lives, which certifications it holds and how it is invoiced.

01NVIDIA B200

NVIDIA B200

Latest-gen accelerators, dedicated, on bare metal. No shared tenant.

02Spain

Spain

Data center on Spanish soil. No international transfers, no sub-processors outside the EU.

03RDMA internal network

RDMA internal network

Low latency between nodes for multi-GPU and sustained throughput in production.

04DDN EXAScaler

DDN EXAScaler

High-speed parallel storage for datasets, checkpoints and persistent volumes.

05ISO 27001 and ENS Medium

ISO 27001 and ENS Medium

Operational certifications. Standard DPA ready to sign. Optional zero retention for enterprise.

06Spanish invoice in EUR

Spanish invoice in EUR

Deductible for EU companies and freelancers, valid for intra-EU reverse charge. No USD/EUR conversion, no month-end FX surprises.

models

Open source models, deployed and ready.

Open models from several providers, all deployed on our own infrastructure in Spain.

Support

It all happens here, with real people behind it.

You send an email and a person replies, within four working hours. For critical incidents, a call with an expert.

Write to us
Machines in Spain
A data centre on Spanish soil, with no copies outside the EU.
Invoiced in euros
Ready for your accountant, with VAT included and no currency conversion.
ISO 27001 and ENS Media
Downloadable certifications and a DPA ready to sign.
Human support
In Spanish and English, in under 4 working hours.

frequently asked questions

What people almost always ask before starting.

Where is my data processed?

All data is processed at a datacenter in Spain. It never leaves the European Union and is not replicated to other regions. The platform is ISO 27001 and ENS Medium certified, and we have a standard DPA ready to sign.

Is it compatible with the OpenAI or Anthropic SDK?

Yes. The API exposes /v1/chat/completions (OpenAI) and /v1/messages (Anthropic). Point your SDK at api.gpuflow.ai with your API key and the catalog models are available directly. We support SSE streaming and function calling where the model allows it.

Do I need an account to try? Do I need a card?

Account yes, 1 minute with email or SSO. Add your card and get 50 € of credit to call the API or spin up a Sandbox. The 1 € verification charge is refunded.

How does billing work?

Prepaid model. Top up credit when you want and consume by real usage, measured in tokens (Token Factory) or active hours (Sandboxes). Automatic notification when balance drops below 20%. Deductible Spanish invoice on every top-up.

What happens if the environment goes idle?

We only charge for the hours the sandbox is active. After the idle time you configured (from 10 min to 4 h, or manual stop only) with no SSH activity or user processes, it suspends automatically and stops billing compute. Your persistent volume survives intact and is available again in 5-8 seconds on reconnect.

Can I cancel or delete my sandbox?

Yes, anytime from the dashboard. Billing is prorated to the exact moment of cancellation. If you delete, the persistent volume and all snapshots are removed; you have up to 7 days to undo the deletion if it was accidental.

Do you have a public SLA?

Yes, we publish real-time status at gpuflow.ai/status. The contractual SLA with support response times is agreed in enterprise plans.

What does it mean that GPU Flow is in beta?

That the service is already in production and billed on real usage, but we are still polishing the platform: there may be the occasional interruption and the interface changes from one week to the next. Prices, your balance and data residency in Spain do not change because of the beta. The state of every service is published on the status page, and if something breaks you write to us at [email protected] and we sort it out.

Your AI infrastructure, sovereign and in euros.

Data that never leaves Europe, billing in euros and one balance across API, Sandbox and GPU. Create an account, add your card and start with €50 free.

Building a European startup? Apply for up to €5,000 in credits