What we are offering.
A dedicated inference endpoint is a GPU reserved for one customer, serving an open-weight language model of their choice — up to 8B-class, for example Qwen3-8B — behind a private, OpenAI-compatible API. Existing SDKs work unchanged: change the base URL, keep the code. Streaming and token-usage reporting are included, and every endpoint comes with a data processing agreement. Prompts and completions are processed on our own hardware in Dalsbruk, Finland — not routed through third-party clouds, and never used for training.
Why flat pricing.
We own the site and the hardware, and the site runs on renewable power — so we can price an endpoint like infrastructure rather than like tokens: one flat monthly rate per dedicated GPU, electricity included, no token metering, no overage charges. Pilot pricing starts at €95 for a 14-day evaluation (credited if you continue on a monthly plan) and €295 per month per dedicated GPU. Details are on the AI Inference page.
Who this is for.
The pilot is sized for organisations whose constraint is not raw scale but data control: medical and professional practices with strict confidentiality requirements, SMEs in the EU that want an internal assistant or document pipeline without handing their archive to a hyperscaler, and engineering teams that need a fixed, predictable endpoint for development and staging. An 8B-class model running around the clock covers the everyday text work of a business — drafting, extraction, summarisation, internal Q&A.
How to join.
Tell us your model and workload via the AI Inference page — we come back with a concrete proposal, typically within two business days. Evaluation endpoints are the usual starting point: fourteen days on a dedicated GPU against your own workload, fully credited if you continue.