Open-weight language models behind an OpenAI-compatible API, on GPU infrastructure we own and operate in Dalsbruk. A pilot programme with deliberately limited capacity: one customer per GPU, data residency in Finland, a direct line to the engineers who run the hardware.
No shared queues, no per-token surprises. You get a dedicated inference endpoint on hardware reserved for you alone.
A dedicated GPU serving one open-weight model of your choice (up to 8B-class, e.g. Qwen3-8B) behind an OpenAI-compatible API. Works with every standard SDK — change one base URL, keep your code.
Prompts and completions are processed on our own hardware at the Dalsbruk site — not routed through third-party clouds. Data processing agreement included; your data is never used for training.
€295 per month per dedicated GPU endpoint — electricity included, usage reporting included. Monthly invoice, no minimum term during the pilot phase. See pricing.
Your endpoint runs at our renewable-powered Dalsbruk facility. Electricity is included in the flat price — our cost to manage, not your risk.
Because we own the site and the hardware, we price an endpoint like infrastructure, not like tokens: one flat monthly rate per dedicated GPU, electricity from our renewable-powered Dalsbruk site included. These are pilot rates — customers who join during the pilot keep them.
A real dedicated endpoint to test against your own workload — not a shared demo.
Your own endpoint on hardware reserved for you alone. This is where pilots start.
More throughput or failover: additional dedicated GPUs for the same customer.
Prices exclude VAT, applied according to your billing country. A dedicated GPU is yours around the clock: no token caps, no overage charges. Pilot capacity is limited — endpoints are allocated first come, first served.
An 8B-class model, running around the clock on hardware reserved for you, is sized for the everyday text work of a business — private by construction. These are the scenarios we size pilots for:
Draft referral letters, visit notes and patient correspondence — under a DPA, on hardware where the data never leaves Finland.
Let staff query your own documents and get sourced answers — without handing your archive to a third-party cloud.
Classify incoming requests, extract intent and draft consistent replies — around the clock, at a flat cost.
Turn invoices, forms and reports into clean, structured data your systems can process. Token volume is never a cost factor.
Meeting notes, contracts and long reports condensed to decisions and action points — as many as you need.
A fixed, predictable endpoint for prototyping, CI runs and load tests — before you commit to bigger hardware.
We onboard a limited number of dedicated endpoints — deliberately, one customer at a time, on infrastructure we own and operate. Pilot customers secure preferential rates they keep beyond the pilot, and work directly with the engineers who run their hardware.
Endpoints are allocated one dedicated GPU at a time, first come, first served. Capacity grows with demand — deliberately, not speculatively.
OpenAI-compatible chat completions with streaming and token-usage reporting. Your existing SDKs and tooling work unchanged — only the base URL changes.
Sizing, model selection, latency, compliance — you talk to the people who operate the hardware. No account layer in between.
A GPU at our Dalsbruk site reserved exclusively for you, serving an open-weight language model behind your own private, OpenAI-compatible API URL. No shared queues — your requests are the only requests on the card.
Yes. The endpoint speaks the OpenAI chat-completions protocol, including streaming and token-usage reporting. Point your existing SDK — Python, Node.js or plain cURL — at your base URL and keep your code unchanged.
On GPU hardware we own and operate in Dalsbruk, Finland. Prompts and completions are not routed through third-party clouds, a data processing agreement is included, and your data is never used for training.
Any open-weight model up to 8B-class — for example Qwen3-8B (Apache-2.0). Larger models on bigger GPUs are available on request.
Flat: €295 per month per dedicated GPU, additional GPUs €245 each, electricity included. There is no token metering and there are no overage charges — the GPU is yours around the clock. A 14-day evaluation is €95 and fully credited if you continue. Prices exclude VAT.
About 30 tokens per second of sustained generation per GPU, measured on our Dalsbruk hardware on 12 August 2026. That is sized for the everyday text workloads of a business — drafting, extraction, summarisation, internal assistants — not for frontier-scale training.
Monthly, throughout the entire pilot phase. No setup fee, no minimum term. Customers who join during the pilot keep their pilot rates.
Tell us your model and workload; we come back with a concrete proposal — typically within two business days.