Dedicated model · Available as managed deployment
DeepSeek's reasoning model — a 671B mixture-of-experts trained with reinforcement learning to think before it answers, with a 164k-token context, under MIT. Validated on AxForge hardware and deployed on a multi-GPU system sourced against your request for your traffic only — an OpenAI-compatible endpoint on hardware only you use, operated by AxForge in the EU.
Why AxForge
| Reasoning at the top of the open field | R1 made long-form chain-of-thought reasoning available as open weights — maths, code, analysis and planning tasks where a step-by-step model wins. |
|---|---|
| MIT licence | The most permissive licence in the field: no usage conditions, distillation allowed, commercial use allowed. |
| Your own R1, in the EU | A multi-GPU system reserved for you, with prompts processed in memory and never retained — the reasoning stays in Europe. |
Specifications
| Model | DeepSeek-R1 — deepseek-ai |
|---|---|
| Modalities | Text |
| Sizes | 684.5B |
| Context window | 163,840 tokens |
| Licence | Open weights — mit; commercial use permitted |
| Hardware | Multi-GPU system (H100, H200 or B200 class) sourced against your request |
| Managed service | Quoted per deployment |
| Region | European region — placement confirmed with your request |
Full details, benchmarks and FAQ on the DeepSeek-R1 page. Prices exclude VAT.
How it works
| 1 | Request deployment — describe your traffic, context needs and rental term. |
|---|---|
| 2 | You receive the configuration, hardware rental and managed-service price in writing before anything is billed. |
| 3 | AxForge deploys DeepSeek-R1 on a multi-GPU system sourced against your request reserved for you. |
| 4 | Point your OpenAI SDK at your own endpoint with the model name you receive. |
| 5 | Adjust the term — hour, week, month or year — as your workload settles. |
Request deployment or sign in to start.
FAQ
Not on the serverless API — it is available as a managed deployment: validated on AxForge hardware and deployed on a multi-GPU system sourced against your request for your traffic only. The serverless API serves Qwen3.8 27B.
A multi-GPU system — H100, H200 or B200 class — which AxForge sources against your request and confirms in writing before anything is billed.
Yes — the R1 distilled models (Qwen and Llama based, 1.5B to 70B) carry much of the reasoning behaviour and fit a single dedicated machine; ask for them in the request.
AxForge publishes only numbers it measures itself, and has not benchmarked this model on its nodes yet. For quality benchmarks, see the official model card.
Hardware and the managed service are quoted per deployment — both confirmed in writing before anything is billed.