- Home
- Private endpoints
Your own nodes.
We run the models on them.
A private inference endpoint built on single-tenant hardware and operated by us. You get the isolation of a dedicated cluster without taking on the job of running the stack.
What is included.
- Single-tenant nodes — Physical nodes are assigned to your account alone. GPUs are allocated whole and are never split between customers.
- Model operations — Weight placement, the inference server, batching, quantisation choices and version management are ours to run.
- Bring your own weights — Bring your own weights, or a continued-pretrained or fine-tuned model. The weights remain yours.
- A dedicated endpoint — An OpenAI-compatible URL and API key. Rate limits follow your reserved capacity and are unaffected by anyone else's load.
- No content inspection — On single-tenant hardware, compliance monitoring is metadata-level only. We do not inspect prompts or completions.
- The same SLA — ≥99.5% monthly availability per allocation unit, with 24×7 monitoring and first response.
Billing and contract.
Per reserved GPU
You pay monthly for the GPUs you reserve; the capacity is held for you regardless of utilisation. JPY, billed monthly in advance.
From one year
Shorter terms than a dedicated cluster. Longer terms improve the terms you are offered.
Node by node
Add capacity a node at a time, subject to window availability. Moving up to a dedicated cluster is handled inside the same agreement.
Pricing is determined by delivery window, contract term, reserved GPU count and reservation level. Qualified customers receive an allocation offer.
Tell us what you need. We reply within 48 hours.
Qualified enquiries receive proposed times for a technical meeting within 48 hours, and an allocation offer after that meeting. Information is handled under NDA.