- Home
- Platform
Three ways to consume it. One platform.
One power source, one data centre and one operations team, with three entry points at different granularities. You can start with a single API key and move to private endpoints and then to a dedicated cluster. Capacity is allocated by window and fixed by a Capacity Reservation.
Managed API
An OpenAI-compatible endpoint. Billed per token, from a single API key. Models are offered by capability class.
Private endpoints
A private inference endpoint we run for you on single-tenant nodes. Billed per reserved GPU, from one node of 8 × B300.
Dedicated clusters
Single-tenant bare metal — you run your own stack. From 32 nodes (256 GPUs) to 256 nodes (2,048 GPUs).
It comes down to who runs the stack.
| Managed API | Private endpoints | Dedicated clusters | |
|---|---|---|---|
| Who runs the stack | Nachster AI | Nachster AI | You |
| Hardware sharing | Shared capacity | Single tenant | Single tenant, bare metal |
| Billing | Per token | Per reserved GPU | Committed capacity |
| Minimum | One API key | 1 node (8 GPUs) | 32 nodes (256 GPUs) |
| Lead time | Q2 2027 | ~3–4 months | ~6 months standard |
| Best for | Evaluation, internal tools, variable inference | Steady inference, serving your own weights | Training, sustained large-scale inference |
If you need dedicated hardware below 32 nodes, private endpoints are offered per node. Talk to us.
The NVIDIA B300 standard node.
Every tier runs on the same standard node: eight B300s in a 10U air-cooled chassis, drawing roughly 15 kW per node as a reference figure.
Nachster AI B300 — standard node and first cluster
| Platform | HGX B300 NVL16 · 8-GPU node · 10U air-cooled |
|---|---|
| GPU | 8 × NVIDIA B300 · 288 GB HBM3e each · ~2.3 TB / node |
| CPU · memory | 2 × Intel Xeon 6 · 3 TB DDR5-6400 |
| Local storage | 2 × 960 GB M.2 (OS) · 8 × 3.84 TB NVMe |
| Networking | 8 × ConnectX-8 800G (integrated) · BlueField-3 DPU |
| Cluster · fabric | 32 nodes · 256 GPUs · Quantum-X800 InfiniBand · rail-optimized · 1:1 |
| Shared storage | Dedicated high-performance storage |
| Power | ~15 kW per node (reference) |
| Location | Japan · dedicated, single tenant |
| Operations · SLA | Managed 24×7 · ≥99.5% per allocation unit / month |
| Contract | Multi-year committed capacity · 5-year standard |
Standard node configuration; the final BOM is attached to the allocation offer. Later generations — GB300 · Vera Rubin — are allocated on subsequent windows.
Four cluster sizes.
Every size is built on the same 1:1 non-blocking rail-optimised fabric. GPU memory and load are derived from the per-node reference figures; confirmed values appear in the allocation offer.
| Size | GPU | Total GPU memory | IT load (approx.) |
|---|---|---|---|
| 32 nodes | 256 | ~74 TB | ~0.5 MW |
| 64 nodes | 512 | ~147 TB | ~1.0 MW |
| 128 nodes | 1,024 | ~295 TB | ~1.9 MW |
| 256 nodes | 2,048 | ~590 TB | ~3.8 MW |
The first delivery window (Q2 2027) is 32 nodes / 256 GPUs. Larger sizes are allocated on subsequent windows.
Generations and windows.
NVIDIA B300
HGX B300 NVL16, allocated by the 8-GPU node. Air-cooled, closed aisle.
NVIDIA GB300
Allocated by the rack, direct liquid cooling. Once a qualified site is in place.
NVIDIA Vera Rubin
Allocated by the rack, direct liquid cooling.
Custom cluster
Built to your design. Bring it to the technical meeting.
The allocation unit changes with the generation — a node for B300, a rack for the rack-scale generations. All capacity figures on this site are stated in GPUs; per-rack GPU counts and power density are confirmed at contracting.
Tell us what you need. We reply within 48 hours.
Qualified enquiries receive proposed times for a technical meeting within 48 hours, and an allocation offer after that meeting. Information is handled under NDA.