Who it is for

Teams that know the workload but need the right capacity form.

Training, inference, and batch work are qualified through runtime shape instead of a hardware name alone.

  • Training teamsRuns measured in daysA fixed term holds the accelerators, and the interconnect is chosen before the run starts, not after a failure.
  • Inference teamsTraffic that movesSteady load sits on reserved capacity; peaks go to on-demand, and idle periods do not hold hardware.
  • Batch and evaluationWork that can waitInterruptible capacity is acceptable when the job can restart, and that tolerance is stated up front.

What the account receives

Capacity decisions tied to the way work runs.

Workload qualification

Define duration, traffic shape, and interruption tolerance.

The capacity form fits the workload from the first run.

Memory fit

Match working-set needs to a realistic accelerator tier.

Memory limits enter the decision before a hardware name does.

Scale path

Review single-node and coordinated work through the interconnect need.

A design is proven against the workload before it scales.

Continuity choice

Select on-demand, reserved, or serverless capacity by operating pattern.

Continuity expectations are written down before the term starts.

B200 SXM 192 GB Large training · Reserved
H200 SXM 141 GB Memory-heavy inference · Reserved
H100 SXM 80 GB Training · Flexible forms
L40S 48 GB Inference · Batch

Ready to secure capacity for your workload?

Tell us how the workload runs: how long, how tolerant of interruption, how much memory. A member of the team replies within one business day.

Request access