Quick Comparison Table
| Feature | On-Demand GPUs | Reserved Clusters | Private Cloud | Serverless Inference |
|---|---|---|---|---|
| Best for | Training, fine-tuning, experiments, burst compute | Sustained 24/7 workloads, predictable capacity | Production-scale, lowest pricing, long-term | Production APIs, prototypes |
| Tenancy | Shared aggregated pool | Dedicated GPUs | Single-tenant, customer-isolated | Multi-tenant API |
| Setup time | < 5 minutes | 24–48 hours | Custom (days); capacity confirmed in 24h | Instant |
| Minimum commitment | None (hourly) | 1 month | Long-term (multi-month to multi-year) | None (pay-per-use) |
| Pricing model | $/GPU/hour | Discounted prepaid $/GPU/hour | Custom contract | $/1M tokens |
| GPU access | Full SSH / root (VM or bare metal) | Full SSH / root, dedicated | Full root, single-tenant; optional managed k8s / Slurm | API only |
| Scaling | Manual; self-serve multi-node | Pre-provisioned | Pre-provisioned, custom topology | Automatic |
| Networking | Ethernet or InfiniBand (multi-node) | Ethernet (100 Gb/s) or InfiniBand (3.2 Tb/s) | InfiniBand, private subnets, network isolation | Managed |
| Support level | Standard (sub-24h; P0 < 1h) | Priority | Dedicated, direct-to-engineer (24×7) | Standard |
| SLA | 99.5% uptime | 99.5% min (10% bill credit if breached) | Custom, negotiated per contract | 99.9% |
GPU rates vary by region and availability. Always check app.hyperbolic.ai/gpus for current rates and real-time availability.
Detailed Comparison
On-Demand GPUs
Self-serve, pay-as-you-go GPU instances drawn from an aggregated supplier network. Spin up a single GPU or an interconnected multi-node cluster in minutes — no contracts, no upfront commitment, no sales calls.When to use:
- Training and fine-tuning custom models
- Running experiments, notebooks, and evaluations
- Burst and batch processing jobs
- Short training runs where you need full control of the environment
- Full root access via SSH
- Virtual Machine or Bare Metal configurations
- Deploy interconnected H100, H200, or B200 clusters (8, 16, 32, 64, 128+ GPUs) instantly
- Up to 24 TB local NVMe storage; attachable network storage (~$0.0766/TB/month)
- Hourly billing with no hidden or egress fees — failed instances are never charged
Reserved Clusters
Reserve dedicated GPUs at discounted prepaid pricing with guaranteed uptime — ideal for teams that want predictable, isolated capacity without competing for on-demand supply.When to use:
- 24/7 production inference and LLM tooling
- Sustained or scheduled training runs
- High-volume, predictable usage
- Teams that want capacity certainty plus a discount on what they already run on-demand
- Dedicated GPUs with guaranteed availability
- Cluster sizes from 8 to 1,024 GPUs, co-located for distributed training
- Networking: Ethernet (100 Gb/s) standard, or InfiniBand (3.2 Tb/s, NDR) for large-scale multi-node training (+$0.20/GPU/hour)
- Management options: managed Kubernetes, managed Slurm, or bare-metal access
- Discounted prepaid $/GPU/hour, billed monthly
- Volume tiers: 8–32 GPUs (standard), 32–64 GPUs (volume discount), 64+ GPUs (custom enterprise pricing)
- Term options from 1 month to 1 year with locked pricing
- Get an instant quote
Private Cloud
Dedicated, single-tenant GPU infrastructure for production-scale AI workloads. Get an isolated environment with custom networking, storage, and SLAs — plus the best unit economics on long-term commitments.When to use:
- Enterprise production workloads at sustained scale
- Security-conscious teams that require single-tenant isolation
- Large, long-term commitments (typically $1M+ deployments)
- Workloads that need custom SLAs, custom networking, or custom storage
- Single-tenant, customer-isolated GPU clusters (H100, H200, B200)
- Network isolation with private subnets and firewall rules
- High-bandwidth InfiniBand interconnect for distributed training
- Choice of operating model: full platform visibility, managed Kubernetes, managed Slurm, or a fully isolated environment where only your team holds credentials
- Custom storage architecture (node-local NVMe and shared/parallel filesystems)
- Dedicated, direct-to-engineer support with 24×7 coverage for critical issues
- Custom contract based on cluster size, GPU type, term length, and configuration
- Long-term commitments with the best unit economics
- Talk to sales for a tailored quote
Serverless Inference
An OpenAI-compatible API for running the latest open models with zero infrastructure to manage. Pay only for the tokens you use.When to use:
- Production API endpoints
- Rapid prototyping
- Variable or unpredictable traffic
- Cost-sensitive applications using standard models
- Instant deployment — no provisioning
- Automatic scaling
- OpenAI-compatible API
- Pay-per-use token billing
Available GPUs
Hyperbolic offers H100, H200, and B200 capacity across On-Demand GPUs, Reserved Clusters, and Private Cloud. Availability may vary by product, region, supply, and configuration. If you need a specific GPU type, cluster configuration, or capacity profile, contact sales@hyperbolic.ai and our team can help evaluate what is available or what can be sourced for your workload.| GPU | Architecture | Memory | Memory Bandwidth | NVLink | TDP | Best for |
|---|---|---|---|---|---|---|
| B200 | Blackwell | 192 GB HBM3e | 8 TB/s | 1.8 TB/s | ~1000 W | Frontier-scale training and highest-throughput inference |
| H200 | Hopper | 141 GB HBM3e | 4.8 TB/s | 900 GB/s | 700 W | Large-model training and memory-bound inference |
| H100 | Hopper | 80 GB HBM3 | 3.35 TB/s | 900 GB/s | 700 W | Mainstream training, fine-tuning, and inference |
Need Help Deciding?
Talk to Sales
Get personalized recommendations and custom quotes for Reserved Clusters and Private Cloud
Contact Support
Get help from our support team
Get an Instant Quote
Price a reserved cluster in-app in minutes
Get Started
Launch your first GPU instance in minutes

