Skip to main content

Quick Comparison Table

FeatureOn-Demand GPUsReserved ClustersPrivate CloudServerless Inference
Best forTraining, fine-tuning, experiments, burst computeSustained 24/7 workloads, predictable capacityProduction-scale, lowest pricing, long-termProduction APIs, prototypes
TenancyShared aggregated poolDedicated GPUsSingle-tenant, customer-isolatedMulti-tenant API
Setup time< 5 minutes24–48 hoursCustom (days); capacity confirmed in 24hInstant
Minimum commitmentNone (hourly)1 monthLong-term (multi-month to multi-year)None (pay-per-use)
Pricing model$/GPU/hourDiscounted prepaid $/GPU/hourCustom contract$/1M tokens
GPU accessFull SSH / root (VM or bare metal)Full SSH / root, dedicatedFull root, single-tenant; optional managed k8s / SlurmAPI only
ScalingManual; self-serve multi-nodePre-provisionedPre-provisioned, custom topologyAutomatic
NetworkingEthernet or InfiniBand (multi-node)Ethernet (100 Gb/s) or InfiniBand (3.2 Tb/s)InfiniBand, private subnets, network isolationManaged
Support levelStandard (sub-24h; P0 < 1h)PriorityDedicated, direct-to-engineer (24×7)Standard
SLA99.5% uptime99.5% min (10% bill credit if breached)Custom, negotiated per contract99.9%
GPU rates vary by region and availability. Always check app.hyperbolic.ai/gpus for current rates and real-time availability.

Detailed Comparison

On-Demand GPUs

Self-serve, pay-as-you-go GPU instances drawn from an aggregated supplier network. Spin up a single GPU or an interconnected multi-node cluster in minutes — no contracts, no upfront commitment, no sales calls.When to use:
  • Training and fine-tuning custom models
  • Running experiments, notebooks, and evaluations
  • Burst and batch processing jobs
  • Short training runs where you need full control of the environment
Key features:
  • Full root access via SSH
  • Virtual Machine or Bare Metal configurations
  • Deploy interconnected H100, H200, or B200 clusters (8, 16, 32, 64, 128+ GPUs) instantly
  • Up to 24 TB local NVMe storage; attachable network storage (~$0.0766/TB/month)
  • Hourly billing with no hidden or egress fees — failed instances are never charged

Reserved Clusters

Reserve dedicated GPUs at discounted prepaid pricing with guaranteed uptime — ideal for teams that want predictable, isolated capacity without competing for on-demand supply.When to use:
  • 24/7 production inference and LLM tooling
  • Sustained or scheduled training runs
  • High-volume, predictable usage
  • Teams that want capacity certainty plus a discount on what they already run on-demand
Key features:
  • Dedicated GPUs with guaranteed availability
  • Cluster sizes from 8 to 1,024 GPUs, co-located for distributed training
  • Networking: Ethernet (100 Gb/s) standard, or InfiniBand (3.2 Tb/s, NDR) for large-scale multi-node training (+$0.20/GPU/hour)
  • Management options: managed Kubernetes, managed Slurm, or bare-metal access
Pricing structure:
  • Discounted prepaid $/GPU/hour, billed monthly
  • Volume tiers: 8–32 GPUs (standard), 32–64 GPUs (volume discount), 64+ GPUs (custom enterprise pricing)
  • Term options from 1 month to 1 year with locked pricing
  • Get an instant quote

Private Cloud

Dedicated, single-tenant GPU infrastructure for production-scale AI workloads. Get an isolated environment with custom networking, storage, and SLAs — plus the best unit economics on long-term commitments.When to use:
  • Enterprise production workloads at sustained scale
  • Security-conscious teams that require single-tenant isolation
  • Large, long-term commitments (typically $1M+ deployments)
  • Workloads that need custom SLAs, custom networking, or custom storage
Key features:
  • Single-tenant, customer-isolated GPU clusters (H100, H200, B200)
  • Network isolation with private subnets and firewall rules
  • High-bandwidth InfiniBand interconnect for distributed training
  • Choice of operating model: full platform visibility, managed Kubernetes, managed Slurm, or a fully isolated environment where only your team holds credentials
  • Custom storage architecture (node-local NVMe and shared/parallel filesystems)
  • Dedicated, direct-to-engineer support with 24×7 coverage for critical issues
Pricing structure:
  • Custom contract based on cluster size, GPU type, term length, and configuration
  • Long-term commitments with the best unit economics
  • Talk to sales for a tailored quote

Serverless Inference

An OpenAI-compatible API for running the latest open models with zero infrastructure to manage. Pay only for the tokens you use.When to use:
  • Production API endpoints
  • Rapid prototyping
  • Variable or unpredictable traffic
  • Cost-sensitive applications using standard models
Key features:
  • Instant deployment — no provisioning
  • Automatic scaling
  • OpenAI-compatible API
  • Pay-per-use token billing

Available GPUs

Hyperbolic offers H100, H200, and B200 capacity across On-Demand GPUs, Reserved Clusters, and Private Cloud. Availability may vary by product, region, supply, and configuration. If you need a specific GPU type, cluster configuration, or capacity profile, contact sales@hyperbolic.ai and our team can help evaluate what is available or what can be sourced for your workload.
GPUArchitectureMemoryMemory BandwidthNVLinkTDPBest for
B200Blackwell192 GB HBM3e8 TB/s1.8 TB/s~1000 WFrontier-scale training and highest-throughput inference
H200Hopper141 GB HBM3e4.8 TB/s900 GB/s700 WLarge-model training and memory-bound inference
H100Hopper80 GB HBM33.35 TB/s900 GB/s700 WMainstream training, fine-tuning, and inference
Node configuration (H100 / H200): 8 GPUs per node with up to 160 vCPUs and up to 1.5 TB RAM per node, and up to 24 TB of local NVMe storage. Available as Virtual Machines (fast, flexible, ideal for development and small-to-medium training) or Bare Metal (full root access and system-level control, ideal for large-scale distributed training and custom CUDA configurations). Interconnect: Single-node instances use standard networking; multi-node clusters can use InfiniBand (up to 3.2 Tb/s, NDR / ConnectX-7) for low-latency GPU-to-GPU communication — essential for distributed training at 32+ GPUs.
Choosing between H100, H200, and B200? Use H100 for mainstream training, fine-tuning, and inference. Choose H200 for larger models, longer context, or memory-bound inference. Choose B200 for maximum per-GPU throughput on frontier-scale training and inference.

Need Help Deciding?

Talk to Sales

Get personalized recommendations and custom quotes for Reserved Clusters and Private Cloud

Contact Support

Get help from our support team

Get an Instant Quote

Price a reserved cluster in-app in minutes

Get Started

Launch your first GPU instance in minutes