How Much Does an NVIDIA H200 Cost in 2026?

In 2026, NVIDIA H200 pricing sits between H100 and Blackwell, making it a strong option for teams that need more memory without moving straight to the newest, highest-cost GPUs. It uses the same Hopper architecture class as the H100, but adds 76% more memory: 141GB of HBM3e versus 80GB on the H100 SXM. It also increases memory bandwidth from 3.35 TB/s on H100 SXM to 4.8 TB/s on H200 SXM, a major advantage for memory-bound LLM inference, fine-tuning, long-context workloads, and high-throughput serving.

The H200 is not NVIDIA’s newest GPU, but it remains one of the most practical options for teams that need more memory than the H100 without paying Blackwell-level rates.

This guide breaks down NVIDIA H200 pricing in 2026, including hardware purchase costs, hourly rental rates in a market where demand continues to outpace supply, and when the H200 is worth the premium over the H100.

H200 purchase price

Buying H200s outright means taking on the full hardware cost upfront. NVIDIA does not publish a standard retail price for H200 GPUs, so real quotes vary by OEM, reseller, form factor, warranty, region, availability, and whether you are buying a standalone GPU or a full HGX system.

Configuration

Typical 2026 Price Range

Notes

H200 SXM (per GPU, in HGX boards)

$30,000 - $44,000

SXM is sold on 4- and 8-GPU HGX boards, at about $44K and $38-39K per GPU respectively; bare modules occasionally list lower (~$29,500) but require an HGX baseboard.

H200 NVL (PCIe)

$32,000 - $33,500

Retail listings fall between $32,000$33,500. Available as an individually purchasable dual-slot PCIe card for compatible, adequately powered and cooled enterprise servers.

8-GPU HGX H200 Server

$300,000–$330,000

Full systems depend on CPU, RAM, networking, and support margins.

The purchase price is only the entry cost. An H200 SXM has a configurable TDP up to 700W, so an 8-GPU server can draw several kilowatts from the GPUs alone before CPUs, memory, networking, storage, fans, and cooling overhead.

For teams without existing data center infrastructure, the all-in cost includes power, cooling, rack space or colocation, high-speed networking, spare parts, monitoring, and the engineering time required to keep the system available. That is why many teams consume H200 capacity through cloud rental instead of buying hardware directly.

H200 rental prices in 2026

H200 rental pricing varies widely by provider. The biggest differences come from provider type, whether single-GPU rental is available, whether the instance is reserved or on demand, and whether you are buying raw GPU time or a fully managed cloud environment.

Provider / category

Current public H200 price

Notes

Hyperbolic 

$3.49/GPU-hour

Listed in the Hyperbolic dashboard as of July 16, 2026.

RunPod

~$4.39/GPU-hour

Public H200 141GB pricing.

CoreWeave

~$6.31/GPU-hour

CoreWeave lists H200 as an 8-GPU HGX node at $50.44/hr, or about $6.31/GPU-hour equivalent. The complete node is the minimum configuration shown.

AWS Capacity Blocks, p5e / p5en

~$5.97-$6.87/GPU-hour

AWS lists p5e.48xlarge Capacity Blocks at $47.76 per 8-GPU instance, or $5.97 per H200, and p5en.48xlarge at $54.92 per 8-GPU instance, or $6.87 per H200, in several U.S. regions.  Capacity Blocks are reserved in advance and charged through an upfront reservation fee. 

Oracle Cloud BM.GPU.H200.8

$10.00/GPU-hour

OCI lists H200 141GB Tensor Core GPU pricing at $10.00 per GPU-hour.

Azure ND96isr H200 v5

~$13.78/GPU-hour

Vantage lists ND96isr H200 v5 at $110.24/hr for the 8-GPU VM, or $13.78 per GPU-hour.

Prices are normalized to a per-GPU-hour equivalent for comparison. Minimum configurations, reservation requirements, included CPU and networking resources, regional availability, and billing terms vary by provider. Some providers require customers to rent a complete eight-GPU node.

H200 pricing varies widely by provider, and that variance can materially change the cost of running the same workload.

As of July 2026, the providers reviewed here range from $3.49 per GPU-hour to approximately $13.78 per GPU-hour on a normalized basis. Specialized GPU clouds generally post lower rates than the hyperscaler configurations included in this comparison, although minimum configurations and purchasing terms differ. Relative to Hyperbolic’s $3.49 rate, the hyperscaler configurations shown here are approximately 1.7x to 4x more expensive.

Marketplace pricing moves with real-time supply and demand. With on-demand H200 capacity operating near full utilization, available supply can be bid up as demand increases. Hyperbolic’s current $3.49 per GPU-hour rate reflects today’s market conditions, not a fixed pricing tier. 

For raw GPU workloads, that gap can dominate the total cost of training, fine-tuning, or serving. If your workload mainly needs H200 memory and compute, and you do not need a specific hyperscaler environment for credits, compliance, networking, or existing infrastructure, the higher hourly rate may not translate into better price-performance.

For sustained workloads, committed or reserved capacity can reduce the effective hourly rate. Public GPU pricing data shows that H100 and H200 one-year reservation discounts generally fall in the single digits to low teens, while some providers advertise discounts of up to 35% for large-scale clusters reserved for multiple months. The actual rate depends on the provider, commitment length, GPU quantity, region, and availability. 

What you get with the H200: key specs

Specification

NVIDIA H200 SXM Value

Architecture

Hopper

VRAM

141GB HBM3e

Memory bandwidth

4.8 TB/s

FP16 / BF16 Tensor performance

1,979 TFLOPS with sparsity

FP8 Tensor performance

3,958 TFLOPS with sparsity

NVLink

900GB/s

TDP

Up to 700W configurable

Form factor

SXM

NVIDIA’s official H200 specs list 141GB of HBM3e memory, 4.8 TB/s memory bandwidth, 3,958 TFLOPS of FP8 Tensor Core performance with sparsity, 900GB/s NVLink, and up to 700W configurable TDP for the SXM version.

The main advantage is memory. Compared with H100 SXM, the H200 gives you 76% more GPU memory and about 43% more memory bandwidth. That matters most when model weights, KV cache, batch size, or context length are the bottleneck.

H200 vs H100: is the price difference worth it?

The H100 and H200 have very similar compute specs, but the H200 has much more memory and bandwidth.

Spec

H100 SXM

H200 SXM

VRAM

80GB HBM3

141GB HBM3e

Memory bandwidth

3.35 TB/s

4.8 TB/s

FP8 Tensor performance

3,958 TFLOPS with sparsity

3,958 TFLOPS with sparsity

TDP

Up to 700W

Up to 700W

Hyperbolic marketplace rental

~$2.89/hr

~$3.49/hr

NVIDIA lists H100 SXM and H200 SXM with the same peak FP8 Tensor Core throughput, but H200 increases the memory subsystem from 80GB of HBM3 at 3.35 TB/s to 141GB of HBM3e at 4.8 TB/s. The performance gain comes from higher memory capacity and bandwidth, not additional Tensor Core compute. H200 is most valuable when model weights, KV cache, batch size, or context length are constrained by memory rather than raw FLOPS.

H200 is best understood as a memory-focused Hopper upgrade. Its advantage appears when memory capacity or bandwidth limits throughput, such as with larger models, longer contexts, bigger batches, or KV-cache-heavy inference. Compute-bound workloads that already fit comfortably within 80GB may see less benefit. Because H100 and H200 both use Hopper, and NVIDIA describes HGX H200 as drop-in compatible with HGX H100, teams can gain additional memory headroom without the broader platform transition required for Blackwell.

Choose H200 when:

  • You are serving models that exceed 80 GB or need more KV-cache headroom 

  • You need batch sizes that are constrained by H100 memory.

  • You are running long-context workloads.

  • Your workload is memory-bandwidth-bound.

  • You want to reduce tensor parallelism or fit a workload on fewer GPUs, provided it fits within 141GB.

Choose H100 when:

  • Your workload fits comfortably in 80GB.

  • You are running smaller models.

  • Your job is more compute-bound than memory-bound.

  • You care more about lowest hourly cost than memory headroom.

Compare cost per million tokens or cost per training run instead of cost per GPU-hour. Memory-bound inference is limited by HBM bandwidth and KV-cache capacity, not just FLOPS. Compare cost per million tokens or cost per training run, not cost per GPU-hour. H200’s advantage shows up when the workload is memory-bound, particularly at long context or high concurrency, where H100’s smaller memory capacity can constrain the KV cache and throughput. In those cases, H200 can deliver better cost per token. When the model fits comfortably in 80GB and the workload is compute-bound, H100 usually wins on unit economics. 

Buy vs. rent: breakeven math

Assume an H200 is available for $35,000 through a reseller and costs $3.49/hr to rent. This simplified comparison excludes the server platform, financing, depreciation, resale value, maintenance, downtime, storage, and data transfer.

  • Hardware-only breakeven: $35,000 ÷ $3.49/hr = about 10,000 hours, or roughly 14 months of 24/7 use.

  • One year of 24/7 rental: $3.49 × 8,760 hours = about $30,600.

  • At 50% utilization: annual rental cost is about $15,300, pushing hardware-only breakeven past two years.

  • Additional ownership costs: The server platform, power, cooling, hosting, networking, maintenance, and operations generally extend the ownership breakeven period, although resale value can partially offset the cost.

Renting is generally more attractive for startups, research teams, and product teams when demand is variable, utilization is uncertain, or the team does not already operate suitable data-center infrastructure. It avoids capex, removes much of the infrastructure burden, and lets teams move between A100, H100, H200, B200, or future GPUs as workload requirements change. Buying can make sense for consistently utilized, multi-year workloads when the organization can secure favorable hardware pricing and operate the system efficiently.

Where H200 makes sense in 2026 

The H200 is most compelling when you need more than 80GB of VRAM or when memory bandwidth is limiting throughput. It gives teams a strong middle option between the lower hourly cost of H100 and the higher memory and next-generation performance of Blackwell GPUs.

Where the H200 still shines:

  • Inference on large language models 

  • Long-context inference

  • Large-batch serving

  • RAG workloads with large context windows

  • Fine-tuning jobs that benefit from extra memory headroom

  • Teams that want Hopper maturity without jumping straight to Blackwell pricing

Where you may want a different GPU:

  • Use H100 if your workload fits in 80GB and is not memory-bound.

  • Use A100 if you need cheaper experimentation and can tolerate lower throughput.

  • Use B200 if you need more memory, FP4 inference support, or stronger next-generation performance.

Ready to use H200s without buying a server?

Rent H200 GPUs on demand through Hyperbolic’s GPU marketplace. Get market-driven pricing, hourly billing, and live availability without the capital expense of buying an HGX H200 server.

Browse live H200 pricing

Frequently asked questions

NVIDIA does not publish a standard H200 retail price. In 2026, example reseller listings for individual H200 configurations range from approximately $31,000-$44,000, depending on form factor, supplier, warranty, and availability. Example listings for full 8-GPU HGX H200 systems range from roughly $300,000 to $500,000+ depending on configuration. These are indicative reseller listings, not NVIDIA MSRP or guaranteed transaction prices.

As of July 2026, the public H200 configurations reviewed here range from about $3.49/GPU-hour on Hyperbolic’s marketplace to approximately $13.78/GPU-hour on a normalized basis. RunPod lists single-GPU H200 access at $4.39/hr. CoreWeave lists an 8-GPU H200 node at $50.44/hr, or about $6.31 per GPU-hour equivalent. AWS lists scheduled p5e Capacity Blocks at $47.76/hr for an 8-GPU instance, or $5.97 per GPU-hour equivalent. Oracle lists H200 at $10/GPU-hour, and Vantage lists Azure ND96isr H200 v5 at approximately $110.24/hr for an 8-GPU VM, or $13.78 per GPU-hour equivalent. Minimum configurations, regions, and purchasing terms differ.

The NVIDIA H200 has 141GB of HBM3e memory and 4.8 TB/s of memory bandwidth.

For memory-bound workloads, yes. The H200 has the same listed FP8 Tensor Core performance as H100 SXM, but it has 76% more memory and 43% more memory bandwidth. For compute-bound workloads that fit in 80GB, H100 can still be the better value.

Renting is generally more attractive when demand is variable, utilization is uncertain, or the team does not already operate suitable data-center infrastructure. At $35,000 to buy versus $3.49/hr to rent, the simplified hardware-only breakeven is about 14 months of 24/7 utilization. At 50% utilization, it moves past two years before the server platform, financing, power, cooling, hosting, networking, maintenance, and operations are included. Buying can make sense for consistently utilized, multi-year workloads when the organization can secure favorable hardware pricing and operate the system efficiently.

Choose H200 if your workload fits in 141GB and you want strong Hopper price-performance. Choose B200 if you need more memory, Blackwell features, or stronger performance for workloads that can benefit from the newer architecture.