NVIDIA H200 NVL GPU Server | RedSwitches
// coming soon: hopper pcie

NVIDIA H200 NVL GPU Server

Hopper-generation bare metal for LLM inference, RAG and HPC in air-cooled PCIe servers: 141 GB HBM3e, 4.8 TB/s, FP8 Transformer Engine, and a 2-way or 4-way NVLink bridge at 900 GB/s per GPU. Reserve capacity now.

  • Coming soon, reservations open
  • 141 GB HBM3e, 4-way NVLink bridge
  • 20+ Tier III global data centers
  • Crypto payments & 24/7 support included
New HardwareComing Soon
Reserve your H200 NVL server

H200 NVL hardware is landing in our data centers soon. Spec your build now and our engineers will reserve it and confirm the lead time.

GPUH200 NVL

Reserve NVIDIA H200 NVL Capacity

H200 NVL servers are landing in our data centers soon. Leave your details and our engineers will reserve a configuration for you, confirm lead times, and quote single-GPU, 2x, 4x NVLink-bridged and 8x builds with volume and committed-term discounts.

Fastest channel for quick deploys

What You Get With Every NVIDIA H200 NVL Server

Every RedSwitches NVIDIA H200 NVL server is single-tenant bare metal with full root and IPMI access, unmetered 1, 10 or 25 Gbps bandwidth, no setup fee, no bandwidth overage and a flat monthly price. Stocked builds are online in about 1 hour, larger configurations up to 8 GPUs per node are built to order, volume and 6/12-month committed-term discounts apply, payments include crypto, and engineers answer 24/7.

  • Free

    Reserve at No Cost

    Reservations for NVIDIA H200 NVL are free and non-binding. An engineer confirms the lead time, holds the hardware for you, and sends a quote before you commit to anything.

  • $0

    Setup Fee, Flat Monthly

    No setup fee and a flat monthly price for the whole server. No per-hour meter and no surprise line items, so GPU spend is forecastable.

  • Unmetered

    Bandwidth, No Egress Bills

    Unmetered 1, 10 or 25 Gbps uplinks are included, and whatever port speed you choose you can use all of it: no overage charges and no egress bills, ever. Moving datasets, checkpoints and model weights in and out costs nothing extra.

  • 1 tenant

    Bare Metal, Full Control

    Single-tenant hardware with root and IPMI access. You choose the OS, drivers, CUDA or ROCm version, and the NVLink or MIG layout, and a private VLAN can link your RedSwitches servers.

  • Up to 8x

    Multi-GPU, Built to Order

    Single and multi-GPU nodes, up to 8 GPUs per server with NVLink where the card supports it. Volume discounts on multi-GPU and multi-server orders, plus 6 and 12-month term savings.

  • 24/7

    Engineers, Not Bots

    Live chat, Telegram and email answered by engineers around the clock. Pay by card, PayPal, bank wire or crypto with no KYC, in 20+ Tier III data centers across the EU, US and Asia.

RedSwitches H200 NVL Server vs a Typical Cloud GPU Instance

Same accelerator, different economics: what changes when the GPU sits in a dedicated server you control instead of a metered instance.

RedSwitches NVIDIA H200 NVL dedicated server versus a typical hyperscale cloud GPU instance
RedSwitches H200 NVL serverTypical cloud GPU instance
BillingFlat monthly price per serverPer-hour or per-second metering
BandwidthUnmetered 1/10/25 Gbps, use the full port, no overage or egress feesEgress billed per GB
Private networkingPrivate VLAN between your servers on requestPaid VPC and peering constructs
Setup fee$0Varies by instance and region
TenancySingle-tenant bare metalShared, virtualised hosts
AccessRoot and IPMI, your OS and driversHypervisor-managed images
Multi-GPUUp to 8 per node, NVLink where supported, built to orderFixed instance shapes
PaymentsCard, PayPal, bank wire, cryptoCard or invoice
Support24/7 engineers on chat, Telegram and emailTicket tiers, paid support plans

NVIDIA H200 NVL Key Specifications

Hopper silicon on a dual-slot PCIe Gen 5 card: 141 GB HBM3e, 4.8 TB/s bandwidth, FP8 Transformer Engine, and a 4-way NVLink bridge at 900 GB/s per GPU.

Architecture
NVIDIA Hopper (GH100), the same silicon family as H100 and H200 SXM, packaged as a PCIe card
Memory
141 GB HBM3e, 4.8 TB/s bandwidth per GPU
Computeper GPU, Tensor figures with sparsity
FP8 Tensor Core
3,341 TFLOPS (1,671 TFLOPS dense)
INT8 Tensor Core
3,341 TOPS (1,671 TOPS dense)
FP16/BF16 Tensor Core
1,671 TFLOPS (835 TFLOPS dense)
TF32 Tensor Core
835 TFLOPS (418 TFLOPS dense)
FP32
60 TFLOPS
FP64 / FP64 Tensor Core
30 / 60 TFLOPS
Form Factor & TDP
Dual-slot PCIe Gen 5 card, air-cooled, up to 600 W (configurable)
Interconnects
NVLink bridge
2-way or 4-way, 900 GB/s per GPU
PCIe
Gen 5 to the host
Multi-Instance GPU
Up to 7 MIG instances at 16.5 GB each
Engines
Transformer Engine (FP8), Confidential Computing
Software
NVIDIA lists a 5-year NVIDIA AI Enterprise subscription with H200 NVL; CUDA 12.x, NIM microservices, TensorRT-LLM, NeMo, PyTorch, JAX, Triton Inference Server

Why Choose H200 NVL

H200 memory and bandwidth in an air-cooled PCIe card: 141 GB per GPU, 564 GB across a 4-way NVLink bridge, for racks that cannot take SXM or HGX.

141 GB on a standard PCIe card

141 GB of HBM3e at 4.8 TB/s on a dual-slot PCIe Gen 5 card, 1.5x the 94 GB of H100 NVL, keeps a 70B-parameter model in FP8 on a single GPU with room for long KV caches, no offloading to host memory.

4-way NVLink bridge, 900 GB/s per GPU

NVLink bridges link two or four H200 NVL cards at 900 GB/s per GPU, the same per-GPU NVLink bandwidth as H200 SXM, so a 4-way group pools 564 GB of HBM3e for 70B to 400B class models without crossing the PCIe bus.

Hopper FP8 Transformer Engine

1,671 TFLOPS FP16/BF16 and 3,341 TFLOPS FP8 Tensor performance with sparsity; the Transformer Engine manages FP8 and FP16 precision layer by layer for more tokens per second without retraining your model.

Air-cooled, fits enterprise racks

A dual-slot PCIe card at up to 600 W configurable TDP drops into air-cooled servers and racks that cannot take SXM modules or HGX baseboards; we build nodes with 1 to 8 cards to order.

MIG and Confidential Computing

Partition each card into up to 7 MIG instances of 16.5 GB for many concurrent models, or run Confidential Computing to keep data and model weights protected in GPU memory for regulated workloads.

Bare metal, not shared cloud

Single-tenant servers with full root and IPMI access, dual AMD EPYC or Intel Xeon hosts, NVMe storage, unmetered 1/10/25 Gbps uplinks, and no noisy neighbours, with NVLink bridge topology you control.

Ideal Use Cases

From 70B to 400B model inference and RAG to FP64 HPC, where air-cooled H200 NVL nodes earn their place.

LLM Inference for 70B to 400B Models

Serve Llama, DeepSeek, Mistral and proprietary models with TensorRT-LLM or vLLM; one card holds a 70B model in FP8, a 4-way NVLink bridged group holds 400B-class models.

Retrieval-Augmented Generation

Keep embedding, reranker and generator models on one node; 141 GB per card holds long contexts and large KV caches so RAG answers stay fast under load.

Fine-Tuning and LoRA

Fine-tune 7B to 70B models in BF16 or FP8 on 1 to 4 NVLink-bridged cards, on the same CUDA stack as your SXM clusters, without moving to HGX infrastructure.

HPC and Scientific Computing

30 TFLOPS FP64 and 60 TFLOPS FP64 Tensor per card, with 141 GB for large meshes and datasets, for CFD, genomics, seismic processing and molecular dynamics in air-cooled racks.

Enterprise AI in Standard Racks

Run Hopper-class inference where only air-cooled PCIe servers fit: colocation cages, enterprise data halls, and edge sites provisioned for standard 2U and 4U chassis.

Multi-Tenant Inference Platforms

Slice each card into up to 7 MIG instances of 16.5 GB for many small models, or use Confidential Computing for private inference services and regulated workloads.

Trusted by Enterprise Teams Worldwide

  • Check Point
  • German Football Association
  • Mubi
  • Pluxee
  • Zeeve
  • University of Malta
  • mSpy
  • RevX
  • Turbo VPN
  • Athos Commerce
  • Heckyl
  • GenXAI
  • WLVPN
  • FMS
  • Contaque
  • Monotek
  • EasyGo VPN
  • Neopool
  • InfyGlobe Technologies
  • ALFA University College
  • Stief Group
  • SSH Invest Holding
  • Rhysley
  • Spirit of Math
  • VideoShip
  • ORB VPN
  • Ping VPN

Deep Dive & FAQs

Availability and reservations, H200 NVL vs H200 SXM and H100 NVL, 4-way NVLink bridges, pricing, software compatibility, power, and MIG.

When will H200 NVL servers be available and how does reservation work?

H200 NVL capacity is being installed across our data centers now. Submit the reservation form with your preferred configuration and region; our engineers confirm the lead time for that build, hold the hardware for you, and send a quote. Reservations are free and non-binding until you approve the quote.

H200 NVL vs H200 SXM: which should I choose?

Both carry 141 GB of HBM3e at 4.8 TB/s on Hopper silicon. H200 SXM runs at 700 W with 1,979 TFLOPS FP16 Tensor and links 8 GPUs through NVLink and NVSwitch at 900 GB/s on an HGX baseboard; H200 NVL is a dual-slot PCIe Gen 5 card at up to 600 W with 1,671 TFLOPS FP16 Tensor and a 2-way or 4-way NVLink bridge at 900 GB/s per GPU. Choose H200 SXM for 8-GPU training and the highest per-GPU throughput; choose H200 NVL for inference, RAG and HPC in air-cooled PCIe servers or racks that cannot take HGX systems.

H200 NVL vs H100 NVL: what changes?

H200 NVL raises memory from 94 GB HBM3 to 141 GB HBM3e and bandwidth from 3.9 TB/s to 4.8 TB/s, on the same Hopper architecture and PCIe form factor. The extra 47 GB per card lets a single GPU hold larger models and longer contexts, and the 4-way NVLink bridge pools 564 GB across four cards. TDP moves from 350 to 400 W on H100 NVL to up to 600 W configurable on H200 NVL, still air-cooled.

How is H200 NVL priced and are there discounts?

Pricing is quoted per configuration because card count, host CPU, memory, storage and uplink choices vary widely. Volume discounts apply to multi-GPU and multi-server orders, and every GPU server qualifies for committed-term discounts on 6 and 12-month billing. Crypto, card, PayPal and bank wire are accepted.

How does the 4-way NVLink bridge work and how far does it scale?

NVLink bridges connect two or four H200 NVL cards in the same server at 900 GB/s per GPU, so tensor-parallel inference across a bridged group does not cross the PCIe bus. A 4-way group pools 564 GB of HBM3e, enough for 400B-class models in FP8 or 70B-class models with very long contexts. An 8-card node runs as two 4-way bridged groups; if you need all-to-all NVSwitch bandwidth across 8 GPUs, choose H200 SXM.

Is my existing CUDA / PyTorch code compatible?

Yes. H200 NVL is Hopper, the same architecture as H100 and H200 SXM, so CUDA 12 builds, PyTorch, JAX, TensorFlow, TensorRT-LLM and vLLM run unchanged. NVIDIA lists a 5-year NVIDIA AI Enterprise subscription with H200 NVL, which adds NIM microservices and enterprise support for the software stack. We install Ubuntu, Debian, Rocky/AlmaLinux, Windows Server or your own ISO, with the NVIDIA driver, CUDA toolkit and container runtime preinstalled on request.

What about power and cooling for H200 NVL servers?

H200 NVL is a dual-slot PCIe card with a configurable TDP of up to 600 W and air cooling, so it fits standard enterprise servers rather than HGX baseboards. Our GPU racks are provisioned for high-density PCIe servers; an 8x H200 NVL node is deployed with the power feeds and airflow it needs. You do not need to plan facilities, only the workload.

Does H200 NVL support MIG and Confidential Computing?

Yes. Each H200 NVL can be partitioned into up to 7 MIG instances of 16.5 GB, each with its own memory and compute, for many concurrent models on one card. Hopper Confidential Computing keeps data and model weights protected in GPU memory, useful for private inference services and regulated workloads.

Not Sure Exactly What You Need

No problem. Our talented engineers will consult, architect, migrate, manage, and do whatever it takes to help your business grow and succeed.