141 GB on a standard PCIe card
141 GB of HBM3e at 4.8 TB/s on a dual-slot PCIe Gen 5 card, 1.5x the 94 GB of H100 NVL, keeps a 70B-parameter model in FP8 on a single GPU with room for long KV caches, no offloading to host memory.
Hopper-generation bare metal for LLM inference, RAG and HPC in air-cooled PCIe servers: 141 GB HBM3e, 4.8 TB/s, FP8 Transformer Engine, and a 2-way or 4-way NVLink bridge at 900 GB/s per GPU. Reserve capacity now.
Thanks. An engineer will spec your H200 NVL build and reach out shortly.
Something went wrong. Please try again or email sales@redswitches.com.
H200 NVL servers are landing in our data centers soon. Leave your details and our engineers will reserve a configuration for you, confirm lead times, and quote single-GPU, 2x, 4x NVLink-bridged and 8x builds with volume and committed-term discounts.
Thanks for reaching out. An engineer will spec your build and get back to you shortly.
Something went wrong. Please try again or email sales@redswitches.com.
Every RedSwitches NVIDIA H200 NVL server is single-tenant bare metal with full root and IPMI access, unmetered 1, 10 or 25 Gbps bandwidth, no setup fee, no bandwidth overage and a flat monthly price. Stocked builds are online in about 1 hour, larger configurations up to 8 GPUs per node are built to order, volume and 6/12-month committed-term discounts apply, payments include crypto, and engineers answer 24/7.
Reservations for NVIDIA H200 NVL are free and non-binding. An engineer confirms the lead time, holds the hardware for you, and sends a quote before you commit to anything.
No setup fee and a flat monthly price for the whole server. No per-hour meter and no surprise line items, so GPU spend is forecastable.
Unmetered 1, 10 or 25 Gbps uplinks are included, and whatever port speed you choose you can use all of it: no overage charges and no egress bills, ever. Moving datasets, checkpoints and model weights in and out costs nothing extra.
Single-tenant hardware with root and IPMI access. You choose the OS, drivers, CUDA or ROCm version, and the NVLink or MIG layout, and a private VLAN can link your RedSwitches servers.
Single and multi-GPU nodes, up to 8 GPUs per server with NVLink where the card supports it. Volume discounts on multi-GPU and multi-server orders, plus 6 and 12-month term savings.
Live chat, Telegram and email answered by engineers around the clock. Pay by card, PayPal, bank wire or crypto with no KYC, in 20+ Tier III data centers across the EU, US and Asia.
Same accelerator, different economics: what changes when the GPU sits in a dedicated server you control instead of a metered instance.
| RedSwitches H200 NVL server | Typical cloud GPU instance | |
|---|---|---|
| Billing | Flat monthly price per server | Per-hour or per-second metering |
| Bandwidth | Unmetered 1/10/25 Gbps, use the full port, no overage or egress fees | Egress billed per GB |
| Private networking | Private VLAN between your servers on request | Paid VPC and peering constructs |
| Setup fee | $0 | Varies by instance and region |
| Tenancy | Single-tenant bare metal | Shared, virtualised hosts |
| Access | Root and IPMI, your OS and drivers | Hypervisor-managed images |
| Multi-GPU | Up to 8 per node, NVLink where supported, built to order | Fixed instance shapes |
| Payments | Card, PayPal, bank wire, crypto | Card or invoice |
| Support | 24/7 engineers on chat, Telegram and email | Ticket tiers, paid support plans |
Hopper silicon on a dual-slot PCIe Gen 5 card: 141 GB HBM3e, 4.8 TB/s bandwidth, FP8 Transformer Engine, and a 4-way NVLink bridge at 900 GB/s per GPU.
H200 memory and bandwidth in an air-cooled PCIe card: 141 GB per GPU, 564 GB across a 4-way NVLink bridge, for racks that cannot take SXM or HGX.
141 GB of HBM3e at 4.8 TB/s on a dual-slot PCIe Gen 5 card, 1.5x the 94 GB of H100 NVL, keeps a 70B-parameter model in FP8 on a single GPU with room for long KV caches, no offloading to host memory.
NVLink bridges link two or four H200 NVL cards at 900 GB/s per GPU, the same per-GPU NVLink bandwidth as H200 SXM, so a 4-way group pools 564 GB of HBM3e for 70B to 400B class models without crossing the PCIe bus.
1,671 TFLOPS FP16/BF16 and 3,341 TFLOPS FP8 Tensor performance with sparsity; the Transformer Engine manages FP8 and FP16 precision layer by layer for more tokens per second without retraining your model.
A dual-slot PCIe card at up to 600 W configurable TDP drops into air-cooled servers and racks that cannot take SXM modules or HGX baseboards; we build nodes with 1 to 8 cards to order.
Partition each card into up to 7 MIG instances of 16.5 GB for many concurrent models, or run Confidential Computing to keep data and model weights protected in GPU memory for regulated workloads.
Single-tenant servers with full root and IPMI access, dual AMD EPYC or Intel Xeon hosts, NVMe storage, unmetered 1/10/25 Gbps uplinks, and no noisy neighbours, with NVLink bridge topology you control.
From 70B to 400B model inference and RAG to FP64 HPC, where air-cooled H200 NVL nodes earn their place.
Serve Llama, DeepSeek, Mistral and proprietary models with TensorRT-LLM or vLLM; one card holds a 70B model in FP8, a 4-way NVLink bridged group holds 400B-class models.
Keep embedding, reranker and generator models on one node; 141 GB per card holds long contexts and large KV caches so RAG answers stay fast under load.
Fine-tune 7B to 70B models in BF16 or FP8 on 1 to 4 NVLink-bridged cards, on the same CUDA stack as your SXM clusters, without moving to HGX infrastructure.
30 TFLOPS FP64 and 60 TFLOPS FP64 Tensor per card, with 141 GB for large meshes and datasets, for CFD, genomics, seismic processing and molecular dynamics in air-cooled racks.
Run Hopper-class inference where only air-cooled PCIe servers fit: colocation cages, enterprise data halls, and edge sites provisioned for standard 2U and 4U chassis.
Slice each card into up to 7 MIG instances of 16.5 GB for many small models, or use Confidential Computing for private inference services and regulated workloads.
Trusted by Enterprise Teams Worldwide
Read all RedSwitches reviews, or see them on Google, HostAdvice, Cryptwerk and Trustpilot.
Availability and reservations, H200 NVL vs H200 SXM and H100 NVL, 4-way NVLink bridges, pricing, software compatibility, power, and MIG.
H200 NVL capacity is being installed across our data centers now. Submit the reservation form with your preferred configuration and region; our engineers confirm the lead time for that build, hold the hardware for you, and send a quote. Reservations are free and non-binding until you approve the quote.
Both carry 141 GB of HBM3e at 4.8 TB/s on Hopper silicon. H200 SXM runs at 700 W with 1,979 TFLOPS FP16 Tensor and links 8 GPUs through NVLink and NVSwitch at 900 GB/s on an HGX baseboard; H200 NVL is a dual-slot PCIe Gen 5 card at up to 600 W with 1,671 TFLOPS FP16 Tensor and a 2-way or 4-way NVLink bridge at 900 GB/s per GPU. Choose H200 SXM for 8-GPU training and the highest per-GPU throughput; choose H200 NVL for inference, RAG and HPC in air-cooled PCIe servers or racks that cannot take HGX systems.
H200 NVL raises memory from 94 GB HBM3 to 141 GB HBM3e and bandwidth from 3.9 TB/s to 4.8 TB/s, on the same Hopper architecture and PCIe form factor. The extra 47 GB per card lets a single GPU hold larger models and longer contexts, and the 4-way NVLink bridge pools 564 GB across four cards. TDP moves from 350 to 400 W on H100 NVL to up to 600 W configurable on H200 NVL, still air-cooled.
Pricing is quoted per configuration because card count, host CPU, memory, storage and uplink choices vary widely. Volume discounts apply to multi-GPU and multi-server orders, and every GPU server qualifies for committed-term discounts on 6 and 12-month billing. Crypto, card, PayPal and bank wire are accepted.
NVLink bridges connect two or four H200 NVL cards in the same server at 900 GB/s per GPU, so tensor-parallel inference across a bridged group does not cross the PCIe bus. A 4-way group pools 564 GB of HBM3e, enough for 400B-class models in FP8 or 70B-class models with very long contexts. An 8-card node runs as two 4-way bridged groups; if you need all-to-all NVSwitch bandwidth across 8 GPUs, choose H200 SXM.
Yes. H200 NVL is Hopper, the same architecture as H100 and H200 SXM, so CUDA 12 builds, PyTorch, JAX, TensorFlow, TensorRT-LLM and vLLM run unchanged. NVIDIA lists a 5-year NVIDIA AI Enterprise subscription with H200 NVL, which adds NIM microservices and enterprise support for the software stack. We install Ubuntu, Debian, Rocky/AlmaLinux, Windows Server or your own ISO, with the NVIDIA driver, CUDA toolkit and container runtime preinstalled on request.
H200 NVL is a dual-slot PCIe card with a configurable TDP of up to 600 W and air cooling, so it fits standard enterprise servers rather than HGX baseboards. Our GPU racks are provisioned for high-density PCIe servers; an 8x H200 NVL node is deployed with the power feeds and airflow it needs. You do not need to plan facilities, only the workload.
Yes. Each H200 NVL can be partitioned into up to 7 MIG instances of 16.5 GB, each with its own memory and compute, for many concurrent models on one card. Hopper Confidential Computing keeps data and model weights protected in GPU memory, useful for private inference services and regulated workloads.
No problem. Our talented engineers will consult, architect, migrate, manage, and do whatever it takes to help your business grow and succeed.