NVIDIA H100 NVL GPU Server | RedSwitches
// coming soon: hopper pcie

NVIDIA H100 NVL GPU Server

Hopper bare metal for mainstream PCIe servers: 94 GB HBM3 at 3.9 TB/s, 1,671 TFLOPS FP16 Tensor, FP8 Transformer Engine, and a 600 GB/s NVLink bridge between pairs of cards, 1 to 8 GPUs per node in air-cooled racks. Reserve capacity now.

  • Coming soon, reservations open
  • 94 GB HBM3, PCIe Gen5 with NVLink bridge
  • 20+ Tier III global data centers
  • Crypto payments & 24/7 support included
New HardwareComing Soon
Reserve your H100 NVL server

H100 NVL hardware is landing in our data centers soon. Spec your build now and our engineers will reserve it and confirm the lead time.

GPUH100 NVL

Reserve NVIDIA H100 NVL Capacity

H100 NVL servers are landing in our data centers soon. Leave your details and our engineers will reserve a configuration for you, confirm lead times, and quote single-card, bridged-pair, 4x and 8x PCIe builds with volume and committed-term discounts.

Fastest channel for quick deploys

What You Get With Every NVIDIA H100 NVL Server

Every RedSwitches NVIDIA H100 NVL server is single-tenant bare metal with full root and IPMI access, unmetered 1, 10 or 25 Gbps bandwidth, no setup fee, no bandwidth overage and a flat monthly price. Stocked builds are online in about 1 hour, larger configurations up to 8 GPUs per node are built to order, volume and 6/12-month committed-term discounts apply, payments include crypto, and engineers answer 24/7.

  • Free

    Reserve at No Cost

    Reservations for NVIDIA H100 NVL are free and non-binding. An engineer confirms the lead time, holds the hardware for you, and sends a quote before you commit to anything.

  • $0

    Setup Fee, Flat Monthly

    No setup fee and a flat monthly price for the whole server. No per-hour meter and no surprise line items, so GPU spend is forecastable.

  • Unmetered

    Bandwidth, No Egress Bills

    Unmetered 1, 10 or 25 Gbps uplinks are included, and whatever port speed you choose you can use all of it: no overage charges and no egress bills, ever. Moving datasets, checkpoints and model weights in and out costs nothing extra.

  • 1 tenant

    Bare Metal, Full Control

    Single-tenant hardware with root and IPMI access. You choose the OS, drivers, CUDA or ROCm version, and the NVLink or MIG layout, and a private VLAN can link your RedSwitches servers.

  • Up to 8x

    Multi-GPU, Built to Order

    Single and multi-GPU nodes, up to 8 GPUs per server with NVLink where the card supports it. Volume discounts on multi-GPU and multi-server orders, plus 6 and 12-month term savings.

  • 24/7

    Engineers, Not Bots

    Live chat, Telegram and email answered by engineers around the clock. Pay by card, PayPal, bank wire or crypto with no KYC, in 20+ Tier III data centers across the EU, US and Asia.

RedSwitches H100 NVL Server vs a Typical Cloud GPU Instance

Same accelerator, different economics: what changes when the GPU sits in a dedicated server you control instead of a metered instance.

RedSwitches NVIDIA H100 NVL dedicated server versus a typical hyperscale cloud GPU instance
RedSwitches H100 NVL serverTypical cloud GPU instance
BillingFlat monthly price per serverPer-hour or per-second metering
BandwidthUnmetered 1/10/25 Gbps, use the full port, no overage or egress feesEgress billed per GB
Private networkingPrivate VLAN between your servers on requestPaid VPC and peering constructs
Setup fee$0Varies by instance and region
TenancySingle-tenant bare metalShared, virtualised hosts
AccessRoot and IPMI, your OS and driversHypervisor-managed images
Multi-GPUUp to 8 per node, NVLink where supported, built to orderFixed instance shapes
PaymentsCard, PayPal, bank wire, cryptoCard or invoice
Support24/7 engineers on chat, Telegram and emailTicket tiers, paid support plans

NVIDIA H100 NVL Key Specifications

Hopper GH100 on a dual-slot PCIe Gen5 card: 94 GB HBM3, 3.9 TB/s bandwidth, FP8 Transformer Engine, and a 600 GB/s NVLink bridge between pairs.

Architecture
NVIDIA Hopper (GH100), the PCIe Gen5 member of the H100 family
Memory
94 GB HBM3, 3.9 TB/s bandwidth per GPU
Computeper GPU, with sparsity; dense is half
FP8 Tensor Core
3,341 TFLOPS
FP16/BF16 Tensor Core
1,671 TFLOPS
TF32 Tensor Core
835 TFLOPS
INT8 Tensor Core
3,341 TOPS
FP32
60 TFLOPS
FP64 / FP64 Tensor Core
30 / 60 TFLOPS
Form Factor & TDP
PCIe Gen 5 dual-slot, air-cooled, 350 to 400 W (configurable)
Interconnects
NVLink bridge
600 GB/s between a pair of cards
PCIe
Gen 5, 128 GB/s to the host
Multi-Instance GPU
Up to 7 MIG instances at 12 GB each
Engines
Transformer Engine (FP8), Confidential Computing
Software
NVIDIA AI Enterprise subscription (included with H100 NVL per NVIDIA), CUDA 12.x, TensorRT-LLM, NIM microservices, NeMo, PyTorch, JAX, Triton Inference Server

Why Choose H100 NVL

More memory and bandwidth than H100 SXM on a 350 to 400 W air-cooled PCIe card, with an NVLink bridge for paired inference and 1 to 8 cards per node.

94 GB of HBM3 on a PCIe card

94 GB of HBM3 at 3.9 TB/s per GPU, more memory and more bandwidth than the 80 GB H100 SXM, keeps larger weights, longer KV caches and bigger batches resident on a single card, and a bridged pair pools 188 GB.

Hopper Tensor Cores with FP8 Transformer Engine

1,671 TFLOPS FP16/BF16 and 3,341 TFLOPS FP8 Tensor performance with sparsity; the Transformer Engine selects FP8 or FP16 per layer automatically, doubling inference throughput over FP16 without retraining.

NVLink bridge for paired cards

A 600 GB/s NVLink bridge joins two H100 NVL cards, nearly 5x the 128 GB/s of PCIe Gen5, so tensor-parallel serving of Llama 70B class models in FP8 runs on one bridged pair with minimal interconnect overhead.

Air-cooled, 350 to 400 W

A dual-slot PCIe Gen5 card at 350 to 400 W, roughly half the 700 W of H100 SXM, fits mainstream air-cooled servers with 1 to 8 cards per node: no liquid cooling, no HGX baseboard, no special facilities.

Partition it or protect it

Up to 7 MIG instances of 12 GB each turn one card into several isolated GPUs for smaller models and tenants, and Hopper Confidential Computing shields data and models in GPU memory inside a trusted execution environment.

Bare metal, not shared cloud

Single-tenant servers with full root and IPMI access, host CPU, RAM and storage sized to your workload, unmetered 1/10/25 Gbps uplinks, no setup fee and no noisy neighbours, built to order with the card count you choose.

Ideal Use Cases

From Llama 70B class inference on a bridged pair to RAG, fine-tuning and MIG-partitioned platforms, where H100 NVL nodes pay for themselves.

LLM Inference and Serving

Serve Llama 70B class models in FP8 on a bridged pair, or 7B to 30B class models on a single 94 GB card, with TensorRT-LLM, NIM microservices or Triton Inference Server.

Retrieval-Augmented Generation

Keep embeddings, reranker and generator on one node; 94 GB per card and 3.9 TB/s bandwidth serve long-context RAG with predictable latency.

Fine-Tuning and Continued Pre-Training

Run LoRA, QLoRA and full fine-tunes of 7B to 70B parameter models on 2x to 8x nodes, with the FP8 Transformer Engine cutting time per epoch.

Multimodal and Diffusion Models

Train and serve vision-language, image and video-generation models whose activations outgrow 48 GB and 80 GB cards.

HPC and Scientific Computing

60 TFLOPS FP64 Tensor Core and 30 TFLOPS FP64 per card for CFD, genomics, molecular dynamics and simulation in standard PCIe servers.

Multi-Tenant AI Platforms

Slice each card into up to 7 MIG instances and run many concurrent model instances, with Confidential Computing for private inference services on dedicated hardware.

Trusted by Enterprise Teams Worldwide

  • Check Point
  • German Football Association
  • Mubi
  • Pluxee
  • Zeeve
  • University of Malta
  • mSpy
  • RevX
  • Turbo VPN
  • Athos Commerce
  • Heckyl
  • GenXAI
  • WLVPN
  • FMS
  • Contaque
  • Monotek
  • EasyGo VPN
  • Neopool
  • InfyGlobe Technologies
  • ALFA University College
  • Stief Group
  • SSH Invest Holding
  • Rhysley
  • Spirit of Math
  • VideoShip
  • ORB VPN
  • Ping VPN

Trusted in Production

Real reviews from Google, HostAdvice, and Cryptwerk.

  • 4.8/525Gbps Bare Metal
    Port speeds on cloud were capped and shady. With RS I actually get the 25Gbps they say. No throttle bs.
    Felix MartinsenDenmark
    HostAdvice
  • 4.8/5Validators
    Set up multiple Solana + Avalanche validators through them. Hardware was clean, latency was low (especially in Europe), and uptime’s been 100% so far.
    Peppe TerranovaSouth Korea
    HostAdvice
  • 5.0/5Full Root Access
    As a sysadmin, I care more about control than flashy dashboards. RedSwitches gives me root access, IPMI, and actual hardware specs I can configure. Good for serious users.
    Eun-Woo GimSouth Korea
    Cryptwerk
  • 5.0/5Cloud Exit
    RedSwitches gave us full control over our infrastructure without the vendor lock-in we kept running into with cloud hosts. Customizable builds, fast provisioning, and actual humans handling support tickets. It’s refreshing.
    Princeton OttUnited States
    Cryptwerk
  • 5.0/599.99% Uptime
    My website works smoothly, thanks to their 99.99% uptime guarantee. The support team is another plus for me, as I always have been able to get help whenever I needed it in just 5 minutes on average.
  • 5.0/5ETH & BTC Nodes
    Deployed a few Ethereum and Bitcoin nodes here. Uptime’s been flawless, and sync speed was great thanks to their storage config.
    Ervin KaruGermany
    HostAdvice
  • 5.0/5Game Servers
    Been using RedSwitches for 6 months now for my small game server biz. Uptime has been great, and I haven’t run into any hidden charges.
    Jerold PerkinsUnited States
    Cryptwerk
  • 5.0/5AI Storage
    Using the storage servers to archive logs and snapshots from our AI pipeline… I also love that I could pick the datacenter closest to our team.
    Eino KoppelEstonia
    Cryptwerk
  • 4.8/5Support
    Support is super responsive. I had an issue with an OS reinstall and they jumped in within 10 minutes… Transparent pricing = win.
    Boniface LeandreUnited States
    HostAdvice
  • 5.0/5Instant Delivery
    They have an instant delivery section… they delivered it within 120 mins with all my requirements fulfilled (OS/RAID/Software configured etc).
    Daniel StuartSri Lanka
    HostAdvice
  • 4.8/5Global Locations
    With servers available in numerous strategic locations, RedSwitches offers exceptional versatility and performance for our company’s diverse hosting needs. Plus, their no setup fee policy really helps keep costs down.
    Tiberiu RomanUnited States
    HostAdvice
  • 5.0/5Bare Metal Cloud
    The dedicated server I got from RedSwitches has been incredibly reliable and fast. Their bare metal cloud solutions offer excellent performance, and the cloud VPS options are perfect for scaling. Highly recommended!
  • 5.0/5Easy Onboarding
    Absolutely delighted with RedSwitches! The setup was quick and free, and the fact that they accept all major payment gateways made the process seamless.
    Birk SpillumNorway
    Cryptwerk
  • 5.0/5Bare Metal
    Very good experience using their bare metal servers. Their customer service is one of the finest I have experienced - always prompt at resolving troubles. Highly recommend.

Deep Dive & FAQs

Availability and reservations, H100 NVL vs H100 SXM and L40S, NVLink bridge scaling, pricing, software, power, and MIG.

When will H100 NVL servers be available and how does reservation work?

H100 NVL capacity is being installed across our 20+ Tier III data centers in the EU, US and Asia now. Submit the reservation form with your preferred card count and region; our engineers confirm the lead time for that build, hold the hardware for you, and send a quote. Reservations are free and non-binding until you approve the quote.

H100 NVL vs H100 SXM: which should I pick?

Both use the Hopper GH100 GPU, but the packaging differs. H100 NVL is a dual-slot PCIe Gen5 card with 94 GB HBM3 at 3.9 TB/s, 350 to 400 W, and a 600 GB/s NVLink bridge between pairs of cards; H100 SXM is an HGX module with 80 GB HBM3 at 3.35 TB/s, 700 W, 1,979 TFLOPS FP16 Tensor with sparsity, and 900 GB/s NVLink across all 8 GPUs. Pick H100 NVL for inference, RAG and fine-tuning in air-cooled PCIe servers where 1 to 8 cards and more memory per card matter; pick H100 SXM when you train across all 8 GPUs at once and need the full NVSwitch fabric.

H100 NVL vs L40S: when is the step up worth it?

L40S has 48 GB GDDR6 at 864 GB/s and 733 TFLOPS FP16 Tensor at 350 W; H100 NVL roughly doubles the memory to 94 GB HBM3, delivers 4.5x the bandwidth at 3.9 TB/s and 2.3x the FP16 Tensor throughput at 1,671 TFLOPS, at a similar TDP. Memory bandwidth bounds LLM decode speed, so H100 NVL serves large models with far more tokens per second per card. L40S remains the better value for graphics, rendering and smaller models that fit comfortably in 48 GB.

How is H100 NVL priced and are there discounts?

Pricing is quoted per configuration because card count, host CPU, memory, storage and uplink choices vary widely. Volume discounts apply to multi-GPU and multi-server orders, and every GPU server qualifies for committed-term discounts on 6 and 12-month billing. Crypto, card, PayPal and bank wire are accepted.

How does the NVLink bridge work and how far does multi-card scaling go?

The NVLink bridge joins two H100 NVL cards at 600 GB/s, compared with 128 GB/s over PCIe Gen5, so a pair shares 188 GB of HBM3 for tensor-parallel inference of Llama 70B class models in FP8 with minimal interconnect overhead. Bridges link cards in pairs; a 4x or 8x node is built as two or four bridged pairs that communicate over PCIe, which suits data-parallel serving and fine-tuning well. If your workload needs all-to-all NVLink across 8 GPUs, choose H100 SXM instead.

Is my existing CUDA / PyTorch code compatible, and what software is included?

Yes. H100 NVL runs the standard CUDA programming model, so CUDA 12 builds, PyTorch, JAX, TensorFlow, TensorRT-LLM and Triton workloads run unchanged, and the Transformer Engine adds FP8 when you enable it. NVIDIA states that H100 NVL includes an NVIDIA AI Enterprise subscription, covering NIM microservices, NeMo and enterprise-supported frameworks. We install Ubuntu, Debian, Rocky/AlmaLinux, Windows Server or your own ISO, with the NVIDIA data center driver, CUDA toolkit and container runtime preinstalled on request.

What about power and cooling?

Each H100 NVL card draws 350 to 400 W (configurable) and is air-cooled in a dual-slot PCIe form factor, roughly half the 700 W of an H100 SXM module. An 8x node therefore fits standard high-density racks without liquid cooling, and our GPU racks are provisioned with the power feeds and airflow each configuration needs. You plan the workload, not the facilities.

Does H100 NVL support MIG and Confidential Computing?

Yes. Each H100 NVL can be partitioned into up to 7 MIG instances of 12 GB each, with isolated compute and memory, so one card can serve several smaller models or tenants. Hopper Confidential Computing protects data and models in GPU memory inside a trusted execution environment, useful for private inference and regulated workloads.

Not Sure Exactly What You Need

No problem. Our talented engineers will consult, architect, migrate, manage, and do whatever it takes to help your business grow and succeed.