NVIDIA L40S Tensor Core GPU Server | RedSwitches
// gpu compute

NVIDIA L40S Tensor Core GPU Server

Universal GPU Power for AI Workloads, Real-Time Rendering, VDI & Edge Graphics. All in a Single-Slot PCIe Gen 4 Card.

  • Online in 1 hour when in stock
  • 20+ Tier III global data centers
  • Unmetered 1/10/25 Gbps bandwidth
  • Crypto payments & 24/7 support included
Starting FromLive Pricing
Low-Cost L40S BuildNVIDIA L40S
CPU
2x AMD EPYC 7413
Cores
48C / 96T
RAM
128 GB
Storage
2x960GB SSD
Network
1 Gbps · 100 TB
Location
Montreal, Canada
$760.79/moDeploy Now

NVIDIA L40S Tensor Core GPU Server Price

Single and multi-GPU L40S builds on bare metal. No setup fee, full root access, and unmetered 1/10/25 Gbps uplinks.

Filters

GPU Dedicated Servers

Enterprise NVIDIA GPUs on bare metal for AI and HPC.

GPUNVIDIA L40S24 GB · 7,424 CUDAVRAM24 GBCUDA Cores7,424Tensor Cores240TFLOPS30.3AI, Video
Server
RecommendedTop Deal · 50% Off1 Hour
2x AMD EPYC 741348C / 96T · 2.65 GHz
Memory128 GB
Storage2x960GB SSD
Network1 Gbps100 TB
LocationMontrealCanada
Price
$760.79/mo
Deploy Now
GPUNVIDIA L40S24 GB · 7,424 CUDAVRAM24 GBCUDA Cores7,424Tensor Cores240TFLOPS30.3AI, Video
Server
RecommendedTop Deal · 50% Off1 Hour
2x AMD EPYC 754364C / 128T · 2.8 GHz
Memory128 GB
Storage2x960GB SSD
Network1 Gbps100 TB
LocationMontrealCanada
Price
$841.41/mo
Deploy Now
GPUNVIDIA L40S24 GB · 7,424 CUDAVRAM24 GBCUDA Cores7,424Tensor Cores240TFLOPS30.3AI, Video
Server
RecommendedTop Deal · 50% OffNew Gen
2x AMD EPYC 922448C / 96T · 2.5 GHz
Memory128 GB
Storage2x960GB SSD
Network1 Gbps100 TB
LocationMontrealCanada
Price
$1003.78/mo
Deploy Now
GPUNVIDIA L40S24 GB · 7,424 CUDAVRAM24 GBCUDA Cores7,424Tensor Cores240TFLOPS30.3AI, Video
Server
RecommendedTop Deal · 50% OffNew Gen
2x AMD EPYC 933464C / 128T · 2.7 GHz
Memory128 GB
Storage2x960GB SSD
Network1 Gbps100 TB
LocationMontrealCanada
Price
$1083.27/mo
Deploy Now
GPUNVIDIA L40S24 GB · 7,424 CUDAVRAM24 GBCUDA Cores7,424Tensor Cores240TFLOPS30.3AI, Video
Server
RecommendedTop Deal · 50% Off
2x AMD EPYC 754364C / 128T · 2.8 GHz
Memory128 GB
Storage2x960GB SSD
Network1 Gbps30 TB
LocationLondonUnited Kingdom
Price
$1209.31/mo
Deploy Now
GPUNVIDIA L40S24 GB · 7,424 CUDAVRAM24 GBCUDA Cores7,424Tensor Cores240TFLOPS30.3AI, Video
Server
RecommendedTop Deal · 50% OffNew Gen
2x AMD EPYC 922448C / 96T · 2.5 GHz
Memory128 GB
Storage2x960GB SSD
Network1 Gbps100 TB
LocationFrankfurtGermany
Price
$1426.19/mo
Deploy Now
GPUNVIDIA L40S24 GB · 7,424 CUDAVRAM24 GBCUDA Cores7,424Tensor Cores240TFLOPS30.3AI, Video
Server
RecommendedTop Deal · 50% Off
2x AMD EPYC 741348C / 96T · 2.65 GHz
Memory128 GB
Storage2x960GB SSD
Network1 Gbps30 TB
LocationSingaporeSingapore
Price
$1545.42/mo
Deploy Now
GPU2x NVIDIA L40S24 GB/GPU · 7,424 CUDAVRAM24 GB / GPUCUDA Cores7,424Tensor Cores240TFLOPS30.3AI, Video
Server
RecommendedTop Deal · 50% OffNew Gen
2x AMD EPYC 933464C / 128T · 2.7 GHz
Memory128 GB
Storage2x960GB SSD
Network1 Gbps100 TB
LocationMontrealCanada
Price
$1555.64/mo
Deploy Now
GPUNVIDIA L40S24 GB · 7,424 CUDAVRAM24 GBCUDA Cores7,424Tensor Cores240TFLOPS30.3AI, Video
Server
RecommendedTop Deal · 50% OffNew Gen
2x AMD EPYC 933464C / 128T · 2.7 GHz
Memory128 GB
Storage2x960GB SSD
Network1 Gbps30 TB
LocationSingaporeSingapore
Price
$1906.50/mo
Deploy Now
GPU2x NVIDIA L40S24 GB/GPU · 7,424 CUDAVRAM24 GB / GPUCUDA Cores7,424Tensor Cores240TFLOPS30.3AI, Video
Server
RecommendedTop Deal · 50% OffNew Gen
2x AMD EPYC 933464C / 128T · 2.7 GHz
Memory128 GB
Storage2x960GB SSD
Network1 Gbps30 TB
LocationSingaporeSingapore
Price
$2784.25/mo
Deploy Now
Bare Metal // Standard Equipment

All Bare Metal Plans Include

Every server ships fully dedicated: no shared resources, no usage meters, no surprises.

Setup Cost
Free
Provisioning
Instant & Automated
Access
KVM, IPMI, Root
Protection
DDoS Shield Included
Cores
Up to 128 (Dual Socket)
Memory
Up to 2TB RAM
Storage
Enterprise NVMe & SSD
Support
24/7/365 Human Engineers
OS & Panels
  • Ubuntu
  • Debian
  • AlmaLinux
  • Rocky Linux
  • CentOS
  • Windows Server
  • cPanel
  • Plesk
  • Proxmox
  • Docker

* Select your OS and panel at checkout

What You Get With Every NVIDIA L40S Server

Every RedSwitches NVIDIA L40S server is single-tenant bare metal with full root and IPMI access, unmetered 1, 10 or 25 Gbps bandwidth, no setup fee, no bandwidth overage and a flat monthly price. Stocked builds are online in about 1 hour, larger configurations up to 8 GPUs per node are built to order, volume and 6/12-month committed-term discounts apply, payments include crypto, and engineers answer 24/7.

  • 1 hr

    Online When in Stock

    Stocked NVIDIA L40S builds are online in about 1 hour. Configurations not in stock are built to order, and an engineer confirms the lead time before you commit.

  • $0

    Setup Fee, Flat Monthly

    No setup fee and a flat monthly price for the whole server. No per-hour meter and no surprise line items, so GPU spend is forecastable.

  • Unmetered

    Bandwidth, No Egress Bills

    Unmetered 1, 10 or 25 Gbps uplinks are included, and whatever port speed you choose you can use all of it: no overage charges and no egress bills, ever. Moving datasets, checkpoints and model weights in and out costs nothing extra.

  • 1 tenant

    Bare Metal, Full Control

    Single-tenant hardware with root and IPMI access. You choose the OS, drivers, CUDA or ROCm version, and the NVLink or MIG layout, and a private VLAN can link your RedSwitches servers.

  • Up to 8x

    Multi-GPU, Built to Order

    Single and multi-GPU nodes, up to 8 GPUs per server with NVLink where the card supports it. Volume discounts on multi-GPU and multi-server orders, plus 6 and 12-month term savings.

  • 24/7

    Engineers, Not Bots

    Live chat, Telegram and email answered by engineers around the clock. Pay by card, PayPal, bank wire or crypto with no KYC, in 20+ Tier III data centers across the EU, US and Asia.

RedSwitches L40S Server vs a Typical Cloud GPU Instance

Same accelerator, different economics: what changes when the GPU sits in a dedicated server you control instead of a metered instance.

RedSwitches NVIDIA L40S dedicated server versus a typical hyperscale cloud GPU instance
RedSwitches L40S serverTypical cloud GPU instance
BillingFlat monthly price per serverPer-hour or per-second metering
BandwidthUnmetered 1/10/25 Gbps, use the full port, no overage or egress feesEgress billed per GB
Private networkingPrivate VLAN between your servers on requestPaid VPC and peering constructs
Setup fee$0Varies by instance and region
TenancySingle-tenant bare metalShared, virtualised hosts
AccessRoot and IPMI, your OS and driversHypervisor-managed images
Multi-GPUUp to 8 per node, NVLink where supported, built to orderFixed instance shapes
PaymentsCard, PayPal, bank wire, cryptoCard or invoice
Support24/7 engineers on chat, Telegram and emailTicket tiers, paid support plans

NVIDIA L40S Key Specifications

Ada Lovelace silicon with 48 GB ECC GDDR6, 4th-gen Tensor Cores, and 3rd-gen RT Cores in a single-slot card.

Architecture
NVIDIA Ada Lovelace with 4th-gen Tensor Cores & 3rd-gen RT Cores
Compute Performance
FP32
91.6 TFLOPS
Tensor (FP8)
1,466 TFLOPS (peak w/ sparsity)
RT Cores
212 TFLOPS
Memory
48 GB GDDR6, 864 GB/s, ECC, 384-bit
CUDA Cores
18,176; Tensor Cores: 568; RT Cores: 142
Clock Speeds
Base ~1,110 MHz → Boost up to ~2,520 MHz
Formfactor
Dual-slot PCIe 4.0 ×16, passive cooling, 300 W TDP
Virtualization & Security
ECC, SR-IOV (256 VFs), NEBS-3, Secure Boot, Root of Trust

Why Choose L40S

One universal GPU for AI, graphics, media, and VDI, with top-tier inference and real-time ray tracing.

Universal performance

Designed for AI, graphics, media, and VDI, one GPU for many workloads.

Top-tier AI inference

Up to 1.5× faster than A100 for inference tasks and 5× over A40.

Real-time ray tracing & DLSS

Ideal for Omniverse, CAD, simulation, and virtual workstations.

High-density deployment

PCIe single-slot design enables up to 8 GPUs per server.

Enterprise reliability

24/7-ready with ECC memory, passive cooling, and compliance standards.

Targeted Use Cases

From generative AI to virtual workstations, see where the L40S changes what your team can ship.

Generative AI & LLM Inference

Ideal for BERT, GPT, Stable Diffusion, embeddings, with fast FP8 throughput.

High-Performance Graphics & Rendering

Supports Omniverse, RT/VR workloads, 2× faster ray tracing vs A40.

AI Video & Media Pipelines

Features strong NVENC/NVDEC support for real-time AV1/H.264/HEVC streaming.

VDI & Virtual Workstations

48 GB memory + RT cores + DLSS deliver seamless user experience.

Edge & Vision AI

Dense inference clusters for computer vision pipelines, efficient and low-power.

Mixed-AI/HPC Workloads

Great fit for small AI training, data analytics, and scientific acceleration

Trusted by Enterprise Teams Worldwide

  • Check Point
  • German Football Association
  • Mubi
  • Pluxee
  • Zeeve
  • University of Malta
  • mSpy
  • RevX
  • Turbo VPN
  • Athos Commerce
  • Heckyl
  • GenXAI
  • WLVPN
  • FMS
  • Contaque
  • Monotek
  • EasyGo VPN
  • Neopool
  • InfyGlobe Technologies
  • ALFA University College
  • Stief Group
  • SSH Invest Holding
  • Rhysley
  • Spirit of Math
  • VideoShip
  • ORB VPN
  • Ping VPN

FAQs & Deep Dive

Common questions about L40S density, thermals, virtualization, and PCIe compatibility.

How many L40S GPUs per server?

Supports up to 8 x L40S in PCIe-dense chassis, perfect for compact GPU clusters.

Is 300 W TDP manageable?

Yes, passive cooling suits data-center racks; airflow-optimized design handles heat efficiently.

Can it virtualize workloads?

Yes, supports SR-IOV with up to 256 virtual functions for VDI or tenant-based use.

Better than L4 or T4?

L40S offers ~3× more compute and ~12× memory vs L4, and ~1.5× inference over A100.

Does it fit PCIe4 environments?

Yes, requires PCIe4 ×16; backward compatible with PCIe3 at reduced speed.

Not Sure Exactly What You Need

No problem. Our talented engineers will consult, architect, migrate, manage, and do whatever it takes to help your business grow and succeed.