Skip to main content

GPU Product

Rent NVIDIA L40S GPUs

48 GB of GDDR6 ECC, 733 TFLOPS of FP8 Tensor throughput, and full vGPU support in a passive 350 W dual-slot card. The Ada Lovelace workhorse for AI inference, fine-tuning, and virtual workstations.

Technical Specifications

ArchitectureNVIDIA Ada Lovelace
Memory Size48 GB GDDR6 ECC
Memory Bandwidth864 GB/s
Ray Tracing Cores142 (3rd gen)
Tensor Cores568 (4th gen)
NVIDIA L40S datacenter GPU
Rental Options

L40S Rental Options

Rent NVIDIA L40S GPUs for production AI workloads, virtualized graphics, and enterprise applications. Get the reliability and features of data center GPUs with flexible on-demand pricing.

Comparison

L40S vs RTX 4090

L40SRTX 4090% Diff
ArchitectureAda LovelaceAda LovelaceN/A
CUDA Cores18 17616 384+10.9%
Tensor Cores568 (4th gen)512 (4th gen)+10.9%
RT Cores142 (3rd gen)128 (3rd gen)+10.9%
Memory TypeGDDR6 ECCGDDR6XN/A
VRAM48 GB24 GB+100%
Bus Width384-bit384-bit0%
Bandwidth864 GB/s1 010 GB/s−14.5%
FP32 Performance91.6 TFLOPS82.6 TFLOPS+10.9%
BF16/FP16 Tensor362 TFLOPS330 TFLOPS+9.6%
FP8 Tensor733 TFLOPS661 TFLOPS+11.0%
TDP350 W450 W−22.2%
PCIePCIe 4.0 ×16PCIe 4.0 ×16N/A
Form FactorDual-slotTriple-slotN/A
CoolingPassiveActiveN/A
Display Outputs4× DP 1.4a3× DP 1.4aN/A
vGPU SupportYesNoN/A

Performance

Key performance metrics

Enterprise AI Workloads

Optimized for AI inference and training with 18,176 CUDA cores delivering ~90.5 TFLOPS FP32 performance. Perfect for production-scale AI deployments.

Virtualization Ready

Built for multi-tenant cloud environments with NVIDIA vGPU support, enabling secure workstation virtualization and remote graphics workloads.

Professional Graphics

48GB ECC memory and advanced encoding (3x NVENC/NVDEC with AV1) enable real-time ray tracing, 8K video workflows, and complex 3D scene rendering.

Use Cases

What the L40S Is Good At

AI Inference

Deploy production AI models with high-throughput inference for real-time apps.

Workstation VMs

Power remote design and engineering teams with GPU-accelerated graphics pipelines.

Rendering & VFX

Real-time ray tracing and high-resolution rendering for visual-effects production.

Multi-Tenancy

Serve multiple users with secure, isolated GPU resources at high concurrency.

NVIDIA L40S FAQ

Common Questions About the L40S

The NVIDIA L40S is a data center GPU designed for AI inference, model fine-tuning, virtual workstations, 3D rendering, and multi-tenant cloud deployments. It features 48GB GDDR6 ECC memory, 733 TFLOPS of FP8 Tensor performance, vGPU support, and enterprise-grade reliability.
The L40S and L40 share the same Ada Lovelace silicon: 18,176 CUDA cores, 568 4th-gen Tensor cores, 142 3rd-gen RT cores, and 48GB of GDDR6 ECC memory at 864 GB/s. The L40S is tuned for AI rather than graphics — it runs at higher clocks with a 350W TDP (versus 300W) and delivers roughly double the Tensor throughput, at 362 TFLOPS BF16/FP16 and 733 TFLOPS FP8 against the L40's 181 and 362. Choose the L40S for inference and fine-tuning; the L40 was positioned more for visualization and vGPU density.
The NVIDIA L40S has 48 GB of GDDR6 memory with ECC (Error-Correcting Code) on a 384-bit bus at 864 GB/s. That is enough to serve most 70B-parameter models quantized to FP8, or to fine-tune mid-size models without sharding across GPUs.
The A100 80GB has more memory (80GB vs 48GB) and far more bandwidth (2,039 GB/s vs 864 GB/s), which makes it stronger for large-model training and memory-bound workloads. The L40S wins on FP8 — a format the Ampere-based A100 does not support at all — and adds RT cores plus 3× NVENC/NVDEC with AV1 for rendering and video pipelines. For FP8 inference and mixed graphics/AI workloads the L40S is usually the better value; for large-scale training the A100 is the stronger choice.
Yes, the L40S supports NVIDIA vGPU technology for virtual workstations (vWS) and virtual PCs/apps, making it ideal for multi-tenant cloud environments and remote graphics workloads. Note that unlike the A100 it does not support MIG, so partitioning is done through vGPU rather than hardware instance slicing.
Key features include 18,176 CUDA cores, 568 4th-gen Tensor cores, 142 3rd-gen RT cores, 48GB GDDR6 ECC memory, 864 GB/s bandwidth, 91.6 TFLOPS FP32, 733 TFLOPS FP8 Tensor, PCIe 4.0 x16, a passive dual-slot 350W design, and 3x NVENC/NVDEC with AV1 support.
Yes — the L40S is in our catalog with flexible hourly pricing. On-demand stock varies, so check the Console for current availability, or contact us to reserve capacity.
Get in touch

Ready to get started?

Get in touch with our team to discuss your requirements and find the right solution for your infrastructure.