Skip to main content

GPU Product

Rent NVIDIA V100 GPUs

16 GB or 32 GB of HBM2 at 900 GB/s — the cost-leader for memory-bound AI fine-tuning, HPC, and batch inference. From $0.25/hour on sustainable, second-life Volta silicon.

Technical Specifications

ArchitectureNVIDIA Volta
Memory Size16 or 32 GB HBM2
Memory Bandwidth900 GB/s
CUDA Cores5 120
Tensor Cores640 (1st gen)
NVIDIA V100 SXM datacenter GPU module
Rental Options

V100 Rental Options

Rent NVIDIA V100 GPUs from $0.25/hour — 16GB SXM2 or 32GB SXM3. Hourly billing with no minimum commitment, plus 1-month and 3-month reservations for committed workloads. Capacity sourced from sustainable second-life datacenter hardware.

Comparison

V100 32GB vs L40S

Both are datacenter GPUs in the CloudRift fleet, positioned for different workloads. The table below uses the 32 GB V100; the 16 GB part is the same Volta silicon with half the memory, at $0.25/hr. The V100 wins on price (roughly half the hourly rate), HBM2 memory bandwidth, FP64 throughput, and NVLink for multi-GPU scale-out. The L40S wins on raw FP32, modern Ada Tensor Cores, and total VRAM — choose it when you need maximum inference throughput per card.

V100 32GBL40S% Diff
CloudRift Price$0.29 / hr$0.57 / hr reserved−49%
ArchitectureVoltaAda LovelaceN/A
Memory TypeHBM2 ECCGDDR6 ECCN/A
VRAM32 GB48 GB−33.3%
Bus Width4 096-bit384-bit+966%
Memory Bandwidth900 GB/s864 GB/s+4.2%
FP64 Performance~7.8 TFLOPS~1.4 TFLOPS+457%
FP32 Performance~15.7 TFLOPS~90.5 TFLOPS−82.7%
Tensor Performance~125 TFLOPS~362 TFLOPS−65.5%
CUDA Cores5 12018 176−71.8%
Tensor Cores640 (1st gen)568 (4th gen)N/A
Multi-GPU InterconnectNVLink 2.0PCIe 4.0 onlyN/A
Form FactorSXM3Dual-slot PCIeN/A
TDP350 W350 W0%

Performance

Key performance metrics

HBM2 Memory Bandwidth

900 GB/s of HBM2 bandwidth on a 4 096-bit bus — the cost-leader for memory-bound workloads like scientific compute, mid-sized LLM fine-tuning, and large-batch inference.

Tensor Core Acceleration

640 first-generation Tensor Cores deliver up to ~125 TFLOPS of mixed-precision throughput — proven silicon for CUDA, cuDNN, and mature ML toolchains.

A Fraction of Hyperscaler Pricing

$0.25/hr on CloudRift for the 16GB part and $0.29/hr for the 32GB, vs ~$3.90/hr per GPU on AWS p3dn.24xlarge and ~$2.75/hr per GPU on Azure NDv2 — up to 13× cheaper than hyperscalers for the same V100 silicon.

NVIDIA V100 FAQ

Common Questions About the V100

V100 rentals on CloudRift start at $0.25/hour for a 16GB SXM2 GPU, or $0.29/hour for a 32GB SXM3. Pay-as-you-go, no minimum commitment.
The V100 is a Volta-architecture datacenter GPU used for AI fine-tuning, batch inference, and HPC workloads like CFD, molecular dynamics, and computational chemistry. Its HBM2 memory — 16 GB or 32 GB depending on the configuration — and strong FP64 performance make it cost-efficient for memory-bound and double-precision workloads.
CloudRift runs two V100 configurations: 16 GB HBM2 on the SXM2 part and 32 GB HBM2 on the SXM3. Both sit on a 4 096-bit bus with 900 GB/s of bandwidth — meaningfully higher than GDDR-class cards in the same price tier.
Yes — for the right workload. The V100 has 640 first-generation Tensor Cores and HBM2 memory: the 16GB part suits 7B-class fine-tuning and batch inference, and the 32GB part stretches to 13B. Both are strong on memory-bandwidth-bound HPC. For training the largest frontier models you want H100/H200 or MI350X, but the V100 remains the cost-leader for many production workloads.
A V100 32GB on CloudRift is $0.29/hr, and the 16GB part is $0.25/hr. The same 32GB silicon on AWS p3dn.24xlarge runs ~$3.90/hr per GPU, and on Azure NDv2 ~$2.75/hr per GPU — CloudRift is roughly 10–13× cheaper. Capacity comes from sustainable second-life datacenter hardware, redeployed instead of replaced.
The V100 32GB is $0.29/hr on demand. The L40S is reserved-only, from $0.57/hr on a 1-month reservation, so the V100 is roughly half the price. It has HBM2 memory (900 vs 864 GB/s), a much wider 4 096-bit bus, NVLink 2.0 for multi-GPU, and about 5× the FP64 throughput (~7.8 vs ~1.4 TFLOPS). The L40S has 50% more VRAM (48 GB), newer Ada Tensor Cores, and roughly 6× the FP32 throughput, so it is the better pick when you need maximum inference throughput per card or modern Ada-only features. Pick the V100 for memory-bandwidth-bound workloads, FP64-heavy HPC, multi-GPU training, or the lowest hourly rate. Pick the V100 for memory-bandwidth-bound workloads, FP64-heavy HPC, multi-GPU training, or the lowest hourly rate.
Yes. Open the Console, pick a V100 instance, and launch — most instances are running in under a minute. Hourly billing, no long-term commitment.
Get in touch

Ready to get started?

Get in touch with our team to discuss your requirements and find the right solution for your infrastructure.