Skip to main content

GPU Product

Rent NVIDIA A100 80GB GPUs

80 GB of HBM2e at 2,039 GB/s, up to seven hardware-isolated MIG instances, and NVLink 3.0 for multi-GPU scale-out. The proven workhorse for large-model training, high-throughput inference, and HPC — from $1.05/hour.

Technical Specifications

ArchitectureNVIDIA Ampere
Memory Size80 GB HBM2e
Memory Bandwidth2 039 GB/s
CUDA Cores6 912
Tensor Cores432 (3rd gen)
NVIDIA A100 80GB datacenter GPUs
Rental Options

A100 Rental Options

Rent NVIDIA A100 80GB SXM4 GPUs from $1.05/hour. Each GPU ships with 15 CPU cores, 210 GB of RAM, and 800 GB of NVMe, and nodes scale to 8 GPUs over NVLink 3.0. Hourly billing with no minimum commitment, plus 1-month and 3-month reservations for committed workloads.

Comparison

A100 80GB vs V100 32GB

The A100 is the direct generational successor to the V100, and both live in the same CloudRift datacenter — so this is the upgrade decision most tenants actually face. For roughly 3.6× the hourly rate you get 2.5× the memory, 2.3× the bandwidth, and 2.5× the Tensor throughput, plus TF32, BF16, MIG partitioning, and NVLink 3.0. The V100 remains the better value when your model fits in 32 GB; the A100 earns its premium the moment it does not.

Comparison dimensionA100 80GBV100 32GB% Diff
ArchitectureAmpere (GA100)Volta (GV100)N/A
Memory TypeHBM2e ECCHBM2 ECCN/A
VRAM80 GB32 GB+150%
Bus Width5 120-bit4 096-bit+25%
Memory Bandwidth2 039 GB/s900 GB/s+126.6%
FP64 Performance9.7 TFLOPS~7.8 TFLOPS+24.4%
FP64 Tensor Core19.5 TFLOPSNot supportedN/A
FP32 Performance19.5 TFLOPS~15.7 TFLOPS+24.2%
TF32 Tensor Core156 TFLOPSNot supportedN/A
BF16/FP16 Tensor312 TFLOPS~125 TFLOPS+149.6%
CUDA Cores6 9125 120+35%
Tensor Cores432 (3rd gen)640 (1st gen)−32.5%
Multi-Instance GPUUp to 7 MIGNot supportedN/A
Multi-GPU InterconnectNVLink 3.0NVLink 2.0N/A
Form FactorSXM4SXM3N/A
TDP400 W350 W+14.3%

Performance

Key performance metrics

HBM2e Memory Bandwidth

2 039 GB/s of HBM2e bandwidth across an 5 120-bit bus — more than double the V100 and enough to keep 70B-class models fed during high-batch inference and training.

Third-Generation Tensor Cores

432 third-generation Tensor Cores deliver 312 TFLOPS of BF16/FP16 and 156 TFLOPS of TF32 — TF32 accelerates existing FP32 training code with no changes to your model.

Multi-Instance GPU (MIG)

Partition a single A100 into up to seven hardware-isolated instances, each with dedicated memory, cache, and compute — run separate tenants or jobs without noisy-neighbour interference.

Use Cases

What the A100 Is Built For

Large-Model Training & Fine-Tuning

80 GB of HBM2e holds 30B–70B parameter models for LoRA and QLoRA runs, or full fine-tunes of 13B-class models — with NVLink 3.0 for multi-GPU scale-out.

High-Throughput LLM Inference

Token generation is memory-bandwidth bound. At 2 039 GB/s the A100 sustains high concurrency on vLLM and TensorRT-LLM without spilling to host memory.

Multi-Tenant Serving

MIG carves one A100 into up to seven isolated instances — serve several models or customers from a single card with guaranteed quality of service.

HPC & Scientific Computing

9.7 TFLOPS of native FP64 — and 19.5 TFLOPS through the FP64 Tensor Cores — for CFD, molecular dynamics, genomics, and computational chemistry.

NVIDIA A100 FAQ

Common Questions About the A100

A100 80GB rentals on CloudRift start at $1.05/hour for a single GPU. Pay-as-you-go, no minimum commitment, and nodes scale up to 8 GPUs.
The A100 is an Ampere-architecture datacenter GPU used for large-model training and fine-tuning, high-throughput LLM inference, multi-tenant model serving via MIG, and HPC workloads such as CFD, molecular dynamics, and genomics. Its 80 GB of HBM2e and 2,039 GB/s of bandwidth make it a strong fit for memory-bound and bandwidth-bound work.
The A100 80GB has 80 GB of HBM2e memory on a 5,120-bit bus delivering 2,039 GB/s of bandwidth — more than twice the bandwidth of the V100 32GB and roughly 2.4× that of a GDDR6 card like the L40.
MIG lets you partition a single A100 into as many as seven hardware-isolated GPU instances, each with its own memory, cache, and streaming multiprocessors. Because the isolation is enforced in hardware rather than by the scheduler, one tenant cannot affect another’s throughput — which makes it well suited to multi-tenant inference and shared research clusters.
The A100 80GB has 2.5× the VRAM (80 vs 32 GB), 2.3× the memory bandwidth (2,039 vs 900 GB/s), roughly 2.5× the Tensor throughput (312 vs ~125 TFLOPS), and adds TF32, BF16, MIG, and NVLink 3.0. The V100 costs about a third as much per hour at $0.29 vs $1.05. Choose the V100 for smaller memory-bound jobs on a tight budget, and the A100 when the model does not fit in 32 GB or when memory bandwidth is your bottleneck.
An A100 80GB on CloudRift is $1.05/hr. The same silicon on AWS p4de.24xlarge runs roughly $5.12/hr per GPU and on Azure NDm A100 v4 roughly $4.10/hr per GPU, so CloudRift is about 4–5× cheaper for equivalent hardware.
A single A100 80GB comfortably serves models up to about 70B parameters in 4-bit quantization, or roughly 34B in FP16. For full-precision fine-tuning it handles 13B-class models on one card, and multi-GPU nodes with NVLink 3.0 scale to larger runs.
Yes. Open the Console, pick an A100 instance, and launch — most instances are running in under a minute. Hourly billing, no long-term commitment.
Get in touch

Ready to get started?

Get in touch with our team to discuss your requirements and find the right solution for your infrastructure.