Large language model inference is racing through a quantization gauntlet. NVIDIA's Blackwell chips reportedly deliver three times the token throughput of their Hopper predecessors by dropping from 8-bit floating-point (FP8) to 4-bit floating-point (FP4) precision, according to TensorFoundry's 2026 field guide. Apple's M5 silicon added native FP8 and INT4 support to its GPU cores the same year. The shift is not about Moore's Law. It is about memory bandwidth: moving weights and activations from DRAM to compute units burns more energy and time than the matrix multiplications themselves, so shrinking the data payload by half or three-quarters changes the economics of every inference call.
The transition from FP8 to 4-bit integer (INT4) quantization is uneven. FP8 delivers near-lossless accuracy with modest speedups, while INT4 cuts memory usage by 75 percent but degrades some tasks sharply. Production deployments must now choose between three incompatible optimization targets: minimize latency, maximize throughput, or maximize concurrent users. The answer depends on whether your bottleneck is compute, memory bandwidth, or memory capacity, and the frameworks racing to support every format are converging on a messy middle ground where the optimal choice changes with model size, hardware generation, and workload mix.
The quantization ladder and where each rung lands
Quantization compresses model weights and activations from 16-bit or 32-bit representations down to 8-bit or 4-bit formats. NVIDIA's TensorRT-LLM documentation establishes the performance hierarchy: FP8 quantization delivers 1.4 to 1.5 times the throughput of FP16 on Hopper GPUs, while INT4 weight-only quantization (W4A16, where weights are 4-bit but activations remain 16-bit) reaches 2.7 times baseline speed. The gap widens under memory pressure. When first-token latency is constrained to 500 milliseconds on a LLaMA-v2-7B model at batch size 16, FP8 quantization achieves a 2.3 times speedup by keeping more of the model resident in fast memory.
Accuracy costs vary by format and calibration method. Digital Applied's survey of six 2026 open-weight 70-billion-parameter models, including Llama 4 70B, Qwen 3 72B, and DeepSeek V4-Flash, found FP8 quantization landed within 0.4 points of FP16 on MMLU-Pro and HumanEval+ benchmarks. INT4 quantization using the AWQ (Activation-aware Weight Quantization) method showed a 1.6-point delta, while GPTQ (Generative Pre-trained Transformer Quantization) showed 1.9 points. NVIDIA's Falcon-180B tests recorded a 0.14 percent MMLU loss with FP8 (70.4 to 70.3), a 2.56 percent loss with INT8 SmoothQuant (70.4 to 68.6), and a 0.85 percent loss with INT4-AWQ (70.4 to 69.8).
Task sensitivity complicates the picture. AIMultiple's Qwen3-32B benchmark showed INT4 quantization caused less than a 2-point MMLU-Pro drop but an 8-point HumanEval decline (39.02 percent to 31.02 percent). Code generation degrades more sharply than knowledge retrieval under aggressive quantization, a pattern that forces production teams to maintain multiple quantized checkpoints for different workload classes.
Memory capacity versus memory bandwidth
The quantization roadmap splits along two axes: memory bandwidth (how fast data moves) and memory capacity (how much fits). Weight-only INT4 quantization (W4A16) accelerates memory-bound decode phases by shrinking the weight tensor from 2 bytes per parameter to 0.5 bytes, but leaves 16-bit activations untouched. Full INT4 or FP8 quantization (W4A8 or W8A8) compresses both weights and activations, speeding up compute-bound prefill phases where all input tokens are processed in parallel.
AIMultiple's Qwen3-32B tests on an H100 GPU demonstrated that INT4 quantization freed 47.3 gigabytes of VRAM compared to 4.4 gigabytes in BF16 mode, enabling a 12-times increase in concurrent users (47 versus 4 at 4-kilotoken context length). The throughput gain was 2.7 times, but the concurrency gain mattered more in production because it allowed the same hardware to serve an order of magnitude more requests without queuing. Memory capacity, not compute speed, determined the deployment's revenue per GPU.
KV cache quantization extends the same logic to the attention mechanism's key-value pairs, which grow linearly with sequence length and batch size. NVIDIA's documentation reports that switching from FP16 to FP8 KV cache on H100 GPUs enables 2 to 3 times larger batch sizes for models like GPT-J, delivering an additional 1.5 times performance benefit on top of weight and activation quantization. The recommendation is explicit: use FP8 KV cache over INT8 on Hopper and Ada GPUs because FP8 shows lower accuracy impact in most tested cases.
Framework support and the hybrid kernel race
Inference frameworks are converging on a common set of quantization formats but diverging in implementation quality. LLM Pioneer Hub's production guide documents that vLLM supports AWQ, GPTQ, bitsandbytes, FP8, INT4, INT8, GGUF, TorchAO, and quantized KV cache as of 2026, while TensorRT-LLM supports FP4, FP8, AWQ, GPTQ, and NVFP4 KV cache. The format zoo reflects competing optimization strategies: AWQ and GPTQ calibrate quantization parameters using small datasets to minimize accuracy loss, while simpler round-to-nearest methods trade calibration time for faster deployment.
Hybrid quantization kernels are emerging as a middle ground. FireQ's research paper introduced an INT4-FP8 hybrid kernel that achieved 1.68 times the speed of QServe on Llama2-7B feed-forward network layers and 1.26 times on Llama3-8B prefill (batch size 16, sequence length 1024) on H100 GPUs. The technique uses RoPE-aware outlier smoothing to handle activation spikes that destabilize low-bit quantization, a problem that becomes acute in rotary position embeddings. The result is a kernel that runs INT4 weights against FP8 activations, splitting the difference between W4A16 and W8A8 formats.
TensorFoundry's field guide identifies MXFP8 (microscaling FP8) as the converging datacenter baseline, a format that uses per-block scaling factors to preserve dynamic range without per-tensor calibration. NVIDIA's Blackwell architecture reportedly delivers approximately three times the token throughput of Hopper's FP8 by dropping to native FP4, though the claim cites DeepSeek-R1 workloads without primary benchmark data.
Production decision trees and the memory vendor angle
Choosing a quantization format in production depends on three variables: model size, hardware generation, and workload mix. Dreaming Press's quantization comparison recommends FP8 as the default for general-purpose serving on native hardware (Hopper, Blackwell, or Apple M5), INT4 weight-only when VRAM is the binding constraint, and BF16 only when accuracy cannot be compromised. DEV Community's overview provides similar defaults: BF16 for quality-constrained workloads, FP8 for general-purpose serving, INT4 when memory or cost binds.
The roadmap's real beneficiaries are memory vendors. High-bandwidth memory (HBM) suppliers like SK Hynix and Samsung capture more value as quantization pushes the bottleneck from compute to memory bandwidth. A 4-bit quantized model moves four times less data per inference call than a 16-bit model, but the memory subsystem must still deliver that data at multi-terabyte-per-second rates to keep tensor cores fed. NVIDIA's Blackwell architecture pairs native FP4 support with HBM3e memory running at 8 terabytes per second, a combination that shifts the cost structure of inference from floating-point units to DRAM dies.
The quantization ladder is not a one-way descent. Digital Applied's 70B-class model survey shows that FP8 delivers 90 percent of FP16 accuracy with 40 percent speedup, while INT4 delivers 85 percent of FP16 accuracy with 170 percent speedup. The gap between FP8 and INT4 is narrowing as calibration methods improve, but the gap between INT4 and FP4 remains wide. Native FP4 support in Blackwell suggests the industry expects another halving, but the accuracy floor is rising. Models quantized below 4 bits show catastrophic degradation on reasoning tasks, a hard limit that no amount of calibration has yet overcome.
Sources & further reading
- NVIDIA TensorRT-LLM quantization documentation and benchmarks Read →
- FireQ INT4-FP8 hybrid kernel research paper Read →
- TensorFoundry LLM quantization field guide Read →
- AIMultiple Qwen3-32B quantization benchmark Read →
- Digital Applied 70B-class model quantization survey Read →
- LLM Pioneer Hub production quantization guide Read →
- Dreaming Press quantization format comparison Read →
- DEV Community quantization overview Read →