Skip to main content
Cloud & AI Hub
Browse
Glossary AI Directory Playgrounds Models Prompts Explainers Strategy Matrix Benchmark Decoder

FlashAttention-3

An optimized self-attention algorithm designed for Hopper GPUs, exploiting asynchronous execution and FP8 precision.

Last reviewed: July 25, 2026

Technical Overview of FlashAttention-3

FlashAttention-3 is an optimized self-attention algorithm designed for high-performance GPUs (specifically targeting NVIDIA Hopper architectures like the H100).

Standard attention calculations are memory-bound, bottlenecked by reading and writing intermediate attention matrices to high-bandwidth memory (HBM). FlashAttention addresses this by tiling inputs and executing softmax normalization incrementally in fast GPU local memory (SRAM). FlashAttention-3 introduces support for FP8 precision and asynchronous memory execution, increasing processing speeds.

Key Architecture & Implementation

FlashAttention-3 introduces architectural improvements to exploit Hopper GPU capabilities:

  1. Asynchronous Execution:
    • Uses Hopper’s Tensor Memory Accelerator (TMA) to overlap data transfers between global memory (HBM) and shared memory (SRAM) with Tensor Core computations.
  2. FP8 Quantization:
    • Supports low-precision FP8 operations, doubling the theoretical throughput of Tensor Cores compared to 16-bit precisions.
  3. Warp-Specialization:
    • Dedicates specific GPU warps to handling data loading while others execute arithmetic operations, minimizing hardware idle cycles.

Core Parameters

  • Speedup: Up to 3x faster than FlashAttention-2 on H100 GPUs.
  • Numerical Precision: Maintains model output quality even when quantizing intermediate attention matrices to FP8.

Real-world Applications

  • Accelerating LLM training runs and high-throughput inference deployments.
  • Integrated into PyTorch and high-performance hosting engines like TensorRT-LLM.

Why It Matters Beyond FlashAttention-2

FlashAttention-3 specifically targets NVIDIA’s Hopper architecture (H100 GPUs), exploiting hardware features that earlier FlashAttention versions couldn’t use: asynchronous execution that overlaps matrix multiplication with softmax computation, and native FP8 support that lets attention run at even lower precision than FP16/BF16 with minimal accuracy loss. Together these changes push attention closer to the theoretical maximum throughput Hopper hardware can deliver, which matters enormously for training and serving the long-context models that have become standard since 2024. Because attention’s compute cost scales quadratically with sequence length, even a modest percentage improvement in attention efficiency compounds into a large real-world difference for 128k+ token contexts. Most production LLM serving frameworks (vLLM, TensorRT-LLM) adopted FlashAttention-3 kernels specifically for this reason once H100 deployment became widespread, since it directly reduces both training cost and inference latency without any change to model architecture or output quality.

This kernel-level optimization work matters because attention has historically been one of the least hardware-efficient parts of running a transformer, spending a disproportionate share of its time on memory movement rather than actual computation — FlashAttention’s whole lineage of improvements (v1 through v3) has been about closing that gap rather than changing what attention mathematically computes.

Advertisement (In-Content)

Historical figures and technical concepts for informational purposes only. Not technical, professional, legal, or financial advice. Sources: Official Documentation.