Self-Attention Math
Attention Softmax Simulator
Interactive mathematical step-through. Play with Query-Key dot products and dimension scaling sizes to observe how softmax prevents vanishing gradients.
Score & Dimension Input
Configure vector dimension size and raw dot product alignment values.
64
Vector length of key/query projections. Higher dimensions result in higher dot products. 12.0
8.0
4.0
Attention Step Trace
Step 1: Raw Similarity Scores (Q · KT)
-
Step 2: Scaling by √dk (Division by -)
-
Step 3: Softmax Probabilities (Attention Map Weights)
-