Hacker News
Transformers Explained Visually
Transformers are neural network architectures introduced in 2017’s “Attention is All You Need” that use self-attention to process entire sequences and predict the next token. GPT-2 (small) has 124 million parameters, a 50,257-token vocabulary, and 768-dimensional embeddings stored in a 50,257 × 768 matrix. The model’s architecture consists of embedding, attention-based transformer blocks with MLP layers, and final linear-softmax output layers that generate token probabilities.