Transformer Architecture
Transformers replaced recurrence with self-attention and became the default backbone for language, vision, and multimodal models.
Before diving into details, read my note on Attention Mechanisms — it covers the core operation everything else builds on. For how representations are stored and searched, see Embeddings and Retrieval.
Why Transformers Won
The key insight is parallelizable sequence modeling: every token can attend to every other token in one pass, which maps cleanly to modern hardware.
- Self-attention replaces fixed receptive fields with content-based routing
- Positional information is injected (sinusoidal, learned, or RoPE)
- Stacked blocks of attention + feed-forward with residual connections and normalization
Stack at a Glance
graph TD %% Node Definitions A[Input Tokens] --> B[Embedding Layer] C[Positional Encoding] --> D[Combine: Embedding + Position] B --> D %% Stacked Layers Loop subgraph Stack[Stacked Blocks × N Layers] D --> E[Self-Attention] E --> F[Feed-Forward Network - FFN] end %% Output F --> G[Output Head] %% Styling style Stack fill:#f9f9f9,stroke:#333,stroke-width:2px,stroke-dasharray: 5 5
Related
<- All writing