Attention Is All You Need — How Transformers Actually Work
The 2017 paper behind every model you use now. A plain-English walk through each piece, from tokenization and embeddings up to multi-head attention, and why it beat the recurrent models it replaced.