Portfolio/ PromitDocument
Promit / Reference LibraryThird Epoch · 16,400 words
Document preface

The Transformer Bible

The complete transformer architecture explained from the ground up. It runs from tokenization and embeddings through attention, the transformer block, and output generation, all the way to Mixture of Experts — with every mechanism, formula, and number worked by hand.

Written & compiled byPromit
PublishedJune 30, 2026

Transformers become mystical when explanations jump from tokens to intelligence and hide the arithmetic in between. This book refuses that jump. It follows information through the machine one shape, rotation, dot product, residual path, and probability at a time.

The journey begins before attention—with vocabularies, tokenization, embeddings, and position—and continues through RoPE, masking, normalization, feed-forward networks, decoding, and Mixture of Experts. The mechanisms are not merely named. The numbers are worked by hand so that the architecture can be seen rather than memorised.

This is for the reader who is no longer satisfied saying that attention lets tokens look at other tokens. By the end, the familiar block diagram should feel less like an icon and more like a machine you could dismantle and rebuild.

A look inside

What this edition gives you

  1. 01

    A token's complete journey from text to next-token probability

  2. 02

    Attention, RoPE, normalization, and residuals worked numerically

  3. 03

    Modern scaling ideas through grouped queries and Mixture of Experts

Continue reading

Open it with a pencil nearby. This is not a glossary of transformer terms; it is an invitation to take the machine apart.