How a chain of failures produced the LLM (part 2) — from attention to the scaling revolution and alignment
From context to the finished LLM: the limits of RNNs, the birth of attention, the Transformer, GPT-3 scaling, RLHF alignment, and the era of efficiency.
