an interactive course in thirteen chapters · 1804 – 2026
from gears
to gradients
Every computer is a stack of switches arranged to do arithmetic, and every modern AI is arithmetic arranged to learn. This course walks that story in order, opens up each component to show what it actually does, follows the branch where computing turned into learning — and ends on the machines you own.
master timeline · tap a dot to open its chapter
how this is built to be learned
Chapter 1 · Before electricity
Arithmetic you could turn with a crank
The first computers had no electricity, no memory chips and no screens. They had gears, cards with holes in them, and one idea that still runs everything: a machine can follow a sequence of instructions it does not understand.
Three separate inventions had to meet before a "computer" was even thinkable. The first was mechanised arithmetic: Pascal's adding machine (1642) and Leibniz's stepped reckoner (1673) proved that carrying a digit could be done by a tooth on a wheel. The second was the stored instruction: in 1804 Joseph-Marie Jacquard controlled a silk loom with a chain of punched cards, so the pattern lived on the cards rather than in the weaver's head. The third was logic as algebra: George Boole showed in 1847 that reasoning with true and false could be written as equations with 1 and 0.
Charles Babbage put the first two together. His Difference Engine (designed 1822) computed tables of polynomials using nothing but repeated addition, and his never-built Analytical Engine (1837) had every part of a modern computer in brass: a "store" (memory), a "mill" (processor), punched-card input and a printer for output. Ada Lovelace, translating a paper about it in 1843, wrote out a step-by-step procedure for computing Bernoulli numbers, and observed that the machine could in principle manipulate any symbols, not just numbers. That note is the reason she is called the first programmer.
- 1804Jacquard loom: punched cards carry the weaving pattern. Instructions become a physical, replaceable object.
- 1822Babbage proposes the Difference Engine; a working fragment is demonstrated in 1832.
- 1837Analytical Engine designed: store, mill, conditional branching and loops, all mechanical.
- 1843Lovelace's Notes describe the first published algorithm intended for a machine.
- 1847Boole's Mathematical Analysis of Logic: AND, OR and NOT as arithmetic on 0 and 1.
- 1890Herman Hollerith's electric tabulator reads punched cards for the US census. His company later becomes IBM.
How the Difference Engine computed without multiplying
Demo · method of differencesBabbage's favourite example was x² + x + 41, which produces primes for x = 0…39. For any polynomial the second difference is constant, so after the first two rows every new value is just two additions. Each column is one stack of number wheels; the crank adds a column into its neighbour.
| x | f(x) = x²+x+41 | Δ¹ (first difference) | Δ² (second difference) |
|---|
Chapter 2 · Logic becomes electric
Bits, gates and the switch
Between 1936 and 1945 three ideas fused: any computation can be reduced to a few simple steps (Turing), those steps are Boolean logic (Shannon), and Boolean logic can be built from electrical switches. After that, building a computer was an engineering problem.
In 1936 Alan Turing described an imaginary machine with a tape, a read/write head and a table of rules, and proved that such a machine could compute anything that can be computed by following rules at all. The result set the ceiling: no cleverer architecture would ever compute more, only faster. A year later Claude Shannon's master's thesis noticed that relay circuits obey Boole's algebra, so any logical expression can be wired up directly. A relay is an electromagnet that closes a contact: one current controls another. That is the whole trick, and every later generation, vacuum tube, transistor, CMOS, is a faster and smaller way of doing exactly that.
Wartime forced the ideas into hardware. Konrad Zuse's Z3 (1941) ran programs from punched film on 2,600 relays. Britain's Colossus (1944) used 1,500+ vacuum tubes to break the Lorenz cipher and was the first electronic digital machine, though it could not be reprogrammed for other jobs. ENIAC (1945) at the University of Pennsylvania used about 17,500 tubes, weighed 30 tons, and computed artillery tables a thousand times faster than a mechanical calculator. But it was "programmed" by rewiring plugboards, which took days.
Why binary?
Not because computers "think in 0 and 1," but because a switch has two reliable states and ten unreliable ones. Decimal machines like ENIAC existed; binary won because a tube or transistor that only needs to be "on" or "off" tolerates noise, ageing and manufacturing variation. Everything else — numbers, text, images, model weights — is an agreed-upon encoding on top of those two states.
One byte, three readings
Demo · encodingGates, then an adder
Demo · Boolean logicA half adder is just XOR (the sum bit) next to AND (the carry bit). Chain four full adders and you can add two 4-bit numbers. This is the circuit at the heart of every ALU ever built.
- 1936Turing, "On Computable Numbers": the universal machine, and the limits of computation.
- 1937Shannon's thesis: relay circuits implement Boolean algebra. Logic design is born.
- 1941Zuse Z3, Berlin: first working programmable, fully automatic digital computer (electromechanical relays).
- 1942Atanasoff–Berry Computer, Iowa: first to use vacuum tubes for binary arithmetic.
- 1944Colossus at Bletchley Park; Harvard Mark I (relays) at IBM/Harvard.
- 1945ENIAC completes: ~17,500 tubes, 5,000 additions per second, programmed by cable.
Chapter 3 · The stored program
The architecture we still use
In 1945 John von Neumann wrote up a design in which the program lives in the same memory as the data. A machine could now modify its own instructions, load a new program in seconds, and — crucially — be built once and used for anything.
The First Draft of a Report on the EDVAC described five parts: an arithmetic unit, a control unit, memory, input and output. The control unit repeats one loop forever: fetch the next instruction from memory, decode what it means, execute it, and move the program counter forward. The Manchester "Baby" ran the first stored program in June 1948 (it took 52 minutes to find the highest factor of 2¹⁸); Cambridge's EDSAC (1949) was the first practical one, and UNIVAC I (1951) was the first sold commercially. Software as a separate discipline — assemblers, then FORTRAN in 1957 — appeared because the program was now data that other programs could manipulate.
The design has a famous cost: instructions and data share one path to memory, so the processor can only ever be as fast as that path lets it be. The "von Neumann bottleneck" is why caches (chapter 6) and GPUs (chapter 10) exist.
Fetch · decode · execute
Demo · the CPU loopA tiny stored-program machine with 12 memory cells. The program adds two numbers stored in memory and writes the result back. Step through it and watch the program counter (PC) and instruction register (IR).
Memory (address · contents)
Control unit & registers
What just happened
- 1945Von Neumann's EDVAC report circulates: the stored-program design.
- 1948Manchester Baby executes the first stored program, 21 June.
- 1949EDSAC, Cambridge: first practical stored-program computer; mercury delay-line memory.
- 1951UNIVAC I: first commercial computer in the US; 46 built.
- 1953Magnetic-core memory (Forrester, MIT Whirlwind) makes RAM reliable for two decades.
- 1957FORTRAN, the first widely used high-level language, compiles maths into machine code.
Chapter 4 · The branch begins
Can a switch learn?
While engineers were building the first computers, a neurophysiologist and a logician asked a stranger question: is a neuron a logic gate? The answer launched a field that would spend seventy years alternating between euphoria and winter.
In 1943 Warren McCulloch and Walter Pitts modelled a neuron as a unit that sums its inputs and fires if the sum crosses a threshold, and showed that networks of these units can compute any logical function. In 1950 Turing proposed judging a machine by conversation rather than mechanism (the "imitation game"), and in 1956 a summer workshop at Dartmouth gave the field its name: artificial intelligence.
Frank Rosenblatt's perceptron (1958) turned the McCulloch–Pitts neuron into something that learned. Each input has a weight; the unit outputs 1 if the weighted sum plus a bias is positive. Training is almost embarrassingly simple: show an example; if the answer is wrong, nudge each weight toward the direction that would have fixed it. Rosenblatt proved this converges whenever a straight line can separate the two classes. The press promised thinking machines within years.
Then in 1969 Marvin Minsky and Seymour Papert showed just how much a single perceptron cannot do: it cannot learn XOR, the simplest function whose classes are not separable by one line. Funding for neural networks collapsed for fifteen years. The fix — stacking layers — was known, but nobody had a practical way to train the hidden layer. That arrives in chapter 7.
Train a perceptron by hand, or let it learn
Demo · learning rule- 1943McCulloch & Pitts: the threshold neuron as a logical unit.
- 1949Donald Hebb: cells that fire together wire together — the first learning rule.
- 1950Turing, "Computing Machinery and Intelligence": the imitation game.
- 1956Dartmouth workshop (McCarthy, Minsky, Shannon, Rochester) names the field.
- 1958Rosenblatt's Mark I Perceptron, built in hardware with motor-driven potentiometers as weights.
- 1966ELIZA (Weizenbaum): pattern-matching "therapist" that people confided in anyway.
- 1969Minsky & Papert, Perceptrons: the XOR limit; neural-network funding dries up.
Chapter 5 · Silicon
The switch shrinks a billion-fold
The transistor did not change what a computer does. It changed how many switches you can afford, how little power each one burns, and how fast they flip. Everything after 1947 is that curve compounding.
On 16 December 1947, at Bell Labs, John Bardeen and Walter Brattain made a sliver of germanium amplify a signal; William Shockley's junction transistor followed in 1948. A transistor does what a relay and a tube do — one signal controls another — but it is solid, has no filament to burn out, and switches in nanoseconds. In 1958 Jack Kilby (Texas Instruments) built several components on one piece of semiconductor; in 1959 Robert Noyce (Fairchild) showed how to connect them with metal deposited on the surface. The integrated circuit meant that wiring, the most failure-prone part of any machine, could be printed.
The same year Mohamed Atalla and Dawon Kahng at Bell Labs demonstrated the MOSFET: a transistor where the controlling "gate" is insulated from the current path by a film of silicon dioxide, so it draws almost no current to hold its state. Pair a p-type and an n-type MOSFET (CMOS, 1963) and a logic gate burns power only while it switches. That is why a phone with billions of transistors does not melt. In 1965 Gordon Moore observed that the number of components per chip was doubling roughly every year (revised to every two years in 1975), and in 1971 Intel's 4004 put an entire processor — 2,300 transistors — on one chip.
One switch, three generations
Demo · relay · tube · MOSFETMoore's law, 1971 → 2023
Chart · transistors per chip- 1947Point-contact transistor, Bell Labs (Bardeen, Brattain; Shockley's junction design 1948).
- 1954First silicon transistor (Texas Instruments); TRADIC, the first transistorised computer, at Bell Labs.
- 1958Kilby's integrated circuit; Noyce's planar version follows in 1959.
- 1959Atalla & Kahng demonstrate the MOSFET.
- 1963CMOS (Wanlass, Fairchild): near-zero static power.
- 1965Moore's article in Electronics.
- 1971Intel 4004: a 4-bit CPU, 2,300 transistors, 740 kHz.
Chapter 6 · Memory and storage
Where the bits live
A processor is only as fast as its next byte. Memory is a hierarchy because no single technology is both fast and cheap, so computers keep a little of the fast kind close and a lot of the slow kind far away.
Early memories were exotic: mercury delay lines that stored bits as sound pulses, cathode-ray tubes that stored them as charge spots on glass. Magnetic-core memory (1953) — tiny ferrite rings threaded on wires, each holding one bit by the direction of its magnetisation — was the workhorse until DRAM arrived. Robert Dennard's 1966 idea was the smallest possible memory cell: one transistor and one capacitor. A charged capacitor is a 1, an empty one a 0. The charge leaks in milliseconds, so the chip rewrites ("refreshes") every cell thousands of times per second — the D is for dynamic. Intel's 1103 (1970) made it commercial, and DRAM is still main memory today.
SRAM uses six transistors per bit as a stable flip-flop: faster, no refresh, but far larger, which is why it is used for the small caches on the CPU die rather than main memory. Storage must survive with the power off: magnetic tape, then hard disks (IBM RAMAC, 1956, 5 MB in a cabinet the size of two refrigerators), and since Fujio Masuoka's flash memory (1980s) charge trapped on a floating gate, which is what an SSD and your phone use.
The hierarchy
Registers inside the CPU respond in a fraction of a nanosecond. Three levels of SRAM cache sit within a few nanoseconds. DRAM is ~100 ns away — long enough for a modern core to have executed several hundred instructions. An SSD is tens of microseconds; a spinning disk milliseconds. Programs work because of locality: what you just used, you will probably use again soon, so caches keep it near.
Latency at human scale
Demo · memory hierarchy- 1949EDSAC's mercury delay lines; Williams–Kilburn CRT tubes on the Manchester machines.
- 1953Magnetic-core memory in MIT's Whirlwind.
- 1956IBM 305 RAMAC: the first hard disk drive, 5 MB on fifty 24-inch platters.
- 1966Dennard's one-transistor DRAM cell (IBM); Intel 1103 ships in 1970.
- 1971The 8-inch floppy disk (IBM).
- 1980sMasuoka's flash memory at Toshiba; first commercial NOR flash 1988, NAND later.
- 1991First commercial flash SSD (SunDisk, later SanDisk), 20 MB.
Chapter 7 · Rules, winters and backpropagation
Two ways to be smart
With neural networks out of favour, AI spent the 1970s and 80s writing down what experts know as explicit rules. It worked well enough to sell, and in one case — a program that recommended antibiotics — to match the specialists. Then it hit a wall, and the networks came back with the trick they had been missing.
MYCIN (Stanford, 1972–76, Edward Shortliffe's doctoral work) held about 600 rules of the form IF the organism is gram-positive AND grows in chains THEN there is suggestive evidence (0.7) that it is a streptococcus. It asked the clinician questions, chained the rules backward from the goal ("what therapy?") and attached a certainty factor to each conclusion. In a 1979 evaluation its antibiotic recommendations for meningitis were judged acceptable about as often as those of infectious-disease faculty — and more often than those of the residents. It was never used on a ward: no legal framework, no integration with records, and a session took half an hour of typing. But it proved that clinical reasoning could be written down, and its rule engine, stripped of the medicine, became the template for hundreds of commercial expert systems.
The trouble was the knowledge acquisition bottleneck. Every rule had to be extracted from a human by interview, rules interacted in ways nobody could predict, and each system knew nothing outside its narrow domain. When the specialised Lisp-machine hardware market collapsed in 1987, the second "AI winter" set in.
Backpropagation
Meanwhile, in 1986, David Rumelhart, Geoffrey Hinton and Ronald Williams popularised an algorithm that trained multi-layer networks. The idea: measure the error at the output, then use the chain rule of calculus to work out how much each weight in each earlier layer contributed to that error, and nudge every weight downhill. Because the units use a smooth activation (a sigmoid instead of a hard threshold) the error surface is differentiable, and gradient descent can walk down it. A hidden layer lets the network bend its boundary, and XOR — the 1969 counterexample — falls in a few hundred steps. The same algorithm, with more layers and more data, trains every model in chapters 10 and 11.
A MYCIN-style rule chain
Demo · backward chaining with certainty factorsA simplified fragment in MYCIN's style (illustrative rules, not clinical advice). Answer the questions the engine asks; watch it chain rules toward the goal and combine certainty factors.
Backpropagation learns XOR
Demo · live trainingA 2 → 4 → 1 network with sigmoid units, trained here in your browser by gradient descent on the four XOR examples. The background shows the network's output for every point; the four dots are the training data.
- 1972Shortliffe begins MYCIN at Stanford; DENDRAL (chemistry) had pioneered the approach from 1965.
- 1973The Lighthill Report ends most UK AI funding: the first AI winter.
- 1980XCON at Digital Equipment configures VAX orders; saves an estimated $25M a year.
- 1982Japan's Fifth Generation project; Hopfield networks revive interest in neural computation.
- 1986Rumelhart, Hinton & Williams, "Learning representations by back-propagating errors."
- 1987Lisp-machine market collapses; second AI winter begins.
- 1989Yann LeCun trains a convolutional network to read handwritten postcodes for the US Postal Service.
Chapter 8 · The personal computer
A computer on every desk
The microprocessor made a whole CPU cost a few dollars. Hobbyists put one in a box, a spreadsheet gave businesses a reason to buy the box, and a graphical interface let everyone else use it. By 1995 the components of a PC were standardised well enough that you could assemble one from parts — and people did.
The Altair 8800 (1975) was a kit with an Intel 8080 and toggle switches for input; Bill Gates and Paul Allen wrote a BASIC interpreter for it and founded Microsoft. The Apple II (1977) shipped assembled with colour graphics, and VisiCalc (1979) — the first spreadsheet — turned it into a business tool. IBM's PC (1981) used off-the-shelf parts and an open bus, which let other companies build compatible machines and made Intel and Microsoft the platform. The Macintosh (1984) brought the graphical interface developed at Xerox PARC — windows, icons, mouse — to the mass market, and Windows 95 made it universal.
Under the hood, the PC settled into a pattern that still holds: a CPU, RAM on a fast bus, a "chipset" that fans out slower connections, expansion slots (ISA, then PCI, then PCIe) for graphics and other cards, and storage on its own interface. The diagram below is the modern version; hover a part.
Anatomy of a modern PC
Demo · tap a componentThe PC as a system
Fast things attach directly to the CPU; slower things go through the chipset. Widths of the connecting lines are drawn roughly in proportion to bandwidth.
- 1971Intel 4004; Unix and the C language mature at Bell Labs over the next two years.
- 1975MITS Altair 8800 on the cover of Popular Electronics; Microsoft founded.
- 1977Apple II, Commodore PET and Tandy TRS-80: the "1977 trinity."
- 1979VisiCalc. Businesses buy computers to run one program.
- 1981IBM PC with MS-DOS; the 8088 CPU and open architecture spawn the "clone" industry.
- 1984Macintosh: the GUI goes mainstream.
- 1991Linus Torvalds posts the first Linux kernel; the free software stack for servers is born.
- 1993Intel Pentium (3.1M transistors); 1995 Windows 95 and PCI slots standardise the desktop.
Chapter 9 · The networked world
Computers that talk
A network is a computer's way of borrowing another computer. The idea that made it scale was to chop every message into small packets and let each one find its own way.
Telephone networks reserved a circuit for each call. Paul Baran (RAND) and Donald Davies (NPL) independently proposed in the 1960s that data should instead be split into packets, each carrying its destination address, and forwarded hop by hop by routers that need to know only the next hop. No central switch, no reserved line, and a broken link just means packets take another route. ARPANET sent its first packet in October 1969. Vint Cerf and Bob Kahn's TCP/IP (1974; adopted network-wide on 1 January 1983) layered the job: IP delivers packets unreliably, TCP on top numbers them, retransmits losses and reassembles the message. Ethernet (1973) handled the local wire, and Tim Berners-Lee's World Wide Web (1989–91) put a document format, an address scheme and a protocol on top so that anyone could publish. Search (Google, 1998) made it navigable; the iPhone (2007) put it in every pocket.
For AI the network mattered twice over: it created the corpus — the web is what large language models are trained on — and it made it possible to rent a warehouse of GPUs by the hour.
Packet switching
Demo · send a message- 1969ARPANET's first message, UCLA to Stanford ("LO" — the system crashed before "GIN").
- 1973Ethernet (Metcalfe, Xerox PARC).
- 1974Cerf & Kahn publish the TCP design; 1983 ARPANET switches to TCP/IP.
- 1989Berners-Lee proposes the Web at CERN; first site online 1991.
- 1993Mosaic browser; the Web becomes graphical and popular.
- 1998Google founded on PageRank; 2006 Amazon launches EC2, renting computers by the hour.
- 2007iPhone: a networked computer in every pocket.
Chapter 10 · Parallel hardware meets deep networks
The graphics card that learned to see
Backpropagation had worked since 1986. What it lacked was data and arithmetic. The web supplied the data; a chip designed to draw video games supplied the arithmetic. In 2012 the two met, and a decade of "AI winter" ended in one afternoon.
A CPU is built to run one thread of unpredictable instructions as fast as possible: big caches, branch prediction, out-of-order execution. A GPU is built to do the same simple operation on millions of pixels at once: thousands of small arithmetic units, each doing the same thing to different data (SIMD). Nvidia's CUDA (2007) let programmers use that for anything, and it turned out that a neural network — a huge stack of matrix multiplications — is exactly a "same thing to different data" problem.
Fei-Fei Li's ImageNet (2009) gathered 14 million labelled images and ran an annual competition. In 2012 Alex Krizhevsky, Ilya Sutskever and Geoffrey Hinton entered AlexNet, an eight-layer convolutional network trained on two consumer GPUs for a week. Its error rate was 15.3%; the runner-up's was 26.2%. Within two years every entrant was a deep network. Convolutional networks work because they learn small filters — edge detectors, texture detectors — and slide them across the image, sharing the same weights everywhere; deeper layers combine those into parts and objects. ResNet (2015) showed that with "skip connections" you could train networks 150 layers deep, and in 2016 DeepMind's AlphaGo, combining deep networks with tree search, beat Lee Sedol at Go, a game thought to be a decade away.
Serial versus parallel
Demo · 64 identical operationsCPU · 4 wide cores
GPU · 64 narrow lanes
What a convolution does
Demo · slide a 3×3 filter- 1997IBM's Deep Blue beats Kasparov (search, not learning); Hochreiter & Schmidhuber publish the LSTM.
- 1998LeCun's LeNet-5 reads cheques in production; convolutional nets are practical but niche.
- 2006Hinton's "deep belief nets" and the word deep learning.
- 2007Nvidia releases CUDA: general-purpose programming on GPUs.
- 2009ImageNet published; Raina, Madhavan & Ng train networks on GPUs.
- 2012AlexNet wins ImageNet by a landslide.
- 2013word2vec: words as vectors, where king − man + woman ≈ queen.
- 2014Generative adversarial networks; sequence-to-sequence translation with attention (Bahdanau).
- 2015ResNet; 2016 AlphaGo defeats Lee Sedol 4–1.
Chapter 11 · Attention, scale and language
Predict the next token
A large language model is a very large neural network trained on one deceptively simple task: given some text, guess the next piece. The architecture that made this scale is the transformer, and its central idea is that every word should be able to look at every other word.
Recurrent networks read text one token at a time, carrying a summary forward, which made them slow to train and forgetful over long spans. The 2017 paper Attention Is All You Need (Vaswani et al., Google) dropped recurrence entirely. Each token is turned into a vector (an embedding). In an attention layer, every token computes three vectors from its own — a query ("what am I looking for?"), a key ("what do I contain?") and a value ("what do I pass on?"). The score between two tokens is the dot product of one's query with the other's key; a softmax turns the scores into weights summing to 1; the token's new representation is the weighted sum of the values. Because all tokens do this simultaneously it is one big matrix multiplication — a GPU's favourite meal — and because the weights are data-dependent the model can learn that in "the pharmacist checked the order because it looked wrong," it should attend to order.
Stack dozens of these layers with small feed-forward networks between them, train on trillions of tokens to predict the next one, and something unexpected happens: the model must learn grammar, facts, and a fair amount of reasoning simply because they help it predict. GPT-2 (2019, 1.5B parameters) wrote passable paragraphs; GPT-3 (2020, 175B) could follow instructions given as examples in the prompt. RLHF — fine-tuning on human preferences between candidate answers — turned a next-token predictor into an assistant, and ChatGPT (November 2022) put that in front of a hundred million people in two months. Open-weight models (Meta's Llama, 2023; Mistral; and others) plus quantisation tools like llama.cpp made it possible to run a capable model on a laptop or a home server.
What the model is, physically
A file of numbers. A 7-billion-parameter model is 7 billion learned weights; at 16 bits each that is 14 GB, which must sit in fast memory during inference because every token generated touches every weight. Quantising to 4 bits cuts that to ~4 GB with a modest quality loss, which is what makes local inference practical. Generation is memory-bandwidth-bound: tokens per second ≈ bandwidth ÷ model size, which is why a GPU's fast VRAM matters more than its clock.
Attention weights, step by step
Demo · softmax over scoresPick which token is asking (the query). The bars show an illustrative attention pattern for that token — the shape a trained head produces for pronoun resolution. Below, drag raw scores to see how softmax turns them into weights.
A language model you can watch think
Demo · n-gram next-token predictionA transformer predicts the next token from context using learned weights. This tiny model does the same job by counting: it learned which word follows which pair of words from the paragraph below, and samples the next word from those counts. Same task, no neural network — which is exactly what makes the difference visible.
Training text (about 190 words)
the pharmacist checked the order before the dose was sent to the floor . the dose was adjusted for renal function because the patient had a low clearance . the patient was started on an antibiotic after the culture came back . the culture grew a gram positive organism so the team narrowed therapy . the team asked the pharmacist to verify the interaction between the two drugs . the two drugs share a metabolic pathway so the level was monitored . the level was drawn before the next dose and the dose was held . the nurse called the pharmacist about the order because the rate looked wrong . the rate was corrected and the infusion was restarted on the floor . the model predicts the next word from the words before it . the words before it are the context and the context is all the model can see . the model was trained on text from the web and the text was tokenized . the weights of the model were adjusted by gradient descent until the loss was low . the loss was low so the predictions were good and the team was pleased with the model .
Will it fit? Memory for local inference
Calculator · weights in memory- 2017"Attention Is All You Need": the transformer.
- 2018BERT (Google) and GPT-1 (OpenAI): pre-train on raw text, fine-tune for tasks.
- 2019GPT-2, 1.5B parameters; 2020 GPT-3, 175B, and the scaling-law papers.
- 2021AlphaFold 2 predicts protein structures at near-experimental accuracy; Codex writes code.
- 2022InstructGPT/RLHF; Stable Diffusion; ChatGPT launches 30 November.
- 2023GPT-4, Claude, Llama (open weights), llama.cpp brings inference to consumer hardware.
- 2024–26Multimodal models, long contexts, agents that use tools, and small models that run on a phone.
Chapter 12 · Where the two lines meet
From transistor to token
Type a question into a chat model and every chapter of this page runs in order, in about a second. Read the stack from the bottom up.
What has not changed
Turing's ceiling still holds: a language model computes nothing a 1948 machine could not, given enough tape and time. What changed is the price of a switch (from a dollar to a trillionth of a cent), the price of a byte, and the discovery that if you make a function large enough and differentiable, you can fit it to the world instead of programming it. The expert systems of chapter 7 tried to write medicine down as rules; the models of chapter 11 read the literature and inferred the rules, and neither approach removes the need for someone who can tell when the answer is wrong.
Where the open questions are
- Hardware: transistors are near atomic limits; gains now come from 3D stacking, chiplets, and purpose-built matrix engines, with memory bandwidth the binding constraint.
- Efficiency: training frontier models costs gigawatt-hours; small distilled models and better quantisation are moving capability toward local devices.
- Trust: a next-token predictor has no built-in notion of truth. Retrieval, tool use, and verification against sources are the current answers — the same discipline as checking an order.
Chapter 13 · The stack on your own hardware
What happens when you ask pharm-gemma a question
Every chapter above runs, in order, on machines you own. Here is the path a question takes from your phone to the OptiPlex and back, and — more usefully — the arithmetic that predicts how fast it will be, so you can decide what to run locally and what to send to a frontier model before you download anything. (Everything that enters any of these systems is de-identified; the local box is about cost, availability and control, not a PHI boundary.)
Two phases, two bottlenecks
Answering has two distinct phases, and they hit different walls. Prefill reads your prompt: all the prompt's tokens go through the network at once as one big matrix multiplication, so it is compute-bound — more cores, wider SIMD, or a GPU make it faster. Decode generates the reply one token at a time; each token needs the whole weight file streamed from memory, so it is bandwidth-bound — only faster memory helps. This is why a GPU with 288 GB/s of VRAM beats the OptiPlex's CPU by ~10× on generation even though the chip does far more arithmetic than that ratio suggests, and why Apple's unified memory (120 GB/s on an M4 Air) makes a laptop a respectable inference box.
The rule of thumb, then: tokens/s ≈ usable bandwidth ÷ model size in bytes. Quantisation (chapter 11) helps twice — the model fits, and there are fewer bytes to stream. The estimator below applies that arithmetic to your machines.
Will it fit, and how fast? Your devices
Estimator · bandwidth ÷ bytes| device | memory for model | needs | fits? | generation | read a 2k-token prompt | 300-token reply |
|---|
What the numbers say about your setup
The OptiPlex's 32 GB is generous for fitting models — a 27B model at Q4 loads with room to spare — but its DDR4 bandwidth means anything above ~9B generates slower than you can read. The sweet spot for an always-on box like this is a 4–9B model at Q4–Q6: that is exactly where Gemma E4B landed when it passed your vancomycin AUC test, and the physics agrees with the clinical result. The MacBook Air, with 3–4× the bandwidth, is the machine for trying larger models interactively; the phone and iPad have the bandwidth for a small model but only ~5 GB to hold one, which is why they work best as the thin head. The two GPU rows show what a card would buy: roughly 10× on the OptiPlex, 30× with a high-end card. But look at what that 10× actually is — an 8B model answering faster — against a frontier model that answers better, on demand, for less per month than a card costs. That is the arithmetic behind staying with what you have: run the small always-on model locally for the cheap, private, repetitive work, send the hard reasoning to the cloud, and let any future PC be a gaming machine that happens to run LLMs well as a side effect.
-t 4 to the 4 physical cores of the 6700T, not the 8 logical ones. Hyper-threads share the same AVX2 units and the same memory bus; extra threads just fight over them.-c to what the task needs, not the maximum the model allows.llama-server with --mlock (so the OS never pages the weights out) holds it in RAM; with 32 GB you can afford to leave pharm-gemma loaded permanently.llama-server); the MacBook for interactive trials of bigger models; a frontier model for anything where answer quality matters more than where it ran.- ch 9Tailscale is packet switching plus WireGuard encryption plus a coordination server that tells your nodes each other's addresses; MagicDNS is the Web's naming idea applied to your own machines.
- ch 3llama.cpp is a stored program: the CPU fetches, decodes and executes its instructions; the GGUF model file is data it reads. Same memory, two roles — the von Neumann idea.
- ch 6Model loading is the memory hierarchy in action: SSD → DRAM once, then DRAM → cache → registers billions of times per second during decode.
- ch 10AVX2 is SIMD on the CPU: 8 float32 multiplies per instruction per core. A GPU has thousands of such lanes; that difference is the whole gap in the table.
- ch 11Temperature 0.3 in the pharm-gemma Modelfile is the softmax knob from the attention demo — sharpening the distribution so clinical answers stay consistent run to run.
review · the whole course on one screen
thirteen sentences
Each chapter reduced to the one line worth keeping, with your check result beside it. Reread the ones that did not land, then run the practice below — the same questions, shuffled, graded by you.
key terms · all chapters
practice run · 39 retrieval questions, shuffled
Answer each in your head, reveal, and grade yourself honestly. Missed questions are kept so you can run just those next time — the spacing is up to you: come back tomorrow, then next week.
your notes
About the demos. All demos run entirely in this page. The perceptron and XOR trainers use real learning rules; the n-gram generator is a real (tiny) language model; the attention bars are an illustrative pattern; the MYCIN fragment is a teaching simplification and not clinical advice. Dates follow the standard histories (Ceruzzi, A History of Modern Computing; Russell & Norvig, AI: A Modern Approach; Computer History Museum timelines).
Scores and notes live only in this browser.