gears → gradientsoverview

an interactive course in thirteen chapters · 1804 – 2026

from gears
to gradients

Every computer is a stack of switches arranged to do arithmetic, and every modern AI is arithmetic arranged to learn. This course walks that story in order, opens up each component to show what it actually does, follows the branch where computing turned into learning — and ends on the machines you own.

master timeline · tap a dot to open its chapter

hover or tap a milestone
hardwareartificial intelligence

how this is built to be learned

one chapter at a timeEach chapter is a single screen: orient, story, open it up, check yourself, recap. About four minutes of reading plus whatever you spend on the demos.
objectives firstEvery chapter opens with what you'll be able to do afterwards, and which earlier chapters it builds on. If a prerequisite is unread, go back — the ideas stack.
recall before you continueChapters begin by asking you to recall the previous one from memory before showing the answer. Retrieval, not rereading, is what makes it stick.
demos are the worked examplesSeventeen working mechanisms. Change the inputs until you can predict the output; that is the point at which you understand it.
recap, terms, questionsEach chapter ends in one sentence, its key terms, and three retrieval questions. The review screen shuffles all thirty-nine into a self-graded practice run and remembers what you missed.
your four bucketsWhat clicked · what confused me · connections · questions for next session. A notes box per chapter, saved in this browser, exportable as one markdown file from the review page.
1804– 1890

Chapter 1 · Before electricity

Arithmetic you could turn with a crank

The first computers had no electricity, no memory chips and no screens. They had gears, cards with holes in them, and one idea that still runs everything: a machine can follow a sequence of instructions it does not understand.

Three separate inventions had to meet before a "computer" was even thinkable. The first was mechanised arithmetic: Pascal's adding machine (1642) and Leibniz's stepped reckoner (1673) proved that carrying a digit could be done by a tooth on a wheel. The second was the stored instruction: in 1804 Joseph-Marie Jacquard controlled a silk loom with a chain of punched cards, so the pattern lived on the cards rather than in the weaver's head. The third was logic as algebra: George Boole showed in 1847 that reasoning with true and false could be written as equations with 1 and 0.

Charles Babbage put the first two together. His Difference Engine (designed 1822) computed tables of polynomials using nothing but repeated addition, and his never-built Analytical Engine (1837) had every part of a modern computer in brass: a "store" (memory), a "mill" (processor), punched-card input and a printer for output. Ada Lovelace, translating a paper about it in 1843, wrote out a step-by-step procedure for computing Bernoulli numbers, and observed that the machine could in principle manipulate any symbols, not just numbers. That note is the reason she is called the first programmer.

  • 1804Jacquard loom: punched cards carry the weaving pattern. Instructions become a physical, replaceable object.
  • 1822Babbage proposes the Difference Engine; a working fragment is demonstrated in 1832.
  • 1837Analytical Engine designed: store, mill, conditional branching and loops, all mechanical.
  • 1843Lovelace's Notes describe the first published algorithm intended for a machine.
  • 1847Boole's Mathematical Analysis of Logic: AND, OR and NOT as arithmetic on 0 and 1.
  • 1890Herman Hollerith's electric tabulator reads punched cards for the US census. His company later becomes IBM.

How the Difference Engine computed without multiplying

Demo · method of differences

Babbage's favourite example was x² + x + 41, which produces primes for x = 0…39. For any polynomial the second difference is constant, so after the first two rows every new value is just two additions. Each column is one stack of number wheels; the crank adds a column into its neighbour.

xf(x) = x²+x+41Δ¹ (first difference)Δ² (second difference)
Each turn does two additions and no multiplications: Δ¹ grows by the constant Δ² = 2, then f(x) grows by the new Δ¹. The engine only ever needed to add, which is why it could be built from carry wheels.
1936– 1945

Chapter 2 · Logic becomes electric

Bits, gates and the switch

Between 1936 and 1945 three ideas fused: any computation can be reduced to a few simple steps (Turing), those steps are Boolean logic (Shannon), and Boolean logic can be built from electrical switches. After that, building a computer was an engineering problem.

In 1936 Alan Turing described an imaginary machine with a tape, a read/write head and a table of rules, and proved that such a machine could compute anything that can be computed by following rules at all. The result set the ceiling: no cleverer architecture would ever compute more, only faster. A year later Claude Shannon's master's thesis noticed that relay circuits obey Boole's algebra, so any logical expression can be wired up directly. A relay is an electromagnet that closes a contact: one current controls another. That is the whole trick, and every later generation, vacuum tube, transistor, CMOS, is a faster and smaller way of doing exactly that.

Wartime forced the ideas into hardware. Konrad Zuse's Z3 (1941) ran programs from punched film on 2,600 relays. Britain's Colossus (1944) used 1,500+ vacuum tubes to break the Lorenz cipher and was the first electronic digital machine, though it could not be reprogrammed for other jobs. ENIAC (1945) at the University of Pennsylvania used about 17,500 tubes, weighed 30 tons, and computed artillery tables a thousand times faster than a mechanical calculator. But it was "programmed" by rewiring plugboards, which took days.

Why binary?

Not because computers "think in 0 and 1," but because a switch has two reliable states and ten unreliable ones. Decimal machines like ENIAC existed; binary won because a tube or transistor that only needs to be "on" or "off" tolerates noise, ageing and manufacturing variation. Everything else — numbers, text, images, model weights — is an agreed-upon encoding on top of those two states.

One byte, three readings

Demo · encoding
Eight switches give 256 patterns. The same pattern is an unsigned number, a signed number (two's complement, where the top bit is worth −128), and a text character in ASCII. The hardware never knows which reading you meant.

Gates, then an adder

Demo · Boolean logic
Toggle the inputs; every gate updates.

A half adder is just XOR (the sum bit) next to AND (the carry bit). Chain four full adders and you can add two 4-bit numbers. This is the circuit at the heart of every ALU ever built.

A (0–15)
B (0–15)
Sum (carry + 4 bits)
Each stage takes two input bits plus the carry from the stage to its right and produces a sum bit and a carry out. Modern CPUs use faster carry-lookahead designs, but the logic is the same.
  • 1936Turing, "On Computable Numbers": the universal machine, and the limits of computation.
  • 1937Shannon's thesis: relay circuits implement Boolean algebra. Logic design is born.
  • 1941Zuse Z3, Berlin: first working programmable, fully automatic digital computer (electromechanical relays).
  • 1942Atanasoff–Berry Computer, Iowa: first to use vacuum tubes for binary arithmetic.
  • 1944Colossus at Bletchley Park; Harvard Mark I (relays) at IBM/Harvard.
  • 1945ENIAC completes: ~17,500 tubes, 5,000 additions per second, programmed by cable.
1945– 1955

Chapter 3 · The stored program

The architecture we still use

In 1945 John von Neumann wrote up a design in which the program lives in the same memory as the data. A machine could now modify its own instructions, load a new program in seconds, and — crucially — be built once and used for anything.

The First Draft of a Report on the EDVAC described five parts: an arithmetic unit, a control unit, memory, input and output. The control unit repeats one loop forever: fetch the next instruction from memory, decode what it means, execute it, and move the program counter forward. The Manchester "Baby" ran the first stored program in June 1948 (it took 52 minutes to find the highest factor of 2¹⁸); Cambridge's EDSAC (1949) was the first practical one, and UNIVAC I (1951) was the first sold commercially. Software as a separate discipline — assemblers, then FORTRAN in 1957 — appeared because the program was now data that other programs could manipulate.

The design has a famous cost: instructions and data share one path to memory, so the processor can only ever be as fast as that path lets it be. The "von Neumann bottleneck" is why caches (chapter 6) and GPUs (chapter 10) exist.

Fetch · decode · execute

Demo · the CPU loop

A tiny stored-program machine with 12 memory cells. The program adds two numbers stored in memory and writes the result back. Step through it and watch the program counter (PC) and instruction register (IR).

Memory (address · contents)
Control unit & registers
FetchDecodeExecute
What just happened
Press Step to begin. Nothing has run yet.
Instruction and data sit in the same memory: addresses 0–4 hold the program, 8–10 hold numbers. The only thing that makes cell 0 an "instruction" is that the PC pointed at it. A real CPU does this loop billions of times per second, several instructions in flight at once.
CPU Control unit ALU + registers Memory instructions and data, one address space address data (the bottleneck) Input Output
The von Neumann layout. The single data path between CPU and memory is why nearly every speed-up since 1950 has been a way of touching memory less often.
  • 1945Von Neumann's EDVAC report circulates: the stored-program design.
  • 1948Manchester Baby executes the first stored program, 21 June.
  • 1949EDSAC, Cambridge: first practical stored-program computer; mercury delay-line memory.
  • 1951UNIVAC I: first commercial computer in the US; 46 built.
  • 1953Magnetic-core memory (Forrester, MIT Whirlwind) makes RAM reliable for two decades.
  • 1957FORTRAN, the first widely used high-level language, compiles maths into machine code.
1943– 1969

Chapter 4 · The branch begins

Can a switch learn?

While engineers were building the first computers, a neurophysiologist and a logician asked a stranger question: is a neuron a logic gate? The answer launched a field that would spend seventy years alternating between euphoria and winter.

In 1943 Warren McCulloch and Walter Pitts modelled a neuron as a unit that sums its inputs and fires if the sum crosses a threshold, and showed that networks of these units can compute any logical function. In 1950 Turing proposed judging a machine by conversation rather than mechanism (the "imitation game"), and in 1956 a summer workshop at Dartmouth gave the field its name: artificial intelligence.

Frank Rosenblatt's perceptron (1958) turned the McCulloch–Pitts neuron into something that learned. Each input has a weight; the unit outputs 1 if the weighted sum plus a bias is positive. Training is almost embarrassingly simple: show an example; if the answer is wrong, nudge each weight toward the direction that would have fixed it. Rosenblatt proved this converges whenever a straight line can separate the two classes. The press promised thinking machines within years.

Then in 1969 Marvin Minsky and Seymour Papert showed just how much a single perceptron cannot do: it cannot learn XOR, the simplest function whose classes are not separable by one line. Funding for neural networks collapsed for fifteen years. The fix — stacking layers — was known, but nobody had a practical way to train the hidden layer. That arrives in chapter 7.

Train a perceptron by hand, or let it learn

Demo · learning rule
Output = 1 if w₁x₁ + w₂x₂ + b > 0. The learning rule: for each misclassified point, w ← w + η·(target − output)·x. On separable data the line converges; on XOR data it thrashes forever, exactly as Minsky and Papert proved.
  • 1943McCulloch & Pitts: the threshold neuron as a logical unit.
  • 1949Donald Hebb: cells that fire together wire together — the first learning rule.
  • 1950Turing, "Computing Machinery and Intelligence": the imitation game.
  • 1956Dartmouth workshop (McCarthy, Minsky, Shannon, Rochester) names the field.
  • 1958Rosenblatt's Mark I Perceptron, built in hardware with motor-driven potentiometers as weights.
  • 1966ELIZA (Weizenbaum): pattern-matching "therapist" that people confided in anyway.
  • 1969Minsky & Papert, Perceptrons: the XOR limit; neural-network funding dries up.
1947– 1971

Chapter 5 · Silicon

The switch shrinks a billion-fold

The transistor did not change what a computer does. It changed how many switches you can afford, how little power each one burns, and how fast they flip. Everything after 1947 is that curve compounding.

On 16 December 1947, at Bell Labs, John Bardeen and Walter Brattain made a sliver of germanium amplify a signal; William Shockley's junction transistor followed in 1948. A transistor does what a relay and a tube do — one signal controls another — but it is solid, has no filament to burn out, and switches in nanoseconds. In 1958 Jack Kilby (Texas Instruments) built several components on one piece of semiconductor; in 1959 Robert Noyce (Fairchild) showed how to connect them with metal deposited on the surface. The integrated circuit meant that wiring, the most failure-prone part of any machine, could be printed.

The same year Mohamed Atalla and Dawon Kahng at Bell Labs demonstrated the MOSFET: a transistor where the controlling "gate" is insulated from the current path by a film of silicon dioxide, so it draws almost no current to hold its state. Pair a p-type and an n-type MOSFET (CMOS, 1963) and a logic gate burns power only while it switches. That is why a phone with billions of transistors does not melt. In 1965 Gordon Moore observed that the number of components per chip was doubling roughly every year (revised to every two years in 1975), and in 1971 Intel's 4004 put an entire processor — 2,300 transistors — on one chip.

One switch, three generations

Demo · relay · tube · MOSFET
Flip it. All three devices do the same job; only the physics differs.
Relay (1930s) coil control lamp magnet pulls a contact · ~10 ms Triode tube (1940s) cathode (heated) grid plate grid voltage gates electrons · ~1 µs MOSFET (1959 →) p-type silicon body n n gate source drain control field forms a channel · ~10 ps oxide insulates the gate: ~zero current to hold
Relay: current in a coil pulls a contact closed. Tube: a positive grid lets electrons cross the vacuum from cathode to plate. MOSFET: a positive gate attracts electrons under the oxide, forming a conductive channel between source and drain. Same truth table, a billion times faster and smaller.

Moore's law, 1971 → 2023

Chart · transistors per chip
Hover a point for the chip.
Logarithmic vertical axis: each gridline is 10×. A straight line here means exponential growth — about 2× every two years for five decades. Recent gains come increasingly from packaging several dies together rather than shrinking one.
  • 1947Point-contact transistor, Bell Labs (Bardeen, Brattain; Shockley's junction design 1948).
  • 1954First silicon transistor (Texas Instruments); TRADIC, the first transistorised computer, at Bell Labs.
  • 1958Kilby's integrated circuit; Noyce's planar version follows in 1959.
  • 1959Atalla & Kahng demonstrate the MOSFET.
  • 1963CMOS (Wanlass, Fairchild): near-zero static power.
  • 1965Moore's article in Electronics.
  • 1971Intel 4004: a 4-bit CPU, 2,300 transistors, 740 kHz.
1949– 1991

Chapter 6 · Memory and storage

Where the bits live

A processor is only as fast as its next byte. Memory is a hierarchy because no single technology is both fast and cheap, so computers keep a little of the fast kind close and a lot of the slow kind far away.

Early memories were exotic: mercury delay lines that stored bits as sound pulses, cathode-ray tubes that stored them as charge spots on glass. Magnetic-core memory (1953) — tiny ferrite rings threaded on wires, each holding one bit by the direction of its magnetisation — was the workhorse until DRAM arrived. Robert Dennard's 1966 idea was the smallest possible memory cell: one transistor and one capacitor. A charged capacitor is a 1, an empty one a 0. The charge leaks in milliseconds, so the chip rewrites ("refreshes") every cell thousands of times per second — the D is for dynamic. Intel's 1103 (1970) made it commercial, and DRAM is still main memory today.

SRAM uses six transistors per bit as a stable flip-flop: faster, no refresh, but far larger, which is why it is used for the small caches on the CPU die rather than main memory. Storage must survive with the power off: magnetic tape, then hard disks (IBM RAMAC, 1956, 5 MB in a cabinet the size of two refrigerators), and since Fujio Masuoka's flash memory (1980s) charge trapped on a floating gate, which is what an SSD and your phone use.

The hierarchy

Registers inside the CPU respond in a fraction of a nanosecond. Three levels of SRAM cache sit within a few nanoseconds. DRAM is ~100 ns away — long enough for a modern core to have executed several hundred instructions. An SSD is tens of microseconds; a spinning disk milliseconds. Programs work because of locality: what you just used, you will probably use again soon, so caches keep it near.

Latency at human scale

Demo · memory hierarchy
…then the others would take:
Typical figures for a 2020s desktop; the bars use a logarithmic scale, so each gridline step is 10×. A cache miss to DRAM is like a ten-minute walk when you expected a glance; a disk seek is a year. This gap is why data layout and batch size matter so much when running models on your own hardware.
One DRAM bit = 1 transistor + 1 capacitor bit line word line (row select) gate ground capacitor charged = 1 · leaks in ms · refreshed constantly read: raise word line, transistor opens, charge spills onto the bit line, a sense amplifier detects the tiny swing, then the cell is rewritten.
Reading a DRAM cell destroys its contents (the charge drains onto the bit line), so every read is followed by a rewrite. Billions of these cells, arranged in rows and columns, make up a memory module.
  • 1949EDSAC's mercury delay lines; Williams–Kilburn CRT tubes on the Manchester machines.
  • 1953Magnetic-core memory in MIT's Whirlwind.
  • 1956IBM 305 RAMAC: the first hard disk drive, 5 MB on fifty 24-inch platters.
  • 1966Dennard's one-transistor DRAM cell (IBM); Intel 1103 ships in 1970.
  • 1971The 8-inch floppy disk (IBM).
  • 1980sMasuoka's flash memory at Toshiba; first commercial NOR flash 1988, NAND later.
  • 1991First commercial flash SSD (SunDisk, later SanDisk), 20 MB.
1970– 1989

Chapter 7 · Rules, winters and backpropagation

Two ways to be smart

With neural networks out of favour, AI spent the 1970s and 80s writing down what experts know as explicit rules. It worked well enough to sell, and in one case — a program that recommended antibiotics — to match the specialists. Then it hit a wall, and the networks came back with the trick they had been missing.

MYCIN (Stanford, 1972–76, Edward Shortliffe's doctoral work) held about 600 rules of the form IF the organism is gram-positive AND grows in chains THEN there is suggestive evidence (0.7) that it is a streptococcus. It asked the clinician questions, chained the rules backward from the goal ("what therapy?") and attached a certainty factor to each conclusion. In a 1979 evaluation its antibiotic recommendations for meningitis were judged acceptable about as often as those of infectious-disease faculty — and more often than those of the residents. It was never used on a ward: no legal framework, no integration with records, and a session took half an hour of typing. But it proved that clinical reasoning could be written down, and its rule engine, stripped of the medicine, became the template for hundreds of commercial expert systems.

The trouble was the knowledge acquisition bottleneck. Every rule had to be extracted from a human by interview, rules interacted in ways nobody could predict, and each system knew nothing outside its narrow domain. When the specialised Lisp-machine hardware market collapsed in 1987, the second "AI winter" set in.

Backpropagation

Meanwhile, in 1986, David Rumelhart, Geoffrey Hinton and Ronald Williams popularised an algorithm that trained multi-layer networks. The idea: measure the error at the output, then use the chain rule of calculus to work out how much each weight in each earlier layer contributed to that error, and nudge every weight downhill. Because the units use a smooth activation (a sigmoid instead of a hard threshold) the error surface is differentiable, and gradient descent can walk down it. A hidden layer lets the network bend its boundary, and XOR — the 1969 counterexample — falls in a few hundred steps. The same algorithm, with more layers and more data, trains every model in chapters 10 and 11.

A MYCIN-style rule chain

Demo · backward chaining with certainty factors

A simplified fragment in MYCIN's style (illustrative rules, not clinical advice). Answer the questions the engine asks; watch it chain rules toward the goal and combine certainty factors.

Inference trace
Certainty factors combine as CF = a + b·(1−a) for two supporting pieces of evidence, and a rule's conclusion is capped by the weakest of its premises. This is not probability theory — Shortliffe knew it — but clinicians found it matched how they talked about evidence.

Backpropagation learns XOR

Demo · live training

A 2 → 4 → 1 network with sigmoid units, trained here in your browser by gradient descent on the four XOR examples. The background shows the network's output for every point; the four dots are the training data.

0.80
Each step: forward pass (compute outputs), compute the mean squared error, backward pass (chain rule gives ∂error/∂weight for every weight), then weight ← weight − lr × gradient. Sometimes it gets stuck in a poor local minimum — re-initialise and try again; that too is faithful to 1986.
  • 1972Shortliffe begins MYCIN at Stanford; DENDRAL (chemistry) had pioneered the approach from 1965.
  • 1973The Lighthill Report ends most UK AI funding: the first AI winter.
  • 1980XCON at Digital Equipment configures VAX orders; saves an estimated $25M a year.
  • 1982Japan's Fifth Generation project; Hopfield networks revive interest in neural computation.
  • 1986Rumelhart, Hinton & Williams, "Learning representations by back-propagating errors."
  • 1987Lisp-machine market collapses; second AI winter begins.
  • 1989Yann LeCun trains a convolutional network to read handwritten postcodes for the US Postal Service.
1971– 1995

Chapter 8 · The personal computer

A computer on every desk

The microprocessor made a whole CPU cost a few dollars. Hobbyists put one in a box, a spreadsheet gave businesses a reason to buy the box, and a graphical interface let everyone else use it. By 1995 the components of a PC were standardised well enough that you could assemble one from parts — and people did.

The Altair 8800 (1975) was a kit with an Intel 8080 and toggle switches for input; Bill Gates and Paul Allen wrote a BASIC interpreter for it and founded Microsoft. The Apple II (1977) shipped assembled with colour graphics, and VisiCalc (1979) — the first spreadsheet — turned it into a business tool. IBM's PC (1981) used off-the-shelf parts and an open bus, which let other companies build compatible machines and made Intel and Microsoft the platform. The Macintosh (1984) brought the graphical interface developed at Xerox PARC — windows, icons, mouse — to the mass market, and Windows 95 made it universal.

Under the hood, the PC settled into a pattern that still holds: a CPU, RAM on a fast bus, a "chipset" that fans out slower connections, expansion slots (ISA, then PCI, then PCIe) for graphics and other cards, and storage on its own interface. The diagram below is the modern version; hover a part.

Anatomy of a modern PC

Demo · tap a component
Tap a part
The PC as a system

Fast things attach directly to the CPU; slower things go through the chipset. Widths of the connecting lines are drawn roughly in proportion to bandwidth.

Since ~2008 the memory controller and the fastest PCIe lanes have lived inside the CPU package, so the old "northbridge" chip is gone. The chipset (once the "southbridge") handles USB, SATA, audio and extra PCIe lanes over a narrower link.
  • 1971Intel 4004; Unix and the C language mature at Bell Labs over the next two years.
  • 1975MITS Altair 8800 on the cover of Popular Electronics; Microsoft founded.
  • 1977Apple II, Commodore PET and Tandy TRS-80: the "1977 trinity."
  • 1979VisiCalc. Businesses buy computers to run one program.
  • 1981IBM PC with MS-DOS; the 8088 CPU and open architecture spawn the "clone" industry.
  • 1984Macintosh: the GUI goes mainstream.
  • 1991Linus Torvalds posts the first Linux kernel; the free software stack for servers is born.
  • 1993Intel Pentium (3.1M transistors); 1995 Windows 95 and PCI slots standardise the desktop.
1969– 2007

Chapter 9 · The networked world

Computers that talk

A network is a computer's way of borrowing another computer. The idea that made it scale was to chop every message into small packets and let each one find its own way.

Telephone networks reserved a circuit for each call. Paul Baran (RAND) and Donald Davies (NPL) independently proposed in the 1960s that data should instead be split into packets, each carrying its destination address, and forwarded hop by hop by routers that need to know only the next hop. No central switch, no reserved line, and a broken link just means packets take another route. ARPANET sent its first packet in October 1969. Vint Cerf and Bob Kahn's TCP/IP (1974; adopted network-wide on 1 January 1983) layered the job: IP delivers packets unreliably, TCP on top numbers them, retransmits losses and reassembles the message. Ethernet (1973) handled the local wire, and Tim Berners-Lee's World Wide Web (1989–91) put a document format, an address scheme and a protocol on top so that anyone could publish. Search (Google, 1998) made it navigable; the iPhone (2007) put it in every pocket.

For AI the network mattered twice over: it created the corpus — the web is what large language models are trained on — and it made it possible to rent a warehouse of GPUs by the hour.

Packet switching

Demo · send a message
Each packet carries a sequence number and the destination address. Routers forward to whichever neighbour is currently best. The receiver sorts by sequence number and asks for any that never arrived. Cut a link and the network routes around it without anyone reconfiguring anything.
  • 1969ARPANET's first message, UCLA to Stanford ("LO" — the system crashed before "GIN").
  • 1973Ethernet (Metcalfe, Xerox PARC).
  • 1974Cerf & Kahn publish the TCP design; 1983 ARPANET switches to TCP/IP.
  • 1989Berners-Lee proposes the Web at CERN; first site online 1991.
  • 1993Mosaic browser; the Web becomes graphical and popular.
  • 1998Google founded on PageRank; 2006 Amazon launches EC2, renting computers by the hour.
  • 2007iPhone: a networked computer in every pocket.
1997– 2016

Chapter 10 · Parallel hardware meets deep networks

The graphics card that learned to see

Backpropagation had worked since 1986. What it lacked was data and arithmetic. The web supplied the data; a chip designed to draw video games supplied the arithmetic. In 2012 the two met, and a decade of "AI winter" ended in one afternoon.

A CPU is built to run one thread of unpredictable instructions as fast as possible: big caches, branch prediction, out-of-order execution. A GPU is built to do the same simple operation on millions of pixels at once: thousands of small arithmetic units, each doing the same thing to different data (SIMD). Nvidia's CUDA (2007) let programmers use that for anything, and it turned out that a neural network — a huge stack of matrix multiplications — is exactly a "same thing to different data" problem.

Fei-Fei Li's ImageNet (2009) gathered 14 million labelled images and ran an annual competition. In 2012 Alex Krizhevsky, Ilya Sutskever and Geoffrey Hinton entered AlexNet, an eight-layer convolutional network trained on two consumer GPUs for a week. Its error rate was 15.3%; the runner-up's was 26.2%. Within two years every entrant was a deep network. Convolutional networks work because they learn small filters — edge detectors, texture detectors — and slide them across the image, sharing the same weights everywhere; deeper layers combine those into parts and objects. ResNet (2015) showed that with "skip connections" you could train networks 150 layers deep, and in 2016 DeepMind's AlphaGo, combining deep networks with tree search, beat Lee Sedol at Go, a game thought to be a decade away.

Serial versus parallel

Demo · 64 identical operations
Each tick, every free lane processes one element.
CPU · 4 wide cores
ticks: 0
GPU · 64 narrow lanes
ticks: 0
A real GPU has tens of thousands of lanes and a matrix of them can be fed from memory at terabytes per second. The catch: every lane must want the same instruction. Branchy, unpredictable code stays on the CPU; matrix maths moves to the GPU.

What a convolution does

Demo · slide a 3×3 filter
Input 8×8 (0 = black, 9 = white)
Kernel
Output 6×6
At each position the filter's nine numbers are multiplied by the nine pixels beneath and summed into one output pixel. A convolutional network does not hand-write these filters: it starts them random and learns them with backpropagation. Its first layer usually rediscovers edge detectors on its own.
  • 1997IBM's Deep Blue beats Kasparov (search, not learning); Hochreiter & Schmidhuber publish the LSTM.
  • 1998LeCun's LeNet-5 reads cheques in production; convolutional nets are practical but niche.
  • 2006Hinton's "deep belief nets" and the word deep learning.
  • 2007Nvidia releases CUDA: general-purpose programming on GPUs.
  • 2009ImageNet published; Raina, Madhavan & Ng train networks on GPUs.
  • 2012AlexNet wins ImageNet by a landslide.
  • 2013word2vec: words as vectors, where king − man + woman ≈ queen.
  • 2014Generative adversarial networks; sequence-to-sequence translation with attention (Bahdanau).
  • 2015ResNet; 2016 AlphaGo defeats Lee Sedol 4–1.
2017– 2026

Chapter 11 · Attention, scale and language

Predict the next token

A large language model is a very large neural network trained on one deceptively simple task: given some text, guess the next piece. The architecture that made this scale is the transformer, and its central idea is that every word should be able to look at every other word.

Recurrent networks read text one token at a time, carrying a summary forward, which made them slow to train and forgetful over long spans. The 2017 paper Attention Is All You Need (Vaswani et al., Google) dropped recurrence entirely. Each token is turned into a vector (an embedding). In an attention layer, every token computes three vectors from its own — a query ("what am I looking for?"), a key ("what do I contain?") and a value ("what do I pass on?"). The score between two tokens is the dot product of one's query with the other's key; a softmax turns the scores into weights summing to 1; the token's new representation is the weighted sum of the values. Because all tokens do this simultaneously it is one big matrix multiplication — a GPU's favourite meal — and because the weights are data-dependent the model can learn that in "the pharmacist checked the order because it looked wrong," it should attend to order.

Stack dozens of these layers with small feed-forward networks between them, train on trillions of tokens to predict the next one, and something unexpected happens: the model must learn grammar, facts, and a fair amount of reasoning simply because they help it predict. GPT-2 (2019, 1.5B parameters) wrote passable paragraphs; GPT-3 (2020, 175B) could follow instructions given as examples in the prompt. RLHF — fine-tuning on human preferences between candidate answers — turned a next-token predictor into an assistant, and ChatGPT (November 2022) put that in front of a hundred million people in two months. Open-weight models (Meta's Llama, 2023; Mistral; and others) plus quantisation tools like llama.cpp made it possible to run a capable model on a laptop or a home server.

What the model is, physically

A file of numbers. A 7-billion-parameter model is 7 billion learned weights; at 16 bits each that is 14 GB, which must sit in fast memory during inference because every token generated touches every weight. Quantising to 4 bits cuts that to ~4 GB with a modest quality loss, which is what makes local inference practical. Generation is memory-bandwidth-bound: tokens per second ≈ bandwidth ÷ model size, which is why a GPU's fast VRAM matters more than its clock.

Attention weights, step by step

Demo · softmax over scores

Pick which token is asking (the query). The bars show an illustrative attention pattern for that token — the shape a trained head produces for pronoun resolution. Below, drag raw scores to see how softmax turns them into weights.

softmax(s)ᵢ = eˢⁱ ⁄ Σ eˢʲ — drag the scores
Softmax exaggerates differences: raising one score by 2 multiplies its share by e² ≈ 7.4. "Temperature" in a chat model divides the scores before this step — lower temperature sharpens the distribution, higher flattens it.

A language model you can watch think

Demo · n-gram next-token prediction

A transformer predicts the next token from context using learned weights. This tiny model does the same job by counting: it learned which word follows which pair of words from the paragraph below, and samples the next word from those counts. Same task, no neural network — which is exactly what makes the difference visible.

0.8
Candidates for the next word (probability after temperature)
Training text (about 190 words)

the pharmacist checked the order before the dose was sent to the floor . the dose was adjusted for renal function because the patient had a low clearance . the patient was started on an antibiotic after the culture came back . the culture grew a gram positive organism so the team narrowed therapy . the team asked the pharmacist to verify the interaction between the two drugs . the two drugs share a metabolic pathway so the level was monitored . the level was drawn before the next dose and the dose was held . the nurse called the pharmacist about the order because the rate looked wrong . the rate was corrected and the infusion was restarted on the floor . the model predicts the next word from the words before it . the words before it are the context and the context is all the model can see . the model was trained on text from the web and the text was tokenized . the weights of the model were adjusted by gradient descent until the loss was low . the loss was low so the predictions were good and the team was pleased with the model .

With two words of context and a few hundred words of training text the model repeats its sources. Scale that to a context of hundreds of thousands of tokens, trillions of training words, and weights that generalise instead of counting, and the same next-word loop produces a conversation.

Will it fit? Memory for local inference

Calculator · weights in memory
Estimated memory needed
— GB
Weights = parameters × bits ÷ 8. The KV cache (what attention remembers about the context) is estimated for a typical 8B-class architecture with grouped-query attention and grows with context length. Real runtimes add a few hundred MB of overhead. Rough guide, not a spec sheet.
  • 2017"Attention Is All You Need": the transformer.
  • 2018BERT (Google) and GPT-1 (OpenAI): pre-train on raw text, fine-tune for tasks.
  • 2019GPT-2, 1.5B parameters; 2020 GPT-3, 175B, and the scaling-law papers.
  • 2021AlphaFold 2 predicts protein structures at near-experimental accuracy; Codex writes code.
  • 2022InstructGPT/RLHF; Stable Diffusion; ChatGPT launches 30 November.
  • 2023GPT-4, Claude, Llama (open weights), llama.cpp brings inference to consumer hardware.
  • 2024–26Multimodal models, long contexts, agents that use tools, and small models that run on a phone.
2026today

Chapter 12 · Where the two lines meet

From transistor to token

Type a question into a chat model and every chapter of this page runs in order, in about a second. Read the stack from the bottom up.

Ch 5 · transistorsBillions of MOSFETs, each an insulated-gate switch, patterned onto silicon at a few nanometres.
Ch 2 · gates & addersSwitches wired into AND/OR/XOR, adders and multipliers — the arithmetic of matrix multiplication.
Ch 3 · CPUFetch–decode–execute runs the operating system, the tokenizer and the inference program.
Ch 6 · memoryModel weights stream from HBM or DRAM through caches; the KV cache holds the conversation so far.
Ch 10 · GPUThousands of lanes multiply the input by the weight matrices, layer after layer.
Ch 9 · networkPackets carry your prompt to a data centre and the tokens back — or nowhere, if the model runs at home.
Ch 4 · neuronsEach unit is still a weighted sum and a nonlinearity, as in 1958.
Ch 7 · backpropThe weights were set by gradient descent over trillions of tokens; inference just uses them.
Ch 11 · attentionEvery token weighs every other; softmax picks what matters; the next token is sampled.
Ch 1 · the programAnd the whole thing is a stored sequence of instructions the machine follows without understanding — Jacquard's cards, at scale.

What has not changed

Turing's ceiling still holds: a language model computes nothing a 1948 machine could not, given enough tape and time. What changed is the price of a switch (from a dollar to a trillionth of a cent), the price of a byte, and the discovery that if you make a function large enough and differentiable, you can fit it to the world instead of programming it. The expert systems of chapter 7 tried to write medicine down as rules; the models of chapter 11 read the literature and inferred the rules, and neither approach removes the need for someone who can tell when the answer is wrong.

Where the open questions are

  • Hardware: transistors are near atomic limits; gains now come from 3D stacking, chiplets, and purpose-built matrix engines, with memory bandwidth the binding constraint.
  • Efficiency: training frontier models costs gigawatt-hours; small distilled models and better quantisation are moving capability toward local devices.
  • Trust: a next-token predictor has no built-in notion of truth. Retrieval, tool use, and verification against sources are the current answers — the same discipline as checking an order.
2026your lab

Chapter 13 · The stack on your own hardware

What happens when you ask pharm-gemma a question

Every chapter above runs, in order, on machines you own. Here is the path a question takes from your phone to the OptiPlex and back, and — more usefully — the arithmetic that predicts how fast it will be, so you can decide what to run locally and what to send to a frontier model before you download anything. (Everything that enters any of these systems is de-identified; the local box is about cost, availability and control, not a PHI boundary.)

iPhone / iPad Open WebUI · Shortcuts tokenizer runs server-side prompt · ch 9 tokens stream back Tailscale (WireGuard): encrypted, on your own network optiplex · Ubuntu 24.04 · always on llama.cpp a stored program · ch 3 SSD gemma weights file flash · ch 6 32 GB DDR4 weights + KV cache · ch 6 i7-6700T 4 cores · AVX2 (8-wide SIMD) ch 2 adders · ch 5 MOSFETs ch 10 lanes, but only 32 HD 530 iGPU: not used for inference per token embed · ch 11 × N layers: attention feed-forward softmax · ch 11 sample (T = 0.3) every weight read from RAM, once, for every token load once (mmap) ~25 GB/s
The request path in your lab. The copper double arrow is the one that sets the speed: for each generated token, every byte of the model crosses the CPU–RAM link once. On the OptiPlex that link is dual-channel DDR4-2133, about 34 GB/s on paper and ~25 GB/s in practice.

Two phases, two bottlenecks

Answering has two distinct phases, and they hit different walls. Prefill reads your prompt: all the prompt's tokens go through the network at once as one big matrix multiplication, so it is compute-bound — more cores, wider SIMD, or a GPU make it faster. Decode generates the reply one token at a time; each token needs the whole weight file streamed from memory, so it is bandwidth-bound — only faster memory helps. This is why a GPU with 288 GB/s of VRAM beats the OptiPlex's CPU by ~10× on generation even though the chip does far more arithmetic than that ratio suggests, and why Apple's unified memory (120 GB/s on an M4 Air) makes a laptop a respectable inference box.

The rule of thumb, then: tokens/s ≈ usable bandwidth ÷ model size in bytes. Quantisation (chapter 11) helps twice — the model fits, and there are fewer bytes to stream. The estimator below applies that arithmetic to your machines.

Will it fit, and how fast? Your devices

Estimator · bandwidth ÷ bytes
devicememory for modelneedsfits?generationread a 2k-token prompt300-token reply
Generation = bandwidth × ~65–75 % efficiency ÷ (weights + KV bytes touched). Prefill assumes a realistic fraction of each chip's peak arithmetic. "Memory for model" is what the runtime can actually use, not the sticker size — macOS reserves roughly a quarter of unified memory from the GPU, and the OptiPlex needs headroom for Ubuntu, Docker, Jellyfin and Nextcloud. Shaded rows are the machines you have today; the last two are reference points for what a discrete GPU changes, not a plan. Treat all figures as ±30 %.

What the numbers say about your setup

The OptiPlex's 32 GB is generous for fitting models — a 27B model at Q4 loads with room to spare — but its DDR4 bandwidth means anything above ~9B generates slower than you can read. The sweet spot for an always-on box like this is a 4–9B model at Q4–Q6: that is exactly where Gemma E4B landed when it passed your vancomycin AUC test, and the physics agrees with the clinical result. The MacBook Air, with 3–4× the bandwidth, is the machine for trying larger models interactively; the phone and iPad have the bandwidth for a small model but only ~5 GB to hold one, which is why they work best as the thin head. The two GPU rows show what a card would buy: roughly 10× on the OptiPlex, 30× with a high-end card. But look at what that 10× actually is — an 8B model answering faster — against a frontier model that answers better, on demand, for less per month than a card costs. That is the arithmetic behind staying with what you have: run the small always-on model locally for the cheap, private, repetitive work, send the hard reasoning to the cloud, and let any future PC be a gaming machine that happens to run LLMs well as a side effect.

threadsSet -t 4 to the 4 physical cores of the 6700T, not the 8 logical ones. Hyper-threads share the same AVX2 units and the same memory bus; extra threads just fight over them.
context lengthThe KV cache (chapter 11) grows with context and is streamed alongside the weights on every token. Set -c to what the task needs, not the maximum the model allows.
quantisationQ4_K_M is the usual balance; Q6_K is nearly lossless at 1.5× the bytes; Q8 is for when you suspect the quantised model is dropping clinical nuance — test it the way you tested the base model.
keep it residentA model that has to be reloaded costs a full read of the weights file from the SSD each time. A long-running llama-server with --mlock (so the OS never pages the weights out) holds it in RAM; with 32 GB you can afford to leave pharm-gemma loaded permanently.
batch the promptPrefill is the phase your 4 cores are worst at. A long system prompt is re-read on every request unless the runtime caches it; keep the clinical system prompt tight and let the retrieval layer supply detail.
where to run whatSince everything is de-identified before it goes anywhere, the split is by fit, not by sensitivity: the OptiPlex for the always-on, repetitive, zero-marginal-cost jobs (drafts, extraction, the "Ask Pharm" shortcut against llama-server); the MacBook for interactive trials of bigger models; a frontier model for anything where answer quality matters more than where it ran.
  • ch 9Tailscale is packet switching plus WireGuard encryption plus a coordination server that tells your nodes each other's addresses; MagicDNS is the Web's naming idea applied to your own machines.
  • ch 3llama.cpp is a stored program: the CPU fetches, decodes and executes its instructions; the GGUF model file is data it reads. Same memory, two roles — the von Neumann idea.
  • ch 6Model loading is the memory hierarchy in action: SSD → DRAM once, then DRAM → cache → registers billions of times per second during decode.
  • ch 10AVX2 is SIMD on the CPU: 8 float32 multiplies per instruction per core. A GPU has thousands of such lanes; that difference is the whole gap in the table.
  • ch 11Temperature 0.3 in the pharm-gemma Modelfile is the softmax knob from the attention demo — sharpening the distribution so clinical answers stay consistent run to run.

review · the whole course on one screen

thirteen sentences

Each chapter reduced to the one line worth keeping, with your check result beside it. Reread the ones that did not land, then run the practice below — the same questions, shuffled, graded by you.

key terms · all chapters

practice run · 39 retrieval questions, shuffled

Answer each in your head, reveal, and grade yourself honestly. Missed questions are kept so you can run just those next time — the spacing is up to you: come back tomorrow, then next week.

your notes

About the demos. All demos run entirely in this page. The perceptron and XOR trainers use real learning rules; the n-gram generator is a real (tiny) language model; the attention bars are an illustrative pattern; the MYCIN fragment is a teaching simplification and not clinical advice. Dates follow the standard histories (Ceruzzi, A History of Modern Computing; Russell & Norvig, AI: A Modern Approach; Computer History Museum timelines).

Scores and notes live only in this browser.