Computer Architecture Roadmap

Start at Level 1 and follow the nine levels in order, from building arithmetic out of logic gates to reading performance counters on a loaded multicore machine. Every stage names what it needs first and what you should be able to do before moving on. Progress is stored locally in your browser.

Where to start

0 / 138 lessons masteredNot started 138Learning 0Practicing 0Mastered 0
  1. 1

    Level 1 · How arithmetic becomes hardware

    Start here
    0/11

    Binary and two's complement, why floating point approximates, and the path from a truth table to an adder. It comes first because every later level treats a register, an ALU and a memory cell as given; here you see they are made of gates. By the end, arithmetic is no longer a primitive — it is something you could build.

    Before moving on: Convert between decimal and two's complement by hand, explain why 0.1 + 0.2 is not 0.3, and draw a one-bit full adder out of gates.

  2. 2

    Level 2 · What a CPU is made of

    0/7

    Registers, the ALU, the datapath they sit on and the control unit that steers it, wired together from the adders and latches of Level 1. Plus the first myth to lose: that clock speed is performance.

    Before moving on: Name what each part of a CPU does while one instruction runs, and explain why two CPUs at the same clock can finish the same program in different times.

  3. 3

    Level 3 · The contract, and executing against it

    0/14

    The ISA as the interface between software and hardware, the distinction from microarchitecture that most performance arguments miss, and the fetch-decode-execute cycle that pipelining then overlaps. It needs the datapath from Level 2, because a pipeline is that datapath cut into stages.

    Before moving on: Read a short assembly listing, say what the ISA fixes and what a microarchitecture is free to change, and trace an instruction through a five-stage pipeline including a stall and a forward.

  4. 4

    Level 4 · Guessing, reordering and doing several things at once

    0/16

    A modern core is a throughput machine: it predicts branches, executes speculatively, runs instructions as their inputs arrive, and only commits in program order. Nearly every mechanism here is a fix for a pipeline hazard from Level 3; SIMD is the other way to do several things at once, one instruction over many elements. This is where the line-by-line mental model dies.

    Before moving on: Explain why sorted input runs faster, why IPC rather than clock speed measures a core's work, and say whether a given loop can be vectorized and what stops the compiler when it cannot.

  5. 5

    Level 5 · The memory hierarchy

    0/14

    The deepest level, because this is where most real programs spend their time. Lines, locality, associativity, replacement, thrashing and prefetching — everything behind "why is this loop slow". It starts from the load and store instructions of Level 3 and asks what happens after they leave the core.

    Before moving on: Classify a miss as compulsory, capacity or conflict, predict the cliff when a working set outgrows a cache level, and explain why a power-of-two stride can be pathological.

  6. 6

    Level 6 · Past the cache, and where the bytes sit

    0/13

    DRAM organization, latency against bandwidth as two separate resources, and the layout decisions — alignment, padding, contiguity, AoS versus SoA — that decide how much of each fetched line you actually use. It assumes the cache line from Level 5, because layout only matters in units of lines.

    Before moving on: Tell a bandwidth-bound loop from a latency-bound one, compute a struct's padded size from its field order, and explain why an array beats a linked list at equal complexity.

  7. 7

    Level 7 · Address translation and protection

    0/9

    The hardware half of virtual memory: the MMU on the path of every access, the TLB that makes it affordable, and the privilege and permission bits that process isolation is actually built from. It comes after the caches and the layout level because a page-table walk is itself a chain of dependent memory accesses — pointer chasing in the hardware — and the TLB is one more cache with its own misses.

    Before moving on: Trace a load through the MMU and TLB, say what a TLB miss costs relative to a cache miss, and explain what changes in the hardware when a system call crosses into the kernel.

  8. 8

    Level 8 · More than one core

    0/18

    Coherence keeping caches consistent, false sharing as its most expensive surprise, NUMA and affinity — then the ordering rules and atomic instructions that every lock and every lock-free structure is built on. It needs both the caches of Level 5 (coherence is about lines) and the out-of-order core of Level 4 (reordering and store buffers are why a memory model is needed).

    Before moving on: Spot false sharing from the layout of a struct two threads update, explain why a store can become visible late on one ISA and not on another, and build a spinlock from compare-and-swap with the right barriers.

  9. 9

    Level 9 · The rest of the machine, and reading it

    0/36

    Interrupts, DMA and the path to devices; GPUs and accelerators as throughput hardware with a transfer bill; the counters that tell you what the machine is doing; and where all of it shows up in code you already write. It comes last because reading a counter, a side channel or a virtual machine only makes sense once you know the speculating core, the translation hardware and the coherent caches it is reporting on.

    Before moving on: Read performance counters to say whether a program is compute-bound or memory-bound, decide whether a kernel is worth moving to a GPU once transfers are counted, and explain Spectre as a consequence of speculation.