What Were Spectre and Meltdown? One of Them Got Fixed
The processor does the work, decides it should not have, and throws the result away. Architecturally, nothing happened. No register changed. No instruction retired. There is no trace of it in any state the programmer is allowed to see.
Except the cache is faster now for the thing it touched.
That is the entire finding, and it took twenty-three years for anyone to publish it.
What were Spectre and Meltdown?
Two families of attack, disclosed together on 3 January 2018, both exploiting speculative execution — the technique by which every fast processor built since the mid-1990s runs ahead on guesses rather than waiting to be certain.
-
Meltdown —
CVE-2017-5754. Reads kernel memory from an ordinary unprivileged process. Breaks the isolation between user space and the operating system. -
Spectre —
CVE-2017-5753(bounds check bypass) andCVE-2017-5715(branch target injection). Tricks a program into speculatively executing code that leaks its own secrets. Breaks isolation between processes, and between a program and its own sandbox.
Jann Horn of Google's Project Zero reported the problem to Intel, AMD and ARM in June 2017. He was not alone. Werner Haas and Thomas Prescher at Cyberus Technology found Meltdown independently, as did Daniel Gruss, Moritz Lipp, Stefan Mangard and Michael Schwarz at Graz University of Technology. Paul Kocher, with Daniel Genkin, Mike Hamburg, Moritz Lipp and Yuval Yarom, arrived at Spectre separately.
Three independent groups, converging inside a few months on something that had been sitting in the silicon since 1995. That is usually a sign that a field's tooling has just crossed a threshold, not that anybody got lucky.
Why processors guess
A modern core can retire several instructions per cycle. A read from main memory costs on the order of two hundred cycles. If the core stopped and waited every time it hit a branch whose condition depended on a pending load, it would spend most of its life idle.
So it does not wait. It predicts which way the branch will go, and starts executing down that path immediately, on the assumption it will probably be right. When the condition finally resolves: if the guess was correct, the work is already done and gets committed. If it was wrong, the results are discarded and execution restarts from the correct path.
This is not an optimisation bolted onto the side. It is the reason a processor from 2018 is faster than one from 1995 by rather more than the clock speed accounts for. Remove it and you do not get a slightly slower chip; you get a fundamentally different and much worse one.
What the rollback does not undo
The discard is thorough with respect to architectural state — registers, memory, flags, everything the instruction set defines.
The cache is not architectural state. It is a performance structure, deliberately invisible to the programming model, and nothing in the rollback touches it.
So if the speculatively executed code read address X, then X is now in cache. Later, timing a read of X tells you whether it is cached, and therefore whether the discarded path touched it. Do that across a range of addresses and you recover a value the processor was never supposed to let you have — not by reading it, but by observing which door the wrong path opened.
The secret is not leaked through the result. It is leaked through the timing of everything the result caused.
Meltdown, and why it was fixable
Meltdown is the more dramatic of the two and the less interesting, because it has a fix.
On the affected hardware, a user process could issue a read of a kernel address. That read is illegal and will fault — but on an out-of-order core the permission check and the data fetch are not strictly ordered, so the value was momentarily available to subsequent speculative instructions before the fault was delivered. Long enough to encode it in the cache.
The fix is architectural rather than clever. Operating systems had, for good performance reasons, mapped the entire kernel into every process's address space, marked inaccessible. Meltdown made mapped but inaccessible insufficient. So the mapping was removed: KPTI, kernel page-table isolation, which grew out of the earlier KAISER work at Graz.
It works, and it costs. Every system call now involves switching page tables and the TLB flushing that implies. Workloads that spend their time in user space barely notice; workloads that make heavy use of system calls — databases, network-intensive servers — notice considerably. The commonly cited range at the time was a few per cent to around thirty, and the honest answer was always it depends entirely on your workload.
AMD was not affected by Meltdown. Its permission checks happen early enough that the value never becomes speculatively available.
Spectre, and why it was not
Spectre has no KPTI, and this is the part worth understanding properly.
Meltdown is a mistake — a specific implementation decision, on specific hardware, that turns out to be wrong. Fix the decision and the attack is gone.
Spectre is not a mistake. It is the observable consequence of speculation working correctly. A branch predictor that can be trained is a branch predictor that can be trained by an attacker; a cache that makes recent accesses faster is a cache that reports which accesses happened. There is nothing to correct, because nothing is behaving incorrectly.
It was demonstrated on Intel, AMD and ARM alike — not because all three made the same error, but because all three build fast processors, and this is what fast processors do.
So the response has been mitigation rather than repair: retpolines, microcode updates exposing new barrier instructions, compilers inserting speculation fences, browsers coarsening their timers and moving origins into separate processes. Each addresses a known pattern. Each new pattern — and there have been many, through the whole family that followed — needs its own.
Eight years on, that is still the state of it.
The 1995 claim
The scope figure everyone quotes deserves reading slowly, because it comes from the researchers rather than a vendor, and it is carefully worded.
The Meltdown paper's claim is that every Intel processor implementing out-of-order execution is potentially affected — effectively every processor since 1995, excepting Intel Itanium and Intel Atom before 2013.
1995 is the Pentium Pro, the first out-of-order x86. So the claim is not that a bug was introduced and went unnoticed for twenty-three years. It is that the property was there from the moment the architecture became speculative, in every generation since, in hundreds of millions of machines, doing precisely what it was designed to do.
Nobody looked, because there was no reason to. Performance structures were not part of the threat model. They were beneath the abstraction, and the whole point of an abstraction is that you do not look underneath it.
The embargo that broke
The coordinated disclosure date was 9 January 2018. It did not hold.
The patches were being developed in public. The Linux kernel had visibly urgent page-table isolation work landing over the Christmas period, with unusual backports and conspicuously vague commit messages. Anyone who read kernel mailing lists for a living could see the shape of the thing without being told, and on 2 January The Register published the story. The formal disclosure was brought forward to 3 January.
The consequence was that the first days were confused: partial information, mitigations shipping ahead of explanations, Intel's initial statements read as minimising, and a set of microcode updates that caused unexpected reboots and had to be withdrawn. Congress later asked the companies involved why nobody had told the US government.
The defensible version of the argument is that six months of embargo produced coordinated patches for three vendor architectures and every major operating system, and that this was genuinely difficult. The uncomfortable version is that a vulnerability class affecting essentially all computing was known to a handful of large companies for six months while everyone else ran unpatched.
Further reading
- Lipp, Schwarz, Gruss, Prescher, Haas, Fogh, Horn, Mangard, Kocher, Genkin, Yarom, Hamburg — Meltdown: Reading Kernel Memory from User Space, 2018.
- Kocher, Horn, Fogh, Genkin, Gruss, Haas, Hamburg, Lipp, Mangard, Prescher, Schwarz, Yarom — Spectre Attacks: Exploiting Speculative Execution, 2018.
- Jann Horn — Reading Privileged Memory with a Side-Channel, Google Project Zero, 3 January 2018.
- Gruss et al. — KASLR is Dead: Long Live KASLR (the KAISER paper), 2017.
- LWN.net — A look at the handling of Meltdown and Spectre, January 2018, on the disclosure process.
EXHIBIT 0025 — SPECTRE + MELTDOWN
We put it on a shirt. A slate blackboard in a wooden frame on an easel, drawn as an engraving, wiped clean with broad sweeping strokes, felt eraser resting in the tray.
The board has been erased. The writing is still readable: 0xFFFF880000000000 — the base of the kernel's direct map of physical memory on x86-64, which is to say the exact region Meltdown was demonstrated reading.
Callouts on the left: THE WORK WAS DISCARDED, THE CACHE WAS NOT. On the right, the asymmetry that is the actual point: MELTDOWN GOT A PATCH, SPECTRE GOT MITIGATIONS.
The processor took it back. The cache remembered.
That's the whole idea.