←  Machine Learning 1  ·  Ideal Machine Intelligence

Watching a posterior tighten

Machine Learning 1 — a companion to lecture 4, §4.3–4.4

A straight line has two parameters, so the space of all possible models is a plane and we can draw the whole of it. The three left-hand panels are that plane three times over — what we believed before this observation, what the observation alone says, and what the two of them come to — and they read left to right as one statement. The fourth is the usual picture: data, and lines through it. The arrow underneath is the recursion: the posterior on the right becomes the prior on the left when the next observation arrives.

observations N = 0

What the parameter panels are

Every point (w0, w1) in each square is a complete model: the line y = w0 + w1x. The origin is the flat line through zero; the far corners are steep lines with a large offset. Shading is how plausible each model is, darker being more plausible, and the cross marks the parameters the data was generated from — w0 = −0.3, w1 = 0.5.

The six small rings in that panel are the six lines in the right-hand one. Each is a w drawn from the distribution being shaded, so the two panels are showing the same six models in the two different ways — a scatter of points here, a fan of lines there. Before any data they are strewn across the whole square and the fan goes everywhere; by twenty observations they are a tight knot and the lines are nearly indistinguishable. The knot and the fan are the same fact.

The middle of the three is the same square, shaded by how plausible the single most recent observation makes each model. It is a ridge, not a blob: an observation (x, t) is consistent with every line through it, and those lines are a line's worth of parameter pairs rather than a single pair. So one observation cannot pick out a model. Two cross two ridges, and a crossing is a blob.

Where the ridge points

The ridge for an observation at x runs along w0 + w1x = t, so its slope in the parameter plane is −1/x. Drag the observation about in the right-hand panel and watch that happen:

observation atridge in the planewhat it pins down
x ≈ 0nearly horizontal the intercept, and almost nothing about the slope
x = ±1diagonal, at 45° a trade-off: intercept and slope in equal measure
large |x|nearly vertical the slope, and almost nothing about the intercept

So if you get to choose where to measure, measure at the ends: two observations at x = ±1 pin a line far better than two near the origin, which between them say the same thing twice.

How the shading is built

The frame called how the likelihood gets filled in takes the middle panel apart, and holds the posterior back while it does — that panel reads not yet, because the likelihood it would be built from is still being drawn. Pick one candidate model, a single point in the square drawn as a black outlined square, and its line appears in data space.

The line is the mean of a distribution over what t could be at each x — the target distribution of §4.3, a Gaussian riding along the line with the same width everywhere:

p(t | x, w, β) = N(t | w0 + w1x, β−1) = √(β/2π) · exp(−½β(t − w0 − w1x)²)

and the likelihood of the observation is the height of that Gaussian at the observed value of t. The dashed drop in data space is r = t − (w0 + w1x), how far the observation falls from the line, so

p(t | x, w, β) = √(β/2π) · exp(−βr²/2)

and that number is printed as the brush moves. It is a density, not a probability, so it is not bounded by one: at r = 0 it is √(β/2π) ≈ 1.99, and that is the largest value anywhere in the square. And it falls away fast — β = 25 here, so β−1 = 0.04 and the noise has standard deviation 0.2; a line missing by 0.5 is two and a half standard deviations out and scores about twenty-three times less than one passing straight through. That steepness is why the dark band arrives almost all at once as the panel paints, and why it is so narrow.

One cell of the middle panel is one evaluation of that formula. Do it for every model in the square and you have the whole thing; press Sweep and watch it fill in, one raster row at a time, with the candidate line swinging round in data space as the brush moves. Drag inside the panel to take the brush over and probe wherever you like. The posterior appears on the frame after, once the sweep has finished.

However long you sweep, the band stays a line rather than a spot: the ridge again.

The posterior is the prior

The left-hand panel is labelled prior for observation n, and underneath it says what it is: the posterior after n − 1. After the first observation there is no separate prior in this problem — only the previous posterior, playing that part. Step N and whatever was in the third panel appears in the first.

So nothing in the machinery cares whether the observations arrive together or one at a time. On the frame where you step through them, the green line prints the difference between the posterior built one observation at a time and the same posterior built by handing all n of them to the formula at once: around 3 × 10−16 over the whole run, which is machine precision. Reversing the order of arrival changes nothing either, so it can be run on a stream.

Why this model and not a curve. Two parameters fit in a square, so the three left-hand panels show the whole space of models. With the nine Gaussian basis functions of §4.5 that space has nine dimensions and any picture of it is a projection. The straight line is the one case in the lecture where both views are complete, and after this the right-hand panel is the only one that can be drawn.

Where this goes next

The right-hand panel draws lines sampled from the posterior, which is one way to see the spread but not a summary of it. Averaging over the posterior instead of sampling from it gives the predictive distribution of §4.5: a mean line with a band around it, and the band is wide where the data is thin. That is the picture lecture 3 promised when it noted that a single fitted w gives the same error bar everywhere, however far from the data you go.