←  Machine Learning 1  ·  Ideal Machine Intelligence

Sliding a Gaussian over the data

Machine Learning 1 — a companion to lecture 3, recapping maximum likelihood

Maximum likelihood asks one question: of all the Gaussians we could have drawn this data from, which makes it least surprising? Move the two sliders and watch the answer. The left panel is the data and your Gaussian; the right is the whole landscape you are searching, with your position marked. The last four frames ask a second question — what if we already believed something before the data arrived?

μ = 2.00
σ = 1.00

What the numbers mean

The likelihood is a product over the data: one factor per point, each of them a number below one. Even at the best possible fit, sixty such factors multiply out to about 10−34. Push the sliders into a corner — σ small and μ off to one side — and it falls past 10−1400. That is not a sign anything has gone wrong; it is what multiplying sixty small numbers does.

It is also past what a computer can hold. Double-precision arithmetic bottoms out around 10−308, so the product would come back as exactly zero and every position in that corner would look equally bad. This page never forms it: it sums logarithms and converts to a mantissa and an exponent only for display.

That is the practical half of the argument for the negative log likelihood. The other half is that it costs nothing: a logarithm turns the product into a sum, and because it is an increasing function it leaves the location of the optimum exactly where it was. Maximizing the likelihood and minimizing the NLL are the same act.

Two things worth trying. Put μ in the right place and then make σ very large: every stem gets short, the curve flattens, and the NLL climbs. A Gaussian that is too wide explains the data no better than one in the wrong place — it just spreads its bets.

Then make σ very small. Points near the mean get tall stems, but the ones in the tails get essentially zero density, and because the NLL is a sum of −log terms a single near-zero factor dominates everything. The landscape shows this as a steep wall on the left.

A prior, and what multiplication looks like

The last four frames add one thing: a belief held before the data arrives. Maximum likelihood asks which Gaussian makes the data least surprising. Bayes’ rule asks something else — given what we thought beforehand and what we have now seen, what should we think? The answer is the posterior, and it is the prior times the likelihood:

p(μ, σ | D)  ∝  p(D | μ, σ)  p(μ, σ)

The three panels are that formula. All three are densities over the same two axes, so the multiplication happens cell by cell — and in logs it is an addition, which is how the page computes it, for exactly the reason the likelihood was never formed directly. The posterior is dark only where the prior and the likelihood are both dark. What has been left out is p(D), the evidence. It does not depend on μ or σ, so it rescales the third panel without moving anything inside it; that is why the middle sign is a ∝ and not an equals.

One caveat. The prior here is an independent Gaussian on each parameter, of width τ. A Gaussian on σ is not a prior — it puts mass below zero — so what is drawn is that Gaussian restricted to the half plane σ > 0. The conjugate choice for a variance is an inverse gamma; swapping it in changes the shape of the first panel and changes nothing these frames are for. The one exact result on this page is the last frame, where σ is held fixed and only μ is inferred. That case is conjugate, and the numbers printed there are the closed form.

What N = 1 is for

Drag N down to one and look at the middle panel. There is no minimum. A single observation is explained perfectly by a Gaussian centered on it with σ as small as you like, and the likelihood grows without bound as σ → 0. Maximum likelihood fails here: the estimate σ = 0 puts infinite density on the one point we saw and zero everywhere else. The green ring slides to the floor of the panel and stays there.

The prior is what rescues it. The posterior in the third panel has its mode in a reasonable place — near (2.7, 1.7) with the prior as it loads — because the prior’s own decay beats the likelihood’s blow-up. Regularization is not a trick bolted on to make the optimizer behave; it is what a prior does. Slide N back up and the likelihood takes over; by sixty observations the prior barely registers, and the posterior mode and the maximum likelihood estimate are almost the same point.

The conjugate case

The last frame is the one to write out. Take σ as known, a prior μ ~ N(μ0, τ²), and N observations with sample mean x. Then the posterior over μ is Gaussian, and

1/σN² = 1/τ² + N/σ²
μN = (μ0/τ² + N x/σ²)  ÷  (1/τ² + N/σ²)

Precisions add, and the mean is the precision-weighted average of what you believed and what you measured. Everything the earlier frames show by eye is in those two lines. With N small the prior term dominates and the posterior sits near μ0; each further observation adds another 1/σ² of precision, so the posterior narrows like 1/√N and slides onto the sample mean. On the data as it loads the posterior mean is 31% of the way from the prior mean to the sample mean at N = 1, 70% at N = 5, and 96% at N = 60. Those are the closed form above, and they agree to six decimals with a posterior formed numerically on a 200,000-point grid.

Where this goes next

Nothing here is specific to a Gaussian. Write down a model with parameters, write down the probability it assigns to the data, take the log, and minimize the negative of it. That recipe is the whole of §3.2 in the lecture, where the model is a line through (x, t) pairs and the answer comes out as least squares.

Put the prior back and the same substitution carries you to §3.5. The negative log posterior is the sum of squares plus a term in the squared length of the weights, which is ridge regression — and the penalty strength λ = α/β is nothing but the ratio of the two precisions that were added together above. The three panels and the ridge penalty are the same statement written twice.