←  Machine Learning 1  ·  Ideal Machine Intelligence

What basis functions look like

Machine Learning 1 — a companion to lecture 3, §3.1

A linear model can only draw a straight line. Basis functions are the trick that lets it draw almost anything — without giving up any of the mathematics that made the straight line easy. This page is the picture behind that sentence. Step through it, then play with the controls.

0 of 7

Why this works

Every curve on the plot is a sum of the same fixed shapes, each multiplied by one number. Change the numbers and you change the curve; the shapes never move. That is the whole idea, and it is why the model is still called linear — it is linear in the weights w, even though it is thoroughly nonlinear in the input x.

That distinction is not cosmetic. Because the model is linear in w, the squared error stays a convex quadratic in w, and the best weights can be written down in closed form — the normal equations. Press Fit to the data and you are watching that solution being computed, not searched for.

Watch what breaks. Push M up towards 14 with the noise turned up. The curve starts chasing individual points, and the weighted bumps below the axis grow huge and start cancelling each other out. That is overfitting, and those exploding weights are exactly the quantity ridge regression penalizes in §3.5.

And what a bad basis looks like. Switch to the polynomial basis and raise M. Polynomials are global: every basis function is nonzero everywhere, so moving one weight changes the fit across the whole plot, and the ends fly off. Gaussians and sigmoids are local, which is usually much better behaved.