Welcome


The first post of any new blog has a lot of pressure on it. For one thing, it has to be good inasmuch as it sets the tone for subsequent posts. But more importantly, someone reading the first post of a new blog probably is asking a very reasonable question: is this the first of many? Or more likely, is it merely the first of one?

Most blogs have one post

We've probably all seen the plots which show that number of posts follows a power law distribution with a fat tail: by far the most blogs have exactly one post. Very, very few have hundreds or thousands. We can't all be Tyler Cowen and Alex Tabarrok!

I do a lot of machine learning in my day job. One of the most important curves we look at is loss (on the y-axis) as a function of training time (on the x-axis). You generally see the training loss (which you're trying to minimize) go down as training proceeds.

But what you care about is building a model that generalizes outside of the training data. That's why you also plot the loss on a dataset (the "test" data) that your training process hasn't seen. And if you're doing things properly, the loss on your test set also goes down.

For a while.

Then the test loss starts to tick back up as your training process starts to overfit to the training data. It's getting better by memorizing your training data in ways that don't generalize outside of it.

These plots look a little bit like... the logo in the top left corner.

A more general phenomenon

I believe this phenomenon actually shows up in a lot of places. In any new domain, you start with learning the easy stuff, then later the more difficult parts. Learning and getting better means, in large part, reducing the error in your predictions about the domain. But there's a natural limit beyond which you can't keep reducing error without ending up overfitting to your training data. This is easiest to see in games and sports:

They've all overfitted to the training environment and it's hurt their "test time" performance. But this phenomenon happens everywhere else too:

Again, what they all have in common is developing narrow skills that don't generalize to new situations.

So why "Reducible Errors"?

In machine learning, reducible errors are the ones your learning system does have the possibility of eliminating. It's the difference between the error you started with and the error left at the bottom of the test loss curve.

One of my biggest goals in life is to learn things that generalize. To be competent and knowledgeable across domains, and to use that knowledge to build new things. Since one of the best ways to learn is to write, this site is part of that effort. My only agenda is to write things I think are interesting, though you'll be the ultimate judge of that.

Let's see where this goes.