Our intuition for small numbers of dimensions leads us astray in very large dimensions.
For example, why don't gradient descent algorithms get "stuck" at local minima constantly?
In three dimensions, it's quite common to get stuck in a local minima, to have nowhere to advance incrementally without having to go back up.
But this happens vanishingly rarely in very large dimensional spaces.
In that situation, each point is more likely to be a "saddle point."
That is, a local minima in some dimensions, but not in others.
The intuition is about probability.
Imagine in each dimension, you have a 50/50 chance of being at a local minima in that dimension.
The chance you're at a local minimum with 3 dimensions is: 0.5^3= 0.125.
The chance you're at a local minimum with 1000 dimensions is effectively zero.
Another odd thing in high dimensions is how sparse everything is.
This is called the "curse of dimensionality."
To cover a space at fixed density, the number of samples you need scales exponentially with dimensions.
Also, they all become roughly the same distance away from each other.
Thanks to the law of large numbers.
Similar to summing up the numbers of 1000 six-sided dice.
Each die roll is random… but over many, many of them the outcome tracks arbitrarily close to the average.
The blessed operation is "descend," and the cursed operation is "cover."
So why does deep learning work despite this?
The manifold hypothesis is that real, meaningful high-dimensional data does not evenly cover the space: it lies on a much lower-dimensional curved manifold.
Imagine all of the random images of a certain size–the vast, vast majority are just radio static, and the meaningful images are a tiny subset.
So even if the dataset has 50,000 dimensions, the intrinsic dimension of the domain might be, say 50 dimensions.
Thanks to Claude for helping me develop and strengthen these intuitions!