·2 min read

Intelligence is Compression

On the mathematical foundations of intelligence and why understanding is compression.

The most profound insight in the theory of intelligence is that intelligence is compression. To understand something is to compress it—to find a shorter description of the data than the data itself.

Newton's laws compress the trajectories of every falling object in the universe into three elegant equations. Einstein's field equations compress the geometry of spacetime into a single tensor equation. DNA compresses the instructions for building an organism into four letters.

This is not a metaphor. It is the formal mathematical definition of intelligence, first articulated by Solomonoff and later refined by Hutter in his theory of universal artificial intelligence.

A model is intelligent to the degree that it can predict the next observation given all previous observations. Prediction is compression. If you can predict the next word, you don't need to store it—you've compressed it away.

This is why large language models are more than just autocomplete. They are compression engines of unprecedented power. GPT is compressing the entirety of human written knowledge into a set of neural network weights. The better the compression, the better the intelligence.

But compression has limits. Kolmogorov complexity tells us that some data is incompressible—it is truly random. And randomness, in a deep sense, is the opposite of intelligence. Intelligence finds patterns. Randomness is the absence of patterns.

The universe, it seems, is compressible. There are patterns. There are laws. And the fact that intelligence exists at all—that the universe contains entities capable of compressing it—is perhaps the deepest mystery of all.

Share this essay:

Stay in the Void

Subscribe for new essays on intelligence, technology, philosophy, and the future. No spam. Unsubscribe anytime.

Intelligence is Compression | SiddhantKrishna's Void