Compression is prediction

#recommend for the Programming friends or the Math inclined. Does a very simple sketch of compressing text, and then a simple sketch of making a model to predict text. Explains entropy, the average number of bits needed to encode a symbol (ie the next letter in a compressed string, or the next word in an AI model), and how high entropy leads to worse compression and worse prediction.

GPT-2 actually fares better (smaller) than an order-1 model on a small Charles Dickens paragraph in compression output size. However, GPT-2’s model is huge and requires huge compute to run, whereas gzip is tiny and requires tiny compute.

Compression ratios is a solved problem, but an open question is how to minimise entropy.

I find it interesting how collections of frequency maps are used to create a lossless alternative representation of text (or bits) that is much smaller. Similarly, JPEG, mp3 and lossy compression finds a close-enough representation of bits (music, images) in a highly efficient vector space through a Fourier transform. Or in essence, you can view something as bits over time (a sequence of bits encoding the raw text), or a collection of ratio of frequencies at a particular time; but in text compression, it seems that “frequency ratios at time t” are actually “frequency ratios preceded by n characters” (in an order-n model).

Ultimately, learning the similarities of text prediction models (aka ShatGPT) and compression strengthens my distrust of generative AI. Compression is not used to predict, it’s used to reconstruct, and maybe that’s where we should limit it’s usage.