Pangram verdict · v3.3
We believe that this document is fully human-written
AI likelihood · overall
HumanArticle text · 1,988 words · 5 segments analyzed
Winslow Homer’s most famous watercolor rendered as a child’s drawing.Lately I’ve been asking myself: what might artificial intelligence be good for besides answering questions and writing code? My answer is the latent spaces within AIs themselves will become a new medium for creativity. I will first explain what I mean by latent space, and then at the end of this explanation, I offer possible ways scientists and artists may use the latent spaces inherent in neural nets to serve as a new platform for creativity.A Large Language Model (LLM) is like a small zip file that contains all human knowledge. It takes massive arrays of 100,000 GPU chips working in the cloud, and costing billions of dollars, to compress all of human writing into a small working model that could run on one single GPU chip. Even the biggest frontier models compress down to several hundred gigs, which is small enough it can fit on a card in your palm. In a strange but real way the resulting tiny file contains all the information that is on the internet and in our libraries. This tiny card holds a significant proportion of what humans collectively know. Of all the remarkable aspects of AI, this astounding feat of compression may be the least appreciated. This dense, high order compression of human knowledge — called “latent space” — may also be a new medium itself.This extreme compression of knowledge within latent spaces was not the original intention of the researchers who invented LLMs. The book smartness they contain came as sort of a surprise to the people training them, and we are still trying to figure out how they actually work. What we can say for sure is that the LLM does not contain copies of everything it knows. For instance it knows all Shakespeare plays, and it could create a new play that sounded exactly like Shakespeare, and can even quote famous lines in his plays, but nowhere in the model are the actual texts of Shakespeare. Instead there is simply the abstract information about all the plays, the plots, the characters, the words, the style, the references. Likewise, the LLM could recognize the face of almost any person, and it could generate any possible human face, but nowhere in its code are copies of human faces. Rather, the model is storing all the information about human faces, without storing any faces.This is weird. Until recently we might have thought that all the information about a thing would take up more storage space than the thing itself.
That may be true for a single thing, but not for the aggregate of all things. That is because most things share a lot of common attributes with other things. The neural nets of an LLM do a magic trick by abstracting the information of everything at once, so that it uses the myriad common relationships between things and ideas to compress and abstract them into this virtual “latent” or hidden space.All three terms in “Large Language Model” are key. For “Large”, the models contain all the knowledge in, say, Wikipedia, and all the text from decades of the internet, all webpages and online discussions, and all the scanned books and journals in most libraries. So far, the power of the model keeps increasing as it gets scaled up in size. The more information it is trained on, the more connections, the better it gets.The “Language” part of LLMs turns out to be the secret sauce. LLMs were originally invented to do automatic language translation, that is all. But instead of teaching it the rules of language, which is what earlier AI researchers did, this time no language expertise was required. Instead, a neural net absorbed a very large database of human written language (the internet), with the goal of having the neural net (AI) extract out all the hidden patterns of language below our awareness contained within those billions of documents. The goal of the program was to replicate, imitate and synthesize the patterns of language as it is used everyday by humans.The results shocked everyone. Sure the LLMs could translate language like a human, but the AI also displayed glimpses of human-like intelligence. They could also be creative with language, like they could write up a sales pitch in the style of a sonnet. Some early researchers were spooked by this emergent behavior, including a Google researcher who felt Google’s LLM had an internal intelligence that should not be turned off. We now understand that the intelligence we see in LLMs comes from the logic within the language they were trained on. (See my Why Are LLMs Smart?)The form of this new mindfulness — the “Model” part of LLMs — is a latent space. Latent space is an abstraction, a map built not in two dimensions, but in billions of dimensions. Imagine a brain made up of billions of straight long arrows going in all directions. Each arrow is dedicated to one idea or one thing. There is an arrow for dogs and an arrow for cats.
Related arrows are located next to each other. So the map shows cats and dogs sharing a nearby arrow for fluffy fur. They also share an arrow for ears, and one for tails. Those two attributes are also shared by other animals (other locations) as well. Most of what a dog is is shared by mammals, so this overlap is one source of the compression.You can think of every concept that we can put into words as being a direction in this space. The dog arrow is really a direction of dogness. Catness is a direction, and so is fluffiness. Anything can become more catlike, or fluffier. You start with a shoe, or a chimney, or a fern, and you can push it along the cat direction and make it more catlike. Or you can push it in the direction of apple toward more appleness, or of smoothness, or in the direction of reddish, or excitement, or more circular. You can also reverse direction and make it less catlike, or less red, less atomic. There are billions of directions in this space.Related things are near each other in this space. Cats and dogs share many attributes so they intersect many common arrows, such as tails, whiskers, ears, four legs, animals, short, life, etc. But because they hear, they also intersect the microphone vector; because they can jump, they intersect with basketball. Cats are stealthy and intersect with spies. Because dogs are loyal they intersect the vector of patriotism.Every thing, every concept has a specific location in the map of this huge space, but instead of having just two coordinates (x,y) each thing has a billion-long coordinate. So an old rusty gasoline lawnmower buried in weeds is a very specific intersection with a very long address. Each of its thousands of attributes (rust, gas, lawn, cut, weeds, push, red, dirt, clippings, roar, etc.) has its own direction intersected. Nearby in latent space is a lawn mower that is more in the rust direction, or less red, but also more catlike, or more doglike, or less spaceship-like, or more like whipped cream. That point may represent a real thing or only a virtual or theoretical thing. This mapping works for not just nouns, but any idea, any sound, any image.
The whoosh of a splash of water is a direction in latent space. The aha moment in invention. The fright seeing a snake on a path. The notion of a prime number. All these are contained within a single map. This is one of the most astounding, yet underappreciated aspects of an LLM latent space: Everything — everything! — appears on just one map. We’ve never had a system to integrate everything we know and everything we can imagine. One map for all! This has long been a holy grail.Just to be clear, no human action is doing the mapping. The system itself, the LLM, is mapping each bit of the world, all things, all attributes, all art, all words, all ideas. And astoundingly it creates this map, this latent space, not piecemeal, but all at once simultaneously. (To do so requires an immense, energy-hungry, massive cluster of chips, all connected together with miles of wires — the famous data center now in short supply.)While training, the LLM is fed millions of books, billions of web pages, and billions of pages of text from social media. It reads every word on each of them, and once this entire library of material is loaded into its mind, it massively calculates all the interconnecting vectors, all the relative directions pointing to each other. The scale of this vast synchronized parallel calculation is staggering. It then throws away the books, the text, the images, and only keeps this tangled web of directions and vectors. These billions of directions are called its parameters. As we build larger and larger models, mapping more and more material, the parameters increase. The latest models on the frontier of AI contain trillions of parameters, meaning there are trillions of directions, or trillions of attributes that it uses to map every idea or thing it has seen.Something as complicated as a book winds up as both a point in latent space and a journey through latent space. All the notions encountered in a story (window, mid-day stroll, street, vendor, chat, anger, fight, forgiveness) are directions, and as sentences pile up, the directions shift around, going one way and then intersecting in another. The story is really a journey through latent space, which very much mirrors the journey-like experience we have when we read.So a book contains a sequence of vectors in latent space.
But the sum meaning of a book is also just a single point or direction in itself. For instance if I reference the book The Iliad, I’m referring to the whole book, and its vector is closely related, and therefore “nearby” to the other epic war narratives like Beowulf, The Mahabharata, or even Apocalypse Now, even though many parts of them only tangentially intersect. The more related a thing or idea is, the more directions (vectors) it shares with similar things. This is in part how LLMs know stuff. They search for patterns nearby.When you ask an LLM a question, it will find the answer in latent space. Your question itself begins as a direction, which points to the answer. The LLM addresses each word in your prompt one by one, with each new word shifting the direction of where it goes. The model travels through latent space with each word of the prompt, searching for its answer, step by step. In this way the answer is grown, rather than found.We naively imagine that an LLM has a mind that thinks a thought and then expresses it. But the LLM finds the answer as it writes the words. There’s no pre-formed thought “behind” the words that then gets translated into language. The words are the thinking. The path through latent space and the answer are the same thing happening simultaneously. In the most modern versions of an LLM, the model will proceed through a “chain of thought” intermediate stage, which jots down words and ideas as it thinks about a problem. Even here the chain of thought is the thinking, not a report of thinking that happened elsewhere. The model isn’t reasoning privately and then writing it down — the writing-it-down is the reasoning.As an answer grows along the direction of the prompt, the natural question is how does the LLM know when to stop? How does it know when it is correct? The astounding answer is that “correctness” and “completeness” and “cohesiveness” are vectors in this space, too. Any correct answer shares the same “correctness” direction with all other factually true statements. In other words, correctness, truth, cohesiveness, completeness, comprehension, etc are all essentially patterns that are mapped in this space.