Pangram verdict · v3.3
We believe that this text is a mix of AI and human-written content.
AI likelihood · overall
MixedArticle text · 889 words · 2 segments analyzed
#Artificial Intelligence#World ModelsTable of contentsThe PlanWhat Even is a World Model?How the World Model Learns to PredictFrom screenshots to embeddingsPrediction loss and collapseSIGReg: keeping the embedding space usefulThe Pokémon Training DataPlanning in the Learned WorldWhy the First Plan FailedRollout Fine-Tuning and the Second AttemptThe End…?I think some of the most fascinating work happening in the field of Artificial Intelligence surrounds world models. There are many kinds of world models (and the term itself has become a bit overloaded), but one of the most exciting architectures for world models is Yann LeCun’s JEPA (Joint Embedding Predictive Architecture). There are many variants, and one that caught my eye in particular was LeWorldModel1. It seemed small enough to train locally on my RTX 3080 Ti and had a simpler design than many of the previous JEPA architectures.San Francisco has been covered in Pokémon memorabilia–big advertisements featuring many of the 1026 pocket monsters plastered across the subway lines and bus stops on Market Street. SF was the site of the Pokémon World Championships this year, and many eclectic and joyous Pokémon trainers could be found wandering the temperate hills and concrete financial district of downtown SF. Maybe all the advertisements had subliminally controlled me, but I had decided that a good “world” for our world model would be the 1996 game that started it all, Pokémon Red.2The mechanics of PokémonPokémon Red is a game where you explore a world, collect creatures called Pokémon, and use them in battles.You use the direction buttons to walk around or move through menus. The A button interacts with things, advances dialogue, and confirms choices; the B button usually cancels or backs out. Choosing a starter means approaching a Poké Ball (a container that holds Pokémon) and getting through the dialogue that confirms your choice.Near the beginning of the game, the player is inside Professor Oak’s lab, where he offers you your first Pokémon: Bulbasaur, Charmander, or Squirtle. These are the three “starters,” each waiting in a Poké Ball on a table in his lab.The GIF above is the final result of the model training: the model planned a sequence of button presses that selected Squirtle. There were more difficulties and setbacks than I expected, even though the model had looked promising in simpler tests.The PlanThe goal was relatively straightforward: defeat Professor Oak’s grandson. Breaking it down into multiple steps, I came up with:Get to Professor Oak’s LabFinish Dialogue with OakSelect a Starter PokémonTry to exit the labFace and beat Oak’s grandsonIt quickly came to my attention that this might be ambitious for a first experiment–so I narrowed it down to selecting a starter from a saved state in Oak’s Lab, where acquiring any of the three starters would count as a successful attempt.From the saved position in the lab, having the model press A twelve times is enough. But can the model learn that? The model could also just wander around aimlessly, maybe in perpetual torment inside Oak’s lab. Or the model could press B after every few sequences of A, cancelling its effort when it almost reached its goal.What Even is a World Model?A world model is a model in which we begin with some current state or observation, and some action happens that modifies this state, producing a new observation. Realistically, there should be some sort of correlation or hopefully causation between the action on some state and the new state it produces. The goal of the world model is to learn this correlation. Let’s say the current observation is a screenshot of Professor Oak’s lab, with the player in front of the Poké Balls, and the action is pressing left, moving the player one tile to the left–producing the new end screenshot. The goal would be for the world model to develop some form of intuition that pressing left moves the player left.$$ o_{t+1}\approx F(o_t,a_t). $$Here \(o_t\) is the current screenshot, \(a_t\) is the button pressed, and \(F\) is the function the world model learns to predict the next screenshot \(o_{t+1}\).But the goal should be to generalize–the model shouldn’t learn that pressing left in Professor Oak’s lab moves the player left, but that pressing left anywhere should move the player to the left.3As an aside, this setup might sound somewhat similar to reinforcement learning. But crucially, the world model learns to understand the state dynamics without a reward, aka being reward-free (more on this later).How the World Model Learns to PredictFrom screenshots to embeddingsSo if what we are really interested in is the state-transitions and the dynamics of the environment4 and all we have are observations (in this case, in the form of screenshots from our game), does the model learn to predict screenshots?No, what the model actually learns is to predict within its latent space, also known as its embedding space. We first need to take an encoder that, when given a screenshot, produces an embedding.What is an embedding? ScreenshotEncoder0.24−1.070.63⋮0.18192 numbersAn embedding is a learned representation of an input as a vector of numbers.
An encoder can turn an image, a sentence, or a sound into such a vector, giving another model something it can compare, predict, or use as input.The entries are not hand-labeled features: one coordinate does not have to mean “color” or “position.” Information can be spread across many entries, and some details of the original input can disappear altogether.