Pangram verdict · v3.3
We believe that this entire text is human-written.
AI likelihood · overall
HumanArticle text · 1,722 words · 1 segments analyzed
The new season of Survivor airs next week. My favorite thing about watching Survivor is arguing about the strategy of the show: who should have won the season? Is it better to play under the radar? Or does making big moves set you up to win? Who is the greatest player of all time: Cirie? Tony? Boston Rob? I thought it’d be fun to build a machine learning model to help answer these questions. On Survivor a bunch of real-life strangers are stranded somewhere remote. They manage camp life, compete in challenges for “immunity,” hunt for advantages, and vote each other out every three days at “tribal council.” Once it’s down to the final two or three, the players who have been voted out (now called “the jury”) vote for the winner. The winner takes home a million dollars and the title of “Sole Survivor.” Ultimately it’s a social strategy game. It’s messy and feels in some way outside the scope of what you can do with a mathematical model. But when we watch and argue about the show, we’re implicitly building mental models of what it takes to win. An ML model is just a way to systematize those theories and put them to the test. And maybe, hopefully, it can give a meaningful preview of who will take the title of “Sole Survivor” (assuming, like me, you’re the type to ignore the betting market leaks). I loved using it while watching Season 50, and it identified several interesting threads early on. I’ll share some of those insights with you here as well. The model predicts, episode by episode, who’s most likely to win, and who’s most likely to go home next. You can play with the model’s results here. An overview of the Survivor ML dashboard. The website loads the most recent season by default, but you can explore previous ones too. The top row shows win and elimination probabilities across a season. Below that is a cumulative plot of who’s been ranked #1 most often, which is a cleaner way to see the real contenders. A “Player Breakdown” section lets you pick a player to understand which features bump their odds up or down. More on all of this below. (And if you want all the nerdy ML details, there’s a notes section at the end.) Caution: Spoilers for the TV show Survivor ahead! Building the model The data comes from the survivoR GitHub repo, which has a ton of nicely structured Survivor data like voting history, challenge results, advantages, demographics, edit metrics, etc. I trained two separate models: one predicts who will go home next episode, and one predicts who will win the whole season. This isn’t a lot of data (especially for the Win model, since there are only 50 winners across US Survivor’s history), so take all the results with a grain of salt. Both models use logistic regression. I tried fancier models too, but they performed worse (details in Methodology notes if you want them). I built out a bunch of candidate features and let the model pick which ones mattered. The ones that survived into the Win model are: Times in danger (number of tribal councils where this person received any votes) N (a control for how many players are left. It’s the same for everyone, so it’s not shown in the dashboard, but the model uses it) Confessional share (rolling share of confessionals over the last 3 episodes; this one is more about the editorial storytelling than the other features) Age (this worked best as a curve rather than a straight line, meaning the model includes both a linear and a squared term. So, you’re most likely to win if you’re around 30, and your chances are lower if you’re much older or younger) Number of previous seasons Has advantage (yes/no) The features for the Elimination model are slightly different: N (again, a control for how many players remain) Advantages held (how many advantages, rather than yes/no) Votes against (last 3 episodes) Individual immunity win rate Age (also a curve, but its effect is weaker than in the Win model) An interesting outcome here: the two models kept different features, which means winning is not the same thing as surviving vote-outs, presumably because jury appeal is key and relies on different things. This lines up with what I think is the general take here. For instance, the show Rob Has a Podcast did a recent episode breaking down how Survivor 50 was won and lost, where Rob Cesternino concludes that a low likelihood of elimination doesn’t necessarily give you high odds of winning. How good is the model? Not bad, depending on your perspective. For simplicity I’ll talk mostly about the Win model. I’m not going to pretend you can use this to know who’s going to win. But on average it picks the eventual winner roughly 2x better than chance. What that means: if you pick a player at random, you’ll get the winner right about 10% of the time (this varies with how many people are left, but 10% is an average). My model’s #1 pick is right just over 20% of the time. That 2x improvement holds for most of the game, basically after two episodes have elapsed (once the features have enough data). Being right ~20% of the time might not sound like much. It does mean it’s wrong most of the time. That’s fine to me, there’s a huge amount of unquantifiable strategy and luck in any season that no model like this can capture. But 2x over baseline isn’t bad! It’s also worth separating two things: how the model ranks players, and its exact probabilities. The episode-to-episode bumps are tiny, especially early, when a big cast is splitting up the win equity. And it turns out the ordering may be more informative than the numerical probability of winning — the model’s #1 pick actually wins more often than the probability suggests (for instance, the #1 pick wins 40% of the time going into the finale even though that pick has an average probability of only 29%). Just being the model’s favorite has value. I think that discrepancy is likely an odd quirk of how I built this. More discussion on this in the notes section. Anecdotally, I’ve found it pretty compelling how often the numbers match the assessment of people watching (or at least my own). It heavily favors a lot of the best winners, moves players’ odds up and down in ways that make sense given what’s happening in the game, and often picks up on trajectories before they’re obvious in the main narrative. I got to watch that happen live, since I ran the model alongside Season 50. How did Season 50 go? Pretty well. The model was fun to have as a companion while I watched. A few things I noticed, to show how it can be used: The model flagged Jonathan (the runner-up) as a contender before I took him seriously. He ended up making a lot of strong endgame moves and had the highest P(Win) going into the finale (the wrong choice in the end, but not an unreasonable one, given how close he came). The “Win Probability by Episode” plot shows the relative P(Win) for every player, and you can see his trajectory rising mid-game. The model flagged Jonathan as a strong contender early. Aubry, the actual winner, was one of the model’s main contenders all season. By the finale it did give Jonathan a higher P(Win) than Aubry. But she had a strong early game and was the #1 pick for much of the season. This is where the second metric gives you a different perspective. The “Fraction of Episodes Ranked at #1” is a smoothed-out way to see who the real contenders are, and it only appears on the dashboard once a few episodes have aired. In S50, the only players who ever reach #1 at all are Jonathan, Aubry, Ozzy, and Christian, who are all reasonable players to take seriously. The main contenders for S50. This chart shows the fraction of episodes where each player was ranked #1 for P(Win). It's a smoothed view of the main plot, and it helps identify the real contenders. Aubry was the most consistent pick earlier in the season. The model responds to in-game events realistically. In episode 11, Cirie sensed danger and (correctly) played her extra vote advantage to save herself. She avoided elimination, but she had now received votes (bad) and spent her advantage (also bad). The model reflected this change in her situation accurately, with her elimination probability shooting up and her relative win chances dropping. She went home the next episode. Probability of elimination (left) and probability of winning (right) for Cirie's game. For most of the season, Cirie had a low probability of both elimination and winning. Her P(Win) rose toward the end, but her big move that saved her also left her vulnerable: her P(Elimination) shot up and her relative rank in P(Win) sank, going into episode 12 when she was eliminated. Why was the model so down on Cirie? Cirie is a beloved fan favorite, and I think a lot of people were thrilled to see her play in S50. In my view, she has a mastery of the social game unlike anyone else. Her odds crept up steadily over the game, but the model never had her as high as I personally wanted. Her player breakdown shows her age was capping her win probability. I don’t know whether Cirie’s low P(Win) is a result of the model missing Cirie’s idiosyncratic strengths, or a result of it capturing something real that fans would rather ignore. Probably some of both. Player Breakdown for Cirie in episode 12 (the episode she was voted out), explaining the model's prediction. The left two plots are the Elimination model. The right two are the Win model. The top row is "Model Contributions," where the width of each bar shows that feature's contribution. The bottom row is "Compared to Field," a way to see how the player's values compare to everyone else remaining (bar width here represents the feature's coefficient). There are lots of stories like these! You can explore previous seasons on the dashboard too. What is important for winning?