I'm not sold on your flops comparison between the kid and the machine on Pokemon. The internet is replete with Pokemon FireRed strategies, and Fable has clearly seen them during pre-training. So the kid is spending their 50 hours learning the details of Pokemon FireRed's world, and Fable is spending its 50 hours on perception and navigation issues. So not comparable for two different reasons: (1) pretraining helps Fable a lot *on* PokemonRed, and we should account for that, and (2) the flops are spent so differently (Fable itself guesses that if its navigation and perception issues were generally better, it could do it in 6-10 hours).
(I queried Fable on whether it knows a strategy: "The recall is extensive. Without looking anything up: the full type chart; the gym order (Brock through Giovanni) and each leader's specialty; that Charmander is the hard-mode starter because of Brock and Misty; where the HMs are (Cut from the S.S. Anne captain, Fly from the Celadon rooftop area, Surf and the Gold Teeth in the Safari Zone, Strength from the Fuchsia warden); the Silph Scope → Pokémon Tower → Poké Flute → Snorlax dependency chain; the Rocket Hideout under the Celadon Game Corner with the spinner-tile maze; the Elite Four lineup and rough teams. That's not "a good strategy" in some abstract sense — it's most of the actual solution to the game.")
(I'm separately not sold on your numeric comparison there; I think it's plausible that all-in the LLM used about as many flops as the kid. But that's a different issue.)
The larger claim that the milestone is solved is fine.
(1). In terms of Pokemon, after looking at the GPT chat it seems to me that an AI cannot yet beat 10 years old at Pokemon. This probably will happen by the end of the year. Organizing a birthday I would say is 55-45 though I would think this definitely gets falsified. Making a 4 minute video seems extremely close to becoming falsified by the end of the year. Making a 2 hour visual novel I think close to 50-50. Solving IMO’s hardest problem I think is 100% falsifiable.
(2). In terms of both an AI beating a top human at STS 2 and organizing a wedding, I don’t think either will become falsified by the end of the year. Making a 10 minute or more film and a 10 hour visual novel will probably be both falsifiable by the end of the year. As for writing a math paper for a top journal, I don’t think this has been falsified and I am leaning towards this getting resolved by the end of the year.
(3). I know you’re a collaborator with METR. Do know if METR is planning to update their time horizon suite with new tasks and evaluate new AI models?
Epoch recently streamed 5.6 Sol playing STS2. It seems much worse at it than STS. I interpret this to mean that pretraining knowledge is helping a lot. It’s clearly nowhere near a top player.
That said, plenty of top STS players had relatively low win rates when STS2 launched. There’s a lot of learning that happens across games (reflecting on failures) and outside of games (learning from others) that they have benefited from.
I'm not sold on your flops comparison between the kid and the machine on Pokemon. The internet is replete with Pokemon FireRed strategies, and Fable has clearly seen them during pre-training. So the kid is spending their 50 hours learning the details of Pokemon FireRed's world, and Fable is spending its 50 hours on perception and navigation issues. So not comparable for two different reasons: (1) pretraining helps Fable a lot *on* PokemonRed, and we should account for that, and (2) the flops are spent so differently (Fable itself guesses that if its navigation and perception issues were generally better, it could do it in 6-10 hours).
(I queried Fable on whether it knows a strategy: "The recall is extensive. Without looking anything up: the full type chart; the gym order (Brock through Giovanni) and each leader's specialty; that Charmander is the hard-mode starter because of Brock and Misty; where the HMs are (Cut from the S.S. Anne captain, Fly from the Celadon rooftop area, Surf and the Gold Teeth in the Safari Zone, Strength from the Fuchsia warden); the Silph Scope → Pokémon Tower → Poké Flute → Snorlax dependency chain; the Rocket Hideout under the Celadon Game Corner with the spinner-tile maze; the Elite Four lineup and rough teams. That's not "a good strategy" in some abstract sense — it's most of the actual solution to the game.")
(I'm separately not sold on your numeric comparison there; I think it's plausible that all-in the LLM used about as many flops as the kid. But that's a different issue.)
The larger claim that the milestone is solved is fine.
Great post Ajeya, thanks!
(1). In terms of Pokemon, after looking at the GPT chat it seems to me that an AI cannot yet beat 10 years old at Pokemon. This probably will happen by the end of the year. Organizing a birthday I would say is 55-45 though I would think this definitely gets falsified. Making a 4 minute video seems extremely close to becoming falsified by the end of the year. Making a 2 hour visual novel I think close to 50-50. Solving IMO’s hardest problem I think is 100% falsifiable.
(2). In terms of both an AI beating a top human at STS 2 and organizing a wedding, I don’t think either will become falsified by the end of the year. Making a 10 minute or more film and a 10 hour visual novel will probably be both falsifiable by the end of the year. As for writing a math paper for a top journal, I don’t think this has been falsified and I am leaning towards this getting resolved by the end of the year.
(3). I know you’re a collaborator with METR. Do know if METR is planning to update their time horizon suite with new tasks and evaluate new AI models?
Epoch recently streamed 5.6 Sol playing STS2. It seems much worse at it than STS. I interpret this to mean that pretraining knowledge is helping a lot. It’s clearly nowhere near a top player.
That said, plenty of top STS players had relatively low win rates when STS2 launched. There’s a lot of learning that happens across games (reflecting on failures) and outside of games (learning from others) that they have benefited from.
What is the meaning of probability of “self-sufficient AI” having the value of 2.5% this year?