Parker Rex
All videos
Parker Rex

I Built Flappy Bird Using Every AI Model (You Wont Guess Who Won)

April 26, 2025

Watch on YouTube@parkerrex

I ran a hands-on benchmark by rebuilding Flappy Bird in Python with Pygame and feeding the same prompt to a lineup of AI models. The goal was to see which model delivers the most usable, responsive game loop, and Grok 3 turns out to be the surprise winner.

Setup and approach#

  • Same prompt used across all models to keep the comparison fair.
  • Each model was run via a separate folder with a small, self-contained Python script (main.py) that uses Pygame.
  • The workflow: open multiple terminals, cd into each model folder (01, 03, 04 mini, etc.), and run python main.py to test.
  • Observations often involved rendering quirks (bird shapes: square, circle, triangle) and control responsiveness (gravity, speed, and spacebar behavior).
  • Tools noted:
  • VS Code for editing and quick syntax fixes
  • Copilot used live to fix syntax issues when needed
  • Pygame as the game engine required by the prompt

Example prompt approach (paraphrased):

  • A fixed prompt was used for all models to request a self-contained Flappy Bird implementation using Pygame, with gravity, collision, and a responsive spacebar control.

Models tested#

  • Gemini family
  • Gemini 2 flashing
  • 01
  • 03 mini
  • 03 mini high
  • Sonnet / Anthropic family
  • Sonnet 37
  • Sonnet 35
  • EnAnthropic (Anthropic’s style prompt)
  • Other Open AI family
  • Gemini 25
  • OpenAI’s latest model (as of the test period)
  • 04 mini family
  • 04 mini
  • 04 mini high
  • Grock family
  • Grock 3 (the sleeper hit)

Observations and quick verdicts#

  • Gemini 2 flashing
  • Verdict: rough start, slow and unimpressive
  • Gemini 01
  • Verdict: poor baseline
  • Gemini 03 mini
  • Verdict: underwhelming behavior, gravity and control feel off
  • Gemini 03 mini high
  • Verdict: notably better; one of the better performers in the run
  • Sonnet 37
  • Verdict: strong potential, good thinking/adjustments; bugged reset at one point
  • Sonnet 35
  • Verdict: hit-or-miss; some syntax/implementation issues
  • EnAnthropic
  • Verdict: mixed results; reliability varied across runs
  • Gemini 25
  • Verdict: mixed-to-average outcomes; not the standout
  • OpenAI latest
  • Verdict: included in the mix, but not clearly the best in this batch
  • 04 mini
  • Verdict: fast start, but encountered stability/syntax problems; not consistently reliable
  • 04 mini high
  • Verdict: one of the faster, more responsive attempts; still some quirks
  • Grock 3
  • Verdict: the clear winner in this test; finally delivered a reliable, playable result
  • Noted as the “line 129” moment where the implementation aligned cleanly with the prompt
  • Final assessment: Grock 3 beat the others for playable performance and stability

The winner and what it means#

  • Grock 3 emerged as the winner of this Flappy Bird benchmark.
  • Why it stood out:
  • More stable gravity and collision behavior
  • More reliable spacebar responsiveness
  • Fewer quirky rendering anomalies compared to other models
  • Takeaway: even with a single shared prompt and the same task, model performance can be all over the map. The best result often comes from a model that handles the control flow and timing more predictably, not just raw language quality.

Notable moments and gotchas#

  • The experiment wasn’t just about accuracy; it was about “playability” and responsiveness in a simple game loop. Some models produced nice text but struggled to drive a real-time PyGame window smoothly.
  • Syntax hiccups and code integration issues were common (hence the Copilot mentions). Having a quick fix path can save lots of time but doesn’t change the underlying model performance.
  • Visual quirks (bird shapes, velocities) reminded that rendering decisions can skew perceived performance even when logic is correct.

How you can reproduce#

  • Use a fixed prompt across all models to ensure comparability.
  • Create a separate folder per model and place a minimal PyGame-based main.py in each.
  • Run: python main.py from each model’s folder and observe:
  • Responsiveness to spacebar
  • Consistency of gravity and bird movement
  • Stability of the game loop (no freezes or crashes)
  • Grade each model by how playable the result is, not just whether it runs.

Actionable takeaways#

  • When benchmarking AI models on code tasks, prioritize runtime behavior and user experience (e.g., input responsiveness, frame rate stability) in addition to correctness.
  • A single “best” model can be elusive; you may need to test multiple generations within a provider’s lineup to find a playable option.
  • Don’t rely on a model’s language quality alone—test its ability to produce real-time, deterministic results in an executable environment.
  • Having quick tooling fixes (e.g., Copilot or editor assist) is helpful, but keep the core tests model-centric rather than tool-centric.
  • Pygame: https://www.pygame.org
  • Visual Studio Code: https://code.visualstudio.com
  • GitHub Copilot: https://github.com/features/copilot
  • Google Gemini: https://ai.google/discover/gemini
  • Anthropic Claude: https://www.anthropic.com/claude

If you want more AI/battery-tested benchmarks like this, I’ll run additional rounds and share the breakdowns with a focus on real-time performance and robustness.

Transcript

Everyone has their own benchmarks. I want to do a benchmark by making Flappy Bird. If you're not familiar, really popular game back in the day where you hit the space bar and you make the bird fly, right? Most of you have seen it. Went super viral very long time ago. But I had done this with models from a while ago. And you're like, "Whoa, VS Code." Yeah, because I could put augment over here. I'll do a video about that in another video. That's why you should subscribe. So, I have done this with Gemini 2 flashing 01, 03 mini, 03 mini high. What we can do real quick is we'll open up a few terminals. We'll close out this thing on the right and we will CD into 01 and then we will CD into 03. And we'll just do a quick run so you guys can see, you know, what state-of-the-art a couple months ago was I did this in January and we have all them here. So then if we just run Python main pie, this will be Gemini 2 flashing. Whoa. So this is the cheapo cheap boy. Pretty cool. Kind of boring, right? But it works. And I have the same prompt that I'm using for all of them. So if you're curious, that is the prompt right there. So now let's try this next one. So, I'd give uh let's rate these. Gemini 2 flashing. I'd give it a D. Just look bad. Let's do a one. This back in the day should be fantastic. Give us a square. Uh the window size is off. A little bit slower. It feels chunky and it doesn't like these are really close. So, it's actually starting out a lot harder. So, kind of surprised that 01 was actually worse. So, I'll give 01 and F. Let's go over to 03. And I'm curious once we get to the new ones. Whoa. 03 Mini. Really horrible gravity. Very laggy. Honestly, the first one, the cheap one, actually feels the best so far, which is kind of wild. Okay, and let's go get 03 mini high, which should be one of the best. Maybe the best one we've ever seen. Whoa. Okay. Yeah, this is snappier. The gravity's kind of wild, but cool. Whoa, now it's a triangle. Okay. So, we're going to give 03 mini a D. We're going to give 03 mini high A B. Now, up next tests for April 23rd. We're going to give a couple. We're going to try Sonnet 37 thinking. We're going to do Sonnet 35. So, Enanthropic, we're going to do how about Gemini 25 and then whatever the newest one from Open AI is. I honestly forgot. So, let's go in. We're going to open up 04. Let's go 04 mini a rip. So, we'll grab the same exact You must use Pygame. The background should be I can read that to you in a second where I'm actually letting it fly. So, we've got that. Cool. We can pop this open. We'll create a new one and we'll go to 04 mini high. Do that. Those two are now going. We've already gotten the 04 mini very quick. Holy smokes. Wow, that was very fast. Okay, so let's pop open our VS Code and let's see. How did we do this before? So, what was that one? 04 mini. Cool. Let's kill these. And let's go back one and let's do make dur 04 mini. Make dur mini high. Make dur. Oop. Sorry. Yeah. No, we're already back. Make dur. Uh, Sonnet 3.7 Sonnet. I feel like 3.5 is kind of cheating. I don't want to do that one. I want to do Let's do Brock three. That I feel like that's the sleeper. It's not even a sleep. I feel like everyone kind of thinks that's going to be the one where you're like, "Whoa, dude. You literally made the thing." Plus, it's Son's been around for a while. O4 mini. Hi. Okay. O4 mini. And then what's the last one we want is Gemini. I'm rooting on Gemini. Curious what you guys think. Drop a comment right now. Actually, this is how good are you at testing models? How good are you at guessing? This has nothing to do with testing. This is a total gamble. Go ahead and gamble below and tell me, is it going to be 25 or is it going to be Grock 3? I don't know. Find out this summer. Uh, lastly, Gemini 25. Cool. So, let's pop open all of those and we'll start pasting stuff. I actually I find Grock the most helpful for like random stuff I need to do with Python. because I've been doing a lot of it lately. Is it cheating if you do think? Who knows? Let's go into enthropic. Uh yeah, console keys workbench. Now we're in. We are in. Ladies and gentlemen, we're going to make none of that. We're just going to put this here. We're going to see with 37 sonnet. Where's the thinking button? I think I don't see it. Boom. Let's see. Let's see what it comes back with since we're already in here. Then we'll need a vertex cloud console home server. No, we want automations and create prompt. That's good. And let's switch over to 25 Pro Preview 325. No, the newest. No, that's Flash. Okay, let's do this one. Okay. Boom. Exciting. Let's grab this boy here. Pop it into our sonnet. We need Is there anything beyond that one thing? No, it should have should be self-containing. So, where's set touch main.py echo open param close param. Oh, did it freeze? Okay, cool. And then main.py. Cool. That's in there. If VS Code can handle a little file. Good. Dead. Good. Let's grab our mini high. Good. Touch. Just doing the same thing. Touch main.py echo open pram close param gator tail bay.py. Cool. Mini touch.py. Echo. I missed it. Echo. There we go. Come on. Two more. Stick with me. I know the anticipation is just killing you right now. We're going to Sonic. Touch made up. Need a snippet at this point. Go to the bottom. That's annoying. Come on. There we go. Maypod. And then finally, Grock, where you at? And this one. Holy smokes. I feel like Yeah, Gemini is definitely definitely taking the dub here. I just feel it. Absolutely feel it. Look at all that good code. Cut. It's amazing to me how slow that is. I just don't know why. What do you mean bird? No. No. Oh gosh. What do I do? Let's go back. Faster. Okay. And then Gro three. I'm leaned in. Hey, the moment of truth is upon us. My money's on Grock. Feel like that's kind of cheating, but let's see. Go to the glorious leaderboards here. I have no idea how to use co-pilot because it's like grandpa's favorite LM to use or code assistant. So I I don't even know. But we're going to do 37 O three 04 mini 04 mini high five candidates. Let's let rip python main.py up number one. I feel like there's not that much variance. Just honestly, it's going to be like a feel thing. Let's see that. No way. Gemini doesn't even have the space bar. I'm distraught. This is supposed to be the best one. You're supposed to be Google. You were the chosen one. What do you mean? This is I'm cutting the video. Hey guys, thanks for watching. Um, no. But that's so stupid. How do they let the team down like that? Did we paste something wrong? The whole video was supposed to be how I'm so smart that I knew. I mean, does it even run? Press the space bar. It's It's so clear. Oh, I can't get emotional about it. I might start just crying. So, we're gonna forget that ever happened. Gemini, you're dead to me. Google, I know you want to sponsor me. You're still great. I know you you've already vectorized these sentences. Actually, Google, you're so great. Good job, Gemini wins. A+ they get they get a like a whatever below. They get a G. That's after F. And that's the wrong one. So, we're going to put it in the right place. That's such an upset. Okay, let's do main uh Python main.py. Copy that. So, I never have to type it again in my whole life. Cool. 04 mini with the biggest upset ever. Okay, never mind. It won't even start. It won't even start. 04 mini. Bro, you're supposed to be the chosenish one. 04 mini. Did I P? Was it a skill issue? I have Py game. And it's got a big old syntax. What? Let's try it again. Come on, Sam Alman. Well, okay. Let's just get it to run by using this ancient thing from Narnia called Copilot. I By the way, I'm doing this in here. Don't hate on me. I know you're like, "Wow, we're going to throw tomatoes at him because he's using this thing called VS Code." Go ahead, throw the tomatoes. I really don't care. I'm using it because I just want augment and then I'll use Bim. Get at me. But no. So, okay. Fix the syntax in 04 mini or Flappy Bird dies. I'll be set. This is called psychology. This thing is going to crush it now. Okay, so send. And yeah, I could go fix the syntax, but what's the fun in that? I have a little bot that can do it for me. Yeah, I know. I know it's got Whoa, that's kind of cool. Maybe they maybe Microsoft did a good job. Okay, so how do I Grandpa's trying to figure out Grandpa's app. Okay, there we got it. Now we're still in the right place. It runs. Oh my gosh. Okay, the space bar works. Yeah, this is great. No, it's got a little too much of a an accelerator on it. Like I hit it and then it just really zings. You can't even It's impossible. I'm one of the most skilled video game players on the planet Earth. Maybe even Mars. And that's not possible. All right. They get a horrible grade. I'm sorry, Sam. But someone's got to put you in your place. 04 mini. You get an F. Sorry, man. It's just not in the cards. We literally can't. Like, kids are gonna cry if they played this game. Let's run 04 mini pi or mini high and see if our It doesn't run. Wow, man. They they show all these benchmarks. Oh, we can make a travel agent that literally is like a dead job. Why is that the freaking benchmark? We can make a travel agent. No one cares about that. Everyone cares about Flappy Bird. This is the state-of-the-art, most high quality test you could possibly have. Okay. Do we have up arrow? Oh, we do. Wow. Okay. In 04. You can't do ats. I know it sounds like I'm complaining a lot. It's cuz I am. Cuz this is like this is a blood bath. Okay, let's see. Everything's on the line. Open AAI might just close up shop after those few months. Okay, that's good. You did a good job. Self-contained. We can go ahead and close that. What? What's Whoa, whoa, whoa, whoa. Why is it going again? It's trying to cheat. Okay, that's fine. For many Gemini, I'm sorry, Gemini. I just can't believe you let me down. I had so much like my whole all my pride. Okay, run. Whoa. All right, here we go. Why was it a square? Now it's a circle. Why can't they draw a bird? I honestly came into this being like, "Wow, this this is going to be a bird. They're going to draw a bird cuz it's called Flappy Bird." I thought thought there going to be wings instead. We have a triangle sometimes. This is possible. Really hard, but possible. And I know there's going to be a bunch of Python stands game devs in here if they watch it like, "Oh my gosh, this guy's an idiot." Well, guess what? I don't care what you think because this is just the ultimate benchmark. And this isn't This one's still I honestly think the ones before were a little better. I don't know about you guys. I think maybe Yeah, this is broken now. We broke it. Another one bites the dust. If I had an editor, I'd totally We'd put that in right now. 04 mini high. You get a G. We're moving Gemini down to Z. Sorry, Gemini. RIP Gemini. Just kidding. And because Google's listening. So, let's go on to Sonnet 37. Do something for us, please. For the love of computer science, just pull through. What? No way. They can't get syntax. I'm floored. We've had one get it right off the rip. This is fix the syntax. Can't believe this. And here or I get fired. Always put a little bit of a stake on it. I am sitting there like, man, I really got to get this right. And I'm not making that up. There's an an arvix paper that backs that. Oh, look at that. They got these nice things in here and maybe even a little uh little logging. Is that what we see? Yeah. I feel like Sonnet 3 is just This is the one here, fellas. Here we go. Oh, yeah. I mean, this is this is great. This could be a blockbuster hit. This video might be four hours because I'm just going to be able to play this forever. almost too easy. I don't know chat, what do you guys think? You're not chat, but you know, the metaphorical chat, the understood you. This looks good. There's, you know, Sonet 37. It's got that word thinking in it. It thought about this. It's like, you know what? We want this game to be fun. You know, we want everyone in the world to play it. I'm having fun here. I'm at 17. I got to one on the others. We ladies and gentlemen, I would wait a second. It broke. Sonnet, you you're doing such a good job, but then you can't figure out how to reset. I was going to give it an A, but think of all the kids who would just burst out crying with this kind of behavior out of the model. you get a good, you know, we weren't going to give out B+es because I ran up the flag pole to my board and they said, "No, but I'm breaking the rules." Telling the board to f off. You're going to get a B plus. Crowd goes wild. Everyone's freaking out. I know you are. You're sitting there in the chair and you're like, I don't even know what to do today. I'm going to go tell my wife that I saw the coolest thing on the internet today. And I get that because that was I mean, it's just life-changing. Okay, so now we can't get syntax, right? We We still haven't Grock, dude. You literally Elon Musk's like thing like how will it will it run? Let's see. This is maybe as important as getting like a peace prize if this works. Fix the syntax where I snap Grock 3 in half. Come on, you can do it. Line 129. There's not much missing. Wow. Great. It did it. Okay. This is the demo. It's a French word for the unraveling of the story. Everything in your life has led up to this moment. Parker, show them what they what, dude. Oh, okay. So, it's like super super sensitive croc. Okay, video's over. Grock, you are so I'm I gotta just bite my tongue. That's a D. We have a winner here. We're going to put this in the MoMA. Everybody, there's no There's literally if I got five views, I'm guessing I'm forecasting this video is going to get six views, 6,000. And no one got that right. There's no one who watched this video that got that right. There's no way. If you did, you're lying. There's just it doesn't make any sense. I appreciate you watching. Make sure you like the video, even if you hated it because that would just be awesome so other people can just see this intense information that we covered. And of course, subscribe. And if you want to learn more AI stuff that's not Flappy Bench, if you wanted to learn how to make Flappy Bench with your eyes closed with a bandana on, then you can learn that from me on my YouTube or in our community. And that's it for this video. Like and subscribe.