Parker Rex
All videos
Parker Rex

How I Discovered New AI Models Before Everyone Else (copy this)

February 24, 2025

Watch on YouTube@parkerrex

Rumors and early signals point to a major shift for coders: Anthropic’s Sunnet 3.7 could bring stronger standard and extended thinking, better complex reasoning, and bigger context windows. This video breaks down what’s being teased, how it could change coding and workflows, and what to watch for next.

What’s being talked about (Sunnet 3.7 rumors)#

  • Anthropic Sunnet 3.7 is appearing in model catalogs and on leaked feeds; some users report seeing it, others don’t. Release is rumored within a couple of days.
  • The model is pitched as stronger across writing, coding, and general tasks, with notable emphasis on complex reasoning and extended thinking.
  • Context window targets and tool use are in scope, with hints of richer agentic capabilities and potential image-generation tie-ins (Stable Diffusion mentions).

New capabilities and how they differ#

  • Standard thinking vs extended thinking
  • Standard thinking: typical reasoning and prompting patterns.
  • Extended thinking: deeper Chain-of-Thought-style reasoning, more meticulous multi-step analysis.
  • Complex reasoning
  • Promises improved problem solving on difficult tasks without heavy prompt engineering.
  • Tools and actions
  • RAG, product recommendations, forecasting, targeted marketing, code generation, quality control, parsing text from images.
  • Possible inclusion of more robust “computer use” / agentic capabilities.
  • Context and data handling
  • Hints of a larger context window (possible >200k tokens) and broader knowledge integration.
  • Unclear about fine-tuning support at launch.

What this could mean for coders#

  • Coding workflows
  • Enhanced code generation and debugging help with less specialized prompt crafting.
  • Potentially stronger integration with code reasoning tasks and multi-step solutions.
  • Workflow impact
  • Could shift how we approach RAG, testing, and tool use in coding pipelines.
  • Expect improvements in tasks like project scaffolding, error diagnosis, and large-context code reviews.
  • Tooling ecosystem
  • Cursor and other assistants may need new prompt strategies to align with extended thinking.
  • Early-access experiments likely; expect competing approaches to evolve quickly.

Benchmarks and signals to watch#

  • AER benchmark
  • Sunnet 3.7 is expected to push results in conjunction with latest tooling (e.g., R1 + Claw combos).
  • Web and developer benchmarks
  • Web Dev Arena leaderboard remains a reference point for top-performing models in coding tasks.
  • Grok vs other research tools
  • Grok is highlighted as strong for deep web research; comparisons with Kimy and Perplexity will be informative once new model data lands.
  • Context window vs real-world use
  • Larger context windows help with long files, complex reasoning, and multi-step tasks; verify whether the rumored 200k+ token window becomes practical.

What to monitor next#

  • Official rollout details
  • Confirm release timing, regional availability, and integration with Bedrock, Vertex, or other platforms.
  • Fine-tuning and customization
  • Check whether fine-tuning is supported at launch or if it’s reserved for later.
  • Integration with image generation
  • If Stable Diffusion-related attributes show up, watch for image generation capabilities tied to Sunnet 3.7.
  • Competitive landscape
  • How Cursor, Kimy, Perplexity, and Grok adapt to Sunnet 3.7; any new “specialized prompt” strategies emerge.

Takeaways for developers#

  • Be prepared for bigger context handling and deeper reasoning out of the box.
  • Start experimenting with extended-thinking prompts and track where they save time or reduce debugging cycles.
  • Keep an eye on access and rollout timelines; early-access signals can guide when to allocate time for testing.
  • Watch for real-world prompts and workflow changes: how code-generation tasks, RAG workflows, and large-scale reasoning tasks perform in practice.

Actions you can take now#

  • Track rumors into an action list: note release timing, platform availability, and any official announcements.
  • Prepare your toolchain
  • Consider how you’d adapt Cursor or other assistants to leverage extended thinking and larger context windows.
  • Plan tests around code generation, debugging, and long-context tasks.
  • Benchmark planning
  • Decide which internal tasks to re-test (e.g., large codebases, multi-file reasoning, long dependency graphs) once Sunnet 3.7 lands.

If new details drop, I’ll break down the specifics and what they mean for real-world coding work. Like and subscribe to stay updated, and I’ll see you in the next video.

Transcript

we're on the verge of something here something big a new release anthropic is coming out with their new sunit which is a huge deal for coders specifically but in general it's going to be probably way better at writing and all sorts of tasks so I came across this x poost this morning and it's showing that 3.7 Sunnet is on the model catalog now I went to AWS where this is being leaked and I don't see it but showing up for other people this is also a newer account CU I don't use AWS and if you uh don't use it for a while they can even shut you down so I'm trusting that other people are seeing it and it's scheduled for release in two days and why is that a big deal well if you look at the previous kind of changes in the model when we jumped from the older version of sunet to 3.5 it had a pretty drastic improvement where it went from 33 to 49% on the swe bench verified Benchmark and then obviously tool use and so this is what it is going to look like there's a lot of stuff in here so yeah that's kind of wild I'm hoping that comes out everywhere in a couple days not just on Bedrock but we don't know that for sure and you can also find I made a couple of slides that just kind of like break it down so the biggest difference is it will have standard thinking and then extended thinking so it's going in the direction of the Chain of Thought stuff and it will obviously include the computer use all the normal agenta capabilities but then the biggest one is the complex reasoning now some would say oh we can do this already with putting in recursive prompts and step by step and critique yourself but it's nice not to have to know how to prompt that and so a couple of the things that are mentioned in it is obviously you can do the normal stuff of rag and product recommendations and all these kind of different ones but I think what will be a big deal is how it plays out for coders because if we look at this Benchmark which I think is a really good one it's by AER it's still holding the strongest when you combined uh deep seek R1 with the latest claw and I think even with the latest releases from Chachi BT kind of consensus is that CLA has held strong through all this is just wild and we don't know how it ranks against Croc but we will soon find out if you look at the web dev Arena leaderboard this one's still at the top which is great and I think that we're going to see a lot of big changes as soon as this comes out and um see what else I'm hoping that the context window is larger and here I had grock do a little bit of research on it so hopefully it's over 200,000 that'd be awesome I doubt that they'll launch with the fine-tuning support but who knows again the biggest difference is just that enhanced complex problem solving don't know if they're going to bring search but that would be awfully helpful too some people are questioning whether this is a joke or not I hope it's not a joke I like I said I couldn't find it on AWS and I also couldn't find it on vertex but that doesn't mean that it's not coming out I'm interested as to why they went with AWS but it's probably because they received a massive investment from our boy Jeffy B this was how the person found it was they got this output and you can see that it self-describes as the ideal choice for powering AI agents especially customer facing agents in complex AI workflows supported you the supported use cases are rag over vast amounts of knowledge product recommendations forecasting targeted marketing code generation quality control parsing text from images gentic computer use content generation so it says that the attributes are reasoning text gen code gen Rich Text formatting and agentic computer use I see something about stable diffusion in here which is cool so this could be the first time that we see anthropic combining with image generation which would be great I'm hoping that we see like I mentioned earlier some of the web and kind of deep researchy stuff anthropic has yet to put out Chain of Thought I think if you put this in a timeline it was cool we had chat and then we had it was like if we had a rough timeline it was one we have chat two we have Chain of Thought then three we get the kind of deep research where it's grocking through a bunch of websites and you see that with Kimmy the Chinese model R1 kind of blew up obviously of chat GPT Pro but I do believe that Gro seems to be the best and this is a pretty cool workflow for how someone's doing competitor research I personally think grock is great because I was able to just type in you know what's included in this one that's coming up and it searched across 27 if I had put that same thing into let's say kimy we can test that out and I'll do long thinking and we can put that one by side with anthropic or sorry perplexity you can see this one did 30 pages you see all of the news of people talking about this but it is weird cuz it's pulling from an article from January so that's not that great but this is besides the point I think that this will be coming out in the next week and I'm very very excited everyone that is a coder should also be very excited I'm sure cursor will launch something the same day that rolls it in what I kind of worried about and I think something that we've all run into as coders is how a new model responds with cursor because cursor has its own set of system prompts that it's doing and we can see that it's nailed it with 3.5 but I'm sure that they get Early Access I'm sure that they've already been testing and trying to figure it out Chain of Thought stuff typically doesn't work that well I found in cursor but it's gotten a little bit better with these specific prompting that you can do rather than just one cursor rules file you can have specific ones so let's stay tuned on what's Happening any sort of news that comes out I'll make a video and if not make sure you like the video that's your way to pay the YouTube algorithm do all that fun YouTube stuff and I'll see you in the next one