Anthropic's New Claude is Kind of Absurd (Best one by far)
February 24, 2025
Claude 3.7 Sunnet is the headline here: hybrid reasoning with near-instant, visible step-by-step thinking, plus stronger coding and GitHub tooling. Parker dives into hands-on testing, prompts, and what this means for builders and workflows.
Quick take#
- Claude 3.7 Sunnet brings hybrid reasoning and visible chain-of-thought-style thinking.
- Major updates for developers: GitHub integration, a more capable Claude Code CLI, and broader agent tooling.
- Real-world testing: using 3.7 for coding tasks, summarization workflows, and prompt-based pipelines (YouTube transcripts to blog posts).
Claude 3.7 Sunnet: what’s new#
- Hybrid reasoning with visible step-by-step thinking
- Users can see the model’s reasoning path and adjust how long it should think.
- Expanded “thinking budget”
- You can control how much the model should think before delivering results.
- Notable lift in coding capabilities
- Improved performance for coding tasks and front-end/CLI workflows.
- Availability across tooling
- Claude 3.7 Sunnet is accessible via multiple interfaces (baseline API, tooling partners, and IDE integrations mentioned like Cursor).
- Extended thinking concept
- A more formal take on long-form reasoning with a practical, usable interface for developers.
GitHub integration and IDE vs in-browser tooling#
- GitHub beta integration
- Connect your repository and pick exactly what to sync or surface (files, folders, or entire repos).
- View model updates and leverage the integration for in-repo prompts and actions.
- IDE vs browser tooling
- Expect continued competition between in-browser tooling and IDE integrations; both sides pushing for a smoother developer experience.
- Enterprise vs new surfaces
- Some features (like in-browser analysis and code tooling) are extending beyond traditional enterprise tiers into broader plans.
Claude Code and coding workflows#
- Claude Code is now more prominent
- Code generation and agent-driven coding improvements are highlighted as a core use case.
- Availability via API and across more plans (including some references to broader access in previews).
- CLI-focused coding demonstrations
- The CLI’s integration with Claude enables code tasks to proceed closer to “the metal,” reducing UI friction.
- Agent-based coding previews
- Early previews for autonomous agent workflows are being touted; expect more on this soon.
Hybrid reasoning and “visible thinking”#
- The model makes its step-by-step thinking visible
- Users can see the reasoning path and how the model arrives at conclusions.
- Budgeted thinking
- You can dial how long the model should think before replying, balancing speed and depth.
- Practical implications
- Better auditing of outputs, easier debugging of complex tasks, and more controllable reasoning depth for developers.
Real-world testing notes and workflows#
- Copywriting and prompt libraries
- Claude 3.7 performed well on writing tasks; the speaker used a library of top prompts to keep outputs concise and clean.
- YouTube-to-blog post pipeline (prompt-based)
- Two-file prompt approach: attach a prompt file and a transcript, then apply “YouTube script to blog post” instructions to generate a structured post.
- Steps demonstrated:
- Upload a video transcript (or 40-minute script) as a text file.
- Use a prompt designed to transform subtitles/transcripts into a comprehensive, user-friendly blog post.
- Iterate with edits to improve structure and readability.
- Transcript summarization experiments
- Experimented with summarizing long-form video content into blog-friendly formats; verified faster processing and robust results.
Practical takeaways for builders and teams#
- Leverage GitHub beta integration
- Start by linking your repo, then select the parts of the codebase you want Claude to surface or modify.
- Tap into Claude Code and the CLI
- Use the terminal-enabled capabilities for faster, more direct code generation and agent tasks.
- Enable hybrid thinking for tricky tasks
- Turn on extended thinking or hybrid mode for complex refactors, architecture decisions, or multi-step tasks.
- Build pipelines around transcripts and media
- Use the two-file prompt approach to convert long transcripts into blog posts, summaries, or show notes automatically.
- Experiment with agent-mode previews
- If you’re a builder, join the agent previews to test autonomous workflows and tool-calling capabilities.
What Parker’s planning to test next#
- Deeper testing with long-form prompts and video scripts
- More agent-enabled tasks (especially around tooling and code changes)
- Comparisons against other models (e.g., Grock 3) on specific tasks like coding and reasoning
- More hands-on demos with the GitHub integration and CLI workflows
Known questions and caveats#
- Availability across providers
- Some features (e.g., specific API coverage or cloud integrations) may vary by account or provider; check the latest docs for your setup.
- AWS page alignment
- There are occasional discrepancies in feature listings across documentation pages; verify feature availability in your own console.
- Progress on agent previews
- Agent previews are evolving; expect incremental improvements and more tutorials as the feature matures.
Final take#
Claude 3.7 Sunnet marks a meaningful jump in practical AI tooling for developers and creators. The combination of visible thinking, stronger coding aids, and tighter GitHub integration lays groundwork for faster iteration and more autonomous workflows. If you’re building AI-assisted tooling or content pipelines, this is a release worth hands-on testing and benchmarking in your own stack.
Links#
- Claude 3.7 Sonnet release notes (overview of features and thinking improvements)
- Claude Code (coding enhancements and API access)
- GitHub beta integration docs (connecting repositories and surface options)
- Cursor integration updates (IDE support and testing)
- Grok comparison notes (for context on how 3.7 Sonnet stacks up on certain tasks)
Transcript
well well well we did get a new model and I made a video about this earlier today I wasn't sure if it was going to come out but I was stating that 3.7 was going to get dropped by anthropic because there was some leak going on and it happened there's a few things we're going to cover First Impressions the GitHub Integrations for the coders this is a huge deal because now you can connect your repo and you can see it was updated today this used to just be on the Enterprise plans and you can select all the different things that you want in there from your GitHub whether it's just a couple files or folders and then you can do all the stuff in here which is great I think uh this will continue to be a trend where you're fighting IDE versus in browser I think that we have a bunch of different companies obviously in this space that are fighting one another which is great for users so let's see when you come in here now you have this GitHub beta you also can see that there's this analysis tool so you can upload csvs looks like that's new my plan ironically ended yesterday so I'm going to be getting a new one but you can see that you now have Claude 3.7 Sunnet and I was just using it to write some copy for the agency that I run alongside my SAS and it's doing really good job so I have this big old prompt library in here and it's got a bunch of the top one% prompts that I found and that I use on the rag so I just came in here and I used this standard one that I have around making stuff concise and so I took a transcription of a one minute video script and I popped it in and I'll just say First Impressions is that the writing is fantastic I think a lot people know that about Claud that's something that they love is that the writing is really good and um you can just see that in this prompt as for three versions it came back and it just sounds clean it's concise it's nice and that is awesome so very very exciting looking forward to using this more I actually asked uh grock I was like what would be a good thing for us to to test here and um it says that I should try text image I don't think that that exists so what I want to do is I'm going to take a script that is from a 30 I think it's like a 35 or 40 minute video on the on chat GPT basically that I did and I'm going to take this big text file I'm going to drop it into here so now that's attached then I'm going to find a good prompt for summarizing and kind of converting that video script into a blog post so I'll open up YouTube to blog post this would typically be pulling the subtitles but instead we have actual transcript so let's take the H how about YouTube script no how to yeah let's do this this how to one and we'll run it so I'm going to give it an objective titles instructions here's all this stuff and special considerations fin instructions okay so we're going to say that I'm going to say two files are attached one is the prompt it's to transform a YouTube DIY tutorial subtitles into comprehensive userfriendly step-by-step blog post use that file as your set of instructions and then apply that instruction onto the other file which is a transcript okay and hopefully I don't hit a ceiling so after to the races it feels faster and again this is like a 40 minute video so it's including these Pro tips this is also a really good prompt so that's probably helpful but I noticed that it's faster and I'm want to kind of read through it what you'll learn base versus premium when to use each yeah having used this model a bunch of different times I noticed that it's better but let's just see what they say let's go to the release notes um CLA add updates no that's not it API updates no how about on the homepage news 3.7 and Claude code here we go there's a lot to cover so today we announced 3.7 Sunnet our most intelligent model today in the first hybrid reasoning model on the market hybrid reasoning Market on the okay that's different um near instant step by step thinking that's visible to the user o there it is extended okay cool I feel like this is going to be an absolute Game Changer I mean we've seen that before but I feel like they're going to do it better just because their last Model had such good staying power okay strong improvements in coding and front end command line tool for agented coding Cloud code is available oh I need to get on that okay so from the terminal you can use it it's in every plan for those it is in vertex let's just check I was looking at this earlier today and it was not there so I'm assuming it's back boom yep available via API cool let's look at AWS and see this still not on that page weird just want to see if it's available via API in here still not in there huh maybe it's just my account who knows but writing seems good the frontier reasoning made practical bodies of philosophy you have the toggle to say if you wanted to think longer self-reflect yep nothing new there can control the budget for thinking that's pretty cool early testing and coding whoa okay so from 49 to 62 whoa okay what is the percent increase from 49% to 62.3% okay that's a big jump 27% increase that is big that is a big jump grock 3 is on here interesting so grock 3 is still better at graduate level reasoning it's still better at visual reasoning we don't know on the coding stuff obviously let's see what this should we be doing like big smile or uh exp funny how they try to make I'm an engineer I'm cat I'm a product manager we love seeing what people build with repository we don't know much about this code base it looks like an app for chatting with a customer support agent let's get Claude to help explain this code base to us that's a lot of Diet Coke Claud starts by reading the higher level files and then it Dives in deeper now it's going through all the components in the project whoa okay so it's indexing it in the CLI here's its final [Music] analysis so say I was asked to replace this left sidebar with a chat history and I'm also going to add a new chat button I'm going to ask quad to help me out here this reminds me a lot of Aer we haven't specified any files or paths and claude's already finding the right files to update by itself Claude can also show its thinking and we can see how it's decided to tackle this problem this is bad news for a or any CLI next it's up so reing a new chat [Music] button Cod identified the build errors and is now fixing them okay I know exactly what this does this is awesome this is yeah having it be in the CLI makes a ton of sense because it's closer to the the metal so to speak where you won't have to deal with any of the UI like chugging along I think the more unified the model provider is with the tools for the code gen the better there's just less stuff you want less stuff and um it's got tool calling I definitely am joining that preview research preview you can just join it okay definitely doing that after this we'll have a video on that up on the channel soon we yep you can do that it's safer okay pioneer years challenging breakthroughs so this is just them saying hey we have this awesome model and we're showing you early stages of full autonomous agents taking more agency with their agents and they're going to learn what's awesome about this whole thing is as they release this and get more users using the platform then they have more data obviously to see what works what doesn't how are users getting positive results um probably an error rate will be assigned to this like how close was the output to what the final um you know like the the conversation closer like what were the things they were commonly running into that were errant but yeah this is all really exciting I think there was another blog post in here yeah the extended thinking so that's what it sounds like it's just basically Chain of Thought They said that they're the first hybrid one which I want to giv an impressive boost I want to see what the difference is there several downsides visible thought process they had it playing Pokemon that's nice let's see the word hybrid in here don't know what that means produ your instance step by step thinking this made visible to the user yeah so we already we spoke about that already but now let's see what the team at cursor if they've brought that in already I'm going to just type in uh sunet 3.7 yep it's already available in cursor cool very impressed by its coding ability especially on real world agentic tasks appears to be new state the art cool so they have a little thought thing highest level of thinking try it out select that and enable agent mode okay so there's sunet thinking and then sunet normal yeah if you can't select it update it let's see what the first impressions are from people Yep they're going after coding cool well that's the first impressions I'm going to start playing with it some more I just wanted to get a quick video out kind of covering the announcements very very exciting for everybody in the space and uh get on it start building building more agents and do all that fancy YouTube stuff like the video so that I show up more and this gets out to more people have a good one