Claude Releases New Model (Haiku 3.5) | Code Testing & Impressions
November 5, 2024
Claude Haiku 3.5 is put to the test in Parker’s hands-on look at coding and UI prompts, with a practical lens on cost, performance, and real-world workflows. Here are the key takeaways and hands-on findings.
Benchmark landscape at a glance#
- Claude Haiku 3.5 performance
- On the AER (Agent for Coding) leaderboard: around 4th place with ~75% in coding tasks, trailing the top model.
- Pricing vs. capability: marketed as a cheaper option, but user feedback flags the price as high for a model that isn’t the absolute leader.
- Reasoning benchmarks
- Demonstrates solid chain-of-thought reasoning, but still not on par with the flagship, top-tier models.
- Compared to Gemini/other premium models, it’s competitive in some aspects but notably costlier.
- Practical takeaway
- If you’re optimizing for price-to-performance in coding tasks, Haiku 3.5 is a mixed bag: better than many but expensive for what you get relative to the best-in-class.
Hands-on testing: prompts and results#
- Test setup (summary)
- Two primary prompts/tests: a Pixel-Perfect clone prompt (component copy) and a small web app UI task to compare Sunet vs Hau/HaCou outputs.
- Same prompts run against the two models to compare results and speed.
- Test 1: Pixel-Perfect clone (component copy)
- Sunet: produced more usable, structured results with clearer code blocks and narrative guidance.
- Hau/HaCou: struggled to start or render clean outputs; outputs were less consistent and harder to translate into workable code.
- Takeaway: Sunet tends to be stronger for direct coding tasks that require clean structure and stepwise delivery.
- Test 2: One-shot web app with an ingredient history UI
- Prompt involved generating a UI that shows ingredient history changes over time.
- Sunet: produced a more complete artifact (code + architecture notes) but acknowledged gaps needing refinement.
- Hau/HaCou: often failed to render a coherent TSX flow or skipped key pieces, making it less reliable for this kind of task.
- Takeaway: For frontend app generation, Sunet more consistently delivers usable scaffolds; HaCou lags behind in this test.
Prompt engineering and architect-style workflow#
- Architect-focused workflow
- The idea from AER-style practice: start with planning, then define data models, APIs, and TS/React stack steps.
- Example high-level approach (condensed prompt pattern):
- Act as the software architect
- Define data models and API endpoints
- Output a structured plan in Markdown with code blocks and explanations
- Model behavior differences
- Sunet tends to produce clean markdown with distinct sections, code blocks, and step-by-step guidance.
- HaCou tends to mix formats and can require more manual post-processing to extract usable code.
- Practical note
- If the goal is documentation-to-code or vice versa, Sunet’s output tends to align better with developer workflows and tooling.
Practical workflows and takeaways#
- Low-code evaluation workflow
- Make.com (Integromat) is a practical path to automate model testing without writing code.
- Build a workflow that sends repeated prompts, collects outputs, and stores results in a spreadsheet for quick comparison.
- Documentation-ready AI
- For turning a website or doc into well-formatted Markdown with snippets, you can feed a URL and have the AI generate structured docs, checklists, and examples.
- Quick-start prompts for coding
- Use an architect-style prompt to define scope, then drill down with TS/React/Tailwind specifics.
- Example, minimal prompt snippet:
- Act as the software architect. Define data models, API endpoints, and a TypeScript/React plan. Output in Markdown with code blocks for key parts.
- What to test with your team
- Run side-by-side comparisons on your own coding tasks (component scaffolds, API schemas, UI flows) to see which model handles your typical patterns best.
- Track both quality and speed to judge whether the cost aligns with your productivity gains.
Final takeaways#
- Haiku 3.5 is not the top dog on coding benchmarks, but it’s a credible option with strong consistency in some tasks.
- Sunet generally outperforms HaCou for coding-focused prompts and architecture-style workflows, often at a better price-performance point.
- If you’re building workflows that rely on repeated AI tasks, consider low-code automation (Make.com) to test and compare models efficiently.
- Use architecture-first prompts to drive clearer outputs (data models, APIs, and structured steps) before diving into code generation.
Actionable next steps#
- Do a two-model test for your key coding tasks and measure time-to-useful-output versus cost.
- Set up a Make.com workflow to automate iterative prompts and export results for quick side-by-side analysis.
- Experiment with an architecture prompt to see how each model formats its plan; use Sunet’s markdown/code structure to scaffold your project.
Links#
- AI Coding Leaderboards (coding benchmarks and discussion)
- Claude 3.5 Haiku (Anthropic)
- Claude Models (Sonnet and other models)
- Make.com (low-code automation platform)
- Framer Motion (animation/UI tooling)
Transcript
[Music] anthropic just released their new model let's talk about it and how it'll affect you I'm not going to get into like the super nerdy every single Benchmark stuff uh I think that we just care will this make me better at generating code will this make our agents better so we'll go through some of the opinions that people are having online of their first impressions and then I'm going to do a little test so first off AER came out with their updated leaderboard so if you're not familiar with AER you can check out my other videos but it's really good Agent for coding and it looks like clae 35 came in fourth on the coding leaderboard 75% so just behind 35 on it now you're probably wondering why would they release a model it is actually worse than the previous ones it's because it's supposed to be cheaper so you have in the open AI world you have 40 40 mini there's typically kind of like this sidekick uh smaller model that is more cost effective now what's ironic about that is you know this thing's performing well it's better than a lot of the models out there but people are getting a lot of complaint around the price being too high for this type of you know not not leading model so we obviously know that 35 sunit West no doubt there um but yeah so that's what they were saying and then on the llm for reasoning Benchmark which reasoning is Chain of Thought so you can think of like a one preview it actually did really well um you know you still would expect it to not be as good as the flagship really expensive models um and then up next you you can see some of the nice stuff that the Reet folks said you know they said it's it's validating that it's making these really awesome leaps so yeah it's better than the previous one and then underperforms maze Gemini Flash and is 13 times costlier so that's sort of what I was getting at there was a lot of complaints around that and you can see how people are benchmarking it I'll put these links in the in the description so you can check them out on your own but let's do our own little test so what I want to do is I have a bunch of these different prompts and I'll put them in the description below but let's use one for generating a Pixel Perfect clone so I have this prompt and let's just try that out and we'll do it with two different um components or one one different comp one component two different models and then after that we'll actually go through and see see how fast it can build this so the first test we'll do a component copy and then the second test we'll actually try to oneshot this web app that allows you to see different foods and how the ingredients set has changed over time with the fact that it took about nine tries to get it to where this is using sunet the latest one um so I don't really expect it to you know out perform that one um and we'll use the same prompt there so let's just say oh this was actually someone else's test was build me a floating city above you make it vibrant and beautiful you know what we're doing a third test so let's try this we're going to build a floating city above you and see huh okay never mind we're not doing that doesn't even want to so let's try that carbon copy so what I'll do is with sunet I'll grab this and then let's grab something cool looking um how about the v.d okay yeah let's find a decent looking hero section in dark mode yeah let's try this so wow took 22 on v0 dev let's grab this and we going to see what Sun it comes up with for that and below it what we can do is open another one switch it over to ha cou paste the image paste the prompt and see so again on the bottom we have we have Hau and on the top we have the latest so wow I do not have Okay interesting so it wouldn't even start with Hau that's kind of kind of weird but here we have these nice animations let's see if I can't make this bigger so it definitely did a better job as you would expect with sunet but what is strange and it has this kind of nice uh animation they probably used framer I'm guessing yeah framer but with Hau it won't even try um clear sections I'm happy to discuss this is not interesting H just do it anyways let's see huh weird okay well that's a fail on their part very weird I won't even do it h okay uh let's try something else let's uh see how it does what what I like to do when I'm building something uh new like a little web app like I I showed you here uh is let's see for example with this one um what I'll do is I'll actually I have a little process so I'll come in and I'll start by saying hey act as the software architect and this is coming from AER and from really just the best pattern that I found with AI coding is you start with planning if you do want the AER one you can just type in AER chat and you can go to their GitHub and then you can actually get theirs so if I just type in architect time so this is a really good one and in this case what I would typically do is I'd come in and if I want to make something I'll act as the architect and then it'll come up with some ideas it'll give me the data structures all this um let's see if it can go through this same process with Hau so let's see if I give the exact same promp comps how it will operate oops Hi C interesting okay it's going to different route with this you can kind of compare the examples right so on Hau you get this markdown on sunit it breaks it between both markdown uh code Snippets and explanations so again you would expect it to be better but let's keep going so then I'll ask for the markdown actually break out all these steps we need break out all the steps we need to take to implement this step by step your output is a guide for a an editing or a programmer so I want to kind of skip all this fluff because it went straight to the markdown so I just want to see like how it thinks so so it's going for this it's saying okay we need to make this data model build this API okay pretty good let's stick to typescript next 15 react Tailwind update the plan super base for postgress DB One update the plan and we're going off script from the original one right because it went totally sideways on the way not necessarily sideways doesn't mean it's bad that it's going this way but it's definitely different okay so we got a network connection let's try that again so generate the step-by-step instructions okay it's definitely quicker yeah so I can already tell I mean I wouldn't use this I don't I don't have a need for this one um I think it it's just if you have like a ton of documentation like this would probably be better for what I call like a a documentation uh agent so you could ideally give it access in cursor to a URL and say hey go read through this don't miss anything and turn it from a website of documentation into a really well uh formatted markdown file with helpful Snippets and all that um but how about this now build it okay saying it can't make the components for it and on TSX file will it do it okay it's doing something uh omit next router and super base stuff so we can see it in the artifact viewer again very quick which is nice but it didn't render it huh let's look at the code so it's defining the products it's defining the ingredient history if we comp that against ours here okay again bottom one works yeah this is just not nearly not even close to as good um so that those are my first impressions and a bunch of different people's opinions on it uh what I would maybe do next is there's a workflow that you can do with this product called make.com and it would just allow you as a non developer to go in and basically send a bunch of requests oops you could you can just do a drag and drop it's kind of like zappier but you can say hey I want to call this model and do this task and try a writing task and have it run on a loop like 20 times and have it spit into a spreadsheet and I find that that is a really easy way to do evales because I'm not a big python guy I don't have a background in that but makes pretty great for that and so that would be something that I would do but those are my first impressions on it hope this video is helpful and make sure you like the video and subscribe and I'll see you in the next one