This AI Agent Plans, Codes, Tests & Integrates Everything (5 Models + Context Engine)
May 30, 2025
Augment arrives as a full-stack AI assistant for codebases: it packs five models, a powerful context engine, and agentic planning that can handle planning, coding, testing, and integration. Parker walks through his real-world setup, from a fresh deployment on a tight budget to CI/CD and observability, with practical workflow notes you can apply.
What Augment is and why it matters#
- A multi-model AI agent suite built for enterprise-scale codebases, with built-in context management.
- Key differentiators:
- Agentic capabilities for planning and execution without constant prompting.
- A context engine that indexes millions of lines of code and uses markdown-style memories.
- Integration-friendly with tools you already use (GitHub, etc.).
- Avoids the “model selector” clutter by benchmarking and running automated agent-driven code generation.
Core features and differentiators#
- Five baked-in models + no need for a separate model picker; the system runs its own benchmarks and optimizes agent behavior.
- Context engine powered by GCP:
- Vector indexing that stays in sync with code changes.
- Memory/knowledge base stored locally (not purely in the cloud) for speed and reliability.
- Mermaid diagrams and structured outputs help you visualize the plan and progress.
- Enhanced prompt workflow:
- Enhance Prompt button backfills context to produce richer prompts.
- Use prompts and a backlog (treated like PRDs) to shape scope before execution.
- Agent modes and workflow:
- Chat for questions and guidance.
- Agent for active planning and code work.
- Remote agents concept for parallelized, role-specific tasks (future potential).
- Tight integration with your dev stack:
- VS Code-friendly, no fork required, background indexing and ongoing context updates.
- Real-time checks and feedback during CI/CD, including SSH keys and container registry considerations.
Workflows and how to use it in practice#
- Start and index a new codebase quickly:
- Open as a workspace in Augment; it begins indexing via cloud4 by default.
- You’ll be prompted to set what you want to know about the project (e.g., database, auth).
- Use cases that matter:
- Read and refactor guidance, feature suggestions, and plan creation for large codebases.
- Treat tasks as formal backlog items (fleets of PRDs) with timestamps and a conventional-commit workflow when done.
- Memory and task management:
- Memories collect notes, decisions, and steps as you work; you can review and prune as needed.
- Time-stamp memories to create a running log of progress and decisions.
- Prompts and task flow:
- A repository (AISDLC prompts) provides structured prompts for ideation, planning, and implementation.
- Example flow: idea → plan → code with agent → update docs → push changes with conventional commits.
- Outputs and follow-ups:
- When a task is complete, Augment can generate a continuation prompt for the next steps and push a conventional commit to prod.
- Outputs are rich but still require human review to ensure correctness and alignment with project goals.
Practical setup notes and roadmap#
- Architecture you can realistically deploy on a small VPS:
- A quintessential $12 VPS setup to deploy and run end-to-end with CI/CD and observability.
- Cloud4 is used for cost efficiency; the $50/month plan plus per-thread usage is competitive, with memory and context handling baked in.
- Codebase examples Parker mentions:
- A FastAPI + Next.js project with multiple bots (e.g., Discord bots) is a typical use case; migration from older stacks (Astro, TanStack, Express) is feasible in days, not weeks.
- Tools and integrations:
- Direct GitHub integration for code syncing and automation.
- SSH key handling for container registries and CI/CD pipelines.
- Reminders to not rely solely on outputs; the engineer’s “muscles” of reading and guiding outputs are still essential.
- Remote agents concept:
- The idea is to run multiple role-based agents in parallel, with read-only and write access controls, to cover frontend design, backend logic, etc. This is promising but still maturing in Parker’s use.
Pros, caveats, and best practices#
- Pros:
- Great for large, complex codebases; effectively acts as a project manager that can also implement.
- Context engine plus memory system dramatically reduces context loss as code evolves.
- Practical for developers who want to ramp up quickly with CI/CD and observability.
- Caveats:
- Memories aren’t perfect out of the box—verify and prune to keep them accurate.
- It’s not a black box replacement for human judgment; you still read outputs and guide the process.
- No dedicated model selector UI by design; relies on built-in benchmarking and agent-driven optimization.
- Best practices:
- Treat the memory bank as the workflow backbone; timestamp and prune as you go.
- Use the enhanced prompt feature to keep prompts aligned with evolving context.
- Start with a focused, smaller scope and gradually scale to more of your codebase.
Getting started tips#
- In VS Code, open Augment on the right; let it index and then define your questions (e.g., database architecture, auth flow).
- Use agent mode for planning and execution; switch to chat for questions or refactoring advice.
- Build a backlog of tasks as markdown docs (PRD-like) and let the agent follow the memory-driven plan to completion.
- Pair with your existing tooling (GitHub, SSH keys, container registries) to streamline CI/CD.
Final takeaways#
- Augment can dramatically shorten the loop from planning to production for sizable codebases, especially when you leverage its context engine and memories.
- The combination of agent-based planning, code-aware context, and integrated workflows can be a game changer for teams looking to automate large parts of development and maintenance.
- You still need to read outputs, guide decisions, and maintain quality, but the tool handles the heavy lifting of planning and execution scaffolding.
Links#
- Augment (official overview and docs)
- Augment documentation on model benchmarking and context engine
- GitHub for setup and CI/CD integration with SSH keys and containers
Transcript
Everybody's arguing about windsurf versus cursor versus client and rue code and then which model you use for what? How do I gain context? There's a company that kind of solves all of them for you and has five models baked in and a context engine that works on millions of lines of code. Let's talk about it. We're talking about Augment and this is a tool that someone in our community talked about maybe a month or two ago and I was really skeptical because it can't be that good but it turns out it is. And I want to explain how I went from using an entirely different architecture two days ago to having it fully deployed, fully tested with CI/CD, observability, monitoring, and all of that on your quintessential $12 VPS. So, a couple of things we'll cover are the features and then some of my workflows and differentiators. So when it comes to features, you have the typical thing over on the right hand side. Now I've stopped using cursor because I've noticed since their background agents came out in beta that it's been very buggy and I wanted to give this a shot. So I'm running it in VS Code. Now VS Code there's no fork. It just works. There's no performance issues. And so I just put augment over on the right side that kind of mimics the same workflow that I was using in cursor. Couple of the core differentiators are first of all you have your threads so you can see them all in here. You see some of these chat previews and then you have different ways you can work with it. So chat is what it sounds like. Agent is what I'm using right now. Remote agent we'll get to. But the core difference is they actually have agentic capabilities for planning and for coming up with how you're going to go and execute on that work without you even prompting around it. So that's the magic that I've found with it and I want to pull this doc up. So to differentiate when to use what because I found myself not using chat as much. But when I come in here, it's basically like you can use chat for asking questions, getting advice on how to refactor it, adding new features to selected lines of code. Their next edit, the biggest difference between cursor and augment is next edit is command and semicolon versus the tab thing. But I think tab is actually a symptom of something worse, which is if you're making one variable change or function change or a string change across a lot of areas, it already missed the boat by not just reducing the complexity of the code. So, Augman is positioning itself more as the partner for big code bases in the enterprise, which means that it's almost overpowered for individual people like myself. And I find that really helpful. So in my codebase I have a fast API and just a Nex.js and then all these Discord bots that are running in it. So not too complex. Before that I had an architecture that included Astro and Tanstack start and an express server. So switching from that to what I'm doing now took two days to deploy. And these are the differences you can do here. when you start a new codebase and let's say open up midday. We'll open it as a workspace. We'll make it a little bigger for the people and then let's pop this open. So, command L pops that open. I always just open augment and immediately it's going to start indexing the codebase. Now, by default, this uses cloud 4. And that's great because from a price perspective, if I go and I use Claude Code and I don't have the max plan, which is I think there's a $200 plan, which is great, but that's still four times more than using Augment. Augment's 50 bucks a month and I've been hammering it and I might reach the 600 request threshold, but it's for an entire thread. When this finishes indexing, you'll see that it'll actually ask me a couple of questions about what is it that I want to know about it. So, it gives me the summary on it. It's now baked into the context engine. And so, it's also really good at ramping you up on code bases. So, if I selected one of these, I can ask what's the database and the authentication system. It gives me the details on that. Right. Another great feature about it is this enhance prompt button. Now, I feel like it's a meme because I'm gonna break it, but every time I write a prompt, I just click that button and then it's backfilling the context that it already has in its engine to make the prompt a lot better. I've talked about this a lot on this channel, but if you use a tool like Taskmaster, it's great for simpler things, but as soon as you're getting complex and you have a bunch of relative paths and snippets and file names and to-dos and all these things like baked into basically a PRD, then you're going to degrade the quality of the task items, which isn't good. And so I find myself basically coming up with a plan in agent mode with auto off and then once I feel good about the plan which is following this folder pattern that I teach of having and then having a backlog that's full of PRDS and then doing which is what I'm currently doing. This is a summary of what I'm currently doing the actual what I'm doing and I treat them like migration files. Alo has like time stamps in there. But then when I'm done with one, it will follow the memories that I've get given it. So if you click on to memories, you can see that you got this file and I can come and basically comb through here to see if there's anything that I don't want to have in here. But it's really taking all the best parts from memory bank on Klein and then just throwing it into essentially a markdown file. That's what it is. It's markdown. But what powers this the other side of the coin and this is something that I was like prototyping and trying to build myself, but I'm happy they just built it so I don't have to is they use GCP on their stack. So their founder basically was like, hey, we need to have a context engine that's paired with these markdown files such that it's kind of like a state machine where it's, okay, as soon as something in the codebase changes, now it's running a vector index and it's indexing everything in the codebase. And so if you just use client, one, it's going to be more expensive, and then two, it doesn't have this context engine, which is the sauce behind all of this. Now, I've seen complaints about the memories not being great. It's not listening. Just make sure after a coding session that you come through and make sure like it's accurate to what you want. So, I found it to be pretty much flawless. And then in my memories for my workflow, what I'm doing is I'm just timestamping them. Like I said, then when we get to the end of the list, it's going to move it into done and then write a conventional commit and push it up and update the docs. So, it's using the memories as the workflow to help you churn through these pieces of work. And so, you can see here I have all these that are done. And about halfway through doing this migration, I realized, oh, I should like time stamp these because that's kind of it. And then you have the running log of everything that you've done. Now, some of the things that I'm pairing augment with are the different prompts that I like to use. And so, I have a repo that's called AISDLC. Totally free. And if you just go to the repo of Vibe withAI, you can see I finally got it all deployed, right? Never done that on this stack before. But if you go to github.com/vi and go to s or sorry aisdlc and then you can click into prompts. Let's make it bigger for the people. Then you can see all the different prompts. So this tasks prompt you just go like sequentially. So if you're going 0 to one, you'd start at the idea. You can skip ahead to the prompt. But I'm just using a chat window in augment because it has context of the whole project to help me kind of shape that idea up and then ultimately come up with the strategy and then I just start going in agent mode. So in this case, let's open up an example. If I got to the end of this and I want to just close this thread out, it's so funny having this so big, but I know I'm going to get complaints if I don't make it nice and big. But if I scooch this out, you can see I had a documentation update for it. I have the next priority of what we need to do. I can see that all these things were completed. I can see that the critical blocker was resolved. It is funny seeing how enthusiastic Claude is about all this stuff, but it just does all these things for me and it's really nice. And I say update our task list. And then I actually bound this next command which is cool. Write a continuation prompt for the next LLM agent to read for coding this up. After that, write a conventional commit and push to prod. That's what I was doing at first, but then Augment just picked up on it and then started adding timestamps to it. So, I've noticed it thinks on its own, which is really nice cuz like I've just been using keyboard expanders and the markdown files. I'll still use the markdown files, but having it know what we want and then plusing it up is pretty dope. Other things that I really like are the fact that it has direct integrations with tools. So, if you're like me and you're allergic to MCPs because they're you start sneezing because they're so inconsistent, then this solves that problem. So, I don't use any of these tools, but I use GitHub and I use Superbase. So having the GitHub in there is really nice. My use case for it was I needed to have an SSH key so that I can do the container registry and connect it to my VPS. And it noticed an error when I was doing my CI/CD pipeline where there was an off issue. So it's oh you don't have the right pub key for this. Let's go check. Okay, cool. Do you want this? Great. Awesome. Off to the races. Now it's connected and now it just doesn't miss. can also just check the status of them and all that as you would expect. And other things you can change the startup scripts. There's user guidelines. I haven't touched them. And then there's context. I can't say enough how good the context thing is. Like people are pairing this with other tools. I think if you're an avid cursor user, like I was for since it came out, I was hesitant cuz I was like, "Oh no, I'm going to miss the tabby thing." But I didn't edit many files. I was basically just reading. So you still want to read the code. You don't want to vive code. You want to make sure the outputs are good and then guide the absolute crap out of them with markdown files and research. So augment's not going to do the whole thing for you. There's still the job to be done, which is make sure that you're building the right thing. Make sure you're selecting the right tools. Make sure you're actually reading the outputs. Otherwise, those muscles that you've gained from learning how to code and being an engineer are going to atrophy. But I find it super duper helpful that it has all these options. So, other things I take some notes on this. Yeah, it uses all these different models for you. There's no model selector. And people would be like, "Wait, why can't I have the model selector?" And it's actually, they wrote about this in a blog, and at first I was hesitant, but now I get it. It's a bad design pattern because it's just another variable for you to try to figure out on your own, but they run their own benchmarks and they've have agents that are writing 90% of the code for the agent. They have a really good talk on this actually. They benchmark what is best. They do use Cloud 4 for some of the heavy lifting. They have a deep partnership there which means reduced costs. I made notes here. Let me just see um how the agent actually works cuz I want to sound smart and actually know what I'm talking about. Context engine power. Yeah, it's powered by the context engine. Powerful LM architecture. It's built for high complexity, high stakes environments. So for me and for smaller teams, it's just like I'm saying like ridiculous. And then acts as a project manager as well. Like I was mentioning with the timestamping thing. I just came up with that. But yeah, it breaks them down into functional plans. It can also execute on them and then goes systematically. I also pair those with prompts and then it will enthusiastically let you know what's going on. The memory system is stored not in the cloud. So that's what you saw before. And all agent requests have the accumulated knowledge base. So it's kind of like memory bank but on Turbo. It's got built-in mermaid too which is really helpful. And I teased this idea of the remote agents, but I have not built them out to where I want them to be to like confidently talk about it. But it's basically like in remote agents, you can set up different branches for each one like cloud code. But then I want to actually have different scripts that tie to different roles in the codebase. So, I want to have one that's like the front-end kind of designer and it'll be trained on the Vzero style prompts and then have a bunch of examples of those. So, I can actually write a bash script that then goes and gets and loads in context. It has its own kind of memory bank for that and has read only for certain things and then write access to other things. So, I think of it like your org tree. If you had a flat or tree with six people on the team, I want a remote agent for each one of these and then have them running in parallel. And this would be nice where if I wanted work done while I'm away, then I can just kick these off and let it go. Again, I don't have it working the way I want yet. I haven't set it up, but that'll be in a future video. When it comes to the actual commands, they have a bunch of different commands in here. So, you can do slash commands and at@mention, so you can find things in your codebase, explain, fix, or test. You have next ededit, which is what I was talking about, but I really haven't had to touch that much because it's just basing it on a lot of code that was written for me. But there's ways that you can edit this. So, if you want to see like the changes and they're highlighted, that's really nice. And then here's some of the stuff for the code completions. So, I highly recommend using it. I've been using it the last week and a half, pretty much every single day. I'm maybe a third of the way through the $50 limit that they give you per month dollar for dollar. It just blows everything out of the water. And you get access to the context engine which is running on GCP. You get access to all the different stuff. The remote agents which I have run them and they don't crash which is a major upgrade from cursor. You can go and check them out. If you learned anything in this video, make sure you like the video and if you want to see more videos like this, make sure you subscribe. If you're curious in learning more about the community, the AI, and all the stuff that I built and I was referencing in this video, you can check that out in the description. And I'll see you in the next one.