I Built Multi Agent AI Coding System This Weekend (Aider, Augment, Grok, Gemini, Docker)
June 9, 2025
I built a weekend multi-agent orchestration stack to connect the dots in the software lifecycle. Here are the key takeaways, what I built, and how you can apply this approach.
Why building a multi-agent orchestrator makes sense#
- There’s a gap between writing code and executing the full SDLC. Orchestrating steps can unlock speed and consistency.
- Turning knowledge into repeatable workflows helps teams scale with agents doing the execution rather than individuals.
Continuation prompts and live workflow#
- Use continuation prompts after finishing a thread to drive the next agent.
- Tie prompts to concrete outputs (e.g., conventional commits and branch names) to keep context and language consistent.
- Scrum-inspired pacing with story points helps the system stay objective and predictable, even when humans aren’t involved in day-to-day writing.
PRDs, story splitting, and the Spider framework#
- Long, cross-cutting PRDs degrade when you scale with agents. Break them up.
- Try spider-style splitting: spike-based discovery to decide how to split work.
- Five splitting techniques (path, interfaces, data, rules, etc.) help turn big ideas into manageable chunks.
- When you have a spike, extract it to shrink the original story and inform the rest of the backlog.
Architecture plus docs: using Midday-style scaffolding#
- Architecture is more than a diagram: document data flows, middleware, and decision points.
- Core components mentioned: Next.js, Supabase, TRPC, Hono, real-time/storage pieces, and how they connect.
- Build a decision tree for data access, security, tokens, and cross-origin concerns to guide agents.
- Create templates and diagrams that employees can reuse on every project; docs become part of the product.
Idea processing pipeline: from idea to PRD to backlog#
- Drop ideas into an ideas directory; a watcher kicks off the pipeline.
- Generate PRDs from templates, then run multiple agents to produce outputs in parallel.
- Use XML-based PRDs for depth and performance, plus smart chunking to fit token budgets.
- Context modules (MCPS) and documentation references ensure agents have the right sources.
Tools, prompts, and templates in use#
- PRD generation prompts map ideas to structured outputs; token budgeting is intentional.
- Smart chunking uses available context modules (e.g., documentation and API refs) to keep outputs relevant.
- Refs folder stores tool descriptions, API calls, and parameter guidance for consistency.
- Runtime uses Docker and Ader to spin up ephemeral agents; scale up/down as needed.
The agent roles and runtime flow#
- Roles: Planner, Builder, Reviewer, Fixer, Product Owner, Critic.
- Stakeholder validation, backlog refinement, capacity planning, and sprint planning are encoded into the flow.
- Guard rails: read-only paths for generated artifacts; prevent unintended writes.
- Execution is fragmented into sprints, with activities tracked and cross-cutting concerns identified early.
Status, future directions, and what’s next#
- Current mode: build-focused, with plans for debug and knowledge-management templates.
- The plan is to generalize this on the AISDLC repo branch and keep building templates for docs, diagrams, and templates.
- Grok-based prompts outperformed older Gemini prompts in the latest pass; expect ongoing prompt tuning.
How to participate or follow along#
- Fork the AISDLC branch and contribute ideas, prompts, or templates.
- Expect future content around build logs, with Q&A sprinkled in as the project evolves.
- VI platform and community involvement are on the roadmap for broader collaboration.
Takeaways and practical tips#
- Treat architecture and documentation as code: codify decisions, flows, and data access models.
- Keep PRDs modular; avoid single massive documents.
- Use spikes to validate options before broadening scope.
- Chain work with continuation prompts to keep the agent workflow cohesive.
- Watch token budgets with smart chunking; chunk aggressively where possible.
- Use ephemeral Docker-backed agents to experiment safely and cheaply.
Links#
- Mountain Goat Software — Five simple but powerful ways to split user stories
- IndyDevDan (video on multiple agents working in parallel)
- Midday project (architecture scaffold example referenced)
- Aider (model execution environment)
- Docker (containerized agent runtime)
- tRPC, Next.js, Supabase (stack references)
- MCP (Model Context Protocol) (context and tool documentation concepts)
Transcript
I rolled my own multi- aent orchestration software this weekend and you're probably thinking, "Wow, that's stupid." Maybe it is. But I learned a lot. And so I want to talk about the things that I've picked up from building a ton of different agents and agentic workflows and then cover what I actually made. So first of all, why do this? It's because there's a disconnect between the steps in the software development life cycle. And there's more to the story than just having your code written. I think a lot of people are obsessed with that one part, but you have other pieces and so why not string them together? So, one thing that I've learned is this pattern of continuation prompts when you are using one of these threads is pretty nice and you can tie that to a hotkey. So, I've done all these processes manually with augment. And I was just curious what could I do if I switched it up. Let's see. That's all live. Great. So, when I finish a thread, just quick tip before we get into the rest of the stuff. When I finish a thread, I will basically use a continuation prompt for the next agent to read for coding this up. After that, write a conventional commit and push to name of branch. And so it already is speaking the way I want it to speak. Yours may not look like this because I have story points assigned. I'm really going after this kind of scrum style thing. Now, I hated Scrum in real life, but because it's very objective and it has very strict rules around it, the LMS tend to like it. And so, I'll just grab this continuation prompt. This looks fine to me. And I will go into here. And then I'll just click this and have it go. So, that's going to keep going. Now, when it comes to other things that I've noticed with making my own is the variation that everybody has with their PRDS. So, if you make a massive one, something that's over 20,000 tokens and it's all over the place and it's crosscutting, then you're going to have a bad time because it's going to just struggle. it will start to deteriorate in terms of quality the larger that it gets. So that comes down to not knowing how to write PRDs to begin with. A lot of people just don't know what a PRD is and haven't done the job of a product manager. So you're trying to pick that up really fast and that's great, but a lot of the time you will just make decisions to break things apart even more and even more. And there's frameworks for that. One of them is called the spider framework. So I tested that out. I didn't think it was that great. So spider five examples to split user stories. So this guy talks about spider where you split user stories using a spike. So a spike is a research activity that a team owner takes to learn more about some backlog item. And spikes can also give team the knowledge they need to split that story. Take YouTube for example. go back in time to when YouTube added automated captioning. The team doing that might have faced a build versus buy decision. Did they use commercially available software to generate the captions or their needs so unique that they need to build something from scratch? The way to settle that would be a spike to test out one or more commercially available captioning products. Extracting a spike makes the original story smaller because some or all of the research research included in the original story is removed. Absolutely essential split stories. So extracting a spike is one of the five splitting techniques you should use. So this is all about splitting. And you wouldn't want to use this if you were doing your PRD like in a one shot. It's just when you actually get some version of that, then you could choose whether or not you want to go and further split. So P is for path. This talks about how if you had a share button, when that pops open, some people might say build a share button. use this thing, but there's 14 buttons in there. There's more to the story. You got to build more stuff, right? And so he ends up finding 16 paths. Then you can split the user stories by interfaces. We don't have to think about this as much, luckily, because of Vzero and these tools that are really good at thinking about it. But you still, if you're building production quality software for a bunch of users, it probably makes sense to have it not just oneshotted to look cookie cutter. And then you can split the user stories by data. And then you can split the user stories by rules. So this is all about splitting. This is just one framework. If you want to, you go to Mountain Goat Software and then it's in their blog. Five simple but powerful ways to split user stories. So, I've really leaned into the agile slcrum stuff because I've recognized that in my past of moving fast and breaking things in startup land that it doesn't work with LLMs because you're missing the talent that each person on that team had, which was to be able to think on their feet and not need the documentation because it's actually a waste of time. But if we're moving to a world where all of the execution of the code is done by these then we are becoming orchestrators and we need to take all of that inherent knowledge that we had in our head and put it into documentation into diagrams. So you can see in this example, I really like this project Midday and I'm using it as a scaffold for my project VI, which is the platform that is for our private network with people from Microsoft and Google and fast growing startups to learn how to use AI. And so I wanted to build out something for them and for our community. And as a part of that, I wanted to follow this architecture. But when people think architecture, they don't realize there's more to the story than just, oh, what are all the individual components or individual pieces of technology that are used? So this is one of the many pieces of the puzzle, but this has a midday dashboard, which is Nex.js. This has superbase, which has different pieces within it, off storage, real time, a database. It has an attachment to react query cache. There's a TRPC client. There's a hono setup so you can work with external thirdparty APIs. So you're using TRPC as a server internally so that you get that nice t type safety. And then for this middleware stack, how does that work? And then there's replicas. I wouldn't I'm not doing that part. I'm doing most of this. But then you have the request flow diagram. So how do requests work? How do you deal with cache or transactions or security and headers or cores? How do you deal with context? How do you deal with verifying a JWT token? Extracting extracting a team ID. All these things come into this diagram. How do you deal with O? Because there's so many players within a system if you have multiple clients and multiple servers. Then how do you deal with that? And your projects will look simpler. I'm doing this because it's going to open up the door for a mobile app and that's what I need. But this same pattern would apply to your projects. How do you deal with data access? If you have different ways of dealing with data, what's the decision tree that the person on the team would inherently have in their head? But now that they're not writing the code and some agents are doing it, you need to get that out and onto paper and in this case onto docs. So this is the data access pattern decision tree. So is O required? Is there real-time updates? How did we deal with that? Where's log in and log out? In this case, you can see I'm putting the majority of it onto TRPC procedures or if it's external, then it's a rest endpoint. And so you can also see the middleware, right? There's all these pieces. So a lot of people think of middleware if they're just coming from next world, middleware and next is a misnomer. I can do a whole video on that, but there's more to the story there. How do you deal with middleware when it comes to security, rate limiting, scope validation, your queries, all that? And then what about for database routing? So, there's all these things that go on. And then I'm just trying to templatize it and then put it into this system so that I can use whatever agent that I want. So, if it's an augment remote agent, I'd love to be able to use it there and have just the outputs of my CLI tool give me the right things to say because I still want to lean on their context engine and have the option to or what about all the components just as a summary. So, you can see everything in here to ramp up on a new project. And then I've put in an FAQ. So this is all for this particular project and it goes on and on. Why not use this, why not use that and all that. So having this really helpful and then other ones this kind of explainer tree. So this is just more depth on how the whole system works. So what does the rest API layer look like? What does the TRPC layer look like? How does schema validation look in this particular project? So this was me analyzing a codebase and then generating all these documents. And then the next natural evolution to that was okay, how could I take this same thing that I just spent all this time on and then systemize it so that I can do it on every project and move quicker. So that led me to building this idea processing pipeline. So I want to have my idea that I write and I spend time doing research. If it's a zero to one project, so something brand new, then it's going to look a little bit different, I can maybe go and do some deep research and figure some stuff out. But then I'll bring all that information in and there's a folder where as soon as I drop it into this ideas directory, it's being watched and that kicks off the pipeline. Now, this is a work in progress, but I just want to speak through it and get some feedback and then you guys can also go and contribute to it if you'd like or just fork it. It's available on a branch right now. As soon as that starts, then it starts to generate the PRD and it can generate the PRD with multiple different outputs. So, there's a good indie dev Dan video that came out recently where he explains this nice thought around the benefit of having multiple agents work on the same thing to get multiple outputs at the same time. And so I'm applying that to all these steps or selectively to these steps rather. So for PRD generation, it will go and write those PRDs based on a template that I have. So you can see that there's different prompts that are in here. So you have the idea to PRD and this is an XML format has all the stuff that you need. So less human readable but better in terms of performance. If you read any of the enthropic or open AAI documentation on agents XML is it. So this will run that and then it will output it and we have a token limit on it. But at the same time I wanted to be able to have it be crazy in depth. So I'm exploring this idea of smart chunking. So smart chunking is the ability for it to use the MCPS that are available that I've selected. So I have context 7 for documentation on APIs and libraries verifying technical details the GitHub MCP. So this can be useful for postf facto updates but also looking in the past to see how things were done before it doing actually the project management as well. So milestones tying the milestones to the issues because we just want everything to be down to a science where it's like there's percentages, there's data values, there's validated inputs, validated outputs so it can know how far it's going. So eventually this could just be a thin wrapper on top of it where you're like go do this thing. Where are we at in the sprint? Where are we at in the stories? And then sequential thinking MCP. So this will be for analyzing the PRD structure and identifying cross cutting functionality and then dividing it logically into the chunks and that's because the quality of the work goes up when you have less tokens and conversely it's vice or vice versa the longer the context the worse the outputs. So I have all the tools that are available to this particular part and then in this refs folder I actually have all the tools and then what they do and their use cases. So you can see I took the documentation from those separate MTPs made sure that it had access to all the different tool function calls and descriptions of each and then the parameters for each. So it's going to reference that and then come back and then it will go make sure it's not crosscutting and then it ends up making these different chunks. Right? So as we move through this then you have the next step which in scrum it is the stakeholder validation. So it validates the chunk for the team which is defined here of the planner, the builder, the reviewer, the fixer, the product owner and the critic. Just quickly since I have that one up, the fixer is the one that's called if there's an issue in production. So that's for later. That's if something's going on. Then we'll log an issue immediately have that thing spin up a VM because this is using Docker and it'll be using Ader, which we have a config for that that allows you to select which model you want for each type of task. And then it can go fix itself and then spin down the docker once it's done. So these things are ephemeral. They spin up, they spin down, they send spin sideways. And yeah, so you just move through this. Then you have stakeholder validation. So are there any ambiguities in these requirements? Can we get any clarification? Can we then figure out who's going to do and validate this chunk of work for issues with cross cutting stuff? because I want to be able to have five of the agents that are coding it go and do all five sprints at once. That's the idea where it's like this thing goes. It's going to take a lot of work to get it right, I'm sure. But for basic stuff I've been testing, it's working. And then you have the actual backlog refinement. So in Scrum World, that means generating the user stories from the validated PRD chunk by extracting the features, creating the stories of acceptance criteria, technical task, and estimating effort. That's probably what you're wondering. It's like why do you have the effort estimates in it? But I think that also is just more data for it to think around. Like I've said in the past, don't ever have time estimates in there. But I'm realizing that it's nice because it's just another milestone for it. So if it says it's going to be like, oh, on day one do this, day two do that, I've done 5 days worth of stuff before hitting the warning basically in one of these individual threads. So it does this for each story list the three to five atomic tasks the libraries and APIs which are plused up by context 7 and then the readonly and writable file paths. So that is following a pattern that will then pass into ader so that if there's things like the types that are autogenerated or the database schema it will mark those as read only so that it can't go. It's basically putting guard rails on it. Then it prioritizes the story so it knows what to do to get ahead of that question where it's here's all the things it thinks you should do. What's first? No, you already know. Go do the thing. And then you get the output of that. It has its own sort of like front matter. So maps to the story type, the status and then a chunk ID that's passed back so it can reference it. And then you get into capacity planning. So, how many agents? How many of these things do we need to spin up as VMs? Not even separate VMs because you can just do multiple shells in one. But I'm still obviously figuring this out. It's not done. Spent about a half day on this. I'm really excited about it. You get all these things that are passed in and then ideally in the future you can see I want to have a running log of how this thing performs so they can see what capacity looks like. what are the diffs between what it's thinking it's going to take versus how long it takes and then error percentage. So that way we have data to then use for future either prompt enhancement or basically having a brain to know how this whole system works. And then you have the sprint planning. So that's actually taking all those different stories and then putting them together, making sure they're not crosscutting all that. And so pretty excited about this. This was the original kind of executive summary slash PRD for it. This is 16K tokens and this was done with a bunch of exploration in Grock. I found Grock actually had a majority of the sauce compared to 03. I used 03 Gemini 25 Pro on the latest one from 4 days ago. Today's June 9th and then Grock and just iterated through it. But yeah, the future for this is I'm going to just keep working on that branch inside of the AISDLC repo and get this part working. So this is one of three modes. So this is build mode and then you'll have debug mode and then you'll have documentation knowledge management mode and that will include all of those templates that I tease at the beginning of the video where I talked about the different charts and stuff like diagrams. That's one of 10 different templated essential documentation pieces. So that's it for today. If you guys found this interesting, if you think this is helpful, then like the video, subscribe, of course, if you want to keep updates on this. I'm going to be doing the kind of content road map for this channel is just straight up build logs like this. And I'll be peppering in the questions and answers as well. I'm already at 20 minutes. And on the main channel, three videos a week, Monday, Wednesday, Friday. And that's actually more build stuff, but it's as I'm doing it. So, it's a little more scripted where it's like on the fly like we're going to build it together whereas this is like postfacto build log. And hope you guys find this helpful. If you want to join us in VI, you should check it out. We are launching this new platform very soon, hopefully this week. and I'll see you the next one.