CEO of Augment: Building a $1BN AI Coding Empire (Ex Google Research Scientist)
June 27, 2025
In this conversation, Guy Gerari, co-founder of Augment and former Google research scientist, shares how Augment is accelerating software development with AI-powered tools, plus his views on planning, testing, and the future of AI-driven coding.
Guy’s Background and Augment’s Mission#
- Ex-Google research scientist focused on vision tasks, optimization, and scaling with generative models.
- Physics background (Wiseman Institute, Stanford, IAS) and a career centered on pushing AI research into real products.
- Co-founded Augment to bring AI-assisted coding to developers via chat, code completions, guided edits, and a suite of agents (auto, background, remote).
- Core belief: rapid experimentation and scaling AI in coding can redefine how software is built.
Augment: The Platform and How It Fits Into a Dev Workflow#
- Core tools: chat, code completions, guided edits, and a family of agents (auto and background) to accelerate tasks.
- Agents help with both inner-loop coding and outer-loop software development lifecycle (planning, reviews, triage, deployments).
- Practice tip: humans still supervise code; CI, unit tests, and code reviews remain essential even with AI agents.
Prompting Best Practices and Memory#
- Context and precise instructions matter: more detailed, unambiguous prompts yield better results.
- Calibrate task size: too big a task can derail the model; break work into bite-sized pieces for reliability.
- Memory feature: agents learn from mistakes over time, reducing the need for constant guideline updates.
- If a recurring issue happens, update the system prompt to address it globally rather than micromanaging prompts.
Actionable takeaways:
- Start tasks by loading relevant code context and asking the agent to summarize before proceeding.
- Break problems into smaller chunks and iterate; reserve larger, end-to-end tasks for later.
- Leverage memory to minimize repetitive prompt tuning.
Guideline Evolution: Memories and System Prompts#
- Guidelines are used less over time as memories improve accuracy.
- When you notice a widespread problem, updating the system prompt is often more effective than iterating individual prompts.
- The approach emphasizes building reliable, repeatable behavior rather than hand-crafting prompts for every case.
Actionable takeaways:
- Rely on memory to reduce friction; use system prompts to codify broad fixes for recurring issues.
- Periodically audit and refresh system prompts to reflect the current best practices and learnings.
Planning vs. Coding with Augment#
- Start by familiarizing the agent with the codebase to load relevant context.
- Prefer implementing a solution directly with the agent to learn its approach; you’ll often get valuable implementation insights.
- Then switch to task lists to organize work into manageable PRs.
- Work in parallel with separate workspaces; map tasks to PR boundaries to avoid conflicts.
Actionable takeaways:
- Use the agent to implement a baseline, then organize the remainder with a task-list workflow.
- Maintain parallel workspaces for multiple PRs to scale throughput without constant context switching.
Day-to-Day at Google and the Research Mindset#
- Google work centered on scientific discovery and rigorous experimentation, with metrics like loss and evaluation results guiding progress.
- Big question: how to reason and whether we can train models to improve reasoning by scaling up.
- The “shut up and calculate” mindset from physics guided practical experimentation and rapid iteration.
Takeaway:
- In fast-moving AI R&D, focus on measurable experiments and scale bets that yield actionable insights, even if the problem is abstract.
Shipping, Testing, and Trust in AI-Powered Coding#
- Augment employs traditional software development quality controls: code reviews, CI, unit tests.
- AI-generated code is treated like human-written code—supervised, reviewed, and tested.
- The AI tools accelerate work, but do not eliminate the need for human oversight and robust workflows.
Takeaway:
- Build AI-assisted workflows that integrate with existing engineering practices rather than replacing them.
Agent Organization: Tasks, Sessions, and Workspaces#
- Agents are organized by task and PR; multiple agents can run in parallel across separate workspaces.
- History compression allows continuing work across sessions, but you can keep agents focused on discrete PRs.
- Remote agents can handle some tasks autonomously, while more complex or sensitive work remains supervised in IDEs.
Takeaway:
- Structure agent usage around PR boundaries and code ownership; use parallel workspaces to maximize throughput.
Outer Loop and the Future of AI in Coding#
- The big frontier is the software development lifecycle outside the IDE: code reviews, ticket management, production alerts triage.
- Remote agents are being positioned as both a feature and a platform to automate the outer loop tasks.
- The vision is to extend AI-driven coding beyond the editor into end-to-end lifecycle automation.
Takeaway:
- Expect AI to handle more outer-loop tasks; design tooling to act as a platform for automating lifecycle processes, not just code generation.
Three-Year Outlook: AI Coding and Multi-Agent Systems#
- Strongly bullish: engineers and companies will increasingly rely on AI to automate substantial portions of engineering work.
- Multi-agent systems are expected to be the next big leap, enabling coordinated AI-driven workflows.
- Progress is rapid; the tooling and capabilities will outpace today’s expectations, opening up more creative possibilities.
Takeaway:
- Stay lean and experiment aggressively; plan for a future where AI robots many of your engineering tasks but require solid governance, reviews, and safety nets.
Final Takeaways#
- AI is a practical tool that requires new workflows, not a magic replacement for engineers.
- The fastest path to value is to start implementing with agents, then organize and scale with task lists and proper supervision.
- The next few quarters will push AI further into the outer loop of software development, not just the code-writing inner loop.
Links#
- Augment Code (AI-powered platform for software development)
- Modern CTO Podcast (reference for the "AI as a tool" mindset)
- Augment Documentation (planning and orchestration in workflows)
Transcript
Guy Gerari, welcome to the show, man. I am very excited to have you on this. I know your tool very well, and my friends and I actually we joke about how we wish AI companies had like day passes for just how fast everything moves. But I feel like with Augment, it's a little bit different. They're not just gassing you up because you're here. Um, but just because of how fast everything moves and um, haven't had a problem with it. Also kind of terrified by the idea of having dozens of agents blow up my codebase. So everything that you're involved in from completions to chat to auto agents, the background agents, it's just fascinating to me. So thank you for doing this and it's going to be very very educational for me. Yeah, thanks for having me. uh excited to talk about all that stuff and glad you're enjoying the tool. Yeah. So, everybody starts with an introduction. So, let's just rip it real quick and um after I do it, let me know if I missed anything. So, Guy Gerari, co-founder of Augment, an AI powered platform for software development that provides services for chat, code completions, and guided edits. Former research scientist who was part of pushing the company's move to using generative language models. master's in PhD graduate in physics from the Wiseman Institute of Science, a postdoal fellow at Stanford University, and a member of Princeton's Institute for Advanced Study, also known as the IAS. It's a home to early giants like Oenheimer and Einstein. Maybe you've heard of them, but that's over a decade of study on topics spanning string theory, high energy theoretical physics, and of course, machine learning. guy's been cited in over 15,000 research pieces according to Google Scholar. Ended up going all in on machine learning and then ended up pitching Google to bring a whole group of researchers to help push the company forward. That is a crazy background. Did I miss anything? That was pretty comprehensive. Thank you. Yeah. Okay, cool. I was doing the research and I'm like, "Oh man, I'm gonna have to like recite this because there's a lot in there." So cool. So, the last thing and then we'll get into it. So, I founded an AI community in April to surround myself with a bunch of motivated engineers who are obsessed with all this coding stuff just as much as I am. And so, as a part of that, for them being day ones, I offer them the ability to ask questions to people. And so, the first one is from Brian. He says, "What characteristics of a prompt do you find work the best that no one is talking about?" Yeah. So, with prompting, it's it's interesting how deceptively simple it is to have a prompt box where you can just type anything, but the results vary wildly depending on how much goes in there. So, I'd say certainly the more context the better. I think one thing that people tend to underestimate is how much the models benefit from very detailed and correct and unambiguous instructions on what you want it to do. Really, the more you can tell it and the more you can tell it how you want things done, the better the better it gets. Um, I think the other side of it is it this takes time to kind of calibrate the size of task you want to give the model to chew on. um if it's too big it will tend to go off the rails and I think the challenge not not just go off the rails but also produce a lot of artifacts that are then pretty tedious to review. So there's a bit of a balance and taking a task and cutting it up to bite-sized pieces to hand off to the model so that it can do a good job and also you don't go crazy supervising and like reading too much code. Yeah, I feel that it's this tough balance. That's great advice and especially coming from you. So the next one is from Josh and he said I'd love to hear guys take on he think on how he thinks about the evolution of the augment guidelines and so we see this pattern with every tool. How do you set yours up? Are you still using them? How do you think it evolves in a codebase? Yeah, that's a good question. So I I actually use them less and less over time. So my guidelines are actually currently extremely short. Uh one reason for that is we have the memories feature which learns from the agents mistakes. It gets updated automatically. I mean, I found that that feature has been able to capture I mean, when it makes an error once or twice, that error kind of tends to disappear or or not happen again because of memories. And so, I haven't had to guide it with user guidelines as much. I think the other advantage is when I see a problem, if I sense it's not just a problem for me, but a general problem, I actually go and update our system prompt. That's best way to do it. So, I have a bit of an unfair Oh my gosh. situation there. Yeah. Yes. Cheat codes. That answered a question for later then because I I figured that would be it, right? Is you hit a same wall a couple times, you're like, let me just make a quick commit here. I got this one. Yeah. Okay. And uh the last one is from Mike from Australia. He says, "What are the best practices for planning versus writing code with augment? Are you starting on a whiteboard and then you go to linear and then you go to augment? This is kind of an all-enccapsulating open-ended question. Yeah, that's a great question. So, I the way I like to start is I like to start brainstorming with the agent. Typically, I will even start with telling it like familiarize yourself with this part of the codebase and summarize it for me before we even start talking about what the task is just so it has like the relevant context loaded up. Um, and then I will often skip the specking part and we'll just go straight for implementation. Not because it gets it right, but because I find it very illuminating to see its take on the implementation. So I prefer to read the code than to read a spec it writes for me. I find that more useful. So that's kind of my planning step with it is let's just try to implement something and see how it goes. Maybe I'll chuck it after or maybe I'll adopt it and iterate on it depending on how that is. Um but then once I have something basic that works and the design seems right, I will jump into the task list which is a feature we uh released recently. Previously before we had the task list, I would just do it in a markdown file and I will list start listing off all the tasks I wanted to do and then click play. Maybe I'll click play on a single task. Maybe I'll click the global play so it starts crunching through them and usually the rest of the work for me and that could span like multiple PRs will happen in the task list. I find it like a very useful way to organize the work. Yeah, it works. It's a great addition. Our uh Discord channel is popping off. We can tell when there's new releases from augment is just like you just see like a block. But anyways, let's jump into some some background. So in your four years, is it four years at Google? I want to get that right. Yeah. You worked on vision tasks, optimization algorithms, and ultimately you tested your like scale theory with generative language models. Could you explain from a high level what the daytoday looked like for a normie? Like just break it down where you went to work and that's what you were doing every day. What did what did a day in the life look like when you're researching researcher at Google? Yeah. Yes. So right so the my my goal there was very different right now I'm building a product but there it was um doing research and for the purpose of scientific discovery really not for the purpose even directly of improving products like there was always this hope that we will discover things that will end up making their way into products but the goal was really scientific um scientific research so the the dayto-day involved D I would say yeah it's interesting so I would say sometimes there were experiments running and we got results and then we would like need to analyze them and see what's wrong especially if when like when you're training models there's just a lot of you know is the loss going down quickly enough and if not is it working down yes I mean that's like just classic uh machine learning research like you try to get the loss to go down and sometimes the loss goes down but you test it in some evaluation and then evaluation is not good. So uh research just like um bread and butter research I would say and then the there was a lot of discussion on what should we be researching I think that's kind of one of the hardest things in research and this was true in physics it's true in machine learning it's always true they're like with with the ultimate freedom to work on whatever you want you then have to somehow choose what problems to work on which is it's certainly difficult when a field is moving so fast and you're trying to understand what what is the best thing to work on. So we at some point ended up narrowing in on well we want to understand reasoning uh and then we spent way too much time debating what that even means and what is reasoning uh and how do we research reasoning like super meta yes reasoning is clearly maybe the key part clearly an important part on the road to AGI but what does that even mean and how can we tell if a model his reasoning or not. But I think eventually we kind of adopted the uh we kind of adopted this uh fineman approach of shut up and calculate from physics. So we were like let's just shut up and legend and do experiments. Yes. And and our t I mean our approach to that was okay let's try to see if we just scale things up will the models get better at reasoning. Um, that was a pretty clarifying moment for me and for the group. It was like, okay, we don't know if this is going to work or not, but we at least arrived at a bet we're going to make. That's like half the battle is understanding we're going to make a bet. We're going to do an experiment, and whatever the result of that experiment is, we're going to learn something. So, a lot a lot of work is like figuring out what experiment you want to do. Um, and then of course there's work on like, okay, we don't have enough TPUs, we need to get more TPUs, we need to convince someone to give us TPUs. Oh man. And that's played out quite nicely for them. But yeah, I I asked all that because I know it kind of blossomed into this aha moment and that led you to ultimately starting this awesome company that a ton of developers are excited about. And um so it it happened around when GPT3 came out, I'm correct. Yes. Like whoa, something happened. Yes, certainly. Before that, we were working on um vision tasks and optimization algorithms. And when GPT3 came out, it was clear that this was a fundamental shift in how we're going to be doing machine learning. Especially, I mean, especially because with language models, you get a whole new interface into the models, which vision models didn't have. You get to talk to them using language. And so, that opens up many new experiments that are doable. Uh fshot prompting was another like big thing because previously if you wanted to get a model to do anything you would have to go collect a data set train it evaluate it uh which all took a lot of time with few shot prompting at least for simple things you could show the model a few examples of what you wanted and it would match the pattern without any additional training without any fine-tuning. Uh and so we kind of pivoted the team to working on language models. Uh so that was certainly very influential for us. Uh and then in the beginning we worked more on evaluation tasks because we saw evaluation benchmarks just getting saturated by all these new uh excellent models. Um and then we landed on yeah let's let's try to understand reasoning and within that we worked mostly on math and science questions. Uh but the other big reasoning task is code which is what attracted me to working on uh yeah AI for code after. Yeah. So then you make the leap and I listened to a bunch of your talks and I was like, "Oh man, he obviously gets it." I mean, I am a just a young Padawan in comparison to the the Master Jedi here. So I had a bunch of questions specific to when you're on the Modern CTO podcast. You mentioned something that really resonated, which is that AI is a tool that you need to learn how to get the hang of. And there's really it's this, it's not really a metaarning skill. its own set of skills of how do I know when to grab the wheel and let it take the wheel. What do you think those like how do you tell developers and teach developers on your team and others how to flatten that steep learning curve? Yeah. So, it's it's really like a list of little tricks and techniques. None of it is too deep. I think it starts with just an appreciation that yes, this is a tool. It's a tool like any other. Again, that prompt box is deceptive in that you can get garbage out of it and you can get magic out of it and it like it's not always intuitive why you got the magic. You just kind of need to get the hang of it. I've seen that I've seen work repeatedly for this is find someone who is a power user of agents. Someone who's like usually it's pretty clear that oh suddenly their productivity has skyrocketed somehow uh and it's because they're using agents but it's not clear exactly what happened. And just watch them work. If they let you like watch over their shoulder for like 10 20 minutes you'll kind of see you'll see oh the they're writing like pretty detailed prompts. Oh, they're using this trick for like quickly reading the code and like iterating and this is how they're giving the agent feedback on what's wrong and this is when they decided to stop it and this is the bite-sized chunks they're giving the model and this is how they're organizing their work and like is there a spec, is there not a spec, all that stuff. Again, it's like it's like a list of tricks that you see you you get the hang of pretty quickly if you see someone doing it. That's the best hack uh I've seen. The other way is just like grind through it. I've seen people like augment starting to use agents and have bad experience after bad experience for like a week or two and then suddenly they're like, "Oh, I just had my first magical agent experience and then they're hooked. It just takes longer that way." Yeah, totally. And that kind of just takes the whole traditional software development life cycle and compresses it. I think everyone's trying to figure out how do we move quicker? How does the product team bring these teammates in and make them reliable? Which brings up this this kind of like a trust issue, which is when do you put enough safeguards in place? When do you not put safeguards in place? So, I'm curious at Augment with access to all of these tools, how are you guys thinking about shipping? Are you including background agents for certain tasks? Is there a testing framework? Just kind of in general, how's the factory work? So, from what I've seen so far, I'd say the the normal best practices of software development apply here as well. We haven't seen new trust issues come up solely because the code was written by agents. It's still all supervised by humans. We still do code review. We have CI. We have unit tests. all all that stuff still exists. Um, and so yeah, I would say we look like a pretty normal software development shop except things go faster because every other conversation is like, "Oh, we want to do this. Okay, I'll get an agent on it." That's like a normal a normal conversation. But at the end of the day, all that code still gets reviewed and so on and still gets tested. Uh so I found so I guess I found that the all the battle tested practices that humans have come up with for developing software they still apply for agents and they're still like as useful and I'm sure we'll develop new ones and there are interesting ways we can talk about about how this is all going to evolve but for now the best practices they apply and they work. Yeah, that's kind of what I've been doing with my client projects and just different things I'm working on is it's still the same. It's just it's compressed in each step. And that brings me to when you think about both the auto agent, the background agent. Do you think about it as separate concerns for each one where you're mapping them to specific areas of context, specific areas of environment? This one's for Python on this area of the codebase. This one's for CI/CD optimization, or this one's for looping back, reading through hotel and making fixes. Do you kind of think it of it mapping to JDs like jobber uh descriptions or is it just more okay each time we get in it's a vanilla background agent? How do you think about it? So personally here different people use these things in different ways. For for me personally I I split agents up by task. Um a task typically results in a PR. Sometimes I'll keep going in the same session and do a few a few PRs and stack them. Um, and at this point our history compression is is good enough that I can just actually keep going in the same session. But typically I will have I guess a a session for like a good size a good size PR. Uh, and then if I think the agent can kind of do it in one shot, if it's fairly simple, I'll probably launch a remote agents for it and just look at the PR that comes out. Uh, but for for the non-trivial PRs, I'll want to supervise it more and then I'll typically do it in the ID uh, and work with it. But I like splitting things up by PRs. And then what I do is I actually have um several copies of the repository on the machine and I have an ID window open to each and so usually I will have an agent working on a PR in between like three and five of those workspaces in parallel and then I will kind of jump around between them but the work that they're doing is independent of each other so I don't have to synchronize them and I don't merge confidence. Got it. So it's not work trees. It's simpler than that. It's hey, I understand that this is going to be in a different area of the codebase. Got it. Could be even. Yeah, often it's the same area. Just takes some care to not have too many merge conflicts. But if it's like a big feature I'm working on, I have a bunch of tasks in it. I'm going to have a few agents working on them. Got it. And then when you think about just the agents in general, the ones you guys are working on, background agents, remote, what's your best guess on kind of the vision for these over the next couple quarters? Yeah. So I think what we're what we've seen as we've we've made a lot of progress on the inner loop of software development. This is what uh the tool is very good at. Um we are now looking at the outer loop. So the rest of the software development life cycle um and so that means many tasks that are happening outside of the ID. So this could be code reviews, this could be ticket management, this could be alerts from production that come in and you have to triage them and address them. All of that stuff I think is very ripe for AI agents to to tackle. Um and so we are thinking of our remote agents both as a feature where you can yourself paralyze tasks in the background but also as a platform uh on which we're going to build new features that are going to address the the rest of the software development life cycle the whole outer loop. And I think that's going to be the theme for the next yeah couple of quarters is probably right. Oh man that's huge. That's such a big deal. Yeah, because I was wondering how does Augment get out of the sidebar and into everything else. Yeah, that's that's good to hear. It's for remote agents. Exactly. Yeah. So, we're wrapping up the interview here, guy, and I'm wondering where do you think this is for all the founders, technical people that watch this, where do you think that AI coding in general will be in three years? I know it's going to be an educated guess, but you're much more educated than most. I mean, I have to say I've I've kind of changed my mind on this. I'm I'm a lot more bullish on how much we'll be able to automate. Um I think in 3 years we are going to see large very successful companies come up that where engineering is mostly or fully automated by AI is my uh is my sense just I'm just extrapolating from the rate of progress on the models and the tooling that we're seeing. It's been it's been incredible. Even if you track you track model releases and you see how far things come in terms of long context behavior in terms of tool calling the next big thing is going to be very likely multi- aent systems that's going to I mean I expect that's the next like big jump in capabilities are going to come up and I expect to see that in the next few months. Yeah. So just extrapolating from that rate of progress, if that's going to continue or even be like a bit slower for the next three years, we're going to see a lot of engineering get automated. Um, and so I I'd expect we'll be able to build I think the effect of that is going to be we're going to be able to build a lot more cool stuff than we can today. Um, so I think overall that so and I I actually I think three years is a very long time for that to happen. I actually expect it to happen faster. That's such an exciting future. I I think there's a lot of people get all doom and gloom about it, but the reality is we're going to push the edge of what's possible. The better the tools get, the more stuff you can make, the more creative output. That's just that's really really good insight. Um, well, perfect. Well, Guy, I want to thank you again for doing this. That was fascinating just to learn everything that you're up to at Augment and I hope to see you again and I want to wish you the best of luck because it's very exciting stuff. Thank you so much. This was great. Thanks a lot for the insightful questions and for having me on. Of course.