Parker Rex
All videos
Parker Rex

Microsoft KILLED Every Prompt Tool Company With This ONE GitHub Feature

June 3, 2025

Watch on YouTube@parkerrex

Microsoft is integrating AI model experimentation directly into GitHub, throwing a big lever for enterprise AI. Here’s what you need to know and how to use it to move from endless prompts to real, testable results.

What GitHub Models is#

  • A workspace inside a private GitHub repo that lets you find and experiment with AI models for free.
  • You get a versioned, playground-like environment (similar to Vercel's playground) embedded into your GitHub workflows.
  • Aims to lower barriers to enterprise-grade AI adoption by tying AI development to familiar GitHub processes.

How it works (quick setup outline)#

  • In your repo, you’ll use a prompt.yaml file to define prompts, models, and parameters.
  • When the repo is private, you get a secure, sandboxed environment for model experimentation.
  • A new Models tab in GitHub Copilot UI exposes:
  • Prompt name, description, and the chosen model
  • Parameters and message stacks (system/user messages)
  • A place to store and iterate on prompts directly alongside your code
  • The ability to provide test messages or test inputs

Sample workflow:

  • Create and edit prompts with system instructions, user prompts, and test inputs.
  • Save prompts, model choices, and parameter settings to prompt.yaml.
  • Run side-by-side evaluations of multiple models using identical prompts.

Code blocks and UI are designed to support back-and-forth prompt iteration, including a message stack for realistic dialog.

Core features you’ll actually use#

  • Prompt storage under version control
  • Model selection and parameter tuning in one place
  • Structured prompts with system instructions, test inputs, and variables
  • Model comparison: run multiple models side by side with the same prompts and inputs
  • Evaluators: scoring metrics like similarity, relevance, and groundedness to analyze outputs
  • Prompt/model/parameter triage saved in a single file (prompt.yaml)
  • Private repo safety: work remains in your org, with governance around model access

Evaluators and metrics#

  • Built-in evaluators let you score outputs to guide cost and quality:
  • Similarity: how closely output matches expected ideas
  • Relevance: alignment with the task
  • Groundedness: factual alignment with inputs
  • Use these scores to:
  • Filter which model/prompts to adopt
  • Optimize prompts for the best return
  • Control costs by preferring cheaper models when quality is acceptable

Collaboration, governance, and admin controls#

  • Organization admins can allow all models or create a allow/deny list
  • Makes cross-team prompt sharing practical while keeping governance tight
  • Great for coordinating prompts across large teams and ensuring consistency

Tooling and the bigger picture#

  • GitHub Models is positioned to work with tool-ability in Copilot:
  • API/tool-calls in the generative stack (coming via Microsoft’s tools)
  • Possibility to define personal tools or company-specific toolsets within Copilot
  • Plans around open-source Copilot concepts (C-Pilot) that could enable custom tool integration
  • Practical implication: you could ship your own tools that Copilot can call, all managed inside your private repo

Note: While some of these tooling capabilities are still evolving, the direction is clearly toward embedded tool calls, extensible prompts, and self-contained tool ecosystems inside GitHub.

UI/Workflow highlights#

  • The models workspace shows a side-by-side comparison UI for models and prompts
  • You can commit prompt changes like you would code changes, keeping prompts traceable
  • The models page exposes code snippets you can drop into projects, enabling quick adoption
  • Test data can be embedded in prompt.yaml so samples exist alongside your prompts

Real-world takeaways#

  • If you’re building prompts for products, GitHub Models lets you iterate faster with real evaluation metrics (no more guesswork)
  • You can compare multiple models with identical prompts to see which one meets your needs at a given cost
  • The ability to save prompts, models, and settings in a single YAML file makes it easier to version and share within teams
  • Admin controls help you manage risk and keep teams aligned on which models are allowed

Practical tips and cautions#

  • Start simple (KISS): use a minimal prompt.yaml to get a feel for iteration before scaling
  • Leverage evaluators early to quantify gains and justify costs
  • Weigh privacy carefully: even in private repos, consider what data is used for training or feedback
  • Expect the tooling to evolve quickly — stay flexible and keep prompts modular

Getting started (actionable steps)#

  1. Create a private repo or use an existing one and enable the GitHub Models workspace
  2. Add a prompt.yaml with a basic prompt, a chosen model, and simple test inputs
  3. Use the Models tab to create side-by-side experiments (e.g., GPT-4 vs. another model)
  4. Enable evaluators (similarity, relevance, groundedness) and compare results
  5. Save promising combinations and incrementally add more prompts, tests, and parameters
  6. Explore tool-calls concepts and await upcoming Copilot/C-Pilot capabilities to add personal tools

Takeaways#

  • GitHub Models turns prompt experimentation into a first-class, version-controlled workflow
  • You can compare models, use evaluators, and iterate quickly within GitHub
  • Governance and private repos help teams adopt AI safely at scale
  • This approach aligns with the broader trend toward embedded tool calls and customizable copilots
Transcript

The focus has been on Enthropic for I don't know like two days, but Microsoft's cooking. They are really starting to push on areas that other companies just can't because they have GitHub. And GitHub, that's leverage. So, let's talk about it. What am I insinuating? What am I getting at here? I'm getting at the new models release. Now, it's been out for a little bit, but I have not been testing it until today. And I wanted to share because as we developers try to figure out all the tips and tactics and new extension here and new prompt there, we can get overwhelmed. And I've been shouting this for the last couple of weeks of how do we keep it simple, right? Kiss. And so the closer that it gets to GitHub, the better in my opinion. And this is a great leap forward with their new GitHub models. So find and experiment with AI models for free. About GitHub models. Essentially what it does is it allows you to store a file type with a certain.prompt.yamel or YML either one. And if your repository is private, then you essentially get a Versel playground like thing. So, if you've played with the Verscell playground, then you know what I'm talking about where you have the different models and you can fiddle around with them. You've probably built one of these yourself. And this being under version control is just such a game changer. So, let's read about it. It's a workspace lowering the barrier to enterprisegrade AI adoption. Helps you move beyond isolated experimentation by embedding AI development directly into familiar GitHub workflows. Yeah. If you've watched this channel, then you know I was talking about cursor forever. But what's crazy is the more that cursor ships, the more buggy I find it. I still want to give it another shot next week with the remote agent stuff, but I'm finding augment actually to be far more performant and reliable. But if you've watched this before, you know that I've mentioned this instructions thing here. And that file extension allows you to then tag those prompts within your copilot. But because I'm using augment because their context engine is just bananas, I don't have a real need for the instructions thing. But now they've added another one. So if you haveprompty and you pop one of these open, then this will enable a new tab in your GitHub called models. And in here you can do the name. This is just the example one. The name, the description, the model that you're using, the parameters within that. So everything that you'd expect to see in a playground is just right in here. And then you can actually provide a string of messages. So a lot of time if you've made prompts and you've worked in playgrounds before, you know that it's helpful to have some back and forth. So that way if this is called, it has that message stack in there. So this is really important for anyone that is using prompts, which is most people watching this channel. If you're using them in your products especially, they're very helpful. And it is fun. Like I'm just This is just going right. You guys probably noticed that I have old Auggie just ripping here, but it's it's fantastic. We've seemed it's going to have this one issue. I'll go in there and fix it. But back to here. So, it allows you to do a set of features to support prompt iteration eval. So, did they just delete a bunch of eval software? Because it's literally tied in there. I think the only Achilles tendon thing here is if you have privacy nuts, but it's on a private repo. But still, people that sign some terms of service that are a neck beard are going to be like, I can't have that. I can't have you training on my stuff. But you have prompt development directly in a structured editor that supports system instructions, test inputs, and variable configuration. You have model comparison. So test multiple models side by side with identical prompts and inputs. Experiment with different outputs. Evaluators. Use scoring metrics such as similarity, relevance, and groundedness to analyze outputs and track performance. That's a huge one because that's how you get cost reduction while also increasing the output of the desired return prompt configuration. So you can save prompt, model, and parameter inside that file. Pretty dope. And then something I find really interesting with Microsoft's approach lately where they have APIs for all these things, but also the open source element of C-pilot coming that's going to be next month. It's the first month of their financial quarter is next month. Pretty excited for that because then you can actually do tool calls within their generative language API. So it looks like I couldn't figure this out. So I'll go back and fix this. But what that would mean in the case of Copilot is if I pop this open and then I did an at I can actually write my own tool which is really cool. Very cool. And it would be actually in here. It's not the ad, it's a slash. But that's what you can see here with these mermaid and like mermaid's one that supports extensions, but you could write something in there that is your personal tool set. And that's what I think bigger companies are going to be doing is Zach mentioned that majority of their code is being written by their company co-pilots, but those are custom tools that are trained on their codebase. You can think of it as like a giant mixture of a state machine with a bunch of embeddings, but also rerankers and all that stuff. Go watch a cursor video of the founders explaining it. But if you go read the augment code blog, it explains it pretty well. And so to use that, you flip this on and then you have to have a private repository. So yeah, put that in there. You get all these things. Pretty freaking cool. And then another thing that I wanted to talk about on their blog. Let's just see here. So I finally found it. This would be a UI of how it actually works. So then you can see a comparison side by side. Obviously GPT41 is going to be more expensive. So you would expect it off the bat to be better. It's running on the main branch. And then you can see that similar to different types of workflows. You have the status of it. So it would be running that eval and then providing you with the score that you can yourself dictate. So you might have a relevancy score. We might have something different and let's just watch this quick video to the prompt and commit the changes like any other code. The compare tab lets you run experiments to understand how variations in your prompts and which models you use affect the output. You may notice that this page already has some sample data and that's because we included it in the prompt.yaml file as test data. I can add additional prompts to see side by side how different models, parameters, system, and user prompts perform. You can also apply evaluators to quickly see what combinations meet your desired output and see it directly in the data set grid, which we'll look at in just a moment. The existing evaluator uses an LLM as a judge, resulting in either a pass or a fail. Let's see how the evaluators are shown when running two different variations. In this case, two different models. What's funny about this is I literally was let's see where's my sky draw was drawing this to make it. I don't know where it is, but it had inputs, outputs, and intent. So awesome that we don't have to build that because it's just in there now, which is very cool. Skip back. You can see here that all four requests, both prompts against both inputs have all had that one evaluator pass as well as some other metadata about each request. Organization admins can allow all models or specify a list to allow or disallow. This is a lovely way to collaborate on prompts with everyone in your team. And once you're ready to implement a prompt in your project, you can use code snippets provided in the models page and use over 40 models with just your GitHub account. That's bananas. So start playing around with that and I think it's going to change the way that a lot of us build. If you learned one thing in this video, make sure you like it. And if you want to see more videos like this in the future, then you should definitely subscribe and consider joining VI. It's a private builder network where we get together, we learn together, we have workshops, and we build together. I will see you in the next one.