Parker Rex
All videos
Parker Rex Daily

Automating YouTube Content with AI: Tools, Costs, and Strategy

April 19, 2025

Watch on YouTube@parkerrex2

In today’s daily update, Parker dives into automating YouTube content with AI: the latest tooling, cost realities, and a practical strategy to build a production pipeline without getting bogged down in hype.

News and quick demos#

  • Cursor and OpenAI integration is evolving: newer capabilities move beyond OCR, with edge-function workflows and easier page changes. One-tap sign-in and auto-diff application are in the mix, plus visible diffs and quick reverts.
  • Ader polyglot leaderboard: contrasts architect mode (read-only) with coding mode (doer). Top performers can be pricey (GPT-4o-type costs), so cost-performance tradeoffs matter. Klein and Codeex get a mention as notable players.
  • Vertex AI experiment notes: Parker tried analyzing a YouTube video via Vertex, aiming for a structured, point-extracting prompt. It required a transcript and some workaround to fetch details. Demonstrates real potential but also friction around transcript access and workflow speed.
  • Emphasis on practical, modular AI workflows over “monolithic agents” for production-grade apps (12-factor style; more on this in Strategy).

Tools, models, and costs you should know#

  • AI tooling mix:
  • Looker dashboards for marketing data and trends (Google Trends, YouTube trends).
  • Vector search for comments (up to thousands daily) and content discovery.
  • Image generation and thumbnail creation using Vertex Gemini (and image-text overlay piping).
  • Practical cost takeaways:
  • Self-hosted VPS (example: 8 cores, 16 GB RAM, 512 GB SSD) around $14/month — compelling for a lean startup automation stack.
  • Cloud/serverless costs can scale quickly: rough figures discussed show around $18/month for ~3 million 1-second executions on some clouds, with Azure/AWS similar ranges and storage costs adding up.
  • Bottom line: for smaller ops, a well-architected VPS can be cheaper; as scale grows, cloud can win but you’ll need DevOps discipline to keep costs sane.
  • Other references mentioned:
  • 12-factor apps (Heroku-era guidelines) for building robust AI services.
  • LangChain and other agent frameworks (the talk favors modular LLM loops over full agent stacks).
  • Code and tooling ecosystems like Vertex AI, Looker, and vector databases for content and marketing workflows.

Production pipeline: level 1 and level 2 orchestration#

  • Level 1: Content ingestion and post-processing
  • Ingest video to storage, run transcription, generate subtitles (VTT), and auto-edit to remove filler words and pauses.
  • Post-processed assets are funneled to main and daily channel pipelines.
  • Objective: fast, repeatable, testable post-production with minimal manual steps.
  • Level 2: Image and thumbnail automation
  • Generate multiple thumbnail options with Gemini; text overlays on top via a follow-on step.
  • Use a Discord hook to pick “1–10” options, then trigger an automated thumbnail upload and share the link for review.
  • Long-term: you can run this on a VPS or a beefy container stack with optional cloud-backed storage.
  • Data and insights layer
  • Use Looker (or similar) to visualize marketing data: audience trends, emerging topics, and cross-channel performance.
  • Leverage vectorized comments, trend signals, and content search to steer future videos and thumbnails.

Strategy and best practices#

  • Embrace modular, small LLM loops rather than chasing monolithic agents.
  • Focus on natural language prompts that map directly to tool calls (read email, update CRM, check package status, etc.).
  • Build deterministic control flows (one prompt decides next step, then a switch/call pattern ensures predictable outcomes).
  • Treat the prompt as close to the metal as possible: reduce ambiguity, favor repeatable steps, and keep critical decisions locked to structured logic.
  • For YouTube automation, start with data and orchestration first (transcripts, post-edit, captions), then layer on visuals (thumbnails, overlays) and distribution (cross-posting, metadata enrichment).
  • Have a clear hosting plan early:
  • VPS for cost discipline and control.
  • Cloud where needed for scale, with a DevOps mindset to keep costs predictable.

Q&A and practical takeaways#

  • Is this overkill for a daily update? It can be, but a lean, modular setup pays off as you scale. Start small, prove the ROI, then layer complexity.
  • Costing sanity check:
  • VPS ≈ $14/month as a baseline for a multi-app automation stack.
  • Cloud serverless can be cheap at small scales, but spend grows quickly if you don’t optimize functions, storage, and data transfer.
  • Expect storage and egress to push monthly bills up; plan for cost monitoring and simple dashboards to track usage.

What Parker is building next#

  • A two-stage orchestration framework that:
  • Automates video post-processing (transcripts, edits, captions) with a clear, cost-conscious hosting plan.
  • Generates and tests multiple thumbnails, then uses a feedback loop (Discord-based selection) to finalize assets before publishing.
  • The goal: save time, improve consistency, and make video content more scalable without blowing up costs.
Transcript

Hey there, I'm Parker Rex. I led tech for a startup that sold for 23 million bucks. After that, tried my hand at building Airbnb for music. Did not work. And on this channel, this is a daily upload. I share some of the news if it's actually worth talking about an AI. Try to be that filter. And then I go through some of the questions that we get on our channels. We have the daily channel and we have main channel. And then I go into some of the strategy things that I'm doing, things that I'm building that I'm excited about. and people tend to learn a couple things and it's fun. So, I moved myself to the bottom left after some feedback because I realized I was kind of in the way. And let's get into the news first, then we'll do questions, then we'll do strategy. So, first off, AI, open AI is definitely making its way into cursor, which is really interesting because they're also thinking about buying winter. probably already saw this news, but it's interesting using it now because if I have a change that I need to have made, they started sneaking in by putting this little thing here that works with Oops, I got to move this over. Works with, but now it actually can go and make the changes beyond just OCR. I don't know how they're doing it, but it it works. It's kind of crazy. You can see I'm debugging these edge functions to make sure everyone's rolls are correct. It got a lot of rolls and I'm trying to do I added one tap sign in with Google. So if you've been on a website before like Xedia or something and it says continue as Parker then that's what that is. So setting that up with an edge function and that's just what I'm doing right now. But yeah, it's interesting because if I just pulled up the companion and I said, let's make this page better. See what it does. We have 04 mini running. So, it's reasoning. It's looking at cursor. Before you could only do from line zero to about line 40, which is what this would be. But now it can go the full way. So, let's say implement. Thanks. See it in action. And it says cool. It's writing code. And then it has auto apply code on. I don't know. That's And you'll see it just go. Yeah. So it includes things not just on one line that it's visible. Oops. Where did we go? Oh, I already applied it. See, that's interesting because it just does it. And so you actually see your diffs in here. Big win for OpenAI, honestly. But also big win for Xcode users. They're probably loving this right now. And how I can just hit revert is kind of wild. So, no clue how this works. I'd love to see maybe like a man-in-the-middle proxy between these two apps. Not going to try to go do that right now, but just to know what in this code base allows it to access all the contents of that. And then how does cursor feel about that? Because they obviously had to access or allow this to happen. But it sure seems like they're eating cursors lunch ever so slowly. What a strange thing. But I noticed that today and I just wanted to call it out. Then up next, I wanted to talk about the Ader polyglot leaderboard. Actually, yeah, it is the coding polygot. And on here, you see how they rate the models when combined for architect mode and coding mode. Architect mode is read only and coding mode is the doer. So 03 high and GPT41 is at the top, but it costs what? 10 times about 10 times more than if you just did 25 pro preview the whole way. So that's kind of nuts, right? Yeah, that's nuts. Literally do this six times in a row and then you'd still be winning. So good job on them, but the cost just doesn't make a ton of sense. And if you guys aren't familiar with Ader, I used this for the longest time until Cursor kept up. But Ader and Klein are awesome. Like they just they just innovate like crazy. I think Ader probably lost a lot of popularity after the Cloud Code and now Codeex came out. I haven't used Codeex at all. If anyone anyone's used Codeex, let me know how you're liking it or not liking it. Next up, I was thinking about doing this as a demo. I'll show you what the idea is just real quick. But it's like if I see sir, I guess we'd have to use this. Now I'm a little confused. Is it just the model? Let's see. Let's We're getting to the bottom of this. I want to make a prompt and vertex to have it analyze a YouTube video, watch it, and then return the best points vertex. example. What do I do to make this reality? Here's the prompt. You know, sometimes things just don't go as planned. But let's put that over here while we obtain it. Never mind. It's too fast. Okay, so manually go to YouTube show trans. So, it needs the transcript. I guess you can't do it via the API. I don't love that. Let's see if we can't do it in AI Studio. I know you can do it in here. It's just weird that it's weird to me that you can't do it in Vertex. We can do it in there. Again, I keep closing this out because I think we're going to solve it. Oh, here. Got boom. All right, we just tried using vertex and failed miserably. But now we're in here and it seems fine. And then I do want to have a structured output. So, let's I guess it doesn't matter. Let's just run this. We're going to crank the temperature down. We don't need any tools. Run. Right at the million mark. About that. Cool. That should take some time. Like I would expect that to actually take a fair amount of time. So I'll put it there and I'll save you with posts by not making you watch review fight or text for so long. I wanted to look into this article. This one should be good. So may have hit a nerve here. It all started by trying to understand how production AI systems actually work. Folks, I've tried every agent framework out there and talked to many strong founders building impressive things with AI. But I sort of surprised to find out that most successful AI systems aren't following the here's your prompt, here's a bag of tool powder. They're mostly just well-engineered software with LLM capabilities integrated key points. The companies shipping high quality AI aren't building monolithic agents from scratch. I knew it. Failed to generate. What a tough L. We're going to plop in a shorter one. Let's go to YouTube. Come on, buddy. I don't want to be debugging with browser MCP right now. Let's go back. Let's get one of these. Let's grab this. I didn't think we could. That's That's an hour, too. How about this? Yes. Take a If he can't take a fire ship, then this thing I'm going to be upset. Put that there. There. Delete this. Oh, man. Recast. Now it won't even not even show in the YouTube video. What? Come on. Buggy buggy. Okay. 5m minutee video. Please be able to do this. 87,000 tokens. You got it, dog. Okay. So, this guy was talking about agents. So they're incorporating small focused LLM loops that do one thing well. Yeah, I've been saying that. So I set out to document the principles for building production grade LM applications and 12 factor agents. We're on the front page of Hacker News all day on Wednesday with great discussion with the community. Check it out below. Let me know what factors you think are missing. Is that still going? Yes. Scoot that over. And I've been Okay. In the spirit of Heroku's 12 factor apps, this worked by the way. Cool. Great demo. In the spirit of Heroku's 12factor apps, I don't even know what that means. I'm an idiot. 12 factor. It's this 12 factor app in the Oh, no. I have read this a very long time. Yeah, very long time ago. I've seen many SAS builders try to pivot towards building AI or towards AI by building green field new projects on agent frameworks only to find that they couldn't get past the 70 80% reliability with out of box tools. Yep, shout out lang chain. And the ones that did succeed tended to take small modular concepts from agent building and incorporate them into their existing product. Yeah, that's literally it. It's putting them on to little parts of the app to handle the very well-defined task. So factor one natural language to tool API calls. Add a ticket to restock the file cabinet. Read this email thread and update the CRM for Acne Corp. What's the status of my package? That's wrong. Use the XYZ project. Benefits feels like magic. So contact comes in. Determine next step. Read, write data. Yeah, that makes sense. Unfortunately, we learned this the hard way. The frameworks are useless if you're building anything novel. I've asked other founders about specific capabilities. It's always, oh, we built that ourselves. Yeah, definitely. Another working title is going to be agents the hard way. Yeah, this is literally what I struggle with map is you have to even think of perplexity. Perplexity has so many different tools baked into it, but you just don't think about it. Like it's a lot. Active two, owner prompts, agent interface, ro as a string, goal as a string, personality string, task, instruction string, expected output, type base model. Yeah, tools, list of tools, blackbox versus full control. Yeah, you got to get closer to the prompt. That brings me to the thing we're going to talk about when I get into the strategy stuff because I just want to own I want to be as close to the metal the metal the prompt as possible. So in this case let's see one prompt determine next step two switch statement do we either call the API yes we want to have deterministic code do we kick off a pipeline to be update the DB yeah and then we get to the final answer versus if you had this and it was wrong then you keep no I get it okay so you keep looping provide the context I don't totally understand this okay one prompt do this It's going to do one of these. Yep. This makes a lot of sense. This is how I think about like agentic workflows. If you were to do auto replies on things, if you were to scrape all the YouTube comments off of someone's video and then classify them, hey GBT, is this a question? Is this one that you can handle that's a question based on your context? Is this a comment? Is this a spicy take? Each one of those, it can read through and figure that out. And then it can give a number. So it's like, yeah, it's this one. You don't have to give it a number, but that's if you're doing it like N, but it would classify it and then it would go and call something. But it's just literally like NLP on turbo and then you want Yeah. deterministic stuff. So this is like music to my ears. So I feel like been yapping about this for a while. Having done it full agent mode, it just doesn't work. Especially with like health information. Oh my gosh. Yeah. Here we go. Generative AI for marketing. Now, this is really interesting to me because I just spent a bunch of time starting to build out my orchestration layer where when I open up my Finder, first of all, this is cool. I showed this yesterday, but GCS views gives me access to my Google Cloud bucket right here. And then idea being as soon as I finish a video, I just drop it in here. And this is my test one. And when it's in there, then it invokes a serverless function, does a trigger, and then it calls out to service. It does a bunch of post-processing, and it does the transcription, it does the subtitles, it does the captions, it does all this stuff, and then it makes it sound really good, and it also edits it. So, it'll clip out the pauses and the filler words, and then it'll drop it into processed. So, you can see that there, which is nice. So I can see it has VTT which I think is subtitles is a JSON and then it has this thing. So pretty cool. And then what you can do with that is have another file, another Python file that does the uploads. So the upload to the main channel, the upload to the daily channel. And it's just going to save me so much time. And that's level one. So, I'm gonna go to level two, which is how do I use imagining for you guessed image generation. So, let's go to and that's vertex imaging. Go into our Gemini. Oh gosh, that's not it. Let's go into our console. Let's go into vertex. Let's go into where create a prompt agent garden agent. Where are you? Mala garden. go into image in and we can use this which is going to be so cool. So we can call out to this and then make different images based on what we want. And I actually am curious. I want to see make a studio jibli representation of a duck in a pond. I have no idea if this is going to work, but let's give it a try. I don't know if it can do jibli. Maybe it can. But I actually started to Oh my gosh. Yeah, that bottom one. Are you kidding? It's perfect. Very good. Cool. Cool. Cool. So, I'm excited about this because what I can do then is I'll get the first pass at it and then I can use something on top of that to put the text on. So, I can have titles that are generated and then I can get a Discord hook that lets me know here's the titles, which one do you want? And then I just text back and I'm like that one. you know, do you want one through 10 based on the transcript, based on all this stuff? Which one do you want? And that's gonna be sick because literally I just record, I drag and drop, I get texted the thumbnails, and I say, "Which one I want?" And then another Python function runs when I do that, and it just uploads the thumbnail. Then it can send me the link and say, "Does everything look good to you?" And that's just so sick. So really excited about that. And I think just the versatility you have, it doesn't have to be on GCP, by the way. I I realize that cuz I have a beefy VPS that I am planning on setting up that can handle all of this because it's just at the end of the day a bunch of Python functions and it would not make sense to host it from a cost perspective on GCP, but I want to learn it. So anyways, that that was interesting to me. And then if I show you, there's some logs of me rolling this thing. Where did our Google GitHub go? Here we go. So, this is really cool. It's powerful. So, take a look. The idea being that you have a big old set of like kind of like marketing orchestration stuff going on. So you can do news scraping, you can do translation API for other news in other areas of the world. But in my my case, I'll just read it off to you because it's a little easier. But you utilize the Looker dashboards to access and visualize marketing data. Marketers can access and visualize marketing data to make data driven blah blah blah. But really the one that gets interesting is this audience and insight finder. So when you did have data, then that wasn't it. It's the trend spotting. Identify emerging trends in the market by analyzing Google trends data on a looker dashboard. Same thing goes with YouTube. So you can see what's trending and then not only across the market of what people are searching, what people are interested in, but within YouTube. And then also you can vectorize all the comments on your videos as well as others videos up to 10,000 times a day. So it's kind of bananas. And then content search. You can use vector AI search. What's this do? Oh, this is internal. Reduce time with vector or vertex foundation models to create email copy, website articles, social media posts, and assets for Pmax. What's this? Performance max campaigns. No idea. It looks just like a ad campaign for Google ads, but whatever. But this is the one that really makes me excited. And then this one has lang chain. I don't know why they'd use that, but this is a new summarization one. Image gen fine-tuning. I have to take a look at that. Tuning Gemini with your data set to improve his model's response to get your brand voice document summarization techniques. Ve vertex AI web search. How to search through corpus of documents. I don't know what web I guess they call it. I just don't understand what the difference between web search and document searches. Hi there, my name is Holt SK. the data that you've imported in already on the cloud console or something like that. So, I have a few different examples here of what those look like. We have, for example, the Google Cloud website. I can search for things like Vertex AI or Vertex AI search or I don't know, document AI or something. So, there we go. And it'll show all sorts of stuff. It's from cloud.google.com. So, anything that's on there is fair game. Oh. Um, there's a few other examples on here. Some of these are for websites like Wikipedia. I have one that you can just search all Wikipedia. I have one. I have one for University of Missouri, my ELM. This one's actually a blended engine, so it uses a few different sources on there. It uses some from, I think, the University of Missouri subreddit, the actual University of Missouri website. So, if I type in like computer science, I can probably see. Yep, here we go. So, it shows stuff. This one's using a I want to find a notebook, same category, advanced website indexing. I have one here from Stack Overflow that uses custom embeddings. And I have a notebook that shows how to do that as well. a website here. Let's search for Well, we can search for Vertex AI again or I don't know. Let's search for tab here. The Reddit ones. What's interesting to me? And this right here is an image that I just took myself. This is not from text search for this socks. There we go. Is a Yeah. So, basically could let you bundle together the news sources and then find stuff in there. Don't totally understand that, but I just think these are so valuable. So, I'm excited to fiddle around with some of these. And for marketing insights, let's see what this one looks like. No, that's not it. Yeah. So, that actually plays into the strategy stuff that I was talking about. I kind of like mentioned it, but it's basically like, cool, I did the first step, which is this, and then other things that I'm building are this, which is I vibe with AI with the thing that I added today is going to be super helpful for converting users, getting them to have saved prompts and all of that jazz. And it's like a really simple setup. I think I just was splitting my time between a bunch of different projects at the same time between the content and the code. So it's taking a little bit longer than expected, but you know, used Taskmaster literally the whole time, which is nice. That one was for I vibe with AI. Yeah. And then I'm just excited to keep kind of hammering away at this stuff. My goal is I'm going to put this on a VPS because someone asked about the pricing. So on the Q&A side, I go to our questions. Someone was like, "Yo, this thing seems a little overkill," which it is. Can you share any general cost figures? I've explored GCP stuff while trying to use their free tiers. Yeah, there's free up to I think it's like 300,000 serverless functions, but kind of kind of crazy. Let's see. They get you on like the storage is what I was trying to say. I was like half reading that other comment about But if you look at it, so I gave it a bunch of specs and I said I want to have let's see what did I say? GCP versus Versell and self-hosted on a VPS with the following specs. And so I said thousands of operations a day with n or thousands on serverless functions and then five to 10 apps with a mixture of next containers, express containers, python containers, whatever. And then I also I mean I did say I think GCP is going to be superior but come up with the price structure. And then I gave it the specs of my VPS. So yeah, I said there's my VPS. I got 8 core, 16 gigs of RAM, 512 SSD for $14 from Netcup, which is crazy. And then as you go down the list, this is where it's kind of interesting. So you get this compute number. Kind of crazy. Free tier, 400,000 gigabyte seconds a month. So not 400,000 requests. Silly. And that would come out to 300,000 executions, aka 150,000 GB seconds. It's so confusing, right? But it basically says if you got to 3 million executions a month, then it'd be $18. That's from this price comparison website. Maybe this is better. It's probably less confusing. I mean, serverless can it's honestly just really confusing. 3 million at 1 second is $18 a month if they're just 1 second and they're not longstanding. But some of mine are definitely longer. That'd be 18 bucks for that. How much is it on Azure for serless? 18 bucks. 25 for GCP. How about it? So Azure and AWS similar monthly. Lowest unit price per GB. They seem to have the highest price per month because cost incurred for CPU. Okay, good to know. I guess. But yeah, basically what I got out of it is this seems kind of crazy to have 87 a month for SSD as a persistent disc versus, you know, the $12 all in 12. I think it's like $12.80 in euros. And it ended up like GCP is just obviously going to be more, but I think it's cuz you could go to the moon with it. And I was like, how do I make it less expensive? because it ended with saying, "Yeah, that is definitely superior, but it's going to be $261 to do what you want on GCP versus your $14 VPS. You just need to learn some DevOps." So, I'm looking at that and I'm like, "Okay, I'm not going to use Nad for all this stuff. I'll use Python functions. I'll use Nad for random things." But yeah, kind of wild. So then I just ended it because I've never run the VPS on Debian and I've only run a VPS with one app, not doing Docker Swarm for if things mess up or whatever. And it says that it can handle up to 10,000 operations a day or and 3,333 requests per app per day for how many apps did I say? Seven. That's kind of kind of crazy, but yeah. I said, "Yo, I don't have that much experience. I got Teras." And I found that Grock's actually really good at this kind of stuff. Like I'll run with Grock for initial like setup and then as soon as I can, I get into a an IDE. So yeah. Are there any other questions? Keep them coming. I think what you're looking for is WP sell the community and you can build native apps that live inside and extend the app. I don't know if W can native apps developer W developers allows you to sell your applications simple and secure way to monetize sell access to a website. Use OOTH to sell access. Check out our API. Sell access to software. Build an app. Build an app on WO app extends it. Oh my gosh, that is kind of what it is. You can build an OP. A unique experience hub fizz connects the APIs writes business data. Let's take a look. This actually sounds pretty cool, not gonna lie. Video app. That is so lame. Let's see. So, you have to make a neon Postgress. Oh my gosh. This is a whole guide on how to make one. That's funny. How do we get Okay, here we go. API listing memberships. Oh my gosh. Yeah. See, I wish school had this. I Wow. Yeah, this is way better. List plans. Create a checkout session. Retrieve a company. Yeah. Oh, man. That is such a game changer. There's nothing in school and that's why I spent cumulatively I probably rolled like 10 hours into this thing. Yeah, six today. Let's see. Just adding all the stuff, right? Like it is a fair amount. Why does it keep opening in Safari? See rave waka time. Waka time probably got like 12 hours into I with AI just getting all the stuff ready. Four today. How about across the seven days? 14. Wow. Unbelie Yeah. But problem is I spent a lot of time in that PRD. Problem is if I don't have an independent, I miss out on a lot of things. But this could be another thing you could do. Is how many people are on the website? Let's go to wop.com. Actually, before we do that, let's see how many people are building on here. Because if you're the only one and you build some badass ones, that could be cool. I want to see. That's JavaScript. Let's do that. Let's see if anyone building in here. Let's try this grap GitHub thing. Yeah, here we go. Zero. This. How about this? No, I I don't really understand grap because I don't think that there's zero. Something tells me there's not zero. Yeah. 26. So, this guy made that's just their docs storage. What's this data storage? So, this is his tool. Yeah. Print something. It's got capture. Get homework ID. Try to send that with a license key. Cool. Let's make a request. And then what is this doing? It's a web hook. What's this do? It's a hardware ID. Yeah, that's fine. Must be something that he's selling. Oh gosh. Where's unfinished? See anything in there about this W provider? What's he building? It's dead. What is it supposed to do? Two years ago stack. What is gratitude? Nothing. Go validating a license. So, this guy must sell something and he's selling licenses to this auto FX app. No, that's not it. Copy trading application using the fix. So, he sells Yeah, that's just an access thing. Memberships. Memberships. Memberships. Yeah, these are all just selling you access to another thing. Bet user access get their SDK. Let's take a look at theirs. SDK. W They have much Fallout 56. Let's see their Discord. Actually, discord. Where's their Discord? Oh, anyways, that was a bit rambly today.