← Back to Blog
Application Guide·July 21, 2026·Gabriel Jarrosson

Two AI Agents Directed Full Music Videos for Under $75, Unsupervised. What's Left for a YC F26 AI Video Startup to Sell?

Two AI agents directed full music videos for under $75 each, unsupervised. What that leaves for an AI video startup applying to YC Fall 2026 by July 27.

Share

Two AI agents each directed a full music video for under $75, with no human. None were any good. That gap is your YC F26 AI video pitch.

YC Roaster

A post on the Hacker News front page, published July 16, handed two frontier models a song, a hard dollar budget, and a shell with ffmpeg, then walked away and let each one direct a music video by itself.

The setup was bare: six tools, and the same inputs every run (Mark Ronson and Bruno Mars' "Uptown Funk," a short description, a time-stamped lyric transcript). Claude Fable 5 and GPT-5.6 Sol each ran twice, at $25 and $100.

All four runs finished on their own in 39 to 50 minutes and produced a valid full-length video with the song muxed in. Claude Fable 5 at $100 generated 80 clips and shipped at 1920x1080 for $73.65 all in ($48.60 generation, $25.05 tokens). The cheapest run, GPT-5.6 Sol at $25, came in at $27.45.

Then the author's verdict: "None of the music videos were great."

That sentence is the entire opportunity for anyone drafting a YC Fall 2026 application around AI video.

Does Y Combinator actually fund AI video startups?

Yes, consistently, across batches.

Focal (YC W24) is an AI movie studio from Robert Cunningham and Felix Wang, high school classmates who studied math and physics at MIT and math and CS at Stanford. It turns books and screenplays into movies.

Golpo (YC S25) is two Stanford CS brothers, Shraman and Shreyas Kar, generating whiteboard explainer videos from a prompt or document. They left Stanford and raised $4M.

Absurd (YC F25) makes AI brand and performance ads. Their Kalshi spot did over a million views, their videos average 400K+ organic views, and Hims, Replit, and Brex pay them for 72-hour turnarounds.

Sieve (YC W22) launched as pluggable APIs for video search and is now a research lab selling video datasets to frontier AI labs.

So the category is open. The question is what you can say about it now that isn't already commoditized.

What does a $75 autonomous music video prove?

That the pipeline is finished work. Research the models, generate clips, cut them, sync to audio, mux the file: a language model does all of it unsupervised, in under an hour, for less than dinner costs. If your F26 application says you generate video from a prompt, a partner can run the open-source harness the author published and reproduce your product before your interview slot.

What the experiment did not solve is the interesting part. Read the failure list as a product spec:

  • Recurring characters drift between shots. All four runs. None of the videos held a coherent storyline start to finish either.
  • No model seriously reviewed its own work. Once clips existed, the models concatenated and muxed. The author notes none probed their own footage to confirm it was any good, and one run shipped genuinely low-quality clips anyway.
  • Motion doesn't match tempo. The cuts landed on the beat, because every run used ffmpeg beat detection. The dancing inside the clips did not.
  • Lyrics get read literally. "Make a dragon wanna retire, man" produces an actual dragon.
  • Neither model spent its $100 budget. They stopped at $36.57 and $48.60. With that headroom either could have generated consistent character keyframes first and animated from those. Neither chose to.

Why are YC's video companies already selling into that gap?

Because they scoped themselves around it years ago. Focal's launch post, from March 2024, lists keeping characters consistent across clips and lip-syncing dialogue among the things its product manages. Those read as capabilities, but they map onto the failures these agents hit more than two years later.

Absurd is blunter. Their launch post says there are dozens of AI video models, each good at something different (one nails motion, another faces, another lighting), and that knowing which to use and how to handle their quirks takes serious trial and error. Their answer: agents doing the creative work, humans stepping in to steer.

Golpo solved it by amputation. Whiteboard explainers don't have a character consistency problem, because there are no characters.

Doesn't "human in the loop" fail as an F26 answer?

Yesterday's post here argued that "human in the loop" is getting weaker as a reliability answer, because a partner can do the arithmetic on how many reviewers your margins support. That still holds, and it does not contradict Absurd, but you need to say why.

Where there is machine-checkable ground truth, a human reviewer is a placeholder for a verifier you haven't built yet, and partners will ask when you plan to build it. Taste has no ground truth to check against, which is why nobody has automated it and why the humans in Absurd's loop are not embarrassing. They are also not the moat. If your pitch needs human steering forever, you are selling an agency. If it needs it for now, say so, and say what those humans teach you that becomes the model later.

What should your YC F26 application say instead?

Three positions survive contact with a $73.65 baseline.

Own the judgment layer, not the generation layer

The most damning line in the writeup is that no model seriously reviewed its own output. Automated taste, meaning scoring a clip, catching drift, deciding to regenerate, is unsolved and expensive to solve, which is what makes it fundable. Show a partner a before-and-after where your selection layer turned bad footage usable, and you have a demo the arena repo cannot reproduce.

Pick a genre narrow enough that the hard problem disappears

Golpo's constraint is its moat. Find a format where consistency is structurally easy, own it completely, and let the general-purpose tools flail everywhere else.

Lead with distribution, not the demo

Absurd's application-shaped facts are view counts and named customers, not model choices.

What gets a fast no?

"We use AI to generate video" now describes a weekend project that costs $27 to $74 a run. Claiming raw generation quality as a moat is worse, because you are claiming an edge over labs whose models you rent by the second at published rates: $0.05/s for Wan 2.5, $0.10/s for Veo 3.1 Lite, roughly $0.62 per five-second 1080p clip on Seedance 1.0 Pro. Those are your costs and your competitors' costs, and they fall again before Demo Day.

Sieve is the honest precedent. It launched at video search, watched where value settled, and became the data lab that sells to the labs. Same founders, different layer. That is also the Intuned (YC S22) pattern: get in on evidence you can build, then go where the problem actually is.

When is the YC Fall 2026 application deadline?

July 27, 2026 at 8pm PT, with decisions by August 28. Interviews run through August and September, the batch runs October to December in San Francisco, and Fall 2026 Demo Day is Wednesday, December 2.

If you want a second read on whether your video pitch still sounds differentiated, that is what YC Roaster is for: founders who have been through YC, telling you in plain terms whether your moat paragraph survives someone running the open-source version of your product.

You have six days to make sure the thing you are pitching costs more than $75 to replicate.

Ready to get your YC application roasted?

Get free AI feedback + a review from a YC alumni.

Submit Your Application