← Back to Blog
Application Guide·June 26, 2026·Gabriel Jarrosson

Anthropic Just Accused Alibaba of the Largest Distillation Attack Ever. Can You Build Your YC F26 Startup by Distilling Claude or GPT?

Anthropic accused Alibaba of a 28.8M-query distillation attack on Claude. Here's whether distilling frontier models is safe for your YC F26 startup.

Share

Anthropic Just Accused Alibaba of the Largest Distillation Attack Ever. Can You Build Your YC F26 Startup by Distilling Claude or GPT?

YC Roaster

On June 24, Anthropic publicly accused Alibaba of carrying out what it called "the largest known distillation attack" on Claude. The numbers are staggering: roughly 25,000 fraudulent accounts and 28.8 million model exchanges between April 22 and June 5, all aimed at extracting Claude's most valuable skills — agentic reasoning, software engineering, and long-horizon task execution.

If you're building an AI startup for the YC Fall 2026 batch, this isn't just geopolitics. A huge share of F26 applicants are quietly doing a smaller, legal-grey version of exactly what Anthropic is accusing Alibaba of: training or fine-tuning a cheaper model on outputs from Claude or GPT. So the question lands directly on your application: can you build your YC startup by distilling a frontier model, and will it survive a partner's scrutiny?

What actually happened between Anthropic and Alibaba?

Anthropic laid out its allegations in a June 10 letter to Senate Banking Committee Chairman Tim Scott and Ranking Member Elizabeth Warren. It claims operators tied to Alibaba's Qwen lab used tens of thousands of fake accounts to run 28.8 million exchanges with Claude, deliberately targeting its highest-value capabilities, and that they "ignored the Trump Administration's warnings" in doing so.

The disclosure is already moving policy. Senators Bill Hagerty (R-TN) and Andy Kim (D-NJ) are pushing an amendment to defense legislation that would blacklist or sanction entities running these campaigns. Alibaba disputes the characterization. But the label Anthropic used matters for founders: "adversarial distillation" — training a weaker model on a stronger model's outputs to copy its behavior.

Wait, what is distillation, and why do so many AI startups do it?

Distillation means using a powerful "teacher" model to generate training data, then fine-tuning a smaller, cheaper "student" model on those outputs until it mimics the teacher on a narrow task. Done with a model you're licensed to use, it's a legitimate and widely taught technique.

It's also everywhere in YC's recent batches because it's the cheapest path to a model that feels frontier-grade. Cactus (YC S25) made today's most famous version of the honest path — distilling Gemini into a 26M-parameter model small enough to run on a phone. The reason founders reach for it is obvious: you get GPT-5-class behavior on your specific task at a fraction of the inference cost, and you own weights instead of renting an API.

The problem is that the line between "smart engineering" and "the thing Anthropic just reported to the Senate" is a terms-of-service line, and most founders have never read where it sits.

Is distilling a frontier model against the rules?

For the major US labs, generally yes, when you distill their model to build a competing one. OpenAI, Anthropic, and Google all have usage terms that prohibit using their model outputs to train competing models. That clause is exactly what was invoked when OpenAI raised distillation concerns about DeepSeek in early 2025, and it's the spirit behind Anthropic's Alibaba complaint.

Three distinctions decide which side of the line you're on:

Whose model are you distilling?

Distilling an open-weights model with a permissive license (many Llama, Qwen, and Mistral derivatives) is usually fine and explicitly allowed. Distilling a closed API model like Claude or GPT to build something that competes with it is the prohibited case.

Are you building a competitor or a feature?

Using GPT to generate labeled data for a narrow classifier inside your product is a very different risk profile than cloning a general-purpose assistant. The terms target competing models, not every downstream use.

Are you hiding it?

Anthropic's case isn't only about distillation — it's about 25,000 fraudulent accounts evading detection. Scale plus deception is what turns a gray area into a headline. A founder running thousands of sock-puppet accounts to dodge rate limits has a defensibility problem and an ethics problem.

Will YC reject you for building on distillation?

YC won't reject you for the technique. It will push hard on what the technique implies about your moat. The partner question you should expect in a 10-minute interview is some version of: "If your model is distilled from GPT, what stops the next team from distilling the same thing — or OpenAI from cutting you off?"

That's the real risk for F26. A startup whose entire advantage is "we distilled a frontier model" has two structural weaknesses YC partners are trained to find: a platform-dependency risk (your teacher can ban you, as Anthropic just demonstrated it will) and a non-durable moat (distillation is a commodity skill now). The Alibaba story makes both concerns louder, because it puts "labs actively police and litigate this" into every partner's recent memory.

What should you do instead if your wedge depends on distillation?

Reframe distillation as a cost-optimization tactic, not your moat. The founders who get funded position it the way Cactus did: the distilled model is how they deliver, but the defensibility lives somewhere a competitor can't copy by re-running the same script.

Concretely, build your moat from one of these and let distillation serve it: proprietary data the teacher model never saw (your users' workflows, a labeled dataset you own), distribution and a wedge into a specific buyer, or a system around the model — evals, tooling, integrations — that compounds. Then stay on the right side of the terms: distill from open-weights or properly licensed models, document your data provenance, and never build your demo on a wall of fake accounts.

This is exactly the kind of soft spot worth pressure-testing before you submit. The fastest way to find it is to have someone who has actually sat in a YC interview try to break your defensibility answer. That's what YC Roaster is built for — alumni who've been through the batch will tell you, bluntly, whether your "we distilled GPT" pitch reads as clever or as a moat you don't have, while you still have time to fix it.

The one-line answer

Yes, you can build a YC F26 startup that uses distillation — but distill from models you're allowed to, never make it your whole moat, and be ready for the partner who asks what happens when your teacher model bans you. The Anthropic–Alibaba story just guaranteed that someone will ask.

Ready to get your YC application roasted?

Get free AI feedback + a review from a YC alumni.

Submit Your Application