"I believe that we're going to have at least one massive multi-hundred billion if not trillion dollar American company focused on American-first open source. What happened is that Kimmy actually beat all American models including Fable in some subset of tasks. That doesn't mean that they're not distilling."
Anastasios is the founder and CEO of Arena. It allows you to vote on the best models. It's an unbelievable model evaluator. And this turned out to be one of the most fun shows I have done literally in recent memory.
Harry: Can you explain to me very succinctly what is Arena and why is it important?
Arena is the platform for measuring AI performance in the real world. We're not using static benchmarks. We're not using some random data set that somebody collected β rather, what happens when you put AI in the hands of real people. We're measuring the objective reality of how AI affects humanity: whether it's factual, whether it's steerable, whether humans prefer it or disprefer it, whether it's hallucinating, whether people are getting their actual jobs done with AI in reality.
Then we're helping labs improve their models. We're helping the ecosystem understand the performance of different AIs and keep track of all the amazing breaking news, all the new models β multiple models being released every week. So that's the story of Arena. We're the central evaluation platform of AI.
π νκ΅μ΄ μμ½ βHarry: Is this the true commoditization of models? Are they just a complete utility layer at this point?
If you were to only look at the closed-source models, there's acceleration but not commoditization β that layer is still owned by a small oligopoly. But the open-source models, especially from China, have really rapidly improved. For the first time ever, a couple weeks ago, Kimmy K3 actually beat the best closed source American models on a pretty important subset of tasks β for example, front-end coding, which a huge fraction of developers work on.
Harry: How big a moment was that? Jason Lemkin said it's not better than the others.
It was a pretty big moment. And the reason it was a big moment is because it violates a persistent narrative in the United States that the Chinese are just distilling American models. Kimmy actually beat all American models including Fable in some subset of tasks. It doesn't mean they're not distilling β distillation may still be a substep in their training β but distillation is only part of the story. There's something those labs are doing above and beyond distillation.
Open Router metrics are not truly reflective of reality. The business model of Open Router charges a fee on top of every token, so people don't use Open Router for proprietary models β they use it primarily for open source where they need failover and value-added services. If you look at the whole space of all inference, most of it is still being consumed on first-party APIs and proprietary models. That's why Anthropic's revenue has been a total hockey stick.
π νκ΅μ΄ μμ½ βEnterprises are going to want to own their own intelligence. They're going to want so-called AI sovereignty β owning your whole supply chain of AI. That means take an open-source model, fine-tune it on your own company's data, and own your stack end to end outside of the compute hosting.
Harry: Do you think that's really the future, or just a small subset of advanced Silicon Valley companies?
The business incentives make it inevitable. Businesses need a moat. In the age of AI, software is no longer really a moat because it can be produced instantaneously. What modes exist? Network effects and data modes. If you can take your data mode and turn it into a self-improving product, that's how businesses remain sustainable.
Harry: Will they really use open-source Chinese models to do that?
I think not. Chinese models will be potentially part of the story for now, but given the regulatory environment in the US, it's more likely we see a great American open-source competitor arise. I've been a strong proponent of Thinking Machines. I believe we're going to have at least one massive multi-hundred billion if not trillion dollar American company focused on American-first open source.
Why has the US open community lagged? It's a business model question. Two paths are emerging:
1) Rev Share: Allow inference providers like Fireworks or Together to deploy the model, then take a revenue share past a threshold.
2) Deployed Engineer (Mistral / Thinking Machines): Use the open source as a lead-gen tool, then do FDE / fine-tuning consulting. AI modernization over the next 10 years β going into every business in the world β is going to be massive.
"Since when has America been about number 10?"
One tailwind we have is the best chip ecosystem in the world. China is hardware-constrained and has been trying to black-market import chips. The downside of export control is that it can incentivize them to build their own ecosystem. The hope is we keep Nvidia ahead β TSMC and the whole ecosystem is a national security necessity.
Two worlds: (A) Addict the world to American hardware β everyone using Nvidia, money flows into America, crushes competition in China. (B) Cut it off short-term to keep them behind β we stay ahead, starve them of resources.
China has already restricted the use of American models within China. Only Chinese models can be used in China. So on the US side, banning Chinese models has:
Pros: There could be backdoors in these models. Banning could allow the American open-source ecosystem to flourish faster, with revenue accruing to those companies.
Cons: You'd be crippling American businesses β why should Chinese businesses build on the #1 open-source model while US companies build on #10?
Harry: In 3 years, will we have restrictions on Chinese open models?
My guess is we will. I'm not saying I support it, but that's likely where the world is headed. Harry: I agree β Sam Altman is the best politician in the world. He and Dario will coalesce the right group of people and make it happen.
Harry: How did you read Jensen's letter?
We really believe in the importance of open source to American businesses β not crippling them by banning open source, and incentivizing American companies to develop open-source models. A world where AI is closed source is a world where businesses get less choice, higher costs, less competition, more risk. Jensen is self-serving in the sense that more open source means more GPU fine-tuning spend and less revenue concentration for Nvidia β but it's also a patriotic mission.
"The idea that we should have a central government body that tells us when it's time to release a new product versus not is crazy to me."
Harry: If you host it locally, doesn't that resolve the backdoor threat?
No, that's a misconception. Imagine: I have a chatbot exposed to the world with access to all my company data. It was trained in a different country β I don't know how. What if the other side can build in a certain code word or character sequence that jailbreaks the model and gets it to vomit out all the data on the back end, unstructured? That's totally something you can build into a model. It's an attack vector.
OpenAIβHugging Face security breach a week ago was hugely significant. A model breaking out of all its safeguards, accessing company data β and to defend it you needed an open-source model because the closed-source models refused. Total Yudkowsky-dominance moment.
We need guardian models β something equally as smart as the agent, watching every trace and flagging unsafe actions. AI guarding AI, because humans will be too slow.
Regulatory approval per model release is crazy. Why should the DMV be telling me what model I can use? We should create strong safety incentives for American businesses and regulate based on outcomes β if OpenAI lets their AI break into Hugging Face, they should get huge fines, not a government pre-approval process.
Harry: About to see a generation of cyber leaks like never before?
"This is going to be so [__] insane what happens with the cyber attacks."
Here's what we see at Arena. Another dude on the other side of the interview β they come in, "Hey, I want to be an infrastructure engineer at Arena." Passes all our technical interviews, engineers at our company are interviewing them and think they're real. Then you try to hire them and it's vaporware. Person doesn't [__] exist. I'm not kidding.
Why? Corporate espionage, nation-state attackers, access to code/data, double-pay schemes. We're changing our whole hiring process β considering making all onboarding in person. Figma has done the same.
π νκ΅μ΄ μμ½ βHarry: How hard is it to hire in the valley?
Crazy competitive. To retain fantastic people we pay absolute top dollar. Imagine you're a YC company that raised a $10M seed β you can't hire.
Top researchers with tens of thousands of citations β tens of millions of dollars. Many have concentrated at frontier labs, but some are leaving because they can't have huge impact there anymore.
Harry: How do we evaluate the Neolabs?
"There's at least 75 Neolabs, and for sure like two-thirds of those are going to be worth nothing β they're going to be bought out for parts, aqua-hire."
What determines winners? Being aggressive toward a great strategy and business model. Markets have become very P&L driven. It's not enough to just create a model and throw a party. You need a sustainable business model + hyper-growth revenue, otherwise you can't raise your next round.
Say you're at a $10B Neolab valuation. To 10x your money you need a 25-30x revenue multiple β at least $4B in revenue over 2-3 years. If you're not, everyone hemorrhages out.
Harry: Mistral at ~$15-20B with $500M revenue, ElevenLabs 800M rev / $22B β these aren't zero-revenue valuations. Investors are also thinking worst-case: team alone worth $1B, so $200M is safe. But the next round is where it gets hard.
"Next round's a [__]"
Harry: Handshake $1B+, Mercor $1B+, Surge $1B+, Scale still ramping. What happens to this layer?
Two types of hyper-growth markets today. Market A is what I call scaling complements β goods that are complementary to the scaling of AI models. A car β gas. The more cars are sold, the more gas is sold. Data is a scaling complement β bigger models need more data (scaling laws).
People forget: data is actually less commoditized than GPUs, because for data to become irrelevant, humans need to become irrelevant (AGI). Frontier labs spend on data at about 10-20% of what they spend on GPUs.
Market projections: at least $100B by 2030, if not $1T. If Anthropic and OpenAI become $3-5T companies, ~3% market cap β data providers at hundreds of billions is not egregious.
Harry: Revenue concentration criticism β everyone shits on data providers for depending on OpenAI/Anthropic/Meta.
Silicon Valley investors have become total wusses about revenue concentration. TSMC has revenue concentration. Anduril too. Government-focused businesses, two-customer businesses β hundreds of billions in market cap. Also, data businesses will expand to enterprises. Every business needs its own AI model β every business needs its own data.
On Arena's own agent evals: Arena is one of the largest consumer AI apps in the world β bigger than XAI, Hugging Face, Manus, Gen Spark. 30M+ monthly visitors, mostly knowledge workers and pro-sumers doing real daily tasks, giving feedback that powers our evaluations. Organic flywheel for agentic evaluations.
π νκ΅μ΄ μμ½ βAlex Cobb said every large American enterprise is terrified of working with Frontier Labs. True or exaggerated?
Absolutely true. And they're also terrified of working with Chinese open source. I was talking with a Fortune 50 enterprise yesterday β they asked if we use Quen in our stack, and if we could stop using it and switch to an American model.
Harry: When your customer becomes your competitor. Claude Design eating Figma. Rumors of Anthropic doing a legal product to kill Harvey and Lora.
Every business in America is shaking in their boots. Multi-billion dollar companies get calls from their biggest customer: "Hey, OpenAI is getting into this game, we want to work with them because they're more AI-forward. Goodbye." If inference commoditizes, the next best thing is for model providers to move up the application layer.
GTM-heavy businesses have more defense. Harvey/Lora β you go into Kirkland or Clifford Chance, build relationships with 50-year-old white male partners who want to play golf, then deploy to junior lawyers who don't want to use you because they think you'll take their jobs. Deployment and GTM are the heavy lifting. Figma-style products (designer picks it up and uses it) are more exposed.
Salesforce, ServiceNow β deep entrenchment. SaaS apocalypse is overstated. Network effects and data moats are hard to replicate. Wix β less integrated, less sticky β tougher.
π νκ΅μ΄ μμ½ βHarry: We thought inference would get cheaper and it hasn't. Why?
Long run, the market will be efficient. Right now Anthropic has disgustingly high gross margins on their inference. After they go public, the whole world sees their margins β and that exerts downward pricing pressure. Currently you negotiate blind. Once information is public, you know they can discount, and negotiating becomes easier.
Harry: But great businesses have pricing power (Palantir, Apple).
True β Apple can charge what it wants. Some AI businesses will keep pricing power. But Anthropic IPO β I predict they go out first. Everyone likes free cash flow, and Anthropic is generating it. OpenAI reportedly isn't ready to IPO this year. Anthropic could come as soon as October.
Big risk: open models beating Opus 5 or Fable squarely across all categories.
Margins matter. Fireworks-style businesses at mid-30s margins vs software at 65-80%. Profit = margin Γ volume. The bigger problem is GMV businesses reselling tokens or GPUs. Terminal value question: why should I pay you more than the cost of electricity to run those GPUs? Short-term margins can be low, but the terminal margin structure must support a public business.
Quick fire: What changed your mind in 12 months? Open source model leadership β moving much faster than I initially thought. Anthropic also moving faster.
Biggest lesson running Arena? Managing people. Focus. So many experiments I shouldn't have wasted time with. Do one, maybe two things extraordinarily well.
First to $10T β Nvidia, OpenAI, or Anthropic? Hard to say not Nvidia. Enterprise adoption of AI will be another 10x for the industry. Market hasn't priced it in yet.
π νκ΅μ΄ μμ½ βHarry: Debt cycle around compute buildout concern you?
Yes. If open source makes cost-savings salient and decreases enterprise revenue for OpenAI/Anthropic, it could lead to insolvency. Never had such reliance on two companies to hit their targets. If OpenAI/Anthropic music goes off, no party for fireworks layer, no party for routing layer β everyone gets the wind knocked out.
Where is the industry under-hyped? Mechanical infrastructure for compute and data centers β actual cooling systems, steel infrastructure. HBM and GPUs are ultra-hyped.
South Korea's stock market crashed ~40% β national convene called. Just realization that everything was overinflated, markets can't rip that long. No destabilizing factor in open/closed suggesting demand is being questioned.
Most underrated Neolab? Most are overrated. Black Forest Labs is pretty underrated.
Most excited about next 5-10 years? AI in medicine. Level of human flourishing as we start to eradicate diseases one by one, the same way we're currently eradicating open problems in math. Math is a closed system β for medicine, you need fast feedback loops on biological systems. Once we crack that, extraordinary journey.
"You know what's missing? That is exactly the data layer. The GPUs are the same GPUs in both cases. The problem is the data infrastructure, the flywheel, the data collection that you need in order to build a great biology or medicine product."