We are in the capital intensive game and we competing with the most capitalized companies in the world. Our program this year is 2025 billion. Our competitors hyperscalers have eight times bigger. The AI infrastructure race is on. Capex spend has never been greater. At the center of this, Nebus. Today I'm joined by the co-founder of Nebius, a company that has scaled to a $66 billion market cap, going head-to-head with some of the largest hyperscalers in the world.
Roman: In the next 6 months, the capital cannot help. 6 months is too short time. You have what you have, you need to deliver. The main threat for Nebios as a business is the world will be too much consolidated. Today, we uncover the AI infrastructure bubble and so much more. It's like a shark. You're alive when you move, right? So, we have to move.
Harry: And I'm thrilled to welcome Roman Chernin, who has power against Nvidia. Roman, I am so excited for this, dude. I think Nebius is one of the most unbelievable, incredible stories in terms of what we've seen over the past few years, but also like, holy [ __ ] what an exciting few years we have ahead. So thank you so much for agreeing to do the show.
Roman: Yeah, thank you for inviting and uh glad to be here.
Harry: Now I would love to start with a question that I think is at the top of a lot of people's minds β where are we at on AI infrastructure? A lot of people are seeing the capital going "oh it's a bubble" and a lot of people are going "it's just the start." How do you think about whether we're at an AI infrastructure bubble moment right now?
No, I don't believe it's a bubble. Define the bubble. Do I believe that we will need tens or hundreds times more to build? I thoroughly believe β I'm probably biased, I would probably not be in this business if I didn't believe. We are just at the beginning of this amazing moment when Jensen calls it "useful AI." We have maybe one use case that works β coding β and that started working maybe a few months ago.
So let's put it in perspective: we are just a few months from the moment when we got maybe the first use case that works at scale. If you take any company in the world today and look at their AI adoption, you will actually see they start using AI in the first percent of the volume, in the first percent of the use cases. Even pretty advanced companies, technologically β they are just starting. And I'm taking from that that we are only beginning.
Harry: So we're completely aligned, but it's a boring discussion if I just agree with you on everything. My question on the back of "hey, we've seen coding work for the last 6 to 12 months" β yes, but there is a question that we will move to open source models locally hosted because the cost will be too significant for some of these enterprises to burden. If we do that, isn't that damaging to OpenAI/Anthropic and to Nebius?
Roman: First of all, it's not in the future, it's already in the present. When our customer or product builder gets to scale, they start looking to improve the economics or accelerate growth. And that's where most of them start to look to alternative models. The best way to build today is on frontier models from OpenAI, Anthropic, Google β they provide the best capabilities in the world. But once you figure out the use case, when you see adoption, when you see the customer data loop, you can find a higher-quality way to serve the same use case. You don't need the best universal model in the world β you can create a specialized model that in your particular case will work even better.
Why doesn't it hurt Anthropic and OpenAI? Because in reality they move to the next frontier. There are so many unsolved tasks. Every time we find a way to solve some task more efficiently, we just start solving more complex tasks. We saw it with DeepSeek a year ago, and we continuously see it. You always push the frontier. You always have more complex tasks to figure out how to solve.
Nebius stock went down 40% in one week β February 2025. The same exact week we probably had the best sales week. People on the market were concerned that AI is so much cheaper, maybe it's a bubble. But we never had a better commercial week β so many people figured out they can run inference in production with DeepSeek and the economics work. Every time we get the same unit of intelligence cheaper, we are not reducing consumption β we are increasing it.
Harry: Speaking of Jevons paradox and producing more yielding more demand β where are you not moving fast today where you would like to be moving faster?
Roman: Everywhere. When we think about how we build a company we talk about it in four dimensions. Capacity β how many megawatts, gigawatts, GPUs we deploy. We are an infrastructure company. We need to be large. If you're not large enough, nobody needs us to exist. The team is doing an amazing job, but it's never enough β there are a lot of complications of the real world that prevent you to move fast enough. Supply chain, regulatory, fires and waters β everything that happens in the real world.
Another dimension is product. You want to move fast enough to address new types of workloads, new types of customers coming to the market. We started as an industry from people who first built the models β OpenAI, hyperscalers, large labs. What they need from you is barely compute. Just throw infrastructure β we see a lot of these large bare-metal deals on the market. But this is only the first layer.
I think this is something people underestimate β how important it is to build the foundation for improvements and experimentation. When you understand as a team what is good for you, because you close some use case and it works, but then you want to change the model β how do you know you don't ruin the quality? You need metrics, eval mechanisms, CI/CD for AI development. When customers like Revolut solve these foundational problems, they start growing exponentially.
Roman on Revolut: When we started working with them, 99% of their inference budget was in closed models (OpenAI). They started to crack some use cases that didn't work for them economically β they practically couldn't replace humans or enhance humans in those use cases. So they started moving to open source. But it didn't move fast because they had to build the entire engine internally β focusing on evaluations first. Once they solved the cold-start, they grow exponentially. Their AI budget is growing the same exponential trajectory as their revenue β same as AI-native companies report.
Harry: I always push back on people who proclaimed open source would be a credible threat to the largest model providers, because I said: the biggest enterprises want reliability, security, and ease. They don't want to be tinkering around with all the architecture beneath the surface. What you're telling me is you're able to be all of that to allow them to pipe away from those providers and have a cheaper better experience because you take away the plumbing.
Roman: Correct. Yes. But my point is β closed models with open source models, it's not about reliable or not reliable. It's about capabilities. Closed source frontier models are great and they will become even better β they will solve so many problems we don't solve yet. There will be a market for the smartest models in the world, the fastest models, the in-between (smart enough but cheap enough). You as a customer will be able to pick the right source of token for each particular task. And on the agentic layer, it may not even be the customer's task to choose β it will be the engine that knows all the capabilities underneath.
Roman: The first layer speaks in megawatts β if you read announcements someone signed a large deal with Meta or Microsoft or OpenAI, people speak megawatts. You deliver the megawatts of compute. Then the second layer β managed cloud β people speak GPU hours. This is the key unit. You sell the efficient hours spent on compute with storage, networking, complementary services. People still buy "managed compute."
The next layer we're working on is managed inference. People don't want to think in terms of GPU hours, they don't want to figure out B200s vs H200s vs B300s β what is better for a particular workload. They don't want to manage VLLM or SGLang, do all the optimizations themselves. Our product called Nebius Token Factory β this is a managed inference platform. New customers, mostly people we call vertical AI companies or enterprises. They don't do models, they build products on top. Now we speak in tokens. You don't pay for GPUs, you consume tokens.
But it's also not the final stage. People build agentic applications, agentic workflows. When you build an end-to-end agent, you may not even think in terms of the particular model β you want the end-to-end task to be efficiently executed and provide the expected outcome. The magic the platform can make is actually think for you which model to use in this particular call. Do you need to go to the smarter model, or you can ask two models lighter and get less smart tokens, and then have a judge model that chooses the best result? What size of context you should have, etc. This is the next layer.
Harry: So that fourth layer is a direct competitor to OpenRouter.
Roman: What we would love to bring on that level is the same as what we do on the layers below β the optimization engine. You can build your agent in many open source or proprietary tools, but when you need to scale it, you start thinking about economics, reliability, repeatable execution. It's not just a model choice problem, it's a system problem. You need to make it reliable, repeatable, and economically viable.
You built your great product on OpenAI. You cracked the use case. You don't have enough margin, or you want to apply your data and tune the behavior β you can't in the closed ecosystem. You go to Hugging Face, grab open source weights, take VLLM or SGLang to run it β and it doesn't work. You need optimizations, orchestration, caching, observability. You had all of that on OpenAI because it's a production service. Token Factory gives you managed inference with open source or specialized models, applies optimization techniques, manages the better economics. It's a service.
Roman on how token costs come down: You take the model and optimize it for particular scenarios. You can distill the model β make a smaller model with the same quality. You can do speculative decoding, optimize caching. Out of this model you actually build a system that in your particular case works with your requirements and optimized economics. And one of the things important for customers β the models are changing every week. Today maybe MiniMax 3 was released and there is Ultra announced. Every few weeks. You want flexibility. You want someone to support you on experimenting and adopting the new best models for your use case.
Roman on differentiation vs Coreweave: The principles we build on are full stack. Full stack down and full stack up. Full stack down β we're really deep in the physical world. We build data centers, racks, servers, the platform. When you control these things downstream, you can move faster and squeeze more cost. Full stack up β vertical integration upstream is how you follow customer needs and not be limited by the small population of people who just need infrastructure. Less concentration in the business, more diversified customer portfolio. We believe long-term better positioning for enterprises, where most demand will come from. Enterprises won't buy raw compute β they'll need platforms, tools, respect for their legacy, ability to work with complex environments.
Harry: If you had 10x the capacity today, what would be different? Could you sell it overnight?
Roman: Not overnight, but we would definitely have demand. The key question for us is not "do we have demand" but how we build a portfolio of demand. You can sell bare metal, managed infrastructure, inference, and maybe in the future some new layers. We try to build a diversified portfolio. The higher in the stack we move, the more value potentially we create β and the bigger the population of customers we can serve. On bare metal you have maybe a dozen customers in the world. On managed infrastructure, hundreds. On inference, thousands. On agentic, tens of thousands of new developers.
Harry: Where do you settle on revenue concentration with a Meta or a Microsoft?
Roman: It's the main question of our business. Long-term strategy of Nebius is to serve as much diversified a portfolio as possible. To serve a dozen customers like Meta/Microsoft β they are super advanced, have their entire software stack, they literally need only physical infrastructure. You have a tiny additional value to provide above the physical infrastructure. By the way, to satisfy them with what they need on physical is quite a challenge β they need the most scaled infrastructure in the world that exists. People say it's commodity but it's not really commodity at that scale. From day zero we were building this software stack because we thought it's much more beneficial.
Harry: Given insufficient supply of capacity today, if you doubled pricing would you see any change to demand?
We actually raised prices just a couple of months ago. And we still have fair pipeline pressure on supply. It's not only us being greedy β there is a point, especially in inference. Training is one-off cost, but inference is the cost of serving the customer. There is a level where the economics of our customers' products don't work. They are elastic to some extent. We want to be meaningful and thoughtful. And it's not only GPU-hour cost, it's all the optimizations β what we call TCO, total cost of ownership.
Roman: People are too obsessed about capacity, too obsessed about the nominal price of capacity. You can price GPU $3, $4, $5 β depending on the use case and quality of the platform, it can create completely different outcomes for the customer in real cost. How long it works, what is the effective uninterrupted time you can run there. If you talk about inference β how many tokens you can extract. These optimizations change the price of tokens in order of magnitude. If you build the platform and provide a high level of service, you can extract much more economics β not only from the infrastructure cost structure.
Harry on capacity bottlenecks: Gavin Baker said permitting and regulation and the delayed buildout of data centers has actually helped, because if I enabled you to build 10x the data centers today, it would actually create the glut.
Roman: In the next 6 months the capital cannot help β 6 months is too short, you have what you have, you need to deliver. In the next 12 months you can accelerate something. In 24 months you definitely can unlock so many things. We are not building one data center β we are building a portfolio of capacity. We secure power and land, then build data centers, then fulfill with GPUs. Every next stage requires more capital, but we do as much as possible in advance. The bottlenecks are different on different time spans.
Harry on public backlash: We're seeing more public angst toward AI. Eric Schmidt getting booed off stage. 40 out of 100 data centers not being built when they go through planning. How do you reflect on that?
Roman: This is the environment we work in. We think about it as a portfolio of projects β if one data center is delayed we will still deliver enough capacity. Most customers are not locked in one physical location. Then communities and local authorities require us to work closely with them, explain, address their concerns. It's like Uber when it started β there was pushback, people said "what's happening, it's moving too fast." Sometimes communities are not educated enough, sometimes they have rational concerns. About 70-75% of the new capacity we built mid-term is in the US, so we built a lot of presence on the ground.
Harry: A theme that came up when I spoke to other guests was the relationship with Nvidia. Is a marriage a marriage if one has more power than the other? How do you think about the power dynamics with Nvidia when they have so much power?
We look at this in a very simple manner. We just need to build what we build. Nvidia is still for a big extent an engineers-driven company. The best thing you can do to get respect from Nvidia β if engineers in Nvidia respect your engineers, you will have the right foundation for relations. We managed to prove again and again that we know what we build and we have a strong engineering team.
Harry on Europe: Europe doesn't have anywhere near the model buildout we've seen in the US and China. How important is it that nations have their own sovereign models?
Roman: The world is divided β we may not like it. Having good enough foundational models available for big parts of the world is important. In Europe we should think how we have enough capabilities available. The sovereign AI agenda has been too concentrated around megawatts and power, rather than the builder layer. Megawatts will come β companies like us will build infrastructure if we have demand. Demand comes from builders. We need to care about having more great companies like Lovable, Black Forest Labs, Mistral, and enough people that invest in research and products. They create the demand, the flywheel.
Harry: Where is the most interesting area to invest today? Infrastructure, horizontal model, vertical model, application layer?
Roman: We built infrastructure, so we're happy here. Even though for some extent we are building the easiest part β not that execution is easy, but we kind of know what's needed and our customers help us understand. I think the most amazing people in this industry are those who take the risk to go and build end-user products. They drive most of the growth. People who take the real risk of building something people would need or not need β they are the heroes of our AI journey.
Roman on the hardest part of "just doing your job": Four dimensions. Build scale, build product, work with customers, manage capital. We are in the field business. Cloud is a post-sales business. When you sell, you sell the promise β then you need to satisfy the customer. Working with customers, having a strong customer-facing engineering team (FDE) β go talk to your customers, make sure they know you, you know them. And the most boring but also most exciting is capital. We are in the capital-intensive game competing with the most capitalized companies in the world.
Harry: If I gave you unlimited budget, what would you do differently?
Roman: Build faster. Data centers and fulfill them with GPUs. Our CAPEX program this year is 2025 billion. Our competitors hyperscalers have like 8 times bigger. If I had 10 times bigger capital, I would just build more data centers and serve more customers.
Harry: Data centers on planet Earth is a very difficult logistical buildout. Data centers in space β is that nuts?
Roman: I think everything we see is nuts. So many smart people now working to make it happen β most likely I may be less pessimistic. I don't know β we may build more in space than on Earth in 3 years. So many smart people are trying to solve this and bring compute to space β why wouldn't I believe it will happen? There are still a lot of challenges. But if someone told us 3 years ago that we would build multi-gigawatt data centers with large interconnected compute clusters, I wouldn't have believed it. And we are here β it's routine.
Harry quickfire β what job doesn't exist today that will be common in 5 years?
Roman: One thing that's happening is we are democratizing what people call being a developer. Each of us can be a developer β converting an idea into a digital asset. I hope that giving millions, tens of millions of new people the ability to convert their idea into something that works very easily, we will see a lot of new businesses and ideas just come to life. They will create new work we don't even think exists. What's also challenging is how education will change. When everybody has access to intelligence, what should people learn? You don't need them to learn the facts β everything is available. How do you train people to think when they don't need to think so much? How to teach people to continuously change?
Harry on Roman's two teenage daughters: What do you advise them as they enter the workforce in the next 10 years?
Two things will be needed. One is being able to communicate with people, with empathy β understand humans, communicate with humans, being empathic. The second is creativity β all the art. 10 years ago I thought the most important thing they need to learn is math and engineering. Now I'm far from this belief. I'm quite happy they are much more into soft skills than I was when I was a kid.
Harry: Finish this sentence β the biggest threat to Nebius is not competition butβ¦
Roman: Consolidation in general. The main threat for Nebius as a business is the world will be too consolidated. We try to be diversified β solve problems of different customers on different layers. If you end up in a world where 3-5 super-models, super-companies, super-empires control the world, then Nebius or companies like Nebius will only be needed to serve their needs on the physical layer. The more democratized, the more diversified the world, the more we are needed.
Harry: Do you think that's likely? We're seeing the concentration of value to fewer and fewer players. The opposite of diversification.
Roman: I hope it will not happen. I think it's better for us as humans as well. The world will remain quite diversified in different manners β I'm optimistic. So many people want to build something independently. There is a lot of need to try things and build new things β that organically creates pressure for a more diversified world.
Harry on Aschenbrenner: Leopold Aschenbrenner recently disclosed a 5.3% position in Nebius β about 15% of his portfolio. How do you sit internally? Are you like "yeah, go Leo"?
Roman: We noticed it β everybody noticed it. The stock jumped, it was big news. We take it as justification of what we do. Those people give you credit that you will execute. I come back again and again to: what we do is a post-sales business. Every time we sign a deal, every time someone invests in us, they give us credit and opportunity to deliver. Then go back to your job and deliver. We are in an emotional market β you should keep yourself down to the ground. Remember that all this growth and credit, it's an opportunity to deliver.
Harry: You're such an Israeli. Americans would be like "yeah, go!" You're likeβ¦
Roman: I think I'm Russian in this way. Russians always know you need to look very pragmatically. Russians always with this face like always expect something will happen and you need to be ready. It's a really important part that comes from our CEO and founder. You wake up β it's a new customer, new day, you need to deliver. Nothing is guaranteed. You need to concentrate on the work. We could celebrate a little bit more β we just don't have time. Just give the team more respect for how much is done. It was not easy, it's still not easy, it will not be easy. But we cannot stop. It's like a shark β you're alive when you move. So we have to move.
Harry: On that note, I cannot thank you enough for joining me. You've been fantastic, Roman. Really huge thank you.