The a16z Show

Fei-Fei Li on Spatial Intelligence and Robotics

43 min
Jul 28, 202628 days ago
Listen to Episode
Summary

Fei-Fei Li and Yun-Ju Li discuss World Labs' acquisition of Scenix, a robotics simulation company, to advance spatial intelligence and world models for robotics. They explain how AI systems can learn to perceive, reason about, and act within physical spaces through a real-to-sim-to-real pipeline, positioning simulation and generative models as critical infrastructure for scaling robot learning and evaluation.

Insights
  • Simulation is not an alternative to real-world data but a complementary tool enabling counterfactual reasoning and systematic coverage of scenarios that real-world data collection cannot efficiently provide
  • The robotics industry faces a fundamental data scarcity problem unlike language models, requiring new approaches to scaling laws through synthetic data generation and simulation-based training
  • World models that maintain consistency across space, time, viewpoints, and interactions are essential for robotics, differentiating this approach from video-only prediction models
  • Robotics adoption will progress from fully structured (factories) to semi-structured (warehouses) to unstructured environments, making specialized embodiments more pragmatic than humanoids in the near term
  • Foundation models for robotics must be multimodal and action-conditioned, treating actions as both inputs (predicting environment changes) and outputs (policy generation)
Trends
Shift from video-only models to 3D-consistent world models as the foundation for robot learning and controlSimulation-driven robotics development becoming mainstream with emphasis on systematic randomization and domain randomization for robustnessRobotics companies increasingly adopting foundation model approaches with fine-tuning for specific embodiments and tasks rather than building proprietary end-to-end systemsReal-to-sim-to-real pipelines emerging as standard methodology for accelerating robot development cycles and reducing real-world data collection costsPragmatic, customer-driven robotics development in semi-structured industrial environments (warehouses, assembly) preceding general-purpose humanoid deploymentIntegration of generative AI (image-to-3D) with robotics simulation to enable efficient digital environment creation and scalingEmphasis on evaluation infrastructure and benchmarking in robotics as critical bottleneck, mirroring language model development practicesMulti-modal, action-conditioned foundation models becoming the backbone for robot policy learning across different embodimentsBi-coastal robotics development with remote collaboration capabilities becoming operational necessity for distributed teamsEconomic viability of robotics automation focused on high-frequency, repetitive tasks (cleaning, assembly) rather than general-purpose humanoid applications
Companies
World Labs
AI startup building spatial intelligence and world models; acquired Scenix; developing Marble generative model for 3D...
Scenix
Robotics simulation company acquired by World Labs; develops real-to-sim-to-real pipeline for robot training and eval...
Amazon
Referenced for warehouse automation use cases and as prior employer of Scenix co-founder Sonny Hu in computer vision
Weta
VFX company where Scenix co-founder Changxi Zhen worked before transitioning to robotics and simulation
Tencent
Technology company where Scenix co-founder Changxi Zhen worked in simulation and VFX
Waymo
Self-driving car company referenced for using billions of hours of simulation in autonomous vehicle development
Columbia University
Institution where Scenix co-founder Yun-Ju Li is an assistant professor conducting robotics research
Stanford University
Where Fei-Fei Li conducted postdoctoral research with Yun-Ju Li on robotics and world models
MIT
Institution where Yun-Ju Li completed PhD research in robotics and perception
People
Fei-Fei Li
Discusses spatial intelligence vision, world models, and robotics strategy; explains acquisition rationale and techni...
Yun-Ju Li
Explains real-to-sim-to-real pipeline, robotics challenges, and practical approach to robot development in real envir...
Martin Cassato
Conducts interview, asks clarifying questions about technical approaches and business strategy
Changxi Zhen
Columbia professor with VFX and simulation expertise; brings world-class simulation capabilities to robotics platform
Sonny Hu
Phenomenal engineering leader with Amazon and startup experience; leads technical implementation of robotics stack
Sergey Levine
Referenced for research on simulation-to-real transfer and importance of real-world data collection in robotics
Quotes
"Language models transformed how AI understands words. The next frontier is teaching AI to understand and act within the physical world."
Episode SummaryIntroduction
"My north star is I want the robot to work in the real environment. I'm a very practical person."
Yun-Ju LiMid-episode
"Think about human intelligence. We do a lot of simulation in our head. You know why? There's a very important role simulation plays that real-world data doesn't play, which is counterfactual reasoning."
Fei-Fei LiMid-episode
"We are building a consistent world. Consistency both over space, over time, over different viewpoints, and over different type of interactions."
Yun-Ju LiMid-episode
"The hardest thing in today's AI is to have the right measured optimism."
Martin CassatoLate episode
Full Transcript
We are building the next frontier of AI, which is what we call spatial intelligence. At Cynix, we are developing what we call a real-to-sim-to-real pipeline. We can replace all the data, all the evaluation we need in the real environment by using the data that can generate at a scalable way in our digital world. Think about human intelligence. We do a lot of simulation in our head. You know why? There's a very important role simulation plays that real world data doesn't play, which is counterfactual reasoning. What we are building is a consistent world. Consistence both over space, over time, over different viewpoints, and over different type of interactions. My north star is I want the robot to work. The world we live in can be multiverse, that we create technology to allow people, builders, developers to act within different spaces. Do you believe we'll ever be able to build robots that have the power efficiency of a human being? How far away are we from this? Is this like five years or this is like never? The TLDR is... Language models transformed how AI understands words. The next frontier is teaching AI to understand and act within the physical world. Following World Lab's acquisition of Scenex, Martin Cassato sits down with Fei-Fei Li and Yun-Ju Li to unpack the vision behind the deal. They discuss spatial intelligence, world models, simulation, and why solving robotics will require a new generation of AI built for three-dimensional reasoning, not just language. All right, well, it's great to have you both here. So Fei-Fei, for the listeners that may not have the background, maybe you can give an overview of what World Labs does. Yeah, well, World Labs is a two-year-old startup. I think we should just recognize it's a frontier model lab. We are building the next frontier of AI, which is what we call spatial intelligence. and spatial intelligence is about creating AI that has the ability to generate, understand, reason with, and interact with spaces, whether it's physical or virtual. And of course, a means to an end towards spatial intelligence is building large world models. And that's what World Labs is mostly focused on. Yeah, so you've been saying this since the very beginning, which is the machine's ability to perceive and reason about spaces and act on spaces. But I always had the assumption that the acting on spaces was some long-distance future thing, but now you're acquiring a robotics company. And so maybe talk a little bit about the timeliness of this and the intentions. Yeah. So first of all, it doesn't just take robotics to act within spaces or to interact, right? I mean, look at the creative field, whether it's VFX or gaming or design. Many use cases, you can create and act within virtual spaces. And WorldLabs' thesis has always been that the world we live in can be multiverse, that we create technology to allow people, builders, developers to act within different spaces. Having said that, the ability to act within the physical space is one of the most exciting and most profoundly important capability of the future AI world. So robotics is very much that. So WorldLab has always believed that robotics is an important application as well as use case of spatial intelligence and world modeling. So by joining for us with inviting Scenics and Scenics team to World Labs is part of our long-term vision and mission. We've always committed to that. Amazing. So Yunju, you're the co-founder of Scenics. So maybe provide everyone with a quick overview of your background and what Scenics does. Yeah. So I'm Yunju. So I'm currently co-founder of Scenics and also assistant professor at Columbia University. So my research started from my PhD at MIT and then postdoc with Fei-Pei. Really? Yes. That's great. The world is small. The world is small. It is. Throughout my career, my goal has been very simple, trying to help the robots better perceive and interact with the physical world. So I'm a very practical person. I want my robot to work in the real physical environments. So for Cinex, the unique opportunity we see is that there has been a lot of bottlenecks. Right now, we see faced by the development of general purpose robots, especially around training and also around evaluations. So at Cynics, we are developing what we call a real-to-sim-to-real pipeline. We're going to map the real environments into the digital world that has the best alignments with the real environments. By alignments, we mean that whatever happens in the digital world is also going to happen in the real environment, such that we can replace all the data, all the evaluation we need in the real environment by using the data that can generate at a scalable way in our digital world. So that is how everything started. In Cinex, we put together a very, very strong and best teams around robotics, robot learning, and also simulation and rendering, trying to build this real-to-Cine2real stack to solve some of the key bottlenecks. It's amazing that you two work together. Yeah, and there is a funny story here because you would think because we work together, he was my amazing post-op. We've been talking about this Cynix and WorldLab integration for a long time. It's actually not true. They came into WorldLab as a customer. Really? When we released the first version of our generative model called Marble last winter around November, December, Cynix just signed up. No kidding as a customer. Yes. And I didn't even know what it was. And then I realized this is Yunzhu's company. I called Yunzhu, I'm like, wow, this is your company. And then we realized there's so much synergy. Maybe, Faith, could just quickly describe what Marble is? Yeah, Marble is the codename for the base model that WorldLab is being training and iterating on. The fundamental capability right now of Marble that is publicly released is to take a prompt. It can be an image, it can be a few images or a text, and turn that into a geometrically consistent world that can be represented in 3D geometry, whether it's Gaussian splat or mesh. Really what Cynic's team is doing is trying to solve this extremely difficult problem in robotics, which is the lack of data. The lack of data in training, the lack of data in evaluation, this is very, very different from language models, where data is abundant on the Internet. And we know that in order for robotics to work, we have to somehow unlock the power of scaling law. But where does that come from? This is something that is a profound problem that everybody is battling with in robotics. It'd actually be great to talk about this energy. you have put together a very, very talented team. You have put together a very talented team. And so to what extent is there overlap? To what extent is this an extension? Maybe talk a little bit about that. Yeah, that's actually. How complementary it is. It's actually, the TLDR is very complementary and with a shared mission. So Yunru is one of the three technical co-founders. The other two are Changxi Zhen, another Columbia professor who has been a world-class technologist in simulation. Oh, wow. And Changxi has his background in also VFX. He worked at Weta. He worked at Tencent. He's been an entrepreneur. Then there's Sonny Hu, who is a phenomenal engineering leader who was also in a startup that was acquired by Amazon many years ago. So he worked in many different tech stacks in the computer vision field in Amazon. So when we started talking more seriously, I recognized that a couple of things that Scenics has from a talent point of view is extremely complementary to world labs. One is obviously Andrew's incredible thought leadership and just technical prowess in robotics, right? So from really, from hardware, full stack robotics. And even when he was my postdoc at Stanford, at that time, you already had your faculty offered. So you were there only for one year. I wanted you for more than one year, but he had to go become a, have the real job. So he was a full stack researcher in robotics from modeling to hardware. And, of course, Yunzhu and his students at Cynics was that pool of talent WorldLab hasn't had yet. Then on the Changxi side is just incredible simulation capability, right? He's such a senior researcher and technologist in simulation. And what WorldLab is doing is very much interfacing the world of simulation. So I think what they don't have, obviously, is on the generative model side, as well as the computer vision 3D reconstruction side, we're also very strong at World Labs. So that's a technology that Cynix needs. So together, these two sides come together and make it much more complete. Fei-Fei's motivation in this is like this is an extension and a compliment to get into robotics. You know, having been in your situation, which is deciding when to sell a company, it would be great to hear from you on like how you think about joining World Labs and kind of the fit there and like why you made the decision to do it. Yeah. So at the very beginning, we were deciding, okay, do you want to just keep going? But after chatting with Fei-Fei, after seeing all the synergies that are happening in the middle, it just makes perfect sense for the forces to join each other. So in any sense, as Cynics, what we've been doing is real-to-sim-to-real. is to do dense reconstruction of the environment. So we capture the appearance of the environment, geometry of the environment, and also the dynamics of the environment, meaning how the environment is going to change when you apply actions. So this dense reconstruction right now is still a little bit on the heavier side. And what World Labs right now has been doing involves a lot of profound capabilities around sparse reconstruction and generations So we see a lot of opportunities of leveraging like a marble and other like capabilities as World Labs in order to do very efficient reconstructions and modeling of the environment So can we expect a foundation model for robotics from World Labs? World Lab is building a foundation model, as you know, Martin. We're building a base model, and as the technology has been evolving, some of the most exciting base models are omni models, right? They take multimodal input. They have multimodal outputs. And what is a foundation model for robotics? It's very likely going to involve actions. it's very likely going to involve the output of actions in addition to the state of the world. And we're definitely not ruling this out. Yeah, great. So for example, for the foundation models, it essentially needs to be a multi-modal model. So it has to take into account frame text, image, depth, and different kind of modalities. And action is a very, very important part of that modality. So if you think about frame actions as an input, that essentially affords similarity. that is going to predict how the environment is going to change when you apply a specific action. When the action is output, this is essentially a policy model that is trying to predict, giving a specific goal, like what should be the action you take in the real environment to get you closer to that goal. So this kind of omni models actually can benefit a lot and actually provide huge amount of values for the robotics communities in trying to understand how to model the environments and at the same time how to act in the environments. And this can also act as a backbone for you to fine-tune into specific robotic applications to making sure it's really live up to the reliability and efficiency that's expected by the clients. You know, Yun-Chi, if you don't mind a kind of a lay investor question, I see a lot of robotics companies. And a very popular approach right now for the robotics companies that come in is like, we'll use a video model, you know. And like, you know, that's the predominant method where this is, you know, 3D and simulation. It's a very different approach. And so maybe you could contrast, you know, this popular approach of just using video only versus kind of what the ambition here is. Yeah. So in order to create words with the robot handler, the words, as I mentioned, need to capture the essential structure of the problem. And one of the very important necessary requirements for those words will be consistency. So that is where I actually see there's very, very strong synergies with Marble, because what we are building is a consistent world. Consistence both over space, over time, over different viewpoints, and over different type of interactions. And Marble, the generated world from Marble, also provides an infrastructure, a component of that entire world that we believe is necessary for the robot owner. Imagine if a robot pushes an object forward, the object just magically disappears, which has been a problem. or many of the existing video prediction models. It's one provides good enough signal for the robot to know like what is the right thing to do. But obviously right now, there has been a lot of investigation on building better and better and stronger and stronger like video models. So we actually see a way where some of the infrastructure we build can provide as initial momentums and to go in through this data flywheel of going from this like a more simulation driven models into like a robot policy models, which is going to do the execution in the real environment, collecting new data, the data will come back in. Where the model doesn't necessarily have to be physics only or learning only, but somewhere in the middle, which will be able to capture the essential structure of the problem, but at the same time, be able to scale and become better and better as you accumulate more data. You know, I've worked now, Seifei, very closely for a while, and you've always had this north star, which has driven this, and you've articulated variously as kind of 3D and in a number of other ways. And I'm just wondering, for you, is there also a similar philosophical North Star or you're more the pragmatic, like I am, like build the system, like do the thing? My North Star is to make robots work in the real environment. I'm a very practical person. I want the robot to work. One interesting thing that's actually coming from my collaborations with Phoebe during my postdoc, we are building this kind of benchmark. We actually send out surveys asking the general public what they want their robots to do for them. Among the southern tasks we collected, one-third of the tasks are about cleaning. People just don't like to do those dull and dirty tasks. And those are the scenarios we really want to make sure we have robotic solutions to deal with. One thing I really like about Cinex, Martin, especially continuing your question, there's a lot of robotics companies building models and all that. One thing I truly like about Scenics is Yunju and his co-founders have such an incredibly pragmatic approach to robotics. Especially they come from academia, right? Sunny doesn't, but Yunju and Chanxi come from academia. but their first instinct is work with design partners and customers in real industry, whether it's industry labs or warehouses or electronics assembly. That is such a refreshing, actually, a refreshing way of approaching robotics. And that really made me very excited to work with them. Maybe this is for you and you, but I'll just be, this is personal curiosity, which is, it seems to me that for robotics, you have to be pretty exact. I mean, not perfect, but pretty close. But for the creative use cases, which WorldEps has done a lot of, you kind of don't need to because, you know, I mean, even sometimes like being wrong is stylistic or intentional or whatever. And so from a technical perspective, what is the challenge here for reconciling these two things? Or do they never get reconciled? Like, there'll always be two points in the design space. So they will be, like, reconciled in the long terms, of course. And modeling of the environment doesn't have to be perfect. The model doesn't have to be perfect in robotics. And by the way, again, this is pure curiosity, but is there, like, a bit more formal way to say that? Like, what does that mean not to be perfect? It has to be pretty close. So let me put it this way. For example, models over the developments of all different kinds of robotic applications has been a very important cornerstone. If you look at all the existing robotic applications, like plane, drones, Roomba, or even for quadruped robots, bipedal robots, model has been the way for them to actually work and be able to transfer from simulation to the real robots. But if you look at those locomotion robots, like quadruped robots, bipedal robots, they can work in on snows, they can work in on bushes, but you don't need to have a simulator. They can simulate all the bushes and snows very precisely. You need to have a simulation that captures the essential structure of the problem and do a whole different kind of randomizations inside the digital environments. So that is what we're aiming for. So basically, with Synix and together with WordLabs, we're trying to investigate what is the level of fidelity we need to model the massive, massive worlds besides the robots, such that we'll be able to transfer the robotic systems training the simulated environment and digital worlds back into the real scenarios. As an investor, I've heard other researchers say, like Sergey Levine, say simulation will always eventually deviate from the physical world and real-world data collection is absolutely critical. And so maybe talk a little bit about the viability of this approach where simulation is a cornerstone as opposed to some other approach. So they don't contradict with each other. So if you're thinking about the simulation, simulation is essentially trying to predict how the environment is going to change when you apply the actions. And this is essentially a model of the world. It doesn't necessarily have to be pure physics. It can be a combination between both physics and also learning. We are collecting real-world data. We will be using those real-world data. It's just at different stages of this data flywheel. Maybe at the very beginning, we have stronger emphasis on we have more physics to making sure we have the right consistency and right structure for us to learn the world, for us to train the robot policies. But as we accumulate more and more data, both through data collection and also through the collaboration with our clients, we'll have the data that will be moving towards more learning-based, like modeling of the environments. So this kind of transition and also this kind of data fly-off is really enabling factors of both getting the best of both physics and geometry and consistency, as well as all the power and magics from the data and compute. I want to add to this and be slightly philosophical here is there isn't a binary choice between simulation or no simulation. All this come in together to make robotics work. Think about human intelligence. We do a lot of simulation in our head. You know why? there's a very important role simulation plays that real-world data doesn't play, which is counterfactual reasoning, is that you play out events that hasn't happened or cannot happen, or you don't have enough data to make it happen in real world. And while you play it out, you learn how to act in it. Humans do this all the time. We probably don't, you know, we just, I know you were at the World Cups. I was at the World Cup. Congratulations. to Spain winning. I'm sure in the planning of every game, there is simulation, whether it's digital or on the whiteboard or whatever. That simulation, the role simulation plays is counterfactual reasoning. And that's really important in robotics because we just do not have, cannot possibly have enough real world data for that. Here's a real life example, the industry of self-driving cars. Weibo has officially said they use billions of hours of simulation. And actually, Weibo is more simulation heavy than just real-world data heavy. So these are real examples. And as you know, Martin and Yunju, too, cars are the simplest kind of robots. Yeah, TD. Yeah. So clearly simulation plays a huge role in robotic learning I also want to add to that So like if you put things more specific simulation can provide two levels of benefits The first one is reliability, and the second one is efficiency. So for reliability, if you're thinking about a robotic system working reliably in the real environment, you need data to provide systematic coverage of all the state space and the variations that robots might encounter. That's how you can learn how that is robust. So with simulation, you can do systematic randomizations and control and variations of lighting, frictions, geometries, object types, and also all different kinds of physical parameters to making sure you have sufficient coverage of the state space. So this is what can give the robotic systems reliability. And second is about efficiency. So right now, many people are doing teleoperation. And if you look at many of the teleoperation devices, imagining all the actual skeletons you are using, you are actually collecting the data at a speed that is actually slower than humans actually doing the task. But for many of our clients, human speed to them is not good enough. They want faster than human speed. So for the robot to move faster, it's not as simple as just drives the robot faster because the gravity doesn't change. But in simulation, you can do systematic speed-up of the robot's behaviors to train the robots such that it considers all the dynamics, changes of the environment. So this is what can give our clients, for example, efficiency. So both for the reliability and efficiency, there are some kind of very unique values where simulation can provide. You've talked about the technology and the platform, what it does. Maybe talk about the specific use cases people use it for. There are essential two specific use cases, especially around both training and also around evaluations. Starting from the evaluations. So evaluation is something people often overlook in the robotics. But if you are trimming robotic models, you have to know how well it works. And that is the only source of information for you to iterate. By the way, every AI person really understands what evals are and uses it all the time. Non-AI people, it often means something a little different. So maybe it's even worth just describing specifically what you mean by evaluation. Okay, so what I mean by evaluation is you'll be able to understand for this specific checkpoint, how well does it perform? Does it perform, for example, 95% of the time or 99.9% of the time? And the key criteria people use in industry is how long does it take? How long in work clock time does it take for you to distinguish between a checkpoint that is 90% from a checkpoint that is 92 points? And if you only do that in the real environment, it's just take so long for you to do the distinguishments. And if you really think about also the robotic evaluations right now people are doing in the real environments, the iteration speeds is multiple orders of magnitude slower than iterations of those language models. So not only is the robotic tasks very varied, very diverse. Oh, you actually have to do the thing. Yeah, right, right. Atoms have to move through space. Exactly. But only the laws of physics have to be obeyed. And have you watched those robotics videos? Every video has like 10x, 8x because it moves so slowly. Exactly. So not only is it slow, it's dangerous, it's costly, but at the same time, the speed is also like multiple hours of magnitude, it's like slower. So some of our clients actually need this digital environment that can be used to evaluate their robotic systems. And because our digital environment has proven alignments with the real world, so meaning whatever happens in the sim is also likely to happen in the real environment. If a checkpoint is working better in the simulation, it's also highly likely to also work better in the real environment, as we have also been discussed in the blog post. So that actually gives our clients very strong confidence in actually using the data, using the signal from the digital environment to do scalable, safe, and much faster evaluations of their robotic systems. Great. So that is on the evaluation. Then on the training. So on the training side, so basically, like I also mentioned, it's about controllability. So you want to control all the different possible variations of states, parameters, lighting, frictions, physical parameters, like even object geometry, object types. So you want to make sure you have sufficient coverage of all different kinds of scenarios, such that you will be able to generate informative data for your robots to be robust. And this is just going to be so hard to do just in the real environment. Like we discussed, if you do teleoperation, the speed at which you are collecting data is slow. You're also limited by how many robots you have, how many teleoperation devices you have. There's a whole different kind of challenges around all the data operations around it. But in simulation, everything can be controllable, everything can be systematic, and everything can be understood at a level where you know exactly and making claims about exactly what distribution you have covered. to develop confidence about within the distribution, we know the robot will work. So those kind of confidence and efficiency and scalability is something that our clients also value to use our digital words for the training of robotic systems. Here's a crazy thing. Even before Cynix and we are talking, our inbound customers from Marble were already seeing this kind of demands. We just cannot serve these customers, but we are already getting a lot of phone calls from robotics early stage robotics companies who are developing their models all the way to downstream very pragmatic use cases and we're seeing these needs. When people hear you're going into robotics what they're going to envision is you're pulling out a 3D printer and you're going to be making hardware and then you're going to be programming the brain of a robot and sticking it in the robot and then you've got a robot. And I don't think that's what you guys are talking about here. So maybe talk about where this fits in the life cycle of creating a robot and where you will end and where the rest of the ecosystem will begin. So what we've been building, you can imagine, is an infrastructure with the softwares around these infrastructures for people to, for them, build worlds such that robots can learn and evaluate. And these infrastructures is naturally model agnostic and embodiment agnostic. So I just want to be very clear, just because this is actually a very subtle, I mean, for you it's obvious, but it's a very subtle point, which is from what you said, that's not building a robot. It's building an environment which another company can place their robot brain to navigate and to learn. Yeah. So for our customers right now, they have all different kinds of robots. Some are using, for example, single robot arms. Some are using Bi-Mail. Some are using a fixed arm. Some are using like mobile manipulators. Some are using grippers. Some are using some more elaborate versions of the only factors. So our platform right now is just naturally embodiment agnostic. We can very easily integrate different kinds of robotic embodiments, be able to put them into the worlds we generated, we digitalized, such that we will be able to give those individual robots capabilities of doing the right tasks and the right levels of reliability and efficiency in the real environments. And we are also, for example, model agnostic. So we can just using the data generated by our words to train different models, either from scratch or doing post-training of existing foundation models, like vision language action models or word action models. So to us, it doesn't matter. We just want to make sure we have the infrastructure, we have all the words, such that the robot can work reliably in the real environment. You know, you have told me that you think a lot of the predictions around humanoids were a little bit aggressive and were likely to see more constrained rollouts like warehouses or whatever. Can you talk a little bit about that and how that impacts what you're going to be tackling here at World Labs? So that's a very good question. So if you look at, for example, all the progressions of robotic applications in the real environments, it has always followed the trend from going from fully structured environments into semi-structured environments and then into unstructured environments. For fully structured environments, what we mean is that you have knowledge and the control over all the configurations within the environments. Like factories. Like factories or, for example, car manufacturing lines. Those have been automated for decades. And then you have, for example, semi-structured environments, which you have certain controls over the environments. For example, like the Amazon warehouses. Or, for example, like restaurants, hotels, where you have certain control over the environment to just make the task easier for your robots. But there are obviously many other objects. Or, for example, clothes. Those are the objects you don't have control. And then for the unstructured environments, it's like your home and my home. Those are, I would say, the grand challenge. Especially my house, trust me. Three dogs, five-year-old. Yes, dogs. Exactly. If you're thinking about where the robustness is coming from, robustness is coming from a sufficient coverage of the scenarios that robots might encounter. So it's so much easier and more approachable, at least like right now, to focus more on the semi-structured environments before we move on to fully unstructured environments. So we will move into that direction. It's just we want to take a more sustainable and more realistic approach towards it. I think your point here is that humanoids mimics human body. And evolution has optimized human body for unstructured environment. And so our fingers, our legs are not the best apparatus to do one thing. For example, if our only goal as a species is to climb trees, we will not have this body necessarily, right? So we'll have different kind of fingers. But what humans end up having are evolved into is this body shape that can be very general, but not necessarily best at everything. And that is for the survival of unstructured environment. But from a business point of view, from a pragmatic technology point of view, that this unstructured environment and a generalized body is actually the hardest problem to solve. It not necessarily even the right way to solve the problem We specialize so we take more specialized body to solve a narrower problem But the challenge for Cinex is that to be more body agnostic so that their infrastructure can serve different bodies and different semi-structured environments. A common lens to look at exactly this question is an economic lens, right? which is, if you compare it to, like, generative LLMs, they can create prose or code 10,000 times faster than a human being, a bunch cheaper than a human being. So the economic case makes sense because our brains aren't very efficient at that. However, our brains and our bodies are very efficient at 3D navigation, right? You know, like moving to the world or picking things up. And so this is just a prediction question, but do you believe we'll ever be able to build robots, at least in the foreseeable future, that have the power efficiency of a human being when it comes to menial tasks. So let's say just basically, you know, minimum wage or something like that. Like how far away are we from this? Is this like five years or this is like never? I think it's going to take a very long time. So if you're really thinking about like robots in the real environment, in the end, it will always be a system. So every working robot in the real environment is a system work. You need to be very mindful and thoughtful about how the systems are coming together. The hardware, the software, the brain, even to the details of, for example, what's the friction coefficients of your fingers. So there's a lot of things you have to consider to make these things a reality. And it will take iterations. But what I am excited about is that I have always been at the state of the arts of robot learning and also trying to push the state of the art forward. But the state of the arts always moving faster than I expected. So what I'm focusing on and trying to investigate right now is very different from, for example, when I started my PhD. So this is a speak to how fast the whole ecosystem has been evolving and all the moving pieces started coming together or building these robotic systems. But we also have to be calibrated about our predictions. So we will see a lot of progress. But to achieve, for example, human level efficiency and capabilities, it will take longer. Martin, the hardest thing in today's AI is to have the right measured optimism. Right? Right. It's totally true. Yeah. I mean, even LLMs does not have human brain efficiency. Human brain operates on 30 watts. Yeah, that's true. So we are far from that. But, I mean, performance to power may be close, right? In narrow tasks like software engineering. generating an image or software engineering than it is, right? Yeah, I think so. I don't think we're anywhere close when it comes to robotics. Does this change how you think about your, like strategically the level of ambition that your team can go after? I mean, does it change that or is it still very much in line with what you expected to do when you started? It definitely changed the trajectory in a very profound manner. So we see a lot of unlock in being able to do this whole process, do the modeling of the environments. in a much more efficient and much more scalable manner, especially in partnered together with WordLabs. And I also want to add to Fei-Fei, if you think about, for example, the current states of the language models, so those are models that's with incredible capabilities. But still, you don't just blind trust it to book your flight tickets or make your hotel reservations. Hopefully, there's still a person who's reading the output from those language models. But that is very different from how people will be using, for example, robotic models. Because for robotic models, out of the box, the robot has to work reliably in the real environment. And we don't even have the data. We don't even have all the necessary infrastructures around those for the robots to just out of the box work reliably in the real environment. So for that reason, being able to create these digital worlds, these scalable digital worlds where the robot can learn and evaluate within, Yes, it's going to unlock so much more potentials for being able to replace all the costly and unsafe data in the real environments with the data generated from the words for the robots to be able to do scalable learning and evaluations. You know, I've seen many of these kind of integrations. They actually work very well at this stage when they have this much alignment, which is great. But there's always like this question of, do you integrate now into what's happening now, or do you keep things quite separate and provide kind of like a long-term trajectory that will, you know, be realized, you know, in the year timeframe? How are you thinking about this, Fei-Fei? Is this something that integrates right away, or is this kind of a separate longer term? This is a great question. I think at this point, you know, Yunju, Changxi, Sonny, Justin, Ben, and I have been talking about this. At this point, we are going to take it thoughtfully. We're not rushing to integrate everything from code base to teams because I think Cinex does have a very well-thought, and I wouldn't call it standalone completely, but fairly contained tech stack as well as their customers, as well as the kind of products they're building. We're going to take time. We definitely will, we already on the simulation side, as well as the potential base model, action condition model side, we already are starting to talk. And also they are using Marble as an internal customer So we will be integrating, but we're not rushing to blend the team as like a full salad bowl. How are you thinking about geographies with this? Will the CNX move? Is it going to stay in the same place? Wundru is going to move. Oh, well, welcome. Here? I'm moving to San Francisco. Yeah, Florence to the Renaissance. Perfect. I think Aurora Labs is officially becoming a bi-coastal company where the headquarter is in San Francisco. So I've, you know, I live in Palo Alto. I feel like I'm in a different state. We, we, but we, but I'm actually excited that we're going to have an office in New York that can help us to attract talent on the East Coast. And also we have been talking about making sure that in both offices we set up the robots so that we get to basically test out and mature our engineering stack so that we can work with robots remotely because we have to do that for our customers anyway. So maybe just to be very concrete, Feifei, maybe let's just pencil out, like what is the perfect success case in two years? Like what product do you have? Who's engaging with it? How do they use it? Just to crisp like what this becomes. I'm very happy that Cynix team and WorldLabs team will have validated customers in a small number of important vertical use cases where our system, our infrastructure has proven to be truly beneficial to their automation needs. and these customers became our lighthouse examples to scale our business. How early, let's say someone listening to this is running a robotics company. At what stage should they engage with World Labs? Is it really early on? Is it somewhere in the middle? So right now, for our customers, because we are building this kind of real-to-sim-to-real pipelines, where the simulation is essentially the words we're going to provide the training and evaluation grounds. Some customers, they need only the real-to-sim part. They want to digitalize the tasks they care about and be able to do the evaluations of their robotic systems. Some customers need this real-to-sim-to-this entire pipeline, such as they will be able to have policies running on their hardware. So our platform is also designed in a way that is flexible, depending on what our clients need. And at the same time, the clients we are working with are actually pretty close to the deployments stage. So basically, they are working on very, very practical tasks. Those tasks, when replaced, when we have robotic solutions that are there, can just create value immediately. And they have at least tens or hundreds of these kind of situations they are thinking about to do the automations for. So as like together with WorldLabs, we'll be able to develop reliable solutions for those scenarios. As we have already shown, we have a number of scenarios already instantiated in our blog post. And we'll be able to further our investigation to see how they can actually solve the key requirements and also constraints faced by the real world deployments. I want to be very specific about this. Is it ever too late or too early to call WorldLabs if you're a robotics company? No, we want everybody to call us. We want to learn about your use case. Wonderful. If you're listening to this and you're anywhere close to a robotics project or robotics company, please track World Labs. Yes, thank you. Definitely open for business. Yeah, we're open for business. Not too early. All right. If you're doing robotics, call World Labs. Thank you both very much for coming. Thank you. Thanks for listening to this episode of the A16Z podcast. If you liked this episode, be sure to like, comment, subscribe, leave us a rating or review and share it with your friends and family. For more episodes, go to YouTube, Apple Podcasts and Spotify. Follow us on X at A16Z and subscribe to our sub stack at A16Z.substack.com. Thanks again for listening and I'll see you in the next episode. As a reminder, the content here is for informational purposes only. should not be taken as legal business, tax, or investment advice, or be used to evaluate any investment or security, and is not directed at any investors or potential investors in any A16Z fund. Please note that A16Z and its affiliates may also maintain investments in the companies discussed in this podcast. For more details, including a link to our investments, please see A16Z.com forward slash disclosures. Thank you.