Home AIAI agents create virtual playgrounds to help robots get crucial training data | MIT News

AI agents create virtual playgrounds to help robots get crucial training data | MIT News

by OmarAli
AI agents create virtual playgrounds to help robots get crucial training data | MIT News

Robots are increasingly seen walking down the street, surrounded by astonished onlookers. But these machines are not yet the all-round assistants you want for work in a kitchen or factory, and a major bottleneck is data. Similar to humans, robots learn best through experience. The challenge is that physically teaching these machines so many actions in different environments is labor-intensive and time-consuming.

“One natural idea is to use simulations as a training ground. While there have been significant advances in the physics engines that power robot simulators in recent years, one of the remaining challenges is creating sufficiently rich and diverse simulation content to capture the complexity of the real world,” says Russ Tedrake, Toyota Professor of Electrical Engineering and Computer Science (EECS), Aerospace and Mechanical Engineering at MIT and principal investigator at MIT’s Computer Science and Artificial Intelligence Laboratory (CSAIL).

It turns out that AI agents, or semi-autonomous programs that “think” and perform well-defined tasks, could help create the lifelike virtual environments that robots need. The new “SceneSmith” system, developed by researchers at MIT CSAIL and the Toyota Research Institute, uses three agents to assemble the objects, walls and overall image of a 3D scene. Its recreations of interior spaces such as restaurants, bedrooms and hotels are more realistic and detailed than previous systems, helping robots practice skills and try out different ways of completing tasks before they are turned on. In return, engineers save time during real-world testing.

The agents have a sense of what everyday places should look like because they each rely on a multimodal system called a vision-language model (VLM), specifically the state-of-the-art VLM GPT-5.2. It was trained with lots of text and images from the internet to handle more visual prompts. This advanced model imparts a kind of spatial knowledge to each agent: first, a “designer” agent generates the elements of a scene, then a “critic” advises whether it looks realistic, and finally an “orchestrator” manages their back and forth and decides when the design is ready. Once the three VLMs have completed their creative collaboration, the scene can be loaded directly into the physics simulation software.

“We found that the system can construct 3D scenes the way a human designer would,” says MIT EECS graduate student Nicholas Pfaff, a CSAIL researcher and lead author of a paper introducing Tedrake’s work. “We created over 1,300 scenes with a leading VLM with internet-scale prioritization, and he came up with incredibly creative and diverse arrangements. I didn’t teach the system this in the prompts; it just improvised.”

Talk to my agent

Thanks to VLM agents, you can ask SceneSmith to do things like “create a garage with a car, a workbench, tires stacked in the corner, and a ladder on the wall,” creating a virtual playground full of objects for a robot to tinker with. These rooms are decorated with up to six times more items per scene than previous methods, making them great for helping robots learn skills such as: B. putting a cup in the sink, putting fruit on plates, and moving a soda can from a shelf to a table.

With so many rich virtual environments at your disposal, you can assess whether your robot is ready for use without much trial and error in the physical world. Researchers tested various action plans (also called “guidelines”) in SceneSmith’s digital worlds, generating 100 unique rooms. A VLM agent evaluated each attempt and determined that the robot’s plans were flawed and the machine frequently failed at its tasks. People agreed with the model’s judgments over 99 percent of the time, which could help roboticists weed out erroneous approaches in simulation before a robot moves in the real world.

But how realistic are these virtual worlds really? It can be difficult to prove unequivocally, so researchers approached the question from multiple angles. The most telling test: They inserted a pre-trained robot policy – an AI controller that was trained largely on real data and had never seen a SceneSmith scene – into the generated environments. In one test, users asked the system to “take the apple out of the bowl and put it on the cutting board,” and the simulated robot did just that. If the scenes didn’t match the real-world settings from which politicians learned, it simply wouldn’t have worked.

The team also teleoperated robots through the virtual rooms, guiding them to open cabinets, put away bottles and navigate between rooms. Their experiments showed that the environments can withstand sustained physical interaction and extend beyond visual inspection.

Behind the scenes

The agents used by SceneSmith each play a precisely defined role in the generation process and gradually flesh out scenes. You essentially create a floor plan and bring it to life.

Let’s say you want to create a scene that resembles the first floor of a house. The “Designer” VLM would begin with a general layout that would be reviewed by the “Critic” and then approved by the “Orchestrator.” The agents repeat this approach for each step: adding furniture, placing objects on walls and then ceilings, and finally inserting objects that robots can manipulate. For example, the VLMs can add cabinets that the robots can open and close – a moveable item that previous baselines didn’t often have.

At each stage, the second VLM ensures that the scene is practical, for example by pointing out that a bathtub is being removed from a living room. The third VLM ensures that a high-quality scene is generated and even delays the design process by a few turns if the visual representation does not meet the requirements. Once the three VLMs have completed their creative collaboration, the mechanics of the physical world will be added via simulation software.

With a solid understanding of what rooms should look like, where objects should be placed, and real-world physics, SceneSmith has a noticeable advantage over previous methods. Compared to scene generation baselines like HSM and Holodeck, SceneSmith created environments with more objects, including a private office, a pottery shop, and even a Minecraft-style game room.

SceneSmith was also a favorite with over 200 users. They found that the visual representation of the system was more realistic in over 90 percent of cases. They also found that it generally followed prompts better than other approaches. In other words, it was best at generating the virtual playgrounds that users actually wanted to see.

A system with many talents

Realism, variety and richness are all SceneSmith strengths, even when it comes to creating custom 3D objects. You can ask it to create a rolling trolley and it will create a 2D image, which will then be converted into a detailed model with physical properties such as mass, friction and inertia.

However, such a detailed process comes with a trade-off in speed. A single scene can take several hours to create as agents create and closely inspect each object. With more computing power, the efficiency of the system could be increased dramatically. CSAIL engineers also hope to expand to deformable objects (like sponges) if extensive 3D libraries become available.

“SceneSmith represents a significant advance in this regard by providing an agent framework for generating simulation-ready indoor environments with just simple text input,” says Jeremy Binagia, an applied scientist at Amazon Robotics who was not involved in the research. “It advances the state of the art in several ways, including by pushing the limits of object density in the simulated environment, ensuring that all objects are physically correct (rather than just visually realistic), and creating assets that are not limited to a fixed library as they can be generated via text-to-3D.”

Pfaff and Tedrake co-authored the article with Thomas Cohn SM ’24, an MIT graduate student and CSAIL researcher; and Toyota Research Institute roboticists Sergey Zakharov and Rick Cory SM ’08, PhD ’10. Their work was supported in part by Amazon, the US Office of Naval Research, the Toyota Research Institute and the US National Science Foundation.

The team presented their results in the spotlight at the International Conference on Machine Learning last week.

https://news.mit.edu/2026/ai-agents-create-virtual-playgrounds-to-help-robots-get-crucial-training-data-0713

Viral Trends

This website uses cookies to improve your experience. We'll assume you're ok with this, but you can opt-out if you wish. Accept Read More