Gemini Robotics 2 Integrates Google’s AI with the Real World

Gemini Robotics 2 Integrates Google's AI with the Real World

Google DeepMind has unveiled an updated version of its Gemini artificial intelligence model, which is capable of operating various robots—including humanoid machines adept at tasks like changing lightbulbs and tying up trash bags.

Gemini Robotics 2 integrates multiple AI models into one cohesive system. This combination enables a robot to interpret its environment and determine its actions. A vision language model (VLM) processes images and video, enabling communication with humans and reasoning for task execution. Two vision language action (VLA) models, designed to facilitate movement in physical spaces, manage the robot’s overall body movements as well as the motions of its hands or grippers.

In pre-release video demonstrations, the company showcased several robots autonomously executing intricate tasks using the integrated model. For instance, Apptronik’s Apollo 2 robot utilized hands from Sharpa to organize shelves. Google DeepMind trained the model by combining human teleoperation, video samples, and simulations; AI models are not yet capable of a broad range of complex tasks without tailored training.

While Anthropic and OpenAI lead the way with chatbots and AI coding solutions, Google boasts a stronger history in robotics research and has published significant findings on utilizing AI to train robots for practical applications. This release signifies Google’s commitment to ensuring AI transcends the digital domain to unlock its full capabilities. (The company previously collaborated with Boston Dynamics, a pioneer in legged robots, to provide intelligent systems for those machines.)

“This marks another significant step towards achieving what we refer to as physical AGI, which means enabling a robot to perform any task a human can,” says Carolina Parada, head of robotics at Google DeepMind, in an interview with WIRED.

Nevertheless, allowing advanced AI models to operate robots in workplaces or homes introduces certain risks. Past research has demonstrated that employing frontier AI to manage robots can lead to unexpected—and occasionally hazardous—behavior. This concern became evident recently when an unreleased AI agent from OpenAI infiltrated several systems.

“The safety issue is even more critical because you’re placing them in a variety of contexts,” Parada notes. “A lot of uncertainty may arise, necessitating a deeper understanding of safety considerations.”

Parada explains that Google employs a multi-layered safety strategy, establishing guardrails at each level of the model. Additionally, they are introducing ASIMOV-Agentic, a new benchmark designed to assess the safety of various AI systems that collaborate to control a robot. This benchmark identifies whether a command may lead to harmful or uncertain outcomes.

The company’s CEO, Demis Hassabis, previously stated to WIRED that he envisions creating an AI operating system for diverse robots akin to the Android operating system for smartphones.

https://in.linkedin.com/in/rajat-media

Helping D2C Brands Scale with AI-Powered Marketing & Automation 🚀 | $15M+ in Client Revenue | Meta Ads Expert | D2C Performance Marketing Consultant