Google’s PaLM-E is a generalist robot brain that takes commands

In a demonstration video, a robotic arm controlled by PaLM-E reaches into a bag of potato chips.
Expanding / In a demonstration video, a robotic arm controlled by PaLM-E reaches into a bag of potato chips.

Google research

On Monday, a group of AI researchers from Google and the Technical University of Berlin unveiled PaLM, a multimodal embodied visual language model (VLM) with 562 billion parameters that integrates vision and language for robot control. -E announced. They claim it is the largest of his VLM ever developed and can perform a wide variety of tasks without the need for retraining.

According to Google, given a high-level command such as “bring me a rice ball from the drawer,” PaLM-E generates an action plan for a mobile robotic platform with an arm (developed by Google Robotics). , can be run. action itself.

PaLM-E does this by analyzing data from the robot’s camera without the need for a preprocessed scene representation. This eliminates the need for humans to preprocess and annotate data, enabling more autonomous robot control.

In a demo video provided by Google, PaLM-E “brings rice chips out of the drawer” and does it. It incorporates multiple planning steps and visual feedback from the robot’s camera.

They are also resilient and able to react to their environment. For example, the PaLM-E model can guide a robot to retrieve a bag of chips from the kitchen. Integrating the PaLM-E into the control loop also makes it more tolerant of interruptions that may occur during the task. In the video example, the researcher grabs the chip from the robot and moves it, while the robot finds the chip and grabs it again.

of another example, the same PaLM-E model autonomously controls a robot through a complex sequence of tasks that previously required human guidance. A research paper from Google describes how PaLM-E transforms instructions into actions.

We demonstrate the performance of PaLM-E on challenging and diverse mobile manipulation tasks. We mainly follow the setup of Ahn et al. (2022), robots must plan a sequence of navigation and manipulation actions based on human instructions. For example, given the instruction “I spilled my drink, can you get me something to clean it up?” You need to plan a sequence of . 4. Place the sponge on the user. Inspired by these tasks, we develop his three use cases of affordance prediction, obstacle detection, and long-term planning to test PaLM-E’s embodied reasoning abilities. The low-level policy is from RT-1 (Brohan et al., 2022), a transformation model that takes an RGB image and natural language instructions and outputs end effector control commands.

PaLM-E is a predictor of the next token and is based on Google’s existing Large Language Model (LLM) called ‘PaLM’ (similar to the technology behind ChatGPT), hence ‘PaLM- called E. Google “carved” PaLM by adding sensory information and robotic control.

Based on language models, PaLM-E takes continuous observations such as images or sensor data and encodes them into a set of vectors of the same size as the language tokens. This allows the model to “understand” sensory information in the same way it processes language.

A Google-provided demo video showing the robot being guided by the PaLM-E following the instruction “Bring me a green star”. The researchers say the green star is “an object that this robot is not directly exposed to.”

In addition to the RT-1 Robotics Transformer, PaLM-E draws from Google’s previous work on the ViT-22B, a Vision Transformer model revealed in February. ViT-22B has been trained on a variety of visual tasks, including image classification, object detection, semantic segmentation, and image captioning.

Google Robotics is not the only research group working on robotic control using neural networks. This particular work is similar to Microsoft’s recent paper “ChatGPT for Robotics”. In this paper, we experimented with a similar combination of visual data and a large language model for robot control.

Aside from robotics, Google researchers observed some interesting effects that clearly stem from using a large language model as the core of PaLM-E. One indicates “active transmission”. This means that the knowledge and skills learned from one task can be transferred to another task, resulting in “significantly higher performance” compared to single-task robot models.

Also they Observed Model scale trends: “The larger the language model, the better it retains its language features when training visual language and robotics tasks. Quantitatively, the 562B PaLM-E model retains nearly all of its language features. doing.”

and the researchers Claim PaLM-E supports multimodal thought chain reasoning (which allows models to analyze a set of inputs containing both verbal and visual information) and multi-image inference (which uses multiple images as inputs to make inferences or predictions). ) to demonstrate new features such as ) even though they were trained with only single-image prompts. In that sense, PaLM-E seems to continue a surprising trend that emerges as deep learning models become more complex over time.

Google researchers plan to investigate more applications of PaLM-E in real-world scenarios such as home automation and industrial robotics. And they hope PaLM-E will inspire more research on multimodal reasoning and embodied AI.

“Multimodal” is a buzzword we hear more and more as companies reach for artificial general intelligence that can perform common ostensibly human-like tasks.



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *