CEA-List's smart robotics demonstrator, shown with an end-to-end robotic input demo at the OECD's AI Working Day, highlights generative AI's potential as an enabler of robotic tasks whose instructions are given in natural language. CEA-List researchers designed a robotic handling agent that leverages computer vision and deep learning to accurately execute a grasping task based on a high-level natural-language instruction.
The purpose of the research was to design a software module that would give robots the ability to understand and execute tasks based on instructions given in natural language or provided in images — translating intuitive interactions into specific physical actions. The team integrated a generic foundation transformer AI model that had been pretrained on a large dataset of robot trajectories, then refined it on their own data to improve performance on the target tasks. The model ultimately selected, Octo, adapts efficiently to various robotic configurations, requires relatively little data, and is reasonable in terms of computing resources, thanks to a modular attention structure that lets it adjust to the specificities of target tasks and improve generalization across a wide range of robotic tasks.
To generate data specific to the grasping task, the researchers developed a remote operation mode built on a lightweight six-axis robot controlled using a virtual reality joystick, enabling precise, intuitive handling for quality data acquisition. Volunteers performed robotic grasping tasks involving a dozen objects handled in four distinct spatial configurations, with diversity important to ensure the data represents real-world tasks. CEA-List's PIXANO software was used to clean the gathered data, correcting annotation errors, and the Octo model was then fine-tuned using a cleaned training dataset containing 678 trajectories and a test dataset of 70 trajectories.
Once trained, Octo was successful at identifying and grasping an object from the training dataset, placed either alone or with distractor objects, without a dedicated 3D perception system. Research on more complex tasks, including bimanual object input, is currently underway. "These advances came out of our research on intuitive programming, the purpose of which is to help make robotics more accessible to operators without specialist knowledge or training," said Caroline Vienne, deputy department head at CEA-List. "The goal of our research is to leverage artificial intelligence to develop robotic systems that are robust, accessible, and rapidly deployable in industrial settings," said Jaonary Rabarisoa, research engineer at CEA-List. The work is part of CEA-List's 2024 "Responsible artificial intelligence" research program, which published the results in its 2024 activity report.