Robots Learn Language Like Children Do

Okinawa Institute of Science and Technology (OIST) Graduate University

Do AI models understand the language that they produce?

To help answer this question and elucidate the mechanisms of language acquisition more generally, researchers at the Okinawa Institute of Science and Technology (OIST) have created a virtual robot with a brain-inspired neural network and tested its performance after giving it the remarkably human trait of curiosity.

By exposing the robots to a wide range of verb-adjective-object combinations in a physics-simulated environment, the researchers found a strong explanation for children's mysteriously rapid acquisition of the ability to understand language: a rich and varied language environment combined with childlike wonder. Their findings were published today in Science Advances.

"The curious robots achieve what appears to be a genuine understanding of language in half the time that their indifferent cousins do. And we were amazed by the play-like behavior and exception-handling performance that emerged independently during training," says Theodore Tinker, study first author and PhD student in Cognitive Neurorobotics Research Unit at OIST .

Human-like robots perform better when driven by curiosity

The researcher's AI is inspired by a so-called Predictive coding-inspired Variational Recurrent Neural Network (PV-RNN), which is built on leading theories of how our brains process information. While the rapidly advancing LLMs produce the statistically most probable response to a given input using a vast dataset, PV-RNNs are trained to reduce so-called Free Energy by maximizing accuracy (how well they correctly understand reality) and minimizing complexity (how much they change their internal beliefs after seeing what happened). The robots are trying to be correct while maintaining a stable system of beliefs.

"Imagine you encounter a small, dark square. You eat it, and now you know what dark chocolate tastes like. You then encounter another small, black square. Based on physical appearance and your internal beliefs, you predict it will taste like dark chocolate. If you're correct, your prediction is accurate, and your internal beliefs stay intact, minimizing the complexity term," explains Tinker. "But if it was actually licorice, your prediction is inaccurate, and your internal beliefs have to change to account for your new experience — high complexity."

For this study, the team also incorporated reinforcement learning into their PV-RNN by giving an external reward for completing a task and an internal reward for curiosity, which is directly tied to complexity. This leads to a competing set of impulses: the robot wants to be accurate and keep its current idea of the world intact, yet is also rewarded for taking actions that force it to update its internal beliefs. As such, curiosity gives the robot a way to learn to come up with novel solutions to high-complexity tasks. Tinker continues:

"Maybe you really like dark chocolate but have never tried white. So, one day you try, even though it runs counter to your desire for accuracy and low complexity. You might still think dark is tastier, i.e., it provides a higher external reward than white, but you also get an internal rush for satisfying your curiosity. And, as a bonus, you've updated your understanding of what chocolate is like in general."

The team observed the same behavior for robots mid-training. Even as they got consistently better at solving the task, they still knocked objects over just to experiment. "That's my favorite part, because nothing about any of the tasks requires knocking over objects. We didn't tell the robot to play around like a child. We told it to learn. And yet, apparently, one activity leads to the other — play-like behavior, reinforced by curiosity, helped the robots understand language faster."

Rich language diversity helps solve the Problem of the stimulus problem

The language-learning abilities of children have long puzzled scientists. Since Chomsky presented the 'Poverty of the stimulus' problem in 1980, researchers have struggled to explain exactly how children can rapidly acquire every feature of their language, despite having access only to sparse, often incomplete, and sometimes erroneous fragments.

This study suggests that a key to how toddlers solve this problem is a combination of curiosity and a diverse linguistic environment. The researchers have previously demonstrated that the success rate in completing previously unseen language tasks depends on the diversity of sentence constructions that the robot encounters. In this study, they replicated and extended earlier findings by significantly expanding the range of possible combinations, revealing a clear trend: curious robots exposed to 48 or 180 combinations showed generalization performance rates of 25% and 85%, respectively.

Learning to generalize through mistakes

In addition to play, the researchers were also surprised to discover that the robots appeared to replicate a very specific phenomenon observed in toddlers: the so-called U-shaped exception-handling performance curve.

In English, most past-tense verbs are conjugated by adding "-ed" to the word stem: walked, touched, pushed. But there are also significant exceptions to the rule, like 'ran' and 'went'. If you plot toddlers' success rate of verb conjugation over time, it produces a distinctly U-shaped curve. At first, when their vocabulary is limited, children accurately replicate language in learned chunks, such as "the dog ran" and "the boy went." But later, as they pick up the underlying verb-conjugation rules, their performance plummets, leading them to make mistakes in sentences they had previously been able to replicate successfully, like "the dog runned" and "the boy goed."

This is attributed to overgeneralization: the toddlers are now applying the internalized verb-conjugation rules to novel sentence constructions, rather than just repeating previously heard chunks. As they learn to handle these exceptions, performance rebounds.

The researchers replicated this challenge for the robots by swapping the meanings of two tasks. Where the stated task was "watch magenta pillar," they rewarded the robots for completing the unstated task "be near green pole," and vice versa. At first, they learned to complete these tasks as prescribed. But as they learned to complete more tasks, their performance dropped. Once they learned to handle the exception, their performance rebounded. "The fact that the robots mirrored children's exceptional-handling performance came as a complete surprise to us, because there was nothing in the model to specifically address overgeneralization or exception-handling — they seemed to pick this up on their own," says Tinker.

Clear access to language acquisition

Neural networks provide a good platform for studying language acquisition, as it's both difficult and ethically ambiguous to examine children across their entire language-learning trajectory under various experimental conditions. But while AI development has accelerated at an unprecedented scale over the past decade, LLMs and other sophisticated architectures have trillions of parameters trained on vast datasets, which is far removed from human learning.

In contrast, neural networks such as PV-RNNs are strong candidates for research. On top of reflecting leading theories of human cognition in both architecture and, as this study shows, behavior, they are transparent by design. Their so-called hidden state, which encodes their predictions of the future based on past experiences, is freely accessible, letting researchers track the robots' mental plans in real time as they solve tasks. It also allows the researchers fine-tuned control over which stimuli the robots are curious about, and to what degree.

Professor Jun Tani, study senior author and head of the OIST unit, summarizes: "The robots are not thinking like children, and their situation is very abstract compared to the extremely dense language environments of children. But this abstract environment allows us to study the impact of curiosity within a very convincing model of how humans learn to process language. We're very excited to see what other behaviors we can observe and study with this platform."

/Public Release. This material from the originating organization/author(s) might be of the point-in-time nature, and edited for clarity, style and length. Mirage.News does not take institutional positions or sides, and all views, positions, and conclusions expressed herein are solely those of the author(s).View in full here.