Artificial intelligence (AI) has gradually made its way into many areas of life with abilities that often match or surpass those of humans. But a study in the Cell Press journal iScience publishing on September 17 finds that AI-powered machines have trouble recognizing objects from their overall shapes when aspects of an image are distorted. The findings offer a new way to evaluate and compare computer vision to that of people.
"Current AI models do not accomplish visual object perception in the same way that humans do," says author Biyu J. He of New York University. "We tested more than 200 deep neural networks, and no models fully reproduced humans' object recognition patterns. Humans' ability to leverage the global shape cue for visual object recognition remains unparalleled."
When you're presented with an object, you can recognize it based on multiple features, including its texture, small internal details, and overall shape or silhouette. But, compared to people, machines can get tripped up by even small image distortions.
Earlier studies have shown that computational deep neural network (DNN) models, which are developed by training computers on large amounts of data and images, can rival the abilities of human observers on certain tasks. Over time, the machines "learn" complex patterns that allow them to visually recognize and name objects based on their appearance.
To find out if existing DNN models for machine vision work the same way human vision does, He and her colleagues created a set of images to put the AI models to the test. They systematically altered the global shapes, internal parts, and textures of many images, such as a cat, butterfly, car, and corn on the cob, and then compared humans' ability to correctly identify those objects to that of dozens of the best DNN models.
They found that humans were better at recognizing objects from their overall shapes than the AI algorithms. The computer algorithms consistently underperformed compared to people anytime object recognition depended only on global shape recognition.
"These results show that current computer vision models are not quite human-aligned, despite being trained on a massive number of pictures that humans have taken," says He.
The study offers a new way to evaluate and compare computer vision to that of people. The researchers say that their results may lead to improvements in computer vision in the future. Such improvements could aid in the development of improved assistive devices, including brain-computer interfaces that hold promise for enabling people with disabilities to better see and act in the real world.
The team says that they'll continue making comparisons between AI and the human brain, with the goal to "help build more human-aligned AI."