AI Headphones Single Out Voices in Crowds With One Look

Noise-canceling headphones have gotten very good at creating an auditory blank slate. But allowing certain sounds from a wearer's environment through the erasure still challenges researchers. The latest edition of Apple's AirPods Pro, for instance, automatically adjusts sound levels for wearers - sensing when they're in conversation, for instance - but the user has little control over whom to listen to or when this happens.

A University of Washington team has developed an artificial intelligence system that lets a user wearing headphones look at a person speaking for three to five seconds to "enroll" them. The system, called "Target Speech Hearing," then cancels all other sounds in the environment and plays just the enrolled speaker's voice in real time even as the listener moves around in noisy places and no longer faces the speaker.

The team presented its findings May 14 in Honolulu at the ACM CHI Conference on Human Factors in Computing Systems. The code for the proof-of-concept device is available for others to build on. The system is not commercially available.

"We tend to think of AI now as web-based chatbots that answer questions," said senior author Shyam Gollakota, a UW professor in the Paul G. Allen School of Computer Science & Engineering. "But in this project, we develop AI to modify the auditory perception of anyone wearing headphones, given their preferences. With our devices you can now hear a single speaker clearly even if you are in a noisy environment with lots of other people talking."

To use the system, a person wearing off-the-shelf headphones fitted with microphones taps a button while directing their head at someone talking. The sound waves from that speaker's voice then should reach the microphones on both sides of the headset simultaneously; there's a 16-degree margin of error. The headphones send that signal to an on-board embedded computer, where the team's machine learning software learns the desired speaker's vocal patterns. The system latches onto that speaker's voice and continues to play it back to the listener, even as the pair moves around. The system's ability to focus on the enrolled voice improves as the speaker keeps talking, giving the system more training data.

Related:

/Public Release. This material from the originating organization/author(s) might be of the point-in-time nature, and edited for clarity, style and length. Mirage.News does not take institutional positions or sides, and all views, positions, and conclusions expressed herein are solely those of the author(s).View in full here.