KAIST Unveils AI That Sees Clearly in Dark and Smoke

The Korea Advanced Institute of Science and Technology (KAIST)

Multimodal large language models (MLLMs), which process multiple types of sensory information such as text, images, and audio at the same time, are rapidly expanding the range of applications for artificial intelligence (AI). However, in real-world environments, these models can misinterpret the physical characteristics of sensors, mistakenly identify objects, or claim to hear sounds that are not actually present simply because a certain object appears in a video. These errors are known as hallucinations. A KAIST research team has developed a new technology that corrects such information confusion and physical misperceptions in AI.

KAIST (President Choongsik Bae) announced on the 31st of July that a research team led by Professor Yong Man Ro from the School of Electrical Engineering has developed two core technologies that overcome the tendency of existing large language models to rely too heavily on ordinary camera (RGB) images and enable AI to suppress cross-modal hallucinations that occur when different sensory inputs become mixed.

The first technology developed by the research team is the Diverse Negative Attributes (DNA) optimization method, which helps AI accurately understand the physical characteristics of special camera sensors such as thermal, depth, and X-ray sensors. Existing AI models often failed to understand the physical meaning of such images, for example by mistaking bright areas in thermal images for simple light reflection.

The research team built VS-TDX, the first comprehensive benchmark for evaluating diverse vision sensors, and used the types of wrong answers that AI frequently produces as learning signals to help the model internalize the characteristics of each sensor. As a result, the AI gained a "new eye" that allows it to accurately infer the state of objects even in darkness or smoke.

The second technology is Modality-Adaptive Decoding (MAD), a control method that blocks hallucinations caused by confusion between visual and auditory information at the source. This technology prevents AI from mistakenly claiming that it hears a sound that does not actually exist simply because a certain object appears in a video.

MAD works by having the AI self-assess whether vision or audio is more important for a given task, and then increasing the weight of the more relevant modality in real time. A key advantage of this technology is that it can immediately suppress hallucination errors without costly model retraining, as it is training-free.

Instead of retraining AI models at large scale with massive computing resources, the research team maximized cost efficiency by introducing the DNA method, which enables fine adjustment with only a small amount of data, and the MAD plug-in approach, which requires no additional training at all.

These technologies can be applied to autonomous vehicles operating at night or in bad weather, robots performing missions in smoke-filled environments, and unmanned aerial vehicles using thermal cameras. They are also expected to be useful in fields that process multiple types of sensor information together, such as airport X-ray security screening and medical image analysis.

Professor Yong Man Ro said, "This research is significant because it reduces AI's sensory bias and misperceptions without large-scale retraining," adding, "It will serve as a foundation for building multimodal AI that can be trusted in real-life and industrial settings."

This achievement was notable for its continuity, with Sangyun Chung, a doctoral student in KAIST's School of Electrical Engineering, participating as first author in both studies. Dr. Youngjun Yoo also participated as co-first author in the DNA study.

Among the related papers, the MAD study was presented in June at the Conference on Computer Vision and Pattern Recognition (CVPR), the world's leading international conference in AI and computer vision. The DNA study was published in IEEE Transactions on Image Processing, a leading international journal in the field of image processing.

※ Paper title: Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking, DOI: 10.48550/arXiv.2412.20750

※ Author information: Sangyun Chung (KAIST, co-first author), Youngjun Yoo (KAIST, co-first author), Se Yeon Kim (KAIST, third author), Youngchae Chee (KAIST, fourth author), Yong Man Ro (KAIST, corresponding author)

※ Paper title: MAD: Modality-Adaptive Decoding for Mitigating Cross-Modal Hallucinations in Multimodal Large Language Models, DOI: 10.48550/arXiv.2601.21181

※ Author information: Sangyun Chung (KAIST, first author), Se Yeon Kim (KAIST, second author), Youngchae Chee (KAIST, third author), Yong Man Ro (KAIST, corresponding author)

※ Related demo video: https://youtu.be/VuP9i6Vfk8o

This research was supported by the Institute of Information & Communications Technology Planning & Evaluation's (IITP's) Human-Centered AI Core Technology Development Program and by a Center for Applied Research in Artificial Intelligence (CARAI) grant funded by the Defense Acquisition Program Administration (DAPA) and the Agency for Defense Development (ADD).

/Public Release. This material from the originating organization/author(s) might be of the point-in-time nature, and edited for clarity, style and length. Mirage.News does not take institutional positions or sides, and all views, positions, and conclusions expressed herein are solely those of the author(s).View in full here.