KAIST Develops Vector Database for Up-to-Date Information

Korea Advanced Institute of Science and Technology
The research team. From left: Professor Min-Soo Kim and Darae Lee, a master's graduate of the KAIST School of Computing.
The research team. From left: Professor Min-Soo Kim and Darae Lee, a master's graduate of the KAIST School of Computing.

< The research team. From left: Professor Min-Soo Kim and Darae Lee, a master >

When new information is continually added and old information is removed, an AI system may fail to find the correct material even when that material exists. KAIST researchers have developed a vector database technology that preserves the search paths leading to the needed information as data changes. This allows AI to find up-to-date information more accurately and quickly.

KAIST (President Choongsik Bae) announced on October 5 that a research team led by Professor Min-Soo Kim from the School of Computing has developed CONDA (Connectivity-Aware Dynamic Index). CONDA is a dynamic vector database indexing technology that maintains high search accuracy even as data is continuously added and deleted.

Generative AI models do not inherently have access to information that emerges after training. Retrieval-augmented generation (RAG) has been widely adopted to address this limitation. RAG searches sources such as recent news or internal company documents and uses the results in its answers.

For RAG to work properly, the system needs a search technology that can quickly and accurately find the information relevant to a question among vast amounts of data. Many AI services today convert the meaning of documents or images into vectors, which are numerical representations that computers can process. They then link pieces of information with similar meanings and follow these links to find the material they need.

The challenge is that data in real companies and on the internet changes constantly. New documents, news, and product information keep arriving, while older information is modified or deleted. As these changes repeat, the links built between data points can gradually break, a problem the researchers call "connectivity collapse." The needed information is still in the database, but the AI can no longer reach it, and search accuracy drops. In other words, the information itself has not been lost. The search path that leads to it has been broken.

CONDA considers both distances between data points and whether the search paths needed to reach the information remain connected. When new data is added or existing data is deleted, CONDA preserves the links that are important for search. This prevents specific pieces of information from becoming isolated in the search graph.

Figure 1. The data insertion and deletion processes in CONDA.
Figure 1. The data insertion and deletion processes in CONDA.

< Figure 1. The data insertion and deletion processes in CONDA. >

In experiments where data was constantly added and deleted, CONDA improved search accuracy (recall) by up to 24.5% compared with state-of-the-art methods. It also increased data update throughput by up to 1.90 times.

The team also ran searches and data updates at the same time for six hours on a large-scale dataset containing 100 million vectors. Throughout the test, CONDA recorded the lowest search latency and the highest search accuracy among the compared methods. This result suggests that the technology can be applied to large-scale AI services.

The technology can be used in knowledge search systems where information such as internal company regulations and work documents changes frequently. It can also be used in a range of AI services where data changes continuously, including news search, product search, and recommendation systems. In particular, it is expected to serve as a foundational technology for RAG, in which generative AI searches external information and uses it in its answers. CONDA allows this up-to-date information to be retrieved reliably over extended periods.

The research results are also being applied to a commercial product. CONDA is being integrated into AkasicDB, the database product of GraphAI, an AI data infrastructure company founded by Professor Kim. It is scheduled for commercial release in the fourth quarter of this year.

"The key value of retrieval-augmented generation (RAG) is that it allows large language models (LLMs) to find and use the latest information they were not trained on at the moment it is needed," said Professor Kim. He added that the study is significant because it presents a data infrastructure that lets AI use up-to-date knowledge accurately over long periods, even in real-world environments where data changes constantly.

Darae Lee, who completed a master's degree at the KAIST School of Computing, is the first author of the study, and Professor Kim is the corresponding author. The findings were presented on September 2 at VLDB 2026 (the International Conference on Very Large Data Bases), a leading international conference in the database field.

Paper title: CONDA: A Connectivity-Aware Dynamic Index for Approximate Nearest Neighbor Search over Evolving Data

DOI: 10.14778/3836663.3836694

This research was supported by the Mid-Career Researcher Program of the National Research Foundation of Korea (NRF) and the SW Star Lab program of the Institute of Information & Communications Technology Planning & Evaluation (IITP), both funded by the Ministry of Science and ICT (MSIT).

/Public Release. This material from the originating organization/author(s) might be of the point-in-time nature, and edited for clarity, style and length. Mirage.News does not take institutional positions or sides, and all views, positions, and conclusions expressed herein are solely those of the author(s).View in full here.