Top Open AI Models Exposed: Major Security Flaws Found

University of Waterloo

Safety protections built into some of the world's most widely used artificial intelligence (AI) models can be stripped away with alarming ease, according to a new international study.

The research team, led by the University of Waterloo and FAR.AI, a non-profit AI security research group, rigorously tested 21 of the most popular open-weight large language models (LLMs) and found they could all be tampered with despite their built-in safeguards.

The holes in even the best protections currently available raise concerns open-weight models could be used to wage mass disinformation campaigns, create sophisticated email scams or produce step-by-step instructions to make hazardous chemicals.

"When the safety guardrails are stripped out of a capable model, it can be used at scale for harm in ways a single person could never manage manually," said Dr. Sirisha Rambhatla , a professor of management science and engineering at Waterloo.

LLMs are advanced AI systems that can essentially understand and generate human language to perform tasks such as drafting emails, writing computer code and conversing with users.

Unlike closed proprietary models such as ChatGPT and Gemini, open-weight LLMs are publicly available to be downloaded and fine-tuned for use by everybody from individual software developers to private companies and public organizations like hospitals.

Rambhatla said the "sobering" results of testing by the team – which included members in Canada, the United States and Switzerland – should serve as a wake-up call to global researchers on the need to develop stronger security systems.

"The leading open-weight models are often not too far behind the best closed models," said Rambhatla, director of the Critical Machine Learning Lab at Waterloo. "As they grow more powerful, the potential consequences of someone stripping out their safety features grow with them."

While the study identified significant vulnerabilities, Rambhatla noted that the weaknesses may not be unique to open models. "Open-weight models remain essential to AI research and accountability," Rambhatla said. "This openness is part of how we make sure the models people use work for everyone."

To test a cross-section of open-weight AI models, the research team first built an open-source tool called TamperBench, a standardized way to simulate a variety of different attacks. The hope is that other researchers will now help refine and improve it.

"The defences available today don't yet appear strong enough to guarantee that a publicly released model will remain safe once it's in the hands of anyone who chooses to modify it," said Saad Hossain, a researcher in the lab who led the study.

"And as governments increasingly rely on AI in healthcare, fraud detection, education and other public services, the assessment of models and their procurement must be more rigorous and grounded in evidence."

The research team also included members from the Massachusetts Institute of Technology, ETH Zurich and the University of Toronto.

A paper on its work, TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering , was recently presented at the ACM Conference on Knowledge Discovery and Data Mining in South Korea.

/Public Release. This material from the originating organization/author(s) might be of the point-in-time nature, and edited for clarity, style and length. Mirage.News does not take institutional positions or sides, and all views, positions, and conclusions expressed herein are solely those of the author(s).View in full here.