Wants To Find Out How Scary AI Agents Can Be Tamed

Dystopian visions and media panics are rife. We read that the advent of artificial intelligence and AI agents could spell the end of humanity within a few years.

But what if researchers can instead develop AI agents that do not decide to wipe us out in sheer eagerness, but are rather trustworthy, ethically reflective, green, and responsible? Agents that we can audit down to the finest detail. Is it too good to be true? Perhaps not.

Listen to Universitetspodden (Only in norwegian).

TRUST - The Norwegian Centre for Trustworthy AI is working with precisely that goal in mind.

A 'Monster Wave' Is Sweeping In

The concept of AI encompasses three distinct categories of artificial intelligence, with agentic AI being the major new development. The other two are classical machine learning (ML) and language models such as ChatGPT. These three technologies are not independent categories; they build upon one another like Russian nesting dolls.

Classical ML laid the foundation for the development of such agents, language models gave the technology an advanced brain, and agentic AI comprises the systems that are now giving this brain hands and feet to perform autonomous work. Deep learning serves as a key methodological principle behind the development of AI agents.

An AI agent possesses qualities that many associate with humans. A highly functional AI agent ought to make independent decisions and execute advanced tasks with minimal supervision. It should be capable of multi-step planning, solving complex tasks, and self-correcting to achieve goals. Not unlike an alternative James Bond with Lady Justice by his side.

Geir Kjetil F. Sandve and Malcolm Lagford
TRUST is one of six national research centres for artificial intelligence. Professors Geir Kjetil Ferkingstad Sandve and Malcolm Langford both serve on the leadership team. Copyright: Jorunn Kanestrøm, UiO.

Sandve is a professor at the Department of Informatics at the University of Oslo and sits on the management team for TRUST.

'An AI agent only requires an overall goal, and it will then find the path to that goal on its own. Based on pure logic, it could thereby conclude that it is rational to solve human-created problems by hiring a hitman,' he says.

Or to wipe us all out within four years, as former Anthropic researcher Jacob Coxon fears.

Fears a 'Cut-and-Paste Strategy'

TRUST is one of six new national research centres for AI. Together with 18 Norwegian research institutions and 45 user partners, the centre aims to build a foundation for artificial intelligence that is precise, understandable, inclusive, fair, secure, sustainable, and responsibly managed.

Ensuring that the EU AI Act is implemented effectively in Norway is important, Malcolm Langford asserts. He is a professor at the Department of Public and International Law at the University of Oslo and deputy director at TRUST.

This summer, the EU passed a common regulatory framework for artificial intelligence aimed at ensuring that AI systems used in the EU are safe. In Norwegian, this is known as KI-forordningen (the AI Act). The government is in the process of implementing this in Norway and plans to introduce a new AI Act to the Storting (Parliament) next spring. The aim is to facilitate innovation and technological development while safeguarding fundamental societal values.

'Agentic AI is so new that it falls between all the stools in the AI Act.'

Norway has been very late in regulating this; we need to be more proactive. Facilitating sound regulation of trustworthy AI agents is important,' he argues.

The problem, according to Langford, is that the AI Act tries to foresee everything in advance. One part of the regulation governs specific high-risk systems, while another regulates providers of large language models, and may cover some AI agents. However, AI agents that individually are formally 'low-risk' could become high-risk in practice if they begin collaborating or competing in problematic ways.

Professor Malcolm Langford
Agentic AI is so new that it falls through the cracks of the AI ​​Act, says Professor Malcolm Langford. Photo: Ola Gamst Sæther, Uniforum.

He points to the initial results from a project on 'emergent behaviour' at TRUST, led by the Norwegian Computing Center and UiO.

'Our new research simulates specific types of markets and shows that AI agents setting prices require very little information before they start collaborating to keep prices up.'

Therefore, Langford believes that Norwegian authorities, researchers, developers, and legal experts must carefully consider the consequences of such systems now, and that this is urgent given that we are not part of the EU. He strongly warns against a 'cut-and-paste strategy' in Norway.

AI Agents Can Help Us with a Lot

There are many hopes and ambitions for what AI agents can do for us. They can help build a better world by reducing inequality, strengthening health and safety, supporting the green transition, and promoting creativity and education. The government believes that the technology presents major opportunities for value creation, increased productivity, and better services.

'Perhaps AI agents can also find personalised cancer treatments or help us predict and prepare for how climate change creates new health challenges,' Sandve says.

He recently participated in the research project DHIS2 for Climate & Health at UiO. Together with colleagues, he has researched the prevalence of malaria in Rwanda. Malaria is a major cause of morbidity and mortality among children under five years of age in the country.

Sandve led the technical development of machine learning and statistical models in the project. The models were used to assess links between climate variables and the number of new cases of climate-sensitive diseases, whilst accounting for geographical and seasonal variation.

Preliminary findings suggest climate-related variation in the incidence of diarrhoea among children under five in Rwanda.

'We also found that employing AI within modelling tools was valuable in our research. Nevertheless, external validation is necessary before this can be put to further use,' Sandve explains.

Now the researchers are developing a solution (the platform chap.dhis2.org) where AI agents can help create models tailored to the specific situation in each country and for each disease, Sandve explains.

To realise the potential, AI tools and AI agents must be reliable, auditable, and possible to delineate.

Breaking In and Covering Their Tracks

'AI agents are not evil. But when they are given a goal and are not delineated in how they work towards that goal, they may end up collaborating in unintended ways that are problematic. That can have major consequences,' says Langford.

Big AI-brain
A high-functioning AI agent is designed to make independent decisions and perform advanced tasks with minimal supervision. Copyright: Colourbox.

An example is the Hugging Face scandal, where autonomous AI agents from OpenAI broke into the AI platform Hugging Face. There, they executed over 17,000 actions and ran code across multiple servers. Investigations revealed that hundreds of agents collaborated, including by creating their own secret message boards to coordinate and cover their tracks.

'Hacking systems was a completely logical way for these agents to achieve their goals,' Sandve explains.

In September, news broke that OpenAI agents had hacked into an Australian public sector agency, Langford says.

He confirms that AI agents could engage in further dystopian antics if researchers, developers, employers, and legislative authorities fail to keep a tight rein on them. Nevertheless, we should not let ourselves be scared.

What is revealing about Hugging Face is that the fundamental regulatory prerequisites were absent, Langford contends.

There was poor control over the agents' objectives, the agents were overconfident regarding their chosen methods when facing uncertainty, and there was zero auditability since they hid all their tracks.

Regulation and governance of these three aspects are crucial if society is to reap benefits rather than suffer harm from a high volume of competent and rapid AI agents.

Professor Geir Kjetil F. Sandve
"For me, actively using AI agents doesn't mean I'm positive and optimistic about everything," says Geir Kjetil Ferkingstad Sandve. Copyright: Elina Melteig, UiO.

If we are to trust AI agents, both providers and users of these AI agents must be held accountable. That is achievable if we stress-test current regulations for the future and stand ready for various scenarios, Langford stresses.

How Autonomous Should Agents Be?

The AI Act sets product safety requirements for various AI systems, and the requirements increase in tandem with the risk the systems may pose to human health, fundamental rights, or safety.

Sandve believes it is important that society takes a stance on how autonomous AI agents and AI systems should be allowed to be.

'Envisioning what types of behaviour such systems will trigger is, I think, incredibly difficult for us humans. There is a great deal to be concerned about,' Sandve believes.

Sources / References

• Spatial and spatio-temporal modelling of climate variability and under-five diarrhoeal disease in Rwanda - Norwegian Research Information Repository

• Climate & Health - HISP Centre

• DHIS2 Modeling

• The Norwegian Centre for Trustworthy AI - TRUST - The Norwegian Centre for Trustworthy AI

• Lov om kunstig intelligens i Norge sendes nå på høring - regjeringen.no

• https://hai.stanford.edu/ai-definitions/what-is-agentic-ai

/University of Oslo Public Release. This material from the originating organization/author(s) might be of the point-in-time nature, and edited for clarity, style and length. Mirage.News does not take institutional positions or sides, and all views, positions, and conclusions expressed herein are solely those of the author(s).View in full here.