AI Tackles Rwandan Propaganda, Paves Way for Others

Dartmouth College

A new study from Dartmouth researchers reports the first digital tool for identifying online propaganda in Kinyarwanda, the national language of Rwanda's 15 million people—and possibly the first for any of the Bantu languages spoken by 350 million Africans.

As social media and artificial intelligence accelerate the spread of misinformation, Dartmouth senior Fabrice Niyigaba created the dataset KinyaProp to fill a critical void when it comes to languages like Kinyarwanda that are "substantially overlooked in AI research," he says.

While commercial platforms like Claude and ChatGPT exhibit near-human precision in spotting propaganda in languages such as English and Arabic, Niyigaba says, they are "essentially nonfunctional" at recognizing it in Kinyarwanda, according to results he presented July 6 at the annual meeting of the Association for Computational Linguistics in San Diego.

But when KinyaProp is integrated into these platforms, they can detect misinformation in Kinyarwanda with high levels of accuracy. The Bantu family is so close-knit, Niyigaba says, that his work is a launch pad for developing similar tools for catching misleading rhetoric in languages such as Swahili and Zulu, which have tens of millions of speakers.

"Existing models can generate fluent Kinyarwanda text, so the issue is not that they lack exposure to the language," says Niyigaba, a computer science major and first author of the peer-reviewed study co-authored with Ivory Yang , a PhD candidate in computer science, and Soroush Vosoughi , an associate professor of computer science.

"But if you ask them to complete a task that requires understanding, such as to detect propaganda, they significantly fail," he says.

Raising awareness of propaganda

KinyaProp is an extensive glossary for identifying and classifying misinformation in Kinyarwanda—what manipulative content looks like and the techniques it uses, such as appealing to fear and prejudice or sowing confusion.

While Niyigaba's motivation for KinyaProp is to detect misinformation in underrepresented languages, the extreme repercussions of propaganda loom over his country's recent history, he says.

"If we have models trained to detect misinformation in our language, we can raise awareness and pressure leaders and the media to not mislead people." — Fabrice Niyigaba '27

The 1994 Rwandan genocide claimed more than 800,000 lives in 100 days, mainly Tutsi civilians at the hands of Hutu extremists. The massacres—which followed decades of ethnic tension and warfare—shocked the world and set off a series of deadly conflicts that upended East and Central Africa.

"The news media was one of the platforms used for large-scale manipulation, but in 20 years I never heard anyone say that the news tried to manipulate them. People are not aware that they're being influenced," Niyigaba says.

"In Rwanda, we've seen the worst effects of propaganda, but we still have it. It takes away people's ability to think critically and question what is actually happening," he says. "If we have models trained to detect misinformation in our language, we can raise awareness and pressure leaders and the media to not mislead people. That is one way to reduce this problem."

KinyaProp essentially provides examples of misinformation in Kinyarwanda that the large language models that power platforms like Claude and ChatGPT can learn from and recognize when they see it, Yang says.

"A lot of work has been done in understanding misinformation in high-resource languages, meaning that LLMs are primarily English-centric," Yang says. Her work in Vosoughi's Minds, Machine, and Society Group focuses on training LLMs to recognize Indigenous and endangered languages using relatively small datasets.

"We didn't really know how these models would respond to Kinyarwanda until Fabrice did his work," she says. "We learned that misinformation is expressed much differently in Kinyarwanda and that provides a better understanding of how propaganda can differ across languages."

"What makes this work especially strong is that it fills a genuine research gap rather than simply reporting a small improvement on an existing benchmark," Vosoughi says. He adds that it's uncommon for an undergraduate-led paper to be accepted at ACL, one of the leading venues for research in natural language processing.

An "exceptionally meticulous" dataset

"Fabrice was exceptionally meticulous—the work involved an extended training and calibration process, careful construction of a consensus dataset, and systematic evaluation across multiple models and learning settings," Vosoughi says. "Its strengths are therefore not only its originality, but also its technical rigor, careful execution, and clear social relevance."

Vosoughi notes that "persuasion is not expressed identically across languages as it depends on linguistic structure, idioms, cultural conventions, and broader context. Systems that appear capable in English may leave serious blind spots in underrepresented languages."

To build KinyaProp's dataset, Niyigaba recruited three native speakers based in Rwanda to identify propaganda and misinformation in more than 600 news articles published in Kinyarwanda. The reviewers are all experts in the language who studied advanced grammar and morphology.

For three months, the reviewers were trained by the Dartmouth team on the principal techniques of propaganda and misinformation, then, after showing their proficiency, closely examined and annotated the articles.

If at least two of the reviewers agreed that an article or excerpted content was propaganda or misinformation, it went into KinyaProp. The content is sorted into one or more of 16 categories of propaganda to understand the techniques that are most frequently deployed in the language.

The study found that "appeals to authority," such as praising the actions and statements of the government or its leader, is by far the most common technique in Kinyarwanda.

"Our work is significant because it goes beyond labeling an entire article as propagandistic—it identifies the specific passages and techniques being used. In practice, carefully validated systems like this could help journalists, fact-checkers, and civil-society organizations identify harmful content and let native-language experts make the final judgment," Vosoughi says.

Foundational work in Bantu

The study finds that mainstream platforms struggle with Kinyarwanda's cultural context, complex sentence structure, and frequent use of symbolism. For instance, the phrase "gukura ubwatsi" means "to remove the grass," but is a common expression of gratitude.

As a result, AI models cannot always distinguish ordinary Kinyarwanda from real propaganda, repeatedly flagging normal words, Niyigaba says

For example, the formal form of woman, "igitsinagore," combines the words for sex and female, but an LLM will flag the first word, "igitsina," as harmful language. Similarly, a common phrase for being hungry, "kwicwa n'inzara," which literally means "to die because of hunger," will be marked as loaded language for including the Kinyarwanda word for "to kill."

On the other hand, the articles flagged by the annotation team contain sentences that a native speaker can identify as emotionally charged but that ChatGPT misses.

In one article, a political leader accuses an opposing party of violating human rights by saying, "You should not continue to overlook the violent crushing of human rights." In Kinyarwanda, "crushing" is expressed with the word "ihonyangwa." General AI interprets its meaning as "violation" and treats the sentence as neutral.

But to a native speaker, Niyigaba says, the complex word is extremely dramatic, meaning to physically crush or grind something down like with a mortar and pestle. That makes its use against a political opponent inflammatory.

"There is a difference between an LLM being able to reproduce a language and understanding it deeply enough to detect when it is, or is not, being used manipulatively," Niyigaba says.

The results, Niyigaba says, bode well for other languages in the Bantu family to which Kinyarwanda belongs, including Swahili and Zulu, the two largest.

Bantu consists of more than 600 languages that are spoken by about 30% of Africans and well known among linguists for their shared sentence and grammatical structures. Kinyarwanda is mutually intelligible with Kirundi, the national language of neighboring Burundi.

"A dataset built in Kinyarwanda is by far a closer linguistic neighbor to Swahili or Zulu than any English dataset could ever be," Niyigaba says.

"Our work gives researchers working on other Bantu languages a solid head start, both by using our data directly to train their models and by adapting our annotation framework instead of building from scratch," he says. "That makes KinyaProp foundational in this area."

/Public Release. This material from the originating organization/author(s) might be of the point-in-time nature, and edited for clarity, style and length. Mirage.News does not take institutional positions or sides, and all views, positions, and conclusions expressed herein are solely those of the author(s).View in full here.