KinyaProp: Fine-Grained Propaganda Annotation in Kinyarwanda
Manzi Fabrice Niyigaba, Ivory Yang, Soroush Vosoughi
Abstract
Propaganda is a widely used approach for shaping public opinion and disseminating misinformation in news media. While it has recently gained significant attention within the NLP community, research on fine grained propaganda detection remains heavily concentrated in high resource languages. To bridge this gap, we introduce KinyaProp, the first finegrained propaganda dataset of its kind for Kinyarwanda and, to our knowledge, the first such resource created for a Bantu language. Using this dataset, we evaluate whether state-ofthe-art LLMs can function as reliable annotators in a genuinely low resource and culturally grounded setting. Our results show that current multilingual LLMs do not reliably approximate human annotation behavior. Instead, they behave as conservative annotators whose performance is largely limited to lexically explicit cues, substantially under-identifying propaganda and exhibiting extremely low and unstable performance on discourse-level techniques. Our findings highlight an important limitation of recent successes in LLM based annotation reported for high resource languages, demonstrating that such results do not readily transfer to low resource settings, where scalable annotation would be most valuable. We release KinyaProp to support future research on fine grained propaganda detection and to enable more robust evaluation of multilingual models in underrepresented languages.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a839df4b-ae64-400b-bae7-14886e805abbBuilds on5
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- KinyaBERT: a Morphology-aware Kinyarwanda Language ModelAntoine Nzeyimana, Andre Niyongabo RubungoACL 2022 · 45 citations
- Leveraging Declarative Knowledge in Text and First-Order Logic for Fine-Grained Propaganda DetectionRuize Wang, Duyu Tang, Nan Duan, Wanjun Zhong et al.EMNLP 2020 · 3 citations
- Discourse Structures Guided Fine-grained Propaganda IdentificationYuanyuan Lei, Ruihong HuangEMNLP 2023
- The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMsNitay Calderon, Roi Reichart, Rotem DrorACL 2025
Related papers
- Detecting Propaganda Techniques in Code-Switched Social Media TextMuhammad Umar Salman, Asif Hanif, Shady Shehata, Preslav NakovEMNLP 2023 · 5 citations
- PolyNarrative: A Multilingual, Multilabel, Multi-domain Dataset for Narrative Extraction from News ArticlesNikolaos Nikolaidis, Nicolas Stefanovitch, Purificação Silvano, Dimitar Iliyanov Dimitrov et al.ACL 2025
- Faking Fake News for Real Fake News Detection: Propaganda-Loaded Training Data GenerationKung-Hsiang Huang, Kathleen R. McKeown, Preslav Nakov, Yejin Choi et al.ACL 2023 · 35 citations
- AI 'News' Content Farms Are Easy to Make and Hard to Detect: A Case Study in ItalianGiovanni Puccetti, Anna Rogers, Chiara Alzetta, Felice Dell'Orletta et al.ACL 2024 · 1 citation
- MultiSocial: Multilingual Benchmark of Machine-Generated Text Detection of Social-Media TextsDominik Macko, Jakub Kopal, Róbert Móro, Ivan SrbaACL 2025 · 15 citations
