Learning from AVA: Early Lessons from a Curated and Trustworthy Generative AI for Policy and Development Research
Nimisha Karnatak, Mohamad Chatila, Daniel Alejandro Pinzón Hernández, Reza Yazdanfar, Michelle Dugas, Renos Vakis
Abstract
General-purpose LLMs pose misinformation risks for development and policy experts, lacking epistemic humility for verifiable outputs. We present AVA (AI + Verified Analysis), a GenAI platform built on a curated library of over 4,000 World Bank Reports with multilingual capabilities. AVA’s multi-agent pipeline enables users to query and receive evidence-based syntheses. It operationalizes epistemic humility through two mechanisms: citation verifiability (tracing claims to sources) and reasoned abstention (declining unsupported queries with justification and redirection). We conducted an in-the-wild evaluation with over 2,200 individuals from heterogeneous organisations and roles in 116 countries, via log analysis, surveys, and 20 interviews. Difference-in-Differences estimates associate sustained engagement with 2.4–3.9 hours saved weekly. Qualitatively, participants used AVA as a specialized “evidence engine”; reasoned abstention clarified scope boundaries, and trust was calibrated through institutional provenance and page-anchored citations. We contribute design guidelines for specialized AI and articulate a vision for ‘ecosystem-aware’ Humble AI.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ab679107-5ec7-468f-94b5-6f894a0f3fa1Builds on31
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model SocietyGuohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin et al.NeurIPS 2023 · 1,975 citations
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka et al.NeurIPS 2024 · 1,166 citations
Related papers
- Accommodation and Epistemic Vigilance: A Pragmatic Account of Why LLMs Fail to Challenge Harmful BeliefsMyra Cheng, Robert D. Hawkins, Dan JurafskyACL 2026 · 6 citations
- Behavioral Indicators of Overreliance During Interaction with Conversational Language ModelsChang Liu, Qinyi Zhou, Xinjie Shen, Xingyu Bruce Liu et al.CHI 2026 · 4 citations
- ClaimDB: A Fact Verification Benchmark over Large Structured DataMichael Theologitis, Preetam Prabhu Srikar Dammu, Chirag Shah, Dan SuciuACL 2026 · 2 citations
- LiveFact: A Dynamic, Time-Aware Benchmark for LLM-Driven Fake News DetectionCheng Xu, Changhong Jin, Yingjie Niu, Nan Yan et al.ACL 2026
- CGRiC: Compositional Risk Certification for Structured LLM OutputsIbne Farabi Shihab, SANJEDA AKTER, Anuj SharmaICML 2026
