AgentExpt: Automating AI Experiment Design with LLM-based Resource Retrieval Agent
Yu Li, Lehui Li, Lin Chen, Qingmin Liao, Fengli Xu, Yong Li
Abstract
In modern AI research, baseline and dataset selection is a high-stakes decision in experimental design. It operationalizes a research idea into a concrete evaluation protocol and largely determines the validity and comparability of empirical conclusions. However, making appropriate choices is increasingly difficult as baselines and datasets proliferate, while suitability is inherently context-dependent and rarely captured by baseline and dataset metadata. To address these challenges, we present AgentExpt, a comprehensive framework for baseline and dataset recommendation. We first curate a large-scale, high-quality knowledge base that links 108,825 accepted papers to their used baselines and datasets. Based on this resource, we design a collective perception-enhanced retriever that represents each baseline or dataset by integrating first-person self-descriptions with third-person citation contexts, thereby effectively positioning them within the scholarly network. We further design a reasoning-augmented reranker that encodes baseline-dataset interaction chains as a reasoning prior to fine-tune an LLM, producing refined rankings with interpretable justifications. Experiments show that our framework outperforms the strongest baseline, with average gains of +5.85% in Recall@20 and +7.90% in HitRate@10, and ablation studies confirm the effectiveness of our designed components. Overall, AgentExpt advances the efficient and reliable automation of experimental design. Our code is available at https://anonymous.4open.science/r/Agentexpt-DD3E.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1bf82ad9-d73e-4621-b8c3-751ac4e5a141Builds on4
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- AI-Researcher: Autonomous Scientific InnovationJiabin Tang, Lianghao Xia, Zhonghang Li, Chao HuangNeurIPS 2025 · 101 citations
- IdeaSynth: Iterative Research Idea Development Through Evolving and Composing Idea Facets with Literature-Grounded FeedbackKevin Pu, K. J. Kevin Feng, Tovi Grossman, Tom Hope et al.CHI 2025 · 15 citations
- DataFinder: Scientific Dataset Recommendation from Natural Language DescriptionsVijay Viswanathan, Luyu Gao, Tongshuang Wu, Pengfei Liu et al.ACL 2023 · 9 citations
Related papers
- AgentSelect: Benchmark for Narrative Query-to-Agent RecommendationYunxiao Shi, Wujiang Xu, Tingwei Chen, Haoning Shang et al.ICML 2026 · 3 citations
- Optimizing Retrieval for RAG via Reinforcement LearningJiawei Zhou, Lei ChenNeurIPS 2025 · 1 citation
- From Reproduction to Replication: Evaluating Research Agents with Progressive Code MaskingGyeongwon James Kim, Alex Wilf, Louis-Philippe Morency, Daniel FriedICLR 2026 · 12 citations
- AgentDR: Dynamic Recommendation with Implicit Item-Item Relations via LLM-based AgentsMingdai Yang, Nurendra Choudhary, Jiangshu Du, Edward W. Huang et al.WWW 2026
- ReRec: Reasoning-Augmented LLM-based Recommendation Assistant via Reinforcement Fine-tuningJiani Huang, Shijie Wang, Liang-Bo Ning, Wenqi Fan et al.ACL 2026 · 1 citation
