Leveraging semantic similarity for experimentation with AI-generated treatments
Lei Shi, David Arbour, Raghavendra Addanki, Ritwik Sinha, Avi Feller
Abstract
Large Language Models (LLMs) enable a new form of digital experimentation where treatments combine human and model-generated content in increasingly sophisticated ways. The main methodological challenge in this setting is representing these high-dimensional treatments without losing their semantic meaning or rendering analysis intractable. Here, we address this problem by focusing on learning low-dimensional representations that capture the underlying structure of such treatments. These representations enable downstream applications such as guiding generative models to produce meaningful treatment variants and facilitating adaptive assignment in online experiments. We propose double kernel representation learning, which models the causal effect through the inner product of kernel-based representations of treatments and user covariates. We develop an alternating-minimization algorithm that learns these representations efficiently from data and provides convergence guarantees under a low-rank factor model. As an application of this framework, we introduce an adaptive design strategy for online experimentation and demonstrate the method's effectiveness through numerical experiments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on12
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao et al.NeurIPS 2020 · 2,727 citations
- MIND: A Large-scale Dataset for News RecommendationFangzhao Wu, Ying Qiao, Jiun-Hung Chen, Chuhan Wu et al.ACL 2020 · 454 citations
- Rank-N-Contrast: Learning Continuous Representations for RegressionKaiwen Zha, Peng Cao, Jeany Son, Yuzhe Yang et al.NeurIPS 2023 · 129 citations
- High-Dimensional Sparse Linear BanditsBotao Hao, Tor Lattimore, Mengdi WangNeurIPS 2020 · 77 citations
Related papers
- AI-Assisted Variance Reduction in Randomized ExperimentsDavid Arbour, Eli Ben-Michael, Avi Feller, Apoorva Lal et al.KDD 2026 · 4 citations
- Sequences of Logits Reveal the Low Rank Structure of Language ModelsNoah Golowich, Allen Liu, Abhishek ShettyICLR 2026 · 9 citations
- From Causal to Concept-Based Representation LearningGoutham Rajendran, Simon Buchholz, Bryon Aragam, Bernhard Schölkopf et al.NeurIPS 2024 · 37 citations
- LLM-Driven Treatment Effect Estimation Under Inference Time Text ConfoundingYuchen Ma, Dennis Frauen, Jonas Schweisthal, Stefan FeuerriegelNeurIPS 2025 · 7 citations
- End-To-End Causal Effect Estimation from Unstructured Natural Language DataNikita Dhawan, Leonardo Cotta, Karen Ullrich, Rahul G. Krishnan et al.NeurIPS 2024 · 24 citations
