Language Model Augmented Semi-Supervised Statistical Inference
Xinrui Ruan, Yingfei Wang, Waverly Wei, Jingshen Wang
摘要
Semi‑supervised statistical inference plays a key role in biomedical research, where labeled data often have higher quality but are limited due to costly clinical annotation. Yet, existing semi‑supervised statistical inference methods rely heavily on structured variables and strictly matched covariates between labeled and unlabeled datasets -- limitations ill‑suited for the heterogeneity and unstructured nature of real-world biomedical data. Modern biomedical studies increasingly collect unstructured data (clinical notes, patient audio and video recordings), with inconsistent protocols across datasets causing covariate misalignment (for instance, detailed medication histories may be recorded in one study but not another). Recent advances in pre‑trained multimodal large language models (LLMs), which excel at handling unstructured data, present an attractive potential solution. To transform this potential into rigorous semi-supervised statistical inference methods for biomedical research, two key challenges must be addressed: (1) How can we reliably integrate LLMs to enhance semi-supervised inference efficiency without compromising statistical validity? (2) How can those efficiency gains persist despite mismatched covariates between labeled and unlabeled datasets? In this paper, we tackle these challenges by systematically calibrating pseudo-labels provided LLMs with a novel prediction-invariance identification strategy. Our resulting semi‑supervised inference framework improves parameter estimation efficiency while maintaining full statistical validity, as demonstrated through our theoretical results and illustrated in a case study for identifying key biomarkers in Alzheimer’s disease detection with speech data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language ModelsYung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim 等ICLR 2024 · 被引用 354 次
- Using Imperfect Surrogates for Downstream Inference: Design-based Supervised Learning for Social Science Applications of Large Language ModelsNaoki Egami, Musashi Hinck, Brandon M. Stewart, Hanying WeiNeurIPS 2023 · 被引用 74 次
- Discovering Latent Knowledge in Language Models Without SupervisionCollin Burns, Haotian Ye, Dan Klein, Jacob SteinhardtICLR 2023 · 被引用 45 次
相关 Paper
- Medical Vision-Language Pretraining with LLM-Guided Temporal SupervisionLiang Bai, Zhi Wang, Huimin Yan, Xian YangAAAI 2026
- DuSSS: Dual Semantic Similarity-Supervised Vision-Language Model for Semi-Supervised Medical Image SegmentationQingtao Pan, Wenhao Qiao, Jingjiao Lou, Bing Ji 等AAAI 2025 · 被引用 13 次
- LLM-Guided Diagnostic Evidence Alignment for Medical Vision–Language Pretraining under Limited PairingHuimin Yan, Liang Bai, Xian Yang, Long ChenICML 2026 · 被引用 1 次
- Local Patterns Generalize Better for Novel AnomaliesYalong JiangICLR 2025
- Language-Driven Cross-Modal Classifier for Zero-Shot Multi-Label Image RecognitionYicheng Liu, Jie Wen, Chengliang Liu, Xiaozhao Fang 等ICML 2024 · 被引用 7 次
