Efficient Truncated Linear Regression with Unknown Noise Variance
Constantinos Daskalakis, Patroklos Stefanou, Rui Yao, Emmanouil Zampetakis
Abstract
Truncated linear regression is a classical challenge in statistics, wherein a label, 𝑦 = 𝑤 𝑇 𝑥 + 𝜀 , and its corresponding feature vector, 𝑥 ∈ R 𝑘 , are only observed if the label falls in some subset 𝑆 ⊆ R ; otherwise the existence of the pair ( 𝑥, 𝑦 ) is hidden from observation. Linear regression with truncated observations has remained a challenge, in its general form, since the early works of Tobin (1958); Amemiya (1973). When the distribution of the error is normal with known variance, recent work of Daskalakis et al. (2019) provides computationally and statistically efficient estimators of the linear model, 𝑤 . In this paper, we provide the first computationally and statistically efficient estimators for truncated linear regression when the noise variance is unknown, estimating both the linear model and the variance of the noise. Our estimator is based on an efficient implementation of Projected Stochastic Gradient Descent on the negative log-likelihood of the truncated sample. Importantly, we show that the error of our estimates is asymptotically normal, and we use this to provide explicit confidence regions for our estimates.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 05c7bc30-aac5-4b41-b9ec-44079e902613Cited by top-tier papers7
- Domain constraints improve risk prediction when outcome data is missingSidhika Balachandar, Nikhil Garg, Emma PiersonICLR 2024 · 11 citations
- Learning Exponential Families from Truncated SamplesJane H. Lee, Andre Wibisono, Emmanouil ZampetakisNeurIPS 2023 · 7 citations
- Smoothed Analysis of Learning from Positive SamplesJane H. Lee, Anay Mehrotra, Manolis ZampetakisSTOC 2026 · 2 citations
- Detecting Low-Degree TruncationAnindya De, Huan Li, Shivam Nadimpalli, Rocco A. ServedioSTOC 2024 · 2 citations
- Linear Regression with Unknown Truncation Beyond Gaussian FeaturesAlexandros Kouridakis, Anay Mehrotra, Alkis Kalavasis, Constantine CaramanisICML 2026 · 2 citations
Related papers
- Truncated Linear Regression in High DimensionsConstantinos Daskalakis, Dhruv Rohatgi, Emmanouil ZampetakisNeurIPS 2020 · 19 citations
- Oracle efficient truncated statisticsKonstantinos Karatapanis, Vasilis Kontonis, Christos TzamosICLR 2025
- Sparse Linear Regression Is Easy on Random SupportsGautam Chandrasekaran, Raghu Meka, Konstantinos StavropoulosSTOC 2026
- Information-Computation Tradeoffs for Noiseless Linear Regression with Oblivious ContaminationIlias Diakonikolas, Chao Gao, Daniel Kane, John D. Lafferty et al.NeurIPS 2025
- Label Robust and Differentially Private Linear Regression: Computational and Statistical EfficiencyXiyang Liu, Prateek Jain, Weihao Kong, Sewoong Oh et al.NeurIPS 2023 · 10 citations
