On Aligning Tuples for Regression
Chenguang Fang, Shaoxu Song, Yinan Mei, Ye Yuan, Jianmin Wang
摘要
Regression models are learned over multiple variables, e.g., using engine torque and speed to predict its fuel consumption. In practice, the values of these variables are often collected separately, e.g., by different sensors in a vehicle, and need to be aligned first in a tuple before learning. Unfortunately, flowing to various issues like network delays, values generated at the same time could be recorded with different timestamps, making the alignment diffcult. According to our study in a vehicle manufacturer, engine torque, speed and fuel consumption values are mostly not recorded with the same timestamps. Aligning tuples by simply concatenating values of variables with equal timestamps leads to limited data for learning regression model. To deal with timestamp variations, existing time series matching techniques rely on the similarity of values and timestamps, which unfortunately are very likely to be absent among the variables in regression (no similarity between engine torque and speed values). In this sense, we propose to bridge tuple alignment and regression. Rather than similar values and timestamps, we align the values of different variables in a tuple that (i) are recorded in a short period, i.e., time constraint, and more importantly (ii) coincide well with the regression model, known as model constraint. Our theoretical and technical contributions include (1) formulating the problem of tuple alignment with time and model constraints, (2) proving NP-completeness of the problem, (3) devising an approximation algorithm with performance guarantee, and (4) proposing efficient pruning strategies for the algorithm. Experiments over real world datasets, including the aforesaid engine data collected by a vehicle manufacturer, demonstrate that our proposal outperforms the existing methods on alignment accuracy and improves regression precision.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper1
相关 Paper
- The Best of Both Worlds: On Repairing Timestamps and Attribute Values for Multivariate Time SeriesJingyu Zhu, Weiwei Deng, Yu Sun, Shaoxu Song 等SIGMOD 2025
- Imputing Various Incomplete Attributes via Distance Likelihood MaximizationShaoxu Song, Yu SunKDD 2020 · 被引用 15 次
- Multivariate Time Series Cleaning under Speed ConstraintsAoqian Zhang, Zexue Wu, Yifeng Gong, Ye Yuan 等SIGMOD 2025 · 被引用 4 次
- Representing Temporal Attributes for Schema MatchingYinan Mei, Shaoxu Song, Yunsu Lee, Jungho Park 等KDD 2020 · 被引用 2 次
- Cleaning Time Series under Seasonal and Trend ConstraintsZijie Chen, Aoqian Zhang, Shaoxu SongSIGMOD 2026
