Systematically Exploring Associations among Multivariate Data
Lifeng Zhang
摘要
Detecting relationships among multivariate data is often of great importance in the analysis of high-dimensional data sets, and has received growing attention for decades from both academic and industrial fields. In this study, we propose a statistical tool named the neighbor correlation coefficient (nCor), which is based on a new idea that measures the local continuity of the reordered data points to quantify the strength of the global association between variables. With sufficient sample size, the new method is able to capture a wide range of functional relationship, whether it is linear or nonlinear, bivariate or multivariate, main effect or interaction. The score of nCor roughly approximates the coefficient of determination (R) of the data which implies the proportion of variance in one variable that is predictable from one or more other variables. On this basis, three nCor based statistics are also proposed here to further characterize the intra and inter structures of the associations from the aspects of nonlinearity, interaction effect, and variable redundancy. The mechanisms of these measures are proved in theory and demonstrated with numerical analyses. Introduction Identifying relationships among variables is one of the most critical issues in data analysis and interpretation (Altman and Krzywinski 2015) with a wide range of applications in diverse fields from data science to neuroscience. Nowadays, however, a large data set may contain a vast number of variable pairs and combinations that are difficult to be examined manually (Reshef et al. 2011). Association measures can be used to quickly find out the significant associations scattered in thousands or even millions of potential relationships without modelling the relationships explicitly, and thereby provide valuable knowledge and promising pointers for future study. Consider a data sample (x(t), y(t))|1≤t≤N that is observed from an underlying functional relationship expressed as follows. y = f(x) + e = ∑
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Absolute Neighbour Difference based Correlation Test for Detecting Heteroscedastic RelationshipsLifeng ZhangNeurIPS 2021 · 被引用 1 次
- Categorical Neighbour Correlation Coefficient (CnCor) for Detecting Relationships between Categorical VariablesLifeng Zhang, Shimo Yang, Hongxun JiangAAAI 2022
相关 Paper
- SHGR: A Generalized Maximal Correlation CoefficientSamuel Stocksieker, Denys PommeretNeurIPS 2025
- Measuring Dependence with Matrix-based Entropy FunctionalShujian Yu, Francesco Alesiani, Xi Yu, Robert Jenssen 等AAAI 2021 · 被引用 36 次
- Adjust Pearson's to Measure Arbitrary Monotone DependenceXinbo AiNeurIPS 2024 · 被引用 1 次
- Multivariate correlations discovery in static and streaming dataKoen Minartz, Jens E. d'Hondt, Odysseas PapapetrouVLDB 2022 · 被引用 4 次
- Intrinsic Dimension Correlation: uncovering nonlinear connections in multimodal representationsLorenzo Basile, Santiago Acevedo, Luca Bortolussi, Fabio Anselmi 等ICLR 2025 · 被引用 1 次
