Systematically Exploring Associations among Multivariate Data
Lifeng Zhang
Abstract
Detecting relationships among multivariate data is often of great importance in the analysis of high-dimensional data sets, and has received growing attention for decades from both academic and industrial fields. In this study, we propose a statistical tool named the neighbor correlation coefficient (nCor), which is based on a new idea that measures the local continuity of the reordered data points to quantify the strength of the global association between variables. With sufficient sample size, the new method is able to capture a wide range of functional relationship, whether it is linear or nonlinear, bivariate or multivariate, main effect or interaction. The score of nCor roughly approximates the coefficient of determination (R) of the data which implies the proportion of variance in one variable that is predictable from one or more other variables. On this basis, three nCor based statistics are also proposed here to further characterize the intra and inter structures of the associations from the aspects of nonlinearity, interaction effect, and variable redundancy. The mechanisms of these measures are proved in theory and demonstrated with numerical analyses. Introduction Identifying relationships among variables is one of the most critical issues in data analysis and interpretation (Altman and Krzywinski 2015) with a wide range of applications in diverse fields from data science to neuroscience. Nowadays, however, a large data set may contain a vast number of variable pairs and combinations that are difficult to be examined manually (Reshef et al. 2011). Association measures can be used to quickly find out the significant associations scattered in thousands or even millions of potential relationships without modelling the relationships explicitly, and thereby provide valuable knowledge and promising pointers for future study. Consider a data sample (x(t), y(t))|1≤t≤N that is observed from an underlying functional relationship expressed as follows. y = f(x) + e = ∑
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 30ca1e1e-b924-4cef-b22f-0efd7b0d1126Cited by top-tier papers2
- Absolute Neighbour Difference based Correlation Test for Detecting Heteroscedastic RelationshipsLifeng ZhangNeurIPS 2021 · 1 citation
- Categorical Neighbour Correlation Coefficient (CnCor) for Detecting Relationships between Categorical VariablesLifeng Zhang, Shimo Yang, Hongxun JiangAAAI 2022
Related papers
- SHGR: A Generalized Maximal Correlation CoefficientSamuel Stocksieker, Denys PommeretNeurIPS 2025
- Measuring Dependence with Matrix-based Entropy FunctionalShujian Yu, Francesco Alesiani, Xi Yu, Robert Jenssen et al.AAAI 2021 · 36 citations
- Adjust Pearson's to Measure Arbitrary Monotone DependenceXinbo AiNeurIPS 2024 · 1 citation
- Multivariate correlations discovery in static and streaming dataKoen Minartz, Jens E. d'Hondt, Odysseas PapapetrouVLDB 2022 · 4 citations
- Intrinsic Dimension Correlation: uncovering nonlinear connections in multimodal representationsLorenzo Basile, Santiago Acevedo, Luca Bortolussi, Fabio Anselmi et al.ICLR 2025 · 1 citation
