FusionQuery: On-demand Fusion Queries over Multi-source Heterogeneous Data
Junhao Zhu, Yuren Mao, Lu Chen, Congcong Ge, Ziheng Wei, Yunjun Gao
Abstract
Centralised data management systems (e.g., data lakes) support queries over multi-source heterogeneous data. However, the query results from multiple sources commonly involve between-source conflicts, which makes query results unreliable and confusing and degrades the usability of centralised data management systems. Therefore, resolving the between-sourced conflicts is one of the most important problems for centralised data management systems. To solve it, many batch data fusion-based methods have been proposed, which require traversing all the data in the centralised data management systems and cause scalability and flexibility issues. To address these issues, this paper explores the problem of on-demand fusion queries, where the between-sourced conflicts are solved with only the query-related data; moreover, we propose an efficient on-demand fusion query framework, FusionQuery, which consists of a query stage and a fusion stage. In the query stage, we frame the heterogeneous data query problem as a knowledge graph matching problem and present a line graph-based method to accelerate it. In the fusion stage, we develop an Expectation Maximization-style algorithm to iteratively updates data veracity and source trustworthiness. Furthermore, we design an incremental estimation method of source trustworthiness to address the lack of sufficient observations. Extensive experiments on two real-world datasets demonstrate that FusionQuery outperforms state-of-the-art data fusion methods in terms of both effectiveness and efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 08f5a2f8-92f3-44e1-b71e-2cf8f61baca8Cited by top-tier papers3
- MultiRAG: A Knowledge-Guided Framework for Mitigating Hallucination in Multi-Source Retrieval Augmented GenerationWenlong Wu, Haofen Wang, Bohan Li, Peixuan Huang et al.ICDE 2025 · 16 citations
- When Traditional Medicine Meets AI: Critical Considerations for AI-Empowered Clinical Support in Traditional MedicineYuling Sun, Wenjing Yue, Xiaofu Jin, Shuai Ma et al.CSCW 2025 · 3 citations
- IncreQueryFusion: On-demand Data Fusion Framework in Dynamic Data LakesWenhao Liu, Sai Wu, Xiu Tang, Yitong Zhang et al.VLDB 2026
Builds on11
- Deep Entity Matching with Pre-Trained Language ModelsYuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan et al.VLDB 2021 · 484 citations
- Deep Learning for Blocking in Entity Matching: A Design Space ExplorationSaravanan Thirumuruganathan, Han Li, Nan Tang, Mourad Ouzzani et al.VLDB 2021 · 109 citations
- GSI: GPU-friendly Subgraph IsomorphismLi Zeng, Lei Zou, M. Tamer Özsu, Lin Hu et al.ICDE 2020 · 62 citations
- PromptEM: Prompt-tuning for Low-resource Generalized Entity MatchingPengfei Wang, Xiaocan Zeng, Lu Chen, Fan Ye et al.VLDB 2023 · 39 citations
- ClusterEA: Scalable Entity Alignment with Stochastic Training and Normalized Mini-batch SimilaritiesYunjun Gao, Xiaoze Liu, Junyang Wu, Tianyi Li et al.KDD 2022 · 36 citations
Related papers
- Subgraph Matching over Graph FederationYe Yuan, Delong Ma, Zhenyu Wen, Zhiwei Zhang et al.VLDB 2022 · 24 citations
- An Effective Framework for Enhancing Query Answering in a Heterogeneous Data LakeQin Yuan, Ye Yuan, Zhenyu Wen, He Wang et al.SIGIR 2023 · 7 citations
- Efficient Online Estimation of Causal Effects by Deciding What to ObserveShantanu Gupta, Zachary C. Lipton, David ChildersNeurIPS 2021 · 19 citations
- MGRAG: Semantic Subgraph Matching and Graph-Aware Caching for Multimodal Retrieval-Augmented GenerationYubo Wang, Haoyang Li, Lei ChenVLDB 2026
- Entity Resolution On-DemandGiovanni Simonini, Luca Zecchini, Sonia Bergamaschi, Felix NaumannVLDB 2022 · 33 citations
