Ember: No-Code Context Enrichment via Similarity-Based Keyless Joins
Sahaana Suri, Ihab F. Ilyas, Christopher Ré, Theodoros Rekatsinas
Abstract
Structured data, or data that adheres to a pre-defined schema, can suffer from fragmented context: information describing a single entity can be scattered across multiple datasets or tables tailored for specific business needs, with no explicit linking keys (e.g., primary key-foreign key relationships or heuristic functions). Context enrichment, or rebuilding fragmented context, using keyless joins is an implicit or explicit step in machine learning (ML) pipelines over structured data sources. This process is tedious, domain-specific, and lacks support in now-prevalent no-code ML systems that let users create ML pipelines using just input data and high-level configuration files. In response, we propose Ember, a system that abstracts and automates keyless joins to generalize context enrichment. Our key insight is that Ember can enable a general keyless join operator by constructing an index populated with task-specific embeddings. Ember learns these embeddings by leveraging Transformer-based representation learning techniques. We describe our core architectural principles and operators when developing Ember, and empirically demonstrate that Ember allows users to develop no-code pipelines for five domains, including search, recommendation and question answering, and can exceed alternatives by up to 39% recall, with as little as a single line configuration change.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- CodexDB: Synthesizing Code for Query Processing from Natural Language Instructions using GPT-3 CodexImmanuel TrummerVLDB 2022 · 77 citations
- DeepJoin: Joinable Table Discovery with Pre-trained Language ModelsYuyang Dong, Chuan Xiao, Takuma Nozawa, Masafumi Enomoto et al.VLDB 2023 · 53 citations
- E2ETune: End-to-End Knob Tuning via Fine-tuned Generative Language ModelXinmei Huang, Haoyang Li, Jing Zhang, Xinxin Zhao et al.VLDB 2025 · 15 citations
- Cross Modal Data Discovery over Structured and Unstructured Data LakesMohamed Y. Eltabakh, Mayuresh Kunjir, Ahmed K. Elmagarmid, Mohammad Shahmeer AhmadVLDB 2023 · 11 citations
- BLEND: A Unified Data Discovery SystemMahdi Esmailoghli, Christoph Schnell, Renée J. Miller, Ziawasch AbedjanICDE 2025 · 4 citations
Builds on7
- TabNet: Attentive Interpretable Tabular LearningSercan Ö. Arik, Tomas PfisterAAAI 2021 · 2,148 citations
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 1,246 citations
- Generalization through Memorization: Nearest Neighbor Language ModelsUrvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer et al.ICLR 2020 · 1,038 citations
- Deep Entity Matching with Pre-Trained Language ModelsYuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan et al.VLDB 2021 · 484 citations
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 417 citations
Related papers
- Optimizing Context-Enhanced Relational JoinsViktor Sanca, Manos Chatzakis, Anastasia AilamakiICDE 2024 · 5 citations
- CtxPipe: Context-aware Data Preparation Pipeline Construction for Machine LearningHaotian Gao, Shaofeng Cai, Tien Tuan Anh Dinh, Zhiyong Huang et al.SIGMOD 2025 · 7 citations
- Multitask Pretraining with Structured Knowledge for Text-to-SQL GenerationRobert Giaquinto, Dejiao Zhang, Benjamin Kleiner, Yang Li et al.ACL 2023 · 6 citations
- A Scalable AutoML Approach Based on Graph Neural NetworksMossad Helali, Essam Mansour, Ibrahim Abdelaziz, Julian Dolby et al.VLDB 2022 · 16 citations
- TabEmb: Joint Semantic-Structure Embedding for Table AnnotationEhsan Hoseinzade, Ke Wang, Anandharaju Durai RajuACL 2026
