The Composable Data Management System Manifesto
Pedro Pedreira, Orri Erling, Konstantinos Karanasos, Scott Schneider, Wes McKinney, Satyanarayana R. Valluri, Mohamed Zaït, Jacques Nadeau
Abstract
The requirement for specialization in data management systems has evolved faster than our software development practices. After decades of organic growth, this situation has created a siloed landscape composed of hundreds of products developed and maintained as monoliths, with limited reuse between systems. This fragmentation has resulted in developers often reinventing the wheel, increased maintenance costs, and slowed down innovation. It has also affected the end users, who are often required to learn the idiosyncrasies of dozens of incompatible SQL and non-SQL API dialects, and settle for systems with incomplete functionality and inconsistent semantics. In this vision paper, considering the recent popularity of open source projects aimed at standardizing different aspects of the data stack, we advocate for a paradigm shift in how data management systems are designed. We believe that by decomposing these into a modular stack of reusable components, development can be streamlined while creating a more consistent experience for users. Towards that goal, we describe the state-of-the-art, principal open source technologies, and highlight open questions and areas where additional research is needed. We hope this work will foster collaboration, motivate further research, and promote a more composable future for data management.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- PilotScope: Steering Databases with Machine Learning DriversRong Zhu, Lianggui Weng, Wenqing Wei, Di Wu et al.VLDB 2024 · 18 citations
- BOSS - An Architecture for Database Kernel CompositionHubert Mohr-Daurat, Xuan Sun, Holger PirkVLDB 2024 · 12 citations
- Active Data Lakes: Regaining Physical Data Independence Without Losing InteroperabilityPascal Ginter, Viktor LeisVLDB 2026 · 4 citations
- Anarchy in the Database: A Survey and Evaluation of Database Management System ExtensibilityAbigale Kim, Marco Slot, David G. Andersen, Andrew PavloVLDB 2025 · 4 citations
- Fast and Scalable Data Transfer Across Data SystemsHaralampos Gavriilidis, Kaustubh Beedkar, Matthias Boehm, Volker MarklSIGMOD 2025 · 4 citations
Builds on1
Related papers
- Towards Designing Future-Proof Data Processing SystemsMichael Jungmair, Jana GicevaVLDB 2025 · 1 citation
- Data Management in Microservices: State of the Practice, Challenges, and Research DirectionsRodrigo N. Laigner, Yongluan Zhou, Marcos Antonio Vaz Salles, Yijian Liu et al.VLDB 2021 · 108 citations
- SparqLog: A System for Efficient Evaluation of SPARQL 1.1 Queries via DatalogRenzo Angles, Georg Gottlob, Aleksandar Pavlovic, Reinhard Pichler et al.VLDB 2023 · 4 citations
- Database Technology for the Masses: Sub-Operators as First-Class EntitiesMaximilian Bandle, Jana GicevaVLDB 2021 · 19 citations
- Declarative Sub-Operators for Universal Data ProcessingMichael Jungmair, Jana GicevaVLDB 2023 · 17 citations
