Lune

ICDE2026Top-tier venue

DistVec: Efficient Distributed Machine Learning in Parallel Database Systems

Xinyi Zhang, Liangzu Liu, Xupeng Miao, Yinjun Wu, Zhen Chen, Wei Lu, Xiaoyong Du, Bin Cui

2026Year

Abstract

Embedding vectors are widely used in various database-related domains. However, efficiently training largescale embeddings directly on DB-resident data remains a challenge, despite the ongoing efforts on supporting machine learning algorithms in database management systems, known as in-DBMS ML. Although current UDAF-based (User-Defined Aggregate Function) approach can support distributed training across parallel DBMS nodes, it enforces strict consistency during model synchronization, leading to high synchronization overhead. Additionally, the incompatibility between DBMS storage architectures and ML workloads further complicates efficient integration. To address these challenges, we propose DistVec, an inDBMS ML framework for efficient distributed training. DistVec introduces multi-query training primitives, that map distributed ML training tasks into concurrent database queries, allowing bounded staleness in model parameters to reduce synchronization costs. Besides, DistVec features an ML-customized model processing pipeline in database, including model materialization, buffer pool eviction, and hotness-aware caching, which minimize disk I/O and remote communication for the ML training workload. Experiments on the Greenplum database show that DistVec achieves an average speedup of 4.53× over UDAF-based training approaches while preserving the model quality.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get 5e14bffe-11cd-442d-b113-b8de30522a8d

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines