Lune

SC2025Top-tier venue

UltraAttn: Efficiently Parallelizing Attention through Hierarchical Context-Tiling

Haoyu Yang, Zan Zong, Yuyang Jin, Kinman Lei, Jiaao He, Qigang Yang, Jidong Zhai

2025Year
1Citations

Abstract

Long-context comprehension is critical for large language models. Context parallelism and irregular block-sparse attention are keyss to accelerating long-context training and inference. Existing context parallelism suffers from poor scalability due to the striped-like partition pattern, which causes high communication traffic, and the ring-based communication pattern, which limits kernel granularity, reduces device utilization, and incurs redundant communication.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get 710bbdb4-0c22-4229-b5de-6c4ecad9f7fd

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines