Attend to Anything: Foundation Model for Unified Human Attention Modeling
Wenzhuo Zhao, Ronghao Xian, Keren Fu, Qijun Zhao
摘要
Existing human attention (saliency) modeling methods persist as highly fragmented across modalities, scenes, and task formulations. Consequently, even with increasing model capacity and data scale, current models predominantly remain scene-dependent and task-specific, failing to practically generalize in real-world applications. To address the fundamental limitations, we present the Attend to Anything Model (AAM), a multi-modal foundation model that unifies attention modeling across various image, video, and audio-visual tasks and scenes. AAM reformulates attention as a cognitive entailment relationship organized in a general-to-specific hierarchy, implemented through language prompts with hierarchical embeddings in hyperbolic space. Furthermore, to unify static image and dynamic video attention, we adopt a fluid-dynamics perspective, formulating video-frame attention as a diffusive temporal evolution governed by the Fokker--Planck equation. Extensive experiments on 16 benchmarks demonstrate that AAM consistently outperforms state-of-the-art methods by an average of 6% across various scenarios, while achieving approximately a 4 speedup in video inference. Overall, these results demonstrate that AAM provides a principled foundation for future research on attention and saliency-related tasks. The dataset and code will be available at https://github.com/wz-zhao/Attend-to-Anything.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- TASED-Net: Temporally-Aggregating Spatial Encoder-Decoder Network for Video Saliency DetectionKyle Min, Jason J. CorsoICCV 2019 · 被引用 189 次
- Hyperbolic Image-text RepresentationsKaran Desai, Maximilian Nickel, Tanmay Rajpurohit, Justin Johnson 等ICML 2023 · 被引用 137 次
- UEyes: Understanding Visual Saliency across User Interface TypesYue Jiang, Luis A. Leiva, Hamed Rezazadegan Tavakoli, Paul R. B. Houssel 等CHI 2023 · 被引用 100 次
- DeepGaze IIE: Calibrated prediction in and out-of-domain for state-of-the-art saliency modelingAkis Linardos, Matthias Kümmerer, Ori Press, Matthias BethgeICCV 2021 · 被引用 98 次
相关 Paper
- Unlocking Slot Attention by Changing Optimal Transport CostsYan Zhang, David W. Zhang, Simon Lacoste-Julien, Gertjan J. Burghouts 等ICML 2023 · 被引用 20 次
- OmniGen-AR: AutoRegressive Any-to-Image GenerationJunke Wang, Xun Wang, Qiushan Guo, Peize Sun 等NeurIPS 2025 · 被引用 7 次
- DiffSal: Joint Audio and Video Learning for Diffusion Saliency PredictionJunwen Xiong, Peng Zhang, Tao You, Chuanyue Li 等CVPR 2024 · 被引用 12 次
- Towards Fine-Grained Interactive Segmentation in Images and VideosYuan Yao, Qiushi Yang, Miaomiao Cui, Liefeng BoICCV 2025 · 被引用 2 次
- Hierarchical Self-Attention: Generalizing Neural Attention Mechanics to Multi-Scale ProblemsSaeed Amizadeh, Sara Abdali, Yinheng Li, Kazuhito KoishidaNeurIPS 2025 · 被引用 2 次
