Generalized Neural Sorting Networks with Error-Free Differentiable Swap Functions
Jungtaek Kim, Jeongbeen Yoon, Minsu Cho
Abstract
Sorting is a fundamental operation of all computer systems, having been a long-standing significant research topic. Beyond the problem formulation of traditional sorting algorithms, we consider sorting problems for more abstract yet expressive inputs, e.g., multi-digit images and image fragments, through a neural sorting network. To learn a mapping from a high-dimensional input to an ordinal variable, the differentiability of sorting networks needs to be guaranteed. In this paper we define a softening error by a differentiable swap function, and develop an error-free swap function that holds a non-decreasing condition and differentiability. Furthermore, a permutation-equivariant Transformer network with multi-head attention is adopted to capture dependency between given inputs and also leverage its model capacity with self-attention. Experiments on diverse sorting benchmarks show that our methods perform better than or comparable to baseline methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4ac49fb2-e736-4642-8a73-343b4890c88bCited by top-tier papers4
- Learning Distributions over Permutations and Rankings with Factorized RepresentationsDaniel Severo, Brian Karrer, Niklas NolteICLR 2026 · 1 citation
- Learning Permutation Distributions via Reflected Diffusion on RanksSizhuang He, Yangtian Zhang, Shiyang Zhang, David van DijkICML 2026
- Learning to Rank by Directly Optimizing Full-Order ProbabilitiesYongxiang Tang, Chao Wang, Jincheng Lu, Yanhua Cheng et al.ICML 2026
- SymmetricDiffusers: Learning Discrete Diffusion on Finite Symmetric GroupsYongxing Zhang, Donglin Yang, Renjie LiaoICLR 2025
Builds on8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- PolyGen: An Autoregressive Generative Model of 3D MeshesCharlie Nash, Yaroslav Ganin, S. M. Ali Eslami, Peter W. BattagliaICML 2020 · 339 citations
- Fast Differentiable Sorting and RankingMathieu Blondel, Olivier Teboul, Quentin Berthet, Josip DjolongaICML 2020 · 285 citations
Related papers
- Differentiable Sorting Networks for Scalable Sorting and Ranking SupervisionFelix Petersen, Christian Borgelt, Hilde Kuehne, Oliver DeussenICML 2021 · 39 citations
- Monotonic Differentiable Sorting NetworksFelix Petersen, Christian Borgelt, Hilde Kuehne, Oliver DeussenICLR 2022 · 32 citations
- OPS: An Order-Preserving Sorting Network for Information RetrievalChao Wang, Yongxiang Tang, Guikai Luan, Kaiyuan Li et al.SIGIR 2026
- Divergence-Free Neural Networks with Application to Image DenoisingSébastien Herbreteau, Etienne MeunierICLR 2026
- LapSum - One Method to Differentiate Them All: Ranking, Sorting and Top-k SelectionLukasz Struski, Michal B. Bednarczyk, Igor T. Podolak, Jacek TaborICML 2025
