STD-Former: Image-Conditioned Texture Dictionary Encoding with Sparse Topological Supervision for Texture Recognition
Bo Peng, Ke Xu, Yurui Pan
Abstract
Texture recognition is often framed as matching an image to a static training-set dictionary or codebook. In practice, this assumption is brittle: label-preserving transformations (illumination, scale, compression, blur) can shift test features away from the fixed training dictionary, producing a training-set codebook misalignment that limits accuracy. We propose STD-Former (Simple Texture Dictionary Transformer), a lightweight framework for image-conditioned texture dictionary encoding. Instead of comparing against a static codebook, STD-Former extracts a compact set of Intrinsic Textons (dictionary atoms / codewords) from the input image itself, yielding self-aligned representations at inference. Our design is intentionally simple and uses a decoupled two-stage recipe. In Stage 1, a Texture Dictionary Extractor (TDE) is pre-trained with a self-supervised Texton Coverage Loss that encourages the learned textons to collectively cover the image patch feature manifold. In Stage 2, a classifier is trained on the encoded dictionary representation; optionally, we add a Sparse Topological Loss derived from 0D persistent homology, which is equivalent to supervising only the (B-1) edges of a minimum spanning tree (MST) in each batch, providing efficient structure regularization. Across six standard texture benchmarks, STD-Former and STD-Former+ achieve new state-of-the-art results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a1cc7f2c-5e8b-4c16-98d7-2745e183aea3Builds on11
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Towards Total Recall in Industrial Anomaly DetectionKarsten Roth, Latha Pemula, Joaquin Zepeda, Bernhard Schölkopf et al.CVPR 2022 · 1,301 citations
- Topological AutoencodersMichael Moor, Max Horn, Bastian Rieck, Karsten M. BorgwardtICML 2020 · 192 citations
- Representation Topology Divergence: A Method for Comparing Neural Network RepresentationsSerguei Barannikov, Ilya Trofimov, Nikita Balabin, Evgeny BurnaevICML 2022 · 69 citations
- Link Prediction with Persistent Homology: An Interactive ViewZuoyu Yan, Tengfei Ma, Liangcai Gao, Zhi Tang et al.ICML 2021 · 59 citations
Related papers
- Learning Deblurring Texture Prior From Unpaired Data with Diffusion ModelChengxu Liu, Lu Qi, Jinshan Pan, Xueming Qian et al.ICCV 2025 · 4 citations
- Texture Reformer: Towards Fast and Universal Interactive Texture TransferZhizhong Wang, Lei Zhao, Haibo Chen, Ailin Li et al.AAAI 2022 · 20 citations
- Geometrized Transformer for Self-Supervised Homography EstimationJiazhen Liu, Xirong LiICCV 2023 · 27 citations
- Wukong's 72 Transformations: High-fidelity Textured 3D Morphing via Flow ModelsMinghao Yin, Yukang Cao, Kai HanNeurIPS 2025 · 5 citations
- Chebyshev Attention Depth Permutation Texture Network with Latent Texture Attribute LossRavishankar Evani, Deepu Rajan, Shangbo MaoCVPR 2025
