Stochastic Interpolants for Revealing Stylistic Flows Across the History of Art
Pingchuan Ma, Ming Gui, Johannes Schusterbauer, Xiaopei Yang, Olga Grebenkova, Vincent Tao Hu, Björn Ommer
Abstract
Generative models have made rapid progress in content creation, particularly in synthesizing artworks and capturing stylistic variation. However, most methods operate at the level of individual images, limiting their ability to reveal broader stylistic trends and temporal transitions. We address this by introducing a framework that models stylistic evolution as an optimal transport problem in a learned style space, using stochastic interpolants and dual diffusion implicit bridges to align artistic distributions across time without requiring paired data. A central contribution is a diverse dataset of over 650,000 artworks spanning 500 years, curated with metadata across multiple genres. Together, our method and dataset enable tracing long-range stylistic transitions and plausible futures of individual artworks, supporting fine-grained temporal analysis. This offers a new tool for modeling historical patterns in visual culture and opens up promising directions in visual understanding. Code and dataset: https://github.com/CompVis/Art-fm.
- Equal Contribution lution of artistic movements and styles over time.
In this paper, we ask: How would a piece of art evolve? What might it look like if created centuries earlier or later? Naturally, no ground-truth pairs exist for such a task, and the artistic style of a creator is inherently dynamic and historically constrained. To address this challenge, we aim to analyze the flow of artistic styles over time. Given a single artwork, we try to determine its closest "projection" onto the manifold of artistic expressions corresponding to a given historical period. While optimal transport provides a natural formulation for such a task, considering only a single isolated artwork is insufficient. Instead, we reformulate the problem as a distribution matching task, where we study the evolution of the underlying distribution of art over time.
Recent advances in diffusion bridges such as DDIB [83], DSB [13,85], and DDBMs [100] have shown that domain translation can be performed implicitly, without the need for paired supervision. We build on this idea in the context of art history by incorporating stochastic interpolants [1], enabling us to model the temporal evolution of artistic styles without requiring explicitly aligned training data. Unlike traditional style transfer methods that transform images based on predefined and discrete styles, our approach projects styles onto their corresponding distributions across time. Operating in a compact latent space allows us to abstract away pixel-level distraction and focus on meaningful stylistic shifts.
To support this temporal modeling, we require a dataset that spans centuries and reflects a wide variety of artistic styles. While datasets like WikiArt [71] offer breadth in genre and metadata, they lack the scale and historical depth needed to model long-term stylistic progression. Others, such as BAM , are either restricted in access or skewed toward modern art. To address this, we curate a new unified dataset of over 650,000 artworks annotated with creation year, location, artist, medium, and other metadata (Fig. 1). This dataset enables the study of artistic evolution at scale and over time. In summary, our key contributions are as follows:
• We propose the first generative framework that explicitly models the temporal progression of artistic styles across This ICCV paper is the Open Access version, provided by the Computer Vision Foundation. Except for this watermark, it is identical to the accepted version; the final published version of the proceedings is available on IEEE Xplore.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a3c277d3-3bca-45ac-aac9-87933f387ba1Builds on31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- Crossing You in Style: Cross-modal Style Transfer from Music to Visual ArtsCheng-Che Lee, Wan-Yi Lin, Yen-Ting Shih, Pei-Yi (Patricia) Kuo et al.ACM MM 2020 · 16 citations
- AnyStyleDiffusion: Flexible Style Transfer with Consistent Content Adaptation Across Diffusion ModelsZhenyu Xu, Junjie Wu, Zhiyan Piao, Xiaoqi Sheng et al.ACM MM 2025
- Diverse Image Style Transfer via Invertible Cross-Space MappingHaibo Chen, Lei Zhao, Huiming Zhang, Zhizhong Wang et al.ICCV 2021 · 42 citations
- Diffusion Bridge or Flow Matching? A Unifying Framework and Comparative AnalysisKaizhen Zhu, Mokai Pan, Zhechuan Yu, Jingya Wang et al.ICML 2026 · 3 citations
- Magic Insert: Style-Aware Drag-And-DropNataniel Ruiz, Yuanzhen Li, Neal Wadhwa, Yael Pritch et al.ICCV 2025
