Synbols: Probing Learning Algorithms with Synthetic Datasets
Alexandre Lacoste, Pau Rodríguez López, Frederic Branchaud-Charron, Parmida Atighehchian, Massimo Caccia, Issam Hadj Laradji, Alexandre Drouin, Matt Craddock, Laurent Charlin, David Vázquez
Abstract
Progress in the field of machine learning has been fueled by the introduction of benchmark datasets pushing the limits of existing algorithms. Enabling the design of datasets to test specific properties and failure modes of learning algorithms is thus a problem of high interest, as it has a direct impact on innovation in the field. In this sense, we introduce Synbols -- Synthetic Symbols -- a tool for rapidly generating new datasets with a rich composition of latent features rendered in low resolution images. Synbols leverages the large amount of symbols available in the Unicode standard and the wide range of artistic font provided by the open font community. Our tool's high-level interface provides a language for rapidly generating new distributions on the latent features, including various types of textures and occlusions. To showcase the versatility of Synbols, we use it to dissect the limitations and flaws in standard learning algorithms in various learning setups including supervised learning, active learning, out of distribution generalization, unsupervised representation learning, and object counting.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b3ccfda3-8763-4215-b06f-a32c8f95713dCited by top-tier papers5
- Online Fast Adaptation and Knowledge Accumulation (OSAKA): a New Approach to Continual LearningMassimo Caccia, Pau Rodríguez, Oleksiy Ostapenko, Fabrice Normandin et al.NeurIPS 2020 · 83 citations
- Beyond Trivial Counterfactual Explanations with Diverse Valuable ExplanationsPau Rodríguez, Massimo Caccia, Alexandre Lacoste, Lee Zamparo et al.ICCV 2021 · 72 citations
- On the Stability-Plasticity Dilemma in Continual Meta-Learning: Theory and AlgorithmQi Chen, Changjian Shui, Ligong Han, Mario MarchandNeurIPS 2023 · 32 citations
- Diverse, Global and Amortised Counterfactual Explanations for Uncertainty EstimatesDan Ley, Umang Bhatt, Adrian WellerAAAI 2022 · 25 citations
- MIND: Multi-Task Incremental Network DistillationJacopo Bonato, Francesco Pelosin, Luigi Sabetta, Alessandro NicolosiAAAI 2024 · 18 citations
Builds on3
- Out-of-Distribution Generalization via Risk Extrapolation (REx)David Krueger, Ethan Caballero, Jörn-Henrik Jacobsen, Amy Zhang et al.ICML 2021 · 1,163 citations
- Invariant Risk Minimization GamesKartik Ahuja, Karthikeyan Shanmugam, Kush R. Varshney, Amit DhurandharICML 2020 · 289 citations
- Weakly Supervised Disentanglement with GuaranteesRui Shu, Yining Chen, Abhishek Kumar, Stefano Ermon et al.ICLR 2020 · 148 citations
Related papers
- SynFER: Towards Boosting Facial Expression Recognition With Synthetic DataXilin He, Cheng Luo, Xiaole Xian, Bing Li et al.ICCV 2025 · 6 citations
- Procedural Image Programs for Representation LearningManel Baradad, Chun-Fu Richard Chen, Jonas Wulff, Tongzhou Wang et al.NeurIPS 2022 · 40 citations
- Arti-PG: A Toolbox for Procedurally Synthesizing Large-Scale and Diverse Articulated Objects with Rich AnnotationsJianhua Sun, Yuxuan Li, Jiude Wei, Longfei Xu et al.ICCV 2025 · 3 citations
- Forte : Finding Outliers with Representation Typicality EstimationDebargha Ganguly, Warren Richard Morningstar, Andrew Seohwan Yu, Vipin ChaudharyICLR 2025
- UniCode: Augmenting Evaluation for Code ReasoningXinyue Zheng, Haowei Lin, Shaofei Cai, Yaodong Yang et al.ICML 2026 · 1 citation
