ACL2026

CaRVE: Critiquing and Refining Visual Elaborations for Figurative Language Illustrations

Manishit Kundu, Tejomay Kishor Padole, Sumit Shekhar, Biplab Banerjee, Pushpak Bhattacharyya

Abstract

Illustrating figurative language remains challenging due to its non-literal semantics, and existing text-to-image frameworks rely heavily on proprietary models or human supervision to achieve adequate alignment.We introduce CaRVE, a lightweight and fully open-source critique-driven framework that employs VLM feedback to refine visual elaborations for figurative image generation.CaRVE bridges the semantic alignment gap even in sub-4B models by correcting visual and conceptual misalignments, reducing over-literalization, and improving robustness to complex figurative expressions.Using only open-source models, CaRVE achieves a 6.49% improvement over prior baselines on intrinsic automatic evaluations and a +0.37 average rank gain in human preference.We further release MetaCaRVE, an enhanced figurative image dataset constructed by refining HAIVMet using CaRVE 1 .1 https://github.com/manishitkundu/CaRVE_ACL_ 2026 language models and smaller, open-source LLMs.We posit that this semantic gap can be effectively bridged in smaller models through a VLM-critiquedriven feedback mechanism.In this work, we introduce Critiquing-and-Refining Visual Elaborations (CaRVE), a VLM-feedback-driven framework for figurative text-to-image generation.CaRVE employs VLM-based critique to iteratively refine visual elaborations for image generation, yielding significantly improved semantic alignment even with smaller sub-4B models.