SAT-RRG: LLM-Guided Self-Adaptive Training for Radiology Report Generation with Token-Level Push-Pull Optimization
Yunyi Liu, Yingshu Li, Tong Chen, Lingqiao Liu, Lei Wang, Luping Zhou
Abstract
Radiology report generators often produce fluent text yet miss crucial details, leading to local semantic conflicts or flipped findings that require stronger penalties. Crossentropy (CE) merely increases the probability of the ground-truth token y * without directly suppressing the model's current wrong choice ŷ, and treats all positions uniformly, so corrections are not prioritized. We introduce a self-adaptive optimization framework that dynamically adjusts token-level gradients based on semantic discrepancy cues derived from a frozen LLM referee. The LLM itself is not the contribution-it merely provides weak supervision to trigger the adaptive learning process. Within this framework, (i) semantic conflicts between the predicted and reference reports are automatically localized and tagged with <e>...</e> (used only during training), and (ii) adaptive, stronger penalties are applied within these sparse but critical spans. Updates follow a push-pull scheme: error spans are pushed down, while non-error tokens are reinforced. The update strength is governed by two complementary signals-normalized entropy (for uncertainty calibration) and focal-style confidence (for handling overand under-confident predictions). On MIMIC-CXR and IU-Xray, our framework consistently improves both language metrics (BLEU-4, ROUGE-L, METEOR) and clinical metrics (RadGraph F1, CheXbert), and remains robust to noisy or imperfect error tags.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on18
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- Generating Radiology Reports via Memory-driven TransformerZhihong Chen, Yan Song, Tsung-Hui Chang, Xiang WanEMNLP 2020 · 552 citations
Related papers
- Rethinking Radiology Report Generation: From Narrative Flow to Topic-Guided FindingsSheng Cheng, Devika SubramanianICLR 2026
- LLM-RG4: Flexible and Factual Radiology Report Generation Across Diverse Input ContextsZhuhao Wang, Yihua Sun, Zihan Li, Xuan Yang et al.AAAI 2025 · 6 citations
- S2D-Align: Shallow-to-Deep Auxiliary Learning for Anatomically-Grounded Radiology Report GenerationJiechao Gao, Chang Liu, Yuangang LiAAAI 2026
- The Double Dilemma in Multi-Task Radiology Report Generation: A Gradient Dynamics Analysis and SolutionErjian Zhang, Yatong Hao, Liejun Wang, Zhiqing GuoICML 2026
- Report-Concept Textual-Prompt Learning for Enhancing X-ray DiagnosisXiongjun Zhao, Zhengyu Liu, Fen Liu, Guanting Li et al.ACM MM 2024 · 3 citations
