Visual Fact Checker: Enabling High-Fidelity Detailed Caption Generation
Yunhao Ge, Xiaohui Zeng, Jacob Samuel Huffman, Tsung-Yi Lin, Ming-Yu Liu, Yin Cui
摘要
A happy little girl is standing in a green field, wearing a plaid shirt and holding onto a string of pink balloons. The balloons are floating in the air, creating a playful and joyful atmosphere. The girl appears to be enjoying the moment, with a smile on her face. In the background, there is a vast expanse of green grass, stretching out as far as the eye can see. The sky above is a bright blue, with a few white clouds scattered about. The overall mood of the scene is one of carefree happiness and joy, with the pink balloons adding a touch of whimsy and playfulness to the atmosphere. The image shows a young girl in a field with her arms joyfully raised. She is holding a bunch of pink balloons in her right hand. She's wearing a plaid shirt and jeans, and her hair is styled in what appears to be a ponytail. The girl is smiling widely, and her pose suggests she is in mid-celebration or play. The field is filled with wildflowers, and trees can be seen in the background under a blue sky with some clouds. The setting appears to be a sunny day, perfect for outdoor activities. The image captures a moment of carefree childhood happiness. A 3D model of a three-level, intricately designed wooden tower with a contrasting blue roof and door, standing on a wooden base. The tower, brown in color, resembles a fusion of a house, a tower, and a castle. At the very top of the tower, there is a crescent moon design. The overall design adds a touch of fantasy to the scene. A 3D model of a small wooden tower with a blue roof.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- ExCap3d: Expressive 3D Scene Understanding via Object Captioning with Varying DetailChandan Yeshwanth, Dávid Rozenberszki, Angela DaiICCV 2025 · 被引用 2 次
- Tera: Rethinking Text-Guided Realistic 3D Avatar GenerationYanwen Wang, Yiyu Zhuang, Jiawei Zhang, Li Wang 等ICCV 2025 · 被引用 2 次
- A Comprehensive Study of Decoder-Only LLMs for Text-to-Image GenerationAndrew Z. Wang, Songwei Ge, Tero Karras, Ming-Yu Liu 等CVPR 2025
- Perception in ReflectionYana Wei, Liang Zhao, Kangheng Lin, En Yu 等ICML 2025
- Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and CoverageSaehyung Lee, Seunghyun Yoon, Trung Bui, Jing Shi 等ICML 2025
它引用的顶会 Paper19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
相关 Paper
- ArtEditor: Learning Customized Instructional Image Editor From Few-Shot ExamplesShijie Huang, Yiren Song, Yuxuan Zhang, Hailong Guo 等ICCV 2025 · 被引用 2 次
- InstanceDiffusion: Instance-Level Control for Image GenerationXudong Wang, Trevor Darrell, Sai Saketh Rambhatla, Rohit Girdhar 等CVPR 2024
- CineMaster: A 3D-Aware and Controllable Framework for Cinematic Text-to-Video GenerationQinghe Wang, Yawen Luo, Xiaoyu Shi, Xu Jia 等SIGGRAPH 2025 · 被引用 13 次
- PneuSeries: 3D Shape Forming with Modularized Serial-Connected InflatablesYu-Wen Chen, Wei-Ju Lin, Yi Chen, Lung-Pan ChengUIST 2021 · 被引用 21 次
- Can Playing with Toy Blocks Reflect Behavior Problems in Children?Xiyue Wang, Kazuki Takashima, Tomoaki Adachi, Yoshifumi KitamuraCHI 2021 · 被引用 8 次
