ACL2026

Putting Captions to the Test: Evaluating Video Caption Quality through Multiple-Choice Question Answering

Zizhen Wang, Bo Feng, Zhengfeng Lai, Shiyu Li, Yang Lu, Meng Cao, Ping Huang, Xiaoming Simon Wang

摘要

Evaluating video captioning remains a critical challenge for Visual Large Language Models (VLLMs). Existing metrics primarily rely on matching generated text against groundtruth references. This paradigm suffers from the "one-to-many" nature of video description, where high-quality captions are often penalized for lexical mismatches or valid shifts in visual focus. Furthermore, such assessments are typically one-dimensional, failing to provide a fine-grained analysis of caption quality. To address this, we redefine caption quality via information fidelity: A caption must maximize the coverage of salient visual information while ensuring strict factuality. We introduce CapQuiz, a novel reference-free benchmark that assesses captions based on their utility in answering human-verified, fine-grained, multiple-choice questions derived from the video. CapQuiz features a hierarchical taxonomy of 10 question types (spanning Descriptive and Inferential categories) across 24 diverse video domains. We further formulate CapF1, a composite metric that synthesizes CapP (measuring factuality) and CapR (measuring coverage). Extensive experiments demonstrate that CapQuiz correlates significantly better with human judgments than existing metrics and offers interpretable insights into model performance. ### 1. Executive Summary & Scene Classification * Core Narrative Synopsis: (A concise, one-to-two-sentence summary encapsulating the primary action, subjects, and outcome.) * Inferred Genre/Context: (e.g., Cooking tutorial, product review, documentary segment, home video, animated short, security footage.) ### 2. Perception Analysis: The Observable Reality * 2.1. Entity & Attribute Identification: * Characters/People: * [Person A]: * Visual Properties: Describe appearance (gender, age est., hair, ethnicity), clothing (type, color, style), and accessories. * Quantity: (e.g., "One person initially, a second person enters at time 00:00:45.") * Key Objects: * [Object A]: * Visual Properties: Describe its type, color, material, shape, and any distinct features. * Quantity: Note the count of similar objects (e.g., "Three blue cups on the table."). * State & State Changes: Describe its condition or status and note any changes (e.g., "Initially closed, opened at time 00:00:30," "Power light is on," "Appears damaged."). * Animals: * [Animal A]: Describe species, breed, color, size. * 2.2. Environment & Setting Analysis: * Location Type: (e.g., Indoor kitchen, outdoor public park, office cubicle, car interior.) * Ambient Details: Note key background elements, furniture, weather conditions, and general state (e.g., "tidy," "cluttered"). * Inferred Time of Day: (e.g., "Bright daylight due to harsh shadows," "Dusk inferred from warm, low light," "Night, lit by artificial sources.") * 2.3. On-Screen Text & Graphics: * [Text/Graphic 1]: Transcribe the text/logo and note its location and the time ranges in which it is visible, e.g. 00:00:15-00:00:30. ### 3. Reasoning Analysis: Connecting the Dots * 3.1. Chronological & Causal Reconstruction: * Time Segment [e.g., 00:00:00-00:00:10]: * Atomic Actions: Describe actions by single entities (e.g., "Person A picks up the red ball."). * Interactions: Describe actions between entities (e.g., "Person A hands the ball to Person B."). * Causal Links: If an action directly causes a result, state it (e.g., "Because the ball was thrown, the window broke."). * Time Segment [e.g., 00:00:10-00:00:20]: (Repeat the structure above.) * 3.2. Relational Analysis: * Spatial Relations: Throughout the sequence, describe the key relative positions (e.g., "The cat is sleeping under the table," "At 00:00:50, Person A moves to stand behind Person B."). * Temporal Relations: Use clear sequential language (e.g., "The phone rings before she opens the book," "While he was cooking, the dog entered the room."). * Comparative Observations: Note any explicit or implicit comparisons (e.g., "The second car is moving faster than the first," "Box A is visibly larger than Box B."). * 3.3. Interpretive & Social Analysis: * Inferred Emotions & Intent: * [Person A]: (e.g., "Appears focused and determined, likely intending to complete the puzzle," "Facial expression shifts from neutral to surprised at 00:01:00."). * Plot & Thematic Reasoning: Describe the overall story, moral, or abstract message being conveyed. * Social & Normative Context: Describe the social dynamics (e.g., "formal interview," "casual conversation," "teacher-student interaction"). Note any actions that align with or deviate from common social norms. ### 4. Prediction & Extrapolation Analysis * 4.1. Immediate Next Action (Short-term Prediction): Based on the final timestamp, what is the single most likely action to occur in the next 1-3 seconds? * 4.2. Plausible Outcome (Long-term Prediction):