Lune

EMNLP2025Top-tier venue

Grounded Semantic Role Labelling from Synthetic Multimodal Data for Situated Robot Commands

Claudiu Daniel Hromei, Antonio Scaiella, Danilo Croce, Roberto Basili

2025Year
1Citations

Abstract

Understanding natural language commands in situated Human-Robot Interaction (HRI) requires linking linguistic input to perceptual context. Traditional symbolic parsers lack the flexibility to operate in complex, dynamic environments. We introduce a novel Multimodal Grounded Semantic Role Labelling (G-SRL) framework that combines frame semantics with perceptual grounding, enabling robots to interpret commands via multimodal logical forms. Our approach leverages modern Vision Language Models (VLMs), which jointly process text and images, and is supported by an automated pipeline that generates high-quality training data. Structured command annotations are converted into photorealistic scenes via LLM-guided prompt engineering and diffusion models, then rigorously validated through object detection and visual question answering. The pipeline produces over 11,000 imagecommand pairs (3,500+ manually validated), while approaching the quality of manually curated datasets at significantly lower cost.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 86225349-06a9-4e57-b0f0-d63d66243097

Builds on7

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines