Seq-SG2SL: Inferring Semantic Layout From Scene Graph Through Sequence to Sequence Learning
Boren Li, Boyu Zhuang, Mingyang Li, Jian Gu
Abstract
Generating semantic layout from scene graph is a crucial intermediate task connecting text to image. We present a conceptually simple, flexible and general framework using sequence to sequence (seq-to-seq) learning for this task. The framework, called Seq-SG2SL, derives sequence proxies for the two modality and a Transformer-based seq-to-seq model learns to transduce one into the other. A scene graph is decomposed into a sequence of semantic fragments (SF), one for each relationship. A semantic layout is represented as the consequence from a series of brick-action code segments (BACS), dictating the position and scale of each object bounding box in the layout. Viewing the two building blocks, SF and BACS, as corresponding terms in two different vocabularies, a seq-to-seq model is fittingly used to translate. A new metric, semantic layout evaluation understudy (SLEU), is devised to evaluate the task of semantic layout prediction inspired by BLEU. SLEU defines relationships within a layout as unigrams and looks at the spatial distribution for n-grams. Unlike the binary precision of BLEU, SLEU allows for some tolerances spatially through thresholding the Jaccard Index and is consequently more adapted to the task. Experimental results on the challenging Visual Genome dataset show improvement over a non-sequential approach based on graph convolution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2e4bc9df-ab55-4669-b5ff-8c27d4f4e537Cited by top-tier papers1
Ask how each one uses itBuilds on1
Related papers
- LayoutTransformer: Scene Layout Generation With Conceptual and Spatial DiversityCheng-Fu Yang, Wan-Cyuan Fan, Fu-En Yang, Yu-Chiang Frank WangCVPR 2021
- Object-Centric Image Generation from LayoutsTristan Sylvain, Pengchuan Zhang, Yoshua Bengio, R. Devon Hjelm et al.AAAI 2021 · 107 citations
- Scene Graph Expansion for Semantics-Guided Image OutpaintingChiao-An Yang, Cheng-Yo Tan, Wan-Cyuan Fan, Cheng-Fu Yang et al.CVPR 2022 · 18 citations
- Exploiting Relationship for Complex-scene Image GenerationTianyu Hua, Hongdong Zheng, Yalong Bai, Wei Zhang et al.AAAI 2021 · 18 citations
- Hierarchical Image Generation via Transformer-Based Sequential Patch SelectionXiaogang Xu, Ning XuAAAI 2022 · 10 citations
