Loading...

Visual Compositional Generation: Origins, Datasets, and Evaluation Metrics

Marioriyad, Arash | 2025

152 Viewed
  1. Type of Document: M.Sc. Thesis
  2. Language: Farsi
  3. Document No: 58553 (19)
  4. University: Sharif University of Technology
  5. Department: Computer Engineering
  6. Advisor(s): Rohban, Mohammad Hossein; Soleymani Baghshah, Mahdieh
  7. Abstract:

  8. The rapid emergence of powerful text-to-image generative models in recent years has attracted significant attention. Trained conditionally on large-scale text–image pairs, these models have achieved remarkable results in terms of quality, realism, and diversity of generated images. Nevertheless, they exhibit substantial weaknesses in visual compositional generation. In this context, compositional generation refers to the model’s ability to correctly interpret and render entities, attributes (such as colors and sizes), inter-object relationships, and object counts specified in the input text, while maintaining overall image quality. In this thesis, we first analyze and categorize the primary failure modes of compositional generation, including entity missing, attribute mis-binding, spatial relationship errors, and numeracy issues. We then survey and classify the major approaches proposed in the literature to address these problems, with particular emphasis on training-free inference-time methods. Furthermore, we review widely used benchmarks for compositional evaluation and discuss assessment metrics ranging from human judgment to automatic methods, including visual question answering, embedding-based similarity, and caption–prompt consistency. Building on this foundation, we focus on two major challenges, entity missing and spatial relationship errors, and propose efficient inference-time solutions that do not require any modification of model parameters. Extensive experiments on diverse datasets, including T2I-CompBench, HRS-Bench, and VISOR, demonstrate that our methods yield significant improvements across multiple compositional alignment metrics such as CLIPScore and VQA-based scores. Moreover, we introduce a novel evaluation metric specifically designed for assessing spatial relationships, which exhibits a stronger correlation with human judgments compared to prior metrics.
  9. Keywords:
  10. Diffusion Model ; Entity Missing ; Spatial Relationship Error ; Attribute Binding ; Visual Compositional Generation ; Numeracy Problem ; Text–Image Alignment

 Digital Object List

 Bookmark

...see more