
## Metadata
- Author: [[Hossein Shahabadi; Niki Sepasian; Arash Marioriyad; Ali Sharifi-Zarchi; Mahdieh Soleymani Baghshah]]
- Full Title:: Infinity and Beyond: Compositional Alignment in VAR and Diffusion T2I Models
- Category:: #🗞️Articles
- URL:: https://arxiv.org/pdf/2512.11542
- Read date:: [[2026-08-13]]
## Highlights
> despite their impressive perceptual quality, contemporary T2I models continue to exhibit substantial limitations in compositional alignment, the ability to faithfully bind objects, attributes, and spatial or non-spatial relations described in natural language into coherent visual outputs. A growing body of work demonstrates that strong visual fidelity does not equate to reliable compositional correctness, particularly in multi-object or attribute-heavy prompts, where models frequently violate attribute bindings, confuse spatial configurations, or hallucinate unintended elements ([View Highlight](https://read.readwise.io/read/01kzx7kfshpmmt426e2rg17eep))