
## Metadata
- Author: [[arxiv.org]]
- Full Title:: ITERCOMP: ITERATIVE COMPOSITION-AWARE
FEEDBACK LEARNING FROM MODEL GALLERY FOR
TEXT-TO-IMAGE GENERATION
- Category:: #🗞️Articles
- URL:: https://arxiv.org/pdf/2410.07171
- Read date:: [[2026-08-13]]
## Highlights
Probably doesn’t apply to Nano Banana and similar models
> these models often struggle to follow complex prompts to achieve precise compositional generation (Omost-Team, 2024; Yang et al., 2024b; Zhang et al., 2024b), which requires the model to possess robust, comprehensive capabilities in various aspects, such as attribute binding, spatial relationships, and non-spatial relationships ([View Highlight](https://read.readwise.io/read/01kzx6t491hpr8ea5bcyf66has))
>  ([View Highlight](https://read.readwise.io/read/01kzx6y75kvz3n7nt69j8s4yyd))