
## Metadata
- Author: [[arxiv.org]]
- Full Title:: EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation
- Category:: #🗞️Articles
- URL:: https://arxiv.org/pdf/2412.18150
- Read date:: [[2026-08-13]]
## Highlights
> Specifically, subject categories are divided into the fol- lowing types: Artifacts, World Knowledge, People, Out- door Scenes, Illustrations, Vehicles, Food &Beverage, Arts, Abstract, Produce & Plants, Indoor Scenes, Animals, and Idioms. Logical relationships are divided into Position Relationship, Number, Color, Writing & Symbols, Per- spective, and Anti-reality. Image style includes Mate- rial, Genre, Design, Photography & Cinema, and Artist & Works. BERT embeddings are used to represent the se- mantic information of the prompts. ([View Highlight](https://read.readwise.io/read/01kzxq0vqx9xa226v5a43tw7p7))