![rw-book-cover](https://d34adp677peecb.cloudfront.net/static/images/article4.6bc1851654a0.png) ## Metadata - Author: [[arxiv.org]] - Full Title:: EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation - Category:: #🗞️Articles - URL:: https://arxiv.org/pdf/2412.18150 - Read date:: [[2026-08-13]] ## Highlights > Specifically, subject categories are divided into the fol- lowing types: Artifacts, World Knowledge, People, Out- door Scenes, Illustrations, Vehicles, Food &Beverage, Arts, Abstract, Produce & Plants, Indoor Scenes, Animals, and Idioms. Logical relationships are divided into Position Relationship, Number, Color, Writing & Symbols, Per- spective, and Anti-reality. Image style includes Mate- rial, Genre, Design, Photography & Cinema, and Artist & Works. BERT embeddings are used to represent the se- mantic information of the prompts. ([View Highlight](https://read.readwise.io/read/01kzxq0vqx9xa226v5a43tw7p7))