The dataset provides a reproducible, multi-model benchmark focused on artistic evaluation rather than traditional photorealism metrics. By publishing 8,000 generated images (1,000 per model) paired with the prompts and three per-image scores, it makes head-to-head comparison of style, affect, and prompt grounding straightforward for researchers and practitioners.
What Sets It Apart
- Shared prompt suite and full cross-model coverage: 1,000 evaluation prompts (500 style-focused, 500 open-ended) that enable direct per-prompt comparisons across eight models, so you can measure relative strengths on identical inputs.
- Human-refined prompts + automated judge: Uses GPT-5.6 Sol to provide consistent per-image scores for aesthetic quality, emotional evocation, and content integrity, reducing variance from ad-hoc human annotations while providing a reproducible baseline.
- Practical release format: Images embedded in 19 Parquet shards (1024×1024 resolution) with explicit columns (image, prompt, model, aesthetic, emotional_evocation, content_integrity), so it plugs into common data pipelines without extra preprocessing.
Who it's for and tradeoffs
Great fit if you need a controlled benchmark to compare artistic and prompt-grounding behavior across modern text-to-image systems, to train or evaluate style-conditioned adapters, or to analyze automated aesthetic metrics. Look elsewhere if you need large-scale natural-image diversity, real-world licensed photographs, or fine-grained human-annotator demographic metadata—the release prioritizes aesthetic evaluation of generated art and reproducibility over crowd-sourced diversity.