Evaluation dataset for comparing eight text-to-image models using 8,000 generated images with source prompts and per-image scores for aesthetic quality, emotional resonance, and content integrity. Includes model labels, shared prompts, GPT-5.6 Sol automated scores, embedded images in Parquet, and an Apache-2.0 license.