AIAny
Icon for item

Moonworks Lunara Art Eval

Evaluation dataset for comparing eight text-to-image models using 8,000 generated images with source prompts and per-image scores for aesthetic quality, emotional resonance, and content integrity. Includes model labels, shared prompts, GPT-5.6 Sol automated scores, embedded images in Parquet, and an Apache-2.0 license.

Introduction

The dataset provides a reproducible, multi-model benchmark focused on artistic evaluation rather than traditional photorealism metrics. By publishing 8,000 generated images (1,000 per model) paired with the prompts and three per-image scores, it makes head-to-head comparison of style, affect, and prompt grounding straightforward for researchers and practitioners.

What Sets It Apart
  • Shared prompt suite and full cross-model coverage: 1,000 evaluation prompts (500 style-focused, 500 open-ended) that enable direct per-prompt comparisons across eight models, so you can measure relative strengths on identical inputs.
  • Human-refined prompts + automated judge: Uses GPT-5.6 Sol to provide consistent per-image scores for aesthetic quality, emotional evocation, and content integrity, reducing variance from ad-hoc human annotations while providing a reproducible baseline.
  • Practical release format: Images embedded in 19 Parquet shards (1024×1024 resolution) with explicit columns (image, prompt, model, aesthetic, emotional_evocation, content_integrity), so it plugs into common data pipelines without extra preprocessing.
Who it's for and tradeoffs

Great fit if you need a controlled benchmark to compare artistic and prompt-grounding behavior across modern text-to-image systems, to train or evaluate style-conditioned adapters, or to analyze automated aesthetic metrics. Look elsewhere if you need large-scale natural-image diversity, real-world licensed photographs, or fine-grained human-annotator demographic metadata—the release prioritizes aesthetic evaluation of generated art and reproducibility over crowd-sourced diversity.

Information

Categories

More Items

Hugging Face

Provides a 1 trillion-token multimodal interleaved dataset (HTML subset updated as data_v1_1 with 742B HTML tokens) and 3.4B images drawn from HTML/PDF/ArXiv sources for multimodal pretraining; released under CC-BY-4.0 with safety and deduplication guidance.

Hugging Face

Provides 369 Harbor sandbox tasks ported from OpenAI's openai/math: each task is a Lean theorem with missing `sorry` proofs that an agent must complete, graded by a strict Comparator exact-match verifier. Includes task definitions, generator, and manifest for RL/code-agent evaluation.

Hugging Face

Curated English Wikipedia text prepared for language-model training and evaluation, provided in WikiText-2 and WikiText-103 variants. Preserves original case, punctuation and numbers; offers raw and tokenized splits for long-range language modeling under a CC BY‑SA license.