AIAny
Icon for item

MiniMax H3 - 1K

Provides 1,000 five-second video clips generated by MiniMax H3 for lightweight evaluation of multimodal generation and understanding. Clips are roughly 768p base resolution with diverse aspect ratios and themes, produced with a pruned int8 minimax_h3_fl2va checkpoint at 30 steps.

Introduction

Synthetic outputs are a practical lens to inspect what multimodal generative models actually learn and where they fail. This 1K collection offers a compact, consistent set of short clips produced by a MiniMax H3 FL2VA pruned int8 checkpoint (30 steps), intended for quick qualitative analysis, probing model biases, and small-scale evaluation workflows.

What Sets It Apart
  • Compact, fixed-length corpus: 1,000 clips of approximately five seconds each, making quick iterations and visual spot checks easy without large storage or compute needs.
  • Consistent generation recipe: all samples were produced with the same pruned int8 minimax_h3_fl2va_convrot checkpoint at 30 steps, which highlights systematic artifacts arising from that specific model/configuration.
  • Diverse presentation: samples cover multiple aspect ratios and a range of visual themes and media types, useful for testing robustness across framing and content variation.
Who It's For and Trade-offs

Great fit if you need a small, reproducible set of MiniMax H3 outputs for qualitative inspection, model-output auditing, or preliminary evaluation of video+multimodal pipelines. Look elsewhere if you need large-scale training data, rigorously curated ground-truth labels, or benchmark-grade diversity — this collection is synthetic, limited in scale, and reflects artifacts and biases from the specific pruned int8 generation setup (30 steps). Licensing is not specified on the dataset card, so verify reuse terms before redistribution.

Information

Categories

More Items

Hugging Face

Benchmark for typed probabilistic decisions: given one shared state, a model answers multiple typed questions at once and returns full probability distributions (noul/choice/score). Contains four workflows with train/test splits and soft gold labels from teacher samples, designed to evaluate accuracy, calibration, and latency trade-offs.

Hugging Face

Provides a complete benchmark and training release for native audio‑visual dialogue: 2,800 synthesized audio‑visual dialogues, tiered rubrics, scoring code, ~28GB of 1080p media, and an RL reward recipe to evaluate and train omni models that take video+audio and return text.

Hugging Face

Maps a state and question to typed probabilistic decisions (choice distributions, yes/no probabilities, or scored/ordinal outputs) across controlled synthetic tasks. Offers multiple frozen configs with train/calibration/validation/test/OOD splits, Parquet + raw JSONL exports, and reproducible manifests and provenance.