AIAny
Icon for item

Alibaba-YuFeng/MMA-SafetyBench

A small image-folder dataset for multimodal/vision model safety benchmarking, containing under 1,000 curated images with annotations to exercise safety-related model behaviours; licensed CC BY 4.0 and hosted on HuggingFace.

Introduction

Why this matters

Multimodal models can exhibit unsafe behaviours on vision inputs, but full-scale benchmarks are slow to run in CI or during rapid prototyping. This dataset offers a compact, curated image-folder collection designed specifically to surface common safety failure modes in vision-capable models, so teams can get fast, repeatable checks before scaling to larger evaluations.

What Sets It Apart
  • Compact size (under 1K images): so what? Enables quick local runs and inclusion in automated test suites without heavy compute or storage costs.
  • Image-folder + simple annotations: so what? Low integration friction with existing evaluation pipelines and model inference scripts.
  • Focused on safety-related scenarios: so what? Prioritizes cases that tend to trigger problematic model outputs rather than broad coverage, making it efficient for regression detection.
  • Clear license (CC BY 4.0): so what? Allows reuse in internal evaluation workflows with attribution requirements clarified.
Who It's For and Tradeoffs

Great fit if you need a fast sanity/regression test to catch vision-safety regressions during model development or CI. Not a substitute for large-scale or domain-specific benchmarks: its small, curated nature means limited coverage and potential sampling biases. Use it as a quick filter before running more comprehensive evaluations or training data collection.

Quick facts
  • Hosted on HuggingFace as an imagefolder dataset
  • Creator handle: Alibaba-YuFeng
  • Created: 2026-05-06
  • Downloads/likes (site metadata): small adoption so far — useful for early-stage checks, not yet widely validated

Information

  • Websitehuggingface.co
  • OrganizationsAlibaba
  • AuthorsAlibaba-YuFeng
  • Published date2026/05/06

Categories

More Items

Hugging Face

Provides a large-scale, multi-speaker Persian speech–text corpus constructed from audiobooks for TTS, ASR, and speaker research. Includes automated alignment and quality scoring, TTS-ready subsets (thousands of hours/1M+ segments) and metadata for speaker IDs and genders — suitable for multi-speaker synthesis and voice cloning research.

Hugging Face

Provides large-scale mathematical problem-solving, rewriting, and dialogue data organized into five Parquet-backed subsets for reasoning-oriented language-model training. Subsets support streaming access, Dataset Viewer inspection, and per-subset provenance metadata; licensed Apache 2.0.

Hugging Face

A multiple-choice benchmark for evaluating LLM understanding in Traditional Chinese across 66 subjects (elementary to professional). Contains ~22K verified questions covering STEM, humanities, social sciences and Taiwan-specific topics, with standardized splits and model leaderboards under an MIT license.