Why this matters
Multimodal models can exhibit unsafe behaviours on vision inputs, but full-scale benchmarks are slow to run in CI or during rapid prototyping. This dataset offers a compact, curated image-folder collection designed specifically to surface common safety failure modes in vision-capable models, so teams can get fast, repeatable checks before scaling to larger evaluations.
What Sets It Apart
- Compact size (under 1K images): so what? Enables quick local runs and inclusion in automated test suites without heavy compute or storage costs.
- Image-folder + simple annotations: so what? Low integration friction with existing evaluation pipelines and model inference scripts.
- Focused on safety-related scenarios: so what? Prioritizes cases that tend to trigger problematic model outputs rather than broad coverage, making it efficient for regression detection.
- Clear license (CC BY 4.0): so what? Allows reuse in internal evaluation workflows with attribution requirements clarified.
Who It's For and Tradeoffs
Great fit if you need a fast sanity/regression test to catch vision-safety regressions during model development or CI. Not a substitute for large-scale or domain-specific benchmarks: its small, curated nature means limited coverage and potential sampling biases. Use it as a quick filter before running more comprehensive evaluations or training data collection.
Quick facts
- Hosted on HuggingFace as an imagefolder dataset
- Creator handle: Alibaba-YuFeng
- Created: 2026-05-06
- Downloads/likes (site metadata): small adoption so far — useful for early-stage checks, not yet widely validated