AIAny
Icon for item

KSAFE-MM

Benchmark for evaluating multimodal LLM safety in Korean cultural contexts — includes KSAFE-MM-G which localizes global safety queries into Korean scenarios and KSAFE-MM-C which targets culture-specific visual-textual vulnerabilities. Provides curated image–text pairs and jailbreak-style prompts to reveal both unsafe behaviors and over-refusal.

Introduction

Most multimodal safety benchmarks are English-centric or focus on generic hazards; KSAFE-MM flips that assumption by centering Korean cultural and institutional contexts so evaluations reflect real-world local risks. The dataset stresses that model safety failures often arise not from raw toxicity but from missing local knowledge, culturally grounded cues, and visual-contextual interplay that enable bypasses or harmful outputs.

Key Findings
  • Two-part design: KSAFE-MM-G transforms globally shared safety queries into Korean-grounded multimodal samples; KSAFE-MM-C uses in-the-wild images and localized visual cues combined with jailbreak-style textual intents to probe culture-dependent vulnerabilities.
  • Reveals asymmetric failure modes: some models show high attack success rates on culturally tailored prompts while others exhibit excessive refusal on benign inputs — indicating a tradeoff between vulnerability and over-sensitivity.
  • Dataset construction emphasizes semantic alignment and privacy filtering: image–query pairs were selected from diverse web sources, de-duplicated, and filtered to avoid references to identifiable individuals or companies.
Who it's for and tradeoffs

Great fit if you evaluate multimodal LLM safety for non-English markets, build culturally robust moderation or alignment layers, or research localized attack vectors. Look elsewhere if you only need generic, English-only toxicity benchmarks or lightweight synthetic tests — KSAFE-MM is designed for contextual, in-the-wild evaluation and thus requires handling image hosting, language-specific annotation, and culturally informed judgment during interpretation.

Information

  • Websitehuggingface.co
  • OrganizationsK-intelligence
  • Published date2026/06/11

Categories

More Items

Hugging Face

Installation-oriented dataset that packages ComfyUI-ready files and instructions for running MiniMax H3 locally — includes pruned/INT8/BF16 checkpoints, matching Qwen3-VL text encoders, video/audio VAEs, and official ComfyUI workflow templates for joint audio+video generation.

Hugging Face

Provides a large-scale, multi-speaker Persian speech–text corpus constructed from audiobooks for TTS, ASR, and speaker research. Includes automated alignment and quality scoring, TTS-ready subsets (thousands of hours/1M+ segments) and metadata for speaker IDs and genders — suitable for multi-speaker synthesis and voice cloning research.

Hugging Face

Provides large-scale mathematical problem-solving, rewriting, and dialogue data organized into five Parquet-backed subsets for reasoning-oriented language-model training. Subsets support streaming access, Dataset Viewer inspection, and per-subset provenance metadata; licensed Apache 2.0.