AIAny
Icon for item

AX-Ray AI/AX Safety Diagnostics Dataset

Provides a machine-readable catalog of 117 AI/AX safety and deployment-readiness diagnostic criteria for assessing model intrinsic and serving/infrastructure risks. Includes MODEL-SCAN and AX-SCAN axes, bilingual source fields, per-item evidence guidance, severity/assurance metadata, and a CC BY-NC 4.0 release-candidate.

Introduction

Most model evaluations focus on capability scores; AX-Ray reframes the question: what must you check before trusting a model in production? The dataset codifies 117 discrete diagnostic criteria that separate a criterion (what to test) from the evidence that would justify a decision, with special attention to causal integrity (e.g., causal leakage) and serving correctness.

What Sets It Apart

AX-Ray is organized into two assessment axes (MODEL-SCAN and AX-SCAN) and eleven operational categories that span causal safety, reliability, robustness, data integrity, internal diagnostics, remediation, serving, infrastructure security, and compliance. Each record pairs a concise technical diagnostic focus with failure rationale, public detection guidance, remediation suggestions, expected evidence class, severity and automation labels, and per-record source fingerprints. The catalog preserves Korean source fields for provenance while providing an English editorial layer; it is deliberately a criteria-and-evidence catalog, not a leaderboard or certification artifact.

Who It's For and Tradeoffs

Great fit if you need an auditable checklist to map model findings to concrete evidence requirements (security teams, model-audit programs, compliance engineers, and deployment gating pipelines). Look elsewhere if you want raw probes, proprietary test prompts, or turnkey certification—AX-Ray intentionally omits exact thresholds, internal scoring recipes, raw outputs, and empirical leaderboard claims. The release-candidate status and retained review gates mean users must perform row-level primary-source verification and legal review before relying on any governance mapping.

Information

  • Websitehuggingface.co
  • OrganizationsFINAL-Bench, VIDRAFT
  • Published date2026/08/14

Categories

More Items

Hugging Face

Provides over 1.1M hours of high-bandwidth, multichannel multilingual speech with segment- and word-level timestamps, English translations, and per-file metadata for ASR, TTS and audio-representation research. Preserves original 48kHz multichannel OPUS audio and is released under CC BY 3.0.

Hugging Face

Synthetic, clinician-verified ChatML dataset of 2,194 doctor–patient encounters covering 2,194 unique human diseases; each JSONL record includes 20 structured fields, verified PubMed references, realistic vitals/labs, and is intended for RAG and model fine-tuning (not medical advice).

Hugging Face

Provides 2,000 synthetic multiple-choice items designed for continuation log-likelihood scoring to evaluate small language models' Theory of Mind (social-cognitive) abilities; 40 constructs, balanced answer positions, and easy/medium difficulty.