Most model evaluations focus on capability scores; AX-Ray reframes the question: what must you check before trusting a model in production? The dataset codifies 117 discrete diagnostic criteria that separate a criterion (what to test) from the evidence that would justify a decision, with special attention to causal integrity (e.g., causal leakage) and serving correctness.
What Sets It Apart
AX-Ray is organized into two assessment axes (MODEL-SCAN and AX-SCAN) and eleven operational categories that span causal safety, reliability, robustness, data integrity, internal diagnostics, remediation, serving, infrastructure security, and compliance. Each record pairs a concise technical diagnostic focus with failure rationale, public detection guidance, remediation suggestions, expected evidence class, severity and automation labels, and per-record source fingerprints. The catalog preserves Korean source fields for provenance while providing an English editorial layer; it is deliberately a criteria-and-evidence catalog, not a leaderboard or certification artifact.
Who It's For and Tradeoffs
Great fit if you need an auditable checklist to map model findings to concrete evidence requirements (security teams, model-audit programs, compliance engineers, and deployment gating pipelines). Look elsewhere if you want raw probes, proprietary test prompts, or turnkey certification—AX-Ray intentionally omits exact thresholds, internal scoring recipes, raw outputs, and empirical leaderboard claims. The release-candidate status and retained review gates mean users must perform row-level primary-source verification and legal review before relying on any governance mapping.