AIAny
Icon for item

AX-Ray AI/AX Safety Diagnostics Dataset

Provides a machine-readable catalog of 117 AI/AX safety and deployment-readiness diagnostic criteria for assessing model intrinsic and serving/infrastructure risks. Includes MODEL-SCAN and AX-SCAN axes, bilingual source fields, per-item evidence guidance, severity/assurance metadata, and a CC BY-NC 4.0 release-candidate.

Introduction

Most model evaluations focus on capability scores; AX-Ray reframes the question: what must you check before trusting a model in production? The dataset codifies 117 discrete diagnostic criteria that separate a criterion (what to test) from the evidence that would justify a decision, with special attention to causal integrity (e.g., causal leakage) and serving correctness.

What Sets It Apart

AX-Ray is organized into two assessment axes (MODEL-SCAN and AX-SCAN) and eleven operational categories that span causal safety, reliability, robustness, data integrity, internal diagnostics, remediation, serving, infrastructure security, and compliance. Each record pairs a concise technical diagnostic focus with failure rationale, public detection guidance, remediation suggestions, expected evidence class, severity and automation labels, and per-record source fingerprints. The catalog preserves Korean source fields for provenance while providing an English editorial layer; it is deliberately a criteria-and-evidence catalog, not a leaderboard or certification artifact.

Who It's For and Tradeoffs

Great fit if you need an auditable checklist to map model findings to concrete evidence requirements (security teams, model-audit programs, compliance engineers, and deployment gating pipelines). Look elsewhere if you want raw probes, proprietary test prompts, or turnkey certification—AX-Ray intentionally omits exact thresholds, internal scoring recipes, raw outputs, and empirical leaderboard claims. The release-candidate status and retained review gates mean users must perform row-level primary-source verification and legal review before relying on any governance mapping.

Information

  • Websitehuggingface.co
  • OrganizationsFINAL-Bench, VIDRAFT
  • Published date2026/08/14

Categories

More Items

Hugging Face

Evaluates schema-guided structured extraction from documents: given a document and a JSON schema, systems must return a schema-valid JSON with page-and-box grounding. Covers 370 documents (4,869 pages) across 8 business domains and 67 document types; scores value accuracy, word/page grounding, and long-list completeness.

Hugging Face

Provides 10 million synchronized egocentric experience episodes with structured 3D/4D multimodal annotations — 2.88B RGB frames, 720M depth frames, 576M pose/mocap frames and ~1PB total. Designed for embodied AI, robotics, and multimodal pretraining; research-only, gated access.

Hugging Face

Provides a reproducible, deduplicated corpus of text extracted from PDFs for LLM pretraining—about 3 trillion tokens from ~475 million documents in 1733 language-script pairs. Includes OCR and text extraction pipelines, per-page language IDs, MinHash deduplication, and is released under ODC‑By 1.0.