AIAny
Icon for item

Real5-OmniDocBench

Provides a physical reconstruction benchmark of OmniDocBench v1.5 by producing five real-world photographic variants (Scanning, Warping, Screen‑Photography, Illumination, Skew) for each of 1,355 pages, inheriting original ground-truth to enable controlled, scenario-wise evaluation of document parsing robustness.

Introduction

Why this matters

Real-world document parsing often fails for reasons that synthetic augmentation cannot isolate: geometric warps, screen moiré, harsh lighting or perspective skew have distinct causal effects. Real5-OmniDocBench converts those uncontrolled confounders into controlled variables by physically recreating every OmniDocBench v1.5 page under five photographic scenarios, so you can measure exactly which physical factor breaks a model and by how much.

What Sets It Apart
  • One-to-one physical reconstruction: each of the 1,355 original test pages has five matched physical captures (totaling 6,775 images), and the dataset reuses OmniDocBench’s JSON annotations without modification. That strict correspondence makes cross-scenario comparisons directly comparable rather than approximate.
  • Scenario-level causal analysis: scenarios are orthogonal (Scanning, Warping, Screen‑Photography, Illumination, Skew) and include diverse sub-conditions (e.g., multiple warping types). This lets researchers attribute performance drops to specific physical factors (geometric vs. optical vs. lighting) rather than aggregate “reality gap” noise.
  • Diagnostic benchmark, not just leaderboard: designed to reveal actionable failure modes—e.g., compact, document-specialized VLMs can outperform much larger generalist models under physical stress—so it guides architecture and data-augmentation choices rather than only ranking models.
Who It's For and Trade-offs

Great fit if you need to evaluate or harden OCR and multimodal document parsers against real photographic distortions, compare augmentation strategies, or perform factor-wise robustness studies. It’s particularly useful for teams developing layout/content/structure parsers and for evaluating VLMs on end-to-end document understanding.

Look elsewhere if you need a broader domain coverage of languages, handwriting-heavy corpora, or dynamic video captures—Real5-OmniDocBench focuses on photographic distortions of printed/digital document pages and intentionally preserves the original OmniDocBench annotations rather than adding new types of ground truth.

Practical notes
  • Evaluation is fully compatible with OmniDocBench metrics and scripts (TextEdit, Formula CDM, Table TEDS, Reading Order Edit, Overall score), enabling plug-and-play benchmarking.
  • Because the dataset emphasizes controlled physical variation, it complements synthetic augmentation: use Real5 to validate whether synthetic methods actually close the real-world gap observed here.

Information

  • Websitehuggingface.co
  • OrganizationsPaddlePaddle
  • AuthorsChangda Zhou, Ziyue Gao, Xueqing Wang, Tingquan Gao, Cheng Cui, Jing Tang, Yi Liu
  • Published date2026/01/26

Categories

More Items

Hugging Face

Provides a 57,937-row, quality-filtered multi-teacher SFT distillation corpus combining outputs from Qwen3.8-Max, GLM-5.2 and Kimi K3 across math, code, reasoning, tool-use and dialogue. Includes 24 parquet training views (including a pre-tokenized GLM-4.7 view), configurable sampling weights (sft_balanced), and explicit tool-call trajectories for agent training.

Hugging Face

A 15,000+ English instruction–response corpus for fine-tuning and evaluating LLM instruction-following behavior. Contains human-authored prompts and answers across categories (closed/open QA, summarization, extraction, classification, brainstorming) and is released under CC BY-SA 3.0.

Hugging Face

Provides 52,000 English instruction–response pairs generated by OpenAI's text-davinci-003 for instruction-tuning language models. Released under CC BY-NC 4.0; low-cost synthetic data useful for research but contains model-generated biases and errors.