AIAny
Icon for item

Real5-OmniDocBench

Provides a physical reconstruction benchmark of OmniDocBench v1.5 by producing five real-world photographic variants (Scanning, Warping, Screen‑Photography, Illumination, Skew) for each of 1,355 pages, inheriting original ground-truth to enable controlled, scenario-wise evaluation of document parsing robustness.

Introduction

Why this matters

Real-world document parsing often fails for reasons that synthetic augmentation cannot isolate: geometric warps, screen moiré, harsh lighting or perspective skew have distinct causal effects. Real5-OmniDocBench converts those uncontrolled confounders into controlled variables by physically recreating every OmniDocBench v1.5 page under five photographic scenarios, so you can measure exactly which physical factor breaks a model and by how much.

What Sets It Apart
  • One-to-one physical reconstruction: each of the 1,355 original test pages has five matched physical captures (totaling 6,775 images), and the dataset reuses OmniDocBench’s JSON annotations without modification. That strict correspondence makes cross-scenario comparisons directly comparable rather than approximate.
  • Scenario-level causal analysis: scenarios are orthogonal (Scanning, Warping, Screen‑Photography, Illumination, Skew) and include diverse sub-conditions (e.g., multiple warping types). This lets researchers attribute performance drops to specific physical factors (geometric vs. optical vs. lighting) rather than aggregate “reality gap” noise.
  • Diagnostic benchmark, not just leaderboard: designed to reveal actionable failure modes—e.g., compact, document-specialized VLMs can outperform much larger generalist models under physical stress—so it guides architecture and data-augmentation choices rather than only ranking models.
Who It's For and Trade-offs

Great fit if you need to evaluate or harden OCR and multimodal document parsers against real photographic distortions, compare augmentation strategies, or perform factor-wise robustness studies. It’s particularly useful for teams developing layout/content/structure parsers and for evaluating VLMs on end-to-end document understanding.

Look elsewhere if you need a broader domain coverage of languages, handwriting-heavy corpora, or dynamic video captures—Real5-OmniDocBench focuses on photographic distortions of printed/digital document pages and intentionally preserves the original OmniDocBench annotations rather than adding new types of ground truth.

Practical notes
  • Evaluation is fully compatible with OmniDocBench metrics and scripts (TextEdit, Formula CDM, Table TEDS, Reading Order Edit, Overall score), enabling plug-and-play benchmarking.
  • Because the dataset emphasizes controlled physical variation, it complements synthetic augmentation: use Real5 to validate whether synthetic methods actually close the real-world gap observed here.

Information

  • Websitehuggingface.co
  • OrganizationsPaddlePaddle
  • AuthorsChangda Zhou, Ziyue Gao, Xueqing Wang, Tingquan Gao, Cheng Cui, Jing Tang, Yi Liu
  • Published date2026/01/26

Categories

More Items

Hugging Face

Provides ~39 TB of pre‑beamformed (channel capture) ultrasound RF data and metadata in zea/HDF5 format for reconstruction, flow, and inverse‑problem tasks. Released under CC‑BY‑4.0 and curated for training and evaluating ultrasound/RF foundation models.

Hugging Face

Provides a bilingual Chinese–English corpus for LLM training covering pretraining, capability-oriented midtraining (16K–256K long contexts), and supervised fine-tuning. Includes ~4.2T pretrain tokens, ~600B midtrain tokens, and ~4.57M SFT samples; sources span web, PDFs/OCR, code, math, QA, and agentic trajectories under mixed upstream licenses.

Hugging Face

Simulation-ready home dataset for embodied AI: CAD-based household scenes with configured physical properties and metadata, plus 1,000 robot trajectory episodes (RGB-D, HDF5/USDZ) for simulation training and evaluation under CC BY-NC-SA 4.0.