AIAny
Icon for item

OpenH-RF

Provides ~39 TB of pre‑beamformed (channel capture) ultrasound RF data and metadata in zea/HDF5 format for reconstruction, flow, and inverse‑problem tasks. Released under CC‑BY‑4.0 and curated for training and evaluating ultrasound/RF foundation models.

Introduction

The dataset aggregates large‑scale pre‑beamformed (channel capture) ultrasound RF recordings and structured metadata to enable raw‑to‑insight model development: reconstruction, flow estimation, quantitative imaging and ultrasound inverse problems. Its scale and raw channel format target foundation models and reconstruction pipelines that need unprocessed sensor measurements rather than beamformed images.

What Sets It Apart
  • Raw channel capture in the zea/HDF5 schema: preserves pre‑beamformed RF waveforms and acquisition metadata so methods can learn physics‑aware reconstructions and device‑specific corrections.
  • Multi‑task scope and scale: ~39 TB across 32 contributed sub‑datasets and ~11,491 HDF5 files, covering b‑mode, flow, localization microscopy, transcranial and other application domains — suitable for training large reconstruction or RF foundation models.
  • Open, permissive license and community stewardship: released under CC‑BY‑4.0 with a community steering group and multi‑institution contributions to encourage reuse and reproducible benchmarking.
Who It's For and Trade-offs

Great fit if you are training or evaluating ultrasound reconstruction models, RF foundation models, or research on sensor‑level inverse problems and flow estimation. Expect substantial storage, I/O, and preprocessing needs (tens of TB); working with this dataset typically requires domain knowledge of ultrasound acquisition, handling device variability, and attention to clinical/privacy constraints. Not ideal if you only need beamformed B‑mode images or small, lightweight example datasets.

Information

  • Websitehuggingface.co
  • OrganizationsNVIDIA Corporation, Stanford University, Eindhoven University of Technology, Tel Aviv University, Siemens Healthineers, Resolve Stroke, University of British Columbia, KAIST, Barreleye Inc., Seoul National University Bundang Hospital, us4us Ltd., University of Colorado Boulder, Vanderbilt University, University of North Carolina at Chapel Hill, Weizmann Institute of Science, University of Oslo, Technical University of Munich, Politecnico di Torino, University of Twente, Weill Cornell Medicine, Worcester Polytechnic Institute, University of Strasbourg, Concordia University, Mosaic Intelligence, PATH Lab - Technion - Israel Institute of Technology, University of Waterloo, Dartmouth College, Provost Ultrasound Lab, Robeauté, OpenH-RF community & Steering Group
  • Published date2026/08/06

Categories

More Items

Hugging Face

Provides a bilingual Chinese–English corpus for LLM training covering pretraining, capability-oriented midtraining (16K–256K long contexts), and supervised fine-tuning. Includes ~4.2T pretrain tokens, ~600B midtrain tokens, and ~4.57M SFT samples; sources span web, PDFs/OCR, code, math, QA, and agentic trajectories under mixed upstream licenses.

Hugging Face

Simulation-ready home dataset for embodied AI: CAD-based household scenes with configured physical properties and metadata, plus 1,000 robot trajectory episodes (RGB-D, HDF5/USDZ) for simulation training and evaluation under CC BY-NC-SA 4.0.

Hugging Face

Synthesizes 234K self-contained, high-difficulty scientific reasoning QA pairs by distilling research papers into compact 'reasoning skeletons'. Emphasizes mechanistic reasoning, hypothesis falsification, quantitative derivation and boundary calibration; built for SFT and reasoning evaluation.