AIAny
Icon for item

Scene2Wave

Provides time-aligned simulated urban driving recordings that pair high-rate CSI/CIR with multi-view cameras, LiDAR, radar, IMU and GNSS for perception-to-channel research; contains 100 validated 1-second samples produced with CARLA and Sionna, but is limited in scene diversity and real-world fidelity.

Introduction

The dataset bridges perception and wireless-channel modeling by synchronizing 2 kHz geometry/CSI streams with modality-native sensor records (cameras, LiDAR, radar, IMU, GNSS). That alignment makes it practical to study how visual and geometric cues correlate with short-timescale channel dynamics and to train models that predict or condition channel estimates on perception inputs.

What Sets It Apart
  • Tight time alignment: CARLA geometry and Sionna-derived CIR/CSI use a 2 kHz clock and per-frame mapping so sensor records can be mapped to high-rate channel frames for cross-modal learning.
  • Controlled propagation variety: four propagation profiles (core, multipath, scattering_mild, scattering_medium) let researchers test sensitivity to scattering, ray-budget, and interaction depth while preserving consistent scenario structure.
  • Multi-sensor urban driving stack: multi-view RGB/depth, high-channel-count LiDAR, roadside radar, IMU and GNSS are packaged per-sample as deterministic tar archives for selective extraction and integrity checking.
  • Reproducible simulated provenance: generated with CARLA 0.9.16 and Sionna RT, and distributed with checksums and generation metadata to support deterministic experiments.
Who It's For and Tradeoffs

Great fit if you need a small, fully documented simulated corpus to prototype perception-to-channel feature extraction, CIR/CSI prediction, synchronization/fusion methods, or robustness studies across speed and propagation profiles. Look elsewhere if you require large-scale real-world channel measurements, broader geographic diversity, long-duration traces, or production-ready wireless datasets—the release is intentionally compact (100 validated 1‑s samples across four CARLA towns) and remains a simulated benchmark with layered third-party licensing.

Practical notes

The dataset emphasizes stratified comparisons (Town × motion-state × profile) rather than treating individual samples as identically configured RF reruns. Users should account for simulation limits, avoid data-split leakage across related runs, and validate model conclusions on real measurements before deployment.

Information

  • Websitehuggingface.co
  • OrganizationsPengcheng Laboratory
  • AuthorsMengfan Zheng, Liwen Jing, Tingting Yang, Li Sun, Yuxuan Shi, Ping Zhang, Leiyang Xu, Jianjun Chen
  • Published date2026/08/03

Categories

More Items

Provides a curated benchmark of 170 real-world, multilingual code-refactoring instances to evaluate AI coding agents on large-scale, behavior-preserving, cross-file refactors. Each task includes rewritten issue descriptions and manually reviewed test suites to avoid over- and under-constraining evaluations.

Hugging Face

Provides 1,000 five-second video clips generated by MiniMax H3 for lightweight evaluation of multimodal generation and understanding. Clips are roughly 768p base resolution with diverse aspect ratios and themes, produced with a pruned int8 minimax_h3_fl2va checkpoint at 30 steps.

Hugging Face

Provides 37,484 validated command-line tasks generated by recursive task synthesis, each paired with searchable metadata and a sanitized, runnable package (instructions, solution, verifier, and optional Dockerfile).