AIAny
Icon for item

DECOMEG — Brain Activity During Typing (MEG & EEG)

Provides de-identified MEG and EEG recordings of 35 native Spanish speakers typing memorized sentences, with synchronized behavioral logs and standardized event tables. Includes raw .fif and BrainVision files plus MATLAB logs (≈262 GB total); released under CC BY-NC 4.0 for non-commercial research on brain-to-text decoding.

Introduction

Non-invasive neural recordings that are precisely time-locked to naturalistic language production are rare but crucial for training and evaluating brain-to-text decoders. This dataset supplies multi-hour, high-sampling-rate MEG and EEG recordings collected while participants memorized and typed short Spanish sentences, with per-keystroke/word/sentence timing and behavioral logs ready for alignment with decoding pipelines.

What Sets It Apart
  • Simultaneous high-density MEG (306 channels, Elekta Neuromag) and 64-channel EEG (BrainVision), both sampled at 1 kHz, enabling comparisons between modalities for decoding tasks.
  • Task design: read → wait (1.5 s fixation) → type from memory on a custom non-ferromagnetic QWERTY keyboard, producing tightly time-locked motor and language signals without on-screen feedback.
  • Data and formats: raw MEG (.fif) and EEG (BrainVision .eeg/.vhdr/.vmrk), MATLAB behavioral logs (.mat), and a standardized event dataframe compatible with the Brain2Qwerty/neuralset tooling — about 262 GB total (≈21.5 h MEG, ≈17.7 h EEG).
  • Reproducibility & privacy: de-identified recordings only (structural MRI, head videos, eye-tracking excluded) and released under CC BY-NC 4.0 for non-commercial research use.
Who It's For and Trade-offs

Great fit if you are a computational neuroscientist, BCI researcher, or ML practitioner building or benchmarking non-invasive brain-to-text decoders who needs time-resolved MEG/EEG aligned to keystrokes and words. The dataset is prepared to work with existing Brain2Qwerty preprocessing and event-building code, reducing integration overhead.

Look elsewhere if you need a purely EEG-only large cohort with clinical populations, an unrestricted commercial license, or a smaller dataset that can be processed on a laptop: this release is large (≈262 GB), modality- and hardware-specific (Elekta Neuromag, BrainVision), and requires EEG/MEG preprocessing expertise and sufficient compute for model training.

Information

  • Websitehuggingface.co
  • OrganizationsBasque Center on Cognition, Brain and Language (BCBL), HybridMojo LLC, PSL University
  • AuthorsJarod Lévy, Mingfang Zhang, Svetlana Pinet, Jérémy Rapin, Hubert Banville, Stéphane d'Ascoli, Jean-Rémi King
  • Published date2026/06/19

Categories

More Items

Hugging Face

Provides a machine-readable catalog of 117 AI/AX safety and deployment-readiness diagnostic criteria for assessing model intrinsic and serving/infrastructure risks. Includes MODEL-SCAN and AX-SCAN axes, bilingual source fields, per-item evidence guidance, severity/assurance metadata, and a CC BY-NC 4.0 release-candidate.

Hugging Face

Evaluates schema-guided structured extraction from documents: given a document and a JSON schema, systems must return a schema-valid JSON with page-and-box grounding. Covers 370 documents (4,869 pages) across 8 business domains and 67 document types; scores value accuracy, word/page grounding, and long-list completeness.

Hugging Face

Provides 10 million synchronized egocentric experience episodes with structured 3D/4D multimodal annotations — 2.88B RGB frames, 720M depth frames, 576M pose/mocap frames and ~1PB total. Designed for embodied AI, robotics, and multimodal pretraining; research-only, gated access.