AIAny
Icon for item

VoxMem: Benchmarking Multimodal Memory in Large Audio Language Models

Evaluates multimodal memory in large audio language models by testing four acoustic evidence types and four memory operations across multi-session spoken histories, stratified across 8K–64K token context budgets.

Introduction

Why this matters

Spoken conversational systems must remember more than words: who spoke, how they spoke, and what ambient sounds were present. Current benchmarks emphasize lexical content and single-session tasks, leaving a gap in measuring whether audio-native models can retain and reason over speaker identity, paralinguistic cues, and environmental sound across long, multi-session histories. VoxMem fills that gap with a taxonomy-driven benchmark that jointly characterizes the acoustic evidence to be retained and the memory operations applied to it.

Key Findings
  • Dataset scale and structure: 3,196 evaluation instances drawn from 34,743 spoken sessions (177 hours), covering four acoustic evidence types (speech semantics, speaker identity, paralinguistic cues, environmental sound) and four memory operations (information extraction, multi-session reasoning, temporal tracking, answer refusal). Samples are stratified by required context budget from 8K to 64K tokens.
  • Measurable gaps by evidence and operation: models retain lexical content far better than speaker identity, paralinguistic cues, or environmental audio; this gap widens for more complex operations (e.g., multi-session reasoning) and longer histories.
  • Strong multi-model baseline result: across 15 evaluated large audio language models, no model exceeded 40% accuracy at 32K tokens, revealing substantial headroom for progress in acoustic memory.
  • Distinct failure modes: errors are not random — models often misattribute speaker-specific facts, ignore subtle prosodic cues, or miss brief ambient events, indicating representation and retrieval weaknesses that depend on evidence type and temporal distance.
Who it's for and tradeoffs

Great fit if you are building or evaluating audio-native LLMs, designing long-horizon voice assistants, or researching multimodal memory mechanisms—VoxMem provides targeted, artifact-minimized tasks and provenance annotations to probe how far back required evidence lies. Look elsewhere if your interest is limited to transcript-only benchmarks, short-session interaction, or purely textual memory: VoxMem intentionally focuses on signals irrecoverable from ASR and emphasizes multi-session accumulation over single-session shortcuts.

Practical implication

VoxMem is a diagnostic and benchmarking tool rather than an engineering recipe: use it to reveal whether a model’s embeddings, retrieval strategy, or attention allocation fail for specific acoustic evidence types or memory operations. Its stratified design makes it easy to track improvements as models scale context length, alter representation learning, or incorporate explicit memory modules.

Information

  • Websitearxiv.org
  • AuthorsYang Xiao, Vidhyasaharan Sethu, Eun-Jung Holden, Ting Dang
  • Published date2026/09/26

More Items

Teaches small reasoning models to diagnose when additional internal thinking is insufficient and to selectively query stronger models; introduces FlyBy, a supervised + cost-aware RL framework that learns when/what to ask. Improves pass@ metrics on hard benchmarks while reducing serving cost.

Shows that private post-training changes leave measurable "behavioral shadows" in task-unrelated single-word outputs and proposes Active Taskless Distillation (ATD) to transfer capabilities to a student using only those single-word teacher responses, without teacher logits or parameters.

Improves test-time scaling of looped transformers by adaptively assigning extra recurrent iterations to tokens that benefit most. TaH2 is a post-training method that jointly trains an iteration decider with the backbone using lookahead depth supervision, boosting the accuracy–compute slope and peak accuracy on challenging benchmarks.