Why this matters
Spoken conversational systems must remember more than words: who spoke, how they spoke, and what ambient sounds were present. Current benchmarks emphasize lexical content and single-session tasks, leaving a gap in measuring whether audio-native models can retain and reason over speaker identity, paralinguistic cues, and environmental sound across long, multi-session histories. VoxMem fills that gap with a taxonomy-driven benchmark that jointly characterizes the acoustic evidence to be retained and the memory operations applied to it.
Key Findings
- Dataset scale and structure: 3,196 evaluation instances drawn from 34,743 spoken sessions (177 hours), covering four acoustic evidence types (speech semantics, speaker identity, paralinguistic cues, environmental sound) and four memory operations (information extraction, multi-session reasoning, temporal tracking, answer refusal). Samples are stratified by required context budget from 8K to 64K tokens.
- Measurable gaps by evidence and operation: models retain lexical content far better than speaker identity, paralinguistic cues, or environmental audio; this gap widens for more complex operations (e.g., multi-session reasoning) and longer histories.
- Strong multi-model baseline result: across 15 evaluated large audio language models, no model exceeded 40% accuracy at 32K tokens, revealing substantial headroom for progress in acoustic memory.
- Distinct failure modes: errors are not random — models often misattribute speaker-specific facts, ignore subtle prosodic cues, or miss brief ambient events, indicating representation and retrieval weaknesses that depend on evidence type and temporal distance.
Who it's for and tradeoffs
Great fit if you are building or evaluating audio-native LLMs, designing long-horizon voice assistants, or researching multimodal memory mechanisms—VoxMem provides targeted, artifact-minimized tasks and provenance annotations to probe how far back required evidence lies. Look elsewhere if your interest is limited to transcript-only benchmarks, short-session interaction, or purely textual memory: VoxMem intentionally focuses on signals irrecoverable from ASR and emphasizes multi-session accumulation over single-session shortcuts.
Practical implication
VoxMem is a diagnostic and benchmarking tool rather than an engineering recipe: use it to reveal whether a model’s embeddings, retrieval strategy, or attention allocation fail for specific acoustic evidence types or memory operations. Its stratified design makes it easy to track improvements as models scale context length, alter representation learning, or incorporate explicit memory modules.