Omnimodal world model that jointly processes and generates text, images, video, audio, and action trajectories for physical AI. Uses a mixture-of-transformers to combine autoregressive reasoning and diffusion-based multimodal generation; released open-source with checkpoints, datasets and benchmarks for robotics and simulation.
Generates synthetic coding-agent session traces by pairing remotely hosted open agent models with local llama.cpp user models across real open-source codebases. Each trace records read/write/edit/bash actions and tool use; the dataset is a reproducible cartesian product (20×3×20×20 = 24,000 sessions) under an MIT license.
Represents episodic memory as a Cue–Tag–Content graph and integrates LLM reasoning into active retrieval so agents iteratively reconstruct and prune evidence paths for long-horizon questions. Reports up to 23% gains on LoCoMo / LongMemEval while reducing token and runtime costs.
Provides a complete, lightly-processed export of AI Village's >1-year multi-agent data: per-agent computer sessions (with screenshots), turn-by-turn computer-use logs, group chats, agent memories, goals, and daily summaries for research into agentic behaviour, multi-agent dynamics, long-horizon memory, and AI safety. Access is manually reviewed.
Measures how coding agents explore repositories by asking them to return a ranked, line-level list of code regions relevant to an issue under a fixed line budget. Covers 848 issues across 203 repos and 10 languages; evaluates coverage, ranking, and context-efficiency to isolate exploration quality.
Removes the subspace of frequent, uninformative tokens that LLMs inject into text embeddings via the model's unembedding matrix. EmbedFilter is a lightweight linear transform that refines LLM-derived embeddings to improve zero‑shot semantic retrieval, enable dimensionality reduction, and speed up indexing; code on GitHub.
Provides a virtual, durable filesystem and pluggable execution runtimes for agents running inside Cloudflare Durable Objects. Offers container, isolate-shell, and isolate-JS backends; preview-stage API with ~10GB Durable Object backing and FUSE-mounted container trade-offs.
Standardizes representation-level evaluation for tabular encoders by exporting row-, column-, and table-level embeddings and probing them with shared lightweight heads across three suites (TRL-CTbench, TRL-Rbench, TRL-DLTE). Supplies curated benchmark assets and task rewrites (50 OpenML tables, 123 targets, a 47,772-table DLTE lake) to enable fair cross-paradigm comparison.
Generates outcome-specific, dialectical rationales with an LLM and derives continuous, calibrated risk scores for irregularly sampled medical time series—mitigating risk polarization. Reports +3.3% average AUPRC and 81% reduction in calibration error across three benchmarks; code released.
Turns raw datasets into verifiable multimodal news features via a multi-agent newsroom pipeline. Key innovations: (1) an Inspector that links each claim to data/code/external references for re-execution and audit; (2) multimodal asset generation (interactive maps, audio, visuals) tailored to the story.
Synthesizes shortcut-resistant search tasks to train deep search agents by controlling four shortcut risks across entity selection, evidence-graph construction, question formulation, and adversarial refinement. Produces training trajectories with longer pre-answer search and fewer shortcut patterns; code will be released on GitHub.
Implements a blockwise sparse attention (MiniMax Sparse Attention) that scores and Top-k selects key-value blocks per Grouped Query Attention group to enable attention over million-token contexts. Paired with an exp-free Top-k GPU kernel and KV-outer sparse execution, it reduces per-token attention compute and yields large prefill/decoding speedups.