Provides the gated, official OSWorld 2.0 Python task class files (task_*.py) required to run the benchmark; distributed via a Hugging Face gated dataset to reduce benchmark leakage. Download requires accepting gated access on Hugging Face.
A 20B retrieval subagent trained with reinforcement learning inside a stateful search harness that externalizes recoverable search state (candidate pool, curated evidence, verification records). The harness lets the policy focus on semantic search decisions, improving curated recall and transfer robustness.
Workflow-aware benchmark for autonomous medical-AI research that splits agent execution into five stages (Plan, Setup, Validate, Inference, Submit) and evaluates long-horizon runs across segmentation, image enhancement, VQA, report generation, and lesion detection with stage-level scoring.
A benchmark for evaluating web-browsing agents in Korean contexts, composed of 400 tasks (300 manually verified by native speakers). Includes a human-verified split and an adversarial synthetic split to probe failure modes; reveals large performance gaps for both frontier and Korean models.
Centralizes indexing and management of local AI coding-agent sessions so you can search, view full context, migrate, resume, and restore conversations across agents and devices. Supports extensible local sources, AI summaries, optional Supabase sync, and Skills management.
Evaluates multimodal LLMs on streaming egocentric video for spatial intelligence using 1,680 human-annotated questions across 348 videos; organizes tasks into four hierarchical levels (perception → tracking → simulation → allocentric mapping) and highlights allocentric mapping as the main bottleneck.
Enables agents to proactively discover multiple hidden problems in a user context and pair each with supporting evidence and concrete actions. Uses iterative discovery (batch rounds conditioned on prior finds) and reusable "thought templates" to expand coverage and ground claims.
Generates synthetic coding-agent session traces by pairing remotely hosted open agent models with local llama.cpp user models across real open-source codebases. Each trace records read/write/edit/bash actions and tool use; the dataset is a reproducible cartesian product (20×3×20×20 = 24,000 sessions) under an MIT license.
Benchmark for evaluating proactive LLM mediators in realistic, multi-domain conflict scenarios by constructing cases from real disputes, probing five socio-cognitive adaptation axes, and using a topic-localized evaluator that achieves 0.82 alignment with human experts.
Open-weight frontier LLM for agentic reasoning and long-context analysis (up to 1M tokens). Uses a LatentMoE + Mamba-2 hybrid with Multi-Token Prediction and NVFP4 efficiency (550B total / 55B active). Suited for multilingual agents, RAG, and heavy tool-use workloads.
Multilingual frontier LLM optimized for long-context reasoning and agentic workflows, combining a LatentMoE (Mamba-2 + MoE) hybrid architecture with Multi-Token Prediction and NVFP4 quantization; targeted for NVIDIA GPU deployments and governed by the OpenMDW-1.1 license.
Benchmark that measures an agent's ability to discriminate fine-grained relational structure in long-term memories. It embeds relation-controlled memory variants into realistic user–agent histories and tests downstream recovery and reasoning, highlighting where current memory systems fail.