Provides a reflexive agentic framework for long-horizon video understanding that replaces costly iterative reasoning with dual contextual states: a consolidated global multimodal script and parametric latent states for fast retrieval and response, improving speed and memory efficiency.
27B multimodal LLM post-trained to prioritize agentic, weight-scaled reasoning over 64K-token contexts. Built on Qwen3.6-27B and released with BF16 weights plus several GGUF quants; aimed at coding, long-document reasoning, tool use and multimodal inspection.
Creates an open-ended interactive world simulator with an unbounded interaction horizon via causal pretraining, a distilled real-time runtime that drives 720p@60fps, a wider action/event repertoire, and a pilot–director agent split for behavior planning and environment synthesis.
Provides a terminal-style benchmark of 46 long-horizon tasks decomposed into fine-grained graded subtasks to produce dense intermediate rewards and partial credit, enabling evaluation of long-horizon planning, long-context management, and iterative debugging. Tasks typically require hundreds of episodes and minutes-to-hours of execution; baseline evaluations report high token and episode consumption with low pass rates, highlighting evaluation headroom.
Simulates a hospital LIMS to benchmark agentic clinical reasoning: agents inspect demographics, medications, lab orders/results and then submit ICD‑10 diagnostic reports scored by deterministic, context‑aware graders. Ships as an OpenEnv/FastAPI runtime with 8 scenarios, step‑level rewards and trajectory capture for RL, tool‑use and evaluation.
Evaluates agents inside a structured hospital workflow via a downloadable FastAPI runtime that enforces role-specific tool permissions, evidence-before-treatment discipline, deterministic grading, dense process rewards, and full trajectory logging. Designed for RL, offline policy learning, multi-agent workflow research and process-supervision datasets; not for real patient care.
Evaluates proactive, multimodal agents on 400 bilingual real‑world tasks across five capability axes (Skill Usage, Exploration, Long‑Context Reasoning, Multimodal Understanding, Cross‑Platform Coordination) using live Docker‑based, stepwise closed‑loop evaluation to separate base model skills from framework design.
Provides a deliberative Agent OS layer for robots that handles scene-conditioned planning, context-isolated skill execution, multi-stage verification, persistent multi-modal graph memory, and edge–cloud collaboration. Introduces EmbodiedWorldBench (16 scenes, 200+ tasks) and a failure-driven self-evolution loop; shows improved task success and strong memory benchmark scores.
A GGUF-format Qwen3.6 35B base model image-text-to-text release repaired via tensor-level SVD/scale correction and packaged with Hermes agent tweaks; multimodal (vision + text), MoE architecture, ready for GGUF runtimes like llama.cpp.
Acquires repository knowledge via a targeted QA loop before generating patches, decoupling knowledge acquisition from repair. A Questioner and Answerer produce evidence-grounded QA pairs that a Resolver uses to generate fixes; improves Pass@1 on SWE-bench Verified with modest overhead.
A GGUF-local variant of Qwen3.6-35B that applies a non-training 'Genesis' tensor-repair process and Hermes-agent fine-tuning to enable uncensored, multimodal (text+image) local inference. Highlights: MoE 35B spec, large native context, Hermes function-calling dataset transfer, and recommended quantization/runtime settings.
Provides live Codex-CLI agent run traces from GPT-5.6 Sol capturing coding, debugging, security reviews, and harness/seed workflows in cumulative next-action prefixes — suitable for supervised fine-tuning and analysis of tool-using coding agents.