Provides 1.7M agent interaction traces in terminus-2 format for training and evaluating agentic LLMs and RL agents. Compiled from 219 source datasets across code repair, shell, math, competitive programming and general tasks; produced with the Harbor harness.
Aggregates 750k+ Harbor-compatible agentic tasks from 100+ public sources (Parquet shards preserved). Includes tasks with and without verifiers for RL evaluation or SFT/datagen workflows, enabling reproducible trace generation.
Terminal-native AI coding assistant optimized for the deepseek-v4 model. Provides configurable "thinking" modes and reasoning-intensity controls, agent skills for extensibility, MCP integration, and a shared config with a VSCode plugin.
An instruct-focused LLM (104B total, 7.4B active) optimized for fast, token-efficient inference in agent workflows. Uses hybrid linear attention plus a sparse MoE to raise throughput and cut token use; suited for high-frequency production agents, with some trade-offs in very deep reasoning.
A trillion-parameter LLM optimized for long-context, low-latency text generation and agentic coding workflows. Combines MLA+Linear Attention and a post-training 'fast thinking' token-suppression strategy to reduce token overhead and improve multi-step execution reliability for production agents.
Distills DeepSeek‑V4's multi-step structured reasoning into a Qwen3.5‑9B model for fast image-text-to-text reasoning and agentic tool workflows. Trades larger teacher size for inference efficiency and improved procedural reasoning — good for low-latency research, evaluation, and agent integration.
Lets developers build stateful, tool-enabled Python AI agents that run on Google's Antigravity runtime. Includes built-in tools (file I/O, shell, image generation), a declarative policy/hook system, multimodal input, and MCP integration.
Early-preview (≈1.2k rows) dataset of agentic coding prompts and unedited model responses generated by DeepSeek‑V4‑Pro, covering real-world programming tasks across many languages. Intended for research, filtering, and model evaluation rather than production training without review.
Performs agent-driven security scans of codebases using LLM coding agents to find and triage vulnerabilities. Combines fast regex discovery, per-file AI investigation and revalidation, with optional sandboxed parallel execution and Vercel AI Gateway integration for large monorepos.
Explains AI coding jargon in plain English, giving concrete engineering meanings for terms like context window, tool call, and attention degradation. Structured as a browsable dictionary with pragmatic examples and guidance for developers building LLM-driven systems.
Converts technical books and document collections into an on-demand agent “skill” that Claude Code, GitHub Copilot CLI, and Amp can load to answer questions from the original content. Produces a compact SKILL.md plus per-chapter files so agents load only the needed sections, cutting token use and reducing hallucination risk.
Evaluates LLM-driven agents on long-horizon, policy-rich U.S. healthcare workflows using 75 clinical task fixtures and a 20-app MCP simulator; includes task fixtures, shared worlds, and leaderboard integration (Managed-Care handbook is gated).