Provides ~12.29M execution‑free agentic coding trajectories (≈112B tokens) sampled from 122K GitHub PRs to mid‑train code and agent models. Uses bash-only actions (grep, git, sed, etc.) so it scales without Docker; trajectories are unverified and intended for mid-training rather than final SFT.
Turns books, long videos, and podcasts into executable, testable AI agent skills using a structured RIA‑TV++ pipeline. Produces multi-file skill packs (BOOK_OVERVIEW.md, SKILL.md, INDEX.md, DIGEST.md), applies triple verification and pressure tests, and can install skills into Claude Code/Cursor for agent use.
Runs multiple AI agents in parallel inside a single macOS browser, giving each agent an isolated Space that can use your real logins without touching your tabs; controllable via an ego-browser JavaScript skill to perform web automation with fewer tokens and faster task completion.
Provides 207k+ LLM-generated agent trajectories of code edits and tool interactions for training and evaluating software-engineering agents. Collected via OpenHands and SWE-agent using Qwen3.5-122B and MiniMax-M2.5, multilingual across nine languages and released under CC BY 4.0.
A healed 64-layer 'frankenmerge' that stacks two Qwen3.5-derived finetunes into an ~18B GGUF model for multilingual text generation, reasoning, and reliable code/frontend output. Healed with a 1000-step QLoRA to reduce layer-boundary artifacts and targeted to run on 12–16 GB GPUs.
Connects an LLM to a real browser over an editable CDP websocket so the agent can drive clicks, navigation, and generate missing helper code during tasks. The harness self-heals by writing reusable helpers, supports local or cloud browsers, and can optionally record sessions for debugging.
Orchestrates multiple LLM-backed agents locally using tmux and per-role git worktrees, converting role prompts into coordinated development workflows. Key features: configurable two-/four-/six-pack workflows, a durable handoff protocol, per-role backend selection and observable terminals.
Provides 34k execution-style agent trajectories (11,766 issues) for supervised fine-tuning of code-focused LLMs. Each instance includes multi-step interactions, tool-call records, and final unified diffs; generated with Qwen3-Coder and released under permissive licenses for commercial use.
Runs goal-driven penetration tests by orchestration of an LLM agent and an MCP toolchain to perform reconnaissance, vulnerability discovery, exploitation, and structured PoC/report generation; supports multiple LLM providers and local MCP integrations; for authorized security testing only.
Monitors and detects risky behavior in enterprise AI agents via high-fidelity telemetry, security benchmarking, and a two-tier detector. Comprises ADR Sensor, ADR-Bench, and ADR Detector; deployed in production at Uber and validated on public benchmarks.
Provides a GGUF-packaged, native-INT4 quantized build of the multimodal Kimi K2.6 model for image-text-to-text inference — packaged for local/self-hosted inference engines (vLLM, SGLang, KTransformers) to reduce footprint while keeping multimodal capabilities.
Terminal-first developer workspace with an agentic AI side-panel that runs against your API keys or local models. Bundles a native PTY terminal, CodeMirror editor with AI edit diffs, file explorer, git history/graph, and a web preview in a ~7–8MB desktop app with no telemetry.