Modular marketplace of focused Claude Code plugins that composes specialized agents and progressive 'agent skills' to orchestrate multi-agent development workflows while minimizing token usage.
An open-source memory layer that turns agent runs and conversations into structured, persistent state recallable across sessions. Captures facts, events, preferences, and relationships automatically; LLM-agnostic with SDK and MCP integration.
Searchable plugin marketplace and curated collections of Claude Code agents, commands, hooks, skills and plugins — lets users discover, browse, and install subagents and tooling via a web UI, plugin commands or a CLI. Indexes community plugins and MCP servers and provides curated packs.
Coordinates specialized AI agents — developer, browser, document, multimodal — running in parallel on your desktop to automate multi-step work. Runs fully local via Ollama, vLLM, or LM Studio, with built-in MCP tools and human-in-the-loop checkpoints.
Parses PDF resumes into structured JSON using LLMs, enriches profiles with GitHub signals, and outputs explainable category scores, evidence, bonuses and deductions. Runs fully local with Ollama or via Google Gemini; designed for reproducible, fairness-constrained resume scoring in hiring workflows.
Compiles an agent's raw chat logs, documents, and tool traces into three persistent layers — index, learned skills, and user memory — so context survives sessions. Claims 92% Locomo-benchmark accuracy and up to 95% lower token cost than replaying history.
Collaborates on web tasks in real time: edit its plan before it runs, pause and grab the browser mid-task, and approve irreversible clicks before they happen. A research prototype for studying human-in-the-loop oversight instead of full autonomy.
Provides ultra-fast, typo-tolerant file search and grep tuned for Neovim and AI agents, with built-in memory (frecency, git status, size, definition matches). It reduces agent token use and speeds developer file discovery in large repos.
Evaluates and optimizes AI agents and language models in containerized environments, supporting large-scale parallel benchmarks and RL rollouts. Integrates with third‑party providers for thousands of parallel environments and serves as the official harness for Terminal‑Bench.
Deploys autonomous AI agents that dynamically attack running apps and return validated proof-of-concept exploits instead of static-analysis noise. Specialized agents cover IDOR, injection, SSRF, XSS, and auth flaws, with HTTP proxy and CI/CD hooks.
Deep research agent for complex, long-horizon research and prediction tasks. Pairs a 256K context window with up to 300 tool calls per query for web search, extraction, and code execution. Ships as open 30B and 235B models scoring 82.7% on GAIA.