A 23-skill Claude Code toolkit that composes an LLM-driven virtual engineering team (CEO, designer, eng manager, QA, security, release) into slash-command workflows — includes real-browser QA, a persistent GBrain memory, multi-agent integrations, and team auto-update semantics.
Scans AI agent skills for security issues—detecting vulnerabilities, malicious patterns, and supply-chain risks before installation. Combines static AST checks (64 patterns across 16 categories) with optional LLM semantic review, OSV live CVE lookups, and JSON/Markdown/SARIF outputs for CI or manual review.
Provides a pytest-native framework to write safety and security tests for agentic AI applications. Defines adversarial attacks, benign-failure suites, and harm-category assertions with evaluation-driven checks and CI-friendly reporting, so red-teaming becomes testable and automatable.
Multimodal image-text-to-text fork of Gemma 4 (31B) using a 'CRACK v2' abliteration — tuned for conversational vision inputs and thinking-mode support in JANG v2 safetensors format. Recommended to run in vMLX; published by dealignai.
Provides hardware-isolated, sub-60ms, ultra-low-overhead sandboxes to run untrusted LLM/agent code. Offers event-level snapshots, kernel-level egress control, credential vaulting, and drop-in E2B SDK compatibility for high-density AI agent deployment.
Removes safety refusals from a Gemma 4 E4B–based model and publishes uncensored, locally runnable GGUF/safetensors variants while preserving all tensors and fixing prior corruption. Intended for red‑teaming and offline research; not recommended for production.
Runs goal-driven penetration tests by orchestration of an LLM agent and an MCP toolchain to perform reconnaissance, vulnerability discovery, exploitation, and structured PoC/report generation; supports multiple LLM providers and local MCP integrations; for authorized security testing only.
Monitors and detects risky behavior in enterprise AI agents via high-fidelity telemetry, security benchmarking, and a two-tier detector. Comprises ADR Sensor, ADR-Bench, and ADR Detector; deployed in production at Uber and validated on public benchmarks.
Performs agent-driven security scans of codebases using LLM coding agents to find and triage vulnerabilities. Combines fast regex discovery, per-file AI investigation and revalidation, with optional sandboxed parallel execution and Vercel AI Gateway integration for large monorepos.
Ingests and normalizes security telemetry, runs multi-model AI agents to produce replayable investigations and automated triage/response; key features include a step-by-step Investigation Ledger, CI-gated eval harness, and self-hostable deployments.
Generates production-ready offensive-security artifacts from prompts—Nuclei templates, CVE PoCs, exploit scripts and pentest tooling—fine-tuned on bug-bounty reports and CVE writeups and quantized for consumer/server GPU deployment.
Routes code-based AI agents through repeatable reverse-engineering and pentesting workflows and orchestrates local and remote tools (jadx, Frida, IDA, BurpSuite) so agents can triage APKs, binaries, JS, firmware, and CTFs without guessing the toolchain. Includes master routing rules, tool-index detection, MCP integration, and a field-journal for reusable lessons.