Chains four swappable open modules — voice activity detection, speech-to-text, an LLM, and text-to-speech — into a local voice agent that needs no proprietary APIs. Runs on CUDA, Apple Silicon, or Docker, with an OpenAI-compatible realtime WebSocket mode.
Developer framework for building AI agents that autonomously trade on Polymarket prediction markets. Bundles the Polymarket and Gamma APIs, a Chroma RAG layer that pulls in news, and a CLI to query markets, reason with an LLM, and execute trades.
Runs a native, extensible AI agent on desktop, CLI, or API to automate code, workflows, research, and writing. Built in Rust, supports 15+ LLM providers and 70+ extensions via the Model Context Protocol — designed for local-first automation and developer workflows.
Converts PDFs, Office files, HTML, images and audio into one structured DoclingDocument, with deep PDF layout, reading order, table-structure and formula recognition, OCR, and native LangChain/LlamaIndex/Haystack integrations for RAG pipelines.
Give an agent a goal and it plans, then executes each step using AI models and your everyday apps. Build agents via chat-driven AutoPilot, a drag-and-drop builder, or self-hosted code, then run them on a schedule across integrations.
Runs autonomous AI-agent workforces where each agent, skill, and company process lives as version-controlled code you own. Agents act in isolated sandboxes and submit deliverables for human review, with 3,000+ connectors plus MCP support.
Enables agents to autonomously operate GUIs and complete complex computer tasks — includes the Agent S papers and the gui-agents SDK, grounding-model support, and runnable S3 agent implementations for Windows/macOS/Linux.
Python web scraping framework that automatically relocates elements when a site's HTML changes, so selectors survive redesigns. Bundles Cloudflare Turnstile bypass, TLS fingerprint impersonation, and a Scrapy-like async spider for full crawls.
VideoCaptioner is an AI-powered video subtitling assistant that combines ASR (local or cloud) with LLM-based subtitle segmentation, correction and translation. It supports offline GPU transcription, concurrent chunk transcription, VAD, speaker-aware processing, batch subtitling and one-click subtitle-to-video synthesis, with both GUI and CLI options.
Turns any website into a structured, text-like interface that LLM agents can read and act on, handling clicks, forms, scraping, anti-detection and CAPTCHAs. Ships as an open-source Python library plus a hosted cloud API for running browser agents at scale.
Keeps the former Windsurf IDE lineage alive as Devin Desktop, a local editor for planning, delegating, reviewing, and shipping code with cloud and local agents from one surface.