Fine-tunes 100+ LLMs and VLMs from one config file or a no-code web UI, unifying LoRA, QLoRA, full tuning, DPO, PPO, KTO and ORPO behind a single interface. Bundles GaLore, Unsloth, FlashAttention-2 and 2-8bit quantization to fit a single 24GB GPU.
Turns local documents into a private, self-hosted ChatGPT-style assistant with no-code agents for web browsing and workflow automation. Runs across LLM providers — OpenAI, Anthropic, Ollama — and routes tools smartly to cut token use.
Provides a RESTful integration layer that connects WhatsApp and other messaging services to external systems; supports both Baileys (Web) and WhatsApp Cloud API, multiple third-party integrations, media storage, and Docker deployment.
Run any open-source LLM, embedding, speech, image, or multimodal model behind one OpenAI-compatible API — swap GPT for an open model in a single line. Routes across vLLM, llama.cpp, GGML, and TensorRT, scaling from a laptop to a multi-node GPU cluster.
Compresses, deploys, and serves LLMs via two engines: TurboMind for raw speed, a PyTorch engine for flexibility. Claims ~1.8x vLLM throughput through persistent batching, blocked KV cache, and split-and-fuse; ships 4-bit AWQ and KV-cache quantization.
AI-assisted database client that generates, explains and optimizes SQL by connecting to your own model. Local-first, cross-platform GUI with metadata browsing, dashboards, 30+ database connectors, and an open-source CLI for automation.
Calls 100+ LLM providers — OpenAI, Anthropic, Gemini, Bedrock, Azure — through one OpenAI-compatible API, as a Python SDK or self-hosted proxy. The proxy adds virtual keys, spend tracking, rate limits, and load balancing across models and providers.
Role-playing LLM agents — CEO, CTO, programmer, tester — collaborate through staged dialogues to turn a one-line prompt into a working software project. Now generalized into a zero-code platform for building custom multi-agent workflows beyond coding.
Converts microphone or streamed audio to text with sub-second latency, pairing WebRTC/Silero voice-activity detection and wake-word activation with swappable local backends — faster-whisper by default, plus whisper.cpp, Moonshine, and sherpa-onnx.
Build and deploy enterprise-grade conversational agents with integrated RAG pipelines, workflow orchestration, multi-modal IO, and model-agnostic integrations (private and public LLMs). Designed for self-hosted production with vector stores and tooling integrations.
Self-hosted browser chat interface for interacting with local or remote LLMs. Supports multiple backends (Ollama, OpenAI-compatible endpoints, llama.cpp), RAG/document chat, plugins/actions, and Docker-based deployment — aimed at teams that need private, customizable LLM UIs.
Turns any website into structured data or an API without code: record clicks once to capture lists and tables, or describe fields in plain language for AI extraction. Also crawls full sites, scrapes pages to Markdown, and runs filtered searches.