Wires retrievers, rerankers, and generators as standalone MCP servers orchestrated in YAML, so iterative RAG logic fits in dozens of lines instead of glue code. Adds loops, conditional branches, one-command web UIs, and shared evaluation benchmarks.
Generates multi-chapter long-form novels with LLMs, automatically linking context and managing foreshadowing for global coherence. Features vector-based retrieval, character/state tracking and a GUI-driven pipeline; requires LLM/embedding API keys.
Performs automated, citation-backed deep research across web, arXiv, PubMed and your private documents using configurable local or cloud LLMs. Runs locally with per-user SQLCipher encryption, Docker/pip installs, LangChain integrations, and an MCP server for assistant integration.
Unifies enterprise knowledge into a permission-aware context layer that delivers citation-backed, explainable search and no-code or SDK-driven agentic workflow automation. Supports 50+ connectors, knowledge-graph retrieval, an MCP server, and bring-your-own-model self-hosting.
Transforms unstructured financial content—papers, news, blogs, and filings—into a queryable semantic knowledge graph for retrieval-augmented research. Combines domain-tuned LLMs, embedding-based search, and modular ingestion pipelines; aimed at quant research teams and institutional workflows.
Measures generative AI inference performance with token-level metrics (TTFT, inter-token latency), latency, and throughput under realistic traffic patterns. Provides a multiprocess engine, real-time TUI dashboard, extensible plugins, and integrations for telemetry and result uploads, aimed at inference benchmarking and capacity planning.
Packages an AI agent's memory — data, embeddings, search indexes, and metadata — into one portable .mv2 file, replacing multi-service RAG stacks. Combines BM25 and HNSW search with temporal queries and sub-millisecond local reads, fully offline.
Wraps Claude Code and Codex with an execution harness that turns one coding agent into coordinated swarms. A single init command adds ~98 agents, an MCP tool server, cross-session vector memory, and cross-machine federation.
Practical, full-stack tutorial for building Retrieval-Augmented Generation (RAG) systems—covers data preprocessing, vector embedding and indexing, hybrid and multimodal retrieval, generation integration, evaluation and production-ready engineering. Includes hands-on projects and examples for developers with Python experience.
Provides semantic code search for AI coding agents by making an entire codebase available as context via hybrid BM25 + vector retrieval, reducing token costs. Uses incremental indexing, AST-based chunking, and Zilliz/Milvus-backed vectors for large-codebase and IDE workflows.
A code-first collection of runnable tutorials for building production-ready generative-AI agents — step-by-step guides covering stateful workflows, vector memory, RAG, tool integrations, Docker/AWS/RunPod deployment, security guardrails, observability, and multi-agent patterns.
Continuously screenshots your screen, feeds the captures to a vision-language model, then pushes back daily summaries, weekly recaps, and todos on its own. Local-first desktop app: data stays on your machine; runs on OpenAI-format or local LLMs.