Offline-first knowledge server that bundles local AI chat (Ollama + vector RAG), offline Wikipedia/education/maps, and utility tools behind a Dockerized management UI — designed to keep searchable knowledge available without cloud access.
Provides a deterministic context and knowledge-graph layer under LLMs and vector stores to record auditable decisions, provenance, and explainable rule-based reasoning. Supports polyglot graph storage, W3C PROV-O export, SHACL governance, and self-hosted enterprise connectors.
An open-source memory layer that turns agent runs and conversations into structured, persistent state recallable across sessions. Captures facts, events, preferences, and relationships automatically; LLM-agnostic with SDK and MCP integration.
Compiles an agent's raw chat logs, documents, and tool traces into three persistent layers — index, learned skills, and user memory — so context survives sessions. Claims 92% Locomo-benchmark accuracy and up to 95% lower token cost than replaying history.
Indexes any repo into a knowledge graph of dependencies, call chains, and execution flows, then feeds it to AI coding agents via MCP so they stop missing context. Ships as a CLI plus a zero-install browser graph explorer with chat.
Provides hierarchical, versioned semantic memory for AI agents with Git-like branching, commits, and rollbacks—using semantic paths and cryptographic provenance instead of opaque vector stores. Designed for branch-aware, auditable memory in multi-agent and production workflows.
A large multi-config collection of query–document pairs assembled to reproduce and extend the mGTE/LateOn data recipe for pre-training text embedding models. Data come in source-specific configs and include per-row drop/duplicate flags and guidance for using cleaned subsets for training.
Provides mined hard negatives and relevance scores for 1.88M queries across seven retrieval datasets, enabling contrastive fine-tuning and nv-retrieve filtering; includes full 2048 mined negatives per query, paired query/document splits, and parquet-formatted files for large-scale training.
Agent memory that learns over time instead of just recalling past chats: retain/recall/reflect primitives turn interactions into facts, experiences, and mental models. Reports top LongMemEval scores; self-hostable with Python and Node SDKs.
Drives penetration testing from chat commands, orchestrating 100+ security tools through an MCP-native multi-agent engine on CloudWeGo Eino. Adds attack-chain graphs, risk scoring, and human-in-the-loop approval gates for authorized use.
Aggregates SEC EDGAR filings into raw files, parsed plaintext, and rich filing metadata for LLM training and retrieval. Includes ~8.05M filings (~590 GB, ~43B tokens), per-filing token counts, and parsed outputs; Apache-2.0.
Combines a vector store, Cypher-style graph queries, and on-device LLM inference in one Rust engine, with a graph neural network that reranks results and adapts to query patterns in under a millisecond. Services ship as self-contained .rvf containers.