Generates 48kHz multilingual speech from text using a tokenizer-free diffusion-autoregressive TTS architecture, supporting natural-language voice design, controllable cloning, and low-latency streaming. Notable for a 2B-parameter backbone and built-in AudioVAE super-resolution (16k→48k).
Maps a codebase plus docs, PDFs, media and configs into a local, queryable knowledge graph; parses code with a local tree-sitter AST (no LLM), uses configurable backends for semantic extraction of non-code, and outputs graph.json, graph.html and a brief report.
Automates scanning and evaluating job listings with LLM-driven agents, then generates ATS-optimized, per-role PDFs and a unified tracker. Supports batch processing and terminal-first workflows with structured A–F scoring and portal scanners.
Lets AI coding agents compile your documents and chat histories into a maintained Obsidian vault: it ingests sources, distills them into interconnected markdown pages, tracks deltas and provenance, and exposes query/lint/export skills across many agents.
Generative-AI-enabled timeline video editor for macOS that lets creators generate and edit videos and images directly inside the timeline. Includes a local MCP server for agent integrations (Claude/Codex/Cursor); editor is open-source while generative processing is closed-source and subscription-based; macOS 26 on Apple Silicon only.
Turns your documents into a persistent, interlinked personal wiki by incrementally reading sources, generating wiki pages, and keeping knowledge up to date. Features two-step chain-of-thought ingest, graph-based relevance with Louvain clustering, optional embedding search (LanceDB), and a local HTTP API for agent integration.
Local-first voice workflows for cloning, multi-engine TTS/ASR, video dubbing, dictation, transcription and audiobook production across 646 languages. Desktop app with a local OpenAI-compatible API, engine catalogue (TTS/ASR/LLM), and explicit opt-ins for remote features to keep audio and projects on-device.
Desktop app for local voice cloning, real-time dictation, and end-to-end video dubbing using zero-shot TTS across 600+ languages; features multi-engine TTS/ASR, speaker diarization, vocal isolation, batch pipelines, and invisible audio watermarking — all run fully offline.
Defines a machine-readable text format that pairs YAML design tokens with human-readable rationale so coding agents can generate, lint, diff, and export UI systems. Bundles a CLI for validating DESIGN.md files and exporting tokens to Tailwind and W3C-compatible formats.
A local-first web and desktop dashboard for Hermes Agent that runs streamed agent chats, manages profiles/providers/models/credentials, schedules cron jobs, and inspects files and terminals across local, Docker, SSH and Singularity backends.
Automates video editing driven by LLM agents: reads word-level transcripts to propose and execute cuts, remove filler words, auto grade color, burn subtitles, and generate animation overlays. Self-evaluates every cut before showing a preview; aimed at talking-heads, tutorials and interviews.
Turns plain-English system or process descriptions into polished, themeable architecture, workflow, sequence, data-flow and lifecycle diagrams as a self-contained HTML file, with one-click theme toggle, copy-to-clipboard and export to PNG/JPEG/WebP/SVG (native up-to-4× rasterization).