Provides an open platform of omnimodal world models, datasets, and tools to build Physical AI — joint perception, generation, and action reasoning for robots, autonomous vehicles, and smart infrastructure. Supports images, video, audio, and action-conditioned workflows.
Connects to Gmail, Calendar, and meeting notes to build a local, Obsidian-compatible Markdown graph it acts on — drafting emails, briefs, and decks. Memory accumulates instead of resetting each session; runs on local or hosted models, extensible via MCP.
Press a configurable shortcut, speak, and have your words transcribed and pasted into the active app. Runs Whisper or the CPU-friendly Parakeet V3 fully offline; a Tauri + Rust build with Silero voice-activity detection and optional GPU acceleration.
Zero-shot, single‑reference voice cloning TTS with multilingual support (ZH/EN/JA/ES/AR), fine-grained emotion and duration control, and pronunciation hooks (Pinyin/CMU/Kana); ships model weights, Web UI and production deployment recipes for local or server use.
Provides real-time, local audio recording and transcription on macOS using Whisper and Parakeet engines, with global hotkeys and hold-to-record behavior. Includes model download, microphone selection, drag-and-drop file transcription, multilingual auto-detection and Asian-language autocorrect; Apple Silicon only.
Connects Claude (via the Model Context Protocol) to Ableton Live so the LLM can create and edit tracks, clips, instruments, and control playback through a socket-based MCP server and an Ableton MIDI Remote Script.
Runs open-source LLMs and multimodal models entirely on mobile devices for offline, private inference. Offers Agent Skills, Thinking Mode, Ask Image, audio scribe, model management and benchmarks, with Gemma 4 and Hugging Face integration.
Open-source TTS that clones a voice from a short reference clip across 23+ languages, with adjustable emotional intensity via exaggeration/cfg controls and a built-in Perth neural watermark on every output.
A toolkit and open-weights system for real-time streaming music generation — offers two model sizes (230M / 2.4B), a Python inference library (JAX/MLX), and a C++ engine optimized for Apple Silicon for embedding into DAWs and apps; real-time streaming requires M‑series chips.
Turns OpenAI Whisper into a live streaming transcriber: audio flows in over WebSocket and text returns word-by-word instead of after full utterances. Adds SimulStreaming and LocalAgreement decoding, Silero VAD, and speaker diarization, all self-hosted.
Synthesizes up to 90 minutes of multi-speaker speech in one pass, with as many as four voices in a single conversation. Pairs continuous acoustic and semantic tokenizers at a 7.5 Hz frame rate with a next-token diffusion head on an LLM backbone.
Isolates any single sound from a complex audio mixture using a text description, a visual cue from a video frame, or a time span, returning both the isolated target and the residual. Released in small, base, and large sizes plus visual-prompt variants.