Fetches multi-source content (webpages, YouTube, PDFs, WeChat, paywalled articles, podcasts), uploads it to Google NotebookLM, and generates outputs such as podcasts, PPTs, mind maps, or quizzes. Differentiators: automatic paywall-bypass pipeline, Claude Code Skill integration, and CLI + MCP components for WeChat and document scraping.
Self-hosted coding assistant that runs frozen local LLMs with constraint-driven planning, energy-based verification, and self-verified repair to produce verified code. Emphasizes offline inference (no cloud), Docker/bare-metal deployment, and requires a 16GB+ GPU.
Self-hosted personal AI agent runtime that runs chats, tools, automations and long-term memory for persistent workflows. Small, readable core with a bundled WebUI, multi-chat integrations, an OpenAI-compatible API and a Python SDK for easy extension and deployment.
Multimodal OCR and document-understanding toolkit for recognizing complex layouts, tables, formulas and code. Uses Multi-Token Prediction and stable RL for better training; ships as a 0.9B-parameter model with a Python SDK and deployment guides for vLLM, SGLang and Ollama.
Provides 100 real-world, open-ended research tasks paired with expert-written rubrics (around 40 weighted criteria per task) to evaluate long-form, web-browsing research agents on factual accuracy, analysis depth, presentation, and citation quality.
Provides L3 refined synthetic training data by converting high-quality web corpora into Q&A pairs and multi-style rewrites; supplies 400B+ English and 200B+ Chinese tokens for late-stage LLM pretraining and decay-phase training.
Generates high‑fidelity, expressive speech and environmental sounds from text. The MOSS‑TTS Family provides specialized models for long‑form TTS, multi‑speaker dialogue, voice design and realtime streaming, plus torch‑free inference paths (llama.cpp / ONNX) and Hugging Face releases.
Provides multi-task long-speech evaluation data for eight speech-understanding tasks (ASR, summarization, QA, translation, emotion, speaker counting, content separation, language detection). Includes 101,822 long audio files and ~204,881 annotated examples with JSONL task splits for easy loading.
Runs a local-first, full AI stack—LLM inference, chat UI, voice, agents, workflows, RAG, and image generation—deployable with one command. Auto-detects hardware and bootstraps a small model for instant chat while larger models download; supports Linux, Windows, macOS and optional cloud/hybrid modes.
A challenge repository for training the best language model that fits inside a 16,000,000‑byte (16MB) submission artifact; provides baseline training code, FineWeb bpb evaluation, a public leaderboard, and compute-grant instructions for short 8×H100 runs.
Turns a PC, Mac, or Linux machine into a private AI server with one-command installers: local LLM inference, a ChatGPT-style web UI, voice, agents, RAG, workflows, image generation, hardware-aware model selection, and optional cloud/hybrid modes.
Cleaned reasoning dataset of problem→thinking→solution triplets derived from Opus 4.6, provided in Parquet with ~2,160 cleaned rows (original 3,305). Filters remove empty/short/refusal/non‑substantive responses; hosted on Hugging Face under Apache‑2.0.