Multilingual streaming ASR that transcribes 40 language-locales using a cache-aware FastConformer‑RNNT architecture. Supports language-ID prompting (or auto-detect), punctuation/capitalization, and configurable chunk sizes to trade latency vs. accuracy for production transcription and streaming voice agents.
A GGUF-format 9B LLM fine-tuned for code generation and agentic tool-calling that uses Multi-Token Prediction (MTP) and draft heads to increase throughput and long-range planning. Intended for local inference and research/experimental coding workflows; Apache‑2.0 license.
Provides a locally runnable 26.9B Qwen3.6 checkpoint that surgically reduces refusal behavior in weight space while preserving capability; ships bfloat16 safetensors and a GGUF quant ladder for local runtimes and red-team evaluation.
Routes LLM API traffic across providers by translating OpenAI, Anthropic, and OpenAI Responses formats, and orchestrates multi-backend routing with typed algorithms and Prometheus metrics. A Rust proxy/library offering launcher, standalone server, and embeddable routing components; experimental (pre-alpha).
Fine-tuned reasoning model that speeds up structured multi-step outputs using Multi-Token Prediction (MTP) from a Qwen3.6-27B base. Produces more concise, faster generations for coding, DevOps, math, and constrained-format tasks; experimental community release for research and evaluation.
Generates temporally coherent MP4 videos from a single input image plus text instructions, with configurable resolution, frame count, and optional AAC audio. Optimized for NVIDIA GPU stacks and integrates with vLLM‑Omni and Hugging Face Diffusers for production inference and research workflows.
Generates audio-driven avatar videos from text, images, or audio inputs with production-grade stability (accurate lip sync, identity consistency) and an 8-step distillation inference mode for faster serving; suitable for broadcasting, virtual hosts, animation, and multi-person scenarios.
A 1.08B-parameter causal LLM engineered for on-device text generation with native long-context (131k tokens) and built-in Think/No-Think modes. It emphasizes tool-calling support, lightweight deployment formats (BF16, GGUF, MLX), and RL+OPD post-training for stronger reasoning and code generation.
A ternary-weight (~1.58-bit) 4B text-to-image diffusion transformer optimized for NVIDIA GPUs using Gemlite INT2 and HQQ; it reduces the transformer to ~1.21 GB (4.55 GB CUDA payload) and targets 1024×1024 generation with a 4-step FlowMatch-Euler sampler.
Processes images and text to produce structured, reasoning-rich text outputs for high-throughput agentic workflows. Sparse MoE design (198B total, ~11B active per token), 256k context window and selectable reasoning levels—optimized for single-pass parsing, verification, and multi-step automation.
Performs hour-scale video understanding and fine-grained temporal localization while exposing agent-style multimodal tool/code/search abilities. Built on a sparse-attention long-context architecture (DSA) and a specialized inference stack—best used in GPU-backed research or production evaluation.
Performs fast, high-quality vision–language grounding: given an image plus a natural-language prompt it returns bounding boxes or points for referred objects. Uses Parallel Box Decoding for parallel coordinate prediction (higher throughput) and targets research/non-commercial use.