Unifies multimodal understanding, reasoning, and image generation in a single end-to-end architecture using the NEO-unify paradigm. Models pixels and words jointly without a separate visual encoder, and provides interleaved image–text generation, infographic editing, and GGUF/low‑VRAM inference options.
Detects and masks personally identifiable information (PII) in text using a bidirectional token-classification model for high-throughput, on‑premises sanitization. Key traits: 1.5B parameters, 128k-token context, Apache 2.0 license, and tunable precision/recall operating points.
An uncensored, fully unlocked GGUF port of Qwen 3.6‑35B‑A3B for local multimodal (text+image) inference, offering K_P 'Perfect' quant variants (Q8–Q2) and an mmproj for vision. Suited for offline research and experimentation; not for use-cases requiring safety filters.
Generates expressive, prompt-driven text-to-speech audio with optional 10-second voice cloning; prompts control speaker identity, emotion, pauses and nonverbal sounds. An IC‑LoRA fine-tune of LTX‑2.3 that applies an imperceptible Resemble Perth watermark.
GGUF quantized files for a Qwen3.6-35B checkpoint fine-tuned with Claude Opus 4.6-style chain-of-thought distillation to improve reasoning. Offers multiple llama.cpp-compatible quant options (Q4/Q5/Q6/Q8) for local text-generation inference.
Fine-tuned Qwen3.6-35B-A3B MoE that reproduces Claude Opus 4.7-style chain-of-thought with explicit <think>…</think> blocks. Offers sparse activation (256 experts, ~3B active params), 64k context, and GGUF builds for local inference; best for long, multi-step reasoning but may emit very long reasoning traces.
An ~18B frankenmerge text-generation model that stacks two 32-layer Qwen3.5-based finetunes and ships as a 9.2GB Q4_K_M GGUF for efficient local inference. A 1000-step QLoRA heal reduces layer-boundary code corruption and targets coding, reasoning, multilingual chat, and 12–16GB GPU compatibility.
Generates English text matching pre-1931 style — a 13B language model trained on ~260B tokens of pre-1931 English, useful for historical-language generation and stylistic research. An instruction-tuned variant exists for interactive tasks.
Provides a GGUF-packaged, native-INT4 quantized build of the multimodal Kimi K2.6 model for image-text-to-text inference — packaged for local/self-hosted inference engines (vLLM, SGLang, KTransformers) to reduce footprint while keeping multimodal capabilities.
Instruction-tuned 13B LLM post-trained on 260B tokens of pre-1931 English and finetuned with online DPO (LLM-as-judge) to improve instruction-following; suited for period-style English generation and etiquette/letter-writing formats, but not optimized for contemporary factual updates.
Unified multimodal LLM for enterprise workflows: ingests video, audio, image and text to perform transcription, OCR, Q&A, summarization and long-context reasoning. Provides BF16/FP8/NVFP4 weights and integrations with vLLM, TensorRT-LLM and other runtimes.
An HDR LoRA fine-tune for Lightricks' LTX-2.3 (22B) that enables image‑conditioned any‑to‑any image-to-video and text-to-video generation. Designed for HDR-aware synthesis workflows; requires the LTX-2.3 base model and a LoRA-capable runtime.