Generates English text matching pre-1931 style — a 13B language model trained on ~260B tokens of pre-1931 English, useful for historical-language generation and stylistic research. An instruction-tuned variant exists for interactive tasks.
Provides 2,405 chain-of-thought reasoning traces generated by Claude Opus 4.7 for hard math, science, and formal problems. Each record pairs a problem with the model's full <think> working and a polished answer; available as parquet splits for non-commercial research under Anthropic's usage policy.
Provides a GGUF-packaged, native-INT4 quantized build of the multimodal Kimi K2.6 model for image-text-to-text inference — packaged for local/self-hosted inference engines (vLLM, SGLang, KTransformers) to reduce footprint while keeping multimodal capabilities.
Synthetic Korean-language persona dataset for training and evaluating conversational and generative models — 1M records (≈7M persona entries) with 26 fields aligned to South Korea’s demographic distributions. Built with NeMo Data Designer and released under CC BY 4.0.
Instruction-tuned 13B LLM post-trained on 260B tokens of pre-1931 English and finetuned with online DPO (LLM-as-judge) to improve instruction-following; suited for period-style English generation and etiquette/letter-writing formats, but not optimized for contemporary factual updates.
Unified multimodal LLM for enterprise workflows: ingests video, audio, image and text to perform transcription, OCR, Q&A, summarization and long-context reasoning. Provides BF16/FP8/NVFP4 weights and integrations with vLLM, TensorRT-LLM and other runtimes.
An HDR LoRA fine-tune for Lightricks' LTX-2.3 (22B) that enables image‑conditioned any‑to‑any image-to-video and text-to-video generation. Designed for HDR-aware synthesis workflows; requires the LTX-2.3 base model and a LoRA-capable runtime.
Produces 384‑dim multilingual (and code) embeddings with up to 32,768 token context, optimized for low‑latency production retrieval. Compact 97M model with ONNX/OpenVINO and vLLM/GGUF deployment options for edge and high‑throughput use.
Provides high-performance CUDA/CUTLASS kernels implementing Kimi Delta Attention (KDA), accelerating KDA prefill on SM90+ (Hopper) GPUs. Integrates as a drop-in backend for flash-linear-attention, supports native variable-length batching, and targets K=V=128; requires CUDA 12.9+/PyTorch 2.4+.
A 27B multimodal causal language model with a vision encoder and native long-context support (262,144 tokens). Optimized for repository-level coding agents and multimodal understanding; includes preserved "thinking" traces, multi-token prediction (MTP), and deployment recipes for vLLM / SGLang / Transformers.
FP8-quantized 27B multimodal Qwen3.6 model weights in Hugging Face Transformers format — supports image/text/video inputs, native 262k token context (extensible to ~1M), and is compatible with vLLM/SGLang/KTransformers for efficient local serving and research.
Benchmark dataset for evaluating clinician-facing chat assistants: physician-authored conversations plus rubric items, use-case and difficulty labels, specialty metadata, and a built-in canary to reduce benchmark contamination. Hosted on Hugging Face under an MIT license.