Provides a 30K+ problem multimodal, multilingual dataset of Olympiad-level math problems with expert solutions and a math-aware retrieval benchmark—includes images, hierarchical topics, provenance from official booklets, and LLM-assisted metadata (v0, CC BY 4.0).
A GGUF-format preview checkpoint derived from Qwen3.6-27B — a multimodal, image-text-to-text reasoning model fine-tuned for more structured reasoning and consistent answer style; packaged for local inference and compatible with engines like vLLM/SGLang/llama.cpp.
A 33B Mixture-of-Experts text-to-text model optimized for local, long-context agentic coding—3B activated params per token, 131k token window, mixed sliding-window and global attention, FP8 KV cache, Apache-2.0 license.
Provides a lightweight assistant (draft) model for Gemma 4 E4B used in speculative-decoding pipelines — it predicts token drafts that the target model verifies in parallel, enabling up to ~2× decoding speedups while preserving identical final outputs. Useful for low-latency, multimodal assistant and on-device scenarios.
A lightweight 'drafter' assistant for Gemma 4 31B that generates speculative token drafts to enable up-to-2× decoding speedups while preserving final output quality; compatible with Hugging Face Transformers and any-to-any pipelines.
Acts as the assistant (drafter) checkpoint for Gemma 4 26B A4B on Hugging Face, used in Speculative Decoding to pre-draft tokens and speed up generation. Designed for long-context, multimodal workflows where lower latency and on-device or edge inference matter.
Provides an end-to-end platform to evaluate, observe, protect, and optimize LLM and AI agent deployments. Integrates OpenTelemetry tracing, 50+ evaluation metrics, agent simulations, an OpenAI‑compatible gateway, and guardrails; self‑hostable under Apache 2.0.
Unifies video, audio, image and text understanding for enterprise Q&A, summarization, transcription and document intelligence. The NVFP4 quantized variant reduces footprint to ~20.9GB for more efficient single‑GPU deployment and is tuned for NVIDIA runtimes (vLLM, TensorRT).
Supervised fine-tuning dataset of 7,716 reasoning-focused Q&A examples distilled from the DeepSeek‑V4‑Flash teacher; provided as a cleaned JSONL train split for distillation and SFT experiments.
1,000 JSONL samples containing full chain-of-thought reasoning traces and final answers produced by DeepSeek‑V4‑Pro for use in student-model distillation and quality checks. Prompts sampled from Jackrong/GLM-5.1-Reasoning-1M-Cleaned; Apache‑2.0 licensed.
Provides unquantized BF16 weights of Qwen3.6-27B with the base model's MTP head grafted in for high-fidelity, uncensored text (and multimodal) generation. Includes deployment guidance and hardware-tuned variants for A100/H100 and Blackwell-class GPUs.
Contains full chain-of-thought traces and final answers generated by DeepSeek-V4-Pro for use as distillation supervision. Key features: full CoT exposure, ~1,000 mixed-domain samples (JSONL/Parquet), Apache-2.0 license — suitable for training student models but watch for source contamination.