A retrieval benchmark suite focused on “oblique queries,” where relevance depends on latent attributes rather than surface keywords. Includes five tasks with large corpora, qrels (and pooled judgments), and task-specific constraints for evaluating embedding-based retrievers and reasoning-augmented retrieval.
Preview of an MoE model family (V4-Pro: 1.6T params, 49B active; V4-Flash: 284B, 13B active) built for 1M-token contexts. A hybrid attention design cuts single-token inference FLOPs to 27% and KV cache to 10% versus V3.2 at million-token length.
A reasoning-enhanced Mixture-of-Experts (MoE) LLM fine-tuned for multimodal image-text-to-text tasks and long-context reasoning; built on Qwen3.6-35B-A3B with LoRA and released as an experimental GGUF community model.
Provides JSON traces from a Codex-driven swebenchpro agentic benchmark, including per-call token counts, cache hit rates, timing, and per-trial outcomes. Useful for research into LLM caching, long-context workloads, and agent evaluation. MIT-licensed and compact.
Provides 1,781 OpenTelemetry execution traces of LLM-powered agents across six benchmarks, including full conversations, token usage, timing, tool calls and model metadata—useful for performance analysis, agent-behavior research, and inference debugging.
Trains reusable natural-language 'skills' for frozen LLM agents by optimizing the skill document in text-space — using trajectory-driven edits, validation-gated updates, and deployable best_skill.md artifacts. Multi-backend, zero inference-time cost at deployment, designed for iterative, validation-led skill improvement.
Contains 4,006 newline-delimited JSONL agent-session traces recording assistant responses and tool calls from deepseek/deepseek-v4-pro — includes a training-ready tools schema snapshot and helpers for conversion to SFT/distillation workflows.
Provides a curriculum-aligned knowledge graph extracted from Chinese K–12 textbooks and accompanying benchmarks and training data to evaluate and train educational LLMs. Releases a 23,640-question multi-select benchmark and a 7,335-sample graph-guided training corpus with multimodal VQA pairs and the full construction pipeline.
Processes text and images to produce conversational, reasoning-focused multilingual outputs for agentic workflows. Built as a sparse MoE decoder (25B active / 218B total parameters) with 128K context and available in BF16/FP8/W4A4 quantizations to balance quality and deployability.
A 30B mixture-of-experts multilingual translation model supporting 33 languages and instruction-following translation. Offers MoE architecture, fast-thinking mode, and quantized/deployment-ready variants for production translation and subtitle tasks.
A family of multilingual translation models optimized for real-world, instruction-following translation across 33 languages. The 1.8B model targets on-device use with extreme quantization (≈440 MB via AngelSlim), while 7B/30B variants trade size for higher accuracy.
7B multilingual translation model optimized for instruction-following and low-latency deployment across 33 languages; provides quantized/FP8/GGUF builds and integrations (vLLM, llama.cpp) for server and on-device inference.