Provides a diagnostic suite that audits video-understanding benchmarks to find samples solvable without visual or temporal input, filters those shortcuts, and produces a distilled video-native testbed that reveals major capability gaps in current Video-LLMs.
Provides a curated collection of DESIGN.md files extracted from real websites so AI coding and design agents can generate visually consistent UIs from a single markdown file. Includes previews, extracted tokens, and ready prompts for quick agent integration.
A dense 128B multimodal model with a 256k context window, configurable reasoning effort, and native function-calling for agentic workflows. Supports text+image input, multilingual output, and is released on Hugging Face under a Modified MIT license with revenue-based exceptions.
Turns natural-language instructions into runnable trading research: data loaders, strategy generation, backtests, reports, and optional broker connectors. Focuses on a tool-driven agent model (36+ MCP tools, 77 finance skills) and an Alpha Zoo of 452 pre-built alphas for reproducible research and gated agentic trading.
Generates and iterates on long‑horizon agentic plans and code — designed to stay productive across many rounds of tool calls and experiments. Emphasizes iterative reasoning, stronger repo/terminal automation and code generation than GLM‑5, and can be served locally for research and autonomous-agent workloads.
Aggregates and deduplicates public Claude distillation datasets into a unified 'messages' format with source attribution; focused on instruction-tuning and reasoning samples for SFT and LLM training, while requiring users to follow original sources' licenses.
Multimodal image-text-to-text fork of Gemma 4 (31B) using a 'CRACK v2' abliteration — tuned for conversational vision inputs and thinking-mode support in JANG v2 safetensors format. Recommended to run in vMLX; published by dealignai.
An 8B-parameter, instruction-tuned long-context LLM optimized for instruction following, tool-calling, and multilingual dialogue — supports 131072-token context and common NLP tasks such as summarization, QA, code, and RAG.
A 30B-parameter, instruction-tuned language model built for long-context text generation, conversational agents, and tool-calling. It combines supervised fine-tuning and RL alignment, supports 131,072-token context, and is optimized for tasks like summarization, code, and RAG.
Provides fully local long-term and symbolic short-term memory for AI agents via a 4-tier layered pipeline and Mermaid canvases, with zero external API dependencies. Key features: lossless drill-down from personas to raw traces, hybrid retrieval, and ready integrations for OpenClaw and Hermes.
Text-generation LLM designed for agentic workflows: supports multi-agent 'Agent Teams', skill stacks and model self-evolution. Ships on Hugging Face with deployment guides (vLLM, Transformers, SGLang) and is positioned for engineering, tool-calling and productivity use cases.
Provides a ~9.2M-instance Japanese multimodal post-training dataset for vision–language models, combining image–text pairs, PDF corpora and generated VQA to improve Japanese VLM performance; access is restricted by Japanese copyright (download via llm-jp GitLab).