Proposes SkillOpt-Lite, a minimal pipeline for optimizing LLM agent skills by treating rollout traces as filesystem files and applying trajectory exploration, consensus mining, and independent validation; integrates as a one-line VSCode Copilot command and reports cross-benchmark improvements that let smaller models sometimes outperform larger ones.
Instruction-tuned compact conversational model (Qwen3-4B-based) that generates short, chat-style replies and is optimized to run on a single mid-range GPU. Uses ChatML prompts, bfloat16 safetensors and is released under Apache-2.0; the model card notes a joke/placeholder disclaimer.
Converts long or messy model reasoning traces into concise, user-facing summaries with optional metadata. 61K cleaned English samples in JSON format, Apache-2.0 licensed, created to train and evaluate reasoning-summarization models and to present safe, readable explanations instead of raw chain-of-thought.
Runs a full 27B-class language model using end-to-end binary (1.125-bit) weights, cutting FP16 size to ~3.9 GB. Key features: 262k-token context, custom 1-bit kernels for Apple MLX and CUDA, and an optional DSpark drafter for faster decoding. Best when memory footprint matters; trades some FP16 accuracy for on-device feasibility.
Provides a 27B-class Qwen3.6-derived language model in GGUF with end-to-end ternary weights (Q2_0_g128), reducing deployed footprint to ~7.2 GB while retaining ~95% of FP16 reasoning ability and enabling on-device 262K-token context inference.
Runs a full 27B-class Qwen3.6-derived language model in a ~3.9 GB 1-bit GGUF pack for on-device inference with a 262K-token context; true 1.125 bits/weight binary representation, DSpark speculative drafter, and llama.cpp (CUDA/Metal/CPU) support.
Runs a full 27B-class Qwen3.6-derived LLM in a ~7.2 GB ternary/2‑bit format for on-device or single‑GPU text generation, retaining ~95% of FP16 performance and supporting a 262K‑token context. Designed for laptop/GPU deployment; exceeds typical phone memory limits.
Decides whether a user prompt should be executed locally on an edge small LLM or routed to a larger cloud model, emitting a deterministic pipe-separated decision string. A 51.7M micro-LLM fine-tuned with multi-task sequence generation to predict domain, complexity and code/math flags, optimized for ultra-low latency edge routing.
Provides labeled prompts with full-reference answers (including chain-of-thought and code blocks) and per-example metadata to train edge routing/orchestrator models that decide whether to handle inputs locally or route them to larger models. Includes complexity scores, coding/math flags, routing justifications, and an automated override rule; suited for fine-tuning small models (50M–1.5B) for edge deployment.
Provides a reusable skill suite for evidence-grounded research ideation: Paper-Search for multi-source literature retrieval, Scoop-Check for prior-art collision checking, and IdeaSpark for pattern-guided idea generation, evidence auditing, and idea-card rendering.
Fine-tuned variant of Qwen3.6-27B that cuts internal reasoning (‘thinking’) token usage by roughly 46% on average while preserving benchmark accuracy and safety behavior. Targets lower latency and inference cost; ships on Hugging Face with GGUF quantizations for local use.
Deployment-optimized hybrid MoE LLM (75B total / 9.3B active) produced via Iterative Puzzle compression and Multi-Token Prediction to double server throughput and raise single-GPU concurrency; designed for multilingual reasoning, long-context generation, and high-volume agentic/chat deployments.