Matches detection paradigms to four stratified attack-surface layers of AI agents — infrastructure, protocol/tool, agent behavior, and model — and presents AI-Infra-Guard: an open-source red-teaming framework with rule-based infra scanning, LLM-driven audits of MCP servers and skill packages, and a jailbreak/attack-operator harness.
Provides quantized GGUF weights and configs for Agents‑A1 — a 35B Mixture-of-Experts agent trained for long-horizon, tool-enabled reasoning; supports 262K-context serving and runtimes like vLLM and SGLang.
Evaluates how long-term memory in LLM agents amplifies sycophantic behavior and when memory should or should not influence decisions. Provides five targeted tasks, 1,550 standardized samples, an evaluation pipeline, and baseline adapters to test memory use, conflicts, scope, updates, and personalization.
A code-agent model for Lean 4 that automates repository-level formal proofs and verification; a Mixture-of-Experts architecture (119B total, 6.5B active) with 256k context, multimodal input and an Apache-2.0 license.
Provides a benchmark and protocol to evaluate agents that iteratively edit executable policies under a fixed interaction budget, recording full execution–feedback–revise trajectories. Built from compact RL environments with trajectory-level diagnostics and hidden held-out validation.
Provides 100,891 JSON-formatted agent conversation examples where each assistant turn includes a short <think> internal reasoning trace before tool/function calls. Human-facing text and tool calls are preserved; intended to fine-tune models to produce concise, cost-efficient chain-of-thought for tool use.
Compiles natural-language function specifications into compact, locally-executable neural programs (PAW) that run on a small frozen interpreter; a 4B compiler emits LoRA adapters for a 0.6B runtime to provide offline, low-memory fuzzy text functions.
Introduces a bounded-memory, typed-retrieval contract for long-horizon LLM agents and evaluates it in Slay the Spire 2 — assembling per-decision prompts from five typed slots rather than appending raw transcripts. Key outputs include ablationable memory layers, 298 labeled trajectories, and reproducible analysis scripts.
Provides a large Mixture-of-Experts instruct LLM (295B total parameters, 21B active, 256K context) optimized for reasoning, long-context retention and agent workflows; open-sourced under Apache-2.0.
Provides a comprehensive benchmark to evaluate LLM-based data agents on realistic, multi-domain data-science workflows. Features skill-level ground-truth labels, 15 vertical domains (including real B2B tasks), and LLM-driven task generation to ensure coverage; includes an open testbed and agent evaluations.
Proposes SkillOpt-Lite, a minimal pipeline for optimizing LLM agent skills by treating rollout traces as filesystem files and applying trajectory exploration, consensus mining, and independent validation; integrates as a one-line VSCode Copilot command and reports cross-benchmark improvements that let smaller models sometimes outperform larger ones.
Trains cross-platform GUI agents by combining a Uni-GUI cross-platform dataset with platform-conditioned multi-teacher on-policy distillation, enabling a shared policy to adapt to new platforms while retaining platform-specific behaviors; suitable for research on continual GUI agent learning and cross-platform adaptation.