A 16 GB, 507-file PhD‑level cybersecurity knowledge base for training and evaluating security-focused LLMs and automation. Covers offensive/defensive/forensics/cloud/iot and AI-security across 30+ domains with real-world labs and framework mappings.
Provides 462 unrestricted long-form chain-of-thought reasoning traces distilled from the full Mythos V2 model (≈104.7M characters); intended for long-context evaluation, trace analysis and process-level supervision. License unknown—verify before reuse.
Proposes TrOPD, a method that restricts token-level on-policy distillation to regions where teacher supervision is reliable to stabilize training under teacher–student distribution mismatch. Adds outlier handling (clipping, masking, forward-KL) and off-policy guidance; shows consistent gains on math reasoning, code generation and general benchmarks.
Studies small trainable adapters (PEFT) used as persistent personal models on top of large foundation models, analyzing three scaling axes—Scale Up, Scale Down, Scale Out—and introducing MinT, an infrastructure for adapter identity, provenance, evaluation, and serving.
Localizes harmful span-level errors inside long research-agent trajectories to show which trajectory segments make final answers unreliable. Provides a 1,000-instance TELBench of annotated spans and DRIFT, a claim-centric auditing method that improves span-level localization and first-error accuracy by up to 30 percentage points.
Analyzes how single-domain RL fine-tuning on LLMs induces cross-domain interference and shows this damage concentrates in a low-dimensional shared conflict subspace; proposes a local perturbation theory and short domain "refresh" procedures that selectively recover earlier domains with minimal collateral loss.
Evaluates multimodal LLMs on streaming egocentric video for spatial intelligence using 1,680 human-annotated questions across 348 videos; organizes tasks into four hierarchical levels (perception → tracking → simulation → allocentric mapping) and highlights allocentric mapping as the main bottleneck.
Studies when and how to combine visual future rollouts from world models with abstract reasoning in multimodal LLMs. Proposes PF-OPSD — a teacher-student distillation that uses ground-truth future videos during training — and evaluates on two human-verified benchmarks, improving accuracy ≈10% while improving robustness to noisy rollouts.
Learns fine-grained preferences over sub-trajectories to identify and penalize redundant steps in long chain-of-thoughts, letting models "fold" reasoning chains into concise paths; reports ~56% token reduction on DeepSeek-R1-Distill-Qwen-7B while keeping accuracy.
Enables agents to proactively discover multiple hidden problems in a user context and pair each with supporting evidence and concrete actions. Uses iterative discovery (batch rounds conditioned on prior finds) and reusable "thought templates" to expand coverage and ground claims.
Provides ~1M synthetic Salvadoran‑Spanish personas (148k records, ~300M tokens) grounded in 2024 census distributions for demographics, occupations and locations; intended for training/evaluating localized LLMs and synthetic-data workflows. CC BY 4.0, adults only.
Agentic LLM for long-horizon, environment-driven workflows: decomposes goals, generates and executes code/tool calls, evaluates outputs, and iterates. The Pro variant emphasizes coding and terminal execution and is published for use with sglang and multi-node H100 deployment.