AIAIAny
  • Search
  • Collection
  • Category
  • Tag
  • Daily AI
AIAIAny

Discover the Best AI Resources

Curated essentials, no noise — just what matters

AIAIAny

Curated AI Resources for Everyone

[email protected]

Powered by airss.app

Product
  • Search
  • Collection
  • Category
  • Tag
Resources
  • Blog
Company
  • Privacy Policy
  • Terms of Service
  • Sitemap
Copyright © 2026 All Rights Reserved.

Contents

Reinforcement Learning Papers·2026
Icon for item

ClawGym II: Exploring Black-Box RL on Agent Harness

Huatong Song, Fei Bai +18

Presents a unified black-box reinforcement learning framework to train and optimize agents running inside complex execution harnesses. Uses sandbox-parallel rollouts, a serving proxy that captures model calls and reconstructs multi-turn trajectories as prefix trees, and adapted GRPO/PPO optimizers to achieve stable, scalable RL across heterogeneous harnesses.

#rl#ai-agent#long-horizon#qwen#claude-code+5
Hugging Face
AI Dataset·2026
Icon for item

Claude protein binder design (data release v1.0)

Anthropic, Adaptyv Bio +1

Provides experimental and in-silico data for 1,440 de novo miniprotein binders designed by Anthropic's Claude models, including per-design kinetics, raw sensorgrams, structure-predictions, and design provenance. Includes two independent wet‑lab assessments and extensive per-design files; data released under CC BY 4.0.

#anthropic#huggingface#benchmark#biology#drug-discovery+1
Large Language Model Papers·2026
Icon for item

FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution

Shuo Yang, Xiaoze Fan +9

Enables interactive serving of large Mixture-of-Experts (MoE) models on personal machines by adapting offload and execution to measured device bandwidth and agentic workload patterns. Key features include bandwidth-adaptive execution, semantic-aware caching of recurrent state, and an elastic GPU expert cache; supports 20+ MoE models and runs models from ~35B to 753B on consumer/workstation GPUs.

#ai-serving#ai-inference#ai-deploy#mLOps#coding-agents+3
AI Agent Papers·2026
Icon for item

Zetta $ζ$: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence

Xin Ding, Liang Mi +13·Institute for AI Industry Research (AIR), Tsinghua University, Z-Trans AI

Enables closed-loop execution for embodied agents by evolving code-based runtime critics and recovery skills online while keeping the base policy frozen. Combines three timescale loops with Z-Infra rollout infrastructure; reports 90.8% on LIBERO-Pro, 93.6% on RoboCasa and an 11.1× inference speedup.

#robotics#agent-skills#ai-agent#RL#mLOps+2
Hugging Face
AI Video·2026
Icon for item

Minimax H3 Latent Upscaler

LBH-123-AI

Upscales Minimax H3 24-channel VAE latents in-place to increase spatial resolution while preserving the time dimension. Replaces the decode→pixel-upscale→encode round-trip with a learned 2D/3D latent upscaler to save compute and avoid interpolation ghosting; supports 1.0–4.0× scaling.

#video#ai-video#safetensors#pytorch#huggingface+2
Hugging Face
AI Model·2026
Icon for item

Thomson-1.0-Small

Shengzhuang Chen, Jerrod Parker +24·Thomson Reuters, Imperial College London +2

Injects proprietary news, regulatory and legal data into an open checkpoint via data-centric continual learning to improve performance on legal, tax and journalism tasks while preserving general capabilities and very long context support.

#foundation-model#qwen#llm#huggingface#safetensors+5
AI Agent Papers·2026
Icon for item

ASI-Bench: At the Dawn of Artificial Superintelligence

Junwei Zhou, Zhen Sun +40

Evaluates whether AI systems can independently carry out project-level scientific research by progressively removing human methodological guidance across 60 tasks in 11 domains. Built with expert review, sandbox execution, and multi-agent–model scoring to measure innovation and autonomous experimental execution.

#benchmark#benchmarks#ai-agent#agent-skills#research+3
Hugging Face
AI Model·2026
Icon for item

Ornith-1.5-9B

Ornith Team

A 9B open-weight reasoning LLM that uses a self-improvement loop to auto-generate tasks, construct scaffolds, and optimize rollouts for stronger agentic coding and long-context reasoning. Single-GPU deployable, supports tool-calling and a 262,144-token context window.

#transformers#safetensors#qwen#reasoning#coding+9
Hugging Face
AI Model·2026
Icon for item

orcarouter/Qwen3.8-27B-Uncensored

orcarouter, Qwen

Provides an abliterated (refusal-removed) build of Qwen3.8-27B for offline research and red‑teaming, keeping multimodal vision, an MTP speculative head, and a 262,144-token context. It has no built-in safety guardrails and is released under Apache‑2.0 for research use only.

#qwen#safetensors#transformers#multimodal#vision+5
Reinforcement Learning Papers·2026
Icon for item

Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL

Yunhao Yang, Yuexin Bian +7·University of Exeter, Independent Researcher +1

Uses cooperative multi-agent RL where multiple decoupled models provide peer-derived pseudo-rewards to each other, enabling unsupervised improvements in reasoning; increases cohort diversity to reduce correlated errors and avoid training collapse, showing consistent gains across text and multimodal benchmarks.

#rl#LLM#multimodal#reasoning#vision+4
Hugging Face
AI Audio·2026
Icon for item

AuK: An Open-Source Foundational Model for Speech Generation and Editing

Ziyang Ma, Zhikang Niu +31·Tencent (Tencent Hunyuan)

Generates and edits speech from natural-language instructions plus optional reference audio, supporting zero-shot TTS, content/acoustic/paralinguistic edits, enhancement, and source separation. Open-source 1.5B-parameter base model with a 4-step distilled AuK‑Flash for faster inference.

#audio#speech#tts#voice#foundation-model+5
Embodied AI·2026
Icon for item

Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation

Hongyan Feng, Sunlai Chen +10

Turns embodied navigation into 2D visual prompting where a vision-language model selects image pixels that are projected to 3D actions; adds selective chain-of-thought, compressed anchor-trajectory memory, and a two-level alignment objective to improve sample and runtime efficiency.

#vision#robotics#rl#multimodal#paper+5
  • Previous
  • 1
  • More pages
  • 180
  • 181
  • 182
  • More pages
  • 205
  • Next