AIAIAny
  • Search
  • Collection
  • Category
  • Tag
  • Daily AI
AIAIAny

Discover the Best AI Resources

Curated essentials, no noise — just what matters

AIAIAny

Curated AI Resources for Everyone

[email protected]

Powered by airss.app

Product
  • Search
  • Collection
  • Category
  • Tag
Resources
  • Blog
Company
  • Privacy Policy
  • Terms of Service
  • Sitemap
Copyright © 2026 All Rights Reserved.

Contents

Computer Vision Papers·2026
Icon for item

ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining

Hao Li, Ganlong Zhao +9

Converts large-scale egocentric human videos into robot-format pseudo-action trajectories and introduces ACE-EGO-0, a VLA pretraining framework that unifies camera-space actions, morphology conditioning, and reliability-aware weighting to jointly learn from noisy human and high-quality robot data for improved robotic manipulation transfer.

#vision#robotics#video#paper#multimodal+1
Hugging Face
AI Dataset·2026
Icon for item

ABC-130k

XDOF, UC Berkeley +3

Provides 130k+ bimanual teleoperation trajectories for robot imitation learning, recorded on low-cost YAM two-arm rigs and shared as MCAP episodes with subtask annotations, training code, and checkpoints.

#robotics#video#huggingface#ai-train#ai-development
GitHub
AI Agent·2026
Icon for item

nodeterm

Enes Kirca

Manages real tmux-backed terminals and AI agents as draggable nodes on an infinite pan/zoom canvas, with a Trello-style kanban view, persistent sessions that survive restarts, mobile companion support, and a browser Server Edition for self-hosting.

#terminal#coding-agents#ai-agent#claude-code#claude+6
Hugging Face
AI Model·2026
Icon for item

GLM-5.2

Z.ai (zai-org)

Provides a large language model optimized for long-horizon agentic tasks and end-to-end coding workflows — with a stable 1,000,000-token context, IndexShare sparse-attention and multi-level thinking-effort modes. MIT-licensed and designed for deployments that need sustained long-context reasoning and coding.

#foundation-model#vibe-coding#deepseek#transformers#llm+4
Large Language Model Papers·2026
Icon for item

LoopCoder-v2: Only Loop Once for Efficient Test-Time Computation Scaling

Jian Yang, Shawn Guo +17

Uses Parallel Looped Transformers (PLT) to make loop count a practical knob for code models, finding two loops give the best test-time gains. Trains 7B models on 18T tokens and attributes saturation beyond two loops to a gain–cost tradeoff from positional mismatch.

#paper#code#LLM#transformers#ai-coding+1
AI Agent Papers·2026
Icon for item

GameCraft-Bench: Can Agents Build Playable Games End-to-End in a Real Game Engine?

Tongxu Luo, Rongsheng Wang +23

Assesses whether coding agents can generate complete, playable games end-to-end inside the Godot engine. Implements an interaction-grounded evaluation (replayed demonstrations + rubric-guided multimodal judging) across 140 tasks and 15 game families; top agents score ~41%.

#evaluation#ai-coding#agent-skills#multimodal#paper+1
Hugging Face
AI Image·2026
Icon for item

Krea 2 (Comfy-Org/Krea-2)

Comfy-Org, Krea

Provides ComfyUI-ready repackaged checkpoints of the Krea 2 image model family for local text-to-image workflows. Includes RAW (undistilled base for fine-tuning and LoRA training) and Turbo (8-step distilled checkpoint for fast inference), using a Qwen Image VAE and Qwen3‑VL encoder.

#qwen#diffusers#huggingface#ai-image#image+2
Hugging Face
AI Model·2026
Icon for item

GLM-5.2-FP8

zai-org, GLM-5 Team +1

Provides FP8-quantized weights of GLM-5.2 — a 744B long-context LLM tuned for sustained 1M-token engineering, coding and agentic workflows; compatible with vLLM, Transformers, SGLang and Ascend NPU deployments.

#foundation-model#vibe-coding#transformers#huggingface#llm+4
Embodied AI·2026
Icon for item

Guava: An Effective and Universal Harness for Embodied Manipulation

Haowen Liu, Xirui Li +6

Provides a harness that lets language models control embodied manipulation via iterative perception–reasoning–action loops, semantic action abstractions, and multimodal observations. Demonstrates distilling capabilities into a 4B open-source model with under 2K simulated trajectories and shows sim-to-real generalization.

#robotics#multimodal#LLM#agent-skills#vision+2
Reinforcement Learning Papers·2026
Icon for item

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients

Byung-Kwan Lee, Ximing Lu +9

Proposes ZPPO, a distillation method that keeps the teacher inside prompts rather than injecting teacher gradients, using binary- and negative-candidate prompts plus a prompt replay buffer to recover learning signal on hard examples; shows gains for small Qwen3.5 students across 31 multimodal benchmarks.

#qwen#RL#llm#multimodal#vision+2
Hugging Face
AI Dataset·2026
Icon for item

Nemotron-Personas-Belgium

Pieter Delobelle, Pierre-Carl Langlais +12·Pleias, NVIDIA Corporation +1

Provides 1.8M synthetic Belgian personas (1.2M records; 300k per language) in Dutch/French/German/English, grounded in Belgian census distributions to improve representativeness for LLM training and evaluation. Includes 23 persona and contextual fields, CC BY 4.0 license, produced with NeMo Data Designer.

#nvidia#huggingface#nlp#LLM#gemma+3
Computer Vision Papers·2026
Icon for item

Reinforcing Dual-Path Reasoning in Spatial Vision Language Models

Yatai Ji, An-Chieh Cheng +14

Provides a dual-path approach for spatial vision-language models: a Language-Only Reasoning (LOR) path for stepwise linguistic deduction and a Detect-Then-Reason (DTR) path that detects 3D cues via region tokens before numerical inference. Trains with chain-of-thought cold-start supervision and reinforcement learning to improve 3D grounding and multi-step spatial reasoning.

#vision#multimodal#RL#paper#depth+1
  • Previous
  • 1
  • More pages
  • 142
  • 143
  • 144
  • More pages
  • 181
  • Next