AIAIAny
  • Search
  • Collection
  • Category
  • Tag
  • Daily AI
AIAIAny

Discover the Best AI Resources

Curated essentials, no noise — just what matters

Contents

AIAIAny

Curated AI Resources for Everyone

[email protected]

Powered by airss.app

Product
  • Search
  • Collection
  • Category
  • Tag
Resources
  • Blog
Company
  • Privacy Policy
  • Terms of Service
  • Sitemap
Copyright © 2026 All Rights Reserved.
Embodied AI·2026
Icon for item

HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone

Yuteng Wei, Jinming Ma +15

Provides a portable, robot-free UMI capture pipeline and shows that policies post-trained only on this high-fidelity data deploy directly on real robots matching teleoperation baselines. Capture achieves ~3 mm end-effector accuracy, microsecond sync, ultra-wide FOV, and releases 2,000h HiFi-UMI-2K.

#robotics#vision#paper#ai-deploy#ai-train+1
Large Language Model Papers·2026
Icon for item

DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space

Jiangwang Chen, Zixin Song +11·Tsinghua University, Qwen Business Unit of Alibaba +2

Co-evolves a solver skill and a rubric-generator skill for text-space LLM optimization under decoupled objectives to avoid rubric gaming without using gold rubrics. Solver updates use criterion-level feedback; generator updates use independent audits of requirement coverage and response discrimination.

#LLM#evaluation#agent-skills#qwen#paper+2
Reinforcement Learning Papers·2026
Icon for item

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization

Bo-Wen Zhang, Junwei He +6

Allocates token-level credit in rubric-conditioned GRPO by counterfactually replaying the same response under rubric and criteria-free prompts, using tokenwise log-likelihood contrasts to compute bounded, response-normalized weights that redistribute GRPO advantages without training an auxiliary scorer.

#RL#LLM#NLP#paper#evaluation
Hugging Face
AI Dataset·2026
Icon for item

ASI-Bench Generated Instances (Seed 42)

Apexintelligence-AI

Provides fixed-seed benchmark instances (prompts and agent-visible inputs) for ASI-Bench to run reproducible evaluations of LLM agents on scientific tasks. Includes four matched prompt levels (B1–B4) across 60 project-level tasks in 11 domains; excludes reference answers and private scorers; Apache-2.0 licensed.

#benchmark#benchmarks#evaluation#ai-agent#agent-skills+4
Computer Vision Papers·2026
Icon for item

CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition

Lai Wei, Chengqi Li +4

Evaluates multimodal context learning across grounding, new information application, and knowledge acquisition using a 3,443-instance benchmark spanning science, finance, long documents, spatial reasoning, and web VQA; finds current multimodal models perform poorly (best score 0.2847) and analyzes failure modes.

#multimodal#benchmark#vision#evaluation#paper+4
Computer Vision Papers·2026
Icon for item

ReDesign: Recovering Editable Design Structures from Images via Agentic Decomposition

Jooyeol Yun, Jintae Park +4

Recovers editable design files from raster images by growing an editable layer hierarchy via an agentic pipeline that selects and composes modality-specific tools. Introduces graceful verification (accept/prune/retry) to prevent error accumulation and presents the Figma Edit Replay Benchmark (909 files, 14,796 edits) to measure editability across layout, color, and text edits.

#vision#image#multimodal#paper#benchmark+3
Hugging Face
AI Dataset·2026
Icon for item

ACE-Data-0

Yukang Cao, Haozhe Xie +14·S-Lab, Nanyang Technological University, Singapore, ACE Robotics

Captures synchronized multimodal embodied-human data in real homes — egocentric and multi-view video, metric body/hand/object motion, audio, and tactile signals. Released under a gated non-commercial research license with identifiable participants and strict non-redistribution/privacy constraints.

#video#robotics#multimodal#audio#long-horizon+3
Hugging Face
AI Dataset·2026
Icon for item

ASI-Bench Generated Instances (Seed 31415)

Apexintelligence-AI

Generated instance set (seed 31415) for ASI‑Bench: includes four matched prompt variants, agent-visible inputs, reference artifacts, and instance metadata for 60 project-scale scientific research tasks across 11 domains; intended for evaluating autonomous research agents. Licensed Apache‑2.0.

#benchmark#benchmarks#ai-agent#agent-skills#research+4
Hugging Face
AI Audio·2026
Icon for item

NVIDIA NemotronLabs VoiceChat 11B

NVIDIA

An end-to-end 11B full-duplex speech model for real-time conversational AI that jointly performs streaming speech understanding and generation, enabling ~450 ms turn-taking, barge‑in and live tool calling in a single unified architecture; research use only.

#nvidia#huggingface#voice#speech#ASR+6
Machine Learning Foundation Papers·2026
Icon for item

Metis: Memory Foundation Model

Zeyu Zhang, Ziliang Guo +15

Presents Metis, a prototype memory foundation model that embeds a persistent native memory state into the backbone so historical experience is compressed and accessed via memory attention. Key features: forward-only, gradient-free online memory updates; memory-specific mid-training objectives; and a dual text/code memory design.

#foundation#llm#ai-agent#agent-skills#multimodal+3
AI Video Papers·2026
Icon for item

VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System

Haodong Li, Tianfei Ren +26

Converts text prompts into physically consistent videos by synthesizing executable Blender programs as a process-level chain-of-thought and using a dual-engine pipeline (deterministic simulation draft + draft-conditioned video editor). Ships with a VideoCoCo-3K draft–instruction–target dataset and shows substantial gains in physical-consistency benchmarks.

#video#ai-video#code#coding#coding-agents+5
Computer Vision Papers·2026
Icon for item

TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM

Hengyi Xie, Chenfei Yao +8

Directly maps visual observations and language instructions to continuous robot actions, replacing LLM-centric V→L→A pipelines. Uses separate visual and language encoders with lightweight bidirectional interaction and a compact decoder to cut inference cost and VRAM, achieving ~31 ms latency and <1 GB VRAM on an RTX 4090; suited for real-time robotic manipulation under tight compute budgets.

#robotics#vision#multimodal#paper#code+4
  • Previous
  • 1
  • More pages
  • 167
  • 168
  • 169
  • More pages
  • 197
  • Next