AIAIAny
  • Search
  • Collection
  • Category
  • Tag
  • Daily AI
AIAIAny

Tag

Explore by tags

AIAIAny

Curated AI Resources for Everyone

[email protected]

Powered by airss.app

Product
  • Search
  • Collection
  • Category
  • Tag
Resources
  • Blog
Company
  • Privacy Policy
  • Terms of Service
  • Sitemap
Copyright © 2026 All Rights Reserved.
  • All

  • 30u30

  • ASR

  • ChatGPT

  • GNN

  • IDE

  • RAG

  • agent-skills

  • ai

  • ai-agent

  • ai-api

  • ai-api-management

  • ai-client

  • ai-coding

  • ai-demos

  • ai-deploy

  • ai-development

  • ai-framework

  • ai-image

  • ai-image-demos

  • ai-inference

  • ai-leaderboard

  • ai-library

  • ai-rank

  • ai-security

  • ai-serving

  • ai-tools

  • ai-train

  • ai-video

  • ai-workflow

  • AIGC

  • algorithms

  • alibaba

  • amazon

  • android

  • anthropic

  • arabic

  • audio

  • aws

  • benchmark

  • benchmarks

  • biology

  • blog

  • book

  • bun

  • bytedance

  • chatbot

  • chatgpt

  • chemistry

  • claude

  • claude-code

  • cli

  • clickhouse

  • code

  • codex

  • coding

  • coding-agents

  • comfyui

  • common-crawl

  • copilot

  • course

  • cpu

  • cuda

  • cursor

  • deepmind

  • deepseek

  • depth

  • devops

  • diffusers

  • distillation

  • docker

  • drug-discovery

  • electron

  • embeddings

  • embodied-ai

  • engineering

  • evaluation

  • facebook

  • finance

  • flow-matching

  • foundation

  • foundation-model

  • fp4

  • fp8

  • gcode

  • gcp

  • gemini

  • gemini-cli

  • gemma

  • genomics

  • gguf

  • gitHub

  • github

  • go

  • google

  • gpu

  • gradient-booting

  • grok

  • groq

  • huggingface

  • hy_v4

  • image

  • imatrix

  • ios

  • java

  • javascript

  • json

  • kimi

  • kotlin

  • kubernetes

  • laion

  • llama.cpp

  • LLM

  • llm

  • long-horizon

  • lora

  • mLOps

  • manipulation

  • math

  • mcap

  • mcp

  • mcp-client

  • mcp-server

  • meta-ai

  • meta-pytorch

  • metal

  • microsoft

  • mlops

  • mobile

  • mocap

  • moe

  • multilingual

  • multimodal

  • mysql

  • nli

  • NLP

  • nlp

  • nodejs

  • numpy

  • nvidia

  • ocr

  • ollama

  • openai

  • opencode

  • pandas

  • paper

  • parquet

  • physics

  • pi

  • plugin

  • polars

  • postgres

  • privacy

  • programming

  • prompt-engineering

  • pwa

  • python

  • pytorch

  • qwen

  • react

  • reasoning

  • red-teaming

  • redis

  • refactoring

  • reranker

  • research

  • retrieval

  • RL

  • rl

  • robotics

  • routing

  • rust

  • safetensors

  • science

  • security

  • segmentation

  • sft

  • shodan

  • skillkit

  • slam

  • software-engineering

  • sora

  • speech

  • sqlite

  • ssh

  • stt

  • supabase

  • swe

  • swift

  • tensorrt

  • terminal

  • thinking

  • trae

  • training-data

  • transformers

  • translation

  • tts

  • tutorial

  • typescript

  • unsloth-dynamic

  • vibe-coding

  • video

  • vision

  • vllm

  • voice

  • vue

  • vulkan

  • vultr

  • web-search

  • webdataset

  • windsurf

  • world-model

  • xAI

  • xai

  • youtube

AI Agent Papers·2026
Icon for item

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

Qiushi Sun, Kanzhi Cheng +21·The University of Hong Kong, Xi’an Jiaotong University +4

Evaluates vision-language model judges on computer-using agent (CUA) trajectories to measure verifier reliability. Provides OSReward-Hard and OSReward-Multi challenge sets, the OS-Shepherd-100K reasoning-annotated corpus, and trained OS-Shepherd reward models that match commercial judges at ~30–60× lower cost.

#evaluation#benchmark#benchmarks#vision#multimodal+4
AI Agent Papers·2026
Icon for item

MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations

Qiming Shi, Yulong Tao +11

A 365-day, order-level simulation benchmark for evaluating long-term coherence of LLM agents in seller-side e-commerce. Grounded in 98,843 real product records and 26 interactive tools, it pairs prompt upstream supplier signals with delayed downstream order outcomes to stress planning, memory, and tool use over long horizons.

#benchmark#benchmarks#long-horizon#agent-skills#LLM+3
AI Agent Papers·2026
Icon for item

OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution

Yunhao Chen, Xin Wang +7·Affiliation: Fudan University, Affiliation: Shanghai Artificial Intelligence Laboratory +1

Evolves persistent, stateful environments to red-team tool-using AI agents — provides 10K+ validated scenarios across 50 domains and a feedback-driven attack policy (EMHA) to surface long‑horizon safety failures.

#paper#benchmark#evaluation#agent-skills#long-horizon+4
AI Video Papers·2026
Icon for item

WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity

Yuxue Yang, Shuyao Shang +14

Introduces WorldExam, a diagnostic benchmark that evaluates controllable video world models across four levels from visual quality to inherent world reactivity. Covers 1,474 cases across eight tasks and supports camera-, action-, and language-driven paradigms, measuring scene-conditioned reactions beyond explicit instructions.

#video#vision#benchmark#benchmarks#evaluation+2
AI Video Papers·2026
Icon for item

JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion

Yicheng Xiao, Wenxun Dai +23·Joy Future Academy, JD

Performs real-time, instruction-guided video-to-video editing on streaming input using a 16B autoregressive diffusion model that preserves subject identity and long-term temporal coherence; achieves end-to-end 720p at ≈30 FPS on a single Nvidia B200 GPU. Key features include chunk-wise autoregressive adaptation, Source-Anchored Distribution Matching Distillation (SA-DMD) that reduces diffusion to a two-step generator, and Long-Horizon Autoregressive Distillation to mitigate temporal drift.

#video#ai-video#distillation#multimodal#vision+5
AI Agent Papers·2026
Icon for item

Beyond Simply Environment Scaling: Designing Effective Environment Distributions for Multimodal Agent Learning

Kejian Zhu, Zhuoran Jin +5

Analyzes how to build effective training environment distributions for multimodal agents and proposes Ability-aware Environment Selection (AES) and Hierarchical Difficulty Curriculum (HDC) to improve diversity and difficulty scheduling, yielding large relative gains in experiments.

#multimodal#agent-skills#RL#paper#ai-train+2
AI Video Papers·2026
Icon for item

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?

Qifeng Zhang, Kaixiang Huang +7

Evaluates VLMs' ability to form global spatial awareness from long-horizon egocentric video. Introduces GST-Bench: a VQA benchmark with human-verified questions from 6,790 minutes of synthetic video, reveals a large gap (best zero-shot 42.68 vs human 79.08) and provides GST-Train dataset.

#video#vision#multimodal#benchmark#evaluation+4
Hugging Face
AI Dataset·2026
Icon for item

ExtractBench

Boyang Zhang, Adrian Lyjak +3·Run Llama, LlamaIndex

Evaluates schema-guided structured extraction from documents: given a document and a JSON schema, systems must return a schema-valid JSON with page-and-box grounding. Covers 370 documents (4,869 pages) across 8 business domains and 67 document types; scores value accuracy, word/page grounding, and long-list completeness.

#benchmark#benchmarks#evaluation#huggingface#pandas+5
Large Language Model Papers·2026
Icon for item

LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers

Tao Feng, Fangxu Yu +10·University of Illinois Urbana-Champaign, University of Maryland, College Park +3

Frames LLM routing as a sequential decision process and introduces LLMRouter plus the xRouteBench benchmark to develop, evaluate, and deploy learned routing policies across heterogeneous LLMs, optimizing response quality versus inference cost.

#llm#benchmark#evaluation#ai-deploy#ai-inference+6
Natural Language Processing Papers·2026
Icon for item

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives

Yingpeng Ma, Jianhao Yan +7·Affiliation: NLP2CT Lab, University of Macau, Macau, China, Affiliation: Westlake University, Hangzhou, China +3

Evaluates whether LLM-driven storyteller agents preserve long-horizon logical consistency under adversarial player interventions. Introduces NCP-Bench (100 movie-derived narrative environments) with structured trajectory/commitments and automatic violation checks; finds strong LLMs often contradict themselves across multi-turn interactions.

#LLM#NLP#evaluation#benchmark#benchmarks+6
Hugging Face
AI Model·2026
Icon for item

Qwen3.8-2.4T-A95B

Qwen Team

A MoE causal large language model for long-horizon agents, coding, and multi-step reasoning: 2.4T parameters (95B activated), native 262,144-token context (extensible to 1,010,000), multi-token prediction, and configurable thinking-mode reasoning controls.

#qwen#foundation-model#LLM#transformers#huggingface+7
Hugging Face
AI Dataset·2026
Icon for item

Slop classifier dataset

bench-labs

Human-annotated text dataset that labels perceived “AI slop” with a continuous human slop_score (-1 / 0 / +1) plus provenance metadata (source_dataset, source_row_id, content_hash). Collected via Bench Labs SlopFinder from public datasets for training classifiers and studying subjective perception.

#huggingface#nlp#evaluation#benchmark#json+2
  • Previous
  • 1
  • More pages
  • 5
  • 6
  • 7
  • More pages
  • 13
  • Next