AIAIAny
  • Search
  • Collection
  • Category
  • Tag
  • Daily AI
AIAIAny

Discover the Best AI Resources

Curated essentials, no noise — just what matters

AIAIAny

Curated AI Resources for Everyone

[email protected]

Powered by airss.app

Product
  • Search
  • Collection
  • Category
  • Tag
Resources
  • Blog
Company
  • Privacy Policy
  • Terms of Service
  • Sitemap
Copyright © 2026 All Rights Reserved.

Contents

Hugging Face
AI Dataset·2026
Icon for item

Last Translation Benchmark

Vilém Zouhar, Niyati Bafna +8·ETH Zurich, Johns Hopkins University +3

A living, crowdsourced dataset of hard-to-translate examples (text, images, audio, video) paired with handcrafted verification rules that flag concrete MT failures. LTBv1 contains 3,456 peer-reviewed examples across many language pairs and accepts ongoing contributions.

#translation#multilingual#benchmark#evaluation#nlp+5
Hugging Face
AI Dataset·2026
Icon for item

UltraData-SFT-Agent-2609

openbmb, MiniCPM Team

Provides ~483K agent instruction‑tuning trajectories for supervised fine‑tuning, including tool calls, environment feedback, errors/retries and verification across search, code, office and general agent workflows; static snapshots for SFT and mix‑ratio studies.

#llm#sft#ai-agent#agent-skills#coding-agents+3
Hugging Face
AI Dataset·2026
Icon for item

UniPhys-Bench

spatialverse, Manycore Tech Inc.

Provides a human-verified benchmark of 1,927 heterogeneous articulated 3D objects with part-level articulation semantics and intrinsic physical-property annotations for evaluating physical grounding and simulation readiness. Includes URDF assemblies, aligned point clouds, per-part JSON annotations, and a curated evaluation protocol; licensed CC BY-NC 4.0 (non-commercial).

#robotics#physics#benchmark#benchmarks#evaluation+3
Hugging Face
AI Video·2026
Icon for item

Viggle-Animate

Viggle Research, MiniMaxAI

Replaces a character in a video using a single repainted frame from the same clip and propagates that edit across the shot while preserving motion, camera and lighting; requires no pose estimator, segmentation, face tracker or text prompt. Key facts: a 33.1B MiniMax-H3 finetune, DMD-distilled to three forward passes, 124 frames in ~26s on one B200 GPU.

#distillation#diffusers#safetensors#video#ai-video+3
Large Language Model Papers·2026
Icon for item

Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement

Yi Ding, Ruqi Zhang·Affiliation: Department of Computer Science, Purdue University, USA

Analyzes on-policy distillation for LLM fine-tuning, shows teacher token-level supervision is often noisy and not the main driver of gains, and introduces OPSA, a supervision-free, entropy-adaptive method that suppresses low-probability tokens to improve downstream accuracy.

#distillation#LLM#rl#nlp#qwen+2
Hugging Face
AI Model·2026
Icon for item

DeepSeek-V4-Flash-Vision-Exp

DeepSeek AI

An experimental multimodal model that adds visual understanding to DeepSeek-V4-Flash: accepts text+image inputs and returns text analyses. Improves vision-dependent agent workflows while maintaining comparable text-only performance; released under an MIT license on Hugging Face.

#deepseek#multimodal#vision#transformers#safetensors+7
Computer Vision Papers·2026
Icon for item

Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling

Minghan Qin, Yuang Wang +7·ByteDance Seed, Peking University +1

Converts posed indoor RGB(-D) video into editable, simulation-ready 3D scene graphs by parsing multi-view evidence into per-object bundles, generating complete object assets from that evidence, and placing them with GizmoAct, a VLM policy that refines 9-DoF poses through closed-loop GUI actions.

#vision#robotics#multimodal#depth#rl+2
AI Agent Papers·2026
Icon for item

PaperGym: Rubric-Centered Evolution for Research-Plan Generation

Yuhan Wang, Zhengxi Lu +7·Zhejiang University, Apple

Turns each research paper into a training environment to generate verifiable research plans by synthesizing questions from goals/background and deriving evaluation criteria from methods/experiments. Key features: four-stage extraction that reduces criterion leakage to 3.7%, a two-stage rubric-centered training (self-distillation then GRPO), and the PaperGym-20k corpus with two held-out benchmarks.

#paper#research#rl#RL#qwen+6
Computer Vision Papers·2026
Icon for item

Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving

Xin Zhou, Zongchuang Zhao +14·Huazhong University of Science and Technology

Develops a vision-language foundation model for autonomous driving that unifies 3D BEV perception, visual question answering, and motion planning without changing the pretrained VLM architecture. Key elements include an external BEV perception head for 3D detection and occupancy, a Planning Expert using flow-matching for trajectory prediction, and a staged training recipe combining driving and general VLM data.

#qwen#vision#multimodal#foundation-model#flow-matching+3
Embodied AI·2026
Icon for item

ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training

Xionghao Wu, Yijun Yang +18·Joy Future Academy

Learns generalizable World Action Models for robotic manipulation by scaling causal egocentric video pretraining and grounding learned dynamics with heterogeneous robot trajectories. Key features: a three-stage curriculum (video pretraining, video-action mid-training with a unified action representation, and target-robot specialization) and a Slow–Fast dual-system for 30 Hz real-time action prediction.

#video#robotics#ai-video#multimodal#paper+3
Hugging Face
AI Dataset·2026
Icon for item

Smart Contract Audit Findings

Zaevlad

Provides 23,625 semi-structured smart-contract audit findings (title, description, PoC, recommendation, normalized severity) for defensive-security research; requires cleaning, deduplication, and PoC filtering before model training.

#security#ai-security#parquet#polars#huggingface+3
Hugging Face
AI Dataset·2026
Icon for item

Text-to-Speech Human Preferences (315K)

Datapoint AI

Provides 315,000 pairwise human-preference votes comparing 15 English TTS models over 300 operational prompts, with 4,500 high‑quality audio renders and structured vote/pair/prompt records for training or evaluating preference/reward models. Metadata under CC-BY-4.0; audio use governed by model providers' terms.

#tts#audio#speech#benchmark#evaluation+4
  • Previous
  • 1
  • More pages
  • 189
  • 190
  • 191
  • More pages
  • 211
  • Next