AIAIAny
  • Search
  • Collection
  • Category
  • Tag
  • Daily AI
AIAIAny

Discover the Best AI Resources

Curated essentials, no noise — just what matters

AIAIAny

Curated AI Resources for Everyone

[email protected]

Powered by airss.app

Product
  • Search
  • Collection
  • Category
  • Tag
Resources
  • Blog
Company
  • Privacy Policy
  • Terms of Service
  • Sitemap
Copyright © 2026 All Rights Reserved.

Contents

AI Video Papers·2026
Icon for item

DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation

DreamX Team, Rui Chen +8

Predicts future video frames conditioned on an observed frame, a language instruction, and a sequence of end-effector poses and gripper states for robot manipulation. Uses per-arm SE(3) geometric encoding (PRoPE-style), a lightweight depth branch, SAM3 masks with a frozen V-JEPA teacher, and distribution-matching distillation for efficient, consistent action-conditioned rollouts.

#robotics#video#depth#distillation#vision+6
Hugging Face
AI Model·2026
Icon for item

unsloth/Qwen3.8-27B-GGUF

unsloth, Qwen Team

Provides a 27B Qwen3.8 GGUF build for local/offline deployment, optimized with Unsloth Dynamic V3.0 quantization. Offers switchable thinking-mode, native vision-language understanding, and native long-context support (262k+ tokens).

#qwen#llm#multimodal#vision#video+5
Hugging Face
AI Model·2026
Icon for item

DeepSeek-V4-Pro-0813

DeepSeek-AI

Provides a Mixture-of-Experts language model tuned for million-token contexts and agentic workflows, with DSpark speculative decoding, FP4/FP8 mixed-precision support, and vLLM/SGLang deployment recipes for low-latency production inference.

#deepseek#safetensors#transformers#vllm#huggingface+6
Hugging Face
AI Model·2026
Icon for item

unsloth/Qwen3.8-27B-NVFP4

unsloth

A 27B Qwen3.8 vision‑language causal transformer quantized to NVFP4 for lower‑memory inference. Provides 262K native context (extensible to 1M), Unsloth Dynamic V3.0 4‑bit quantization and MTP support so Qwen3.8‑class multimodal workloads can run on 24GB‑class GPUs.

#qwen#safetensors#huggingface#llm#multimodal+5
AI Agent Papers·2026
Icon for item

Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development

Yiwei Li, Wanli Yang +11·Affiliation: Meituan, Affiliation: University of Chinese Academy of [email protected]

Systematically evaluates LLM-driven autonomous agents on long-horizon AI research tasks using rule-based within-run metrics (Solution Framing, Execution, Feedback Control). Focuses on experience reuse and harness effects across 36 tasks and seven frontier models, finding agents act more like engineering optimizers than autonomous researchers.

#long-horizon#evaluation#benchmark#ai-agent#agent-skills+1
Computer Vision Papers·2026
Icon for item

HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark

Dairu Liu, Zekun Qi +12

Provides a large-scale benchmark and a human-aligned metric for humanoid whole-body motion tracking — about 153 hours of optical mocap from professional performers plus HumanScore trained on 12K human-labeled preference pairs to reveal contact, timing, and stability failures.

#robotics#mocap#evaluation#benchmarks#long-horizon+1
AI Agent Papers·2026
Icon for item

OmniScientist: An Omni-Modal Omni-Discipline AI Scientist

Bobo Li, Hao Fei +3·1National University of Singapore 2University of Oxford, Project page: https://omni-scientist.github.io +1

Conducts end-to-end multidisciplinary research directly from heterogeneous raw evidence using lifecycle-wide perception and three autonomous agents (Ideation, Experiment, Writeup). Integrates perceptual analysis, execution provenance, and code-enforced checks to produce executable analyses, validated results, and compiled manuscripts across many modalities.

#multimodal#agent-skills#ai-agent#paper#research+5
Hugging Face
AI Model·2026
Icon for item

TeleOCR

Peng Cai, Zhaofan Zou +7·StarDoc-AI, China Telecom

Performs unified parsing of digital and camera-captured documents (layout, text, tables, formulas) using a ~1.2B-parameter vision–language model. Key differences: geometry-aware modeling, curvature-guided sampling, and content-structure decoupled training to handle real-world deformations without separate dewarping.

#ocr#multimodal#transformers#pytorch#huggingface+7
Hugging Face
AI Dataset·2026
Icon for item

Ultra-FineWeb-L1

Junshao Guo, Shuaikang Xue +10

Provides an L1 filtered English web corpus from recent Common Crawl snapshots for LLM pretraining, including main-text extraction, language and heuristic filtering, sensitive-field replacement, customized cleaning, and MinHash deduplication; contains 1T+ tokens across ~1.14B documents with structured metadata fields.

#llm#foundation-model#nlp#huggingface#parquet+4
AI Agent Papers·2026
Icon for item

Agentic Transaction: Towards ACID-Compliant Agent Systems

Zhaoyan Sun, Xiaoxiao Wang +1·Tsinghua University

Defines "agentic transactions" and an ACID-style reliability framework for LLM agents that manage long-horizon tasks over persistent environments. Implements an ACID-compliant data agent using exploration–execution–validation cycles, confidence-divergence checks, semantic isolation, and append-only durable workspaces.

#LLM#ai-agent#coding-agents#long-horizon#reasoning+5
Hugging Face
AI Model·2026
Icon for item

Qwen3.8-27B RVN Heretic Abliterated Uncensored (GGUF)

0bserverx, Tim Rohrbaugh

Provides uncensored variants of Qwen3.8-27B modified with ARA (Arbitrary-Rank Ablation) to surgically remove refusal behavior, packaged as GGUF quant files for local llama.cpp inference. RVN applies two extra ARA passes that reduce harmful-prompt refusals to 0–1/100 with very low KL damage; intended for adult research/creative use and reduces safety guardrails.

#qwen#gguf#huggingface#llm#llama.cpp+3
Hugging Face
AI Dataset·2026
Icon for item

HiPHI

Ji Jiahao, Ma Ji +13·Noitom Robotics, ModalityNet

Provides 617.5 hours of high-precision optical motion-capture with synchronized object trajectories and standardized 55-joint BVH for whole-body and human–object interaction research. Frame‑LU indexed and paired with natural-language descriptions; designed for humanoid learning, motion priors, and interaction-aware benchmarks.

#mocap#robotics#huggingface#benchmarks#benchmark+3
  • Previous
  • 1
  • More pages
  • 177
  • 178
  • 179
  • More pages
  • 203
  • Next