AIAIAny
  • Search
  • Collection
  • Category
  • Tag
  • Daily AI
AIAIAny

Discover the Best AI Resources

Curated essentials, no noise — just what matters

AIAIAny

Curated AI Resources for Everyone

[email protected]

Powered by airss.app

Product
  • Search
  • Collection
  • Category
  • Tag
Resources
  • Blog
Company
  • Privacy Policy
  • Terms of Service
  • Sitemap
Copyright © 2026 All Rights Reserved.

Contents

AI Agent Papers·2026
Icon for item

ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation

Jiahao Zhao, Xiaomin Yu +6

Orchestrates reasoning, external tool use, and native image generation under one unified multimodal agent policy via post-training. Introduces RAD-GRPO for agentic reinforcement fine-tuning and releases training data plus the full post-training infrastructure.

#multimodal#agent-skills#RL#ai-image#vision+3
Natural Language Processing Papers·2026
Icon for item

The Personalization Mirage: How LLMs Fabricate User Profiles, and Why Self-Monitoring Misleads

Yushi Sun, Yanjie Zhang +1

Evaluates how large language models fabricate user attributes in personalization and whether model self-monitoring is a reliable signal. Introduces MirageBench (150 personas, 6 personalization tasks, judge-validated faithfulness taxonomy) and a 12-model leaderboard revealing pervasive over-inference and a 'Self-Monitoring Inversion'.

#NLP#LLM#benchmarks#evaluation#privacy+2
Hugging Face
AI Model·2026
Icon for item

Qwen3.8-27B

Qwen

A 27B-parameter causal language model with a native vision encoder for image/video+text understanding, long-horizon agentic tasks, and tunable thinking-mode reasoning. Native 262,144-token context (extensible to 1,000,000) and production-focused inference recipes.

#qwen#transformers#safetensors#huggingface#multimodal+11
Hugging Face
AI Video·2026
Icon for item

MiniMax-H3 Turbo LoRA

larryvrh

A LoRA adapter for MiniMax-H3 that enables joint video + synchronized stereo audio generation in as few as 4 sampler steps, cutting sampling time roughly ~5×; early prototype under-trained, so 6–8 steps or newer checkpoints give better sharpness.

#lora#ai-video#video#audio#multimodal+3
Computer Vision Papers·2026
Icon for item

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes

Junlin Han, Shengbang Tong +5

Systematically studies how language and vision interact during unified multimodal pretraining, identifies mechanisms that enable modality synergy versus competition, demonstrates the benefit of early joint training, and derives efficient pretraining recipes validated at scale.

#multimodal#foundation-model#vision#paper#llm+2
Hugging Face
AI Video·2026
Icon for item

MiniMax-H3 Turbo 4-Step — ComfyUI Pruned-Model LoRAs

drbaph

Provides ComfyUI-compatible pruned/curve-form LoRA conversions of the MiniMax‑H3 Turbo 4-step audio‑video generation preview, including further-trained ckpt500 EMA and non‑EMA variants and an example ComfyUI workflow for low-step experiments.

#lora#ai-video#audio#video#huggingface+4
AI Video Papers·2026
Icon for item

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?

Qifeng Zhang, Kaixiang Huang +7

Evaluates VLMs' ability to form global spatial awareness from long-horizon egocentric video. Introduces GST-Bench: a VQA benchmark with human-verified questions from 6,790 minutes of synthetic video, reveals a large gap (best zero-shot 42.68 vs human 79.08) and provides GST-Train dataset.

#video#vision#multimodal#benchmark#evaluation+4
Hugging Face
AI Dataset·2026
Icon for item

British Library Book Images

Daniel van Strien·British Library, British Library Labs +3

Provides 1,080,814 images extracted from ~65,000 digitised British Library book volumes (c.1510–c.1900), split into four algorithmic image-type configs and packaged as parquet for image–text multimodal research and retrieval.

#image#book#ocr#parquet#huggingface+5
Reinforcement Learning Papers·2026
Icon for item

EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning

Zishan Xu, Zhiyuan Yao +10

Replaces external environment interaction in agentic RL training with 'world rehearsal': the policy alternates between making tool calls and simulating their environment responses, jointly optimizing both roles so the agent internalizes environment dynamics and improves long-horizon tool use and transfer.

#RL#rl#ai-agent#agent-skills#long-horizon+4
Hugging Face
AI Dataset·2026
Icon for item

ExtractBench

Boyang Zhang, Adrian Lyjak +3·Run Llama, LlamaIndex

Evaluates schema-guided structured extraction from documents: given a document and a JSON schema, systems must return a schema-valid JSON with page-and-box grounding. Covers 370 documents (4,869 pages) across 8 business domains and 67 document types; scores value accuracy, word/page grounding, and long-list completeness.

#benchmark#benchmarks#evaluation#huggingface#pandas+5
Hugging Face
AI Dataset·2026
Icon for item

Nemotron-RL-Agentic-Terminal-Pivot-v1

NVIDIA Corporation

Provides per-decision training samples for RL-driven command-line LLM agents: each record pairs a task prompt plus terminal history with a teacher's next-action in Terminus-2 JSON. Around 31k verifier-passing samples from 630 ATCB tasks, formatted for NeMo Gym's terminus_judge and licensed CC-BY-4.0.

#nvidia#huggingface#terminal#rl#agent-skills+4
Hugging Face
AI Dataset·2026
Icon for item

OpenH-RF

NVIDIA Corporation, Stanford University +28

Provides ~39 TB of pre‑beamformed (channel capture) ultrasound RF data and metadata in zea/HDF5 format for reconstruction, flow, and inverse‑problem tasks. Released under CC‑BY‑4.0 and curated for training and evaluating ultrasound/RF foundation models.

#nvidia#huggingface#foundation-model#training-data#ai-train+1
  • Previous
  • 1
  • More pages
  • 172
  • 173
  • 174
  • More pages
  • 200
  • Next