AIAIAny
  • Search
  • Collection
  • Category
  • Tag
  • Daily AI
AIAIAny

Discover the Best AI Resources

Curated essentials, no noise — just what matters

AIAIAny

Curated AI Resources for Everyone

[email protected]

Powered by airss.app

Product
  • Search
  • Collection
  • Category
  • Tag
Resources
  • Blog
Company
  • Privacy Policy
  • Terms of Service
  • Sitemap
Copyright © 2026 All Rights Reserved.

Contents

Computer Vision Papers·2026
Icon for item

Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing

Junliang Ye, Kenkun Liu +13·Tencent

Provides a unified multimodal framework for large-scale 3D understanding, text-to-3D generation, and instruction-guided 3D editing. Trains on an 87M-sample 3D multimodal corpus (25M understanding, 50M generation, 12M editing) and pairs a vision-language model with a diffusion-based 3D synthesizer to preserve structure and enable part-aware edits; suited for researchers building text-driven 3D asset pipelines but requires large compute and data.

#multimodal#vision#ai-image#AIGC#foundation-model+2
Machine Learning Engineering Papers·2026
Icon for item

Knowledge-Geometry Decoupling: Refreshable Pretrained Transfer for Streaming Recommendation

Zixuan Wang, Yuhong Chen +11·[email protected], [email protected] +3

A pretrain-then-transfer method for streaming recommendation that decouples refreshable behavioral knowledge from task-specific geometry to enable continual model refresh without downstream interference; introduces Behavioral Multi-Token Prediction and Anchored Calibration Residual and shows 4–12% offline gains plus live Shopee A/B lifts.

#paper#foundation-model#embeddings#benchmarks#code+5
Hugging Face
AI Model·2026
Icon for item

Qwen3-VL-32B Ultra Uncensored Heretic — MiniMax-H3 ComfyUI INT8 ConvRot

ethanfel

Provides ComfyUI-ready INT8 MiniMax‑H3 checkpoints (conditioning encoder plus optional generation tail) for a Heretic-edited Qwen3‑VL‑32B source; preserves the vision tower in BF16 and uses row-wise ConvRot INT8 quantization to reduce VRAM needs for ~32GB GPUs. Not a full Transformers generation repository.

#huggingface#qwen#llm#multimodal#pytorch+4
Reinforcement Learning Papers·2026
Icon for item

Progressive Agent Skill Generation via Reinforcement Learning

Junhao Shen, Zhanqiu Zhang +2

Frames skill generation as a sequential editing task and introduces a novel rollback reward to train an RL generator (Skill-α) that evaluates each edit by its downstream execution impact, producing skills that improve agent success rates across document-to-skill and experience-to-skill settings.

#agent-skills#RL#rl#llm#evaluation+3
AI Video Papers·2026
Icon for item

WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity

Yuxue Yang, Shuyao Shang +14

Introduces WorldExam, a diagnostic benchmark that evaluates controllable video world models across four levels from visual quality to inherent world reactivity. Covers 1,474 cases across eight tasks and supports camera-, action-, and language-driven paradigms, measuring scene-conditioned reactions beyond explicit instructions.

#video#vision#benchmark#benchmarks#evaluation+2
Hugging Face
AI Model·2026
Icon for item

Qwen3-VL-32B Ultra Uncensored Heretic — H3 ComfyUI INT8 ConvRot

ethanfel

ComfyUI-ready H3 conditioning encoder builds for Qwen3-VL-32B: a BF16 full-precision checkpoint, an INT8 ConvRot quantized checkpoint, and an optional generation tail (layers 50–63). Retains vision tower in BF16 and targets H3 workflows and lower-VRAM systems.

#qwen#huggingface#pytorch#llm#multimodal+2
Large Language Model Papers·2026
Icon for item

AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling

Jiajun Liang, Yucheng Liao +13

A continuous-latent diffusion language model that preserves a high-capacity, decodable text latent and directly models its distribution via a block-causal diffusion transformer and query-based encoder–decoder; achieves top results on OpenWebText and XSum while scaling to 1B parameters.

#paper#LLM#NLP#flow-matching#diffusers+2
AI Agent Papers·2026
Icon for item

LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks

Ziyu Ma, Hailang Huang +6

Reformulates long-horizon agent execution as explicit task-state management: a manager defines bounded subtasks, fresh-context executors run them, and read-only auditors verify outcomes. Shows large performance gains on WeaveBench, Terminal-Bench and OSWorld.

#long-horizon#llm#benchmarks#evaluation#ai-agent+4
Hugging Face
AI Dataset·2026
Icon for item

stablediffusiontutorials/Minimax-H3

stablediffusiontutorials

Installation-oriented dataset that packages ComfyUI-ready files and instructions for running MiniMax H3 locally — includes pruned/INT8/BF16 checkpoints, matching Qwen3-VL text encoders, video/audio VAEs, and official ComfyUI workflow templates for joint audio+video generation.

#huggingface#ai-video#multimodal#audio#diffusers+5
Hugging Face
AI Dataset·2026
Icon for item

Recursive Task Synthesis

Zhongzhi1228

Provides 37,484 validated command-line tasks generated by recursive task synthesis, each paired with searchable metadata and a sanitized, runnable package (instructions, solution, verifier, and optional Dockerfile).

#huggingface#parquet#terminal#RL#long-horizon+2
Reinforcement Learning Papers·2026
Icon for item

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning

Fangxu Yu, Tao Feng +7

Supervises audio reasoning by generating per-sample, audio-grounded rubrics that evolve with model rollouts and serve as reinforcement-learning rewards, improving perception and adaptive multi-step reasoning while avoiding reward saturation.

#audio#RL#reasoning#benchmarks#paper+3
Hugging Face
AI Dataset·2026
Icon for item

Scene2Wave

Mengfan Zheng, Liwen Jing +6·Pengcheng Laboratory

Provides time-aligned simulated urban driving recordings that pair high-rate CSI/CIR with multi-view cameras, LiDAR, radar, IMU and GNSS for perception-to-channel research; contains 100 validated 1-second samples produced with CARLA and Sionna, but is limited in scene diversity and real-world fidelity.

#multimodal#huggingface#vision#robotics
  • Previous
  • 1
  • More pages
  • 170
  • 171
  • 172
  • More pages
  • 198
  • Next