AIAIAny
  • Search
  • Collection
  • Category
  • Tag
  • Daily AI
AIAIAny

Discover the Best AI Resources

Curated essentials, no noise — just what matters

AIAIAny

Curated AI Resources for Everyone

[email protected]

Powered by airss.app

Product
  • Search
  • Collection
  • Category
  • Tag
Resources
  • Blog
Company
  • Privacy Policy
  • Terms of Service
  • Sitemap
Copyright © 2026 All Rights Reserved.

Contents

Hugging Face
AI Model·2026
Icon for item

Qwythos-9B-v2-GGUF

Empero AI, Alibaba (Qwen team)

Provides GGUF-quantized builds of the Qwythos-9B-v2 LLM for local runtimes, with multiple quant levels, optional MTP-enabled variants, a 1,048,576-token context window, and an optional BF16 vision projector for multimodal use.

#qwen#multimodal#vision#huggingface#llm+2
Hugging Face
AI Model·2026
Icon for item

Qwythos-9B-v2

Empero AI

A 9B-parameter Qwen3.5-based multimodal model tuned to preserve chain-of-thought reasoning while eliminating repetition loops; restores native multi-token prediction, supports 1,048,576-token context, and targets research/red-team use.

#qwen#llm#transformers#huggingface#multimodal+4
AI Video Papers·2026
Icon for item

Video Generation Models are General-Purpose Vision Learners

Letian Wang, Chuhan Zhang +10·Google DeepMind

Uses large-scale text-to-video generative pretraining to create GenCeption, a feed-forward perception model that performs diverse vision tasks from text instructions—depth, surface normals, camera pose, referring segmentation, and 3D keypoints—often matching or surpassing specialized models while requiring far less task-specific data.

#video#vision#foundation-model#deepmind#ai-video+3
Large Language Model Papers·2026
Icon for item

Scalable Visual Pretraining for Language Intelligence

Yiming Zhang, Zhonghan Zhao +14

Explores unsupervised visual pretraining on visually rich documents to improve language-model intelligence; shows visual-pretrained models outperform text-only counterparts on the same corpora. Key aspects: direct use of images/layouts (no OCR-only pipeline), scalable across backbones and benchmarks.

#vision#multimodal#foundation-model#llm#NLP+3
Computer Vision Papers·2026
Icon for item

4D Human-Scene Reconstruction from Low-Overlap Captures

Minhyuk Hwang, Sangmin Kim +3·Seoul National UniversitySeoulRepublic of Korea

Reconstructs 4D dynamic human scenes from sparse, low-overlap multi-camera captures by decoupling background synthesis and human modeling. Synthesizes hundreds of camera-controlled background views with a video diffusion model, initializes deformable Gaussian humans via cross-view identity and triangulated keypoints, then applies motion-adaptive recursive enhancement to reduce artifacts.

#vision#video#ai-video#ai-image#image+2
Hugging Face
AI Video·2026
Icon for item

Wan-Dancer-14B

Mingyang Huang, Peng Zhang +3

Generates minute-scale, temporally coherent dance videos from full music tracks using a hierarchical two-stage approach: global keyframe planning plus local temporal refinement; suitable when long-range musical structure and rhythmic continuity matter.

#diffusers#ai-video#video#AIGC#multimodal+2
Reinforcement Learning Papers·2026
Icon for item

Beyond Euclidean Clipping: Overcoming Exploration Collapse in LLM RL via Riemannian Isometric Policy Optimization

Zhicheng Cai, Xinyuan Guo +5

Proposes Riemannian Isometric Policy Optimization (RIPO) to fix exploration collapse in PPO-style RL for LLMs by aligning policy updates with the policy manifold's Riemannian geometry, improving exploration–exploitation balance and optimization stability across competition benchmarks.

#RL#LLM#paper#nlp#reasoning+2
Hugging Face
AI Video·2026
Icon for item

LTX-Video 2.3 22B — IC-LoRA: CrossView Prompt v0.9

Cseti

Generates a new camera viewpoint from a reference video: an IC‑LoRA adapter for LTX‑Video 2.3 that re‑renders the same scene from a requested discrete camera angle while preserving subject and content. Trained on synthetic multi‑view data, proof‑of‑concept with limited viewpoint range and best for small, chained angle shifts.

#ai-video#video#lora#huggingface#vision+1
AI Agent Papers·2026
Icon for item

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory

Jiayi Tian, Shiao Liu +25

Provides a deliberative Agent OS layer for robots that handles scene-conditioned planning, context-isolated skill execution, multi-stage verification, persistent multi-modal graph memory, and edge–cloud collaboration. Introduces EmbodiedWorldBench (16 scenes, 200+ tasks) and a failure-driven self-evolution loop; shows improved task success and strong memory benchmark scores.

#robotics#multimodal#agent-skills#evaluation#paper+4
AI Agent Papers·2026
Icon for item

ABot-N1: Toward a General Visual Language Navigation Foundation Model

Ruiyan Gong, Yingnan Guo +38

Unifies high-level visual-language reasoning and low-level control for visual navigation by decoupling cognition and control: a slow vision-language reasoner produces pixel goals with explicit chain-of-thought, and a fast action expert converts those anchors into continuous waypoints for robust urban and indoor navigation.

#vision#multimodal#robotics#foundation-model#agent-skills+4
Hugging Face
AI Model·2026
Icon for item

LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V3-GGUF

LuffyTheFox

A GGUF-format Qwen3.6 35B base model image-text-to-text release repaired via tensor-level SVD/scale correction and packaged with Hermes agent tweaks; multimodal (vision + text), MoE architecture, ready for GGUF runtimes like llama.cpp.

#qwen#llm#multimodal#vision#huggingface+4
Hugging Face
AI Model·2026
Icon for item

LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V5-GGUF

LuffyTheFox·Hugging Face, HauhauCS +1

A GGUF-distributed Qwen3.6 35B MoE model variant repaired with a

#qwen#multimodal#vision#llm#multilingual+5
  • Previous
  • 1
  • More pages
  • 156
  • 157
  • 158
  • More pages
  • 191
  • Next