AIAIAny
  • Search
  • Collection
  • Category
  • Tag
  • Daily AI
AIAIAny

Discover the Best AI Resources

Curated essentials, no noise — just what matters

AIAIAny

Curated AI Resources for Everyone

[email protected]

Powered by airss.app

Product
  • Search
  • Collection
  • Category
  • Tag
Resources
  • Blog
Company
  • Privacy Policy
  • Terms of Service
  • Sitemap
Copyright © 2026 All Rights Reserved.

Contents

Hugging Face
AI Model·2026
Icon for item

unsloth/North-Mini-Code-1.0-GGUF

unsloth, Cohere Labs

Provides GGUF quantized weights and runnable instructions to run CohereLabs' North-Mini-Code-1.0 (30B A3B MoE) locally via llama.cpp or vLLM; includes quant files, build/run notes, and recommended sampling and tool-use settings for agentic coding.

#transformers#vllm#llm#code#ai-coding+5
AI Video Papers·2026
Icon for item

JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence

Dingyu Yao, Junhao Zhou +13·Joy Future Academy, JD

Continuously watches live video and autonomously decides each second whether to speak, stay silent, or delegate; released together with an 8B vision-first model, time-aligned interaction data, training recipe, and a deployable real-time system. Designed for vision-triggered, low-latency streaming scenarios and evaluated across six real-world streams.

#video#vision#multimodal#vllm#ai-agent+3
Hugging Face
AI Model·2026
Icon for item

unsloth/diffusiongemma-26B-A4B-it-GGUF

unsloth

A community-distributed GGUF bundle of Google DeepMind’s DiffusionGemma (26B A4B) with multiple quantization variants for local image-text-to-text inference. Targets experimentation and offline deployment via the DiffusionGemma llama.cpp branch and llama-diffusion-cli; choose quantization for GPU memory vs. fidelity trade-offs.

#gemma#google#huggingface#llm#multimodal+2
AI Agent Papers·2026
Icon for item

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application

Jiachun Li, Zhuoran Jin +13

Survey of methods for engineering interactive environments for LLM-based agents, covering environment modeling, symbolic and neural synthesis, evaluation, and agent–environment co-evolution. Identifies evolution paradigms and future directions like Environment-as-a-Service and multi-agent systems.

#llm#LLM#NLP#agent-skills#ai-agent+1
Machine Learning Foundation Papers·2026
Icon for item

Redesign Mixture-of-Experts Routers with Manifold Power Iteration

Songhao Wu, Ang Lv +2

Proposes a router redesign for Mixture-of-Experts (MoE) that aligns each router row with its expert's principal singular direction using Manifold Power Iteration (MPI), improving token–expert affinity. MPI applies a 'power‑then‑retract' step to push router rows toward principal singular vectors while enforcing norm constraints; the paper gives convergence theory and pretraining results on 1B–11B MoE models.

#paper#llm#transformers#foundation-model#nlp
AI Agent Papers·2026
Icon for item

Toward Generalist Autonomous Research via Hypothesis-Tree Refinement

Jiajie Jin, Yuyang Hu +16

Lets an AI agent propose, run, and evaluate multi-step research experiments using a persistent Hypothesis Tree that links hypotheses, artifacts, evidence, and distilled insights. Combines a long-lived coordinator with short-lived executors to carry lessons across time; evaluated on six ML tasks.

#paper#ai-agent#agent-skills#ai-workflow#ai-train+2
AI Agent Papers·2026
Icon for item

FORT-Searcher: Synthesizing Shortcut-Resistant Search Tasks for Training Deep Search Agents

Jia Deng, Yimeng Chen +10

Synthesizes shortcut-resistant search tasks to train deep search agents by controlling four shortcut risks across entity selection, evidence-graph construction, question formulation, and adversarial refinement. Produces training trajectories with longer pre-answer search and fewer shortcut patterns; code will be released on GitHub.

#paper#github#ai-agent#agent-skills#deepseek+2
Hugging Face
AI Audio·2026
Icon for item

ZONOS2

Gabriel Clark, Sofian Mejjoute +3·Zyphra

Multilingual, low-latency text-to-speech model for speech generation and zero-shot voice cloning. Uses an MoE backbone with ECAPA-TDNN speaker embeddings, supports audio prefixes, fine-grained prosody/emotion controls and 44.1kHz output; optimized for Linux + NVIDIA GPUs.

#tts#audio#multilingual#voice#huggingface+4
AI Video Papers·2026
Icon for item

OmniDirector: General Multi-Shot Camera Cloning without Cross-Paired Data

Jiwen Liu, Shujuan Li +9

Encodes and clones camera motion from reference videos to generate multi-shot videos — uses a visual "camera grid" to represent camera parameters, trains on million-scale grid–video pairs, and employs a hierarchical prompt-expansion agent to coordinate camera, subject, and action control for multimodal diffusion models.

#video#multimodal#ai-video#vision#prompt-engineering+2
AI Video Papers·2026
Icon for item

Rethinking RAG in Long Videos: What to Retrieve and How to Use It?

Yuho Lee, Jisu Shin +6·KAIST, Qualcomm AI Research (Qualcomm Korea)

Proposes chunk-level multimodal retrieval and chunk-adaptive reranking for retrieval-augmented generation on long egocentric videos; introduces V-RAGBench to decouple retrieval vs. generation evaluation and CARVE to run parallel retrievers and select per-chunk configurations.

#RAG#video#multimodal#evaluation#vision+2
AI Agent Papers·2026
Icon for item

EvoArena: Tracking Memory Evolution for Robust LLM Agents in Dynamic Environments

Jundong Xu, Qingchuan Li +12

Benchmarks evolving environments as sequences of progressive updates and introduces EvoMem, a patch-based memory that records structured update histories so LLM agents can reason about environment evolution. Demonstrates measurable gains on EvoArena and other benchmarks.

#LLM#ai-agent#agent-skills#paper#nlp
Computer Vision Papers·2026
Icon for item

SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning

Seokju Cho, Ryo Hachiuma +9

Provides a training-free, code-as-action framework that lets VLM-backed agents write and run stateful Python cells to compose perception and geometry primitives for open-ended 3D/4D spatial reasoning. Demonstrates consistent gains across 20 benchmarks and multiple VLM backbones.

#vision#multimodal#ai-agent#agent-skills#paper
  • Previous
  • 1
  • More pages
  • 137
  • 138
  • 139
  • More pages
  • 178
  • Next