AIAIAny
  • Search
  • Collection
  • Category
  • Tag
  • Daily AI
AIAIAny

Discover the Best AI Resources

Curated essentials, no noise — just what matters

AIAIAny

Curated AI Resources for Everyone

[email protected]

Powered by airss.app

Product
  • Search
  • Collection
  • Category
  • Tag
Resources
  • Blog
Company
  • Privacy Policy
  • Terms of Service
  • Sitemap
Copyright © 2026 All Rights Reserved.

Contents

AI Dataset·2026
Icon for item

MobileMem: Learning from a Year of Mobile Experiences

Xinle Deng, Yida Xue +15

Provides a year-scale multimodal benchmark and evaluation framework for on-device long-term memory in personal assistants, built from real mobile user trajectories. Tests memory construction, retrieval, updating, temporal reasoning, and implicit preference inference, and includes a knowledge-grounded synthesis pipeline to form coherent long-horizon trajectories.

#benchmark#long-horizon#multimodal#mobile#evaluation+3
Hugging Face
AI Model·2026
Icon for item

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4

NVIDIA Corporation

Open-weight 30B-parameter Mixture-of-Experts LLM with 3B active params, NVFP4-quantized checkpoint, and speculative-decoding support for long-context (up to 1M tokens) agentic, chat, reasoning and tool-calling workloads optimized for NVIDIA GPUs.

#nvidia#llm#transformers#huggingface#pytorch+9
Computer Vision Papers·2026
Icon for item

AdvFD: Boosting Visual Generation via Adversarial Fr'echet Distance Loss

Mingju Gao, Jingkai Zhou +3

Post-training distribution-level objective that augments static Fréchet-distance losses with an adversarially learned representation and a real-feature whitening step to stabilize min–max optimization and avoid trivial feature amplification; targets one-step image generator post-training.

#vision#paper#research#ai-image#pytorch
Hugging Face
AI Model·2026
Icon for item

S1-mini

Superwhisper

Converts raw ASR transcripts into clean written text: adds punctuation and capitalization, expands spoken numbers/dates/times/currencies/emails, removes fillers and resolves self-corrections. Fine-tuned from Qwen3-0.6B (≈0.6B params), 94.8% token accuracy on a 7,519-case English test set; designed for CPU/edge deployment and deterministic post-processing.

#qwen#transformers#ASR#stt#speech+5
AI Video Papers·2026
Icon for item

AVA-Encoder: Towards Agent-Native Video Representation Learning

Chuyue Li, Jinpeng Yu +8·Affiliation: Qwen Business Unit of Alibaba, Affiliation: ShanghaiTech University +3

Encodes videos into a Film Knowledge Graph and reconstructs them to learn agent-native, editable video representations for agentic reasoning and manipulation. Uses agentic auto-encoding with dual-loop textual-gradient optimization, reports large reconstruction gains, and releases a benchmark and dataset.

#video#ai-video#multimodal#agent-skills#qwen+2
Hugging Face
AI Model·2026
Icon for item

Anima-2.9B (Gazingstars123)

Gazingstars123·Gazingstars123, CircleStone Labs +2

A 2.9B-parameter text-to-image model fine-tuned from CircleStone Labs' Anima for anime and illustration; trained on an additional 1.7M samples with a July 2026 knowledge cutoff. Designed for non-commercial creative image generation and ComfyUI integration; weights released under the CircleStone Labs Non-Commercial (derivative) license.

#huggingface#ai-image#diffusers#lora#ai-train+1
Natural Language Processing Papers·2026
Icon for item

Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill

Zhuoyang Qian, Biao Wu +7·Affiliation: Vast Intelligence Lab University of Technology Sydney, Affiliation: Equal Contribution Corresponding Authorhttps://github.com/Spark-To-Paper-Skills/[email protected]

Generates full publication-format research papers from a short idea by composing 13 coding-assistant skills; it retrieves literature, plans and runs feasible experiments, produces editable vector figures, and enforces deterministic integrity checks so claims are revised to match measured evidence.

#claude-code#agent-skills#ai-workflow#research#paper+3
Computer Vision Papers·2026
Icon for item

StateFlow: Building, Evolving, and Accessing 3D World States for Previsualization

Yuyang Yin, Zixiang Li +12

Builds an editable, persistent 3D world state to drive iterative previsualization for film, games, and design — enabling local edits and recombinations instead of one-shot video regeneration. Uses separate stages for state construction, state evolution, and state access, with render-feedback camera refinement.

#ai-video#video#multimodal#vision#long-horizon+1
AI Agent Papers·2026
Icon for item

AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses

Cheng Qian, Wenting Zhao +7·Salesforce AI Research, University of Illinois Urbana-Champaign

Uses a stronger 'builder' model at inference time to construct executable harnesses that boost weaker target models without parameter updates, mainly by turning unstable reasoning into deterministic code, routing, and strict answer-format enforcement.

#distillation#reasoning#LLM#benchmark#evaluation+4
AI Video Papers·2026
Icon for item

Alaya-EVOKE: From Linear-Scaling Supervision to Endless World

Yuanyang Yin, Gongxuan Wang +4

Externalizes persistent scene state into a camera-indexed world bank and designs a long-horizon teacher whose sparse-attention supervision is distilled into a three-step student, enabling responsive, low-latency interactive long-horizon video generation with bounded denoiser context.

#video#ai-video#long-horizon#distillation#vision+2
Hugging Face
AI Model·2026
Icon for item

Qwen3.8-27B-FP8

Qwen Team

Provides an FP8-post-trained 27B multimodal causal language model with a native vision encoder, large-context support (262,144 native, extensible to 1,000,000), controllable thinking-mode reasoning, and compatibility with common inference engines for deployment.

#qwen#huggingface#safetensors#transformers#vllm+7
AI Agent Papers·2026
Icon for item

Intern-S2-Preview: Scientific Agentic Foundation Model

Lei Bai, Jiaqi Cao +123

Supports multimodal scientific understanding, long-horizon agentic workflows and scientific tool interaction using a unified pipeline of multimodal pretraining, supervised fine-tuning and scalable multi-task reinforcement learning. Distinctive features include time-series modules for signal forecasting and a separate Memory Decoder that enables rapid domain specialization without changing the frozen 397B backbone.

#foundation-model#multimodal#rl#agent-skills#long-horizon+2
  • Previous
  • 1
  • More pages
  • 176
  • 177
  • 178
  • More pages
  • 203
  • Next