AIAIAny
  • Search
  • Collection
  • Category
  • Tag
  • Daily AI
AIAIAny

Discover the Best AI Resources

Curated essentials, no noise — just what matters

AIAIAny

Curated AI Resources for Everyone

[email protected]

Powered by airss.app

Product
  • Search
  • Collection
  • Category
  • Tag
Resources
  • Blog
Company
  • Privacy Policy
  • Terms of Service
  • Sitemap
Copyright © 2026 All Rights Reserved.

Contents

Computer Vision Papers·2026
Icon for item

Colored Noise Diffusion Sampling

Hadar Davidson, Noam Issachar +1

Reallocates injected noise energy across frequency bands to match a diffusion model's spectral bias, improving sampling fidelity without retraining. Uses a timestep- and frequency-dependent colored-noise schedule as a plug-and-play inference-time SDE solver; shows sizable FID drops on ImageNet-256.

#paper#vision#image#ai-image#ai-inference
AI Agent Papers·2026
Icon for item

Masking Stale Observations Helps Search Agents -- Until It Doesn't: A Regime Map and Its Mechanism

Haoxiang Zhang, Qixin Xu +5

Analyzes when masking stale observations improves long-horizon search agents and why, identifying an asymmetric inverted-U relationship between masking benefit, retriever quality, and model capacity; explains a token-for-turn trade-off and releases evaluation scaffolds and trajectories.

#paper#code#github#nlp#llm+4
Speech Technology Papers·2026
Icon for item

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue

Ruiqi Li, Yu Zhang +4

Zero-shot TTS for expressive long-form monologue and multi-speaker dialogue, designed to preserve acoustic consistency, conversational coherence, and affective continuity. Trained on SwanData-Speech and using a 25 Hz VAE, pause-aware text conditioning, and a flow-matching DiT with DiffusionNFT fine-tuning.

#paper#speech#audio#voice#foundation-model
Computer Vision Papers·2026
Icon for item

GGT-100K: Generative Ground Truth for Generalizable Real-World Image Restoration

Xiangtao Kong, Jixin Zhao +3

Synthesizes high-quality targets for real-world image restoration by using multimodal foundation models (MFMs) to convert real low-quality photos into HQ references. Provides GGT-100K (103,707 LQ–HQ training pairs + 500 test pairs) with multi-stage quality control and demonstrates consistent generalization gains for a range of restoration models, especially for finetuning generative restorers.

#paper#vision#image#multimodal#foundation-model+2
Hugging Face
AI Model·2026
Icon for item

unsloth/gemma-4-12b-it-GGUF

unsloth

A GGUF-quantized, locally runnable build of Gemma 4 12B Unified (image-text-to-text) packaged by unsloth; preserves multimodal (image/audio) input support under an Apache-2.0 license and is compatible with common GGUF runtimes and Unsloth Studio.

#gemma#google#deepmind#huggingface#multimodal+7
Speech Technology Papers·2026
Icon for item

Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer

Ke Lei, Yu Zhang +5

Generates synchronized, streaming spatial audio from panoramic video and text prompts using a causal autoregressive diffusion transformer. Combines Spatial Video-Audio Contrastive (SVAC) alignment and online direct preference optimization (ODPO) to improve spatial perception, plus an automated annotation pipeline and public demos.

#paper#audio#speech#multimodal#transformers+3
AI Agent Papers·2026
Icon for item

COLLEAGUE.SKILL: Automated AI Skill Generation via Expert Knowledge Distillation

Tianyi Zhou, Dongrui Liu +3

Automates distillation of heterogeneous traces from a target person or role into versioned, inspectable skill packages for LLM agents — producing separate capability and bounded-behavior tracks that support natural-language corrections, rollback, and cross-host installation. Ships with an open system and a skills gallery.

#agent-skills#skillkit#LLM#nlp#paper+3
Large Language Model Papers·2026
Icon for item

LongTraceRL: Learning Long-Context Reasoning from Search Agent Trajectories with Rubric Rewards

Nianyi Lin, Jiajie Zhang +2

Uses search-agent reading traces and tiered distractors to train LLMs for long-context, multi-hop reasoning, and introduces a rubric reward that supervises entity-level steps (applied only to correct finals). Improves evidence-grounded reasoning and resists reward hacking across 4B–30B models.

#RL#LLM#NLP#paper#code+1
Hugging Face
AI Dataset·2026
Icon for item

Open Spatial Reasoning (Driving 3D Spatial Reasoning)

Anurag Ganguli, Anshuman Lall +5

Evaluates metric 3D spatial reasoning from single driving images via multiple-choice questions that require reconstructing scene geometry rather than relying on image-layout shortcuts. Each sample pairs a numbered-bbox image with a question, four choices, and the correct answer; images come from PlusAI and the dataset is CC BY 4.0.

#vision#image#huggingface#pandas#multimodal+1
Hugging Face
AI Model·2026
Icon for item

ideogram-ai/ideogram-4-fp8

ideogram-ai

Text-to-image model packaged for Diffusers that uses fp8 quantization to lower memory and speed up inference. Delivered as a safetensors checkpoint on Hugging Face with an Ideogram pipeline; created May 30, 2026 — license unspecified.

#diffusers#huggingface#ai-image#image#AIGC+3
Hugging Face
AI Model·2026
Icon for item

ideogram-4-nf4

ideogram-ai

NF4-quantized text-to-image diffusion model released as safetensors and compatible with the Diffusers Ideogram4Pipeline — optimized for lower-memory local inference and faster deployments while preserving the original model's text-to-image capabilities.

#diffusers#ai-image#image#AIGC#foundation-model
GitHub
AI Agent·2026
Icon for item

LoopX

huangruiteng

Maintains a local, durable control-plane state that preserves objectives, typed todos, gates, evidence logs, quotas, and verifiable handoffs for long-running AI agent work. Designed to coordinate multi-day agent loops across Codex, Claude Code, Cursor or custom runners while keeping human judgment, auditability, and safe fallbacks explicit.

#coding-agents#long-horizon#claude-code#codex#cli+5
  • Previous
  • 1
  • More pages
  • 128
  • 129
  • 130
  • More pages
  • 172
  • Next