AIAIAny
  • Search
  • Collection
  • Category
  • Tag
  • Daily AI
AIAIAny

Discover the Best AI Resources

Curated essentials, no noise — just what matters

AIAIAny

Curated AI Resources for Everyone

[email protected]

Powered by airss.app

Product
  • Search
  • Collection
  • Category
  • Tag
Resources
  • Blog
Company
  • Privacy Policy
  • Terms of Service
  • Sitemap
Copyright © 2026 All Rights Reserved.

Contents

Speech Technology Papers·2026
Icon for item

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks

Yu Zhang, Ruiqi Li +4

Generates multi‑speaker speech and environmental audio from textual instructions or a reference clip, supporting zero‑shot voice cloning and detailed scene/specification control. Combines a cleaned, captioned dataset with a VAE-based multimodal generator, reward-conditioned quality control, and staged training to improve expressiveness and multi-audio modeling.

#audio#speech#tts#multimodal#voice+2
Large Language Model Papers·2026
Icon for item

SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs

Kejian Zhu, Zhuoran Jin +6

Analyzes why supervised fine-tuning (SFT) causes severe task conflicts under multi-stage multi-task training while reinforcement learning (RL) enables stable coexistence, attributing the effect to sparse, near-orthogonal RL parameter updates and proposing Parallel-RL to decouple multi-task training.

#RL#llm#NLP#paper#reasoning+2
AI Video Papers·2026
Icon for item

JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion

Yicheng Xiao, Wenxun Dai +23·Joy Future Academy, JD

Performs real-time, instruction-guided video-to-video editing on streaming input using a 16B autoregressive diffusion model that preserves subject identity and long-term temporal coherence; achieves end-to-end 720p at ≈30 FPS on a single Nvidia B200 GPU. Key features include chunk-wise autoregressive adaptation, Source-Anchored Distribution Matching Distillation (SA-DMD) that reduces diffusion to a two-step generator, and Long-Horizon Autoregressive Distillation to mitigate temporal drift.

#video#ai-video#distillation#multimodal#vision+5
Hugging Face
AI Model·2026
Icon for item

MiniMax-H3-TAE

Kijai

A quickly trained 2D "tine" VAE for MiniMax‑H3 that speeds up preview renders of video outputs and typically outperforms latent2rgb for preview use. Currently only compatible with the ModelPreviewOverride node in ComfyUI‑KJNodes and intended for previewing rather than production-grade decoding.

#huggingface#video#ai-video#multimodal#ai-image
Hugging Face
AI Video·2026
Icon for item

TenStrip/10Eros-Max

TenStrip·TenStrip, MiniMaxAI

Experimental MiniMax H3 variant that injects learned stylistic and motion 'character' from LTX 2.3, Wan 2.2 and Krea 2 into H3 by surgically grafting attention and MLP components; preserves H3 modality routing while shifting t2v/i2v aesthetics, with limited audio impact and community-license constraints.

#video#ai-video#multimodal#huggingface#gemma+2
AI Agent Papers·2026
Icon for item

Beyond Simply Environment Scaling: Designing Effective Environment Distributions for Multimodal Agent Learning

Kejian Zhu, Zhuoran Jin +5

Analyzes how to build effective training environment distributions for multimodal agents and proposes Ability-aware Environment Selection (AES) and Hierarchical Difficulty Curriculum (HDC) to improve diversity and difficulty scheduling, yielding large relative gains in experiments.

#multimodal#agent-skills#RL#paper#ai-train+2
Hugging Face
AI Model·2026
Icon for item

Maple-Preview

DeepGrove

A 20B ternary-weight Mixture-of-Experts reasoning LLM optimized for on-device and low-memory inference—delivers high throughput (200+ tok/s on M4) and an extremely long 131k-context for math/logic benchmarks, but is a preview with limited agentic fine-tuning.

#transformers#llm#reasoning#huggingface#benchmarks+1
AI Agent Papers·2026
Icon for item

OneDayAgent: Towards a Long-Horizon Harness for Autonomous Agents

Jingsheng Zheng, Xinyuan Fang +4

Turns open-ended everyday requests into a managed long-horizon execution process that decomposes tasks into bounded subtasks, maintains compact execution memory under context pressure, and verifies and repairs final deliverables. Designed to run unchanged across multiple LLM backends and evaluated on AgentIF-OneDay.

#long-horizon#ai-agent#LLM#agent-skills#benchmarks+1
Large Language Model Papers·2026
Icon for item

When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation

Yinuo Jiang, Yongjie Ye +5

Detects and filters spurious token-level teacher supervision in on-policy distillation by estimating input-groundedness and removing high-impact misleading updates, improving OPD on both LLM and VLM benchmarks.

#distillation#LLM#NLP#vision#multimodal+2
Hugging Face
AI Model·2026
Icon for item

Qwen3-VL-32B Heretic (MiniMax-H3 text encoder) — NVFP4

sakamakismile·Lna-Lab, ethanfel +3

An uncensored NVFP4-quantized text encoder for MiniMax-H3 video generation that fits on a single 16 GB GPU. Mixed-precision bake (mostly NVFP4, embedding left as INT8), preserves ConvRot rotation semantics, and includes the unrotate step required to avoid corrupted conditioning.

#qwen#huggingface#ai-video#video#pytorch+2
Hugging Face
AI Video·2026
Icon for item

SexGod1979/PinkCherry_MiniMax-H3

SexGod1979

Generates short videos with stereo audio from text prompts using a MiniMax‑H3 checkpoint; community‑uploaded on Hugging Face and distributed under Apache‑2.0. Tuned toward stylized creature and floral visuals and updated frequently per the model card.

#transformers#huggingface#ai-video#video#audio+1
AI Agent Papers·2026
Icon for item

WorldClaw: Agentic 3D Open-World Generation at Scale

Chunchao Guo, Jinpeng Li +2

Generates large-scale, explorable 3D open-world scenes from open-ended text prompts, producing editable instance-level assets and a consistent global terrain. Uses agentic planning to convert text into region/terrain/asset specifications and a coarse-to-fine pipeline for terrain construction, mesh reconstruction, and render-based refinement.

#vision#multimodal#agent-skills#ai-agent#ai-image+2
  • Previous
  • 1
  • More pages
  • 171
  • 172
  • 173
  • More pages
  • 199
  • Next