AIAIAny
  • Search
  • Collection
  • Category
  • Tag
  • Daily AI
AIAIAny

Discover the Best AI Resources

Curated essentials, no noise — just what matters

AIAIAny

Curated AI Resources for Everyone

[email protected]

Powered by airss.app

Product
  • Search
  • Collection
  • Category
  • Tag
Resources
  • Blog
Company
  • Privacy Policy
  • Terms of Service
  • Sitemap
Copyright © 2026 All Rights Reserved.

Contents

Hugging Face
AI Dataset·2026
Icon for item

HiFi-UMI-2K

Yuteng Wei, Jinming Ma +15·Simple AI

Provides 2,000 hours of synchronized, high‑fidelity robot‑free bimanual manipulation demonstrations with multi‑view video, calibrated end‑effector trajectories, gripper states, and language annotations. Curated from a 20,000+ hour corpus; features 6 camera views, ~3 mm pose accuracy, <40 µs cross‑sensor sync, and LeRobot v3‑style Parquet+MP4 export under CC BY 4.0.

#robotics#video#multimodal#parquet#huggingface+3
AI Agent Papers·2026
Icon for item

SWE-Pruner Pro: The Coder LLM Already Knows What to Prune

Yuhang Wang, Yuling Shi +7

Prunes tool-output lines inside a coding LLM agent by turning the agent's own internal representations into per-line keep-or-prune labels. Implements a small classification head plus a length-aware embedding, saving up to 39% of tokens across benchmarks while preserving task quality.

#swe#ai-agent#ai-coding#LLM#code+3
AI Video Papers·2026
Icon for item

HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement

Yiyang Cai, Nan Chen +9

Personalizes subject-driven videos to preserve human identity and accurate human–object interactions by integrating multimodal references and MLLM-derived semantics. Introduces global multimodal guidance in self-attention and modality-reference embeddings to align MLLM features with VAE tokens, supporting both inter- and intra-subject inputs (e.g., OCR, multi-view).

#video#multimodal#LLM#ocr#paper+2
AI Video Papers·2026
Icon for item

AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report

AlayaWorld Team, Kaipeng Zhang +16

Generates interactive long-horizon 24-fps video worlds (540p/720p) from text, image, or video inputs. Uses a 15B video diffusion transformer with a bounded visual context (sink frame, compressed temporal history, geometry-aligned spatial memory, recent-frame conditioning) and a discrete autoregressive distillation that cuts inference to ~4 sampling steps per chunk.

#video#ai-video#multimodal#distillation#foundation-model+2
Hugging Face
AI Model·2026
Icon for item

Motif-3-Beta

Motif Technologies

Open preview checkpoint of a sparse Mixture-of-Experts causal LLM with ~314B total params (~13B active per token) and native 256K context for long-context multilingual text generation. Ships with custom modeling code (trust_remote_code) and a research/non-commercial use license.

#llm#transformers#vllm#foundation-model#multilingual+2
Hugging Face
AI Dataset·2026
Icon for item

forensic-refusal

Hugging Face

Provides under-1K JSON agent-trace records documenting model refusal responses and forensic metadata — useful for evaluating refusal-detection, audit pipelines, and safety analysis; small size limits large-scale statistical studies.

#huggingface#json#agent-skills#evaluation#nlp+1
Hugging Face
AI Model·2026
Icon for item

GLM-5.2-Vision (NVFP4)

Baseten, zai-org (GLM-5.2) +2

Adds vision to GLM-5.2 by attaching a MoonViT encoder and a trained 49.5M-parameter PatchMerger projector to enable image→text multimodal reasoning; text and vision backbones are frozen, uses NVFP4 quantized weights and targets Blackwell B200 GPUs.

#multimodal#vision#llm#huggingface#nvidia+1
Hugging Face
AI Model·2026
Icon for item

Cosmos3-Edge

NVIDIA

Generates and reasons about multimodal physical-world content—text, images, video and action trajectories—conditioned on text, images, video and robot/vehicle action inputs. An edge-sized (4B) Mixture‑of‑Transformers omni-model optimized for single‑GPU inference and Physical AI tasks (image→video, action prediction, robot policy).

#nvidia#foundation-model#multimodal#robotics#video+7
Large Language Model Papers·2026
Icon for item

Scaling Laws for Hypernetwork-Based Knowledge Injection in Large Language Models

Nischay Dhankhar, Dos Baha +1

Studies train-time knowledge injection via hypernetworks that generate fixed LoRA adapters from large fact corpora, empirically characterizing power-law scaling across hypernetwork depth, width, and target model size and reporting improved OOD generalization.

#LLM#NLP#paper#lora#reasoning+3
Computer Vision Papers·2026
Icon for item

Generative World Renderer at the Speed of Play

Guixu Lin, Zheng-Hui Huang +4

Synthesizes RGB frames from structured world states exported by physics engines; it reformulates a heavy generative renderer into a few-step autoregressive streaming model and uses lightweight distilled codecs to reach playable ~30 FPS while preserving G-buffer and prompt control.

#vision#video#ai-video#distillation#physics+2
Computer Vision Papers·2026
Icon for item

Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers

Maohua Li, Qirui Li +11

Analyzes internal computation of text-to-image diffusion transformers and shows structural template tokens act as implicit semantic registers that maintain object identity during denoising. Introduces a causal interpretability framework (attention decomposition + targeted interventions) and a training-free pruning rule that cuts ~20% attention FLOPs for a ~1.4-point GenEval drop.

#vision#transformers#diffusers#ai-image#paper+2
Computer Vision Papers·2026
Icon for item

Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing

Xinjie Zhang, Peng Zhang +22

Efficient 4B-scale image generation and editing model family that pairs a lightweight VAE tokenizer (Mage-VAE) with a native-resolution multimodal diffusion transformer, reducing tokenization cost by an order of magnitude and enabling few-step high-resolution generation and editing.

#foundation-model#flow-matching#distillation#multimodal#ai-image+4
  • Previous
  • 1
  • More pages
  • 161
  • 162
  • 163
  • More pages
  • 195
  • Next