AIAIAny
  • Search
  • Collection
  • Category
  • Tag
  • Daily AI
AIAIAny

Discover the Best AI Resources

Curated essentials, no noise — just what matters

AIAIAny

Curated AI Resources for Everyone

[email protected]

Powered by airss.app

Product
  • Search
  • Collection
  • Category
  • Tag
Resources
  • Blog
Company
  • Privacy Policy
  • Terms of Service
  • Sitemap
Copyright © 2026 All Rights Reserved.

Contents

AI Video Papers·2026
Icon for item

Latent Spatial Memory for Video World Models

Weijie Wang, Haoyu Zhao +8

Stores a persistent 3D scene cache directly in a diffusion model's latent space to produce temporally and spatially consistent videos. Constructs memory via depth-guided back-projection and queries it with direct latent-space warping — achieving large speed and memory gains versus pixel-space 3D baselines.

#video#vision#depth#ai-video#paper+1
Computer Vision Papers·2026
Icon for item

ABot-Earth 0.5: Generative 3D Earth Model

Ming Qian, Tianjian Ouyang +26

Synthesizes scalable, photoreal 3D Earth tiles from georeferenced satellite imagery using a generative 3D Gaussian Splatting representation; trained on urban reconstructions, it generates novel scenes at under 10 minutes/km² with hierarchical LOD for real-time web map visualization and Embodied AI use cases.

#vision#paper#ai-image#robotics#depth
AI Agent Papers·2026
Icon for item

Data Journalist Agent: Transforming Data into Verifiable Multimodal Stories

Kevin Qinghong Lin, Batu EI +4·University of Oxford, Stanford University

Turns raw datasets into verifiable multimodal news features via a multi-agent newsroom pipeline. Key innovations: (1) an Inspector that links each claim to data/code/external references for re-execution and audit; (2) multimodal asset generation (interactive maps, audio, visuals) tailored to the story.

#agent-skills#multimodal#ai-agent#paper#code+3
AI Agent Papers·2026
Icon for item

Role-Agent: Bootstrapping LLM Agents via Dual-Role Evolution

Xucong Wang, Ziyu Ma +5

Lets a single LLM simultaneously act as agent and environment to bootstrap co-evolutional training — using state-prediction process rewards (World-In-Agent) and failure-mode retrieval (Agent-In-World) to reshape training data; reports ~4% average benchmark gain.

#LLM#ai-agent#agent-skills#paper#RL+2
Computer Vision Papers·2026
Icon for item

SCAIL-2: Unifying Controlled Character Animation with End-to-end In-Context Conditioning

Wenhao Yan, Fengjia Guo +2

End-to-end framework for controlled character animation that transfers motion from driving videos to reference characters without intermediate pose or background representations. Introduces the MotionPair‑60K end-to-end motion-transfer dataset, in‑context mask conditioning and mode‑specific RoPE for task unification, plus Bias‑Aware DPO to mitigate synthetic-detail errors.

#paper#vision#video#multimodal#ai-video+1
Hugging Face
AI Dataset·2026
Icon for item

HIW-500: Humanoids In-the-Wild Dataset

BitRobot, Unitree +1

Provides 500+ hours of human whole-body teleoperation demonstrations for humanoid robot learning in real homes, with synchronized video, joint states, action traces and language annotations. Includes 23K+ episodes, fine-grained subtask labels, and raw ROS/MCAP plus compressed LeRobot formats.

#robotics#multimodal#vision#huggingface#ai-train+1
Hugging Face
AI Model·2026
Icon for item

DiffusionGemma 26B A4B

Google DeepMind

Generates text from interleaved text, image, and short-video inputs using discrete diffusion and block‑autoregressive multi‑canvas sampling; built on a sparse MoE (8/128) Gemma 4 backbone and optimized for low‑latency inference and very long contexts (up to 256K tokens).

#gemma#foundation-model#multimodal#vision#transformers+5
Hugging Face
AI Dataset·2026
Icon for item

HIW-500: Humanoids In-the-Wild Dataset (LeRobot)

BitRobot, Unitree +1

Provides 500+ hours of human whole-body teleoperation recordings of a Unitree G1 in real homes, packaged in LeRobot v3.0 for robot learning. Contains 23K+ episodes, ~40M frames, multi-view 480p@30 video, 29-DoF states, actions and language annotations; CC BY 4.0 and large download size.

#robotics#video#huggingface#polars#pytorch
Hugging Face
AI Video·2026
Icon for item

SCAIL-2

zai-org

End-to-end pose-driven image-to-video model that animates a reference character from a driving video, supporting cross-identity replacement and multi-character scenarios without intermediate pose representations; performs best at 704p and ships as a diffusers-compatible checkpoint.

#diffusers#video#ai-video#huggingface#image
Hugging Face
AI Dataset·2026
Icon for item

Russian PII NER Benchmark

redmadrobot-rnd

Provides a token-level benchmark for Russian PII detection and NER, with 2,841 sentences and 5,614 annotated spans across 21 fine-grained entity types in BIO format. Mixes sanitized production-log examples, synthetic document templates, and hard negatives to evaluate guardrails and anonymization pipelines.

#huggingface#pandas#nlp#python#security
Hugging Face
AI Model·2026
Icon for item

RazzzHF/Realism_Engine_Ideogram_4

RazzzHF

Fine-tuned Hugging Face image-generation model that biases Ideogram-style prompts toward photorealistic outputs. Emphasizes natural lighting and realistic materials to reduce prompt tweaking; license not specified.

#huggingface#ai-image#image#foundation-model#multimodal+2
Hugging Face
AI Coding·2026
Icon for item

Gemma4-12B-Coder (Composer 2.5 × Fable 5) — GGUF

yuxinlu1

Provides a locally runnable, quantized GGUF release of Gemma 4 12B fine-tuned for Python coding with chain-of-thought distilled from Composer 2.5 and supplemented by Fable 5. Multiple quant options for low‑VRAM setups and execution‑verified training traces. Not safety‑aligned; validate before production.

#gemma#ai-coding#llm#python#code+1
  • Previous
  • 1
  • More pages
  • 135
  • 136
  • 137
  • More pages
  • 175
  • Next