AIAIAny
  • Search
  • Collection
  • Category
  • Tag
  • Daily AI
AIAIAny

Discover the Best AI Resources

Curated essentials, no noise — just what matters

AIAIAny

Curated AI Resources for Everyone

[email protected]

Powered by airss.app

Product
  • Search
  • Collection
  • Category
  • Tag
Resources
  • Blog
Company
  • Privacy Policy
  • Terms of Service
  • Sitemap
Copyright © 2026 All Rights Reserved.

Contents

Hugging Face
AI Model·2026
Icon for item

Cosmos3-Super-Text2Image

NVIDIA

Generates high-fidelity images from text prompts using NVIDIA's 64B Cosmos3-Super multimodal foundation model. Integrates with Hugging Face Diffusers and vLLM‑Omni, is released under OpenMDW1.1 for commercial use, and is optimized for Physical AI workflows (robotics, AV, simulation).

#nvidia#huggingface#diffusers#vllm#ai-image+5
Hugging Face
AI Dataset·2026
Icon for item

SmoothConv

ASLP@NPU, QualiaLabs

Provides ~100 hours of expert-annotated, multi-channel Chinese conversational speech with per-segment timestamps, speaker IDs and paralinguistic labels for turn-taking, overlap/interruption detection and full‑duplex dialogue research. Licensed for academic/non-commercial use (CC BY‑NC 4.0).

#speech#ASR#huggingface#voice#nlp+1
Hugging Face
AI Model·2026
Icon for item

LFM2.5-8B-A1B

Liquid AI

Hybrid LFM2.5 text-generation model optimized for on-device assistants and agentic workflows — 8.3B total / 1.5B active parameters with 131,072-token context. Prioritizes low-latency, high-throughput inference and multilingual instruction-following; not optimized for pure heavy programming or knowledge-heavy QA without retrieval.

#llm#transformers#huggingface#multilingual#vllm+5
Computer Vision Papers·2026
Icon for item

Why Far Looks Up: Probing Spatial Representation in Vision-Language Models

Cheolhong Min, Jaeyun Jung +6

Analyzes spatial representations in vision–language models and reveals a consistent vertical-position ↔ distance entanglement; introduces SpatialTunnel, a synthetic benchmark that exposes this perspective-driven shortcut, and provides code and a project page.

#vision#multimodal#paper#code#depth+1
AI Video Papers·2026
Icon for item

EarlyTom: Early Token Compression Completes Fast Video Understanding

Hesong Wang, Xin Jin +5

Performs training-free early-stage visual token compression inside the vision encoder to cut time-to-first-token (TTFT) and FLOPs for Video-LLMs. Introduces a decoupled spatial token selection strategy and reports up to 2.65× TTFT reduction and 61% FLOPs savings on LLaVA-OneVision-7B (NVIDIA A100) while preserving full-token accuracy — aimed at latency-sensitive video understanding.

#video#vision#ai-video#multimodal#llm+3
Hugging Face
AI Model·2026
Icon for item

Step-3.7-Flash (GGUF quantizations)

stepfun-ai

GGUF quantizations of Step-3.7-Flash: a sparse multimodal Mixture-of-Experts LLM with native image understanding, selectable reasoning levels, and a 256K context window. Ships multiple calibrated Q3/Q4/IQ quant files plus an mmproj vision projector for local llama.cpp inference on high-memory hosts.

#huggingface#llm#vision#multilingual#ai-inference+4
Hugging Face
AI Dataset·2026
Icon for item

Nemotron-Pretraining-Code-v3

NVIDIA Corporation

Metadata-only corpus of 146.3M new GitHub source-code files (commit_id, rel_path, language) intended as an incremental update to Nemotron v1/v2 for LLM code pretraining; CC-BY-4.0 licensed and designed to be used jointly with older versions.

#nvidia#huggingface#code#github#llm+4
Hugging Face
AI Dataset·2026
Icon for item

ResearchMath-14k

amphora

A collection of 14,056 self-contained research-level mathematical problems extracted from papers and open-problem lists, each rewritten with taxonomy labels and open-status metadata for training or evaluating models on research-grade math reasoning.

#paper#huggingface#pandas#python#science+1
Natural Language Processing Papers·2026
Icon for item

UniSteer: Text-Guided Flow Matching in Activation Space for Versatile LLM Steering

Yingdong Shi, Ruiming Zhang +5

Learns a text-conditioned flow (a conditional velocity field) in LLM residual activations to steer frozen models at inference by partially transporting and regenerating activations under target textual conditions — enabling unified control over persona, style, truthfulness, compositional constraints, and activation-space classification.

#LLM#NLP#paper#transformers#foundation-model
AI Video Papers·2026
Icon for item

SANA-Streaming: Real-time Streaming Video Editing with Hybrid Diffusion Transformer

Yuyang Zhao, Yicheng Pan +7

Enables real-time streaming video-to-video editing (1280×704 @24 FPS) on a single RTX 5090 GPU. Uses a Hybrid Diffusion Transformer for balanced local/global modeling, Cycle‑Reverse Regularization for temporal consistency, and system-level mixed-precision and fused kernels to maximize throughput.

#video#ai-video#vision#transformers#nvidia+2
Hugging Face
AI Dataset·2026
Icon for item

ClawHub Security Signals

OpenClaw

Provides a sanitized, MIT‑licensed dataset of scanner evidence and registry verdicts for public ClawHub agent skills — 67k+ latest skill versions with redacted artifacts and structured VirusTotal, static-analysis, and SkillSpector outputs to study scanner disagreement and agent-skill risk governance.

#security#agent-skills#LLM#huggingface#pandas+2
Large Language Model Papers·2026
Icon for item

Draft-OPD: On-Policy Distillation for Speculative Draft Models

Haodi Lei, Yafu Li +9

Introduces Draft-OPD, an on-policy distillation method for training lightweight draft models used in speculative decoding — it focuses learning on draft-induced errors via target-assisted rollouts and replay, improving acceptance length and enabling >5× lossless LLM inference acceleration.

#paper#NLP#llm#ai-inference#ai-serving+2
  • Previous
  • 1
  • More pages
  • 127
  • 128
  • 129
  • More pages
  • 172
  • Next