AIAIAny
  • Search
  • Collection
  • Category
  • Tag
  • Daily AI
AIAIAny

Category

Explore by categories

AIAIAny

Curated AI Resources for Everyone

[email protected]

Powered by airss.app

Product
  • Search
  • Collection
  • Category
  • Tag
Resources
  • Blog
Company
  • Privacy Policy
  • Terms of Service
  • Sitemap
Copyright © 2026 All Rights Reserved.
  • All Categories

  • AI Leaderboard

  • AI Agent Tutorials

  • AI Coding Tutorials

  • AI Model

  • AI Agent Papers

  • Chatbot

  • AI Dataset

  • Machine Learning Foundation Books

  • AI Train

  • AI Deploy

  • AI Client

  • Machine Learning Foundation Papers

  • Machine Learning Foundation Tutorials

  • AI Image Demos

  • AI Agent

  • Large Language Model Tutorials

  • Large Language Model Papers

  • Machine Learning Engineering Papers

  • Computer Vision Tutorials

  • Computer Vision Papers

  • Natural Language Processing Papers

  • Reinforcement Learning Papers

  • Speech Technology Papers

  • AI API

  • AI Coding

  • AI Image

  • AI Video

  • MLOps

  • MCP Client

  • MCP Server

  • AI Video Papers

  • AI Audio

  • AI Others

  • AI Infra

  • Embodied AI

Hugging Face
AI Audio·2026
Icon for item

MOSS-Transcribe-Diarize

OpenMOSS-Team, MOSI.AI +1

Converts long-form multi-speaker audio/video into a compact, speaker-aware transcript with timestamps and anonymous speaker labels in one pass. Combines ASR and diarization in a single model, supports custom prompts/hotwords, and targets meetings, podcasts, interviews and long recordings.

#ASR#audio#speech#stt#transformers+5
GitHub
AI Train·2026
Icon for item

Cosmos-Framework

NVIDIA

End-to-end Python framework for training and serving NVIDIA's Cosmos world models (Cosmos3), integrating distributed training (FSDP/TP/CP/PP), DCP/safetensors checkpoints, dataset adapters, multiple inference backends, online serving, and agent skills.

#nvidia#ai-train#ai-serving#pytorch#cuda+8
Hugging Face
AI Model·2026
Icon for item

Qwopus3.6-27B-v2-MTP

Jack Rong (Jackrong)

Fine-tuned reasoning model that speeds up structured multi-step outputs using Multi-Token Prediction (MTP) from a Qwen3.6-27B base. Produces more concise, faster generations for coding, DevOps, math, and constrained-format tasks; experimental community release for research and evaluation.

#huggingface#transformers#llm#ai-train#ai-inference+5
Hugging Face
AI Model·2026
Icon for item

Fara1.5-27B

Microsoft Research AI Frontiers, Microsoft

Automates end-to-end web workflows from browser screenshots by emitting pixel-grounded actions (click, type, scroll, visit, search). Vision-first multimodal agent fine-tuned from Qwen3.5-27B with critical-point safety checks; intended for sandboxed, human-supervised deployments.

#qwen#multimodal#vision#agent-skills#microsoft+6
Hugging Face
AI Model·2026
Icon for item

Miso TTS 8B

MisoLabs

Generates conversational speech and voice continuation from text and optional audio context, outputting Mimi audio codes. Built on a Sesame-style CSM with an 8B Llama-like backbone plus a smaller autoregressive audio decoder. Suited for local TTS inference and voice-cloning workflows.

#pytorch#audio#voice#speech#huggingface+1
Hugging Face
AI Model·2026
Icon for item

NVIDIA Cosmos3-Super-Image2Video

NVIDIA

Generates temporally coherent MP4 videos from a single input image plus text instructions, with configurable resolution, frame count, and optional AAC audio. Optimized for NVIDIA GPU stacks and integrates with vLLM‑Omni and Hugging Face Diffusers for production inference and research workflows.

#nvidia#huggingface#diffusers#ai-video#video+5
Hugging Face
AI Video·2026
Icon for item

LongCat-Video-Avatar-1.5

Meituan LongCat Team

Generates audio-driven avatar videos from text, images, or audio inputs with production-grade stability (accurate lip sync, identity consistency) and an 8-step distillation inference mode for faster serving; suitable for broadcasting, virtual hosts, animation, and multi-person scenarios.

#ai-video#video#audio#transformers#huggingface+4
Hugging Face
AI Model·2026
Icon for item

MiniCPM5-1B

openbmb

A 1.08B-parameter causal LLM engineered for on-device text generation with native long-context (131k tokens) and built-in Think/No-Think modes. It emphasizes tool-calling support, lightweight deployment formats (BF16, GGUF, MLX), and RL+OPD post-training for stronger reasoning and code generation.

#llm#transformers#huggingface#vllm#ollama+3
Hugging Face
AI Model·2026
Icon for item

Bonsai Image · Ternary 4B (gemlite 2-bit)

Prism ML (prism-ml)

A ternary-weight (~1.58-bit) 4B text-to-image diffusion transformer optimized for NVIDIA GPUs using Gemlite INT2 and HQQ; it reduces the transformer to ~1.21 GB (4.55 GB CUDA payload) and targets 1024×1024 generation with a 4-step FlowMatch-Euler sampler.

#huggingface#ai-image#image#nvidia#ai-inference+3
Hugging Face
AI Model·2026
Icon for item

google/gemma-4-12B-it

Google DeepMind

Instruction-tuned, unified Gemma 4 12B multimodal model that accepts text, image and audio inputs and generates text outputs locally. Encoder-free design reduces multimodal latency and fits on consumer devices while offering long-context support and native thinking/system-prompt features.

#gemma#google#deepmind#multimodal#transformers+5
Hugging Face
AI Model·2026
Icon for item

Gemma 4 12B Unified

Google DeepMind

A 12B unified, encoder-free multimodal model that directly ingests text, images and audio and returns text; supports very long contexts (up to 256K tokens), native function-calling/thinking modes, and small-model deployment for local or on-device use.

#gemma#multimodal#transformers#google#deepmind+8
Hugging Face
AI Model·2026
Icon for item

Step 3.7 Flash

stepfun-ai

Processes images and text to produce structured, reasoning-rich text outputs for high-throughput agentic workflows. Sparse MoE design (198B total, ~11B active per token), 256k context window and selectable reasoning levels—optimized for single-pass parsing, verification, and multi-step automation.

#multimodal#llm#transformers#vllm#ai-inference+4
  • Previous
  • 1
  • More pages
  • 12
  • 13
  • 14
  • More pages
  • 31
  • Next