AIAIAny
  • Search
  • Collection
  • Category
  • Tag
  • Daily AI
AIAIAny

Category

Explore by categories

AIAIAny

Curated AI Resources for Everyone

[email protected]

Powered by airss.app

Product
  • Search
  • Collection
  • Category
  • Tag
Resources
  • Blog
Company
  • Privacy Policy
  • Terms of Service
  • Sitemap
Copyright © 2026 All Rights Reserved.
  • All Categories

  • AI Leaderboard

  • AI Agent Tutorials

  • AI Coding Tutorials

  • AI Model

  • AI Agent Papers

  • Chatbot

  • AI Dataset

  • Machine Learning Foundation Books

  • AI Train

  • AI Deploy

  • AI Client

  • Machine Learning Foundation Papers

  • Machine Learning Foundation Tutorials

  • AI Image Demos

  • AI Agent

  • Large Language Model Tutorials

  • Large Language Model Papers

  • Machine Learning Engineering Papers

  • Computer Vision Tutorials

  • Computer Vision Papers

  • Natural Language Processing Papers

  • Reinforcement Learning Papers

  • Speech Technology Papers

  • AI API

  • AI Coding

  • AI Image

  • AI Video

  • MLOps

  • MCP Client

  • MCP Server

  • AI Video Papers

  • AI Audio

  • AI Others

  • AI Infra

  • Embodied AI

GitHub
AI Model·2026
Icon for item

Bonsai Demo

PrismML-Eng, PrismML

Runs the Bonsai family of quantized LLMs locally (including vision-capable 27B): provides scripts and demo UIs to run 1-bit and ternary Bonsai models on macOS (Metal), Linux/Windows (CUDA/Vulkan/ROCm), or CPU, with long context, tool-calling and an optional Open WebUI agent demo.

#llm#vision#multimodal#huggingface#github+5
Hugging Face
AI Model·2026
Icon for item

OmniVoice

k2-fsa, Han Zhu +9

Converts text to natural-sounding speech across 600+ languages in a zero-shot way, with short-reference voice cloning and fine-grained voice-design controls; uses a diffusion language-model-style architecture to balance quality and very low inference latency.

#multilingual#speech#audio#huggingface#github+2
Hugging Face
AI Model·2026
Icon for item

Mistral Medium 3.5 128B

Mistral AI

A dense 128B multimodal model with a 256k context window, configurable reasoning effort, and native function-calling for agentic workflows. Supports text+image input, multilingual output, and is released on Hugging Face under a Modified MIT license with revenue-based exceptions.

#foundation-model#multimodal#LLM#vllm#huggingface+6
Hugging Face
AI Model·2026
Icon for item

VoxCPM2

OpenBMB

Generates 48kHz multilingual speech from text using a tokenizer-free diffusion-autoregressive TTS architecture, supporting natural-language voice design, controllable cloning, and low-latency streaming. Notable for a 2B-parameter backbone and built-in AudioVAE super-resolution (16k→48k).

#huggingface#audio#pytorch#ai-library#ai-tools+1
Hugging Face
AI Model·2026
Icon for item

GLM-5.1

zai-org

Generates and iterates on long‑horizon agentic plans and code — designed to stay productive across many rounds of tool calls and experiments. Emphasizes iterative reasoning, stronger repo/terminal automation and code generation than GLM‑5, and can be served locally for research and autonomous-agent workloads.

#foundation-model#vibe-coding#llm#ai-coding#ai-agent+4
Hugging Face
AI Model·2026
Icon for item

Gemma 4 31B JANG_4M CRACK (v2) — dealignai

dealignai

Multimodal image-text-to-text fork of Gemma 4 (31B) using a 'CRACK v2' abliteration — tuned for conversational vision inputs and thinking-mode support in JANG v2 safetensors format. Recommended to run in vMLX; published by dealignai.

#huggingface#multimodal#vision#llm#ai-inference+3
Hugging Face
AI Model·2026
Icon for item

Lyra 2.0: Explorable Generative 3D Worlds

NVIDIA

Generates persistent, explorable 3D worlds from a single image by synthesizing long-range, geometry-consistent video and reconstructing it into an explicit 3D Gaussian scene. Intended for internal research use under NVIDIA's research license.

#nvidia#huggingface#paper#github#vision+3
Hugging Face
AI Model·2026
Icon for item

Granite-4.1-8B

Granite Team (IBM)

An 8B-parameter, instruction-tuned long-context LLM optimized for instruction following, tool-calling, and multilingual dialogue — supports 131072-token context and common NLP tasks such as summarization, QA, code, and RAG.

#huggingface#transformers#llm#multilingual#nlp+3
Hugging Face
AI Model·2026
Icon for item

Granite-4.1-30B

Granite Team (IBM)

A 30B-parameter, instruction-tuned language model built for long-context text generation, conversational agents, and tool-calling. It combines supervised fine-tuning and RL alignment, supports 131,072-token context, and is optimized for tasks like summarization, code, and RAG.

#huggingface#transformers#llm#foundation-model#multilingual+4
Hugging Face
AI Model·2026
Icon for item

ERNIE-Image

Baidu

An open text-to-image generation model built on an 8B Diffusion Transformer that focuses on layout-sensitive, text-heavy, and instruction-following image synthesis. Notable for accurate text rendering, structured/compositional generation (posters, comics), and ability to run on consumer 24GB GPUs when paired with prompt enhancement.

#vision#ai-image#huggingface#pytorch#prompt-engineering+3
Hugging Face
AI Model·2026
Icon for item

MiniMax-M2.7

MiniMaxAI

Text-generation LLM designed for agentic workflows: supports multi-agent 'Agent Teams', skill stacks and model self-evolution. Ships on Hugging Face with deployment guides (vLLM, Transformers, SGLang) and is positioned for engineering, tool-calling and productivity use cases.

#llm#huggingface#agent-skills#ai-agent#vllm+3
Hugging Face
AI Model·2026
Icon for item

HY-World 2.0

Tencent

Generates and reconstructs navigable, editable 3D worlds from text, single images, multi-view photos, or video; outputs meshes and Gaussian Splatting assets and includes WorldMirror 2.0 for fast multi-view reconstruction. Suited for research and production pipelines that import assets into engines; requires substantial GPU resources.

#vision#ai-image#huggingface#pytorch#ai-demos+3
  • Previous
  • 1
  • More pages
  • 3
  • 4
  • 5
  • More pages
  • 31
  • Next