AIAIAny
  • Search
  • Collection
  • Category
  • Tag
  • Daily AI
AIAIAny

Tag

Explore by tags

AIAIAny

Curated AI Resources for Everyone

[email protected]

Powered by airss.app

Product
  • Search
  • Collection
  • Category
  • Tag
Resources
  • Blog
Company
  • Privacy Policy
  • Terms of Service
  • Sitemap
Copyright © 2026 All Rights Reserved.
  • All

  • 30u30

  • ASR

  • ChatGPT

  • GNN

  • IDE

  • RAG

  • agent-skills

  • ai

  • ai-agent

  • ai-api

  • ai-api-management

  • ai-client

  • ai-coding

  • ai-demos

  • ai-deploy

  • ai-development

  • ai-framework

  • ai-image

  • ai-image-demos

  • ai-inference

  • ai-leaderboard

  • ai-library

  • ai-rank

  • ai-security

  • ai-serving

  • ai-tools

  • ai-train

  • ai-video

  • ai-workflow

  • AIGC

  • algorithms

  • alibaba

  • amazon

  • android

  • anthropic

  • arabic

  • audio

  • aws

  • benchmark

  • benchmarks

  • biology

  • blog

  • book

  • bun

  • bytedance

  • chatbot

  • chatgpt

  • chemistry

  • claude

  • claude-code

  • cli

  • clickhouse

  • code

  • codex

  • coding

  • coding-agents

  • common-crawl

  • copilot

  • course

  • cpu

  • cuda

  • cursor

  • deepmind

  • deepseek

  • depth

  • devops

  • diffusers

  • distillation

  • docker

  • drug-discovery

  • electron

  • embeddings

  • engineering

  • evaluation

  • facebook

  • finance

  • flow-matching

  • foundation

  • foundation-model

  • fp8

  • gcode

  • gcp

  • gemini

  • gemini-cli

  • gemma

  • genomics

  • gguf

  • gitHub

  • github

  • go

  • google

  • gradient-booting

  • grok

  • groq

  • huggingface

  • hy_v4

  • image

  • imatrix

  • ios

  • java

  • javascript

  • json

  • kimi

  • kotlin

  • kubernetes

  • laion

  • llama.cpp

  • LLM

  • llm

  • long-horizon

  • lora

  • mLOps

  • math

  • mcap

  • mcp

  • mcp-client

  • mcp-server

  • meta-ai

  • meta-pytorch

  • metal

  • microsoft

  • mlops

  • mobile

  • mocap

  • moe

  • multilingual

  • multimodal

  • mysql

  • NLP

  • nlp

  • nodejs

  • numpy

  • nvidia

  • ocr

  • ollama

  • openai

  • opencode

  • pandas

  • paper

  • parquet

  • physics

  • pi

  • plugin

  • polars

  • postgres

  • privacy

  • programming

  • prompt-engineering

  • pwa

  • python

  • pytorch

  • qwen

  • react

  • reasoning

  • redis

  • refactoring

  • research

  • retrieval

  • RL

  • rl

  • robotics

  • rust

  • safetensors

  • science

  • security

  • segmentation

  • sft

  • shodan

  • skillkit

  • software-engineering

  • sora

  • speech

  • sqlite

  • ssh

  • stt

  • supabase

  • swe

  • swift

  • tensorrt

  • terminal

  • thinking

  • trae

  • transformers

  • translation

  • tts

  • tutorial

  • typescript

  • unsloth-dynamic

  • vibe-coding

  • video

  • vision

  • vllm

  • voice

  • vue

  • vulkan

  • web-search

  • windsurf

  • xAI

  • xai

  • youtube

Hugging Face
AI Model·2026
Icon for item

Needle

Henry Ndubuaku, Jakub Mroz +7

A distilled 26M-parameter encoder–decoder LLM for on-device function-calling and tool use. Uses a pure-attention Simple Attention Network, provides open weights and local finetuning, and targets high-throughput inference on the Cactus runtime.

#llm#huggingface#github#ai-inference#ai-serving+5
GitHub
AI Infra·2026
Icon for item

turbovec

RyanCodrai

Compresses high-dimensional embeddings into low-bit TurboQuant indexes for fast, memory-efficient local vector search. Supports online ingest (no train/rebuild), SIMD kernels that match or beat FAISS, per-vector length-renormalization, and runtime allowlists — suited for privacy-sensitive, low-latency RAG.

#rust#python#embeddings#RAG#ai-inference+4
Hugging Face
AI Model·2026
Icon for item

Mistral Medium 3.5 128B

Mistral AI

A dense 128B multimodal model with a 256k context window, configurable reasoning effort, and native function-calling for agentic workflows. Supports text+image input, multilingual output, and is released on Hugging Face under a Modified MIT license with revenue-based exceptions.

#foundation-model#multimodal#LLM#vllm#huggingface+6
GitHub
AI Infra·2026
Icon for item

CubeSandbox

Tencent Cloud

Provides hardware-isolated, sub-60ms, ultra-low-overhead sandboxes to run untrusted LLM/agent code. Offers event-level snapshots, kernel-level egress control, credential vaulting, and drop-in E2B SDK compatibility for high-density AI agent deployment.

#rust#ai-agent#ai-deploy#security#mLOps+1
Hugging Face
AI Model·2026
Icon for item

SuperGemma4-26B Uncensored Fast (GGUF v2)

Jiunsong

Provides a compact GGUF export of a tuned Gemma‑4 26B variant for local inference, optimized for llama.cpp and Apple Silicon to deliver faster, less‑censored chat and coding outputs. Includes Q4_K_M quantization and a neutral embedded template for more reliable local deployments.

#foundation-model#LLM#llm#ai-deploy#ai-inference+3
Hugging Face
AI Model·2026
Icon for item

Nemotron-Labs-TwoTower-30B-A3B-Base-BF16

Fitsum Reda, John Kamalu +4·NVIDIA Corporation

Generates text by iteratively denoising blocks of tokens with a two-tower design: a frozen autoregressive context tower and a trainable diffusion denoiser tower, trading minimal quality loss for higher wall-clock throughput.

#nvidia#huggingface#transformers#pytorch#llm+4
Hugging Face
AI Model·2026
Icon for item

Hy3 preview

Tencent Hy Team

A Mixture-of-Experts instruct-capable LLM (295B total, 21B active) designed for long-context reasoning, code/agent workflows and instruction-following; released by Tencent Hy Team with safetensors weights on Hugging Face.

#llm#huggingface#vllm#ai-train#ai-agent+4
Hugging Face
AI Model·2026
Icon for item

Qwen3.6-35B-A3B-DFlash

z-lab

Drafts multiple tokens in parallel with a lightweight block-diffusion drafter to enable speculative decoding for faster LLM inference. Designed to pair with Qwen3.6-35B-A3B and reports up to ~2.9× throughput improvements on common benchmarks.

#huggingface#llm#vllm#ai-inference#ai-deploy+3
Hugging Face
AI Model·2026
Icon for item

Qwen3.6-35B-A3B-Claude-4.6-Opus-Reasoning-Distilled-GGUF

hesamation

GGUF quantized files for a Qwen3.6-35B checkpoint fine-tuned with Claude Opus 4.6-style chain-of-thought distillation to improve reasoning. Offers multiple llama.cpp-compatible quant options (Q4/Q5/Q6/Q8) for local text-generation inference.

#huggingface#llm#nlp#ai-inference#ai-train+1
Hugging Face
AI Model·2026
Icon for item

Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled

lordx64

Fine-tuned Qwen3.6-35B-A3B MoE that reproduces Claude Opus 4.7-style chain-of-thought with explicit <think>…</think> blocks. Offers sparse activation (256 experts, ~3B active params), 64k context, and GGUF builds for local inference; best for long, multi-step reasoning but may emit very long reasoning traces.

#huggingface#llm#vllm#ai-inference#ai-serving+3
Hugging Face
AI Model·2026
Icon for item

unsloth/Kimi-K2.6-GGUF

unsloth

Provides a GGUF-packaged, native-INT4 quantized build of the multimodal Kimi K2.6 model for image-text-to-text inference — packaged for local/self-hosted inference engines (vLLM, SGLang, KTransformers) to reduce footprint while keeping multimodal capabilities.

#huggingface#vllm#ai-inference#ai-serving#ai-deploy+4
Hugging Face
AI Model·2026
Icon for item

Granite Embedding 97M Multilingual R2

IBM Granite Embedding Team

Produces 384‑dim multilingual (and code) embeddings with up to 32,768 token context, optimized for low‑latency production retrieval. Compact 97M model with ONNX/OpenVINO and vLLM/GGUF deployment options for edge and high‑throughput use.

#embeddings#multilingual#transformers#huggingface#vllm+4
  • Previous
  • 1
  • More pages
  • 9
  • 10
  • 11
  • More pages
  • 20
  • Next