AIAIAny
  • Search
  • Collection
  • Category
  • Tag
  • Daily AI
AIAIAny

Category

Explore by categories

AIAIAny

Curated AI Resources for Everyone

[email protected]

Powered by airss.app

Product
  • Search
  • Collection
  • Category
  • Tag
Resources
  • Blog
Company
  • Privacy Policy
  • Terms of Service
  • Sitemap
Copyright © 2026 All Rights Reserved.
  • All Categories

  • AI Leaderboard

  • AI Agent Tutorials

  • AI Coding Tutorials

  • AI Model

  • AI Agent Papers

  • Chatbot

  • AI Dataset

  • Machine Learning Foundation Books

  • AI Train

  • AI Deploy

  • AI Client

  • Machine Learning Foundation Papers

  • Machine Learning Foundation Tutorials

  • AI Image Demos

  • AI Agent

  • Large Language Model Tutorials

  • Large Language Model Papers

  • Machine Learning Engineering Papers

  • Computer Vision Tutorials

  • Computer Vision Papers

  • Natural Language Processing Papers

  • Reinforcement Learning Papers

  • Speech Technology Papers

  • AI API

  • AI Coding

  • AI Image

  • AI Video

  • MLOps

  • MCP Client

  • MCP Server

  • AI Video Papers

  • AI Audio

  • AI Others

  • AI Infra

  • Embodied AI

AI Deploy·2021
Icon for item

KServe

KServe community·Google, IBM +4

Serves predictive and generative ML models on Kubernetes via a single InferenceService CRD, with scale-to-zero, canary rollouts, and an OpenAI-compatible LLM path on vLLM. One autoscaling abstraction over PyTorch, XGBoost, ONNX, and HuggingFace.

#mlops#ai-inference#ai-serving#ai-deploy#vllm+3
GitHub
AI Infra·2021
Icon for item

xFormers

Facebook Research (Meta)·Meta AI

Drop-in transformer building blocks with custom CUDA kernels: memory-efficient exact attention (up to ~10x faster), block-sparse attention, fused softmax/layernorm/SwiGLU ops. Cuts VRAM and speeds up diffusion and LLM training on Nvidia GPUs.

#pytorch#ai-library#meta-ai#ai-development#foundation-model+1
AI Train·2021
Icon for item

Colossal-AI

HPC-AI Technology Inc. (Colossal-AI team), Shenggui Li +1·HPC-AI Technology Inc., National University of Singapore

Scales a single-GPU training script to thousands of GPUs through a unified interface, combining data, pipeline, tensor, and sequence parallelism. Its Gemini memory manager offloads tensors across GPU, CPU, and NVMe so models far larger than VRAM still fit.

#pytorch#ai-train#ai-inference#ai-serving#mlops+3
GitHub
AI Infra·2022
Icon for item

Self Hosting Guide

mikeroyal

A continuously-updated, categorized self-hosting guide that catalogs tools, deployment notes and resources for containers, networking, home automation and running LLMs/chatbots locally; provided as a long, structured README with links and quick setup tips.

#docker#llm#privacy#pi#chatbot+2
GitHub
AI Infra·2022
Icon for item

NVIDIA Warp

NVIDIA

Compiles plain Python functions into GPU or CPU kernels at runtime via a JIT decorator, with differentiable output that plugs into PyTorch, JAX, and Paddle. Ships physics, robotics, geometry, and FEM primitives — particles, meshes, ray-casting, FFT.

#nvidia#python#ai-framework#pytorch#physics+2
GitHub
AI Infra·2022
Icon for item

Cordis

Cordiverse

Provides runtime support for reversible effects and reactive coeffects so components can be declared, composed, and hot-replaced safely; includes effect tracking, coeffect resolution, a declarative component loader and HMR — aimed at plugin-driven agent harnesses and dynamic systems.

#plugin#deepseek#ai-framework#ai-agent#ai-development+3
GitHub
AI Infra·2022
Icon for item

FlashAttention

Dao-AILab, Tri Dao +4·Dao AI Lab

Fused CUDA kernels that compute exact attention without ever writing the full N×N score matrix to GPU memory, cutting memory from quadratic to linear and speeding up training and inference on A100/H100. Ships FlashAttention-2/3 plus KV-cache decode paths.

#pytorch#ai-inference#ai-train#ai-library#nvidia+3
AI Infra·2022
Icon for item

Instant Observability for Cloud & AI Applications

Yunshan Networks

Collects metrics, distributed traces, and continuous profiles via eBPF with zero code instrumentation, covering apps in any language plus gateways, service meshes, databases, and queues. Profiling adds under 1% overhead.

#mlops#mcp-server#ai-development#github#opencode
GitHub
AI Infra·2022
Icon for item

Manifest

mnfst

Smart model router for personal AI agents that sends each request to the cheapest model capable of handling it — cutting API costs by up to ~70%. Uses a fast 23-dimension scorer, automatic fallbacks, per-tier controls, and supports local Docker self-hosting or a cloud app; ideal for cost-sensitive personal agents.

#llm#ai-agent#ai-api-management#ai-serving#ai-inference+4
AI Agent·2022
Icon for item

LangChain: Observe, Evaluate, and Deploy Reliable AI Agents

Harrison Chase, Ankush Gola +1·LangChain, Inc.

Build, run, and monitor LLM agents across one stack: an open framework for chaining models and tools, LangGraph for stateful agent orchestration, and LangSmith for tracing, evaluation, and deployment in production.

#llm#RAG#embeddings#ai-agent#ai-framework+3
GitHub
AI Infra·2022
Icon for item

Reflex

reflex-dev (GitHub organization)

Build full‑stack web apps entirely in Python — write frontend components and backend state as Python classes with a reactive model. Provides fast refresh, deployment tooling, and AI-focused integrations such as an AI Builder and an Agent Toolkit for connecting LLMs and image models.

#python#ai-tools#ai-development#agent-skills#ai-deploy+3
GitHub
AI Agent·2022
Icon for item

LlamaIndex

LlamaIndex (run-llama), Jerry Liu·LlamaIndex

Connects LLMs to private and domain-specific data with ingestion, indexing, and retrieval primitives for RAG and agentic apps. Centers on document parsing via LlamaParse for 90+ file formats, schema-based extraction, and composable queries.

#ocr#agent-skills#RAG#llm#embeddings+5
  • Previous
  • 1
  • More pages
  • 3
  • 4
  • 5
  • More pages
  • 19
  • Next