AIAIAny
  • Search
  • Collection
  • Category
  • Tag
  • Daily AI
AIAIAny

Tag

Explore by tags

AIAIAny

Curated AI Resources for Everyone

[email protected]

Powered by airss.app

Product
  • Search
  • Collection
  • Category
  • Tag
Resources
  • Blog
Company
  • Privacy Policy
  • Terms of Service
  • Sitemap
Copyright © 2026 All Rights Reserved.
  • All

  • 30u30

  • ASR

  • ChatGPT

  • GNN

  • IDE

  • RAG

  • agent-skills

  • ai

  • ai-agent

  • ai-api

  • ai-api-management

  • ai-client

  • ai-coding

  • ai-demos

  • ai-deploy

  • ai-development

  • ai-framework

  • ai-image

  • ai-image-demos

  • ai-inference

  • ai-leaderboard

  • ai-library

  • ai-rank

  • ai-security

  • ai-serving

  • ai-tools

  • ai-train

  • ai-video

  • ai-workflow

  • AIGC

  • algorithms

  • alibaba

  • amazon

  • android

  • anthropic

  • arabic

  • audio

  • aws

  • benchmark

  • benchmarks

  • biology

  • blog

  • book

  • bun

  • bytedance

  • chatbot

  • chatgpt

  • chemistry

  • claude

  • claude-code

  • cli

  • clickhouse

  • code

  • codex

  • coding

  • coding-agents

  • common-crawl

  • copilot

  • course

  • cpu

  • cuda

  • cursor

  • deepmind

  • deepseek

  • depth

  • devops

  • diffusers

  • distillation

  • docker

  • drug-discovery

  • electron

  • embeddings

  • engineering

  • evaluation

  • facebook

  • finance

  • flow-matching

  • foundation

  • foundation-model

  • fp8

  • gcode

  • gcp

  • gemini

  • gemini-cli

  • gemma

  • genomics

  • gguf

  • gitHub

  • github

  • go

  • google

  • gradient-booting

  • grok

  • groq

  • huggingface

  • hy_v4

  • image

  • imatrix

  • ios

  • java

  • javascript

  • json

  • kimi

  • kotlin

  • kubernetes

  • laion

  • llama.cpp

  • LLM

  • llm

  • long-horizon

  • lora

  • mLOps

  • math

  • mcap

  • mcp

  • mcp-client

  • mcp-server

  • meta-ai

  • meta-pytorch

  • metal

  • microsoft

  • mlops

  • mobile

  • mocap

  • moe

  • multilingual

  • multimodal

  • mysql

  • NLP

  • nlp

  • nodejs

  • numpy

  • nvidia

  • ocr

  • ollama

  • openai

  • opencode

  • pandas

  • paper

  • parquet

  • physics

  • pi

  • plugin

  • polars

  • postgres

  • privacy

  • programming

  • prompt-engineering

  • pwa

  • python

  • pytorch

  • qwen

  • react

  • reasoning

  • redis

  • refactoring

  • research

  • retrieval

  • RL

  • rl

  • robotics

  • rust

  • safetensors

  • science

  • security

  • segmentation

  • sft

  • shodan

  • skillkit

  • software-engineering

  • sora

  • speech

  • sqlite

  • ssh

  • stt

  • supabase

  • swe

  • swift

  • tensorrt

  • terminal

  • thinking

  • trae

  • transformers

  • translation

  • tts

  • tutorial

  • typescript

  • unsloth-dynamic

  • vibe-coding

  • video

  • vision

  • vllm

  • voice

  • vue

  • vulkan

  • web-search

  • windsurf

  • xAI

  • xai

  • youtube

GitHub
Embodied AI·2024
Icon for item

NVIDIA Cosmos

NVIDIA

Provides an open platform of omnimodal world models, datasets, and tools to build Physical AI — joint perception, generation, and action reasoning for robots, autonomous vehicles, and smart infrastructure. Supports images, video, audio, and action-conditioned workflows.

#nvidia#multimodal#foundation-model#diffusers#vllm+9
GitHub
AI Client·2025
Icon for item

Rowboat

Rowboat Labs

Connects to Gmail, Calendar, and meeting notes to build a local, Obsidian-compatible Markdown graph it acts on — drafting emails, briefs, and decks. Memory accumulates instead of resetting each session; runs on local or hosted models, extensible via MCP.

#mcp#mcp-client#ai-agent#agent-skills#RAG+5
GitHub
AI Audio·2025
Icon for item

Handy

cjpais·YNYNG LLC

Press a configurable shortcut, speak, and have your words transcribed and pasted into the active app. Runs Whisper or the CPU-friendly Parakeet V3 fully offline; a Tauri + Rust build with Silero voice-activity detection and optional GPU acceleration.

#ASR#audio#rust#github#ai-tools+1
GitHub
AI Audio·2025
Icon for item

IndexTTS

Yunpei Li, Xun Zhou +12·Bilibili, IndexTeam

Zero-shot, single‑reference voice cloning TTS with multilingual support (ZH/EN/JA/ES/AR), fine-grained emotion and duration control, and pronunciation hooks (Pinyin/CMU/Kana); ships model weights, Web UI and production deployment recipes for local or server use.

#tts#voice#audio#multilingual#speech+4
GitHub
AI Audio·2025
Icon for item

OpenSuperWhisper

Starmel

Provides real-time, local audio recording and transcription on macOS using Whisper and Parakeet engines, with global hotkeys and hold-to-record behavior. Includes model download, microphone selection, drag-and-drop file transcription, multilingual auto-detection and Asian-language autocorrect; Apple Silicon only.

#stt#speech#audio#voice#github+3
GitHub
MCP Client·2025
Icon for item

AbletonMCP

Siddharth Ahuja (ahujasid)

Connects Claude (via the Model Context Protocol) to Ableton Live so the LLM can create and edit tracks, clips, instruments, and control playback through a socket-based MCP server and an Ableton MIDI Remote Script.

#mcp#mcp-server#mcp-client#claude#anthropic+4
GitHub
AI Client·2025
Icon for item

Google AI Edge Gallery

Google AI Edge

Runs open-source LLMs and multimodal models entirely on mobile devices for offline, private inference. Offers Agent Skills, Thinking Mode, Ask Image, audio scribe, model management and benchmarks, with Gemma 4 and Hugging Face integration.

#google#gemini#llm#ai-client#huggingface+6
GitHub
AI Audio·2025
Icon for item

Chatterbox TTS

Resemble AI

Open-source TTS that clones a voice from a short reference clip across 23+ languages, with adjustable emotional intensity via exaggeration/cfg controls and a built-in Perth neural watermark on every output.

#audio#github#ai-tools#pytorch#huggingface
GitHub
AI Audio·2025
Icon for item

Magenta RealTime 2

Magenta (Google)

A toolkit and open-weights system for real-time streaming music generation — offers two model sizes (230M / 2.4B), a Python inference library (JAX/MLX), and a C++ engine optimized for Apple Silicon for embedding into DAWs and apps; real-time streaming requires M‑series chips.

#audio#huggingface#google#python#ai-inference+3
GitHub
AI Deploy·2025
Icon for item

WhisperLiveKit

QuentinFuxa·Independent

Turns OpenAI Whisper into a live streaming transcriber: audio flows in over WebSocket and text returns word-by-word instead of after full utterances. Adds SimulStreaming and LocalAgreement decoding, Silero VAD, and speaker diarization, all self-hosted.

#github#ai-tools#ai-inference#ai-serving#ASR+2
GitHub
AI Audio·2025
Icon for item

VibeVoice

Microsoft·Microsoft Research

Synthesizes up to 90 minutes of multi-speaker speech in one pass, with as many as four voices in a single conversation. Pairs continuous acoustic and semantic tokenizers at a 7.5 Hz frame rate with a next-token diffusion head on an LLM backbone.

#microsoft#audio#github#huggingface#paper+2
GitHub
AI Audio·2025
Icon for item

SAM-Audio

Bowen Shi, Andros Tjandra +13·Meta AI (FAIR)

Isolates any single sound from a complex audio mixture using a text description, a visual cue from a video frame, or a time span, returning both the isolated target and the residual. Released in small, base, and large sizes plus visual-prompt variants.

#audio#foundation-model#meta-ai#github
  • Previous
  • 1
  • More pages
  • 4
  • 5
  • 6
  • More pages
  • 13
  • Next