AIAIAny
  • Search
  • Collection
  • Category
  • Tag
  • Daily AI
AIAIAny

Tag

Explore by tags

AIAIAny

Curated AI Resources for Everyone

[email protected]

Powered by airss.app

Product
  • Search
  • Collection
  • Category
  • Tag
Resources
  • Blog
Company
  • Privacy Policy
  • Terms of Service
  • Sitemap
Copyright © 2026 All Rights Reserved.
  • All

  • 30u30

  • ASR

  • ChatGPT

  • GNN

  • IDE

  • RAG

  • agent-skills

  • ai

  • ai-agent

  • ai-api

  • ai-api-management

  • ai-client

  • ai-coding

  • ai-demos

  • ai-deploy

  • ai-development

  • ai-framework

  • ai-image

  • ai-image-demos

  • ai-inference

  • ai-infra

  • ai-leaderboard

  • ai-library

  • ai-rank

  • ai-security

  • ai-serving

  • ai-tools

  • ai-train

  • ai-video

  • ai-workflow

  • AIGC

  • algorithms

  • alibaba

  • amazon

  • android

  • anthropic

  • arabic

  • audio

  • aws

  • benchmark

  • benchmarks

  • bf16

  • biology

  • blog

  • book

  • bun

  • bytedance

  • chatbot

  • chatgpt

  • chemistry

  • claude

  • claude-code

  • cli

  • clickhouse

  • code

  • codex

  • coding

  • coding-agents

  • comfyui

  • common-crawl

  • copilot

  • course

  • cpu

  • cuda

  • cursor

  • deepmind

  • deepseek

  • depth

  • desktop

  • devops

  • diffusers

  • distillation

  • docker

  • drug-discovery

  • electron

  • embeddings

  • embodied-ai

  • engineering

  • evaluation

  • facebook

  • finance

  • flow-matching

  • foundation

  • foundation-model

  • fp4

  • fp8

  • gcode

  • gcp

  • gemini

  • gemini-cli

  • gemma

  • genomics

  • gguf

  • gitHub

  • github

  • go

  • google

  • gpu

  • gradient-booting

  • grok

  • groq

  • gsq

  • huggingface

  • hy_v4

  • image

  • imatrix

  • ios

  • java

  • javascript

  • json

  • kimi

  • kotlin

  • kubernetes

  • laion

  • llama.cpp

  • LLM

  • llm

  • long-horizon

  • lora

  • mLOps

  • manipulation

  • math

  • mcap

  • mcp

  • mcp-client

  • mcp-server

  • meta-ai

  • meta-pytorch

  • metal

  • microsoft

  • mlops

  • mobile

  • mocap

  • moe

  • multilingual

  • multimodal

  • mysql

  • nli

  • NLP

  • nlp

  • nodejs

  • numpy

  • nvidia

  • ocr

  • ollama

  • openai

  • opencode

  • pandas

  • paper

  • parquet

  • physics

  • pi

  • plugin

  • polars

  • postgres

  • prefix-caching

  • privacy

  • programming

  • prompt-engineering

  • pruning

  • pwa

  • python

  • pytorch

  • quantization

  • qwen

  • rco

  • react

  • reasoning

  • red-teaming

  • redis

  • refactoring

  • reranker

  • research

  • retrieval

  • RL

  • rl

  • robotics

  • routing

  • rust

  • safetensors

  • science

  • security

  • segmentation

  • sft

  • shodan

  • skillkit

  • slam

  • software-engineering

  • sora

  • speech

  • sqlite

  • ssh

  • stt

  • supabase

  • swe

  • swift

  • tensorrt

  • terminal

  • thinking

  • trae

  • training-data

  • transformers

  • translation

  • tts

  • turkish

  • tutorial

  • typescript

  • unsloth-dynamic

  • vibe-coding

  • video

  • vision

  • vllm

  • voice

  • vue

  • vulkan

  • vultr

  • web-search

  • webdataset

  • windsurf

  • world-model

  • writing

  • xAI

  • xai

  • youtube

Speech Technology Papers·2026
Icon for item

Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval

Ilia Semenkov, Daria Kleeva +3

Retrieves short speech segments from MEG recordings with a compact interpretable neural decoder trained against wav2vec 2.0 embeddings, and maps decoder weights to cortical source space to reveal which acoustic and linguistic features drive retrieval.

#paper#speech#audio#retrieval#embeddings+1
Machine Learning Engineering Papers·2026
Icon for item

Knowledge-Geometry Decoupling: Refreshable Pretrained Transfer for Streaming Recommendation

Zixuan Wang, Yuhong Chen +11·[email protected], [email protected] +3

A pretrain-then-transfer method for streaming recommendation that decouples refreshable behavioral knowledge from task-specific geometry to enable continual model refresh without downstream interference; introduces Behavioral Multi-Token Prediction and Anchored Calibration Residual and shows 4–12% offline gains plus live Shopee A/B lifts.

#paper#foundation-model#embeddings#benchmarks#code+5
Hugging Face
AI Dataset·2026
Icon for item

British Library Book Images

Daniel van Strien·British Library, British Library Labs +3

Provides 1,080,814 images extracted from ~65,000 digitised British Library book volumes (c.1510–c.1900), split into four algorithmic image-type configs and packaged as parquet for image–text multimodal research and retrieval.

#image#book#ocr#parquet#huggingface+5
Computer Vision Papers·2026
Icon for item

Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval

Zelong Sun, Jun Wang +4

Generates retrieval-centric Chain-of-Thought (RC-CoT) over initially retrieved candidates to improve unified multimodal retrieval via reranking or full-corpus re-retrieval with a dual-mode embedder. Trains an embedder–adviser framework (UniME-R1) using mined hard negatives, supervised learning, and retrieval-oriented reinforcement learning.

#multimodal#retrieval#embeddings#reasoning#RL+2
Hugging Face
AI Model·2026
Icon for item

WeMM-Embedding-9B

Junjie Zhou, Ke Mei +4·WeChat Vision (Tencent Inc.), Tencent

Generates L2-normalized multimodal embeddings (default 4,096‑D) for text, images, videos and visual documents, supporting interleaved inputs and flexible dimension truncation (Matryoshka). Designed for cross-modal retrieval, ranking and downstream retrieval systems; audio is not supported.

#multimodal#embeddings#qwen#transformers#safetensors+6
Computer Vision Papers·2026
Icon for item

WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report

Junjie Zhou, Ke Mei +4·WeChat Vision (Tencent), Tencent Inc.

Generates unified embeddings for text, images, video, visual documents and interleaved multimodal inputs with configurable output dimensions and Matryoshka truncation to trade accuracy for cost. Model weights and code are released under Apache-2.0; the 9B variant scores 80.6 on MMEB-v2.

#multimodal#embeddings#vision#video#qwen+5
Hugging Face
AI Dataset·2026
Icon for item

Qdrant-FineWeb-10B

Qdrant, Vultr +5

A 10‑billion‑document retrieval benchmark with per‑document 768‑dim unit‑norm dense embeddings and mGTE sparse embeddings, FineWeb text/metadata, and exact top‑1000 MS MARCO ground truth for ~120k queries. Built for large‑scale evaluation of dense/sparse/hybrid retrieval, filtered search, indexing, ANNS algorithms, and embedding compression.

#embeddings#benchmark#huggingface#common-crawl#retrieval+5
Hugging Face
AI Model·2026
Icon for item

EmbeddingGemma 2

Google DeepMind

Generates unified 768‑dimensional embeddings for text (including code), images, video and audio to enable cross‑modal semantic search and retrieval. Supports task instruction prefixes, Matryoshka truncation to 128/256/512/768 dims, and modular encoders for on‑device use under an Apache‑2.0 license.

#embeddings#multimodal#gemma#google#huggingface+10
Hugging Face
AI Model·2026
Icon for item

Needle 3

Henry Ndubuaku, Karen Mosoyan +6·Cactus Compute, Inc.

Runs locally on constrained devices to turn text into guaranteed-parsable JSON tool calls, typed structured extractions, or sentence embeddings. Delivered as a single compact weights file (8–29 MB) with a laddered 2–20-layer design, low-bit quantisation and calibrated confidence scores for on-device apps.

#foundation-model#llm#embeddings#ai-inference#ai-deploy+7
Hugging Face
AI Model·2026
Icon for item

CLM-v0.1-8B

Jacky Kwok, Hangoo Kang +5·Contrastive-LM

Scores candidate actions against a textual state using contrastive state/action embeddings for very fast zero-shot ranking and typed decision-making. Built as two small projection heads on frozen Qwen3-8B; fine-tunable as a verifier for agentic benchmarks.

#qwen#reranker#embeddings#llm#nlp+9
Hugging Face
AI Dataset·2026
Icon for item

Doctor-Patient Conversations — All Human Diseases (Opus 5.5)

Nisten Tahiraj

Synthetic, clinician-verified ChatML dataset of 2,194 doctor–patient encounters covering 2,194 unique human diseases; each JSONL record includes 20 structured fields, verified PubMed references, realistic vitals/labs, and is intended for RAG and model fine-tuning (not medical advice).

#huggingface#claude#sft#RAG#llm+5
Hugging Face
AI Audio·2026
Icon for item

Whistle

Jakub Mroz, Henry Ndubuaku +6·Cactus Compute, Inc.

On-device speech-to-text for short clips (up to 30s) in seven languages, yielding transcripts, word-level timestamps and per-frame speech embeddings in a single 16.9 MB model. Runs on the Needle CPU engine with 2–4 bit quantization, supports keyword biasing and returns empty transcripts for silence.

#quantization#speech#stt#audio#embeddings+6
  • Previous
  • 1
  • 2
  • More pages
  • 11
  • 12
  • 13
  • Next