AIAIAny
  • Search
  • Collection
  • Category
  • Tag
  • Daily AI
AIAIAny

Discover the Best AI Resources

Curated essentials, no noise — just what matters

AIAIAny

Curated AI Resources for Everyone

[email protected]

Powered by airss.app

Product
  • Search
  • Collection
  • Category
  • Tag
Resources
  • Blog
Company
  • Privacy Policy
  • Terms of Service
  • Sitemap
Copyright © 2026 All Rights Reserved.

Contents

Speech Technology Papers·2026
Icon for item

Motion-Omni: End-to-End Joint Speech and Full-Body Motion for Spoken Dialogue

Chengqian Ma, Wei Tao +2

Generates synchronized spoken dialogue and explicit full-body co-speech motion (facial expressions, hands, upper- and lower-body) end-to-end from the same hidden states, replacing the speech-then-motion cascade. Trains with a scalable pseudo-labeling pipeline (422,856 ranked pairs) and supports real-time inference (RTF 0.78) while matching teacher motion metrics within ~2%.

#speech#mocap#multimodal#qwen#ai-video+5
Hugging Face
AI Dataset·2026
Icon for item

SolarWM-Data

Junchao Huang

Provides a unified, camera-conditioned, multi-source video dataset and reproducible processing pipeline for training long-horizon video world models. Key features: a canonical frame-aligned contract (visuals, camera geometry, captions, quality metadata), 1.43M canonical clips with portable releases and reconstruction tools; large download and some backbone assets carry separate licenses.

#video#webdataset#world-model#long-horizon#ai-video+2
AI Agent Papers·2026
Icon for item

ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL

Zhuoshi Pan, Qizhi Pei +5·Tsinghua University, Tencent Youtu Lab +1

Trains LLM agents to proactively edit and manage their working context for long-horizon tasks using an expanded toolset (planning, long-term memory, soft offloading) and a fine-grained RL algorithm that identifies critical edits and assigns action-level credit. Improves accuracy while keeping contexts compact on long-context QA and deep search.

#long-horizon#ai-agent#RL#rl#agent-skills+5
Hugging Face
AI Model·2026
Icon for item

Qwen3.8-27B · GSQ-RCO GGUFs

Deep Algorithms and Systems Lab (DASLab), Institute of Science and Technology Austria, ISTA-DASLab

Provides per-tensor non-uniform GGUF quantizations of Qwen3.8-27B using GSQ and RCO, delivering high accuracy at 2.5–3.5 bits and including a BF16 vision projector for multimodal use. Optimized to run unmodified in llama.cpp, Ollama, and LM Studio.

#gguf#qwen#llm#llama.cpp#huggingface+5
Computer Vision Papers·2026
Icon for item

GenFirst: Generation Before Reconstruction for Stable End-to-End Latent Generative Modeling

Guangting Zheng, Yiyuan Zhang +5

Proposes GenFirst, a generation-before-reconstruction end-to-end training strategy for latent generative models that avoids latent collapse by prioritizing generative objectives and then progressively strengthening reconstruction, validated with strong gFID/GenEval results on ImageNet-256 and text-to-image tasks.

#paper#vision#image#ai-image#research+1
Hugging Face
AI Model·2026
Icon for item

orcarouter/GLM-5.3-Flash-Uncensored-FP8

orcarouter·orcarouter, zai-org (Z.ai / Zhipu AI)

Drop-in abliterated (refusal-removed) build of GLM-5.3-Flash that bakes refusal-direction removal into block-FP8 safetensors, yielding an uncensored 320B (18B active) multimodal MoE model with a 1M-token context. Intended for red-teaming, interpretability, and robustness research; MIT license; not for production without added guardrails.

#moe#multimodal#transformers#safetensors#fp8+6
Hugging Face
AI Model·2026
Icon for item

GLM 5.3 CRACK — Cybersecurity FP8

dealignai

Provides a cybersecurity-focused CRACK variant of GLM-5.3 FP8 that reduces refusals for offensive-security, red-team, exploit-development and malware-analysis queries while retaining native FP8 speed on Hopper GPUs; MIT-licensed for authorized security work.

#safetensors#fp8#moe#llm#red-teaming+5
Hugging Face
AI Dataset·2026
Icon for item

Spark-234K

Yu Li, Wei Li +5·Shanghai AI Laboratory, University of Science and Technology of China +3

Synthesizes 234K self-contained, high-difficulty scientific reasoning QA pairs by distilling research papers into compact 'reasoning skeletons'. Emphasizes mechanistic reasoning, hypothesis falsification, quantitative derivation and boundary calibration; built for SFT and reasoning evaluation.

#science#reasoning#sft#training-data#parquet+2
Hugging Face
AI Dataset·2026
Icon for item

Danbooru 2026 Tag Cleaning Corrections

Grio43

Provides image-level tag correction instructions for a Danbooru anime-image tagging corpus, listing per-post tags to add or remove. Contains 1.74M normalized correction rows (snapshot 2026-08-30); it's a corrections manifest (no images) intended to be applied to existing metadata.

#huggingface#parquet#pandas#polars#training-data+2
AI Video Papers·2026
Icon for item

DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution

Jiashu Zhu, Yanhao Zheng +8

Generates synchronized native 2K audio-video from a single first frame and a text prompt using a compact 7B joint generator. Combines gated cross-modal attention, progressive joint training, audio-video reinforcement learning, and an Autoregressive 1-Step 2K Refinement; releases a 7B generator and 2K Refiner for research use.

#audio#video#ai-video#multimodal#distillation+2
Hugging Face
AI Dataset·2026
Icon for item

Kimi Cyber Reasoning

echel0nn1881

Provides 997 chain-of-thought cybersecurity reasoning records distilled from the Kimi K3 model, each with an explicit <think> trace and a technical resolution or structured tool invocation. Includes verified tool-call objects, diffs, cross-domain coverage, and token-level metadata for fine-tuning and evaluating reasoning models.

#huggingface#reasoning#ai-security#security#kimi+4
Hugging Face
AI Dataset·2026
Icon for item

Last Translation Benchmark

Vilém Zouhar, Niyati Bafna +8·ETH Zurich, Johns Hopkins University +3

A living, crowdsourced dataset of hard-to-translate examples (text, images, audio, video) paired with handcrafted verification rules that flag concrete MT failures. LTBv1 contains 3,456 peer-reviewed examples across many language pairs and accepts ongoing contributions.

#translation#multilingual#benchmark#evaluation#nlp+5
  • Previous
  • 1
  • More pages
  • 188
  • 189
  • 190
  • More pages
  • 211
  • Next