AIAIAny
  • Search
  • Collection
  • Category
  • Tag
  • Daily AI
AIAIAny

Discover the Best AI Resources

Curated essentials, no noise — just what matters

AIAIAny

Curated AI Resources for Everyone

[email protected]

Powered by airss.app

Product
  • Search
  • Collection
  • Category
  • Tag
Resources
  • Blog
Company
  • Privacy Policy
  • Terms of Service
  • Sitemap
Copyright © 2026 All Rights Reserved.

Contents

Computer Vision Papers·2026
Icon for item

An Exam for Active Observers

Jiarui Zhang, Muzi Tao +4

Quantifies active visual observation in multimodal LLMs with ActiveVision, a 17-task benchmark that forces repeated perception rather than one-shot description. Finds frontier MLLMs fail badly (top model 10.6% vs humans 96.1%) and that model-generated vision code does not close the gap.

#multimodal#vision#evaluation#benchmark#reasoning+2
Hugging Face
AI Model·2026
Icon for item

Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF

DavidAU

Provides GGUF-format fine-tuned Qwen3.6-27B weights optimized for consumer hardware, offering NEO IMATRIX and MTP quant variants, vision support, 256k native context, and uncensored 'heretic' traces with published benchmark improvements over the base model.

#qwen#llm#multimodal#vision#huggingface+3
Large Language Model Papers·2026
Icon for item

Loop the Loopies!

Zitian Gao, Yilong Chen +5

A looped-Transformer LLM series using Mixture-of-Experts (20B with 2B active; 6B with 0.6B active) that trades extra pretraining compute for repeated looping. Shows superior compute-efficiency versus matched-compute vanilla baselines and attains gold-medal performance on 2025 IMO and IPhO after a post-training pipeline.

#transformers#LLM#reasoning#evaluation#paper+2
Hugging Face
Embodied AI·2026
Icon for item

MiniCPM-RobotManip

openbmb

Generates robot manipulation actions from visual observations and text instructions using a 1.5B vision-language-action model. Uses streaming context and visual-token compression to cut per-step compute, runs a unified policy across tasks, and is open-sourced on Hugging Face under Apache-2.0.

#robotics#vision#transformers#huggingface#pytorch+2
Hugging Face
Embodied AI·2026
Icon for item

MiniCPM-RobotTrack

openbmb

Predicts eight future [x,y,yaw] waypoints for language-conditioned embodied person-following using fused DINOv3 and SigLIP visual features; trained with quality-driven, DAgger-style self-evolving data and optimized for on-device inference (~5+ FPS, ~180 ms).

#robotics#vision#multimodal#transformers#pytorch+2
AI Agent Papers·2026
Icon for item

DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines

Runming He, Zhen Hao Wong +5

Guides an LLM agent to build persistent, editable DAG-based data pipelines via typed, incremental mutations instead of free-form scripts. Combines DataFlow-Skills, a Model Context Protocol exposing live operator registry and pipeline state, and a synchronized Web UI; achieves 93.3% end-to-end pass rate on a 12-task benchmark while cutting cost and latency versus script baselines.

#agent-skills#mcp#ai-agent#ai-workflow#LLM+4
Hugging Face
AI Dataset·2026
Icon for item

Kimi K3 Coding & Debugging Agent Traces

greghavens·greghavens, moonshiner +3

Provides verified, model-attested end-to-end agent coding and debugging trajectories (JSONL). Each whole-session trace was produced by moonshotai/kimi-k3 on the pi/openrouter runtime, passed acceptance tests and independent model screening — useful for SFT, distillation, and analyzing tool-use behavior.

#huggingface#ai-coding#ai-agent#agent-skills#pi+6
Hugging Face
AI Dataset·2026
Icon for item

Aether-7B-5Attn Intermediate Pretraining Checkpoints

FINAL-Bench, VIDRAFT (주식회사 비드래프트)

Provides intermediate pretraining checkpoints for the Aether-7B-5Attn base model to enable reproducible training-dynamics research. Includes three raw checkpoints (110k, 115k, 162k steps) packaged with model.safetensors, config, and tokenizer; uses a custom aether_v2_7way architecture requiring the aether_pkg loader.

#foundation-model#llm#multilingual#huggingface#ai-train+1
AI Agent Papers·2026
Icon for item

EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World

Qing Zong, Yue Guo +3

Models long-horizon interactive literary simulation where characters and world co-evolve; introduces an open‑schema framework with a Character Agent and an LLM-based World Model, plus seven trainable tasks and a dataset from 57 books for benchmarking persistent narrative state.

#paper#LLM#NLP#ai-agent#agent-skills+2
AI Video Papers·2026
Icon for item

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs

Yuhan Zhu, Changlian Ma +13

Predicts variable-cardinality sets of evidence intervals in videos to temporally ground queries using multimodal large language models. Combines caption-derived multi-span supervision, a temporal Wasserstein matching-free reward, and temporal IoU, yielding strong mIoU gains across multiple benchmarks.

#video#multimodal#LLM#qwen#paper+3
Hugging Face
AI Dataset·2026
Icon for item

Kimi-K3 Codex traces

AletheiaResearch, TeichAI +1

Provides newline-delimited JSON agent session traces (5 files) generated with Teich for moonshotai/kimi-k3, including recovered and embedded tool-schema snapshots so traces remain training-ready even when tools weren't invoked; includes guidance for Teich data preparation and conversion.

#codex#distillation#ai-agent#agent-skills#huggingface+3
Hugging Face
AI Model·2026
Icon for item

Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF

DavidAU

GGUF-quantized releases (NEO IMATRIX + MTP) of a multi-stage fine-tuned, uncensored Qwen3.5-9B model with vision enabled and a native 256k context window—optimized for instruction following, reasoning and image-text-to-text workflows; released under Apache-2.0.

#qwen#llm#multimodal#vision#reasoning+5
  • Previous
  • 1
  • More pages
  • 160
  • 161
  • 162
  • More pages
  • 194
  • Next