AIAIAny
  • Search
  • Collection
  • Category
  • Tag
  • Daily AI
AIAIAny

Discover the Best AI Resources

Curated essentials, no noise — just what matters

AIAIAny

Curated AI Resources for Everyone

[email protected]

Powered by airss.app

Product
  • Search
  • Collection
  • Category
  • Tag
Resources
  • Blog
Company
  • Privacy Policy
  • Terms of Service
  • Sitemap
Copyright © 2026 All Rights Reserved.

Contents

AI Agent Papers·2026
Icon for item

MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks

Yi Zhu, Xiongwei Wu +9

Assesses mobile planning agents' ability to call tools, plan long-horizon workflows, and coordinate sub-agents in realistic, interactive phone scenarios via a stateful executable sandbox. Covers 13 domains, 212 tools, evidence-based verification, and tests memory, skill usage, permission and runtime constraints.

#benchmark#benchmarks#mobile#ai-agent#agent-skills+4
Hugging Face
AI Dataset·2026
Icon for item

ThinkingBox-Bench

Microsoft

Evaluates whether tool-using LLM agents reliably complete stateful business workflows via 507 executable agent–tool–user tasks across retail, travel, auto insurance, neobank, and IT/HR consulting. Provides browsable Parquet tables for tasks, scenarios, and agent instructions; v1.0 is intended for evaluation-only.

#benchmark#evaluation#ai-agent#mcp-server#mcp+4
Hugging Face
AI Dataset·2026
Icon for item

Britannica Illustrated Pages

Daniel van Strien·BigLAM, Hugging Face +8

Provides 115,293 illustrated page images and a 975,345-row manifest sampled from scanned Encyclopaedia Britannica volumes (1768–1929), with per-page classifier probabilities for illustration — ready for image-classification, OCR-aware vision research, and illustration mining.

#huggingface#image#ocr#parquet#polars+2
Hugging Face
AI Dataset·2026
Icon for item

SageBio/mva-hackathon-2026-data

SageBio

Hugging Face dataset for the MVA Hackathon 2026 containing pediatric rare-disease genomic data (~85 GB across 11 files). Access requires accepting dataset conditions; intended for genomic ML, variant analysis, and hackathon submissions, with notable storage and privacy constraints.

#huggingface#genomics#biology#research#privacy
Hugging Face
AI Video·2026
Icon for item

MiniMax-H3-Fun-Controlnet-Union

Alibaba PAI

Conditions a MiniMax‑H3 video generator with a single ControlNet‑Union checkpoint to accept Canny, Depth, HED, MLSD or Pose control videos and run video inpainting. Guidance‑distilled for one‑pass inference; requires the base MiniMax‑H3 weights and specific control-branch config.

#ai-video#video#multimodal#huggingface#qwen+2
AI Agent Papers·2026
Icon for item

AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces

Sungho Park, Wonjoong Kim +11·Affiliation: KAIST, Affiliation: Southern University of Science and Technology +1

Automatically optimizes runtime harnesses for LLM agents by diagnosing failure traces and iteratively applying structured, generalizable patches. Combines batch-based failure diagnosis, code-like patch generation across prompts/tools/middleware, and validation-aware selection to raise long-horizon task success on multiple benchmarks.

#LLM#ai-agent#agent-skills#long-horizon#benchmarks+4
Hugging Face
AI Dataset·2026
Icon for item

IFM/Math-Reasoning

IFM

Provides large-scale mathematical problem-solving, rewriting, and dialogue data organized into five Parquet-backed subsets for reasoning-oriented language-model training. Subsets support streaming access, Dataset Viewer inspection, and per-subset provenance metadata; licensed Apache 2.0.

#parquet#training-data#math#reasoning#huggingface+2
AI Agent Papers·2026
Icon for item

Apodex 1.1: Scaling Agentic Intelligence for Complex Work

Apodex Team, B. An +69

Develops methods to scale agentic AI for sustained, verifiable execution of complex long-horizon work by expanding executable environments and training coordinated agents with a shared execution harness (AgentOS) to maintain state, provenance, and failure recovery.

#ai-agent#long-horizon#agent-skills#ai-workflow#llm+2
Hugging Face
AI Model·2026
Icon for item

GLM-5.3

Z.ai (zai-org)

A large open-weights MoE language model for complex coding, long-horizon agentic workflows, and cyber/security evaluations; post-trained from the GLM-5 family with substantial gains over GLM-5.2. Provides FP8/BF16 checkpoints and native support for very long contexts (up to 1M tokens).

#moe#fp8#safetensors#transformers#coding+6
Hugging Face
AI Audio·2026
Icon for item

Breeze TTS 2

BreezeBlue, RESONIA, INC.

Generates low-latency, instruction-driven English and Chinese speech for voice cloning, voice design, and directed performances; supports real-time streaming, reference-free voice creation, and reference-guided cloning. Open-weight PyTorch model released under a research/non-commercial license with GPU recommendations.

#pytorch#cuda#safetensors#transformers#tts+4
Hugging Face
AI Dataset·2026
Icon for item

Dataset.ET Amharic Speech

Snapwre Technologies PLC, Dataset.ET

Provides 22.7 hours of read Amharic speech (7,405 clips, 320 speakers) for ASR, collected via a crowdsourced Telegram bot and peer-validated; speaker- and prompt-disjoint train/validation/test splits, 16 kHz audio under CC BY 4.0.

#ASR#speech#audio#parquet#huggingface+3
AI Agent Papers·2026
Icon for item

FrontierChallenge: Evaluating Scientific Workflow Completion

Liangcai Su, Zhaopeng Feng +14

Evaluates AI agents' ability to complete end-to-end scientific workflows by releasing and assessing 97 tasks from a 300-task FrontierChallenge suite across chemistry, materials, life science, and electrochemistry. Finds that top agent configurations achieved only a 20.6% pass rate despite high partial scores, revealing a gap between partial progress/confident completion claims and actual complete scientific deliverables.

#benchmark#benchmarks#evaluation#agent-skills#ai-agent+6
  • Previous
  • 1
  • More pages
  • 184
  • 185
  • 186
  • More pages
  • 208
  • Next