AIAIAny
  • Search
  • Collection
  • Category
  • Tag
  • Daily AI
AIAIAny

Discover the Best AI Resources

Curated essentials, no noise — just what matters

AIAIAny

Curated AI Resources for Everyone

[email protected]

Powered by airss.app

Product
  • Search
  • Collection
  • Category
  • Tag
Resources
  • Blog
Company
  • Privacy Policy
  • Terms of Service
  • Sitemap
Copyright © 2026 All Rights Reserved.

Contents

Natural Language Processing Papers·2026
Icon for item

ArcANE: Do Role-Playing Language Agents Stay in Character at the Right Time?

Woojung Song, Nalim Kim +4

Evaluates whether role-playing language agents follow a character's evolving psychological arc rather than a fixed persona, using ArcANE — an automatically constructed benchmark spanning 17 novels and 80 principal characters. Tests both in-text and out-of-text scenarios and compares context strategies and fine-tuned models.

#paper#NLP#LLM#ai-agent#agent-skills
Hugging Face
AI Model·2026
Icon for item

Nex-N2-mini

nex-agi

Provides compact, agentic text-generation for long-horizon, tool-enabled workflows — trading some peak capability for lower latency and easier on-prem deployment. Key features: adaptive/coherent thinking traces, function-calling support, and sglang/docker-ready serving.

#transformers#huggingface#llm#ai-agent#vibe-coding+2
Hugging Face
AI Dataset·2026
Icon for item

EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents

ServiceNow-AI

End-to-end evaluation framework for conversational voice agents that runs bot-to-bot audio simulations and scores agents on task accuracy (EVA-A) and interaction experience (EVA-X). Includes per-scenario backend state, accent/noise perturbations, and 213 scenarios across airline, healthcare HR, and enterprise IT domains.

#huggingface#voice#speech#ASR#tts+4
AI Agent Papers·2026
Icon for item

AdaPlanBench: Evaluating Adaptive Planning in Large Language Model Agents under World and User Constraints

Jiayu Liu, Cheng Qian +9

Dynamic interactive benchmark that tests whether LLM agents can adaptively plan and re-plan when world and user constraints are progressively revealed. Built on 307 household tasks with a multi-turn protocol that exposes hidden constraints only after plan violations, emphasizing iterative revision and constraint inference.

#LLM#NLP#ai-agent#agent-skills#paper
Natural Language Processing Papers·2026
Icon for item

Reinforcement Learning Elicits Contextual Learning of Unseen Language Translation

Hanxu Hu, Zdeněk Šnajdr +3

Trains LLMs with reinforcement learning using a surface chrF reward so models learn to extract and apply linguistic signals from rich context for translating completely unseen languages. Demonstrates better zero-shot translation than in-context learning or supervised fine-tuning, framing outcome-based RL as a meta-skill for language learning from context.

#RL#multilingual#translation#NLP#LLM+1
Hugging Face
AI Dataset·2026
Icon for item

AI Village (HuggingFace dataset)

AI Digest (aidigestorg), Hugging Face

Provides a complete, lightly-processed export of AI Village's >1-year multi-agent data: per-agent computer sessions (with screenshots), turn-by-turn computer-use logs, group chats, agent memories, goals, and daily summaries for research into agentic behaviour, multi-agent dynamics, long-horizon memory, and AI safety. Access is manually reviewed.

#ai-agent#agent-skills#LLM#llm#huggingface+3
Hugging Face
AI Model·2026
Icon for item

google/gemma-4-12B-it-qat-q4_0-gguf

Google DeepMind

Provides a GGUF-ready QAT (Q4_0) quantized build of Gemma 4 12B that preserves near-bfloat16 quality while reducing memory footprint for local inference; compatible with Transformers-based and GGUF runtimes.

#gemma#google#deepmind#huggingface#transformers+3
AI Agent Papers·2026
Icon for item

SWE-Explore: Benchmarking How Coding Agents Explore Repositories

Shaoqiu Zhang, Yuhang Wang +9

Measures how coding agents explore repositories by asking them to return a ranked, line-level list of code regions relevant to an issue under a fixed line budget. Covers 848 issues across 203 repos and 10 languages; evaluates coverage, ranking, and context-efficiency to isolate exploration quality.

#ai-agent#ai-coding#agent-skills#paper#code+1
Hugging Face
AI Model·2026
Icon for item

Gemma-4-12B-OBLITERATED

OBLITERATUS

A surgically modified Gemma 4 (12B) that removes refusal behavior while preserving benchmark parity; released as an uncensored research artifact with GGUF quantizations for local inference and red‑team/alignment evaluation.

#gemma#transformers#huggingface#llm#ai-inference+5
AI Video Papers·2026
Icon for item

MemDreamer: Decoupling Perception and Reasoning for Long Video Understanding via Hierarchical Graph Memory and Agentic Retrieval Mechanism

Cong Chen, Guo Gan +8

Decouples perception and reasoning for hours-long videos by streaming inputs into a three-tier Hierarchical Graph Memory and using an agentic Observation–Reason–Action retrieval loop; reduces reasoning context to ~2% of full video while improving benchmark accuracy.

#paper#ai-video#multimodal#GNN#agent-skills+3
Hugging Face
AI Model·2026
Icon for item

unsloth/gemma-4-12B-it-qat-GGUF

unsloth, Google DeepMind

GGUF-format QAT (quantization-aware training) build of Gemma 4 12B that reduces memory needs for local or lightweight inference while preserving near bfloat16 quality. Ready for any-to-any conversational pipelines and ecosystem deployment.

#gemma#huggingface#google#deepmind#transformers+5
Natural Language Processing Papers·2026
Icon for item

Your UnEmbedding Matrix is Secretly a Feature Lens for Text Embeddings

Songhao Wu, Zhongxin Chen +4

Removes the subspace of frequent, uninformative tokens that LLMs inject into text embeddings via the model's unembedding matrix. EmbedFilter is a lightweight linear transform that refines LLM-derived embeddings to improve zero‑shot semantic retrieval, enable dimensionality reduction, and speed up indexing; code on GitHub.

#embeddings#LLM#NLP#paper#github+3
  • Previous
  • 1
  • More pages
  • 132
  • 133
  • 134
  • More pages
  • 172
  • Next