AIAIAny
  • Search
  • Collection
  • Category
  • Tag
  • Daily AI
AIAIAny

Tag

Explore by tags

AIAIAny

Curated AI Resources for Everyone

[email protected]

Powered by airss.app

Product
  • Search
  • Collection
  • Category
  • Tag
Resources
  • Blog
Company
  • Privacy Policy
  • Terms of Service
  • Sitemap
Copyright © 2026 All Rights Reserved.
  • All

  • 30u30

  • ASR

  • ChatGPT

  • GNN

  • IDE

  • RAG

  • agent-skills

  • ai

  • ai-agent

  • ai-api

  • ai-api-management

  • ai-client

  • ai-coding

  • ai-demos

  • ai-deploy

  • ai-development

  • ai-framework

  • ai-image

  • ai-image-demos

  • ai-inference

  • ai-leaderboard

  • ai-library

  • ai-rank

  • ai-security

  • ai-serving

  • ai-tools

  • ai-train

  • ai-video

  • ai-workflow

  • AIGC

  • algorithms

  • alibaba

  • amazon

  • android

  • anthropic

  • arabic

  • audio

  • aws

  • benchmark

  • benchmarks

  • biology

  • blog

  • book

  • bun

  • bytedance

  • chatbot

  • chatgpt

  • chemistry

  • claude

  • claude-code

  • cli

  • clickhouse

  • code

  • codex

  • coding

  • coding-agents

  • common-crawl

  • copilot

  • course

  • cpu

  • cuda

  • cursor

  • deepmind

  • deepseek

  • depth

  • devops

  • diffusers

  • distillation

  • docker

  • drug-discovery

  • electron

  • embeddings

  • embodied-ai

  • engineering

  • evaluation

  • facebook

  • finance

  • flow-matching

  • foundation

  • foundation-model

  • fp4

  • fp8

  • gcode

  • gcp

  • gemini

  • gemini-cli

  • gemma

  • genomics

  • gguf

  • gitHub

  • github

  • go

  • google

  • gradient-booting

  • grok

  • groq

  • huggingface

  • hy_v4

  • image

  • imatrix

  • ios

  • java

  • javascript

  • json

  • kimi

  • kotlin

  • kubernetes

  • laion

  • llama.cpp

  • LLM

  • llm

  • long-horizon

  • lora

  • mLOps

  • math

  • mcap

  • mcp

  • mcp-client

  • mcp-server

  • meta-ai

  • meta-pytorch

  • metal

  • microsoft

  • mlops

  • mobile

  • mocap

  • moe

  • multilingual

  • multimodal

  • mysql

  • NLP

  • nlp

  • nodejs

  • numpy

  • nvidia

  • ocr

  • ollama

  • openai

  • opencode

  • pandas

  • paper

  • parquet

  • physics

  • pi

  • plugin

  • polars

  • postgres

  • privacy

  • programming

  • prompt-engineering

  • pwa

  • python

  • pytorch

  • qwen

  • react

  • reasoning

  • red-teaming

  • redis

  • refactoring

  • research

  • retrieval

  • RL

  • rl

  • robotics

  • rust

  • safetensors

  • science

  • security

  • segmentation

  • sft

  • shodan

  • skillkit

  • software-engineering

  • sora

  • speech

  • sqlite

  • ssh

  • stt

  • supabase

  • swe

  • swift

  • tensorrt

  • terminal

  • thinking

  • trae

  • training-data

  • transformers

  • translation

  • tts

  • tutorial

  • typescript

  • unsloth-dynamic

  • vibe-coding

  • video

  • vision

  • vllm

  • voice

  • vue

  • vulkan

  • vultr

  • web-search

  • webdataset

  • windsurf

  • world-model

  • xAI

  • xai

  • youtube

Hugging Face
AI Dataset·2018
Icon for item

GLUE (General Language Understanding Evaluation benchmark)

Alex Wang, Amanpreet Singh +4·New York University, Paul G. Allen School of Computer Science & Engineering, University of Washington +1

A multi-task English NLU benchmark for evaluating models across nine tasks (acceptability, sentiment, paraphrase, similarity, and various NLI setups), with a diagnostic evaluation set and an online leaderboard to compare generalization and transfer learning.

#nlp#benchmark#evaluation#huggingface#paper+1
Large Language Model Papers·2021

Codex: Evaluating Large Language Models Trained on Code

Mark Chen, Jerry Tworek +2·OpenAI

Showed that fine-tuning a GPT model on public GitHub code yields a capable program synthesizer, and introduced HumanEval — the docstring-to-function benchmark that still anchors code-generation evaluation. A production variant powers GitHub Copilot.

#openai#code#codex#copilot#evaluation+2
Hugging Face
AI Dataset·2023
Icon for item

GAIA

Grégoire Mialon, Clémentine Fourrier +4·FAIR, Meta +3

Benchmark for evaluating general AI assistants with 466 short, real-world questions that require tool use, multimodality and reasoning; provides a public dev set and a withheld test set used for leaderboard evaluation.

#benchmark#evaluation#huggingface#parquet#multimodal+6
Hugging Face
AI Dataset·2023
Icon for item

TMMLU+

Zhi-Rui Tam, Ya-Ting Pai +5·iKala AI Lab, National Yang Ming Chiao Tung University

A multiple-choice benchmark for evaluating LLM understanding in Traditional Chinese across 66 subjects (elementary to professional). Contains ~22K verified questions covering STEM, humanities, social sciences and Taiwan-specific topics, with standardized splits and model leaderboards under an MIT license.

#benchmark#evaluation#nlp#llm#multilingual+3
Hugging Face
AI Dataset·2024
Icon for item

AnswerCarefully

Hisami Suzuki, Satoru Katsumata +4·LLM-jp (Center for Large Language Model Research and Development, National Institute of Informatics), National Institute of Informatics

Provides manually curated Japanese instruction pairs (questions and safe reference answers) for improving LLM output safety, covering broad harm categories and regionally sensitive cases. Includes English meta-tags and standard splits for benchmarking and fine-tuning.

#nlp#LLM#evaluation#benchmark#multilingual+5
GitHub
AI Agent·2024
Icon for item

Open Deep Research

langchain-ai

Orchestrates configurable deep-research agent workflows that combine LLMs, web search, and MCP tools to produce structured research reports and evaluation outputs. Supports LangGraph Studio, multiple model providers (OpenAI, Anthropic, local models), and Deep Research Bench evaluation for benchmarked comparisons.

#agent-skills#mcp#mcp-server#LLM#evaluation+5
GitHub
AI Infra·2024
Icon for item

AI-Infra-Guard (A.I.G)

Yong Yang, Xing Zheng +7·Tencent Zhuque Lab, Tencent Security Platform Department +1

Full-stack AI red‑teaming platform that fingerprints AI infrastructure for known CVEs, audits MCP servers and agent skills with LLM-driven analysis, and runs cross-model jailbreak evaluations; designed for hands-on security assessment of AI deployments.

#security#mcp#mcp-server#agent-skills#vllm+4
Hugging Face
AI Dataset·2025
Icon for item

Interaction2Code

whale99

A benchmark dataset for evaluating MLLM-driven interactive webpage code generation: provides prototyping screenshots, action.json interaction metadata, and example generation scripts across 127 webpages and 374 interactions to test dynamic UI-to-code capabilities.

#multimodal#image#code#github#LLM+3
GitHub
AI Agent·2025
Icon for item

DeepTeam

Jeffrey Ip·Confident AI

Simulates adversarial attacks against LLMs and AI agents to surface vulnerabilities (e.g., jailbreaks, prompt injection, PII leakage) and ships guardrails to block risky inputs/outputs; runs locally and can be driven from CLI or Python.

#LLM#evaluation#ai-agent#security#privacy+4
Hugging Face
AI Dataset·2025
Icon for item

olmOCR-bench

Jake Poznanski, Jon Borchardt +7·Allen Institute for Artificial Intelligence (AI2), AllenNLP / olmOCR team

Benchmark for evaluating OCR systems that convert PDFs and scans into Markdown and structured text: 1,403 PDFs and 7,010 unit tests covering text presence/absence, reading order, tables, and math formula accuracy. Diverse sources and ODC-BY-1.0 license for research use.

#ocr#evaluation#vision#huggingface#paper+1
GitHub
AI Agent·2025
Icon for item

Biomni: A General-Purpose Biomedical AI Agent

Kexin Huang, Serena Zhang +8

Autonomously executes diverse biomedical research tasks by combining LLM reasoning, retrieval-augmented planning, and code-based execution. Includes a web UI and Gradio demo, a curated Know‑How library, MCP integration, and a biology-tailored reasoning model (Biomni‑R0).

#ai-agent#biology#genomics#drug-discovery#agent-skills+7
AI Others·2025

The Second Half

Shunyu Yao

Argues AI has entered its 'second half': a working recipe (language pre-training priors + scale + reasoning) now generalizes RL across tasks, so the bottleneck shifts from inventing methods to defining problems and rethinking evaluation.

#RL#ai-agent#evaluation#LLM#blog
  • Previous
  • 1
  • 2
  • 3
  • More pages
  • 16
  • 17
  • Next