AIAny
AI Agent2025
Icon for item

Harbor

Evaluates and optimizes AI agents and language models in containerized environments, supporting large-scale parallel benchmarks and RL rollouts. Integrates with third‑party providers for thousands of parallel environments and serves as the official harness for Terminal‑Bench.

Introduction

Benchmarks for AI agents are shifting from small, manual runs to continuous, large‑scale evaluation and RL-driven optimization. Harbor addresses this gap by making it straightforward to run large, reproducible agent evaluations and generate rollouts for RL training across containerized and cloud providers.

What Sets It Apart
  • Scalable parallel execution: designed to run hundreds-to-thousands of environments in parallel via providers like Daytona, Modal, LangSmith, Blaxel, and Novita Sandbox, so evaluations and rollout generation can be parallelized across cloud/container backends.
  • Benchmark-first integration: acts as the official harness for Terminal‑Bench and includes easy access to third‑party datasets and benchmarks, enabling consistent, comparable agent evaluation across models and agents.
  • Agent-agnostic evaluation and RL support: evaluates arbitrary agents (examples include Claude Code, Codex CLI, and others) and can export rollouts suitable for RL optimization workflows.
  • Reproducibility and citation readiness: project provides a citable DOI and a cookbook of end‑to‑end examples to help standardize experiments and results.
Who It's For + Tradeoffs

Great fit if you need repeatable, large‑scale agent evaluations or want to collect rollouts at scale for RL optimization, and you expect to run experiments across cloud/container providers. It benefits benchmarking teams, research groups comparing agent architectures, and MLOps teams operationalizing evaluation.

Look elsewhere if you only need ad‑hoc, single‑machine experiments or a lightweight interactive UI—Harbor is oriented toward automated, containerized workflows and multi‑provider orchestration, which entails a learning curve for orchestration and environment setup.

Information

  • Websitegithub.com
  • OrganizationsHarbor Framework Team
  • Published date2025/08/04

Categories

More Items

Hugging Face
AI Model2026

Multimodal agentic model for long-horizon computer and browser tasks, with visual self-correction and function-calling. The Pro variant is a 397B Mixture-of-Experts (≈17B active) model with a 262,144-token context window, Docker deployment recipes, and weights currently marked “coming soon.”

Hugging Face
AI Model2026

A multimodal, agentic LLM optimized for long‑horizon, visually grounded workflows — capable of operating browsers and terminals and autonomously executing and testing code. Open‑source weights are available and the family ships in mini, Pro and Max variants for different compute/quality tradeoffs.

Hugging Face

Provides ~483K agent instruction‑tuning trajectories for supervised fine‑tuning, including tool calls, environment feedback, errors/retries and verification across search, code, office and general agent workflows; static snapshots for SFT and mix‑ratio studies.