Papers/2609.04298
🧪 Test?View on arXiv

Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation

Not specified in the provided content

benchmarkingevaluationagenticopen-source
2609.04298
Builder Relevance
80%
7h ago

Abstract

The paper introduces Harbor Adapters and Harbor-Index to facilitate the evaluation of agents across diverse benchmarks.

Reality Card

Core Claim

The authors developed a unified evaluation infrastructure that ports over 80 benchmarks for agent evaluation and introduced a curated set of 82 high-quality tasks for comprehensive agentic evaluation.

Method / Result

Conducted a large-scale evaluation of 8 models across 54 benchmarks, with the strongest model achieving a 28.0% pass rate.

Limitations

No evaluated model-harness configuration exceeds a 30% pass rate, indicating potential challenges in achieving reliable performance.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers