Papers/2610.00025
🧪 Test?View on arXiv

Measuring the Microtask Eligibility Gap: When Is an Off-the-Shelf SLM Enough for an Agent Harness?

Author1, Author2, Author3, Author4, Author5

microtaskslanguage modelsbenchmarkingevaluation
2610.00025
Builder Relevance
70%
1h ago

Abstract

This paper investigates the performance of small language models (SLMs) on microtasks in comparison to a baseline, revealing significant eligibility gaps.

Reality Card

Core Claim

Off-the-shelf SLMs fail to meet practitioner-defined thresholds for microtasks, with no configurations passing eligibility tests across multiple models.

Method / Result

0 of 16 configurations passed the eligibility tests for the microtasks evaluated.

Limitations

The findings are sensitive to model size and quantization methods, which may affect reproducibility across different setups.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers