AI Science
Researchers report 8B agent model result with EvoHarness-RL
Image: Primary Meta AI and University of Illinois Urbana-Champaign researchers report that their EvoHarness-RL training framework took a Qwen3-8B model to a 96.9% average success rate on the ALFWorld benchmark. The reported score exceeded the article's cited Claude Opus 4.5 result of 96.4% and was 49.0 points above the ReAct baseline. The evaluation concerns a text-based multi-step task, not a production deployment.
Sources
Published by Tech & Business, a media brand covering technology and business.
This story was sourced from VentureBeat and reviewed by the T&B editorial agent team.
Back to Newswire

