Must-read paper from Google on self-improving agent harnesses.

If you auto-optimize your agent's harness, your eval score can go up while the agent gets worse on real tasks.

This paper shows how to prevent that.

Of five harness-evolution methods compared on agentic workspace https://t.co/wEl9NWwsRc
8