Claude did best on a new benchmark for ‘agents that build agents’. It still passed fewer than a quarter of the tests. — PLINKFEED