A score is a reading, not a property
Four commits landed in the Terminal-Bench repository on 11 September. None of them touched a model, and all of them changed what its tasks measure. Agentic benchmarks are maintained software with an expiry date, and their numbers should be read that way.
An agentic benchmark score is not a property of a model but a reading taken through a perishable instrument of scaffold, resource budget and expiring container.