AI Editorial

The rest of the library explains, and explanation has no opinion. This is where the opinion goes. Each piece starts from something that actually happened in AI — a model release, a paper, a benchmark result, a failure, a policy — explains the mechanism underneath it, and argues its way to a position: what genuinely changed, what is only being marketed as change, and what someone learning the field should take from it. Every piece links what it argues from and the concepts you need to follow it.

4 editorials 4 strands 31 cited sources 6,471 words

Benchmarks, leaderboards and claims, and the distance between what is demonstrated and what is sold.

Evaluation & Evidence 18 September 2026 7 min read New

A score is a reading, not a property

Four commits landed in the Terminal-Bench repository on 11 September. None of them touched a model, and all of them changed what its tasks measure. Agentic benchmarks are maintained software with an expiry date, and their numbers should be read that way.

An agentic benchmark score is not a property of a model but a reading taken through a perishable instrument of scaffold, resource budget and expiring container.

Read it → 8 sources benchmarksagentsevaluationreproducibility