Evidence & Evaluation 28 September 2026 7 min read 1,636 words

A rate cannot say stop

The Green Software Foundation's carbon standard for AI is being edited into ISO shape this month. Its rules for counting emissions are stricter than almost anything in corporate reporting. Its rules for what to divide them by quietly guarantee that the number can never ask anyone to build less.

The argument

SCI for AI divides a provider's emissions by what it consumed rather than by what it delivered, so the carbon of a single answer is the one figure the standard cannot produce.

The rationale document behind the Green Software Foundation's carbon standard for AI contains a small table worth sitting with. It gives three models and what their training is estimated to have cost: roughly 50 tonnes of CO₂e for GPT-2, roughly 1,200 for GPT-3, roughly 20,000 for GPT-4. Then it gives the same three divided by parameter count: about 33.3 tonnes per billion parameters, then 6.86, then 20. Under the first table the newest model is four hundred times the problem the oldest one was. Under the second it sits mid-field, and comfortably ahead of where the series started.

The document is unusually honest about why the second table is the one the specification adopted. Absolute totals, it says, invite the conclusion that newer models are simply worse for the environment, "with no clear path for improvement other than 'make smaller models'". And then the sentence that governs everything downstream of it: "This insight isn't actionable for organizations committed to advancing AI capabilities."

Read that as a design requirement, because that is what it is. A measurement standard has written down a conclusion it does not wish its numbers to support, and then chosen a denominator that cannot support it.

This matters more this month than last, because the specification has stopped being a working group document. It is an extension of ISO/IEC 21031:2024, the international standard for software carbon intensity, and on 24 September a pull request landed against it titled "Align SCI for AI baseline with ISO/IEC Directives Part 2" — the drafting rules you follow when a document is on its way to becoming an ISO standard rather than a foundation's opinion. The specification text itself was last substantively revised in August. What is being tidied now is the grammar of standardisation. Once that finishes, this becomes the thing procurement asks suppliers for.

So it is worth being precise about what it counts, and what it divides by, because those two halves of the same formula were written by different instincts.

The counting half is genuinely tough. SCI is defined as operational plus embodied emissions over a functional unit, and the AI extension insists that training emissions be calculated across the entire training duration — every epoch, every parameter update, intermediate and test runs, early stopping, failed and aborted runs, hyperparameter sweeps, discarded and superseded checkpoints. Anyone who has watched a large training programme knows how much of the electricity goes into runs that produced no model at all. Most published training figures quietly omit them. This specification says you may not.

It is tougher still on the escape hatch that most corporate carbon reporting depends on. Offsets do not count. Nor do renewable energy credits, nor electricity attribute certificates, nor power purchase agreements. The specification says so in terms: a provider's use of PPAs or RECs to claim clean energy for model training does not reduce its score. "We train on one hundred per cent renewables" becomes a statement about contracts rather than about physics, and the score is indifferent to it. That is a harder line than the comparable indices the foundation reviewed, and the review document says so plainly. Whoever wrote the exclusions clause was not looking for friends.

Then you reach the denominator, and the mood changes.

A provider must normalise by one of three units: per FLOP, per training token, or per billion parameters. Every one of them is an input. Not one of them is a thing the model does. There is no term anywhere in the formula for capability, accuracy, task completion, or usefulness of any kind — nothing that would let two models be compared on what they are for. The unit measures how much you spent, and then divides what you emitted by how much you spent.

Follow the arithmetic on per-FLOP. Run the same training job for twice as long on the same hardware, in the same region, on the same grid. Emissions roughly double. FLOPs roughly double. The score does not move. A rate whose denominator scales with its numerator is not insensitive to waste at the margin — a better kernel or a more efficient accelerator will still show up — but it is structurally incapable of noticing that the whole programme got larger. That is not a flaw in the implementation. It is what dividing by your own consumption means.

Per billion parameters is stranger, because it runs the wrong way. Emissions over parameter count improves when parameter count rises. A team that trains a trillion-parameter model scores better than a team that spent the same carbon on a hundred billion. And pruning — genuinely removing weights after the electricity has been spent — shrinks the denominator while the numerator stands still, which makes the score worse. The rationale document nonetheless lists "optimizing the size of models" among the behaviours these units encourage, and the specification's own worked example reports a per-parameter figure as "reflecting pruning of inactive model weights". The claim and the formula point in opposite directions.

There is a further concession in who picks. The provider chooses which of the three units to report, and the specification advises that the choice "should reflect the primary optimization focus of the provider's system design, training strategy, or architecture". The provider also chooses whether to report gross or effective values — total parameters or active ones, raw tokens or deduplicated. This is a standard in which the measured party selects the denominator, and is advised to select the one that flatters what it is already good at. Multiple units may be reported, and the specification encourages it, but nothing requires it.

Meanwhile the consumer side is denominated in exactly the unit that would have been useful: per token, per image, per workflow execution, deliberately matched to how services are billed so that carbon can sit beside cost in the same spreadsheet. And training is excluded from it by construction. Training belongs to the provider boundary; the consumer's score covers operation and monitoring only. The rationale is agency — consumers cannot control how a model was trained, so holding them to it produces a number they cannot act on.

The result is two scores in units that cannot be added, subtracted or traded off. A buyer choosing between two model APIs this quarter can be handed 130 kg CO₂e per billion tokens from one boundary and 4 grams per quadrillion FLOPs from the other, and there is no operation defined anywhere in the document that combines them. The specification's answer is that consumers may "consider" the provider score alongside their own. Consider it how? The single decision a consumer actually controls — which model to call, and therefore whose training bill they are voting to have paid again next year — is precisely where the arithmetic stops.

The agency argument deserves better than dismissal, because it is largely right. Amortising training emissions across inferences requires dividing by lifetime inference volume, and nobody knows that number at report time; worse, the party doing the dividing benefits from guessing high, which hands the industry a denominator it can inflate at will. Holding a buyer accountable for a training run they had no part in is not accountability, it is decoration. And a standard that reports only absolute totals really does collapse to "build less", which no industry body would adopt — and a standard nobody adopts measures nothing at all. Every one of those objections is sound.

They justify keeping training out of the consumer's rate. They do not justify leaving the total undefined and unowned. Nothing prevented this document from requiring providers to publish training emissions as an absolute figure alongside the rate, as a disclosure rather than a score — no forecasting, no amortisation, no gaming, just the number. It chose not to, and the reason given is that absolutes are not actionable for organisations committed to advancing capability.

Everything above is drawn from the foundation's own published documents — the specification, its rationale, and the pull requests currently reshaping both — because they are the only record of these decisions that exists. That is a limit worth stating, and also the point: no outside party has had to check this, and once the ISO grammar is finished, none will be asked to.

This is what makes the piece an architecture story rather than a sustainability one. Architects have spent years learning that what a platform makes countable becomes what it is possible to argue about, and that every metric is a compressed claim about which trade-offs are legitimate. A rate answers one question well: are we getting better at this per unit of it. It cannot answer whether there should be more units. Choose the rate as your only instrument and the second question does not become hard — it becomes unsayable, because there is no field for it in the report, no target to set against it, and no auditor with a mandate to ask.

The same foundation is discovering the shape of this problem in an adjacent metric. Its water handbook, updated a fortnight ago, names the main practical barrier to software water accounting exactly: providers report site-level totals, not workload-level allocation. Carbon has the mirror-image problem. The allocations are specified in fine detail; it is the total that no persona is required to hold.

Which leaves an architect with a genuinely uncomfortable use for this standard. It will tell you, credibly and without offsets, that your inference is cleaner than last year's. It will let a provider tell you their training is more efficient per FLOP than the run before. Both statements can be true while the total rises every quarter, and the document will have nothing to say about it, because it was written not to.

What this is argued from

Reporting and primary material the piece rests on, dated at the time of writing. The interpretation is mine; the facts belong to these.

  1. Software Carbon Intensity for AI Specification Green Software Foundation · 2026-08-18
  2. SCI for AI — Rationale and FAQ Green Software Foundation · 2025-12-08
  3. Pull request 153, Align SCI for AI baseline with ISO/IEC Directives Part 2 Green Software Foundation · 2026-09-24
  4. Existing AI Measurement Metrics (ANCILLARY.md) Green Software Foundation · 2026-08-18
  5. SCI for AI, project README and status Green Software Foundation · 2026-08-18
  6. Software Water Handbook Green Software Foundation · 2026-09-15

Editorials on this site are written to be argued with. If you think the reading is wrong, it probably is in some particular way, and that is the useful part.

carbon accountingmeasurementstandardssustainabilityfunctional units