Per-Request Energy, Water and Carbon
How to put an energy, water and carbon figure on a single inference request, why the two most rigorous published per-request figures differ by a factor of 38 on carbon and 173 on water, and which accounting choices produce the gap.
Two vendors have now published serious environmental figures for one inference request, and they do not agree. Google reports that a median Gemini Apps text prompt consumes 0.24 Wh of energy, emits 0.03 gCO2e and consumes 0.26 mL of water (Google Cloud, 2025, Measuring the environmental impact of AI inference). Mistral, in a life-cycle assessment run with ADEME, Carbone 4, Resilio and Hubblo, reports that a 400-token response from Le Chat carries a marginal impact of 1.14 gCO2e and 45 mL of water (DeepLearning.AI, 2025, French AI Startup Discloses Full Lifecycle Consumption and Emissions for Mistral Large 2). That is 38 times the carbon and 173 times the water for one unit of the same apparent thing.
Neither number is wrong. They answer different questions, and the difference between the questions is the whole subject. Energy and power as the binding constraint treats electricity as a capacity problem at the site. This concept treats it as a per-request accounting problem at the feature, which is where a product team is asked to report it.
What a per-request figure is made of
Energy per request decomposes into what the accelerator draws while computing, what the host machine around it draws, the share of idle reserved capacity the request must carry, and the facility overhead on all of it:
Google published its own figure under two boundaries, which makes the sensitivity legible. Counting only active accelerator draw gives 0.10 Wh, 0.02 gCO2e and 0.12 mL. Adding host CPU and memory, idle machines held in reserve for availability, and data-centre overhead gives 0.24 Wh, 0.03 gCO2e and 0.26 mL. The boundary alone moves energy by 2.4 times.
Carbon is then energy multiplied by a grid intensity, and there are two defensible intensities. The GHG Protocol requires dual reporting: a location-based figure using the average intensity of the local grid, and a market-based figure reflecting the contractual instruments the operator bought (GHG Protocol, Scope 2 Guidance). A market-based figure backed by annual renewable certificates can be a small fraction of the location-based one for the same electrons, which is why a reported per-prompt gCO2e says as much about procurement as about engineering. A proposed revision would require hourly matching and restrict certificates to deliverable sources, with final publication expected in 2027; it would raise many reported figures without changing a single watt.
Water splits the same way: on-site evaporative cooling, measured as water usage effectiveness, plus the off-site water consumed generating the electricity. Reporting only the first understates the total and rewards designs that trade water for a lower PUE.
Why the two published figures disagree
Five choices explain most of the 38-fold gap, and it is worth naming them separately rather than treating the numbers as contradictory.
Google's unit is a median short text prompt; Mistral's is a 400-token response, a longer generation on a larger dense model. Google's inference paper covers serving and does not amortise training into the prompt; Mistral's assessment is a full life cycle including data-centre construction and hardware manufacture, and reports that training and inference together account for 85.5 percent of the model's greenhouse gases and 91 percent of its water. The grids differ. The water boundaries differ. And Google's figures are point-in-time May 2025 data, self-reported and not independently verified, while Mistral's follow ISO 14040/44 and the GHG Protocol Product Standard with third-party involvement.
Attaching it to a feature
The accounting is then multiplication, and the choice of source figure dominates the answer. A feature serving 5 million requests a month at Google's comprehensive figure consumes 1.2 MWh, 150 kg CO2e and 1.3 m³ of water per month. The same volume at Mistral's per-response figure is 5.7 tonnes CO2e and 225 m³. Reporting the first as though it settled the question is a choice, not a measurement, and the honest form of the disclosure names the model, the boundary and the grid method alongside the number.
When it breaks
A median is not a mean, and agents are not medians. Google's figure is explicitly for a median text prompt. A reasoning or agentic request emits far more tokens and resends context on every turn, so a per-prompt constant applied to an agent workload can understate it by an order of magnitude. Meter your own token volumes and scale from an energy-per-token estimate instead.
Marginal figures invite bad comparisons. "Less than nine seconds of television" is true of one prompt and irrelevant to an aggregate that is measured in quadrillions of tokens a month. Marginal intensity falling while total consumption rises is the normal shape of an efficiency story, not a contradiction of it.
Embodied carbon is usually missing. Inference-only accounting omits the manufacture of the accelerators. Mistral's assessment puts 29 percent of materials consumption outside training and inference, which is the part an inference-boundary figure cannot see.
Water and carbon are local and seasonal. A cubic metre in a water-stressed basin is not equivalent to one elsewhere, and grid intensity varies by hour. Annual averages hide both, and hourly accounting changes the ranking of otherwise identical deployments.
References and further reading
Every source this page cites, in the order it cites them. All of them open in a new tab.
6 flashcards for this concept
Click a card to reveal the answer.