By checking twelve points before accepting the conclusion. In the vast majority of cases, the gap claimed between two providers comes from their measurement conventions, not from their efficiency.
The textbook case is everywhere: Google publishes 0.03 gCO2e per request, Mistral 1.14 gCO2e. A factor of 38. Mostly it measures a difference in convention.
The grid, thirteen axes
Bring it out when a client turns up with two numbers and a conclusion. Each axis points to a chapter: boundary, carbon convention, architecture, location.
| # | Axis | Documented magnitude |
|---|---|---|
| 1 | Functional unit | "median prompt" against "400-token answer" against "average request". Not convertible |
| 2 | Median or mean | Google publishes a median and justifies the choice by the skew of the distribution. OpenAI reports a mean. Google's mean is never published |
| 3 | Measurement boundary | 1.72 for the chip-to-production boundary alone, 2.4 including the sample effect |
| 4 | Spread on a single model | several published values coexist for the same model, none of them wrong. MISSING DATA: the range in circulation, 580 to 3,600 prompts per kWh on Llama 3.1 70B, has no identified primary source. Axis 6 gives the sourceable gap |
| 5 | Carbon convention | on the electricity factor: 345 against 94 gCO₂e/kWh at Google, so 3.7. On a whole-company inventory: 4,394 at Meta in 2024, environmental report |
| 6 | Load factor and batching | a factor of around 35 between an unbatched measurement and a production measurement: 1.72 Wh per request on Llama-3-70B at AI Energy Score against 0.0486 Wh on Llama 3.3 70B at ML.Energy v3.0 |
| 7 | Site sample | the most efficient 10% of sites against a fleet average: around 1.4 |
| 8 | Architecture | dense against sparse: Google attributes 10 to 100 to the move to Mixture-of-Experts |
| 9 | Reference year | a 33-fold reduction in twelve months at Google. A 2024 number is not comparable with a 2025 number |
| 10 | Training in or out | Google excludes it, Mistral publishes it separately, nobody amortises it |
| 11 | Task type | the IEA writes that video generation, reasoning and agents "can consume hundreds or thousands of times more energy per query than simple text generation" |
| 12 | Location | 126 between two data centres of the same operator, on Meta 2024 annual inventories, not on intensities. The ratio between grids is 14 to 21 between France and Poland |
| 13 | Intensity or total | NVIDIA reports −24% embodied emissions between HGX H100 and B200. That is per PFLOPS. Per board it is +73.3% |
The Google against Mistral case, worked through
| Mistral | ||
|---|---|---|
| Unit | median text prompt, length not published | 400-token answer |
| Boundary | inference only, training excluded | life cycle, hardware manufacturing included |
| Convention | market-based | location-based |
| Water | on site only | upstream included |
| Quantity published | energy and carbon | carbon, water, resources, never energy |
Google publishes an energy figure without saying how many tokens, Mistral publishes impacts without saying how much energy. The two most-quoted values in the field share no common quantity at all.
And the factor of 38 collapses as soon as you correct a single convention. Google itself publishes both factors in the paper carrying the per-prompt number: 345 gCO2e/kWh location-based for 2024, against a net market-based 94. So the electricity factor is multiplied by 3.7. But the hardware manufacturing term does not move: the per-prompt total goes from about 0.03 to about 0.09 gCO2e, three times more, without a single line of code changing.
The denominator trap, worth knowing by heart
The thirteenth axis deserves unfolding, because it is the hardest to see and the easiest to correct.
A footprint announced "per PFLOPS" is an intensity, not a total. It can fall while absolute consumption rises, and both statements are then true at the same time. In the NVIDIA case, −24% per unit of compute and +73.3% per board coexist without contradiction. If the number of boards sold grows by more than 23%, the total rises.
Same mechanism at Google: energy per request divided by 33 in twelve months, while electricity consumption rose 38% and emissions 18%.
The question to ask takes five words: intensity or total, divided by what? A provider who cannot answer does not know what it is publishing.
The strongest objection to know
It comes from the IEA and targets the reasoning rather than the numbers. It works in both directions.
With a generous average of 1 Wh per request, ten billion requests a day would come to roughly 3.6 TWh a year, under 1% of the 485 TWh consumed by data centres.
The agency draws the conclusion that matters: if large-scale text inference accounts for such a thin slice, then most of the planned capacity is meant for something else. And none of the companies that have published a per-request cost has published how its capacity splits across those uses.
Use it against anyone who infers from a per-request number that AI consumes little. The per-request number is correct. It says nothing about the total, because it does not cover what the capacity is actually being built for.
Question
Provider A reports 0.05 gCO2e per request, provider B 0.60. Which piece of information do you ask for first?
Choisissez une réponse pour voir l'explication.
The regulatory framework settles nothing
Few consultants know this, and it changes the stance you take.
| Text | Status verified on 9 September 2026 | Binding? |
|---|---|---|
| AI Act, annex XI | in force | Yes, on energy consumption alone, with a compute proxy allowed |
| Energy Efficiency Directive, art. 12 | in force, 500 kW threshold | Yes, for data centres |
| French transposition, decree of December 2025 | in force from 1 January 2026 | Yes, fine up to €50,000 per site |
| ISO/IEC 21031 (SCI) | published March 2024 | No, voluntary |
| SCI for AI | ratified November 2025 | No, not yet an ISO standard |
| ITU-T L.1801 | approved February 2026 | No, a recommendation |
| GHG Protocol revised scope 2 | consultation closed, publication coordinated after 2027 | not before 2028 |
The AI Act is worth reading literally. Its annex XI asks for the "known or estimated energy consumption of the model", with this caveat: where consumption is unknown, it "may be based on information about computational resources used".
An obligation on energy, not on emissions. The text imposes no factor and no boundary, and explicitly allows a proxy estimate. And the empowerment meant to harmonise measurement methods, in article 53, has produced no delegated act: so every provider documents its own method.
The form that signatories of the code of practice actually fill in confirms the limit. Its section is titled "Energy consumption (during training and inference)", but inference is documented there in compute only, never in energy. There is no inference energy field at all. And the recipients ticked are the EU AI Office and national authorities, never downstream customers.
So do not promise a client they will get this data from their provider in 2027.
Three conflicts between texts in force
These are not expert disagreements but contradictions written into published texts.
The SCI requires location-based and forbids certificates and purchase agreements by name. The renewable share indicator in EU reporting aggregates them explicitly. Two standards applicable to the same site, incompatible.
Two normative definitions of PUE coexist, and EU law applies the older one.
Hardware amortisation is required by one GHG Protocol text and forbidden by another. The ICT sector guidance asks you to amortise manufacturing emissions over the lifetime. The scope 3 standard forbids it and requires everything to be counted in the year of purchase. The difference is logical, life-cycle assessment on one side and corporate inventory on the other, but no text arbitrates between them.
What a consultant does with this
Since no binding framework imposes a method, the deliverable documents its own boundaries, and that documentation work is the added value, not a preliminary to it.
In practice, on a tender or a provider comparison, the useful clause does not ask for a number. It asks for a declared functional unit, an explicit boundary with justified exclusions, both carbon conventions, and the batch size the measurement was taken at. A provider who cannot answer those four points has not measured, it has estimated, which is acceptable as long as it says so.
FAQ
Can you compare the carbon footprints published by two AI providers?
Rarely as they stand. Thirteen axes of non-comparability are documented, among them the functional unit, the measurement boundary, the carbon convention, the statistic used and the batch size. The factor-of-38 gap often quoted between Google and Mistral comes almost entirely from those conventions, and it falls from 38 to around 13 as soon as you align the carbon convention alone.
Why is Google's number so low compared with the others?
Three choices stack up. It excludes model training. It applies market-based accounting, with a net factor of 94 gCO2e/kWh against 345 location-based. And it covers a median prompt, not a mean, when the distribution is heavily skewed. Recalculated location-based, it goes from about 0.03 to about 0.09 gCO2e.
Does the EU AI Act require providers to publish their models' consumption?
It requires them to document it, not to publish it. Annex XI asks for the known or estimated energy consumption of the model, allows an estimate from compute resources, and that documentation is supplied on request to the EU AI Office and national authorities. The code of practice form contains no inference energy field.
Is there a standard that imposes a method for calculating the footprint of a request?
No, no binding framework imposes one as of 9 September 2026. ISO/IEC 21031 provides a formula but leaves the functional unit free and never mentions inference. ITU-T Recommendation L.1801 is the most complete and remains voluntary. The only binding text, the EU AI Act, standardises neither boundary nor factor.
What should you ask a provider in a tender on this subject?
Four things, not a number: the functional unit used, the measurement boundary with its justified exclusions, both scope 2 carbon conventions, and the batch size the measurement was taken at. A provider unable to answer has estimated rather than measured, which is still acceptable provided it says so.