Chapter 16 of 17 · 13 min read

How do you take apart a comparison between two AI providers?

In this chapter7 sections
  1. The grid, thirteen axes
  2. The Google against Mistral case, worked through
  3. The strongest objection to know
  4. The regulatory framework settles nothing
  5. Three conflicts between texts in force
  6. What a consultant does with this
  7. FAQ

By checking twelve points before accepting the conclusion. In the vast majority of cases, the gap claimed between two providers comes from their measurement conventions, not from their efficiency.

The textbook case is everywhere: Google publishes 0.03 gCO2e per request, Mistral 1.14 gCO2e. A factor of 38. Mostly it measures a difference in convention.

The grid, thirteen axes

Bring it out when a client turns up with two numbers and a conclusion. Each axis points to a chapter: boundary, carbon convention, architecture, location.

#AxisDocumented magnitude
1Functional unit"median prompt" against "400-token answer" against "average request". Not convertible
2Median or meanGoogle publishes a median and justifies the choice by the skew of the distribution. OpenAI reports a mean. Google's mean is never published
3Measurement boundary1.72 for the chip-to-production boundary alone, 2.4 including the sample effect
4Spread on a single modelseveral published values coexist for the same model, none of them wrong. MISSING DATA: the range in circulation, 580 to 3,600 prompts per kWh on Llama 3.1 70B, has no identified primary source. Axis 6 gives the sourceable gap
5Carbon conventionon the electricity factor: 345 against 94 gCO₂e/kWh at Google, so 3.7. On a whole-company inventory: 4,394 at Meta in 2024, environmental report
6Load factor and batchinga factor of around 35 between an unbatched measurement and a production measurement: 1.72 Wh per request on Llama-3-70B at AI Energy Score against 0.0486 Wh on Llama 3.3 70B at ML.Energy v3.0
7Site samplethe most efficient 10% of sites against a fleet average: around 1.4
8Architecturedense against sparse: Google attributes 10 to 100 to the move to Mixture-of-Experts
9Reference yeara 33-fold reduction in twelve months at Google. A 2024 number is not comparable with a 2025 number
10Training in or outGoogle excludes it, Mistral publishes it separately, nobody amortises it
11Task typethe IEA writes that video generation, reasoning and agents "can consume hundreds or thousands of times more energy per query than simple text generation"
12Location126 between two data centres of the same operator, on Meta 2024 annual inventories, not on intensities. The ratio between grids is 14 to 21 between France and Poland
13Intensity or totalNVIDIA reports −24% embodied emissions between HGX H100 and B200. That is per PFLOPS. Per board it is +73.3%

The Google against Mistral case, worked through

GoogleMistral
Unitmedian text prompt, length not published400-token answer
Boundaryinference only, training excludedlife cycle, hardware manufacturing included
Conventionmarket-basedlocation-based
Wateron site onlyupstream included
Quantity publishedenergy and carboncarbon, water, resources, never energy

Google publishes an energy figure without saying how many tokens, Mistral publishes impacts without saying how much energy. The two most-quoted values in the field share no common quantity at all.

And the factor of 38 collapses as soon as you correct a single convention. Google itself publishes both factors in the paper carrying the per-prompt number: 345 gCO2e/kWh location-based for 2024, against a net market-based 94. So the electricity factor is multiplied by 3.7. But the hardware manufacturing term does not move: the per-prompt total goes from about 0.03 to about 0.09 gCO2e, three times more, without a single line of code changing.

The factor-of-38 gap between Google's number and Mistral's comes from conventions, not efficiency

The denominator trap, worth knowing by heart

The thirteenth axis deserves unfolding, because it is the hardest to see and the easiest to correct.

A footprint announced "per PFLOPS" is an intensity, not a total. It can fall while absolute consumption rises, and both statements are then true at the same time. In the NVIDIA case, −24% per unit of compute and +73.3% per board coexist without contradiction. If the number of boards sold grows by more than 23%, the total rises.

Same mechanism at Google: energy per request divided by 33 in twelve months, while electricity consumption rose 38% and emissions 18%.

The question to ask takes five words: intensity or total, divided by what? A provider who cannot answer does not know what it is publishing.

The strongest objection to know

It comes from the IEA and targets the reasoning rather than the numbers. It works in both directions.

With a generous average of 1 Wh per request, ten billion requests a day would come to roughly 3.6 TWh a year, under 1% of the 485 TWh consumed by data centres.

The agency draws the conclusion that matters: if large-scale text inference accounts for such a thin slice, then most of the planned capacity is meant for something else. And none of the companies that have published a per-request cost has published how its capacity splits across those uses.

Use it against anyone who infers from a per-request number that AI consumes little. The per-request number is correct. It says nothing about the total, because it does not cover what the capacity is actually being built for.

Question

Provider A reports 0.05 gCO2e per request, provider B 0.60. Which piece of information do you ask for first?

Choisissez une réponse pour voir l'explication.

The regulatory framework settles nothing

Few consultants know this, and it changes the stance you take.

TextStatus verified on 9 September 2026Binding?
AI Act, annex XIin forceYes, on energy consumption alone, with a compute proxy allowed
Energy Efficiency Directive, art. 12in force, 500 kW thresholdYes, for data centres
French transposition, decree of December 2025in force from 1 January 2026Yes, fine up to €50,000 per site
ISO/IEC 21031 (SCI)published March 2024No, voluntary
SCI for AIratified November 2025No, not yet an ISO standard
ITU-T L.1801approved February 2026No, a recommendation
GHG Protocol revised scope 2consultation closed, publication coordinated after 2027not before 2028

The AI Act is worth reading literally. Its annex XI asks for the "known or estimated energy consumption of the model", with this caveat: where consumption is unknown, it "may be based on information about computational resources used".

An obligation on energy, not on emissions. The text imposes no factor and no boundary, and explicitly allows a proxy estimate. And the empowerment meant to harmonise measurement methods, in article 53, has produced no delegated act: so every provider documents its own method.

The form that signatories of the code of practice actually fill in confirms the limit. Its section is titled "Energy consumption (during training and inference)", but inference is documented there in compute only, never in energy. There is no inference energy field at all. And the recipients ticked are the EU AI Office and national authorities, never downstream customers.

So do not promise a client they will get this data from their provider in 2027.

Three conflicts between texts in force

These are not expert disagreements but contradictions written into published texts.

The SCI requires location-based and forbids certificates and purchase agreements by name. The renewable share indicator in EU reporting aggregates them explicitly. Two standards applicable to the same site, incompatible.

Two normative definitions of PUE coexist, and EU law applies the older one.

Hardware amortisation is required by one GHG Protocol text and forbidden by another. The ICT sector guidance asks you to amortise manufacturing emissions over the lifetime. The scope 3 standard forbids it and requires everything to be counted in the year of purchase. The difference is logical, life-cycle assessment on one side and corporate inventory on the other, but no text arbitrates between them.

What a consultant does with this

Since no binding framework imposes a method, the deliverable documents its own boundaries, and that documentation work is the added value, not a preliminary to it.

In practice, on a tender or a provider comparison, the useful clause does not ask for a number. It asks for a declared functional unit, an explicit boundary with justified exclusions, both carbon conventions, and the batch size the measurement was taken at. A provider who cannot answer those four points has not measured, it has estimated, which is acceptable as long as it says so.

FAQ

Can you compare the carbon footprints published by two AI providers?

Rarely as they stand. Thirteen axes of non-comparability are documented, among them the functional unit, the measurement boundary, the carbon convention, the statistic used and the batch size. The factor-of-38 gap often quoted between Google and Mistral comes almost entirely from those conventions, and it falls from 38 to around 13 as soon as you align the carbon convention alone.

Why is Google's number so low compared with the others?

Three choices stack up. It excludes model training. It applies market-based accounting, with a net factor of 94 gCO2e/kWh against 345 location-based. And it covers a median prompt, not a mean, when the distribution is heavily skewed. Recalculated location-based, it goes from about 0.03 to about 0.09 gCO2e.

Does the EU AI Act require providers to publish their models' consumption?

It requires them to document it, not to publish it. Annex XI asks for the known or estimated energy consumption of the model, allows an estimate from compute resources, and that documentation is supplied on request to the EU AI Office and national authorities. The code of practice form contains no inference energy field.

Is there a standard that imposes a method for calculating the footprint of a request?

No, no binding framework imposes one as of 9 September 2026. ISO/IEC 21031 provides a formula but leaves the functional unit free and never mentions inference. ITU-T Recommendation L.1801 is the most complete and remains voluntary. The only binding text, the EU AI Act, standardises neither boundary nor factor.

What should you ask a provider in a tender on this subject?

Four things, not a number: the functional unit used, the measurement boundary with its justified exclusions, both scope 2 carbon conventions, and the batch size the measurement was taken at. A provider unable to answer has estimated rather than measured, which is still acceptable provided it says so.

Back to the course