Methodology · factor set 2026.08.2

How each number is produced, and what it is worth.

We did not invent a carbon methodology. The Green Software Foundation ratified SCI for AI in late 2025 and submitted it to ISO. Our contribution is to be the reference implementation of it, and to be unusually strict about saying which tier of evidence each figure rests on.

The claim we are not making

Limited assurance under ISAE 3410 does not test whether your number is right. It tests whether your methodology is documented, your figures are traceable to their inputs, your factors are tested, and your uncertainty is disclosed. It explicitly tolerates a range.

This is good news, and most of the market has not internalised it. We do not need to produce a precise carbon figure for a model whose operator publishes nothing. We need to produce a defensible one, show our work, and be honest about the width of the band. That is an engineering problem, and it is winnable.

Evidence tiers

Every figure carries the tier it was earned at.

The tier is not our confidence in the arithmetic. It is a statement about what evidence exists in the world. A composite figure is reported at the weakest tier of anything it depends on, because a chain is only as good as its worst link.

TierBasisWhen it applies
1. Class averageThe model's size class, from measured open-weight models of comparable scale.Any proprietary API model. Nobody publishes their internals, so this is the ceiling, not a shortcut.
2. Model measuredA published measurement for this specific model under a known serving configuration.Open-weight models present in the ML.ENERGY Benchmark: 328 configurations across 24 models.
3. Deployment measuredMeasured energy for this model on the hardware and region it actually ran on.Self-hosted inference where you control and can instrument the serving stack.
4. Measured with dynamic gridDeployment measurement combined with the grid's carbon intensity at the hour of the call.Self-hosted inference with a live grid feed. Cost reaches this tier whenever the provider bills an exact amount.

If you send us traffic that is entirely commercial API models, your energy and carbon will report Tier 1 and there is nothing either of us can do about it. We would rather say that on this page than have you discover it in a report.

Where the width comes from

One step is uncertain. The rest are arithmetic.

The most common objection to a range this wide is that it looks like we have not done the work. The chain says otherwise: token counts are exact, the host overhead is reconciled against a published figure, facility PUE is a measured fleet median, and grid intensity is national annual data.

Nearly the whole band enters at a single link, and it is the one nobody outside the provider can close: how much energy their model spends producing a token. Every step after it multiplies that width through rather than adding to it.

This is also why the composite reads Tier 1. Not an average of the tiers, and not the tier of the best-evidenced input. The weakest link, because that is what the figure actually rests on.

0.1×1×10×TokensCounted, not estimatedtier 4countedModel energyWh per 1k output tokenstier 1ML.ENERGY p10–p90Host + idleChip to whole machinetier 12.23×, reconciledFacility PUEMachine to buildingtier 11.09 fleet medianGrid intensityEnergy to carbontier 1Ember annual meancomposite reported at tier 1 — the weakest link, not the average
The width enters at one step. Token counts are exact, and the grid factor, the facility overhead and the host multiplier are all known to within a few per cent. Nearly the entire band comes from the one quantity no provider publishes: how much energy their model spends producing a token.

Grid signal

Average for the inventory. Marginal for the claim.

The average carbon intensity of a grid answers “what share of the grid’s emissions is mine?” That is the question the GHG Protocol asks for an inventory. The marginal intensity answers “what changed because I stopped?” That is the only question a reduction claim is allowed to answer.

These are different numbers and substituting one for the other is the most common error in this market. Published research finds the same intervention reading 18% savings on one signal, 11% on the other, and negative on a third. A vendor who reports one figure for both purposes is producing an artifact of their choice of signal.

We compute both, label which is which, and refuse to let a reduction claim quote an average.

What we measure

Energy at the accelerator, scaled by an explicit host-and-idle overhead and by facility PUE. Carbon from energy and the grid factor for the region. Water from cooling intensity. Land from the generation mix already needed for the carbon figure. Cost from the provider’s billed amount where they give one, and a versioned price catalogue where they do not.

What we do not

Embodied carbon of the hardware, network transit, and the training run that produced the model. Each is real, none is reliably attributable to a single inference, and inventing an allocation would be exactly the kind of confident fiction this page exists to argue against.

Land is the one that surprises people: renewable grids use more land per kilowatt-hour, not less. We publish that result rather than quietly dropping the resource that makes the green answer look worse.

Restatements

When a coefficient changes, every past figure moves.

In a disclosed inventory that is a restatement event, and it requires documentation. No other tool in this market handles it, so we keep the log inside the library itself, so a report can cite it mechanically rather than a human remembering.

Each entry names the factor, the versions either side, the date applied, the reason, and an estimate of how far affected figures move.

A real entry, from this month

model.class.reasoning · 2026.08.1 → 2026.08.2 · applied 2026-08-01 · materiality −0.66

Model names that negate a capability, such as grok-4-1-fast-non-reasoning, matched our reasoning classifier on the very token that denies it, placing a mid-class model on the reasoning energy curve. Affected figures were overstated roughly threefold. Found in the first production customer fleet rather than in review.

Overstating is not the safe direction. A customer who discovers we inflated their number has the same reason to distrust every other figure as one who discovers we shrank it.

Check it yourself

The engine is open, so none of this has to be taken on trust.

Every coefficient, its source, its version and the date we retrieved it are in the library, along with the tests that pin the behaviour, including the ones that assert uncomfortable results so nobody quietly “fixes” them.