GSF SCI for AI · reference implementation

Four measures of what your AI actually costs.

Energy, carbon, water and land for every call your system makes, with cost sitting outside them as the same consumption, priced. Measured per task, not per query, because an agent does not make one request.

01
Energy
kWh
02
Carbon
gCO₂e
03
Water
litres
04
Land
cm²
·
Cost
the same four, priced

Tetra, four. Meter, measure. The name is the specification: energy, carbon, water and land, with cost outside them as the same consumption priced.

It counts twice. A tetrameter is also a line of verse in four metrical feet — four beats, which is what the mark draws.

LLM observability measures one thing. Carbon tools measure one or two. The largest ESG platform measures three. The number in the name is the difference, and it is not one a competitor can adopt without rebuilding their product.

One report, measured

We show you the error bar.

This is a real recognition report from SiteBeacon (opens in a new tab): 26 model calls across five providers, in one trace, on 1 August 2026. Nothing here is illustrative. The figures come from the same engine the product ships.

Notice that cost is exact and carbon is not. The provider bills us to the hundredth of a cent, so cost is Tier 4 evidence with no band at all. Nobody publishes the energy draw of a proprietary model, so carbon spans two orders of magnitude and says Tier 1 in the corner.

That gap is the product. A tool that printed a single tidy carbon number for the same call would not be more accurate. It would be less honest.

Cost$0.0396
exactno band
central 0.0396 USDtier 4
Carbon0.162 – 18 gCO₂e
0.16218
central 1.09 gCO2etier 1

1.09 gCO₂e is what a single-number tool prints here.

Derived from the same energy figureRangeTier
Energy0.0004 – 0.026 kWh1
Water0.0004 – 0.11 L1
Land0.043 – 30.8 cm²1

These three are one measurement in three units. Each is energy multiplied by a published constant, so they share a band rather than confirming each other.

Tier 1 is a class average. Tier 4 is measured. The tier is not our confidence. It is a statement about what evidence exists, and provider disclosure is the binding constraint on three of these five.

The unit

Measuring every call tells you nothing.

The reasonable objection to trace-level measurement is that per-call already covers everything. It does — and that is the problem. Fifteen requests behind a contract review look exactly like the two behind a document extraction. All ordinary, none remarkable.

One of those tasks costs nineteen times the other. That fact does not exist at the level of a call. It only appears once you group the calls by the thing the business actually asked for.

Which is why a per-query average is not a smaller version of the truth. It is a different number, and it gets further from the truth the more agentic your system becomes.

010k20k30k40k50kPer query17 calls, one dot eachevery call between 1.3k and 3.8k — nothing stands outPer task2 outcomes50,693one contract reviewed2,720one document processed · 19× apartsame tokens, both rows — only the unit changes
Two real traces from the demo organisation. Per call, a contract review is fifteen ordinary requests and a document extraction is two — nothing in the top row suggests one is worth nineteen of the other. That fact only exists at the level of the task, which is why the trace is the unit.

The number nobody shows you

35% of the tokens in that report were billed but never reached a human.

5,416 of 15,289 tokens went to intermediate reasoning, discarded candidates and retries. You paid for all of them and none appeared in the output. Per-query dashboards cannot see this, because the waste lives in the space between the calls, and that space is exactly what a trace is.

Invisible tokens are not a failure. Reasoning is often worth buying. But you cannot decide whether it is worth buying until somebody counts it.

See it working

The dashboard, and the pack an auditor gets.

A demonstration organisation is open without a login: four weeks of traffic, five customers, five models, and the waste findings the detectors actually produced from it.

The traffic is generated rather than measured, and the page says so at the top. Our real capture carries a live customer’s account id, and a product that promises not to leak identifiers should not publish one to make a sales point. Every figure on it is computed by the engine your own data would run through.

Demo Corp · four weeks
Traces254
Calls868
Tokens2,712,503
Billed but never surfaced300,982 · 11%
Waste findings43
Recoverable$0.22 – 0.78

Findings across all four detectors: planning loops, context resend, redundant calls, and reasoning spent on trivial answers.

Seven differences

What this does that the alternatives do not.

LLM observability measures one thing: cost. Carbon tools measure one or two. The largest ESG platform on the market measures three, and none of them can see a token.

The trace is the unit, not the call

Every competitor measures per query. An agentic workload makes one business outcome cost hundreds of times a chat turn, so a per-query average is not a smaller version of the truth; it is a different number. We measure the task.

Our biggest structural edge

Agentic waste, priced

Runaway planning loops, redundant tool calls, quadratic context re-send, multi-agent debate that changed nothing, retries without backoff, reasoning spend on trivial inputs, and provider-billed tokens you cannot inspect. Each one costed in money, grams and litres.

Average and marginal grid signal, used correctly

Average for the inventory, because that is what the GHG Protocol asks for. Marginal for every reduction claim, because that is what is physically true. Published research shows the same intervention reads 18% savings on one signal and 11% on the other, and goes negative measured a third way. Most claims in this market are artifacts of the vendor's chosen signal.

Optimisation you can prove is safe

We shadow-evaluate the cheaper path against your real traffic and report the measured quality delta before recommending a switch. Nobody captures the model-choice lever because of fear of regression, not ignorance. We sell certainty, not insight.

Per-customer attribution

You can hand your customer a real AI carbon number for their account. Structurally hard for cost tools, which think in cost centres, and impossible for ESG platforms, which never see a token.

The auditor's evidence pack

A methodology statement, per-number data lineage, tier classification, uncertainty ranges, every emission factor with its source, version and retrieval date, and a restatement log for when a coefficient changes. Not a feature but a discipline, and the hardest thing here to copy.

The hardest thing here to copy

Metadata only, structurally

No field in our data model can hold a prompt or a completion, and the ingest endpoint rejects a request carrying one rather than dropping it quietly. That opens banks, health, legal and government: buyers that prompt-capturing tools cannot serve at any price.

Where this sits

Between the dashboard and the disclosure.

There is roughly a hundred-fold price gap between LLM observability and enterprise carbon accounting. We serve the budget at the top of that range and cost a fraction of it.

CategoryMeasuresUnitTypical price
LLM observabilityCost, latency, qualityPer call$29 – 2,499 / mo
AI FinOpsCostPer call$30 – 200 / mo, or 1% of spend
Enterprise carbon platformsEmissions across the orgNo token visibility$50,000 – 250,000 / yr
TetrameterEnergy, carbon, water, land + costPer trace$499 / mo – $60,000 / yr

Prices as published by each vendor, July 2026. Ours are on the pricing page in full, because a product that argues for disclosure should start with its own.

Design partners, 2026

California wants Scope 3 with limited assurance from 2027.

A measurement product takes about nine months to become credible enough to sit in front of an auditor. We are taking a small number of design partners who want to arrive at that deadline with their AI footprint already documented.