Comparison
Nobody measures the part you have to file.
Four categories of tool each hold a piece of what AI costs. None of them produces the artifact somebody signs. We compare categories rather than name rivals, because a claim about a competitor cannot be sourced and dated the way we source every coefficient — and it goes stale the moment they ship. Everything below was checked on 31 July 2026.
The map
What each category is actually for.
| Category | Measures | Unit |
|---|---|---|
| LLM observability | Cost, latency, output quality | Per call |
| AI FinOps / cloud cost | Spend, by team or cost centre | Per call |
| LLM carbon libraries | Energy, carbon, and now water | Per call |
| Enterprise carbon platforms | Emissions across the whole organisation | No token visibility |
| Tetrameter | Energy, carbon, water, land, and cost | Per trace |
Each of those is good at its job. The gap is not quality, it is that four different tools each hold one piece of a number somebody has to sign.
One
The unit is the disagreement.
Every category above measures per call. That was the right unit when one request meant one answer. An agentic system spends tens or hundreds of calls producing a single business outcome, and the Green Software Foundation named this directly — its SCI for AI extension calls it the agentic multiplier.
We measured it in our own product on 2 August 2026: one chat turn fanned out across nine providers under a single outcome. A per-call average of that is not a smaller version of the truth. It is the answer to a question nobody asked.
This is checkable against your own traffic in an afternoon, which is the point. Group your calls by the task that caused them and compare the spread to your per-call average. If they agree, you do not need us.
Two
Five things nobody in the market does.
- A filing-grade artifact. Carbon tools output dashboards. ESG platforms output filings but cannot see a token. The bridge between them does not exist, and a dashboard is not evidence.
- Per-customer attribution. Cost tools attribute to teams and cost centres; carbon platforms attribute to legal entities. Neither answers “what did this customer’s usage of this feature emit?” — which is the question a B2B company gets asked by its own buyers.
- Restatement. When a coefficient changes, every historical figure moves — and in a disclosed inventory that is a restatement event requiring documentation. EcoLogits changed its energy benchmark in 2026 and every number it had ever produced shifted. No tooling exists for versioning AI emissions methodology. Ours is public and machine-readable, including the bug we found in our own engine and restated downward.
- Water and land. Two first-party water sources exist and they differ by roughly 173×. Land is measured by nobody. Both have real regulatory teeth as data-centre siting becomes contested.
- Connecting the estimate to the fix. Measurement tools do not optimise and optimisers do not measure carbon. The loop from “here is the number” to “here is what changing it saves, in money and in grams” is open. Our first case study is that loop closing.
Three
Where the open-source libraries sit.
EcoLogits (opens in a new tab) is the best open methodology for LLM emissions, and it is a methodology ally rather than a rival — a non-profit library that will never build an enterprise layer. It has no dashboard, no attribution, no artifact, no cost side. ML CO₂ Impact (opens in a new tab) and CodeCarbon are training-era tools: CodeCarbon measures hardware you own, which is not how anyone consumes a commercial API.
Our engine is open source too, Apache-2.0, for exactly the reason theirs is: a methodology nobody can check is a methodology nobody should file against. What we sell is the operated product around it — attribution, the artifact, the restatement log, and the waste engine.
If the library is all you need, take it. That is a real answer and we say so at more length.
Four
What we do worse.
- We cannot debug a prompt. No field in our data model can hold a prompt or a completion, and the ingest endpoint rejects a request carrying one. That is deliberate — it is what gets us through security review at banks and hospitals — but it means an LLM observability tool will always beat us at the thing it is for.
- We cannot measure whether an answer got worse. Judging output quality needs the completion text we are built never to hold. We can price every consequence of an optimisation except that one.
- We are not a carbon platform. Scope 1, 2 and the rest of Scope 3 are Watershed’s and Persefoni’s job. We are the AI line item they cannot produce, and we would rather feed them than replace them.
- Tier 2 is our ceiling for third-party APIs. No commercial provider discloses per-request energy, so no tool — ours included — can honestly claim measured power for an API call. Anyone claiming otherwise is estimating and not saying so.
Checking this page.
Category claims were verified against vendors’ own documentation and pricing on 31 July 2026. Products change; if something here is out of date, it is wrong rather than clever, and tell us — we will fix it, the same way we restate a coefficient.
We do not publish head-to-head pages against named products. A claim we cannot source and date to the standard we hold our own numbers to is not one we should be making.