Emissions are estimated from token counts using per-token factors derived from published academic research. The approach is intentionally simple: one number per model family per token direction. No real-time data, no per-request tracing.
Jegham et al. (2025), "How Hungry is AI? Benchmarking the Energy, Water, and Carbon Footprint of LLM Inference" (v6, arXiv preprint, not peer-reviewed) arxiv.org/abs/2505.09598
The paper estimates inference energy for a range of models from public API performance data (time to first token, tokens per second) via Monte Carlo simulation over inferred hardware configurations, then converts to CO2e using grid-average carbon intensity. It does not measure hardware power draw: its per-query energies are model outputs resting on assumed GPU nodes and fixed utilization rates, not telemetry. For Claude 3.7 Sonnet it reports three per-query energies (Table 4, PUE included): 0.950 Wh (100 input / 300 output), 2.989 Wh (1k / 1k) and 5.671 Wh (10k input / 1.5k output).
The three estimated points are fit with an ordinary least-squares regression through the origin in (energy per input token, energy per output token), matching the per-token model below (no per-query constant). That gives ~1.35e-4 Wh per input token and ~2.88e-3 Wh per output token. At CIF 0.287 gCO2e/Wh (= 0.287 kgCO2e/kWh) that is ~39 gCO2e/Mtok input and ~826 gCO2e/Mtok output, a marginal ratio of ~21:1. The fit matches the two high-token configurations (1k/1k, 10k/1.5k) to within 1%, the regime real Claude Code sessions live in.
These are the best available public data to date and will keep being refined as estimates improve. Earlier releases used 190/1140, calibrated against an earlier revision of the same preprint; the move to v6's three per-query energies refines that.
EcoLogits (GenAI Impact / Data For Good) estimates inference impact by an independent route: model parameter counts plus a full lifecycle model, rather than API-derived energy estimates. Its estimate for Claude 3.7 Sonnet brackets ~565-1385 gCO2e/Mtok output (parameter range, USA grid, full lifecycle), with a usage-only central value near ~630. The Jegham-derived 826 sits inside that band. That is a useful consistency check rather than strong corroboration: the EcoLogits band spans a factor of ~2.5 (almost any plausible value lands inside it), and its parameter counts for closed models are educated guesses. Note one structural difference: EcoLogits charges zero energy to input tokens, whereas Jegham's estimates show a small but real input cost (the 10k-input query consumes more), which this tool keeps via the input factor.
Simon Couch's January 2026 analysis of AI coding agents derives Claude per-token energy by yet another route: Epoch AI's GPT-4o per-query estimate scaled by Anthropic's API price ratios. His self-described pessimistic figures are ~390 Wh/Mtok input and ~1,950 Wh/Mtok output. The Jegham-derived Sonnet factors used here are equivalent to ~136 Wh/Mtok input and ~2,880 Wh/Mtok output. Three routes (API-based energy estimation, parameter-count lifecycle modeling, price-ratio scaling) landing within ~1.5x of each other on output tokens is encouraging, with caveats. Couch's figures are anchored on Opus-class pricing while the comparison here is against Sonnet factors. The routes share assumptions: price-ratio scaling is the same proxy this tool uses for Opus and Fable, and Epoch AI's anchor is itself a parameter-count model, so the three are not fully independent. And the agreement does not extend to input tokens, where the routes span 0 (EcoLogits) to ~390 Wh/Mtok (Couch).
session_co2_grams = (
(input_tokens + cache_write_tokens) * input_factor
+ cache_read_tokens * (input_factor * cache_read_factor)
+ output_tokens * output_factor
) / 1_000_000
Factors are in gCO2e per million tokens. cache_write_tokens (cache_creation_input_tokens) are a full prefill, so they count at the input factor. cache_read_tokens count at a reduced cache_read_factor (default 0.08) of the input factor (see Cache read energy below).
| Parameter | Value | Description |
|---|---|---|
| PUE | 1.14 | AWS datacenter power usage effectiveness |
| CIF | 0.287 kgCO2e/kWh | AWS region grid CIF, location-based (Jegham et al.); US average is ~380 g/kWh |
| WUE (on-site) | 0.18 L/kWh | Water for datacenter cooling (drives water_ml, on-site term) |
| WUE (off-site) | 5.11 L/kWh | Water for electricity generation (drives water_ml, off-site term) |
| Model family | Input | Output | Source |
|---|---|---|---|
| Fable | 156 | 3304 | Extrapolated (2x Opus) |
| Opus | 78 | 1652 | Extrapolated (2x Sonnet) |
| Sonnet | 39 | 826 | 3-point fit (Jegham et al. v6) |
| Haiku | 20 | 413 | Extrapolated (0.5x Sonnet) |
Output tokens are far more energy-intensive per token than input tokens, but not because they need more math: FLOPs per token are similar in prefill and decode (~2x parameters either way for a dense model). The difference is hardware utilization. Prefill processes the whole prompt in parallel and saturates compute; each decode step re-reads the model weights and the KV cache to produce a single token, so decoding is memory-bandwidth-bound and burns far more energy per token. The exact multiple depends on batch size and serving stack.
The ratio used here is not assumed, it is recovered from the three Jegham Sonnet points (see Deriving the Sonnet input/output factors above): the fit yields a marginal output:input ratio of ~21:1 (826 vs 39 gCO2e/Mtok). Three points and two coefficients leave one degree of freedom, so 21:1 is a point estimate without error bars: the direction is robust, the exact multiple is not. A large input (long context) adds little energy relative to the same number of generated tokens, which a flat low ratio would miss.
The Jegham paper covers Sonnet-class models directly. The other families are estimated by scaling:
- Opus = 2x Sonnet. The current EcoLogits parameter assumptions for Opus 4.5+ (670B vs Sonnet 4.x 440B, active ~133B vs ~88B) and the Anthropic list-price ratio ($5/$25 vs $3/$15) both imply roughly 1.7-2x, not the 3x used in earlier releases. Honest band: 2x-5x Sonnet (Opus is unmeasured; EcoLogits' absolute Opus number is unstable across model generations).
- Haiku = 0.5x Sonnet (smaller model, lighter compute). Wide band: Jegham's 3.5 Haiku estimate reads higher than Sonnet (a serving/latency artifact), while EcoLogits' modern dense Haiku reads far lower; 0.5x is a physically-plausible middle. Taken at face value instead, the Haiku estimate would put Haiku ≈ Sonnet and erase most of the gain from routing tasks to smaller models; the artifact reading is better supported, but the unfavorable reading exists and this is the widest band in the tool.
- Fable = 2x Opus (no published measurement for Fable 5 / Mythos 5; the list-price ratio, $10/$50 vs $5/$25, is used as a compute proxy)
These are order-of-magnitude estimates. Actual values depend on Anthropic's specific hardware configuration and batching strategies, which are not publicly available. Only Sonnet is covered by the Jegham estimates; the others carry the uncertainty bands noted above.
Claude Code can be pointed at non-Anthropic models (e.g. local models behind ANTHROPIC_BASE_URL). Their impact profile is not an AWS datacenter's, so neither the emission factors nor the API pricing apply. Sessions whose dominant model string does not contain claude (including the <synthetic> marker) are stored with their raw token counts but cost_usd = 0, co2_grams = 0 and excluded = 1, and are left out of all report aggregates. Additional models can be excluded by name via the exclude_models patterns in data/factors.json. Because raw tokens are preserved, excluded sessions can be re-priced later by recompute.sh if factors for local models are ever added.
Token counts come from parsing the JSONL transcripts (message.usage). Assistant messages are deduplicated by (message.id, requestId), keeping the last occurrence, before summing. This matters because resumed and compacted sessions replay earlier messages within the same file, and streaming writes the same message multiple times with a growing output_tokens. Without dedup the raw line sum over-counts by roughly 3x (on observed data, 55% of assistant lines are replays). This matches the deduplication ccusage performs.
Claude Code purges JSONL transcripts after about 30 days, so the SQLite DB is the only durable record. Two design choices follow:
-
Capture before purge. The
Stophook (persist-session.sh) writes each session to the DB when it ends, while the JSONL still exists. A throttledSessionStarthook (safety-rescan.sh) re-runsbackfill.shonce a day in the background to catch any session theStophook missed (crash, kill, hook disabled), as long as its transcript is still within the 30-day window. The only unavoidable gaps are history older than the install date and downtime longer than 30 days. -
Store raw tokens, derive on demand. Each row stores the raw token breakdown:
input_tokens(regular input + cache write),cache_creation_tokens(cache write),cache_read_tokens, andoutput_tokens. Cost and CO2 are pure functions of these counts plusdata/factors.jsonanddata/prices.json, so they can be regenerated at any time withrecompute.shwithout re-reading the (purged) JSONL. When a CO2 factor is revised, runrecompute.sh(CO2-only by default); when a price changes, runrecompute.sh --with-cost(cost re-pricing collapses mixed-model rows to the dominant model, ~6% high on subagent sessions, so it is opt-in). Rows are taggedmethodology_version; only version >= 2 carries the full raw-token breakdown, so older rows captured before this change are left untouched as legacy.
recompute.sh recomputes a mixed-model session (subagents on a different model) at the row's dominant model, a small approximation; the original insert is model-accurate per subagent.
A cache_read token is a previously-processed context token whose key/value tensors are reused, so its prefill compute is skipped. It is not free in energy: during decode, every generated token re-reads the entire KV cache from HBM, including the cached tokens (GreenCache, SIGMETRICS: "caching does not reduce computation in the decode phase"). So the energy of a cached token is the decode-phase KV-read residual that survives caching.
No study directly measures the cache_read-to-input energy ratio. The default cache_read_factor of 0.08 (defensible range 0.05-0.15, hard bound 0-0.20) is an engineering estimate derived from adjacent measurements: prefill is ≤ 3.4% of total inference energy for generation workloads and a larger KV cache amplifies per-token decode energy by 1.3-51.8% (both Solovyeva & Castor, measured on 3B-7B models with short prompts), and per-token energy rises ~3x from 2K to 10K context (TokenPowerBench, H100, Llama3 70B). The factor is workload-dependent and grows with context length; a flat constant understates very long reused prefixes.
This factor is not Anthropic's 0.1x cache_read billing ratio. That is a price, not an energy measurement (OpenAI prices the same mechanism at 0.5x). Setting cache_read_factor to 0 is a defensible lower bound but treats a reused 100K-token system prompt as carbon-free, which understates a real memory-bandwidth cost.
Sources: GreenCache (arXiv:2505.23970, SIGMETRICS 2026), TokenPowerBench (arXiv:2512.03024), Solovyeva & Castor (arXiv:2602.05712).
The cost_usd column is the theoretical API list value of the usage (what it would cost on pay-as-you-go), not the subscription price actually paid. It uses current Anthropic list pricing per million tokens (reconfirmed 2026-06-22): Opus 4.6+ at $5 input / $25 output (not the retired $15/$75 of Opus 4.0/4.1), Sonnet at $3/$15, Haiku at $1/$5, Fable 5 at $10/$50. Cache write is billed at 1.25x input (the 5-minute tier, Claude Code's default; the 1-hour tier is 2x) and cache read at 0.1x input. On deduplicated data this reconciles to within a few percent of ccusage.
For EUR, data/prices.json carries eur_per_usd (ECB euro reference rate, 0.8729 as of 2026-06-22); convert cost_usd at display time rather than storing EUR amounts, so a single dated rate stays the only source of truth and historical rows never need re-conversion. The recent USD/EUR range is tight (~±1%), so a monthly refresh keeps EUR accurate well within the CO2/cost uncertainty.
- Order of magnitude only. Do not use these numbers for regulatory reporting or lifecycle assessments. Reports show a single point estimate by design; combining the documented bands (Opus 2x-5x, cache read 0-0.20, the low-end CIF, usage-only scope), the plausible range around any displayed total is roughly 0.7x to 3x, and the excluded items (embodied hardware, amortized training) make the true location-based value more likely above the displayed figure than below it.
- The primary source is an arXiv preprint, and its method has been publicly criticized: a November 2025 analysis argues the paper conflates API speed with hardware efficiency, ignores queueing, and that its absolute values are best used as a relative benchmark between models. The cross-checks above (EcoLogits, Couch) mitigate this but do not remove it.
- Inference only. Training costs, hardware manufacturing, and cooling water are not included.
- Cache read energy is a derived estimate, not a measurement (see Cache read energy below). Cache reads are 90%+ of tokens in Claude Code, and across its hard bound (0-0.20) the chosen factor (default 0.08) swings the headline number from ~0.65x to ~1.5x, the widest single-parameter lever in the tool. At the default, output tokens still dominate the CO2 total.
- Status line is approximate. Claude Code does not expose
cache_read_input_tokensseparately in the statusline hook JSON, and parsing JSONL incrementally at each turn would be too slow. The live display usescontext_window.total_input_tokens(current context size, includes cache reads, no subagents). This is not used in reports. - Grid-average, not real-time. The CIF is the static AWS region grid intensity (location-based, 0.287); the US national average is higher (~380 g/kWh). Actual emissions depend on Anthropic's datacenter location, energy mix, and time of day.
- Single-fleet assumption. Since 2026 Anthropic serves Claude from a mixed fleet: AWS (Trainium and GPUs, the infrastructure Jegham measured), Google Cloud TPUs (1+ GW coming online during 2026), and, per SpaceX's May 2026 S-1 filing, the leased ~300 MW Colossus 1 site in Memphis, largely powered by gas turbines (~350-450 gCO2e/kWh vs the 287 used here). A single CIF and a single per-model energy cannot capture per-request routing across hardware and grids. The weighted effect of Memphis alone is roughly +5-10% on the CIF, within this tool's order-of-magnitude uncertainty; the TPU share pulls the other way. Watch item: revisit the CIF if the fleet mix shifts further toward gas-powered capacity.
| Activity | Emission factor | Source |
|---|---|---|
| Car (thermal, lifecycle) | 142 gCO2e/km | ADEME Impact CO2, 2025 car footprint model |
| Median Gemini prompt | 0.03 gCO2e | Google, Aug 2025 (arXiv:2508.15734) |
| TGV (full scope) | 3.5 gCO2e/km | SNCF open data 2024, per passenger-km |
Car and TGV factors are lifecycle values while this tool's CO2 is usage-only; the equivalences are illustrative, not scope-matched. Refreshed 2026-07-31: earlier releases used 120 g/km (car, close to the EU new-vehicle homologation average, misattributed to ADEME), 0.2 g per Google search (a 2009 blog figure misattributed to the 2023 environmental report), 19 g per email with attachment (ADEME 2011, withdrawn; the current ADEME reference email is ~2.5 g) and 2.4 g/km (TGV, untraceable).
Every session row additionally carries energy_wh, water_ml, and
embodied_gco2e, derived from co2_grams by global constants
(data/factors.json .physics, each value cited there):
energy_wh = co2_grams / 0.287 # CIF identity (exact)
water_ml = energy_wh * (0.18/1.14 + 5.11)
embodied_gco2e = energy_wh * 44.1 / 1000
- CIF identity. The Jegham-derived co2 factors are facility energy times the grid CIF (0.287 g/Wh), so dividing by the CIF recovers facility energy in Wh exactly. There is no second energy model in the ledger.
- PUE-consistent water. Facility energy
Eincludes cooling overhead, so on-site cooling water applies to IT energyE/PUEat WUE 0.18 L/kWh, while off-site generation water applies to the fullEat 5.11 L/kWh (L/kWh = mL/Wh). - Embodied through energy. Hardware embodied carbon is allocated by generation time in EcoLogits, but this ledger stores tokens, not GPU-seconds, so it proxies through energy: an 8xH100 server (5,700 kg server ex-GPU + 8 x 273 kg GPUs, BoaviztAPI / Lees-Perasso et al.) over a 3-year life at full TDP works out to 44.1 gCO2e/kWh. Assuming full utilization is the low-side assumption: real (lower) utilization would raise the per-kWh figure.
scripts/check-crosscheck.sh recomputes the golden-vector token mixes under an
independent EcoLogits-based estimate (factors.json .ecologits: the
ML.ENERGY-fitted per-token function, GPU-count memory model, non-GPU server
power, PUE) and compares blended totals against the Jegham factors. The two
must agree within 3x either way or the factors change is blocked. The comparison
is blended per-session, never per-factor: EcoLogits models input/cache-token
energy as ~0 while Jegham amortizes it into per-token factors, so per-factor
comparison diverges by construction.
Anthropic publishes no per-query energy data (still true as of 2026-08). Every figure this tool shows — co2, energy, water, embodied — is an order-of-magnitude estimate with roughly +/-50% uncertainty, and the Fable factors are a double extrapolation (2x Opus, itself 2x Sonnet). Treat trends and relative comparisons as meaningful, absolute values as indicative.