Skip to content

Development carbon

An estimate of the energy and carbon emitted in developing and validating virgil: AI coding assistants, continuous integration and cluster jobs. It is refreshed by hand with python3 scripts/dev_carbon.py, so it lags the repository.

Estimated development carbon for benjaminpope/virgil from 2026-03-03 to 2026-10-06: 76.5 kg CO₂e (range 59.7–172 kg), from 202 kWh (range 152–447 kWh). Data analysis runs are listed separately below and are not in this total.

By source

Source Items kWh kWh range kg CO₂e kg range
Claude Code 278 153 119–289 50.4 39.1–95.4
OzSTAR/NT Slurm jobs 785 23.2 22.9–58.4 17.1 17–43.2
Copilot cloud agent and review 251 11.7 3.28–45.1 3.89 1.09–14.9
VS Code Copilot Chat 253 9.35 3.82–48.9 3.09 1.26–16.1
GitHub Actions CI 2175 5.25 3.29–5.25 1.94 1.22–1.94

By call and token type

Source Type Items kWh kg CO₂e kg range
Claude Code model calls 278 153 50.4 39.1–95.4
OzSTAR/NT Slurm jobs GPU job 476 21.9 16.2 16–42.3
Copilot cloud agent and review code review 192 10.5 3.47 0.947–13.3
VS Code Copilot Chat local chat 88 6.23 2.06 1.26–10.6
GitHub Actions CI CI 2175 5.25 1.94 1.22–1.94
VS Code Copilot Chat local chat (imputed tokens) 165 3.12 1.03 0–5.49
OzSTAR/NT Slurm jobs CPU job 309 1.25 0.924 0.924–0.924
Copilot cloud agent and review cloud agent 59 1.27 0.424 0.147–1.57

Claude Code by token type (mid estimate):

Token type Tokens kg CO₂e
input 31,604 0.00264
cache write 63,763,507 5.32
cache read 4,277,775,367 30
output 8,001,134 15

By model or workflow:

Source Model or workflow Items kg CO₂e kg range
Claude Code Claude Opus 198 49.3 38.2–93.4
GitHub Actions CI automated tests 645 1.79 1.12–1.79
Claude Code Claude Sonnet 79 1.14 0.914–2.04
VS Code Copilot Chat gpt-5.3-codex 172 1.07 0.0272–5.77
VS Code Copilot Chat gpt-5.6-terra 23 1.02 0.509–5.94
VS Code Copilot Chat claude-sonnet-5 19 0.393 0.393–1.52
VS Code Copilot Chat gpt-5.5-2026-04-23 13 0.285 0.142–1.09
VS Code Copilot Chat gpt-5.6-sol 5 0.124 0.062–1.12
GitHub Actions CI Documentation 645 0.0794 0.0498–0.0794
VS Code Copilot Chat gpt-5.5 2 0.0708 0.0354–0.179
VS Code Copilot Chat claude-opus-5 3 0.0516 0.0516–0.121
GitHub Actions CI lint 601 0.0409 0.0257–0.0409
VS Code Copilot Chat gpt-5.6-luna 7 0.0378 0.0192–0.269
GitHub Actions CI Documentation (Zensical) 240 0.0236 0.0148–0.0236
VS Code Copilot Chat mai-code-1.1-flash 5 0.0196 0.0196–0.0617
VS Code Copilot Chat claude-haiku-4.5 3 0.0113 0.00215–0.0512
VS Code Copilot Chat copilot/auto 1 0.00897 0–0.0179
GitHub Actions CI Dependency Graph 22 0.00572 0.00358–0.00572
Claude Code Claude Haiku 1 0.00468 0.00417–0.00675
GitHub Actions CI copilot-setup-steps 15 <0.001 <0.001–<0.001
GitHub Actions CI pages-build-deployment 7 <0.001 <0.001–<0.001

By feature (top 15 of 218)

Each item is attributed to a pull request through its branch, the commit a job pinned, or the PR a Copilot run served, and labelled with the PR title.

Feature Items kg CO₂e kg range
main / unattributed 451 30.1 22.3–67.3
Validation: imaging contests 472 15.9 15.8–42
#201 virgil docs minor changes 29 4.1 3.2–7.69
#76 Imaging Stage 3: Problem, fit, regularisers, L-curves and diagnose 25 1.43 1.07–2.85
#110 Imaging Stage 5d: a Gauss–Newton NUTS mass matrix, and the sampling tutorial; plan PMOIRED parity 13 1.34 1.07–2.41
#91 Imaging Stage 5c: error-bar scale, sampling basics, and imaging tutorials 2–4 16 1.28 0.977–2.47
#226 fit: optimise in each prior's flat coordinate (LM with Jeffreys priors) 10 0.665 0.481–1.4
#210 Stage 6a PR B follow-up: VISPHI/T3PHI departure stated, Jeffreys note, draw test 23 0.602 0.434–1.25
#83 SPARCO spectra: BlackBody temperatures 17 0.535 0.407–1.02
#120 Stage 6.0: independent, whitened closure phases; INSNAME selection; PHITYP check; no diagonal truncation 9 0.503 0.358–1.07
#118 Merge imaging into main: image reconstruction (milestone 1 and Stage 5) 14 0.479 0.326–1.08
Validation: simulation-based calibration 100 0.466 0.466–0.466
#121 Design: spectro-interferometry, orbits and GRAVITY calibration notes; the plan through Stage 6 27 0.461 0.313–1.02
#244 CI: fix the py3.11 lowest job (orbit doctests, clean hang) 19 0.459 0.398–0.663
#246 Orbit tutorial: joint, hierarchical inference from every epoch's interferometric data 10 0.42 0.277–0.972

Copilot cross-check

Run time times an assumed token rate, against billing (AI Credits or premium requests) attributed to this repository. The headline uses billing where a month has it.

Month Kind Runs Minutes kg (run time) kg (billing) Billing basis
2026-03 cloud agent 12 57.6 0.0446–0.373 – –
2026-03 code review 7 16.5 0.0128–0.107 0.0419–0.165 9.2 premium requests; local VS Code requests share
2026-09 cloud agent 41 167 0.129–1.08 0.069–1.09 2,129 AI Credits; agent PR share
2026-09 code review 24 95.7 0.0741–0.619 0.161–2.55 4,984 AI Credits; local VS Code credits share
2026-10 cloud agent 8 54.1 0.0419–0.35 0.00385–0.0606 119 AI Credits; agent PR share
2026-10 code review 168 647 0.501–4.19 0.693–10.9 21,381 AI Credits; repository-attributed row

Excluded data analysis

Compute for science with virgil (fits to observations, data reduction and archive downloads), and anything whose purpose was unclear, is costed the same way but kept out of the total. It is shown only in aggregate.

Item Records kWh kg CO₂e kg range
Excluded data analysis 555 77.6 28.6 22.9–49

Scope and caveats

  • What is counted. Development and validation of virgil: AI coding assistants, GitHub Actions CI, and jobs on Swinburne's OzSTAR/NT cluster. Validation jobs (simulation-based calibration, the imaging contests, detection false-alarm and orbit checks) are counted. Data analysis is not (see above).
  • Claude Code was used from 2026-09-28, and is counted from its local transcripts on one machine. Sessions run in the cloud (claude.ai, cloud agents) or on other machines are not stored locally and are not captured.
  • Before that, VS Code Copilot Chat was the main assistant (from March 2026). Its logs give prompt and output tokens per request but no cache split, so prompts are costed as uncached input, which errs high.
  • GPT and other non-Claude models have no published per-token energy. Each is assigned an assumed Claude size class (e.g. GPT-5.6 Sol and Terra and GPT-5.5 as Opus class, GPT-5.6 Luna and GPT-5.3-Codex as Sonnet class), with a range across the neighbouring classes.
  • Copilot cloud agent and code review expose no token counts. Their billed AI Credits are converted to tokens at the rate measured on the local chat logs (572 credits per million tokens), with a factor-of-two range; months without credits use run time times a token rate.
  • Data centres and grids. Model inference uses TokenClimate's PUE (1.14) and grid factor; GitHub runners assume a hyperscale PUE of 1.18 and the US average grid; NT assumes Green Algorithms' default PUE of 1.67 and the Victorian grid. None of these is published for the specific facilities.
  • Cache reads dominate the uncertainty of the Claude figure. They are almost all of its tokens, costed at 0.08 times the input energy (range 0.05–0.20); at 1.0 the Claude figure would be several times larger.
  • Model training, local hardware and the embodied carbon of the cluster are not included.

Methodology

Every figure is an estimate with a low, mid and high value; tables show the mid value and the range. Energy is facility energy (server energy times the data centre's PUE).

Source Telemetry Energy model Main assumptions
Claude Code Token counts by type from local transcripts TokenClimate per-token factors by model family Cache reads at 0.08× input energy (range 0.05–0.2); PUE 1.14; 0.376 g CO₂e per server Wh
VS Code Copilot Chat Prompt and output tokens per request from local chat logs TokenClimate factors for an assumed model class GPT and other non-Claude models placed in a Claude size class by assumption; prompts costed as uncached input; the high value counts every model call's prompt; requests without counts take the median of counted ones
Copilot cloud agent and review Workflow run time; billed AI Credits and premium requests Tokens from credits (or run time) on TokenClimate Sonnet factors 572 credits per million tokens, measured on local chat (measured); account-wide billing attributed by this repository's share of local use or agent PRs
GitHub Actions CI Runner time per job Green Algorithms 4 vCPU at 4.38 W each, 16 GB, usage 0.5–1.0, PUE 1.18, 0.37 kg/kWh
OzSTAR/NT Slurm jobs sacct allocation and run time; NT Job Report usage Green Algorithms EPYC 7543 at 7.03 W per core, A100 at 400 W, measured usage where reported, PUE 1.67, Victorian grid 0.74 kg/kWh

Not included: Claude sessions run in the cloud or on other machines, the energy to train the models, and the embodied carbon of local hardware. The cache-read factor dominates the uncertainty of the Claude figure, and Copilot's cloud inference is inferred from billing or run time rather than measured.

Sources

  • TokenClimate, methodology tokenclimate-v3-2026-09, https://tokenclimate.com/en/methodology
  • Lannelongue, L., Grealey, J. & Inouye, M. (2021), Green Algorithms: Quantifying the carbon footprint of computation, Advanced Science 8, 2100707, https://doi.org/10.1002/advs.202100707
  • DCCEEW (2026), Australian National Greenhouse Accounts Factors, Table 1, https://www.dcceew.gov.au/climate-change/publications/national-greenhouse-accounts-factors
  • US EPA, eGRID2022 (US average grid intensity, for GitHub-hosted runners)