Aurxo Baseline v1.0 · Evidence reviewed 27 July 2026

What the evidence supports.

Aurxo separates current product assessments from scoped model references. A dash is not a poor grade. It means the evidence is not complete or comparable enough for a universal product label.

Only Gemini receives a current product-level impact grade in this first baseline. The other rows make missing evidence and available reference studies visible instead of turning them into invented product scores.

Current product assessments

Product / scope
Grade
Confidence
What Aurxo can responsibly say
Gemini Apps
Median production text prompt · May 2025
A
Moderate
Google reports 0.24 Wh, 0.03 gCO₂e and 0.26 mL across the serving stack. This is a provider-authored production study, not an Aurxo audit.
ChatGPT
Current routed product
Low
OpenAI reports average-query energy and water, but no comparable current product carbon value or reproducible routed-product boundary. A scoped historical GPT-4o reference appears below.
Claude
Current routed product
Insufficient
Anthropic provides extensive model and safety disclosures, but Aurxo did not locate complete current product-level energy, carbon and water metrics. A Claude 3.7 reference appears below.
Microsoft Copilot
Current consumer product
Insufficient
Microsoft Research publishes valuable industry-level inference evidence, including a 0.31 Wh median for an optimized frontier-scale distribution, but not a complete Copilot-specific footprint.
Perplexity
Current product
Insufficient
Aurxo did not identify enough comparable public product-level evidence to issue an impact grade in this review.
Le Chat
400-token lifecycle response
Moderate
Mistral discloses 1.14 gCO₂e and 45 mL under a broad lifecycle boundary, but not a separate per-response energy value. The overall grade is therefore withheld.
DeepSeek
Current product route
Insufficient
Independent references show large differences between DeepSeek-hosted and Azure-hosted scenarios. The current product route must be known before a universal product grade can be issued.

Scoped reference labels

These are valid only for the named historical model, prompt size and deployment assumptions. They are not current universal product grades.

Reference model / scope
Label
Evidence
Metrics and limitation
GPT-4o · March 2025
100 input / 300 output tokens
ChatGPT reference
B
Independent model
0.423 Wh · 0.148 gCO₂e · 1.953 mL. Not telemetry for every current ChatGPT route.
Claude 3.7 Sonnet
100 input / 300 output tokens
Claude reference
C
Independent model
0.950 Wh · 0.273 gCO₂e · 5.005 mL. Historical model reference, not telemetry for current Claude routes.
DeepSeek V3
DeepSeek-hosted · 100 input / 300 output
D
Independent model
2.777 Wh · 1.666 gCO₂e · 19.33 mL. A different hosting route requires a different label.
DeepSeek R1
DeepSeek-hosted · 100 input / 300 output
E
Independent model
19.251 Wh · 11.551 gCO₂e · 134.004 mL. Reasoning depth and deployment materially affect the result.