What is the AI Energy Use Per Prompt?
| What it answers | Watt-hours and CO2 behind one prompt. |
|---|---|
| How the answer is produced | The estimate multiplies three quantities: the compute required to generate your tokens, the energy that compute draws including datacentre overhead, and the carbon intensity of the electricity supplying it. |
| What you need to enter | Enter a realistic average response length; most people underestimate it substantially. |
| Where it stops being reliable | Providers do not publish per-query energy figures, so this is a public-research estimate with wide error bars. |
| Cost and sign-up | Free, runs in your browser, no account and no stored inputs. |
How are energy and emissions per prompt estimated?
The estimate multiplies three quantities: the compute required to generate your tokens, the energy that compute draws including datacentre overhead, and the carbon intensity of the electricity supplying it. Each is a published, checkable figure, and each carries real uncertainty.
Inference energy scales with output length far more than input length, because each generated token requires a full forward pass while the prompt is processed in parallel. That is why a long answer costs many times a short one for the same question.
Datacentre overhead is applied through a power usage effectiveness multiplier, typically between 1.1 and 1.3 for large modern facilities. Grid carbon intensity then varies by more than tenfold between regions, which usually dominates every other factor in the calculation.
How do you use the AI Energy Use Per Prompt?
- 1.Enter a realistic average response length; most people underestimate it substantially.
- 2.Pick a model size band rather than a specific product, since providers rarely publish parameter counts.
- 3.Choose the grid region where the provider's datacentres sit, not where you are.
- 4.Compare the result against familiar activities rather than reading the raw figure in isolation.
What can this tool not tell you?
- Providers do not publish per-query energy figures, so this is a public-research estimate with wide error bars.
- Training energy is excluded; this covers inference only.
- Efficiency improves quickly, and figures more than a year old overstate current per-query cost.
Why the honest answer is a range, not a number?
Nobody outside the handful of companies operating these models has access to the true per-query energy figure, which makes any public estimate — including this one — a reconstruction from published research on comparable model sizes rather than a measurement. Presenting that as a precise number would misrepresent the confidence level, which is why the result is best read as an order of magnitude: is this closer to a web search or closer to boiling a kettle, not is it exactly 0.34 watt-hours.
Interpreting the estimate correctly means recognising which input dominates the outcome, and it is rarely the one people assume. Model size matters, but response length usually matters more, because generation cost scales with the number of tokens produced, and people routinely underestimate how long a typical chatbot answer actually runs to. Grid carbon intensity dominates everything else in the emissions figure specifically, since the same energy draw produces wildly different emissions depending on whether the electricity came from a coal plant or a river.
What changes the answer most between two people using the same tool is where the provider's datacentres are located and how long their typical response runs, not which specific product they used, since providers rarely disclose parameter counts precisely enough to distinguish between similarly sized models. The common mistake is comparing a single prompt's energy to a daily human activity and drawing a strong conclusion — the honest use of this figure is relative comparison between habits, such as short answers versus long ones, not an absolute verdict on whether to use AI at all.
Scale is the part of this subject where intuition fails hardest in the other direction. A single prompt is genuinely trivial — comparable to a short web search or a few seconds of video streaming — and no individual's chat habit is a meaningful climate decision. What is meaningful is aggregate: a feature that quietly calls a large model on every page load for millions of users, or a batch job re-summarising an entire archive nightly because nobody checked whether the output was read. The useful place to apply this estimate is therefore an engineering review of automated call volume, not personal guilt about asking a chatbot a question.
What do worked examples look like?
A quick factual question versus a long essay request
A one-sentence factual answer from a mid-sized model estimates to a small fraction of a watt-hour, while asking the same model for a 1,500-word essay on a similar topic can estimate several times higher, purely from the extra tokens generated. The comparison illustrates that phrasing habits — asking for a summary rather than a full essay when a summary would do — affect footprint more than which model is chosen.
The same prompt on two different grids
Running an identical estimated computation against a grid region with high renewable share versus one that is coal-dominant can show a roughly tenfold difference in estimated emissions for the same energy draw. This demonstrates that a provider's choice of datacentre location, something users rarely see, is often a bigger emissions lever than any setting available to the person typing the prompt.
Where the same workload costs ten times less
A support tool classifies incoming tickets into eight categories using a frontier model on every message. Swapping to a small model fine-tuned for that classification, with the large model reserved for the minority of tickets the small one flags as uncertain, keeps accuracy roughly level while cutting estimated energy for the workload by around an order of magnitude. Nothing about the user experience changes; the saving comes entirely from matching model size to task difficulty, which is the lever with by far the largest effect in most production systems.
What do people ask most about this tool?
How much energy does one AI prompt use?
Public estimates for a typical chat response cluster around a few watt-hours — roughly the order of a short web search to a couple of minutes of a laptop's draw, depending heavily on response length and model size.
Is using AI worse than streaming video?
Per minute of use, streaming generally consumes more. Per task, it depends entirely on how many long responses you generate.
Does the region really matter that much?
Yes. The same computation can produce ten times the emissions on a coal-heavy grid compared with a hydro or nuclear-heavy one.
Which related tools should you try next?
Written and reviewed by Jim Vernon, Editor, AI Intelligence International. Published by AI Answer Engine, a service of AI Intelligence International, and checked against our editorial standards.
Lovable Labs Platform