Conversations about AI and energy often jump from a broad estimate to a broad conclusion. That makes action difficult. Product teams have more useful levers close to the work: the model selected, the number of retries, the length of context, the frequency of batch jobs, and whether a result is used at all.

Count meaningful work first

An energy discussion improves when it begins with a bounded activity. Is a model generating internal drafts, classifying a catalogue, or serving a real-time assistant? How many requests complete a user task? What fraction need to be rerun? A workload inventory will rarely be perfect, but it provides a baseline for improvement and a way to avoid treating every token alike.

Efficiency is not a claim about one model. It is a record of useful work delivered with a clear view of the resources needed to deliver it.

Match capability to the task

The largest available model is not automatically the most responsible or useful choice. Simple extraction, routing, and summarisation tasks may work with smaller systems, shorter prompts, or a conventional rule. Evaluating representative work helps a team choose the lowest-complexity method that consistently meets the quality threshold.

Plan for peaks and persistence

A promising pilot can become a persistent service with a very different footprint. Before scaling, teams should look at peak demand, caching opportunities, scheduled work, and the region where compute will run. Capacity choices influence reliability as well as cost and energy use.

Publish the assumptions

Energy estimates depend on hardware, utilisation, grid conditions, cooling, and measurement boundaries. A credible update names those assumptions and distinguishes observed figures from projections. That transparency lets readers understand the result without turning it into a marketing number.

Find waste in the request path

Many efficiency gains do not require a new model. A product can avoid duplicate requests when a user retries after a slow response, cache a stable result, trim context that is never cited, batch non-urgent work, or stop a job that no longer has a consumer. These changes also improve cost and latency, which makes them easier to maintain than an isolated sustainability initiative.

Start with a task inventory that contains the user outcome, request volume, retry rate, context size, model or method used, and whether the response was actually used. It will be imperfect, especially in early development. Its value is that it turns a broad discussion into a set of questions a product owner can answer and revisit after a release.

Use quality thresholds rather than prestige

A model should be selected against a documented requirement. If a classification or extraction task can meet its quality threshold with a simpler method, choosing the largest available system adds resource use without automatically improving the outcome. Conversely, a smaller model that causes a high rate of reviewer rework is not efficient merely because its individual requests are small.

This is why evaluation and workload planning belong together. Test candidate approaches on representative examples, record the errors that matter, and include the human review that the workflow requires. A decision can then be described accurately: a particular approach met a stated quality bar for a bounded task under stated conditions.

Before scaling a workload

  • Identify the user outcome and the requests that contribute to it.
  • Measure retries, unused completions, context length, and scheduled jobs.
  • Compare methods against the same quality and review threshold.
  • Record the compute location, capacity assumptions, and measurement boundary.
  • Publish observed changes separately from forecasts and revisit them after adoption.

Sources and further reading

The International Energy Agency report on energy and AI and the Green Software Foundation resources give useful context. For model-fit decisions, see our guide to reading an AI evaluation beyond one score.