The most common question I get about AI and water has no answer, and the reason it has no answer is the most useful thing about it.
Someone asks how much water a prompt uses. There is a number circulating, roughly half a litre, and it gets repeated in leadership and board packs, community meetings and opinion columns. It comes from real research, carefully done. It is also being used in a way the researchers themselves were careful to say it should not be.
What the paper actually says
The figure traces back to a study called Making AI Less Thirsty, from researchers at UC Riverside and UT Arlington. What it estimated was that GPT-3 consumed roughly 500 millilitres of water for somewhere between 10 and 50 medium-length responses, depending on where and when the request ran.
Read that again, because the range is the finding. Ten to fifty is a fivefold spread on the same model doing the same work. The half-litre-per-prompt version that circulates took the pessimistic end of that range and dropped the variable that produced it.
The authors were explicit that their estimates were conservative and that the true figures could be several times higher, because GPT-3 is now old, third-party colocation facilities run worse water efficiency than the operators they modelled, and the United States electricity water intensity factor they used, 3.14 litres per kilowatt hour, is below more recent estimates of 4.35.
So the number is simultaneously too precise and possibly too low, which is an unusual combination and a sign that the unit itself is wrong.
Two kinds of water, one of them invisible
The mechanism underneath is worth understanding, because it explains the range rather than explaining it away.
Data centers consume water in two distinct ways. On-site water, which the industry calls scope one, is evaporated in cooling towers or evaporative cooling at the facility itself. This is the water the town sees, drawn from the local supply, and it is the water that shows up in a permit fight.
Off-site water, scope two, is consumed at the power plant generating the electricity the facility uses. Thermal generation evaporates a great deal of water. This water is entirely real and it is invisible at the facility, drawn from a different watershed, often in a different county, and it does not appear on any meter the data center operator reads.
Water usage effectiveness, the on-site ratio, runs anywhere from 1 to 9 litres per kilowatt hour depending on weather. Not depending on the model, or the chip, or the efficiency of the code. Depending on the weather.
Which is the point this rests on. There is no litres-per-prompt number because the same prompt, on the same hardware, costs between one and nine times as much water depending on where it lands and what the outside air is doing when it lands there.
The unit that does work
A per-prompt figure fails at both ends. It is unknowable in advance, and even if you knew it, it answers a question nobody is actually asking.
No town has ever objected to litres per prompt. What a community objects to is litres per day, drawn from their aquifer, during their driest month, in a year they are already under restrictions. That is a completely different quantity, it is measurable, it is already reported in permit filings, and it does not require anyone to estimate anything about a language model.
The same substitution works internally. If you are siting compute, the question is not the water footprint of your workload. It is the water stress of the basin you are proposing to put it in, in August, in a dry year, and whether the facility is designed to use evaporative cooling when that basin is short.
The 5.4 million litres the same paper estimated for training GPT-3 makes the case better than any per-prompt figure. That total ranged from 3.4 million litres if trained in Georgia to 15.3 million if trained in Washington. Identical training run. Four and a half times the water. The variable was geography, and geography is a decision somebody made.
The objection from the operators
There is a strong version of the pushback and it deserves stating properly.
Modern facilities increasingly use closed-loop or air-cooled designs that consume very little on-site water. Some run on reclaimed or non-potable supply. Several large operators have committed to water positive targets and are meeting them in aggregate. Treating data centers as a water problem in general is out of date, and the worst examples are not representative of what is being built now.
Most of that is accurate. It also does not survive contact with the two things that actually determine local impact.
The first is that closed-loop and air-cooled designs trade water for electricity, which means they trade on-site water for off-site water at the power plant. The total can go up while the local number goes down. Whether that is an improvement depends entirely on which watershed each one draws from, and almost nobody publishes that comparison.
The second is that aggregate water positive commitments are net figures across a portfolio. Replenishment in one basin does not put water back in another. A community sitting on a stressed aquifer is not made whole by a wetlands project four states away, and telling them otherwise in a public meeting goes badly and deserves to.
What to ask instead
Three questions, all answerable from documents that already exist.
For any facility you operate, lease or are considering, ask what its annual on-site water withdrawal and consumption are, and what its peak month looks like rather than its average. Ask what the water stress level of that basin is, using any of the public indices. And ask what the facility does when the basin is under restriction, specifically whether the cooling design has a mode that does not evaporate.
If the answers come back as a corporate sustainability figure rather than a facility figure, you have received a portfolio number in response to a local question, which is the same substitution the per-prompt statistic makes.
Why this matters beyond the optics
The practical stake is continuity. A facility that depends on evaporative cooling in a basin heading toward restriction has a capacity constraint that will arrive as a surprise, arrive in summer, and arrive at the same time as everyone else's peak. It will not appear in any capacity plan, because capacity plans are written in megawatts and this constraint is denominated in acre-feet.
The second stake is credibility. Executives who arrive at a county meeting with a per-prompt statistic are answering a question nobody asked, using a number the researchers said not to use that way, in front of people who know exactly how much water their own wells are producing. That is a losing position and it is self-inflicted.
Leadership and the board are usually offered a single company-wide water number. The useful version is a short list of facilities, each with its basin, its peak month and its behaviour under restriction. That fits on one page and it survives a hard question.
Stop asking what a prompt costs. Water has never been a global quantity. It is a local one, and it always has been.