The simplest way to use an AI agent is to do the same work faster. The more interesting way is to explore what was out of reach without one. For SUEWS, I think of it as a simulation-to-research ladder, where each rung asks a bolder question than the one below.
Here is that ladder, rung by rung. Each step climbs from a basic model run toward a genuine research question, asking a little more of SUEWS each time.
- Task. “Run last summer for this neighbourhood.” The agent assembles the configuration, checks whether the setup is ready, runs the model and reads back the result. This is what the London example on day one of this series was meant to show.
- Diagnosis. “Why does the afternoon heat peak where it does?” Now SUEWS is not just producing an output; it helps inspect a mechanism: radiation, storage heat, vegetation, water, anthropogenic heat and the timing of exchange with the air.
- Comparison. “How would the same summer look with twenty per cent more tree cover?” Now there is a baseline and a scenario, and the agent keeps the comparison disciplined, changing the intended variable while leaving the rest alone.
- Research question. “Which greening interventions deliver the most cooling per litre of water during a heatwave?” Comparative, mechanistic and falsifiable, and consequential too: water-hungry greening in a drought-prone city is a real planning dilemma, not just a modelling exercise.
The model can be the same at every rung. Much of the input data can be the same too. The difference is the question, and that difference changes the value of the whole exercise.
This is where I think the SUEWS-agent matters most. It should not only make SUEWS easier to run; it should free human time from mechanical setup for the judgement that makes a simulation worth running. At every rung the agent can assemble, check, run, hold a comparison steady, and pull in context from the docs and the literature. What it cannot do is supply the question. Question quality remains stubbornly, reassuringly human.
My rule of thumb is simple. A good research question compares this against that. It points to a mechanism, not only a pattern. And it can be wrong, meaning a result could genuinely surprise you. If a question cannot surprise you, it is decoration.
There is one more move, and it is the boldest of all. Every rung so far climbs within a single kind of expertise: it only ever asks a climate model better climate questions. But a real decision about a city is never settled by physical evidence alone. It is settled in a room full of people who want different things, and who is in that room matters as much as what the model says.
Here is what AI agents make newly possible, and it still gives me a small thrill. A heat decision is inescapably many-voiced: a clinician, a landlord, the water utility, an actuary and an outdoor worker do not see the same city, and no single expert holds the whole of it. I can model the physics, but on my own I could never convene that entire room. With AI agents I can: I can give two dozen perspectives their own voice, seat the climate model among them as just one expert, and watch where the argument actually goes. That is a genuinely new move, not a faster version of an old one.
One line I want to keep visible: this is a provocation, not a finding. I chose the cast, the city and the options, so the room partly reflects my own priors; read it as a thought-experiment about who gets to reason with the evidence.
So I had a roomful of agents role-play one city’s heat-adaptation budget. The city was fictional Solenza, split between a dense, low-income core with almost no air-conditioning and a leafy, low-exposure periphery. Three obvious options sat on the table: a flagship park where the model says surface cooling is greatest, neighbourhood greening and cool roofs across the at-risk core, or direct protection such as cooling centres and help with cooling indoors. The climate model’s own agent was deliberately fenced in: grounded in what SUEWS can actually compute, and instructed to refuse anything beyond its evidence and instead name the experiment that would settle it. Then I let them argue, and this time they began to do what a real panel might do: they challenged each other, formed coalitions, traded conditions, and changed their minds.
The exchanges were the part I did not want to stop reading. The A&E clinician cut straight to it: “a beautiful park where almost nobody lives won’t keep one frail, un-air-conditioned tenant out of emergency care.” The parks lead gave ground to the water manager but not all of it: “cool roofs go first… but I won’t concede the shade.” The youth voice put the timing plainly: “nobody survives August on a tree we plant in June. Direct protection first, but ‘first’ isn’t ‘only.’”
And the climate expert behaved exactly as fenced. Out of more than twenty questions, it answered almost none with a number: “I have not simulated that”, “a surface model does not see indoors”, “that would need the irrigation water costed”. The one substantive thing it offered was not an answer but an instruction: the single experiment, a street-canyon microscale run, that would actually settle the disagreement. It did not decide; it told the room what to measure to decide well.
The room did not choose the option that cools the most. It leaned, fairly consistently, towards protecting the most exposed people first, adding greening only where it could be delivered without a water bill no one had costed, and holding the cooling claims back until they could be tested. Not one of two dozen voices backed the flagship park, and I will be honest that a consensus this clean is itself a caveat: a real chamber fractures more than my agents did.
Still, the shape of it is what I find exciting. The physics did not shrink in that room; it became precise about its own edge, and in doing so it made a large, conflicted argument more honest. The most useful thing the model did was know exactly what it could not yet say.
Inside the panel: the full cast, the mechanism, the to-do list
The city (fictional). Solenza, a hot-summer city with one annual heat-adaptation budget. A dense, low-income core with ~8% air-conditioning and many outdoor workers carries the highest modelled heat-risk; a leafy, low-exposure periphery has the coolest surfaces but few residents.
Who was in the room (about two dozen agents). Public health and an A&E clinician; older-residents’ and disability advocate; parent/school; youth; outdoor worker/small business; low-income tenant and a landlord/housing association; water utility; parks/greening; transport; energy/retrofit; finance; actuary; insurer; chamber of commerce; community organiser; climate-justice NGO; maintenance; legal/procurement; media; a data-ethics adviser; the climate model as the physical-evidence expert; and a chair.
What the physical evidence could say. Greener, wetter areas divert more energy into evaporation than into heating the air (a real, directional lever); and the highest heat-risk is the dense, low-AC core, because risk = hazard × exposure × vulnerability, not the hottest or greenest surface.
One more from the floor. Public health, conceding to finance: “Run it as a capped first-year pilot with a sunset review tied to admissions data.”
The decision. A conditioned, protection-led blend: cooling centres and AC/retrofit for the most vulnerable funded first, a ring-fenced core-only slice for cool roofs and drought-safe greening, and nothing to the flagship park this cycle. Spend followed the risk, not the cooling.
What still needs measuring (the panel’s to-do list). Whether greening’s cooling survives in narrow core streets (canyon-scale runs); indoor temperatures and retrofit effects (building-energy models); how many people each option reaches (uptake data); the irrigation water budget; grid headroom for air-conditioning; and a published harm-averted-per-pound with every coverage gap named.
This is what I want to push at the SUEWS Community Hackathon on 24 June 2026. The interesting test is not only whether people can run a climate model in plain language. The deeper test is whether the time the SUEWS-agent gives back gets reinvested one rung up. So if you are joining us on 24 June, take this as an invitation: do not stop at the obvious run. Pick a question that could surprise you, and reach for the rung above the one you think you are allowed to.
I am looking forward to welcoming many of you to UCL East on 24 June. It is forecast to be a genuinely hot day, so please take care on the way in: keep hydrated, keep cool, and look out for each other. I will come back after the day with what we learned, and what I think it changes.
Figure provenance
The figure was generated in Codex with OpenAI’s gpt-image model. The image metadata identifies the generator as gpt-image version 2.0 and the digital source type as trained algorithmic media. It was then saved for this post as figure-04-ladder-to-civic-v1.png.
Prompt recorded for transparency:
Create a restrained scientific editorial concept figure for a SUEWS-agent post about moving from a climate simulation to a bolder civic research question. Show the left side as a technical simulation-to-research ladder, with model evidence climbing from a basic run toward diagnosis, comparison, and a research question. Show a boundary or doorway in the middle. On the right, show a civic decision room or circular panel where the climate model is only one expert voice among public health, water, finance, legal, community, maintenance, and other perspectives. The visual point is that the ladder is not the destination: after climbing it, the model joins the room but does not rule the room. Style: clean scientific editorial schematic, grounded and public-safe, wide landscape, SUEWS-like teal/deep blue/green with a small orange heat-risk accent, no logos, no identifiable people, no private project names, no private place names.
