I'm an AI weather scientist who came up through observational atmospheric science, spent a few years building the models, and now spends most of my time on the question of whether they hold up when someone has to make a real decision from them.My PhD (Illinois, with Steve Nesbitt and Tim Lang at NASA MSFC) was on tropical mesoscale cold pools using ASCAT and RapidScat scatterometer winds. After postdocs at UVA and Argonne, I went to NVIDIA, where I was co-lead on StormCast, a km-scale generative model for convective-scale forecasting. Most recently I built AI forecasting systems for energy trading at RWE's Commercial AI Lab, and I've just started as Lead AI Weather Scientist at Point72 in commodities. I also co-vice-chair the AMS Committee on AI Applications to Environmental Sciences.
That path — obs to research to GPU vendor to market-facing forecasting — means I've watched the same models get judged four different ways, and I've become somewhat obsessive about the gap between benchmark skill and decision-relevant skill.
Talk to me about:
- Evaluating AI weather models honestly. Our recent arxiv paper benchmarks ten models in a common harness; the short version is that training recipe matters more than architecture, and RMSE leaderboards hide a lot.
- Uncertainty at S2S range — CRPS-trained ensembles on HEALPix grids, bred vectors, and forecast-of-a-forecast approaches to run-to-run instability.
- Generative modeling at convective scale — StormCast, cBottle, and candidly what these are and aren't good for.
- What operational centers look like from the demand side. Energy and commodities are heavy consumers of NOAA output, and I have opinions about which products carry the most value and where the gaps are.
- Career paths across national lab, industry, and finance — especially with students and early-career folks.