Building a Hyperlocal Rain Nowcaster on a $0 Stack: Benchmarking Open-Meteo Against Airport Radar
A rain forecast for Mumbai can tell you to carry an umbrella while your street stays dry. During the monsoon, a useful next-two-hour prediction needs more local evidence than a global model alone can provide. I built Mumbai Rain, also called เคชเคพเคเคธ, around that problem. I collect forecasts, compare them with observations at Mumbai airport, and train a small correction that visitors run in their browsers. First, a qualification about the title: I do not ingest airport radar. I use METAR surface-weather reports from station VABB, Chhatrapati Shivaji Maharaj International Airport. Airport observers and instruments report present weather; I convert their precipitation codes into hourly rain labels. Radar measures precipitation across an area. A METAR report describes conditions at a station. That distinction sets the limits of the entire experiment. The source repository contains the data diary, training code, browser arithmetic, and exported model. You can inspect the method and scoreboard without accepting a headline accuracy claim. Start with a narrower prediction target Engineers who build global models such as ECMWF and GFS solve atmospheric dynamics over grids. A forecast grid cell cannot resolve each Mumbai street, coastal exposure, or passing monsoon shower. Asking for a coordinate does not create street-level observational coverage. For this project, I ask a narrower question: given the forecast inputs available now, how likely is an hourly rain observation at the airport? I use Open-Meteo as the delivery API for two forecast streams: - best_match , the primary forecast feed, including precipitation and relative humidity. - ecmwf_ifs025 , an explicit ECMWF-IFS precipitation benchmark. I do not train a replacement for ECMWF. I train logistic regression to correct the relationship between forecast features and an independent binary observation. Visitors receive a local-coordinate forecast with a correction learned at VABB. They should not read that as a measured rain probability for every neighbourhood. Collect forecasts before the outcome arrives The hourly collection workflow schedules a GitHub Actions run at minute five: schedule: - cron: "5 * * * *" GitHub can delay scheduled jobs. I treat the cron as a collection schedule, not a real-time delivery guarantee. In pipeline/log_snapshot.py , I collect at (19.12, 72.85) , record the forecast issue time, and append forecast-hour rows to data/log.csv . On subsequent runs, I fetch the last 24 hours of METAR reports and fill labels for matching hours. flowchart TD Clock[Hourly GitHub Actions collection] --> Forecast[Open-Meteo best_match and ECMWF-IFS] Forecast --> Diary[data/log.csv: forecasts with blank labels] Airport[VABB METAR present-weather reports] --> Labels[Convert UTC to IST; label RA and DZ] Labels --> Diary Diary --> Filter[Keep labelled forward forecasts] Daily[Daily GitHub Actions retraining] --> Filter Filter --> Train[Fit logistic regression] Train --> Gate[Holdout scores and purged walk-forward folds] Gate --> Decision{Promotion criteria met?} Decision -->|Yes| Export[public/model.json] Decision -->|No| Retain[Retain champion or evaluate raw fallback] Export --> CDN[Astro assets on CDN] CDN --> Browser[Browser loads model] Live[Live Open-Meteo forecast] --> Browser Browser --> Math[Six features and sigmoid] Math --> Verdict[Hourly rain calls and next-two-hour verdict] I record both timestamps because they answer different questions: issued_at Forecast snapshot time, UTC valid_at Forecast target hour, Asia/Kolkata fc_bestmatch_mm Primary forecast precipitation fc_ecmwf_mm ECMWF forecast precipitation fc_rh_bestmatch Forecast relative humidity hour Target local hour recent_rain_mm Recent precipitation feature observed_raining Blank until labelled, then 0 or 1 I can then ask what I knew when I issued a forecast, rather than download a later forecast and pretend I had predicted the outcome. The collector stores the returned day's hourly series, including hours that may precede the issue time. Before training, I filter those rows out. In matured() , I convert issued_at from UTC to IST and require valid_at >= issued_at . I train on forward forecasts because the browser serves forward hours. Label rain without grading a forecast against itself In pipeline/labels.py , I read the METAR wxString field: | Report token | Label interpretation | |---|---| RA , including -RA , +RA , SHRA , TSRA | Rain | DZ | Drizzle | BR , HZ , FG | Mist, haze, fog: not a rain label | I convert each report timestamp from UTC to IST and floor it to the local hour. If I find multiple reports in one hour, I label that hour as rainy if any report includes rain or drizzle. For a missing report, I leave the label blank. I do not equate missing observations with dry weather. This gives me an independent target, with several qualifications. Airport reports can miss a shower between observations. A binary label gives me no measured rainfall amount. An hourly label loses sub-hour timing. And a dry airport cannot establish that Bandra, Thane, or a visitor's GPS coordinate stayed dry. I also keep the recent-rain feature separate from the label. Despite the name recent_rain_mm , I compute that feature from Open-Meteo's precipitation series over the three hours before the current hour. I do not obtain a neighbourhood rain-gauge measurement through that field. Use Git as the database, within its limits I keep the diary in data/log.csv . The bot commits the changed CSV and health metrics after collection. In each row, I retain the forecast and later add the observed label. For one collection point and hourly jobs, I get a readable history and a training dataset that contributors can download without credentials. I also avoid provisioning a database just to collect this experiment. I accept the costs of that choice. The logger reads and rewrites the CSV. Repeated commits increase repository history. Separate collection and training workflows can encounter competing pushes; their individual concurrency groups do not create a shared database transaction. Git suits the current small workload, but I would revisit storage before collecting hundreds of stations. Make the daily model earn deployment The retraining workflow schedules a run at 01:30 UTC . It runs the Python tests, trains the candidate, refreshes health metrics, and commits an artifact if the training decision changes it. In pipeline/train.py , I require at least 200 labelled forward rows. I also require both rainy and dry examples in training and holdout data. I fit six features in this order: x = [best_match_mm, ecmwf_mm, relative_humidity, sin(2π hour / 24), cos(2π hour / 24), recent_rain_mm] I encode the hour as sine and cosine so midnight and 23:00 remain neighbours in the feature space. I keep the feature order identical in Python and src/lib/nowcast.js . Swapping humidity and precipitation would produce valid JavaScript and invalid predictions. For evaluation, I use the Brier score: $$ \operatorname{Brier} = \frac{1}{N}\sum_{i=1}^{N}(p_i-y_i)^2 $$ A confident wrong prediction incurs a large penalty. A lower score means lower squared probability error on the evaluated observations. I compare the candidate against the existing logistic champion, a raw rain call, and climatology: the training set's rain frequency. The main holdout contains the last 20% of rows after time sorting. Purge by time, not row count One target hour can appear in several forecast snapshots. If I shuffle rows, I can train on one prediction of an hour and test on another prediction of the same observed event. I add four expanding-window walk-forward folds. I place boundaries between distinct valid_at timestamps and purge six hours before each test window. I skip folds with fewer than 25 training or test rows, or without both label classes. Time advances to the right Fold 1: [training] [6h purge] [test] Fold 2: [ expanded training ] [6h purge] [test] Fold 3: [ expanded training ] [6h purge] [test] Fold 4: [ expanded training ] [6h purge] [test] For each fold, I calculate Brier skill against the raw call: $$ \operatorname{BSS} = 1 - \frac{\operatorname{Brier}{candidate}}{\operatorname{Brier}{raw}} $$ I require at least two usable folds, a positive median skill, and at least two positive fold scores. For a perfect raw reference, the implementation reports zero skill rather than divide by zero. This excerpt expresses the promotion decision using the functions and score variables from the training module: # Score candidate, champion, raw call, and climatology on the holdout. # Calculate fold_skills with purged expanding-window evaluation first. walk_forward_ok = passes_walk_forward(fold_skills) if walk_forward_ok and passes_gate( cand_brier=cand_b, champ_brier=champ_b, raw_brier=raw_b, clim_brier=clim_b, ): with open(MODEL_PATH, "w") as f: json.dump(candidate, f, indent=2) Two implementation details matter when interpreting the benchmark: - The automated raw gate uses best_match , not ECMWF alone. The code thresholds the first feature at0.3 mm/h to construct its raw binary forecast. I collect ECMWF as a separate benchmark and model input, but I cannot describeraw_brier as an ECMWF-only result. - The holdout gate accepts ties. passes_gate() rejects a candidate whose Brier score exceeds a reference score. The walk-forward median must exceed zero, but the holdout comparisons mean “no worse,” rather than a strict improvement against each reference. I should also avoid promising that future skill can only improve. I evaluate each deployment against the available historical data. Weather regimes change. The code includes a conditional raw-forecast fallback when a champion loses to the raw forecast or climatology and the walk-forward check permits that decision path. The four purged folds reduce temporal leakage, but the separate 80/20 row split does not use the same timestamp-grouping and purge logic. Repeated target hours can straddle that split. I treat the walk-forwa
Comments
No comments yet. Start the discussion.