Acquiring signal .
Acquiring signal .
The El Niño–Southern Oscillation (ENSO) is the dominant mode of interannual climate variability on Earth, modulating seasonal precipitation and temperature regimes across every inhabited continent. Warm ENSO phases (El Niño) are associated with drought in Australia and southern Africa, anomalous rainfall along the South American Pacific coast, and suppressed monsoon intensity across South and Southeast Asia; La Niña events broadly reverse these anomalies. The socioeconomic consequences, affecting agricultural production, water availability, and disaster risk, motivate the development of accurate, extended-range ENSO forecasts.
This project trains a multi-model ensemble on historical tropical Pacific SST anomalies to forecast the Niño3.4 index at lead times of 1–6 months, evaluating the complementary skill profiles of a convolutional neural network, Ridge Regression with PCA, and a random forest baseline.

ENSO is the dominant source of year-to-year climate variability on Earth. Skilful forecasts at lead times of 3–6 months would provide actionable early warning for agricultural, hydrological, and humanitarian planning.
ENSO is the dominant source of year-to-year climate variability on Earth. Traditional forecast systems use expensive dynamical models running the physics of the ocean and atmosphere forward in time. Statistical ML approaches learn the predictor-predictand relationship directly from historical data, offering a much cheaper inference path once trained. A key challenge is constructing a proper temporal train-test split: the tropical Pacific's 1–2 year ocean memory means that a naively random split will silently leak future information into training, producing inflated skill estimates that collapse on a correct holdout.
Three candidate models were independently trained on monthly SST anomalies from COBE-SST2, restricted to the tropical Pacific domain, to predict the Niño3.4 index at each of six lead times:
Ensemble weights were optimised on a validation set via constrained optimisation. To prevent data leakage arising from the tropical Pacific's 1–2 year ocean memory, training was separated from the test period (2007–2017) by gap years (1995 and 2006). Training data were extended back to 1920 (~900 monthly samples) to provide sufficient exposure to diverse ENSO regimes.


A naive random train–test split yielded a test Pearson correlation of 0.92. A temporally correct holdout, ith a full gap year between train and test, gave −0.02. The tropical Pacific's 1–2 year memory renders random splitting a source of severe, silent data leakage in ENSO prediction tasks.
The weighted ensemble achieved a Pearson correlation of 0.94 and RMSE of 0.35 K at 1-month lead, decaying to 0.65 and 0.75 K at 6-month lead. The three component models exhibit markedly different skill decay profiles. The random forest degrades fastest (by 3-month lead its correlation falls to 0.71 and by 6 months to 0.46) as it lacks the spatial inductive bias required to track the slowly-evolving large-scale SST patterns that sustain predictability at longer leads. The CNN degrades more gradually, retaining a correlation of 0.64 at 6 months. Ridge regression with PCA is the strongest individual model beyond 1-month lead, as principal component decomposition efficiently isolates the oceanic modes, including the east Pacific thermocline tilt, that dominate multi-month ENSO predictability.
The three component models have complementary skill profiles: the random forest relies on local SST persistence, the CNN captures distributed spatial patterns, and Ridge with PCA isolates large-scale slowly-evolving modes. The superiority of the ensemble at all lead times reflects genuine diversity in what each model captures about tropical Pacific dynamics.
| Lead (months) | CNN corr | RF corr | Ridge corr | Ensemble corr | Ensemble RMSE (K) |
|---|---|---|---|---|---|
| 1 | 0.926 | 0.941 | 0.949 | 0.940 | 0.353 |
| 2 | 0.759 | 0.846 | 0.901 | 0.907 | 0.471 |
| 3 | 0.803 | 0.711 | 0.841 | 0.835 | 0.564 |
| 4 | 0.763 | 0.623 | 0.776 | 0.789 | 0.609 |
| 5 | 0.671 | 0.493 | 0.709 | 0.709 | 0.715 |
| 6 | 0.640 | 0.460 | 0.628 | 0.646 | 0.750 |
The most interesting result was how the three models have complementary skill profiles. The RF skill collapsed fastest because it had no spatial awareness. The CNN degrades more gracefully, learning spatial SST patterns via convolution. Ridge regression with PCA is most competitive at longer leads because PCA efficiently captures the slowly-evolving large-scale modes, like the eastern Pacific thermocline tilt, that dominate multi-month predictability.
The other key practical lesson was the importance of temporal train-test splits for time series with autocorrelation. A naive random split on the 36-month-lead experiment gave a suspiciously high correlation of 0.92 on the "test" set, which collapsed to −0.02 on a proper temporal holdout, because the ocean's 1–2 year memory was leaking information from train into the randomly-selected test samples.