<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing with OASIS Tables v3.0 20080202//EN" "https://jats.nlm.nih.gov/nlm-dtd/publishing/3.0/journalpub-oasis3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:oasis="http://docs.oasis-open.org/ns/oasis-exchange/table" xml:lang="en" dtd-version="3.0" article-type="review-article"><?xmltex \bartext{Review and perspective paper}?>
  <front>
    <journal-meta><journal-id journal-id-type="publisher">GMD</journal-id><journal-title-group>
    <journal-title>Geoscientific Model Development</journal-title>
    <abbrev-journal-title abbrev-type="publisher">GMD</abbrev-journal-title><abbrev-journal-title abbrev-type="nlm-ta">Geosci. Model Dev.</abbrev-journal-title>
  </journal-title-group><issn pub-type="epub">1991-9603</issn><publisher>
    <publisher-name>Copernicus Publications</publisher-name>
    <publisher-loc>Göttingen, Germany</publisher-loc>
  </publisher></journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.5194/gmd-17-2347-2024</article-id><title-group><article-title>Advances and prospects of deep learning for medium-range<?xmltex \hack{\break}?> extreme weather forecasting</article-title><alt-title>Advances and prospects of deep learning</alt-title>
      </title-group><?xmltex \runningtitle{Advances and prospects of deep learning}?><?xmltex \runningauthor{L.~Olivetti and G.~Messori}?>
      <contrib-group>
        <contrib contrib-type="author" corresp="yes" rid="aff1 aff2">
          <name><surname>Olivetti</surname><given-names>Leonardo</given-names></name>
          <email>leonardo.olivetti@geo.uu.se</email>
        <ext-link>https://orcid.org/0009-0003-4904-4362</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1 aff2 aff3 aff4">
          <name><surname>Messori</surname><given-names>Gabriele</given-names></name>
          
        <ext-link>https://orcid.org/0000-0002-2032-5211</ext-link></contrib>
        <aff id="aff1"><label>1</label><institution>Department of Earth Sciences, Uppsala University, 75236 Uppsala, Sweden</institution>
        </aff>
        <aff id="aff2"><label>2</label><institution>Centre of Natural Hazards and Disaster Science, Uppsala University, 75236 Uppsala, Sweden</institution>
        </aff>
        <aff id="aff3"><label>3</label><institution>Department of Meteorology, Stockholm University, 10691 Stockholm, Sweden</institution>
        </aff>
        <aff id="aff4"><label>4</label><institution>Bolin Centre for Climate Research, Stockholm University, 10691 Stockholm, Sweden</institution>
        </aff>
      </contrib-group>
      <author-notes><corresp id="corr1">Leonardo Olivetti (leonardo.olivetti@geo.uu.se)</corresp></author-notes><pub-date><day>21</day><month>March</month><year>2024</year></pub-date>
      
      <volume>17</volume>
      <issue>6</issue>
      <fpage>2347</fpage><lpage>2358</lpage>
      <history>
        <date date-type="received"><day>27</day><month>October</month><year>2023</year></date>
           <date date-type="rev-request"><day>24</day><month>November</month><year>2023</year></date>
           <date date-type="rev-recd"><day>9</day><month>February</month><year>2024</year></date>
           <date date-type="accepted"><day>19</day><month>February</month><year>2024</year></date>
      </history>
      <permissions>
        <copyright-statement>Copyright: © 2024 </copyright-statement>
        <copyright-year>2024</copyright-year>
      <license license-type="open-access"><license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p></license></permissions><self-uri xlink:href="https://gmd.copernicus.org/articles/.html">This article is available from https://gmd.copernicus.org/articles/.html</self-uri><self-uri xlink:href="https://gmd.copernicus.org/articles/.pdf">The full text article is available as a PDF file from https://gmd.copernicus.org/articles/.pdf</self-uri>
      <abstract><title>Abstract</title>

      <p id="d1e112">In recent years, deep learning models have rapidly emerged as a stand-alone alternative to physics-based numerical models for medium-range weather forecasting. Several independent research groups claim to have developed deep learning weather forecasts that outperform those from state-of-the-art physics-based models, and operational implementation of data-driven forecasts appears to be drawing near. However, questions remain about the capabilities of deep learning models with respect to providing robust forecasts of extreme weather. This paper provides an overview of recent developments in the field of deep learning weather forecasts and scrutinises the challenges that extreme weather events pose to leading deep learning models. Lastly, it argues for the need to tailor data-driven models to forecast extreme events and proposes a foundational workflow to develop such models.</p>
  </abstract>
    
<funding-group>
<award-group id="gs1">
<funding-source>European Research Council</funding-source>
<award-id>948309</award-id>
</award-group>
</funding-group>
</article-meta>
  </front>
<body>
      

<sec id="Ch1.S1" sec-type="intro">
  <label>1</label><title>Introduction</title>
      <p id="d1e124">The very first deep learning models for weather applications date back to the 1990s  <xref ref-type="bibr" rid="bib1.bibx67 bib1.bibx31" id="paren.1"/>, and extensive research on the use of deep learning models for weather forecasting at a local scale <xref ref-type="bibr" rid="bib1.bibx78 bib1.bibx49 bib1.bibx30" id="paren.2"><named-content content-type="pre">e.g.</named-content></xref> and for short-term weather predictions <xref ref-type="bibr" rid="bib1.bibx41 bib1.bibx56" id="paren.3"><named-content content-type="pre">e.g.</named-content></xref> has been ongoing since the mid-2010s. More recently, deep learning models have also been employed successfully as a nowcasting tool for precipitation  <xref ref-type="bibr" rid="bib1.bibx59 bib1.bibx26" id="paren.4"><named-content content-type="pre">e.g.</named-content></xref> and as post-processing tools for numerical weather forecasts  <xref ref-type="bibr" rid="bib1.bibx27 bib1.bibx69" id="paren.5"><named-content content-type="pre">e.g.</named-content></xref>. However, it has only been in the last few years that deep learning models have started to become competitive as self-standing medium-range and subseasonal large-scale forecasting tools. As late as 2021, in a popular review article, <xref ref-type="bibr" rid="bib1.bibx68" id="text.6"/> noted how deep learning research in the field of meteorology “is still in its infancy” and underscored that “a number of fundamental breakthroughs are needed” before deep learning applications may compete with physics-based weather forecasts.</p>
      <p id="d1e154">Much has changed since then. From early 2022, at least seven different research groups <xref ref-type="bibr" rid="bib1.bibx54 bib1.bibx11 bib1.bibx40 bib1.bibx46 bib1.bibx14 bib1.bibx52 bib1.bibx15" id="paren.7"/> claim to have developed deep learning models able to forecast key atmospheric variables with greater accuracy than deterministic forecasts from the European Centre for Medium-Range Weather Forecasts (ECMWF), which are widely regarded as the leading global numerical weather predictions. In addition to technical advances, a key contextual enabler of this explosive development has been the contribution of “Big Tech” – major private actors in the field of information technology <xref ref-type="bibr" rid="bib1.bibx5" id="paren.8"/>. This has contributed to closing the gap between state-of-the-art deep learning, cutting-edge computational resources, and weather practitioners, and it has attracted a larger number of machine learning experts to the field. Although only one of several developments within machine learning, we argue for a crucial role of Big Tech in the advent of the latest generation of large-scale deep learning weather forecast models, which notably require larger<?pagebreak page2348?> computational resources and more specialised knowledge than previous state-of-the-art models <xref ref-type="bibr" rid="bib1.bibx11 bib1.bibx46" id="paren.9"/>.</p>
      <p id="d1e166">Despite this astounding rise, deep learning models for weather forecasting still face a number of challenges. Some of these are well known. For example, deep learning approaches typically do not incorporate physical constraints <xref ref-type="bibr" rid="bib1.bibx60" id="paren.10"/>, which may lead to unphysical forecasts. Furthermore, deep learning models usually produce deterministic forecasts, making it hard to compute reasonable estimates of the uncertainty around their predictions  <xref ref-type="bibr" rid="bib1.bibx68" id="paren.11"/>. A less-studied challenge is that data-driven models have limited capabilities with respect to extrapolating at the edge of their training range or beyond  <xref ref-type="bibr" rid="bib1.bibx29" id="paren.12"/>. Thus, these models may not be as helpful as numerical models for investigating future climates <xref ref-type="bibr" rid="bib1.bibx63" id="paren.13"/> and, more prominently, might struggle with forecasting extreme weather events lying in the tails of a meteorological variable's distribution <xref ref-type="bibr" rid="bib1.bibx73" id="paren.14"/>. If unaddressed, the latter limitation is likely to hold back deep learning models from becoming a credible alternative to numerical, physics-driven forecasting models. Indeed, accurate predictions and early warnings of extreme weather play a key role in disaster prevention and mitigation <xref ref-type="bibr" rid="bib1.bibx75 bib1.bibx50" id="paren.15"/> and are crucial for several economically prominent activities, including but not limited to the energy and insurance sectors <xref ref-type="bibr" rid="bib1.bibx44" id="paren.16"><named-content content-type="pre">e.g.</named-content></xref>. Nonetheless, the pace of development of deep learning weather prediction models continues to be rapid, and a number of promising approaches are being developed to address the above challenges <xref ref-type="bibr" rid="bib1.bibx37 bib1.bibx11 bib1.bibx76 bib1.bibx18 bib1.bibx28 bib1.bibx20 bib1.bibx39" id="paren.17"><named-content content-type="pre">e.g.</named-content></xref>.</p>
      <p id="d1e198">This article reflects on the rise of medium-range weather forecasting with deep learning, the challenges currently being faced when forecasting extreme weather, and the future perspectives opened by the latest research advances. We do not consider in detail issues related to computing forecast uncertainty estimates <xref ref-type="bibr" rid="bib1.bibx65 bib1.bibx20" id="paren.18"/> or incorporating physical reasoning in deep learning models <xref ref-type="bibr" rid="bib1.bibx39 bib1.bibx9" id="paren.19"/>, for which we remand the reader to some recent review articles discussing these topics <xref ref-type="bibr" rid="bib1.bibx51 bib1.bibx22" id="paren.20"/>. We begin with a survey of recent developments in the field of large-scale deep learning weather prediction (DLWP), with a focus on the aforementioned models claiming to outperform deterministic state-of-the-art numerical weather prediction models <xref ref-type="bibr" rid="bib1.bibx54 bib1.bibx40 bib1.bibx11 bib1.bibx46 bib1.bibx14 bib1.bibx52 bib1.bibx15" id="paren.21"/>. Then, we provide a technical justification of why those models might struggle with predictions in the tails of the distribution, namely, weather extremes. Last, we outline alternative approaches that may be employed in order to design deep learning models specifically tailored to extreme weather forecasting.</p>
</sec>
<sec id="Ch1.S2">
  <label>2</label><title>Overview of DLWP models</title>
<sec id="Ch1.S2.SS1">
  <label>2.1</label><title>Early DLWP efforts</title>
      <p id="d1e228">The very first DLWP models were developed in the 1990s <xref ref-type="bibr" rid="bib1.bibx67 bib1.bibx31" id="paren.22"/> and followed a “feed-forward architecture” <xref ref-type="bibr" rid="bib1.bibx38" id="paren.23"/>, a unidirectional, non-recurrent structure in which the input is transmitted through the network sequentially. Feed-forward neural networks (FNNs) are limited in treating spatial data, due to their inability to leverage spatial patterns and their large computational burden, which makes them unsuitable for large datasets. For these reasons, FNNs were soon replaced by convolutional neural networks (CNNs; <xref ref-type="bibr" rid="bib1.bibx47" id="altparen.24"/>), which can learn spatial patterns and display better scalability. Early meteorological applications of CNNs had either a very local character <xref ref-type="bibr" rid="bib1.bibx78 bib1.bibx49 bib1.bibx30" id="paren.25"/> or were aimed at producing nowcasts with lead times from a few minutes to a few hours <xref ref-type="bibr" rid="bib1.bibx41 bib1.bibx56" id="paren.26"/>.</p>
      <p id="d1e246">A further step in the direction of today's medium-range DLWP models was taken with the adoption of recurrent neural network (RNN) architectures (<xref ref-type="bibr" rid="bib1.bibx61 bib1.bibx8 bib1.bibx7" id="altparen.27"/>) and, subsequently, long short-term memory (LSTM) models  (<xref ref-type="bibr" rid="bib1.bibx35" id="altparen.28"/>). These follow the dynamical nature of time-series data by making current observations of the variable of interest depend on previous iterations of that same variable. Thus, they provide an effective framework for accounting for time dependencies in the data and produce predictions on multiple timescales. However, due to their sequential, recursive nature, RNNs and LSTMs are hard to parallelise, preventing an effective exploitation of modern high-resolution climate reanalysis datasets, such as ERA5 <xref ref-type="bibr" rid="bib1.bibx32" id="paren.29"/>.</p>
      <?pagebreak page2349?><p id="d1e258">FNNs, CNNs, and RNNs/LSTMs are the cornerstones of deep learning and were also the dominant supervised learning architectures within DLWP at the time that the review by <xref ref-type="bibr" rid="bib1.bibx68" id="text.30"/> was written. Since then, a number of new architectures have been developed that address the limitations of the classical models via a number of creative innovations, often combining different elements of pre-existing architectures. Here, we focus primarily on those deep learning applications that are most relevant to medium-range, large-scale forecasting of weather extremes. Nonetheless, we acknowledge that data-driven nowcasting and subseasonal forecasting are both thriving fields of research with many potential applications, including for extreme events <xref ref-type="bibr" rid="bib1.bibx16 bib1.bibx3 bib1.bibx19" id="paren.31"><named-content content-type="pre">e.g.</named-content></xref>. <?xmltex \hack{\newpage}?></p>
</sec>
<sec id="Ch1.S2.SS2">
  <label>2.2</label><title>State-of-the-art DLWP</title>
      <p id="d1e278">A common element of current global medium-range DLWP models is the use of a large number of input variables (“features”) at a high temporal and spatial resolution. This is in contrast to older models that, to a large extent, relied on a theoretical understanding of atmospheric dynamics and feature selection for the choice of a few atmospheric variables and pressure levels <xref ref-type="bibr" rid="bib1.bibx24 bib1.bibx74" id="paren.32"><named-content content-type="pre">e.g.</named-content></xref>. For example, <xref ref-type="bibr" rid="bib1.bibx46" id="text.33"/> include 6 input variables on 37 pressure levels, 5 inputs at single levels, and several constant masks. Similarly, <xref ref-type="bibr" rid="bib1.bibx11" id="text.34"/> make use of 4 atmospheric variables at 13 pressure levels, 4 surface variables, and 3 constant masks.</p>
      <p id="d1e292">The fact that the latest DLWP models make use of a larger number of features than prior models may be partly ascribed to (1) computational improvements and (2) deep learning architectural developments. A key factor in this respect is the encoder–decoder architecture <xref ref-type="bibr" rid="bib1.bibx43 bib1.bibx17" id="paren.35"/>. Encoders and decoders can be seen as two separate neural networks connected to each other through a latent encoding vector. Here, “latent” refers to quantities inferred indirectly from the input data. The aim of the first network is to identify and compress (“encode”) into the encoding vector the most important features contained in the input data. The aim of the second network is to upscale (“decode”) the information encoded in the encoding vector until it reaches the dimensionality of the desired output. The target output can then either be the same as the input, perhaps with some small variation (self-supervised problems, e.g. variational autoencoders), or different from the input in terms of timescale, spatial resolution, or even actual features.</p>
      <p id="d1e298">In DLWP, the target output is usually different from the input, and encoders are mostly used to reduce the dimensionality of the input and identify the key latent features. This allows models to use very large input layers, namely, many different atmospheric variables at several pressure levels. Encoder–decoder architectures are used to this effect within cutting-edge global DLWP applications, such as <xref ref-type="bibr" rid="bib1.bibx40" id="text.36"/>, <xref ref-type="bibr" rid="bib1.bibx46" id="text.37"/>, <xref ref-type="bibr" rid="bib1.bibx11" id="text.38"/>, and <xref ref-type="bibr" rid="bib1.bibx14" id="text.39"/>.</p>
      <p id="d1e313">The encoder–decoder structure is also at the core of transformers <xref ref-type="bibr" rid="bib1.bibx72" id="paren.40"/>, a recent architectural innovation allowing for efficient parallelisation of sequential data. Transformers use a so-called attention mechanism <xref ref-type="bibr" rid="bib1.bibx1" id="paren.41"/>, i.e. they compute a score for each element in the input sequence which determines its relevance for the associated decoding step. This removes the need for sequential data intake, thereby enabling an effective utilisation of modern GPUs and TPUs for time-series data. This represents a major improvement over classic RNNs, which instead rely on serial arrangement to learn key features and are, therefore, not easily parallelisable.</p>
      <p id="d1e323">Recently, the use of transformers has been extended to computer vision tasks as an alternative, or complement, to CNNs. <xref ref-type="bibr" rid="bib1.bibx23" id="text.42"/> propose the use of vision transformers, which adapt transformers to visual tasks by introducing an innovative preprocessing step: images are first divided into patches of fixed <inline-formula><mml:math id="M1" display="inline"><mml:mrow><mml:mi>N</mml:mi><mml:mo>×</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:math></inline-formula> size and then run through a flattening layer, so that each patch can be treated as a separate token. Next, transformers are applied just as in sequential tasks. This approach has been applied to weather forecasting, for instance, by <xref ref-type="bibr" rid="bib1.bibx54" id="text.43"/> and <xref ref-type="bibr" rid="bib1.bibx11" id="text.44"/>, who use flattened patches of <inline-formula><mml:math id="M2" display="inline"><mml:mrow><mml:mn mathvariant="normal">4</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">4</mml:mn></mml:mrow></mml:math></inline-formula> pixels to apply transformers to gridded meteorological data.</p>
      <p id="d1e359">A distinct approach featured by several global DLWP models is the use of graph neural networks (GNNs; <xref ref-type="bibr" rid="bib1.bibx62" id="altparen.45"/>). Classic CNNs implicitly assume regular grids, in which the distance between points and the importance of each point is fixed <xref ref-type="bibr" rid="bib1.bibx71" id="paren.46"/>. This assumption is problematic in the case of global forecast models, as climate variables are often provided on regular latitude–longitude or reduced-Gaussian grids. Given that the Earth is quasi-spherical, the distance between degrees of latitude is greater at the Equator than at the poles, and even the length of degrees of longitude varies slightly with latitude. GNNs, unlike CNNs, allow for complex, quasi-spherical shapes. A way to understand this is by drawing a parallel with Cartesian and spherical coordinates. CNNs, similar to Cartesian coordinates, assume a “flat” grid and, in the best case, may introduce a weighting scheme not unlike the one of the <inline-formula><mml:math id="M3" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>-plane, whereas GNNs can model the relation between the nodes as a complex polygon resembling a sphere.</p>
      <p id="d1e375">The bearing of the architectural change from CNNs to GNNs on forecast performance is still the object of debate, and it is likely that it plays a larger role for global rather than local-to-regional applications, given the greater variation in the size of the grid cells in the former case. However, several recent deep learning weather forecasting models have introduced the use of GNNs to good effect. For instance, <xref ref-type="bibr" rid="bib1.bibx40" id="text.47"/> and, more prominently, <xref ref-type="bibr" rid="bib1.bibx46" id="text.48"/> have made use of GNNs to obtain accurate medium-range forecasts of several key atmospheric variables, managing to outperform the most accurate ECMWF deterministic forecasts available at the time of their publication.</p>
      <p id="d1e384">Other current approaches look at ways of accounting for Earth's quasi-spherical nature within a CNN framework, without resorting to GNNs. Examples of this include spherical convolutions <xref ref-type="bibr" rid="bib1.bibx12" id="paren.49"/> and spherical cross-correlations <xref ref-type="bibr" rid="bib1.bibx21" id="paren.50"/>. Recent work by <xref ref-type="bibr" rid="bib1.bibx66" id="text.51"/> showcases the advantages of spherical and hemispheric convolutions over classic CNNs. The authors compare models based on different architectures using the WeatherBench dataset <xref ref-type="bibr" rid="bib1.bibx57" id="paren.52"/> and show that models incorporating spherical or hemispheric convolutions produce more accurate medium-range forecasts of the 500 hPa geopotential height (Z500) and 850 hPa temperature (t850) than models featuring classic CNN architectures. However, feature-rich and high-resolution applications of this kind are still under development.</p>
      <?pagebreak page2350?><p id="d1e399">Finally, we outline how recent medium-range DLWP models treat temporal information. Instead of trying to incorporate the time aspect directly into the model in the form of extra features or channels, they account for the sequential nature of data through a dynamic approach, by using the predictions generated by a given model time step as the input for the next model time step. In other words, as clearly stated by <xref ref-type="bibr" rid="bib1.bibx14" id="text.53"/>, they approximate the forecast at time <inline-formula><mml:math id="M4" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M5" display="inline"><mml:mrow><mml:msup><mml:mi>Y</mml:mi><mml:mi>t</mml:mi></mml:msup><mml:mo>=</mml:mo><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:msup><mml:mi>Y</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mi>Y</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>+</mml:mo><mml:msup><mml:mi>Y</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> through an autoregressive approach of order 1 (AR 1), namely, by sequentially using the forecasts at the previous time steps: <inline-formula><mml:math id="M6" display="inline"><mml:mrow><mml:msup><mml:mi>Y</mml:mi><mml:mi>t</mml:mi></mml:msup><mml:mo>=</mml:mo><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:msup><mml:mi>Y</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:msup><mml:mi>Y</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:msup><mml:mi>Y</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:msup><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mi>Y</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>=</mml:mo><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:msup><mml:mi>Y</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. A similar approach is adopted by <xref ref-type="bibr" rid="bib1.bibx46" id="text.54"/>, with the main difference being that the data generated by the previous two forecasts are used as input for the latest forecast (AR 2).</p>
      <p id="d1e552">If the focus is on a specific lead time, it is, however, not clear whether this iterative approach always outperforms training a model to make a single prediction at the chosen lead time <xref ref-type="bibr" rid="bib1.bibx64" id="paren.55"/>. In this regard, <xref ref-type="bibr" rid="bib1.bibx15" id="text.56"/> show that producing a cascade of models fine-tuned on different timescales can lead to improvements in performance compared with a single model optimised on the whole forecasting window.</p>
      <p id="d1e562">An overview of the DLWP model developments over time is provided in Figs. <xref ref-type="fig" rid="Ch1.F1"/> and <xref ref-type="fig" rid="Ch1.F2"/>. Figure <xref ref-type="fig" rid="Ch1.F1"/> summarises the evolution of model architectures described in this section, while Fig. <xref ref-type="fig" rid="Ch1.F2"/> outlines the continuous improvements in the spatial and temporal domains handled by DLWP models. The current leading global DLWP models are systematically presented in Table <xref ref-type="table" rid="Ch1.T1"/>, where we provide information on the inputs, outputs, main architectural innovations, and performance for extreme weather forecasts of each model.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F1" specific-use="star"><?xmltex \currentcnt{1}?><?xmltex \def\figurename{Figure}?><label>Figure 1</label><caption><p id="d1e577">Evolution of deep learning weather prediction (DLWP) through time: from feed-forward neural networks to graph neural networks and vision transformers.</p></caption>
          <?xmltex \igopts{width=312.980315pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/17/2347/2024/gmd-17-2347-2024-f01.png"/>

        </fig>

      <?xmltex \floatpos{t}?><fig id="Ch1.F2"><?xmltex \currentcnt{2}?><?xmltex \def\figurename{Figure}?><label>Figure 2</label><caption><p id="d1e588">Evolution of the largest geographical and temporal scales of deep learning weather prediction (DLWP) models over time.</p></caption>
          <?xmltex \igopts{width=236.157874pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/17/2347/2024/gmd-17-2347-2024-f02.png"/>

        </fig>

<?xmltex \floatpos{p}?><table-wrap id="Ch1.T1" specific-use="star"><?xmltex \currentcnt{1}?><label>Table 1</label><caption><p id="d1e600">Overview of recent global medium-range DLWP applications.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="5">
     <oasis:colspec colnum="1" colname="col1" align="justify" colwidth="3cm"/>
     <oasis:colspec colnum="2" colname="col2" align="justify" colwidth="2.8cm"/>
     <oasis:colspec colnum="3" colname="col3" align="justify" colwidth="2.5cm"/>
     <oasis:colspec colnum="4" colname="col4" align="justify" colwidth="2.9cm"/>
     <oasis:colspec colnum="5" colname="col5" align="justify" colwidth="4.8cm"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Reference(s)</oasis:entry>
         <oasis:entry colname="col2">Inputs</oasis:entry>
         <oasis:entry colname="col3">Tested outputs</oasis:entry>
         <oasis:entry colname="col4">Main innovation</oasis:entry>
         <oasis:entry colname="col5">Performance on extreme values</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">
                      <xref ref-type="bibr" rid="bib1.bibx54" id="text.57"/>
                    </oasis:entry>
         <oasis:entry colname="col2">Single level – T2m, 10m U, 10m V, SP, MSLP, and IWV; multiple levels – <inline-formula><mml:math id="M16" display="inline"><mml:mi>Z</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M17" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M18" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M19" display="inline"><mml:mi>Q</mml:mi></mml:math></inline-formula>, and <inline-formula><mml:math id="M20" display="inline"><mml:mi>T</mml:mi></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">Z500</oasis:entry>
         <oasis:entry colname="col4">First paper with performance comparable to physics-based numerical models; use of vision transformers <xref ref-type="bibr" rid="bib1.bibx23" id="paren.58"/></oasis:entry>
         <oasis:entry colname="col5">Model evaluation on extreme quantiles, tends to underestimate high quantiles of 10 m zonal wind and total precipitation; source code and trained models available</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">
                      <xref ref-type="bibr" rid="bib1.bibx40" id="text.59"/>
                    </oasis:entry>
         <oasis:entry colname="col2">Static – lsm and orography; single level – solar radiation; multiple levels – <inline-formula><mml:math id="M21" display="inline"><mml:mi>Z</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M22" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M23" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M24" display="inline"><mml:mi>Q</mml:mi></mml:math></inline-formula>, and <inline-formula><mml:math id="M25" display="inline"><mml:mi>T</mml:mi></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">Z500, T850, wind speed 500, and RH700</oasis:entry>
         <oasis:entry colname="col4">Use of GNNs<?xmltex \hack{\hfill\break}?> <xref ref-type="bibr" rid="bib1.bibx4" id="paren.60"/></oasis:entry>
         <oasis:entry colname="col5">Unknown; source code available</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">
                      <xref ref-type="bibr" rid="bib1.bibx10 bib1.bibx11" id="text.61"/>
                    </oasis:entry>
         <oasis:entry colname="col2">Single level – T2m, 10m U, 10m V, and MSLP; multiple levels – <inline-formula><mml:math id="M26" display="inline"><mml:mi>Z</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M27" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M28" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M29" display="inline"><mml:mi>Q</mml:mi></mml:math></inline-formula>, and <inline-formula><mml:math id="M30" display="inline"><mml:mi>T</mml:mi></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">Z500, T500, Q500, U500, V500, Z850, T850, T2m, 10m U, and 10m V; MSLP for cyclone-tracking example</oasis:entry>
         <oasis:entry colname="col4">Three-dimensional vision transformer, hierarchical temporal aggregation to decrease computational burden</oasis:entry>
         <oasis:entry colname="col5">Better than HRES in binary detection of T2m extremes at a 6 d lead time despite the tendency to underestimate their magnitude <xref ref-type="bibr" rid="bib1.bibx6" id="paren.62"/>; more precise tracking of tropical cyclones than HRES in a case study; source code and trained models available</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">
                      <xref ref-type="bibr" rid="bib1.bibx45 bib1.bibx46" id="text.63"/>
                    </oasis:entry>
         <oasis:entry colname="col2">Static – lsm, orography, lat, and long; single level – T2m, 10m U, 10m V, MSLP, TP, solar radiation, h, and elapsed year progress; multiple levels – <inline-formula><mml:math id="M31" display="inline"><mml:mi>Z</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M32" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M33" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M34" display="inline"><mml:mi>W</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M35" display="inline"><mml:mi>Q</mml:mi></mml:math></inline-formula>, and <inline-formula><mml:math id="M36" display="inline"><mml:mi>T</mml:mi></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">Single level – T2m, 10m U, 10m V, and MSLP; multiple levels – <inline-formula><mml:math id="M37" display="inline"><mml:mi>Z</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M38" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M39" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M40" display="inline"><mml:mi>Q</mml:mi></mml:math></inline-formula>, and <inline-formula><mml:math id="M41" display="inline"><mml:mi>T</mml:mi></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">GraphCast,<?xmltex \hack{\hfill\break}?>GNN-based architecture <xref ref-type="bibr" rid="bib1.bibx4" id="paren.64"/>; much larger set of inputs and outputs than predecessors</oasis:entry>
         <oasis:entry colname="col5">More precise tracking of cyclones and atmospheric rivers than HRES at most lead times; better or comparable to HRES in binary detection of T2m extremes at 5 d; source code and trained models available</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">
                      <xref ref-type="bibr" rid="bib1.bibx14" id="text.65"/>
                    </oasis:entry>
         <oasis:entry colname="col2">Single level – T2m, 10m U, 10m V, and MSLP; multiple levels – <inline-formula><mml:math id="M42" display="inline"><mml:mi>Z</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M43" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M44" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula>, RH, and <inline-formula><mml:math id="M45" display="inline"><mml:mi>T</mml:mi></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">Z500, T500, U500, V500, Z850, T850, U850, V850, T2m, 10m U, and MSLP</oasis:entry>
         <oasis:entry colname="col4">Transformer with<?xmltex \hack{\hfill\break}?>encoder–fuse–decoder architecture</oasis:entry>
         <oasis:entry colname="col5">Unknown; trained model available</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">
                      <xref ref-type="bibr" rid="bib1.bibx52" id="text.66"/>
                    </oasis:entry>
         <oasis:entry colname="col2">Static – lsm and orography; single level – T2m, 10m U, and 10m V; multiple levels – <inline-formula><mml:math id="M46" display="inline"><mml:mi>Z</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M47" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M48" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M49" display="inline"><mml:mi>Q</mml:mi></mml:math></inline-formula>, RH, and <inline-formula><mml:math id="M50" display="inline"><mml:mi>T</mml:mi></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">Designed to allow for flexible outputs. Included example – Z500, T2m, T850, and 10m U</oasis:entry>
         <oasis:entry colname="col4">Variable-level<?xmltex \hack{\hfill\break}?>embedding and variable aggregation to allow for heterogenous input datasets</oasis:entry>
         <oasis:entry colname="col5">Unknown; source code available</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">
                      <xref ref-type="bibr" rid="bib1.bibx15" id="text.67"/>
                    </oasis:entry>
         <oasis:entry colname="col2">Single level – T2m, 10m U, 10m V, MSLP, and TP; multiple levels – <inline-formula><mml:math id="M51" display="inline"><mml:mi>Z</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M52" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M53" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula>, RH, and <inline-formula><mml:math id="M54" display="inline"><mml:mi>T</mml:mi></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">Z500, T500, U500, V500, T850, T2m, 10m U, 10m V, and MSLP</oasis:entry>
         <oasis:entry colname="col4">Cascade model architecture with separate fine-tuning for different forecasting windows</oasis:entry>
         <oasis:entry colname="col5">Source code and trained model available; FuXi-Extreme, a version of the model optimised for extreme weather, currently under development <xref ref-type="bibr" rid="bib1.bibx77" id="paren.68"/></oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table><table-wrap-foot><p id="d1e603">The abbreviations used in the table are as follows: 10m U and 10m V denote the <inline-formula><mml:math id="M7" display="inline"><mml:mi>u</mml:mi></mml:math></inline-formula> wind and <inline-formula><mml:math id="M8" display="inline"><mml:mi>v</mml:mi></mml:math></inline-formula> wind at 10 m, respectively; lsm – land–sea mask; MSLP – mean sea-level pressure; RH – relative humidity; <inline-formula><mml:math id="M9" display="inline"><mml:mi>Q</mml:mi></mml:math></inline-formula> – specific humidity; SP – surface pressure; <inline-formula><mml:math id="M10" display="inline"><mml:mi>T</mml:mi></mml:math></inline-formula> – temperature; IWV – integrated total column water vapour; <inline-formula><mml:math id="M11" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula> – <inline-formula><mml:math id="M12" display="inline"><mml:mi>u</mml:mi></mml:math></inline-formula> wind; <inline-formula><mml:math id="M13" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula> – <inline-formula><mml:math id="M14" display="inline"><mml:mi>v</mml:mi></mml:math></inline-formula> wind; <inline-formula><mml:math id="M15" display="inline"><mml:mi>Z</mml:mi></mml:math></inline-formula> – geopotential; TP – total precipitation; lat – latitude; long – longitude; h – hour of the day; 2m – 2 m height; 10m – 10 m height; 500 – 500 hPa; 850 – 850 hPa; and HRES – ECMWF high-resolution deterministic forecast.</p></table-wrap-foot><?xmltex \gdef\@currentlabel{1}?></table-wrap>

</sec>
</sec>
<sec id="Ch1.S3">
  <label>3</label><title>Challenges and opportunities</title>
<sec id="Ch1.S3.SS1">
  <label>3.1</label><title>Current challenges in DLWP</title>
      <p id="d1e1160">A common limitation of most large-scale DLWP applications introduced so far is that they are not targeted in any specific way at extreme weather events; rather, their focus lies on maximising the average skill of the forecasts. Typically, machine learning models struggle to make accurate predictions of extreme values, partly due to (1) the inherently limited training samples for extreme values and (2) the intrinsic inferential challenges related to extrapolation. Given the key role of accurate prediction and of early warnings for extreme weather in disaster prevention and risk mitigation <xref ref-type="bibr" rid="bib1.bibx75 bib1.bibx50" id="paren.69"/>, it would be desirable for current DLWP applications to dedicate greater attention to forecast skill for extreme weather <xref ref-type="bibr" rid="bib1.bibx73" id="paren.70"/>.</p>
      <p id="d1e1169">As highlighted in Table <xref ref-type="table" rid="Ch1.T1"/>, this problem is further exacerbated by the fact that many global DLWP papers provide no or very limited diagnostics on the performance of their models for extreme weather scenarios <xref ref-type="bibr" rid="bib1.bibx40 bib1.bibx14 bib1.bibx52" id="paren.71"><named-content content-type="pre">e.g.</named-content></xref>, making it hard to assess their performance in those situations. Even those that do provide extreme weather diagnostics mostly focus on selected variables and case studies, supplying no systematic overview of how the models perform in the prediction of high-impact surface extremes such as total precipitation or peak wind gusts. Indeed, some state-of-the-art DLWP models, such as <xref ref-type="bibr" rid="bib1.bibx11" id="text.72"/>, do not even produce forecasts for those variables.</p>
      <p id="d1e1182"><xref ref-type="bibr" rid="bib1.bibx73" id="text.73"/> suggests some simple measures that authors could adopt to help readers evaluate whether or not a machine learning model can provide robust forecasts of extreme events: for instance, that all papers should include scatterplots and quantile–quantile plots of forecasted vs. observed values and that performance metrics computed only on extreme values should complement classic metrics of average skill. A positive note since the release of <xref ref-type="bibr" rid="bib1.bibx73" id="text.74"/> is that several research groups have chosen to make the code of their global models publicly available, making it possible for third-party actors with enough computational resources to implement and further test their models. For instance, ECMWF has recently launched an experimental programme running daily 10 d forecasts with 6-hourly time steps of the models introduced by <xref ref-type="bibr" rid="bib1.bibx54" id="text.75"/>, <xref ref-type="bibr" rid="bib1.bibx11" id="text.76"/>, and <xref ref-type="bibr" rid="bib1.bibx46" id="text.77"/>, whose forecasts are available to the general public <xref ref-type="bibr" rid="bib1.bibx25" id="paren.78"/>. Similarly, WeatherBench 2 <xref ref-type="bibr" rid="bib1.bibx58" id="paren.79"/> provides additional scorecards and out-of-sample predictions for several models included in Table <xref ref-type="table" rid="Ch1.T1"/>.</p>
      <?pagebreak page2352?><p id="d1e1208">Some key “inductive biases”, i.e. implicit assumptions of the employed estimation techniques <xref ref-type="bibr" rid="bib1.bibx4" id="paren.80"/>, may also hamper the performance of current DLWP applications for extreme weather forecasting. Most global DLWP models choose to minimise the overall mean-squared error (L2) of the forecast, averaging over all grid points and time steps of interest <xref ref-type="bibr" rid="bib1.bibx54 bib1.bibx40" id="paren.81"/>. The minimisation thus uses the conditional mean of the dependent variable through space and time given the predictors, optimising forecasts for mean rather than extreme values. Furthermore, the use of L2 (and also L1, the mean absolute error, used, for instance, by <xref ref-type="bibr" rid="bib1.bibx11" id="altparen.82"/>, and <xref ref-type="bibr" rid="bib1.bibx15" id="altparen.83"/>) implicitly assumes that, for any given variable, the distribution of the forecast error is symmetric, i.e. that it is possible to obtain both positive and negative errors of the same magnitude, and that deviations from the modelled value in the two directions are equally important. This is seldom the case in weather forecasting. Many weather variables display a high degree of autocorrelation and follow highly asymmetric truncated distributions (e.g. peak wind speed or precipitation), which in combination tend to produce non-asymmetric error distributions  <xref ref-type="bibr" rid="bib1.bibx36" id="paren.84"/>. Moreover, deviations of a variable from its mean in one of the two directions can have larger impacts on human societies than deviations in the other direction (e.g. one would expect that severely underestimating the amount of rain in a flash-flood event would be more harmful than incorrectly predicting rain on a dry day).</p>
      <p id="d1e1227">While the suggestions in <xref ref-type="bibr" rid="bib1.bibx73" id="text.85"/>, if implemented, would go a long way in ensuring greater transparency and credibility for DLWP forecasts of extreme weather, the challenges related to the extrapolation issue and inductive biases still remain. The limited diagnostics provided by <xref ref-type="bibr" rid="bib1.bibx54" id="text.86"/> and <xref ref-type="bibr" rid="bib1.bibx10" id="text.87"/> suggest that their models perform reasonably well on extremes but also that they consistently tend to underestimate their magnitude. Similarly, <xref ref-type="bibr" rid="bib1.bibx6" id="text.88"/> find that Pangu-Weather  <xref ref-type="bibr" rid="bib1.bibx11" id="paren.89"/> can provide high-quality binary forecasts of moderately extreme temperatures but that it also tends to oversmooth the prediction and underestimate the magnitude of the largest cold and hot extremes.</p>
      <p id="d1e1245">Thus, we argue for the need for DLWP models explicitly built to forecast extremes. These should make use of targeted loss functions and produce robust predictions of all relevant variables at or beyond the limits of their training range. In the next section, we propose a schematic framework on which to build such models. The aim here is to provide a general foundation for such approaches, rather than discussing architectural details. Indeed, several of the architectures adopted by the models in Table <xref ref-type="table" rid="Ch1.T1"/> and Sect. 2.1 are, in principle, equally suitable for predictions of extreme weather and average weather. The limiting factors are most likely the choice of the optimisation problem and the lack of a specific treatment of the extremes, rather than the architectures themselves.</p>
</sec>
<?pagebreak page2353?><sec id="Ch1.S3.SS2">
  <label>3.2</label><title>A DLWP workflow for extreme weather</title>
      <p id="d1e1258">A simple way of shifting the focus from the average skill of a deep learning model to its performance in the tails of the distribution is by changing its loss function. Common loss functions, such as the mean absolute error (L1) and the mean-squared error (L2), are minimised by taking the conditional median and mean of the dependent variable, respectively. An alternative loss function is given by the pinball loss, defined as follows <xref ref-type="bibr" rid="bib1.bibx42" id="paren.90"/>:
            <disp-formula id="Ch1.E1" content-type="numbered"><label>1</label><mml:math id="M55" display="block"><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mi mathvariant="normal">pinball</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mi>N</mml:mi></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:munderover><mml:mo movablelimits="false">max⁡</mml:mo><mml:mo>(</mml:mo><mml:mi mathvariant="italic">τ</mml:mi><mml:mo>⋅</mml:mo><mml:mo>(</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>y</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:mo>(</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:mi mathvariant="italic">τ</mml:mi><mml:mo>)</mml:mo><mml:mo>⋅</mml:mo><mml:mo>(</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>y</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mi>i</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
          where <inline-formula><mml:math id="M56" display="inline"><mml:mi mathvariant="italic">τ</mml:mi></mml:math></inline-formula> is the target quantile, <inline-formula><mml:math id="M57" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula> is the number of training observations, <inline-formula><mml:math id="M58" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula> represents a specific observation, <inline-formula><mml:math id="M59" display="inline"><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the actual value of the target variable for that observation, and <inline-formula><mml:math id="M60" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi>y</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the forecast generated by the model. This loss function punishes predictions that are further away from the quantile of interest and is minimised by the conditional quantile of the dependent variable.<fn id="Ch1.Footn1"><p id="d1e1408">As the median is the 50th quantile, the pinball loss is equivalent to L1 when choosing <inline-formula><mml:math id="M61" display="inline"><mml:mrow><mml:mi mathvariant="italic">τ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.5</mml:mn></mml:mrow></mml:math></inline-formula>.</p></fn> By choosing an extreme quantile of interest, it is possible to study, in a regression setting, how different predictors affect the tails of the distribution. Furthermore, models minimising the pinball loss could be used to set approximate confidence intervals around models maximising the average skill of the prediction.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F3" specific-use="star"><?xmltex \currentcnt{3}?><?xmltex \def\figurename{Figure}?><label>Figure 3</label><caption><p id="d1e1426">Extreme event prediction model design workflow. The chosen approach should depend on the information one aims to gather and the return period of the extreme events of interest.</p></caption>
          <?xmltex \igopts{width=312.980315pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/17/2347/2024/gmd-17-2347-2024-f03.png"/>

        </fig>

      <p id="d1e1435">Within a deep learning setting, models minimising the pinball loss often go under the name of deep quantile regression or quantile regression neural networks <xref ref-type="bibr" rid="bib1.bibx70" id="paren.91"/>. A limitation of deep quantile regression is that enough observations below and above the quantile of interest need to be available for the model to work properly. This can sometimes be an issue within a DLWP framework, given that the main interest can lie in very extreme quantiles, i.e. seldom-observed extreme events with long return periods.</p>
      <p id="d1e1442">A solution to this problem has recently been proposed by <xref ref-type="bibr" rid="bib1.bibx53" id="text.92"/>, who, building upon earlier work by <xref ref-type="bibr" rid="bib1.bibx13" id="text.93"/>, suggest using a two-step peak-over-threshold approach. First, a quantile-regression-based estimator, such as linear or deep quantile regression, is used to estimate a conditional threshold of interest. Then, the properties of the distribution of the exceedances are modelled with the help of extreme value theory (EVT). <xref ref-type="bibr" rid="bib1.bibx53" id="text.94"/> assume that, in accordance with <xref ref-type="bibr" rid="bib1.bibx2" id="text.95"/> and <xref ref-type="bibr" rid="bib1.bibx55" id="text.96"/>, independent exceedances approximately follow a generalised Pareto distribution, with parameters depending on the value of the regressors. These parameters can then be estimated with the help of a neural network, and the resulting empirical distribution can be used to derive the properties of the distribution of any extreme event of interest, as is commonly done in EVT.</p>
      <p id="d1e1460">However, even the combination of quantile and EVT-based approaches suffers from a key limitation: it does not provide a deterministic forecast for a given time and place but only return periods or values and a risk ratio of the probability of an event taking place compared to the climatology. In other words, it answers questions such as “Is extreme event `<inline-formula><mml:math id="M62" display="inline"><mml:mi>X</mml:mi></mml:math></inline-formula>' more likely to occur than usual on day `<inline-formula><mml:math id="M63" display="inline"><mml:mi>Y</mml:mi></mml:math></inline-formula>'?” or “How often does an event of a given severity occur given an initial set of atmospheric conditions?”. It does not answer the question typically associated with deterministic weather forecasts, namely, “Is extreme event `<inline-formula><mml:math id="M64" display="inline"><mml:mi>X</mml:mi></mml:math></inline-formula>' going to take place on day `<inline-formula><mml:math id="M65" display="inline"><mml:mi>Y</mml:mi></mml:math></inline-formula>' at location `<inline-formula><mml:math id="M66" display="inline"><mml:mi>Z</mml:mi></mml:math></inline-formula>'?”.</p>
      <p id="d1e1498">A possible alternative for cases in which we are interested in answering the latter question, i.e. we want a deterministic forecast of a given extreme event at a specific time and place, is to use a binary classification model. This can, for instance, minimise a binary cross-entropy loss, defined as follows:
            <disp-formula id="Ch1.E2" content-type="numbered"><label>2</label><mml:math id="M67" display="block"><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mi mathvariant="normal">bincross</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mi>N</mml:mi></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:munderover><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>⋅</mml:mo><mml:mi>log⁡</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>+</mml:mo><mml:mo>(</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>⋅</mml:mo><mml:mi>log⁡</mml:mi><mml:mo>(</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>
          By defining the extreme event on the basis of a threshold or a given quantile of the climatology and minimising Eq. (<xref ref-type="disp-formula" rid="Ch1.E2"/>), we can train a model to estimate the probability of an event of a given magnitude taking place at a specific time and<?pagebreak page2354?> location. The forecasted probability for a specific time and place can then easily be converted into a deterministic forecast by choosing a cutoff probability (e.g. 50 %), where events above that probability are expected to take place and events under that probability are not.</p>
      <p id="d1e1588">Whenever a heavy class imbalance is present, i.e. the interest lies in very extreme quantiles, training a classification neural network may be challenging, as the model may be prone to reverting to the trivial solution of never predicting an extreme event. In those cases, class weights may help: weights are introduced in Eq. (<xref ref-type="disp-formula" rid="Ch1.E2"/>) in order to give greater importance to the loss generated by training samples from the minority class, namely, the extremes. In other words, the model is trained to minimise a weighted cross-entropy loss (Eq. <xref ref-type="disp-formula" rid="Ch1.E3"/>), defined as follows:
            <disp-formula id="Ch1.E3" content-type="numbered"><label>3</label><mml:math id="M68" display="block"><mml:mtable class="split" rowspacing="0.2ex" displaystyle="true" columnalign="right left"><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mi mathvariant="normal">weightedbincross</mml:mi></mml:msub></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mi>N</mml:mi></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:munderover><mml:mo>-</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>⋅</mml:mo><mml:mi>log⁡</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mo>+</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>⋅</mml:mo><mml:mi>log⁡</mml:mi><mml:mo>(</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
          where <inline-formula><mml:math id="M69" display="inline"><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> is the weight assigned to observations in the minority class and <inline-formula><mml:math id="M70" display="inline"><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> is the weight assigned to observations in the majority class.</p>
      <p id="d1e1722">Even after introducing class weights, this approach, like quantile regression, needs enough observations in each of the two classes for the model to work properly. Thus, it is not suitable in isolation for extremes with a very long return period, appearing no or very few times in the training sample. However, as in the previous case, we can build a two-step peak-over-threshold model that addresses this problem. First, we decide through a classification model whether or not an event above a certain not-too-extreme threshold is going to take place. Then, we model the tails of the distribution with the help of the Pickands–Balkema–De Haan theorem <xref ref-type="bibr" rid="bib1.bibx2 bib1.bibx55" id="paren.97"/>, which allows one to make inferences on very extreme cases potentially beyond the model's training range.</p>
      <p id="d1e1728">The different steps introduced above can be combined in order to obtain forecasts providing a rich set of information.<?pagebreak page2355?> For instance, one may jointly implement a classification-based and a quantile-based deep learning model to obtain time- and location-specific forecasts of an extreme as well as information on its return period. Figure <xref ref-type="fig" rid="Ch1.F3"/> summarises the approaches described in this section in a simple framework that can be used to tailor deep learning models to extreme weather forecasting.</p>
</sec>
</sec>
<sec id="Ch1.S4" sec-type="conclusions">
  <label>4</label><title>Conclusions and recommendations</title>
      <p id="d1e1742">Accurate prediction of extreme weather events is a central part of a high-quality medium-range weather forecast and is, thus, of great societal and economic relevance <xref ref-type="bibr" rid="bib1.bibx75 bib1.bibx50" id="paren.98"/>. In order for global end-to-end deep learning models to attain widespread operational use, we argue that achieving greater average skill than physics-based numerical weather prediction models is not sufficient. They additionally need to demonstrate skill for extreme weather events.</p>
      <p id="d1e1748">We identify two key limitations that constrain current state-of-the-art deep learning forecasts of extreme weather. First, current architectures are not optimised to make use of the limited training samples for extreme values. Second, the models are not optimised on extreme event forecasts and make some simplistic assumptions regarding how the forecasting errors are distributed. These issues are compounded by the scant or missing validation of extreme weather forecasts provided by leading global DLWP models.</p>
      <p id="d1e1751">We argue for the urgency of a DLWP workflow targeted to extreme weather forecasts, whereby deep learning models specifically designed to handle extreme events should complement deep learning models maximising the average skill of the forecast. To enable rapid advances, the implementation of such a workflow should rest on adapting existing deep learning architectures, rather than developing radically new and untested approaches. This should be complemented by placing a greater emphasis on assessing the performance of existing and future models in the tails of the distributions of the forecasted variables <xref ref-type="bibr" rid="bib1.bibx73" id="paren.99"/>.</p>
      <p id="d1e1757">Echoing the above recommendations, in this article, we have proposed a foundational workflow to advance deep learning extreme weather forecasts, in which the method of choice depends on the meteorological question to be answered – whether probabilistic or deterministic – and the return period of the extreme events of interest. The workflow is fully enabled by recent architectural advances in deep learning weather forecast models; thus, we envision it as functional to achieve robust deep learning forecasts of extreme weather in the near future.</p>
</sec>

      
      </body>
    <back><notes notes-type="codedataavailability"><title>Code and data availability</title>

      <p id="d1e1765">The authors of some of the global deep learning models presented in Table <xref ref-type="table" rid="Ch1.T1"/> have chosen to make the code used to build their models freely available; whenever this is the case, we mention it in Table <xref ref-type="table" rid="Ch1.T1"/>. Most of these models are trained using the ERA5 reanalysis dataset <xref ref-type="bibr" rid="bib1.bibx32" id="paren.100"/>, which is freely available through the Copernicus Climate Change Service at <ext-link xlink:href="https://doi.org/10.24381/cds.adbb2d47" ext-link-type="DOI">10.24381/cds.adbb2d47</ext-link> <xref ref-type="bibr" rid="bib1.bibx33" id="paren.101"/> and <ext-link xlink:href="https://doi.org/10.24381/cds.bd0915c6" ext-link-type="DOI">10.24381/cds.bd0915c6</ext-link> <xref ref-type="bibr" rid="bib1.bibx34" id="paren.102"/>.</p>
  </notes><notes notes-type="authorcontribution"><title>Author contributions</title>

      <p id="d1e1791">The authors are jointly responsible for the conceptualisation of this work, including the visualisations, and all of the revision and editing of the submitted manuscript. Additionally, LO researched existing literature on the topic, wrote most of the original draft, and created the visualisations. GM acquired the funding and other resources necessary to conduct the research and provided extensive supervision.</p>
  </notes><notes notes-type="competinginterests"><title>Competing interests</title>

      <p id="d1e1797">The contact author has declared that neither of the authors has any competing interests.</p>
  </notes><notes notes-type="disclaimer"><title>Disclaimer</title>

      <p id="d1e1803">Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. While Copernicus Publications makes every effort to include appropriate place names, the final responsibility lies with the authors.</p>
  </notes><ack><title>Acknowledgements</title><p id="d1e1809">The authors thankfully acknowledge support from the European Research Council (ERC) under the European Union's Horizon 2020 Research and Innovation programme (project CENÆ: “Compound Climate Extremes in North America and Europe: from dynamics to predictability”; grant agreement no. 948309). The authors also wish to acknowledge the use of the online tool NN-SWG <xref ref-type="bibr" rid="bib1.bibx48" id="paren.103"/> for plotting some of the neural networks included in Fig. <xref ref-type="fig" rid="Ch1.F1"/>.</p></ack><notes notes-type="financialsupport"><title>Financial support</title>

      <p id="d1e1819">This research has been supported by the European Research Council, Horizon 2020 Research and Innovation programme (grant no. 948309).</p>
  </notes><notes notes-type="reviewstatement"><title>Review statement</title>

      <p id="d1e1825">This paper was edited by Paul Ullrich and reviewed by two anonymous referees.</p>
  </notes><ref-list>
    <title>References</title>

      <ref id="bib1.bibx1"><?xmltex \def\ref@label{{Bahdanau et~al.(2015)}}?><label>Bahdanau et al.(2015)</label><?label bahdanau_neural_2015?><mixed-citation>Bahdanau, D., Cho, K., and Bengio, Y.: Neural Machine Translation by Jointly Learning to Align and Translate, in: 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, 7–9 May 2015, Conference Track Proceedings, edited by: Bengio, Y. and LeCun, Y., <ext-link xlink:href="https://doi.org/10.48550/arXiv.1409.0473" ext-link-type="DOI">10.48550/arXiv.1409.0473</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx2"><?xmltex \def\ref@label{{Balkema and De~Haan(1974)}}?><label>Balkema and De Haan(1974)</label><?label balkema_residual_1974?><mixed-citation>Balkema, A. A. and De Haan, L.: Residual Life Time at Great Age, Ann. Probab., 2, 792–804, <ext-link xlink:href="https://doi.org/10.1214/aop/1176996548" ext-link-type="DOI">10.1214/aop/1176996548</ext-link>, 1974.</mixed-citation></ref>
      <ref id="bib1.bibx3"><?xmltex \def\ref@label{{Barnes et~al.(2023)}}?><label>Barnes et al.(2023)</label><?label barnes_forecasting_2023?><mixed-citation>Barnes, A. P., McCullen, N., and Kjeldsen, T. R.: Forecasting seasonal to sub-seasonal rainfall in Great Britain using convolutional-neural networks, Theor. Appl. Climatol., 151, 421–432, <ext-link xlink:href="https://doi.org/10.1007/s00704-022-04242-x" ext-link-type="DOI">10.1007/s00704-022-04242-x</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx4"><?xmltex \def\ref@label{{Battaglia et~al.(2018)}}?><label>Battaglia et al.(2018)</label><?label battaglia_relational_2018?><mixed-citation>Battaglia, P. W., Hamrick, J. B., Bapst, V., Sanchez-Gonzalez, A., Zambaldi, V., Malinowski, M., Tacchetti, A., Raposo, D., Santoro, A., Faulkner, R., Gulcehre, C., Song, F., Ballard, A., Gilmer, J., Dahl, G., Vaswani, A., Allen, K., Nash, C., Langston, V., Dyer, C., Heess, N., Wierstra, D., Kohli, P., Botvinick, M., Vinyals, O., Li, Y., and Pascanu, R.: Relational inductive biases, deep learning, and graph networks, arXiv, <ext-link xlink:href="https://doi.org/10.48550/arXiv.1806.01261" ext-link-type="DOI">10.48550/arXiv.1806.01261</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx5"><?xmltex \def\ref@label{{Bauer et~al.(2023)}}?><label>Bauer et al.(2023)</label><?label bauer_deep_2023?><mixed-citation>Bauer, P., Dueben, P., Chantry, M., Doblas-Reyes, F., Hoefler, T., McGovern, A., and Stevens, B.: Deep learning and a changing economy in weather and climate prediction, Nat. Rev. Earth Environ., 4, 507–509, <ext-link xlink:href="https://doi.org/10.1038/s43017-023-00468-z" ext-link-type="DOI">10.1038/s43017-023-00468-z</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx6"><?xmltex \def\ref@label{{Ben-Bouallegue et~al.(2023)}}?><label>Ben-Bouallegue et al.(2023)</label><?label ben-bouallegue_rise_2023?><mixed-citation>Ben-Bouallegue, Z., Clare, M. C. A., Magnusson, L., Gascon, E., Maier-Gerber, M., Janousek, M., Rodwell, M., Pinault, F., Dramsch, J. S., Lang, S. T. K., Raoult, B., Rabier, F., Chevallier, M., Sandu, I., Dueben, P., Chantry, M., and Pappenberger, F.: The rise of data-driven weather forecasting, arXiv, <ext-link xlink:href="https://doi.org/10.48550/arXiv.2307.10128" ext-link-type="DOI">10.48550/arXiv.2307.10128</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx7"><?xmltex \def\ref@label{{Bengio and Gingras(1995)}}?><label>Bengio and Gingras(1995)</label><?label bengio_recurrent_1995?><mixed-citation>Bengio, Y. and Gingras, F.: Recurrent Neural Networks for Missing or Asynchronous Data, in: Advances in Neural Information Processing Systems, vol. 8, MIT Press, <uri>https://papers.nips.cc/paper_files/paper/1995/hash/ffeed84c7cb1ae7bf4ec4bd78275bb98-Abstract.html</uri> (last access: 18 March 2024), 1995.</mixed-citation></ref>
      <ref id="bib1.bibx8"><?xmltex \def\ref@label{{Bengio et~al.(1994)}}?><label>Bengio et al.(1994)</label><?label bengio_learning_1994?><mixed-citation>Bengio, Y., Simard, P., and Frasconi, P.: Learning long-term dependencies with gradient descent is difficult, IEEE transactions on neural networks / a publication of the IEEE Neural Networks Council, 5, 157–66, <ext-link xlink:href="https://doi.org/10.1109/72.279181" ext-link-type="DOI">10.1109/72.279181</ext-link>, 1994.</mixed-citation></ref>
      <ref id="bib1.bibx9"><?xmltex \def\ref@label{{Beucler et~al.(2020)}}?><label>Beucler et al.(2020)</label><?label beucler_towards_2020?><mixed-citation>Beucler, T., Pritchard, M., Gentine, P., and Rasp, S.: Towards Physically-Consistent, Data-Driven Models of Convection, in: IGARSS 2020 – 2020 IEEE International Geoscience and Remote Sensing Symposium, 3987–3990, <ext-link xlink:href="https://doi.org/10.1109/IGARSS39084.2020.9324569" ext-link-type="DOI">10.1109/IGARSS39084.2020.9324569</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx10"><?xmltex \def\ref@label{{Bi et~al.(2022)}}?><label>Bi et al.(2022)</label><?label bi_pangu-weather_2022?><mixed-citation>Bi, K., Xie, L., Zhang, H., Chen, X., Gu, X., and Tian, Q.: Pangu-Weather: A 3D High-Resolution Model for Fast and Accurate Global Weather Forecast, arXiv, <ext-link xlink:href="https://doi.org/10.48550/arXiv.2211.02556" ext-link-type="DOI">10.48550/arXiv.2211.02556</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx11"><?xmltex \def\ref@label{{Bi et~al.(2023)}}?><label>Bi et al.(2023)</label><?label bi_accurate_2023?><mixed-citation>Bi, K., Xie, L., Zhang, H., Chen, X., Gu, X., and Tian, Q.: Accurate medium-range global weather forecasting with 3D neural networks, Nature, 619, 1–6, <ext-link xlink:href="https://doi.org/10.1038/s41586-023-06185-3" ext-link-type="DOI">10.1038/s41586-023-06185-3</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx12"><?xmltex \def\ref@label{{Boomsma and Frellsen(2017)}}?><label>Boomsma and Frellsen(2017)</label><?label boomsma_spherical_2017?><mixed-citation>Boomsma, W. and Frellsen, J.: Spherical convolutions and their application in molecular modelling, in: Advances in Neural Information Processing Systems, 30, Curran Associates, Inc., <uri>https://papers.nips.cc/paper_files/paper/2017/hash/1113d7a76ffceca1bb350bfe145467c6-Abstract.html</uri> (last access: 18 March 2024), 2017.</mixed-citation></ref>
      <ref id="bib1.bibx13"><?xmltex \def\ref@label{{Carreau and Bengio(2007)}}?><label>Carreau and Bengio(2007)</label><?label carreau_hybrid_2007?><mixed-citation>Carreau, J. and Bengio, Y.: A Hybrid Pareto Model for Conditional Density Estimation of Asymmetric Fat-Tail Data, in: Proceedings of the Eleventh International Conference on Artificial Intelligence and Statistics,  51–58, PMLR, <uri>https://proceedings.mlr.press/v2/carreau07a.html</uri> (last access: 18 March 2024), 2007.</mixed-citation></ref>
      <ref id="bib1.bibx14"><?xmltex \def\ref@label{{Chen et~al.(2023{\natexlab{a}})}}?><label>Chen et al.(2023a)</label><?label chen_fengwu_2023?><mixed-citation>Chen, K., Han, T., Gong, J., Bai, L., Ling, F., Luo, J.-J., Chen, X., Ma, L., Zhang, T., Su, R., Ci, Y., Li, B., Yang, X., and Ouyang, W.: FengWu: Pushing the Skillful Global Medium-range Weather Forecast beyond 10 Days Lead, arXiv, <ext-link xlink:href="https://doi.org/10.48550/arXiv.2304.02948" ext-link-type="DOI">10.48550/arXiv.2304.02948</ext-link>, 2023a.</mixed-citation></ref>
      <ref id="bib1.bibx15"><?xmltex \def\ref@label{{Chen et~al.(2023{\natexlab{b}})}}?><label>Chen et al.(2023b)</label><?label chen_fuxi_2023?><mixed-citation>Chen, L., Zhong, X., Zhang, F., Cheng, Y., Xu, Y., Qi, Y., and Li, H.: FuXi: a cascade machine learning forecasting system for 15-day global weather forecast, npj Climate and Atmospheric Science, 6, 1–11, <ext-link xlink:href="https://doi.org/10.1038/s41612-023-00512-1" ext-link-type="DOI">10.1038/s41612-023-00512-1</ext-link>, 2023b.</mixed-citation></ref>
      <ref id="bib1.bibx16"><?xmltex \def\ref@label{{Chkeir et~al.(2023)}}?><label>Chkeir et al.(2023)</label><?label chkeir_nowcasting_2023?><mixed-citation>Chkeir, S., Anesiadou, A., Mascitelli, A., and Biondi, R.: Nowcasting extreme rain and extreme wind speed with machine learning techniques applied to different input datasets, Atmos. Res., 282, 106548, <ext-link xlink:href="https://doi.org/10.1016/j.atmosres.2022.106548" ext-link-type="DOI">10.1016/j.atmosres.2022.106548</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx17"><?xmltex \def\ref@label{{Cho et~al.(2014)}}?><label>Cho et al.(2014)</label><?label cho_properties_2014?><mixed-citation>Cho, K., van Merriënboer, B., Bahdanau, D., and Bengio, Y.: On the Properties of Neural Machine Translation: Encoder–Decoder Approaches, in: Proceedings of SSST-8, Eighth Workshop on Syntax, Semantics and Structure in Statistical Translation, 103–111, Association for Computational Linguistics, Doha, Qatar, <ext-link xlink:href="https://doi.org/10.3115/v1/W14-4012" ext-link-type="DOI">10.3115/v1/W14-4012</ext-link>, 2014.</mixed-citation></ref>
      <ref id="bib1.bibx18"><?xmltex \def\ref@label{{Cisneros et~al.(2023)}}?><label>Cisneros et al.(2023)</label><?label cisneros_deep_2023?><mixed-citation>Cisneros, D., Richards, J., Dahal, A., Lombardo, L., and Huser, R.: Deep graphical regression for jointly moderate and extreme Australian wildfires, arXiv, <ext-link xlink:href="https://doi.org/10.48550/arXiv.2308.14547" ext-link-type="DOI">10.48550/arXiv.2308.14547</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx19"><?xmltex \def\ref@label{{Civitarese et~al.(2021)}}?><label>Civitarese et al.(2021)</label><?label civitarese_extreme_2021?><mixed-citation>Civitarese, D. S., Szwarcman, D., Zadrozny, B., and Watson, C.: Extreme Precipitation Seasonal Forecast Using a Transformer Neural Network, arXiv, <ext-link xlink:href="https://doi.org/10.48550/arXiv.2107.06846" ext-link-type="DOI">10.48550/arXiv.2107.06846</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx20"><?xmltex \def\ref@label{{Clare et~al.(2021)}}?><label>Clare et al.(2021)</label><?label clare_combining_2021?><mixed-citation>Clare, M. C., Jamil, O., and Morcrette, C. J.: Combining distribution-based neural networks to predict weather forecast probabilities, Q. J. Roy. Meteor. Soc., 147, 4337–4357, <ext-link xlink:href="https://doi.org/10.1002/qj.4180" ext-link-type="DOI">10.1002/qj.4180</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx21"><?xmltex \def\ref@label{{Cohen et~al.(2018)}}?><label>Cohen et al.(2018)</label><?label cohen_spherical_2018?><mixed-citation>Cohen, T. S., Geiger, M., Koehler, J., and Welling, M.: Spherical CNNs, arXiv, <ext-link xlink:href="https://doi.org/10.48550/arXiv.1801.10130" ext-link-type="DOI">10.48550/arXiv.1801.10130</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx22"><?xmltex \def\ref@label{{de~Burgh-Day and Leeuwenburg(2023)}}?><label>de Burgh-Day and Leeuwenburg(2023)</label><?label de_burgh-day_machine_2023?><mixed-citation>de Burgh-Day, C. O. and Leeuwenburg, T.: Machine learning for numerical weather and climate modelling: a review, Geosci. Model Dev., 16, 6433–6477, <ext-link xlink:href="https://doi.org/10.5194/gmd-16-6433-2023" ext-link-type="DOI">10.5194/gmd-16-6433-2023</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx23"><?xmltex \def\ref@label{{Dosovitskiy et~al.(2020)}}?><label>Dosovitskiy et al.(2020)</label><?label dosovitskiy_image_2020?><mixed-citation>Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N.: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, in: International Conference on Learning Representations, <uri>https://openreview.net/forum?id=YicbFdNTTy</uri> (last access: 18 March 2024), 2020.</mixed-citation></ref>
      <ref id="bib1.bibx24"><?xmltex \def\ref@label{{Dueben and Bauer(2018)}}?><label>Dueben and Bauer(2018)</label><?label dueben_challenges_2018?><mixed-citation>Dueben, P. D. and Bauer, P.: Challenges and design choices for global weather and climate models based on machine learning, Geosci. Model Dev., 11, 3999–4009, <ext-link xlink:href="https://doi.org/10.5194/gmd-11-3999-2018" ext-link-type="DOI">10.5194/gmd-11-3999-2018</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx25"><?xmltex \def\ref@label{{{ECMWF}(2023)}}?><label>ECMWF(2023)</label><?label ecmwf_machine_2023?><mixed-citation>ECMWF: Machine Learning model data, <uri>https://www.ecmwf.int/en/forecasts/dataset/machine-learning-model-data</uri> (last access: 18 March 2024), 2023.</mixed-citation></ref>
      <ref id="bib1.bibx26"><?xmltex \def\ref@label{{Espeholt et~al.(2021)}}?><label>Espeholt et al.(2021)</label><?label espeholt_skillful_2021?><mixed-citation>Espeholt, L., Agrawal, S., Sønderby, C., Kumar, M., Heek, J., Bromberg, C., Gazen, C., Hickey, J., Bell, A., and Kalchbrenner, N.: Skillful Twelve Hour Precipitation Forecasts using Large Context Neural Networks, arxiv, <ext-link xlink:href="https://doi.org/10.48550/arXiv.2111.07470" ext-link-type="DOI">10.48550/arXiv.2111.07470</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx27"><?xmltex \def\ref@label{{Gr\"{o}nquist et~al.(2021)}}?><label>Grönquist et al.(2021)</label><?label gronquist_deep_2021?><mixed-citation>Grönquist, P., Yao, C., Ben-Nun, T., Dryden, N., Dueben, P., Li, S., and Hoefler, T.: Deep learning for post-processing ensemble weather forecasts, Philos. T. Roy. Soc. A, 379, 20200092, <ext-link xlink:href="https://doi.org/10.1098/rsta.2020.0092" ext-link-type="DOI">10.1098/rsta.2020.0092</ext-link>, 2021.</mixed-citation></ref>
      <?pagebreak page2357?><ref id="bib1.bibx28"><?xmltex \def\ref@label{{Guastavino et~al.(2022)}}?><label>Guastavino et al.(2022)</label><?label guastavino_prediction_2022?><mixed-citation>Guastavino, S., Piana, M., Tizzi, M., Cassola, F., Iengo, A., Sacchetti, D., Solazzo, E., and Benvenuto, F.: Prediction of severe thunderstorm events with ensemble deep learning and radar data, Sci. Rep., 12, 20049, <ext-link xlink:href="https://doi.org/10.1038/s41598-022-23306-6" ext-link-type="DOI">10.1038/s41598-022-23306-6</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx29"><?xmltex \def\ref@label{{Gutzwiller and Serno(2023)}}?><label>Gutzwiller and Serno(2023)</label><?label gutzwiller_using_2023?><mixed-citation>Gutzwiller, K. J. and Serno, K. M.: Using the risk of spatial extrapolation by machine-learning models to assess the reliability of model predictions for conservation, Landscape Ecol., 38, 1363–1372, <ext-link xlink:href="https://doi.org/10.1007/s10980-023-01651-9" ext-link-type="DOI">10.1007/s10980-023-01651-9</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx30"><?xmltex \def\ref@label{{Haidar and Verma(2018)}}?><label>Haidar and Verma(2018)</label><?label haidar_monthly_2018?><mixed-citation>Haidar, A. and Verma, B.: Monthly Rainfall Forecasting Using One-Dimensional Deep Convolutional Neural Network, IEEE Access, 6, 69053–69063, <ext-link xlink:href="https://doi.org/10.1109/ACCESS.2018.2880044" ext-link-type="DOI">10.1109/ACCESS.2018.2880044</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx31"><?xmltex \def\ref@label{{Hall et~al.(1999)}}?><label>Hall et al.(1999)</label><?label hall_precipitation_1999?><mixed-citation>Hall, T., Brooks, H. E., and Doswell, C. A.: Precipitation Forecasting Using a Neural Network, Weather Forecast., 14, 338–345, <ext-link xlink:href="https://doi.org/10.1175/1520-0434(1999)014&lt;0338:PFUANN&gt;2.0.CO;2" ext-link-type="DOI">10.1175/1520-0434(1999)014&lt;0338:PFUANN&gt;2.0.CO;2</ext-link>, 1999.</mixed-citation></ref>
      <ref id="bib1.bibx32"><?xmltex \def\ref@label{{Hersbach et~al.(2020)}}?><label>Hersbach et al.(2020)</label><?label hersbach_era5_2020?><mixed-citation>Hersbach, H., Bell, B., Berrisford, P., Hirahara, S., Horányi, A., Muñoz-Sabater, J., Nicolas, J., Peubey, C., Radu, R., Schepers, D., Simmons, A., Soci, C., Abdalla, S., Abellan, X., Balsamo, G., Bechtold, P., Biavati, G., Bidlot, J., Bonavita, M., De Chiara, G., Dahlgren, P., Dee, D., Diamantakis, M., Dragani, R., Flemming, J., Forbes, R., Fuentes, M., Geer, A., Haimberger, L., Healy, S., Hogan, R. J., Hólm, E., Janisková, M., Keeley, S., Laloyaux, P., Lopez, P., Lupu, C., Radnoti, G., de Rosnay, P., Rozum, I., Vamborg, F., Villaume, S., and Thépaut, J.-N.: The ERA5 global reanalysis, Q. J. Roy. Meteor. Soc., 146, 1999–2049, <ext-link xlink:href="https://doi.org/10.1002/qj.3803" ext-link-type="DOI">10.1002/qj.3803</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx33"><?xmltex \def\ref@label{Hersbach et al.(2023a)}?><label>Hersbach et al.(2023a)</label><?label Hersbach2023data?><mixed-citation>Hersbach, H., Bell, B., Berrisford, P., Biavati, G., Horányi, A., Muñoz Sabater, J., Nicolas, J., Peubey, C., Radu, R., Rozum, I., Schepers, D., Simmons, A., Soci, C., Dee, D., and Thépaut, J.-N.: ERA5 hourly data on single levels from 1940 to present, Copernicus Climate Change Service (C3S) Climate Data Store (CDS) [data set], <ext-link xlink:href="https://doi.org/10.24381/cds.adbb2d47" ext-link-type="DOI">10.24381/cds.adbb2d47</ext-link>, 2023a.</mixed-citation></ref>
      <ref id="bib1.bibx34"><?xmltex \def\ref@label{Hersbach et al.(2023b)}?><label>Hersbach et al.(2023b)</label><?label Hersbach2023bdata?><mixed-citation>Hersbach, H., Bell, B., Berrisford, P., Biavati, G., Horányi, A., Muñoz Sabater, J., Nicolas, J., Peubey, C., Radu, R., Rozum, I., Schepers, D., Simmons, A., Soci, C., Dee, D., and Thépaut, J.-N.: ERA5 hourly data on pressure levels from 1940 to present, Copernicus Climate Change Service (C3S) Climate Data Store (CDS) [data set], <ext-link xlink:href="https://doi.org/10.24381/cds.bd0915c6" ext-link-type="DOI">10.24381/cds.bd0915c6</ext-link>, 2023b.</mixed-citation></ref>
      <ref id="bib1.bibx35"><?xmltex \def\ref@label{{Hochreiter and Schmidhuber(1997)}}?><label>Hochreiter and Schmidhuber(1997)</label><?label hochreiter_long_1997?><mixed-citation>Hochreiter, S. and Schmidhuber, J.: Long Short-term Memory, Neural Comput., 9, 1735–80, <ext-link xlink:href="https://doi.org/10.1162/neco.1997.9.8.1735" ext-link-type="DOI">10.1162/neco.1997.9.8.1735</ext-link>, 1997.</mixed-citation></ref>
      <ref id="bib1.bibx36"><?xmltex \def\ref@label{{Hodson(2022)}}?><label>Hodson(2022)</label><?label hodson_root-mean-square_2022?><mixed-citation>Hodson, T. O.: Root-mean-square error (RMSE) or mean absolute error (MAE): when to use them or not, Geosci. Model Dev., 15, 5481–5487, <ext-link xlink:href="https://doi.org/10.5194/gmd-15-5481-2022" ext-link-type="DOI">10.5194/gmd-15-5481-2022</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx37"><?xmltex \def\ref@label{{Hu et~al.(2023)}}?><label>Hu et al.(2023)</label><?label hu_swinvrnn_2023?><mixed-citation>Hu, Y., Chen, L., Wang, Z., and Li, H.: SwinVRNN: A Data-Driven Ensemble Forecasting Model via Learned Distribution Perturbation, J. Adv. Model. Earth Sy., 15, e2022MS003211, <ext-link xlink:href="https://doi.org/10.1029/2022MS003211" ext-link-type="DOI">10.1029/2022MS003211</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx38"><?xmltex \def\ref@label{{Ivakhnenko and Lapa(1965)}}?><label>Ivakhnenko and Lapa(1965)</label><?label ivakhnenko_cybernetic_1965?><mixed-citation> Ivakhnenko, A. G. and Lapa, V. G.: Cybernetic Predicting Devices, Joint Publications Research Service, available from the Clearinghouse for Federal Scientific and Technical Information, 1965.</mixed-citation></ref>
      <ref id="bib1.bibx39"><?xmltex \def\ref@label{{Kashinath et~al.(2021)}}?><label>Kashinath et al.(2021)</label><?label kashinath_physics-informed_2021?><mixed-citation>Kashinath, K., Mustafa, M., Albert, A., Wu, J.-L., Jiang, C., Esmaeilzadeh, S., Azizzadenesheli, K., Wang, R., Chattopadhyay, A., Singh, A., Manepalli, A., Chirila, D., Yu, R., Walters, R., White, B., Xiao, H., Tchelepi, H. A., Marcus, P., Anandkumar, A., Hassanzadeh, P., and Prabhat, n.: Physics-informed machine learning: case studies for weather and climate modelling, Philos. T. Roy. Soc. A, 379, 20200093, <ext-link xlink:href="https://doi.org/10.1098/rsta.2020.0093" ext-link-type="DOI">10.1098/rsta.2020.0093</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx40"><?xmltex \def\ref@label{{Keisler(2022)}}?><label>Keisler(2022)</label><?label keisler_forecasting_2022?><mixed-citation>Keisler, R.: Forecasting Global Weather with Graph Neural Networks, arXiv, <ext-link xlink:href="https://doi.org/10.48550/arXiv.2202.07575" ext-link-type="DOI">10.48550/arXiv.2202.07575</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx41"><?xmltex \def\ref@label{{Klein et~al.(2015)}}?><label>Klein et al.(2015)</label><?label klein_dynamic_2015?><mixed-citation>Klein, B., Wolf, L., and Afek, Y.: A Dynamic Convolutional Layer for short rangeweather prediction, in: 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 4840–4848, <ext-link xlink:href="https://doi.org/10.1109/CVPR.2015.7299117" ext-link-type="DOI">10.1109/CVPR.2015.7299117</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx42"><?xmltex \def\ref@label{{Koenker and Bassett(1978)}}?><label>Koenker and Bassett(1978)</label><?label koenker_regression_1978?><mixed-citation>Koenker, R. and Bassett, G.: Regression Quantiles, Econometrica, 46, 33–50, <ext-link xlink:href="https://doi.org/10.2307/1913643" ext-link-type="DOI">10.2307/1913643</ext-link>, 1978.</mixed-citation></ref>
      <ref id="bib1.bibx43"><?xmltex \def\ref@label{{Kramer(1991)}}?><label>Kramer(1991)</label><?label kramer_nonlinear_1991?><mixed-citation>Kramer, M. A.: Nonlinear principal component analysis using autoassociative neural networks, AIChE Journal, 37, 233–243, <ext-link xlink:href="https://doi.org/10.1002/aic.690370209" ext-link-type="DOI">10.1002/aic.690370209</ext-link>, 1991.</mixed-citation></ref>
      <ref id="bib1.bibx44"><?xmltex \def\ref@label{{Kron et~al.(2019)}}?><label>Kron et al.(2019)</label><?label kron_changes_2019?><mixed-citation>Kron, W., Löw, P., and Kundzewicz, Z. W.: Changes in risk of extreme weather events in Europe, Environ. Sci. Policy, 100, 74–83, <ext-link xlink:href="https://doi.org/10.1016/j.envsci.2019.06.007" ext-link-type="DOI">10.1016/j.envsci.2019.06.007</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx45"><?xmltex \def\ref@label{{Lam et~al.(2022)}}?><label>Lam et al.(2022)</label><?label lam_graphcast_2022?><mixed-citation>Lam, R., Sanchez-Gonzalez, A., Willson, M., Wirnsberger, P., Fortunato, M., Pritzel, A., Ravuri, S., Ewalds, T., Alet, F., Eaton-Rosen, Z., Hu, W., Merose, A., Hoyer, S., Holland, G., Stott, J., Vinyals, O., Mohamed, S., and Battaglia, P.: GraphCast: Learning skillful medium-range global weather forecasting, arXiv, <ext-link xlink:href="https://doi.org/10.48550/arXiv.2212.12794" ext-link-type="DOI">10.48550/arXiv.2212.12794</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx46"><?xmltex \def\ref@label{{Lam et~al.(2023)}}?><label>Lam et al.(2023)</label><?label lam_learning_2023?><mixed-citation>Lam, R., Sanchez-Gonzalez, A., Willson, M., Wirnsberger, P., Fortunato, M., Alet, F., Ravuri, S., Ewalds, T., Eaton-Rosen, Z., Hu, W., Merose, A., Hoyer, S., Holland, G., Vinyals, O., Stott, J., Pritzel, A., Mohamed, S., and Battaglia, P.: Learning skillful medium-range global weather forecasting, Science, 382, 1416–1421, <ext-link xlink:href="https://doi.org/10.1126/science.adi2336" ext-link-type="DOI">10.1126/science.adi2336</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx47"><?xmltex \def\ref@label{{LeCun and Bengio(1995)}}?><label>LeCun and Bengio(1995)</label><?label lecun_convolutional_1995?><mixed-citation> LeCun, Y. and Bengio, Y.: Convolutional networks for images, speech, and time series, The handbook of brain theory and neural networks, Citeseer, 3361, 1995.</mixed-citation></ref>
      <ref id="bib1.bibx48"><?xmltex \def\ref@label{{LeNail(2019)}}?><label>LeNail(2019)</label><?label lenail_nn-svg_2019?><mixed-citation>LeNail, A.: NN-SVG: Publication-Ready Neural Network Architecture Schematics, J. Open Source Softw., 4, 747, <ext-link xlink:href="https://doi.org/10.21105/joss.00747" ext-link-type="DOI">10.21105/joss.00747</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx49"><?xmltex \def\ref@label{{Li et~al.(2018)}}?><label>Li et al.(2018)</label><?label li_method_2018?><mixed-citation>Li, X., Du, Z., and Song, G.: A Method of Rainfall Runoff Forecasting Based on Deep Convolution Neural Networks, in: 2018 Sixth International Conference on Advanced Cloud and Big Data (CBD), Lanzhou, China, 304–310, <ext-link xlink:href="https://doi.org/10.1109/CBD.2018.00061" ext-link-type="DOI">10.1109/CBD.2018.00061</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx50"><?xmltex \def\ref@label{{Merz et~al.(2020)}}?><label>Merz et al.(2020)</label><?label merz_impact_2020?><mixed-citation>Merz, B., Kuhlicke, C., Kunz, M., Pittore, M., Babeyko, A., Bresch, D. N., Domeisen, D. I. V., Feser, F., Koszalka, I., Kreibich, H., Pantillon, F., Parolai, S., Pinto, J. G., Punge, H. J., Rivalta, E., Schröter, K., Strehlow, K., Weisse, R., and Wurpts, A.: Impact Forecasting to Support Emergency Management of Natural Hazards, Rev. Geophys., 58, RG000704, <ext-link xlink:href="https://doi.org/10.1029/2020RG000704" ext-link-type="DOI">10.1029/2020RG000704</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx51"><?xmltex \def\ref@label{{Molina et~al.(2023)}}?><label>Molina et al.(2023)</label><?label molina_review_2023?><mixed-citation>Molina, M. J., O'Brien, T. A., Anderson, G., Ashfaq, M., Bennett, K. E., Collins, W. D., Dagon, K., Restrepo, J. M., and Ullrich, P. A.: A Review of Recent and Emerging Machine Learning Applications for Climate Variability and Weather Phenomena, Artificial Intelligence for the Earth Systems, 2, e220086, <ext-link xlink:href="https://doi.org/10.1175/AIES-D-22-0086.1" ext-link-type="DOI">10.1175/AIES-D-22-0086.1</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx52"><?xmltex \def\ref@label{{Nguyen et~al.(2023)}}?><label>Nguyen et al.(2023)</label><?label nguyen_climax_2023?><mixed-citation>Nguyen, T., Brandstetter, J., Kapoor, A., Gupta, J. K., and Grover, A.: ClimaX: A foundation model for weather and climate, arXiv, <ext-link xlink:href="https://doi.org/10.48550/arXiv.2301.10343" ext-link-type="DOI">10.48550/arXiv.2301.10343</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx53"><?xmltex \def\ref@label{{Pasche and Engelke(2023)}}?><label>Pasche and Engelke(2023)</label><?label pasche_neural_2023?><mixed-citation>Pasche, O. C. and Engelke, S.: Neural Networks for Extreme Quantile Regression with an Application to Forecasting of Flood Risk, arXiv, <ext-link xlink:href="https://doi.org/10.48550/arXiv.2208.07590" ext-link-type="DOI">10.48550/arXiv.2208.07590</ext-link>, 2023.</mixed-citation></ref>
      <?pagebreak page2358?><ref id="bib1.bibx54"><?xmltex \def\ref@label{{Pathak et~al.(2022)}}?><label>Pathak et al.(2022)</label><?label pathak_fourcastnet_2022?><mixed-citation>Pathak, J., Subramanian, S., Harrington, P., Raja, S., Chattopadhyay, A., Mardani, M., Kurth, T., Hall, D., Li, Z., Azizzadenesheli, K., Hassanzadeh, P., Kashinath, K., and Anandkumar, A.: FourCastNet: A Global Data-driven High-resolution Weather Model using Adaptive Fourier Neural Operators, arXiv, <ext-link xlink:href="https://doi.org/10.48550/arXiv.2202.11214" ext-link-type="DOI">10.48550/arXiv.2202.11214</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx55"><?xmltex \def\ref@label{{Pickands(1975)}}?><label>Pickands(1975)</label><?label pickands_statistical_1975?><mixed-citation> Pickands, J.: Statistical Inference Using Extreme Order Statistics, Ann. Stat., 3, 119–131, 1975.</mixed-citation></ref>
      <ref id="bib1.bibx56"><?xmltex \def\ref@label{{Qiu et~al.(2017)}}?><label>Qiu et al.(2017)</label><?label qiu_short-term_2017?><mixed-citation>Qiu, M., Zhao, P., Zhang, K., Huang, J., Shi, X., Wang, X., and Chu, W.: A Short-Term Rainfall Prediction Model Using Multi-task Convolutional Neural Networks, in: 2017 IEEE International Conference on Data Mining (ICDM), New Orleans, LA, USA,  395–404, <ext-link xlink:href="https://doi.org/10.1109/ICDM.2017.49" ext-link-type="DOI">10.1109/ICDM.2017.49</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx57"><?xmltex \def\ref@label{{Rasp et~al.(2020)}}?><label>Rasp et al.(2020)</label><?label rasp_weatherbench_2020?><mixed-citation>Rasp, S., Dueben, P. D., Scher, S., Weyn, J. A., Mouatadid, S., and Thuerey, N.: WeatherBench: A Benchmark Data Set for Data-Driven Weather Forecasting, J. Adv. Model. Earth Sy., 12, e2020MS002203, <ext-link xlink:href="https://doi.org/10.1029/2020MS002203" ext-link-type="DOI">10.1029/2020MS002203</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx58"><?xmltex \def\ref@label{{Rasp et~al.(2024)}}?><label>Rasp et al.(2024)</label><?label rasp_weatherbench_2024?><mixed-citation>Rasp, S., Hoyer, S., Merose, A., Langmore, I., Battaglia, P., Russel, T., Sanchez-Gonzalez, A., Yang, V., Carver, R., Agrawal, S., Chantry, M., Bouallegue, Z. B., Dueben, P., Bromberg, C., Sisk, J., Barrington, L., Bell, A., and Sha, F.: WeatherBench 2: A benchmark for the next generation of data-driven global weather models, arXiv, <ext-link xlink:href="https://doi.org/10.48550/arXiv.2308.15560" ext-link-type="DOI">10.48550/arXiv.2308.15560</ext-link>, 2024.</mixed-citation></ref>
      <ref id="bib1.bibx59"><?xmltex \def\ref@label{{Ravuri et~al.(2021)}}?><label>Ravuri et al.(2021)</label><?label ravuri_skilful_2021?><mixed-citation>Ravuri, S., Lenc, K., Willson, M., Kangin, D., Lam, R., Mirowski, P., Fitzsimons, M., Athanassiadou, M., Kashem, S., Madge, S., Prudden, R., Mandhane, A., Clark, A., Brock, A., Simonyan, K., Hadsell, R., Robinson, N., Clancy, E., Arribas, A., and Mohamed, S.: Skilful precipitation nowcasting using deep generative models of radar, Nature, 597, 672–677, <ext-link xlink:href="https://doi.org/10.1038/s41586-021-03854-z" ext-link-type="DOI">10.1038/s41586-021-03854-z</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx60"><?xmltex \def\ref@label{{Ren et~al.(2021)}}?><label>Ren et al.(2021)</label><?label ren_deep_2021?><mixed-citation>Ren, X., Li, X., Ren, K., Song, J., Xu, Z., Deng, K., and Wang, X.: Deep Learning-Based Weather Prediction: A Survey, Big Data Research, 23, 100178, <ext-link xlink:href="https://doi.org/10.1016/j.bdr.2020.100178" ext-link-type="DOI">10.1016/j.bdr.2020.100178</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx61"><?xmltex \def\ref@label{{Rumelhart et~al.(1986)}}?><label>Rumelhart et al.(1986)</label><?label rumelhart_learning_1986?><mixed-citation>Rumelhart, D. E., Hinton, G. E., and Williams, R. J.: Learning representations by back-propagating errors, Nature, 323, 533–536, <ext-link xlink:href="https://doi.org/10.1038/323533a0" ext-link-type="DOI">10.1038/323533a0</ext-link>, 1986.</mixed-citation></ref>
      <ref id="bib1.bibx62"><?xmltex \def\ref@label{{Scarselli et~al.(2009)}}?><label>Scarselli et al.(2009)</label><?label scarselli_graph_2009?><mixed-citation>Scarselli, F., Gori, M., Tsoi, A. C., Hagenbuchner, M., and Monfardini, G.: The Graph Neural Network Model, IEEE T. Neural Networ., 20, 61–80, <ext-link xlink:href="https://doi.org/10.1109/TNN.2008.2005605" ext-link-type="DOI">10.1109/TNN.2008.2005605</ext-link>, 2009.</mixed-citation></ref>
      <ref id="bib1.bibx63"><?xmltex \def\ref@label{{Scher and Messori(2019{\natexlab{a}})}}?><label>Scher and Messori(2019a)</label><?label scher_generalization_2019?><mixed-citation>Scher, S. and Messori, G.: Generalization properties of feed-forward neural networks trained on Lorenz systems, Nonlin. Processes Geophys., 26, 381–399, <ext-link xlink:href="https://doi.org/10.5194/npg-26-381-2019" ext-link-type="DOI">10.5194/npg-26-381-2019</ext-link>, 2019a.</mixed-citation></ref>
      <ref id="bib1.bibx64"><?xmltex \def\ref@label{{Scher and Messori(2019{\natexlab{b}})}}?><label>Scher and Messori(2019b)</label><?label scher_weather_2019?><mixed-citation>Scher, S. and Messori, G.: Weather and climate forecasting with neural networks: using general circulation models (GCMs) with different complexity as a study ground, Geosci. Model Dev., 12, 2797–2809, <ext-link xlink:href="https://doi.org/10.5194/gmd-12-2797-2019" ext-link-type="DOI">10.5194/gmd-12-2797-2019</ext-link>, 2019b.</mixed-citation></ref>
      <ref id="bib1.bibx65"><?xmltex \def\ref@label{{Scher and Messori(2021)}}?><label>Scher and Messori(2021)</label><?label scher_ensemble_2021?><mixed-citation>Scher, S. and Messori, G.: Ensemble Methods for Neural Network-Based Weather Forecasts, J. Adv. Model. Earth Sy., 13,  MS002331, <ext-link xlink:href="https://doi.org/10.1029/2020MS002331" ext-link-type="DOI">10.1029/2020MS002331</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx66"><?xmltex \def\ref@label{{Scher and Messori(2023)}}?><label>Scher and Messori(2023)</label><?label scher_spherical_2023?><mixed-citation>Scher, S. and Messori, G.: Spherical convolution and other forms of informed machine learning for deep neural network based weather forecasts, arXiv, <ext-link xlink:href="https://doi.org/10.48550/arXiv.2008.13524" ext-link-type="DOI">10.48550/arXiv.2008.13524</ext-link>, 2023. </mixed-citation></ref><?xmltex \hack{\newpage}?>
      <ref id="bib1.bibx67"><?xmltex \def\ref@label{{Schizas et~al.(1991)}}?><label>Schizas et al.(1991)</label><?label schizas_artificial_1991?><mixed-citation> Schizas, C., Michaelides, S., Pattichis, C., and Livesay, R.: Artificial neural networks in forecasting minimum temperature (weather), in: 1991 Second International Conference on Artificial Neural Networks, Bournemouth, UK, 112–114, 1991.</mixed-citation></ref>
      <ref id="bib1.bibx68"><?xmltex \def\ref@label{{Schultz et~al.(2021)}}?><label>Schultz et al.(2021)</label><?label schultz_can_2021?><mixed-citation>Schultz, M. G., Betancourt, C., Gong, B., Kleinert, F., Langguth, M., Leufen, L. H., Mozaffari, A., and Stadtler, S.: Can deep learning beat numerical weather prediction?, Philos. T. Roy. Soc. A, 379, 20200097, <ext-link xlink:href="https://doi.org/10.1098/rsta.2020.0097" ext-link-type="DOI">10.1098/rsta.2020.0097</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx69"><?xmltex \def\ref@label{{Silini et~al.(2022)}}?><label>Silini et al.(2022)</label><?label silini_improving_2022?><mixed-citation>Silini, R., Lerch, S., Mastrantonas, N., Kantz, H., Barreiro, M., and Masoller, C.: Improving the prediction of the Madden–Julian Oscillation of the ECMWF model by post-processing, Earth Syst. Dynam., 13, 1157–1165, <ext-link xlink:href="https://doi.org/10.5194/esd-13-1157-2022" ext-link-type="DOI">10.5194/esd-13-1157-2022</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx70"><?xmltex \def\ref@label{{Taylor(2000)}}?><label>Taylor(2000)</label><?label taylor_quantile_2000?><mixed-citation>Taylor, J. W.: A quantile regression neural network approach to estimating the conditional density of multiperiod returns, J. Forecast., 19, 299–311, <ext-link xlink:href="https://doi.org/10.1002/1099-131X(200007)19:4&lt;299::AID-FOR775&gt;3.0.CO;2-V" ext-link-type="DOI">10.1002/1099-131X(200007)19:4&lt;299::AID-FOR775&gt;3.0.CO;2-V</ext-link>, 2000.</mixed-citation></ref>
      <ref id="bib1.bibx71"><?xmltex \def\ref@label{{Thuemmel et~al.(2023)}}?><label>Thuemmel et al.(2023)</label><?label thuemmel_inductive_2023?><mixed-citation>Thuemmel, J., Karlbauer, M., Otte, S., Zarfl, C., Martius, G., Ludwig, N., Scholten, T., Friedrich, U., Wulfmeyer, V., Goswami, B., and Butz, M. V.: Inductive biases in deep learning models for weather prediction, arXiv, <ext-link xlink:href="https://doi.org/10.48550/arXiv.2304.04664" ext-link-type="DOI">10.48550/arXiv.2304.04664</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx72"><?xmltex \def\ref@label{{Vaswani et~al.(2017)}}?><label>Vaswani et al.(2017)</label><?label vaswani_attention_2017?><mixed-citation>Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Å., and Polosukhin, I.: Attention is All you Need, in: Advances in Neural Information Processing Systems, edited by Guyon, I., Luxburg, U. V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., and Garnett, R., vol. 30, Curran Associates, Inc., <uri>https://proceedings.neurips.cc/paper_files/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf</uri> (last access: 18 March 2024), 2017.</mixed-citation></ref>
      <ref id="bib1.bibx73"><?xmltex \def\ref@label{{Watson(2022)}}?><label>Watson(2022)</label><?label watson_machine_2022?><mixed-citation>Watson, P. A. G.: Machine learning applications for weather and climate need greater focus on extremes, Environ. Res. Lett., 17, 111004, <ext-link xlink:href="https://doi.org/10.1088/1748-9326/ac9d4e" ext-link-type="DOI">10.1088/1748-9326/ac9d4e</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx74"><?xmltex \def\ref@label{{Weyn et~al.(2019)}}?><label>Weyn et al.(2019)</label><?label weyn_can_2019?><mixed-citation>Weyn, J. A., Durran, D. R., and Caruana, R.: Can Machines Learn to Predict Weather? Using Deep Learning to Predict Gridded 500-hPa Geopotential Height From Historical Weather Data, J. Adv. Model. Earth Sy., 11, 2680–2693, <ext-link xlink:href="https://doi.org/10.1029/2019MS001705" ext-link-type="DOI">10.1029/2019MS001705</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx75"><?xmltex \def\ref@label{{{World Meteorological
Organization}(2022)}}?><label>World Meteorological Organization(2022)</label><?label world_meteorological_organization_early_2022?><mixed-citation>World Meteorological Organization: Early warnings for all: Executive action plan 2023-2027, <uri>https://www.preventionweb.net/publication/early-warnings-all-executive-action-plan-2023-2027</uri> (last access: 18 March 2024), 2022.</mixed-citation></ref>
      <ref id="bib1.bibx76"><?xmltex \def\ref@label{{Zhang et~al.(2023)}}?><label>Zhang et al.(2023)</label><?label zhang_skilful_2023?><mixed-citation>Zhang, Y., Long, M., Chen, K., Xing, L., Jin, R., Jordan, M. I., and Wang, J.: Skilful nowcasting of extreme precipitation with NowcastNet, Nature, 619, 526–532, <ext-link xlink:href="https://doi.org/10.1038/s41586-023-06184-4" ext-link-type="DOI">10.1038/s41586-023-06184-4</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx77"><?xmltex \def\ref@label{{Zhong et~al.(2023)}}?><label>Zhong et al.(2023)</label><?label zhong_fuxi-extreme_2023?><mixed-citation>Zhong, X., Chen, L., Liu, J., Lin, C., Qi, Y., and Li, H.: FuXi-Extreme: Improving extreme rainfall and wind forecasts with diffusion model, arXiv, <ext-link xlink:href="https://doi.org/10.48550/arXiv.2310.19822" ext-link-type="DOI">10.48550/arXiv.2310.19822</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx78"><?xmltex \def\ref@label{{Zhu et~al.(2017)}}?><label>Zhu et al.(2017)</label><?label zhu_wind_2017?><mixed-citation>Zhu, A., Li, X., Mo, Z., and Wu, R.: Wind power prediction based on a convolutional neural network, in: 2017 International Conference on Circuits, Devices and Systems (ICCDS), Chengdu, China,  131–135, <ext-link xlink:href="https://doi.org/10.1109/ICCDS.2017.8120465" ext-link-type="DOI">10.1109/ICCDS.2017.8120465</ext-link>, 2017.</mixed-citation></ref>

  </ref-list></back>
    <!--<article-title-html>Advances and prospects of deep learning for medium-range extreme weather forecasting</article-title-html>
<abstract-html/>
<ref-html id="bib1.bib1"><label>Bahdanau et al.(2015)</label><mixed-citation>
      
Bahdanau, D., Cho, K., and Bengio, Y.: Neural Machine Translation by
Jointly Learning to Align and Translate, in: 3rd International
Conference on Learning Representations, ICLR 2015, San Diego,
CA, USA, 7–9 May 2015, Conference Track Proceedings, edited by:
Bengio, Y. and LeCun, Y., <a href="https://doi.org/10.48550/arXiv.1409.0473" target="_blank">https://doi.org/10.48550/arXiv.1409.0473</a>, 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib2"><label>Balkema and De Haan(1974)</label><mixed-citation>
      
Balkema, A. A. and De Haan, L.: Residual Life Time at Great Age,
Ann. Probab., 2, 792–804, <a href="https://doi.org/10.1214/aop/1176996548" target="_blank">https://doi.org/10.1214/aop/1176996548</a>, 1974.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib3"><label>Barnes et al.(2023)</label><mixed-citation>
      
Barnes, A. P., McCullen, N., and Kjeldsen, T. R.: Forecasting seasonal to
sub-seasonal rainfall in Great Britain using convolutional-neural
networks, Theor. Appl. Climatol., 151, 421–432,
<a href="https://doi.org/10.1007/s00704-022-04242-x" target="_blank">https://doi.org/10.1007/s00704-022-04242-x</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib4"><label>Battaglia et al.(2018)</label><mixed-citation>
      
Battaglia, P. W., Hamrick, J. B., Bapst, V., Sanchez-Gonzalez, A., Zambaldi,
V., Malinowski, M., Tacchetti, A., Raposo, D., Santoro, A., Faulkner, R.,
Gulcehre, C., Song, F., Ballard, A., Gilmer, J., Dahl, G., Vaswani, A.,
Allen, K., Nash, C., Langston, V., Dyer, C., Heess, N., Wierstra, D., Kohli,
P., Botvinick, M., Vinyals, O., Li, Y., and Pascanu, R.: Relational inductive
biases, deep learning, and graph networks, arXiv, <a href="https://doi.org/10.48550/arXiv.1806.01261" target="_blank">https://doi.org/10.48550/arXiv.1806.01261</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib5"><label>Bauer et al.(2023)</label><mixed-citation>
      
Bauer, P., Dueben, P., Chantry, M., Doblas-Reyes, F., Hoefler, T., McGovern,
A., and Stevens, B.: Deep learning and a changing economy in weather and
climate prediction, Nat. Rev. Earth Environ., 4, 507–509,
<a href="https://doi.org/10.1038/s43017-023-00468-z" target="_blank">https://doi.org/10.1038/s43017-023-00468-z</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib6"><label>Ben-Bouallegue et al.(2023)</label><mixed-citation>
      
Ben-Bouallegue, Z., Clare, M. C. A., Magnusson, L., Gascon, E., Maier-Gerber,
M., Janousek, M., Rodwell, M., Pinault, F., Dramsch, J. S., Lang, S. T. K.,
Raoult, B., Rabier, F., Chevallier, M., Sandu, I., Dueben, P., Chantry, M.,
and Pappenberger, F.: The rise of data-driven weather forecasting, arXiv,
<a href="https://doi.org/10.48550/arXiv.2307.10128" target="_blank">https://doi.org/10.48550/arXiv.2307.10128</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib7"><label>Bengio and Gingras(1995)</label><mixed-citation>
      
Bengio, Y. and Gingras, F.: Recurrent Neural Networks for Missing or
Asynchronous Data, in: Advances in Neural Information Processing
Systems, vol. 8, MIT Press,
<a href="https://papers.nips.cc/paper_files/paper/1995/hash/ffeed84c7cb1ae7bf4ec4bd78275bb98-Abstract.html" target="_blank"/> (last access: 18 March 2024),
1995.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib8"><label>Bengio et al.(1994)</label><mixed-citation>
      
Bengio, Y., Simard, P., and Frasconi, P.: Learning long-term dependencies with
gradient descent is difficult, IEEE transactions on neural networks / a
publication of the IEEE Neural Networks Council, 5, 157–66,
<a href="https://doi.org/10.1109/72.279181" target="_blank">https://doi.org/10.1109/72.279181</a>, 1994.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib9"><label>Beucler et al.(2020)</label><mixed-citation>
      
Beucler, T., Pritchard, M., Gentine, P., and Rasp, S.: Towards
Physically-Consistent, Data-Driven Models of Convection, in:
IGARSS 2020 – 2020 IEEE International Geoscience and Remote
Sensing Symposium, 3987–3990,
<a href="https://doi.org/10.1109/IGARSS39084.2020.9324569" target="_blank">https://doi.org/10.1109/IGARSS39084.2020.9324569</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib10"><label>Bi et al.(2022)</label><mixed-citation>
      
Bi, K., Xie, L., Zhang, H., Chen, X., Gu, X., and Tian, Q.: Pangu-Weather:
A 3D High-Resolution Model for Fast and Accurate Global
Weather Forecast, arXiv, <a href="https://doi.org/10.48550/arXiv.2211.02556" target="_blank">https://doi.org/10.48550/arXiv.2211.02556</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib11"><label>Bi et al.(2023)</label><mixed-citation>
      
Bi, K., Xie, L., Zhang, H., Chen, X., Gu, X., and Tian, Q.: Accurate
medium-range global weather forecasting with 3D neural networks, Nature,
619, 1–6, <a href="https://doi.org/10.1038/s41586-023-06185-3" target="_blank">https://doi.org/10.1038/s41586-023-06185-3</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib12"><label>Boomsma and Frellsen(2017)</label><mixed-citation>
      
Boomsma, W. and Frellsen, J.: Spherical convolutions and their application in
molecular modelling, in: Advances in Neural Information Processing
Systems, 30, Curran Associates, Inc.,
<a href="https://papers.nips.cc/paper_files/paper/2017/hash/1113d7a76ffceca1bb350bfe145467c6-Abstract.html" target="_blank"/> (last access: 18 March 2024),
2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib13"><label>Carreau and Bengio(2007)</label><mixed-citation>
      
Carreau, J. and Bengio, Y.: A Hybrid Pareto Model for Conditional
Density Estimation of Asymmetric Fat-Tail Data, in: Proceedings
of the Eleventh International Conference on Artificial Intelligence
and Statistics,  51–58, PMLR,
<a href="https://proceedings.mlr.press/v2/carreau07a.html" target="_blank"/> (last access: 18 March 2024), 2007.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib14"><label>Chen et al.(2023a)</label><mixed-citation>
      
Chen, K., Han, T., Gong, J., Bai, L., Ling, F., Luo, J.-J., Chen, X., Ma, L.,
Zhang, T., Su, R., Ci, Y., Li, B., Yang, X., and Ouyang, W.: FengWu:
Pushing the Skillful Global Medium-range Weather Forecast beyond
10 Days Lead, arXiv, <a href="https://doi.org/10.48550/arXiv.2304.02948" target="_blank">https://doi.org/10.48550/arXiv.2304.02948</a>, 2023a.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib15"><label>Chen et al.(2023b)</label><mixed-citation>
      
Chen, L., Zhong, X., Zhang, F., Cheng, Y., Xu, Y., Qi, Y., and Li, H.: FuXi:
a cascade machine learning forecasting system for 15-day global weather
forecast, npj Climate and Atmospheric Science, 6, 1–11,
<a href="https://doi.org/10.1038/s41612-023-00512-1" target="_blank">https://doi.org/10.1038/s41612-023-00512-1</a>, 2023b.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib16"><label>Chkeir et al.(2023)</label><mixed-citation>
      
Chkeir, S., Anesiadou, A., Mascitelli, A., and Biondi, R.: Nowcasting extreme
rain and extreme wind speed with machine learning techniques applied to
different input datasets, Atmos. Res., 282, 106548,
<a href="https://doi.org/10.1016/j.atmosres.2022.106548" target="_blank">https://doi.org/10.1016/j.atmosres.2022.106548</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib17"><label>Cho et al.(2014)</label><mixed-citation>
      
Cho, K., van Merriënboer, B., Bahdanau, D., and Bengio, Y.: On the
Properties of Neural Machine Translation: Encoder–Decoder
Approaches, in: Proceedings of SSST-8, Eighth Workshop on Syntax,
Semantics and Structure in Statistical Translation, 103–111,
Association for Computational Linguistics, Doha, Qatar,
<a href="https://doi.org/10.3115/v1/W14-4012" target="_blank">https://doi.org/10.3115/v1/W14-4012</a>, 2014.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib18"><label>Cisneros et al.(2023)</label><mixed-citation>
      
Cisneros, D., Richards, J., Dahal, A., Lombardo, L., and Huser, R.: Deep
graphical regression for jointly moderate and extreme Australian wildfires, arXiv,
<a href="https://doi.org/10.48550/arXiv.2308.14547" target="_blank">https://doi.org/10.48550/arXiv.2308.14547</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib19"><label>Civitarese et al.(2021)</label><mixed-citation>
      
Civitarese, D. S., Szwarcman, D., Zadrozny, B., and Watson, C.: Extreme
Precipitation Seasonal Forecast Using a Transformer Neural
Network, arXiv, <a href="https://doi.org/10.48550/arXiv.2107.06846" target="_blank">https://doi.org/10.48550/arXiv.2107.06846</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib20"><label>Clare et al.(2021)</label><mixed-citation>
      
Clare, M. C., Jamil, O., and Morcrette, C. J.: Combining distribution-based
neural networks to predict weather forecast probabilities, Q. J. Roy. Meteor. Soc., 147, 4337–4357, <a href="https://doi.org/10.1002/qj.4180" target="_blank">https://doi.org/10.1002/qj.4180</a>,
2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib21"><label>Cohen et al.(2018)</label><mixed-citation>
      
Cohen, T. S., Geiger, M., Koehler, J., and Welling, M.: Spherical CNNs, arXiv,
<a href="https://doi.org/10.48550/arXiv.1801.10130" target="_blank">https://doi.org/10.48550/arXiv.1801.10130</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib22"><label>de Burgh-Day and Leeuwenburg(2023)</label><mixed-citation>
      
de Burgh-Day, C. O. and Leeuwenburg, T.: Machine learning for numerical weather and climate modelling: a review, Geosci. Model Dev., 16, 6433–6477, <a href="https://doi.org/10.5194/gmd-16-6433-2023" target="_blank">https://doi.org/10.5194/gmd-16-6433-2023</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib23"><label>Dosovitskiy et al.(2020)</label><mixed-citation>
      
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X.,
Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S.,
Uszkoreit, J., and Houlsby, N.: An Image is Worth 16x16 Words:
Transformers for Image Recognition at Scale, in: International
Conference on Learning Representations,
<a href="https://openreview.net/forum?id=YicbFdNTTy" target="_blank"/> (last access: 18 March 2024), 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib24"><label>Dueben and Bauer(2018)</label><mixed-citation>
      
Dueben, P. D. and Bauer, P.: Challenges and design choices for global weather and climate models based on machine learning, Geosci. Model Dev., 11, 3999–4009, <a href="https://doi.org/10.5194/gmd-11-3999-2018" target="_blank">https://doi.org/10.5194/gmd-11-3999-2018</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib25"><label>ECMWF(2023)</label><mixed-citation>
      
ECMWF: Machine Learning model data,
<a href="https://www.ecmwf.int/en/forecasts/dataset/machine-learning-model-data" target="_blank"/> (last access: 18 March 2024), 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib26"><label>Espeholt et al.(2021)</label><mixed-citation>
      
Espeholt, L., Agrawal, S., Sønderby, C., Kumar, M., Heek, J., Bromberg, C.,
Gazen, C., Hickey, J., Bell, A., and Kalchbrenner, N.: Skillful Twelve
Hour Precipitation Forecasts using Large Context Neural
Networks, arxiv, <a href="https://doi.org/10.48550/arXiv.2111.07470" target="_blank">https://doi.org/10.48550/arXiv.2111.07470</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib27"><label>Grönquist et al.(2021)</label><mixed-citation>
      
Grönquist, P., Yao, C., Ben-Nun, T., Dryden, N., Dueben, P., Li, S., and
Hoefler, T.: Deep learning for post-processing ensemble weather forecasts,
Philos. T. Roy. Soc. A, 379, 20200092, <a href="https://doi.org/10.1098/rsta.2020.0092" target="_blank">https://doi.org/10.1098/rsta.2020.0092</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib28"><label>Guastavino et al.(2022)</label><mixed-citation>
      
Guastavino, S., Piana, M., Tizzi, M., Cassola, F., Iengo, A., Sacchetti, D.,
Solazzo, E., and Benvenuto, F.: Prediction of severe thunderstorm events with
ensemble deep learning and radar data, Sci. Rep., 12, 20049,
<a href="https://doi.org/10.1038/s41598-022-23306-6" target="_blank">https://doi.org/10.1038/s41598-022-23306-6</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib29"><label>Gutzwiller and Serno(2023)</label><mixed-citation>
      
Gutzwiller, K. J. and Serno, K. M.: Using the risk of spatial extrapolation by
machine-learning models to assess the reliability of model predictions for
conservation, Landscape Ecol., 38, 1363–1372,
<a href="https://doi.org/10.1007/s10980-023-01651-9" target="_blank">https://doi.org/10.1007/s10980-023-01651-9</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib30"><label>Haidar and Verma(2018)</label><mixed-citation>
      
Haidar, A. and Verma, B.: Monthly Rainfall Forecasting Using
One-Dimensional Deep Convolutional Neural Network, IEEE Access,
6, 69053–69063, <a href="https://doi.org/10.1109/ACCESS.2018.2880044" target="_blank">https://doi.org/10.1109/ACCESS.2018.2880044</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib31"><label>Hall et al.(1999)</label><mixed-citation>
      
Hall, T., Brooks, H. E., and Doswell, C. A.: Precipitation Forecasting
Using a Neural Network, Weather Forecast., 14, 338–345,
<a href="https://doi.org/10.1175/1520-0434(1999)014&lt;0338:PFUANN&gt;2.0.CO;2" target="_blank">https://doi.org/10.1175/1520-0434(1999)014&lt;0338:PFUANN&gt;2.0.CO;2</a>, 1999.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib32"><label>Hersbach et al.(2020)</label><mixed-citation>
      
Hersbach, H., Bell, B., Berrisford, P., Hirahara, S., Horányi, A.,
Muñoz-Sabater, J., Nicolas, J., Peubey, C., Radu, R., Schepers, D., Simmons,
A., Soci, C., Abdalla, S., Abellan, X., Balsamo, G., Bechtold, P., Biavati,
G., Bidlot, J., Bonavita, M., De Chiara, G., Dahlgren, P., Dee, D.,
Diamantakis, M., Dragani, R., Flemming, J., Forbes, R., Fuentes, M., Geer,
A., Haimberger, L., Healy, S., Hogan, R. J., Hólm, E., Janisková, M.,
Keeley, S., Laloyaux, P., Lopez, P., Lupu, C., Radnoti, G., de Rosnay, P.,
Rozum, I., Vamborg, F., Villaume, S., and Thépaut, J.-N.: The ERA5 global
reanalysis, Q. J. Roy. Meteor. Soc., 146,
1999–2049, <a href="https://doi.org/10.1002/qj.3803" target="_blank">https://doi.org/10.1002/qj.3803</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib33"><label>Hersbach et al.(2023a)</label><mixed-citation>
      
Hersbach, H., Bell, B., Berrisford, P., Biavati, G., Horányi, A., Muñoz Sabater, J., Nicolas, J., Peubey, C., Radu, R., Rozum, I., Schepers, D., Simmons, A., Soci, C., Dee, D., and Thépaut, J.-N.: ERA5 hourly data on single levels from 1940 to present, Copernicus Climate Change Service (C3S) Climate Data Store (CDS) [data set], <a href="https://doi.org/10.24381/cds.adbb2d47" target="_blank">https://doi.org/10.24381/cds.adbb2d47</a>, 2023a.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib34"><label>Hersbach et al.(2023b)</label><mixed-citation>
      
Hersbach, H., Bell, B., Berrisford, P., Biavati, G., Horányi, A., Muñoz Sabater, J., Nicolas, J., Peubey, C., Radu, R., Rozum, I., Schepers, D., Simmons, A., Soci, C., Dee, D., and Thépaut, J.-N.: ERA5 hourly data on pressure levels from 1940 to present, Copernicus Climate Change Service (C3S) Climate Data Store (CDS) [data set], <a href="https://doi.org/10.24381/cds.bd0915c6" target="_blank">https://doi.org/10.24381/cds.bd0915c6</a>, 2023b.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib35"><label>Hochreiter and Schmidhuber(1997)</label><mixed-citation>
      
Hochreiter, S. and Schmidhuber, J.: Long Short-term Memory, Neural
Comput., 9, 1735–80, <a href="https://doi.org/10.1162/neco.1997.9.8.1735" target="_blank">https://doi.org/10.1162/neco.1997.9.8.1735</a>, 1997.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib36"><label>Hodson(2022)</label><mixed-citation>
      
Hodson, T. O.: Root-mean-square error (RMSE) or mean absolute error (MAE): when to use them or not, Geosci. Model Dev., 15, 5481–5487, <a href="https://doi.org/10.5194/gmd-15-5481-2022" target="_blank">https://doi.org/10.5194/gmd-15-5481-2022</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib37"><label>Hu et al.(2023)</label><mixed-citation>
      
Hu, Y., Chen, L., Wang, Z., and Li, H.: SwinVRNN: A Data-Driven
Ensemble Forecasting Model via Learned Distribution Perturbation,
J. Adv. Model. Earth Sy., 15, e2022MS003211,
<a href="https://doi.org/10.1029/2022MS003211" target="_blank">https://doi.org/10.1029/2022MS003211</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib38"><label>Ivakhnenko and Lapa(1965)</label><mixed-citation>
      
Ivakhnenko, A. G. and Lapa, V. G.: Cybernetic Predicting Devices, Joint
Publications Research Service, available from the Clearinghouse for Federal
Scientific and Technical Information, 1965.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib39"><label>Kashinath et al.(2021)</label><mixed-citation>
      
Kashinath, K., Mustafa, M., Albert, A., Wu, J.-L., Jiang, C., Esmaeilzadeh, S.,
Azizzadenesheli, K., Wang, R., Chattopadhyay, A., Singh, A., Manepalli, A.,
Chirila, D., Yu, R., Walters, R., White, B., Xiao, H., Tchelepi, H. A.,
Marcus, P., Anandkumar, A., Hassanzadeh, P., and Prabhat, n.:
Physics-informed machine learning: case studies for weather and climate
modelling, Philos. T. Roy. Soc. A, 379, 20200093,
<a href="https://doi.org/10.1098/rsta.2020.0093" target="_blank">https://doi.org/10.1098/rsta.2020.0093</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib40"><label>Keisler(2022)</label><mixed-citation>
      
Keisler, R.: Forecasting Global Weather with Graph Neural Networks, arXiv,
<a href="https://doi.org/10.48550/arXiv.2202.07575" target="_blank">https://doi.org/10.48550/arXiv.2202.07575</a>,
2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib41"><label>Klein et al.(2015)</label><mixed-citation>
      
Klein, B., Wolf, L., and Afek, Y.: A Dynamic Convolutional Layer for
short rangeweather prediction, in: 2015 IEEE Conference on Computer
Vision and Pattern Recognition (CVPR), 4840–4848,
<a href="https://doi.org/10.1109/CVPR.2015.7299117" target="_blank">https://doi.org/10.1109/CVPR.2015.7299117</a>, 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib42"><label>Koenker and Bassett(1978)</label><mixed-citation>
      
Koenker, R. and Bassett, G.: Regression Quantiles, Econometrica, 46, 33–50,
<a href="https://doi.org/10.2307/1913643" target="_blank">https://doi.org/10.2307/1913643</a>, 1978.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib43"><label>Kramer(1991)</label><mixed-citation>
      
Kramer, M. A.: Nonlinear principal component analysis using autoassociative
neural networks, AIChE Journal, 37, 233–243, <a href="https://doi.org/10.1002/aic.690370209" target="_blank">https://doi.org/10.1002/aic.690370209</a>,
1991.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib44"><label>Kron et al.(2019)</label><mixed-citation>
      
Kron, W., Löw, P., and Kundzewicz, Z. W.: Changes in risk of extreme weather
events in Europe, Environ. Sci. Policy, 100, 74–83,
<a href="https://doi.org/10.1016/j.envsci.2019.06.007" target="_blank">https://doi.org/10.1016/j.envsci.2019.06.007</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib45"><label>Lam et al.(2022)</label><mixed-citation>
      
Lam, R., Sanchez-Gonzalez, A., Willson, M., Wirnsberger, P., Fortunato, M.,
Pritzel, A., Ravuri, S., Ewalds, T., Alet, F., Eaton-Rosen, Z., Hu, W.,
Merose, A., Hoyer, S., Holland, G., Stott, J., Vinyals, O., Mohamed, S., and
Battaglia, P.: GraphCast: Learning skillful medium-range global weather
forecasting, arXiv, <a href="https://doi.org/10.48550/arXiv.2212.12794" target="_blank">https://doi.org/10.48550/arXiv.2212.12794</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib46"><label>Lam et al.(2023)</label><mixed-citation>
      
Lam, R., Sanchez-Gonzalez, A., Willson, M., Wirnsberger, P., Fortunato, M.,
Alet, F., Ravuri, S., Ewalds, T., Eaton-Rosen, Z., Hu, W., Merose, A., Hoyer,
S., Holland, G., Vinyals, O., Stott, J., Pritzel, A., Mohamed, S., and
Battaglia, P.: Learning skillful medium-range global weather forecasting,
Science, 382, 1416–1421, <a href="https://doi.org/10.1126/science.adi2336" target="_blank">https://doi.org/10.1126/science.adi2336</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib47"><label>LeCun and Bengio(1995)</label><mixed-citation>
      
LeCun, Y. and Bengio, Y.: Convolutional networks for images, speech, and time
series, The handbook of brain theory and neural networks, Citeseer, 3361, 1995.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib48"><label>LeNail(2019)</label><mixed-citation>
      
LeNail, A.: NN-SVG: Publication-Ready Neural Network Architecture
Schematics, J. Open Source Softw., 4, 747,
<a href="https://doi.org/10.21105/joss.00747" target="_blank">https://doi.org/10.21105/joss.00747</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib49"><label>Li et al.(2018)</label><mixed-citation>
      
Li, X., Du, Z., and Song, G.: A Method of Rainfall Runoff Forecasting
Based on Deep Convolution Neural Networks, in: 2018 Sixth
International Conference on Advanced Cloud and Big Data (CBD), Lanzhou, China,
304–310, <a href="https://doi.org/10.1109/CBD.2018.00061" target="_blank">https://doi.org/10.1109/CBD.2018.00061</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib50"><label>Merz et al.(2020)</label><mixed-citation>
      
Merz, B., Kuhlicke, C., Kunz, M., Pittore, M., Babeyko, A., Bresch, D. N.,
Domeisen, D. I. V., Feser, F., Koszalka, I., Kreibich, H., Pantillon, F.,
Parolai, S., Pinto, J. G., Punge, H. J., Rivalta, E., Schröter, K.,
Strehlow, K., Weisse, R., and Wurpts, A.: Impact Forecasting to Support
Emergency Management of Natural Hazards, Rev. Geophys., 58, RG000704,
<a href="https://doi.org/10.1029/2020RG000704" target="_blank">https://doi.org/10.1029/2020RG000704</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib51"><label>Molina et al.(2023)</label><mixed-citation>
      
Molina, M. J., O'Brien, T. A., Anderson, G., Ashfaq, M., Bennett, K. E.,
Collins, W. D., Dagon, K., Restrepo, J. M., and Ullrich, P. A.: A Review of
Recent and Emerging Machine Learning Applications for Climate
Variability and Weather Phenomena, Artificial Intelligence for the
Earth Systems, 2, e220086, <a href="https://doi.org/10.1175/AIES-D-22-0086.1" target="_blank">https://doi.org/10.1175/AIES-D-22-0086.1</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib52"><label>Nguyen et al.(2023)</label><mixed-citation>
      
Nguyen, T., Brandstetter, J., Kapoor, A., Gupta, J. K., and Grover, A.:
ClimaX: A foundation model for weather and climate, arXiv,
<a href="https://doi.org/10.48550/arXiv.2301.10343" target="_blank">https://doi.org/10.48550/arXiv.2301.10343</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib53"><label>Pasche and Engelke(2023)</label><mixed-citation>
      
Pasche, O. C. and Engelke, S.: Neural Networks for Extreme Quantile
Regression with an Application to Forecasting of Flood Risk, arXiv,
<a href="https://doi.org/10.48550/arXiv.2208.07590" target="_blank">https://doi.org/10.48550/arXiv.2208.07590</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib54"><label>Pathak et al.(2022)</label><mixed-citation>
      
Pathak, J., Subramanian, S., Harrington, P., Raja, S., Chattopadhyay, A.,
Mardani, M., Kurth, T., Hall, D., Li, Z., Azizzadenesheli, K., Hassanzadeh,
P., Kashinath, K., and Anandkumar, A.: FourCastNet: A Global
Data-driven High-resolution Weather Model using Adaptive Fourier
Neural Operators, arXiv, <a href="https://doi.org/10.48550/arXiv.2202.11214" target="_blank">https://doi.org/10.48550/arXiv.2202.11214</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib55"><label>Pickands(1975)</label><mixed-citation>
      
Pickands, J.: Statistical Inference Using Extreme Order Statistics,
Ann. Stat., 3, 119–131, 1975.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib56"><label>Qiu et al.(2017)</label><mixed-citation>
      
Qiu, M., Zhao, P., Zhang, K., Huang, J., Shi, X., Wang, X., and Chu, W.: A
Short-Term Rainfall Prediction Model Using Multi-task
Convolutional Neural Networks, in: 2017 IEEE International
Conference on Data Mining (ICDM), New Orleans, LA, USA,  395–404,
<a href="https://doi.org/10.1109/ICDM.2017.49" target="_blank">https://doi.org/10.1109/ICDM.2017.49</a>, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib57"><label>Rasp et al.(2020)</label><mixed-citation>
      
Rasp, S., Dueben, P. D., Scher, S., Weyn, J. A., Mouatadid, S., and Thuerey,
N.: WeatherBench: A Benchmark Data Set for Data-Driven
Weather Forecasting, J. Adv. Model. Earth Sy., 12,
e2020MS002203, <a href="https://doi.org/10.1029/2020MS002203" target="_blank">https://doi.org/10.1029/2020MS002203</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib58"><label>Rasp et al.(2024)</label><mixed-citation>
      
Rasp, S., Hoyer, S., Merose, A., Langmore, I., Battaglia, P., Russel, T.,
Sanchez-Gonzalez, A., Yang, V., Carver, R., Agrawal, S., Chantry, M.,
Bouallegue, Z. B., Dueben, P., Bromberg, C., Sisk, J., Barrington, L., Bell,
A., and Sha, F.: WeatherBench 2: A benchmark for the next generation of
data-driven global weather models, arXiv, <a href="https://doi.org/10.48550/arXiv.2308.15560" target="_blank">https://doi.org/10.48550/arXiv.2308.15560</a>, 2024.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib59"><label>Ravuri et al.(2021)</label><mixed-citation>
      
Ravuri, S., Lenc, K., Willson, M., Kangin, D., Lam, R., Mirowski, P.,
Fitzsimons, M., Athanassiadou, M., Kashem, S., Madge, S., Prudden, R.,
Mandhane, A., Clark, A., Brock, A., Simonyan, K., Hadsell, R., Robinson, N.,
Clancy, E., Arribas, A., and Mohamed, S.: Skilful precipitation nowcasting
using deep generative models of radar, Nature, 597, 672–677,
<a href="https://doi.org/10.1038/s41586-021-03854-z" target="_blank">https://doi.org/10.1038/s41586-021-03854-z</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib60"><label>Ren et al.(2021)</label><mixed-citation>
      
Ren, X., Li, X., Ren, K., Song, J., Xu, Z., Deng, K., and Wang, X.: Deep
Learning-Based Weather Prediction: A Survey, Big Data Research,
23, 100178, <a href="https://doi.org/10.1016/j.bdr.2020.100178" target="_blank">https://doi.org/10.1016/j.bdr.2020.100178</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib61"><label>Rumelhart et al.(1986)</label><mixed-citation>
      
Rumelhart, D. E., Hinton, G. E., and Williams, R. J.: Learning representations
by back-propagating errors, Nature, 323, 533–536, <a href="https://doi.org/10.1038/323533a0" target="_blank">https://doi.org/10.1038/323533a0</a>, 1986.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib62"><label>Scarselli et al.(2009)</label><mixed-citation>
      
Scarselli, F., Gori, M., Tsoi, A. C., Hagenbuchner, M., and Monfardini, G.: The
Graph Neural Network Model, IEEE T. Neural Networ., 20,
61–80, <a href="https://doi.org/10.1109/TNN.2008.2005605" target="_blank">https://doi.org/10.1109/TNN.2008.2005605</a>, 2009.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib63"><label>Scher and Messori(2019a)</label><mixed-citation>
      
Scher, S. and Messori, G.: Generalization properties of feed-forward neural networks trained on Lorenz systems, Nonlin. Processes Geophys., 26, 381–399, <a href="https://doi.org/10.5194/npg-26-381-2019" target="_blank">https://doi.org/10.5194/npg-26-381-2019</a>, 2019a.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib64"><label>Scher and Messori(2019b)</label><mixed-citation>
      
Scher, S. and Messori, G.: Weather and climate forecasting with neural networks: using general circulation models (GCMs) with different complexity as a study ground, Geosci. Model Dev., 12, 2797–2809, <a href="https://doi.org/10.5194/gmd-12-2797-2019" target="_blank">https://doi.org/10.5194/gmd-12-2797-2019</a>, 2019b.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib65"><label>Scher and Messori(2021)</label><mixed-citation>
      
Scher, S. and Messori, G.: Ensemble Methods for Neural Network-Based
Weather Forecasts, J. Adv. Model. Earth Sy., 13,  MS002331,
<a href="https://doi.org/10.1029/2020MS002331" target="_blank">https://doi.org/10.1029/2020MS002331</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib66"><label>Scher and Messori(2023)</label><mixed-citation>
      
Scher, S. and Messori, G.: Spherical convolution and other forms of informed
machine learning for deep neural network based weather forecasts, arXiv,
<a href="https://doi.org/10.48550/arXiv.2008.13524" target="_blank">https://doi.org/10.48550/arXiv.2008.13524</a>, 2023.


    </mixed-citation></ref-html>
<ref-html id="bib1.bib67"><label>Schizas et al.(1991)</label><mixed-citation>
      
Schizas, C., Michaelides, S., Pattichis, C., and Livesay, R.: Artificial neural
networks in forecasting minimum temperature (weather), in: 1991 Second
International Conference on Artificial Neural Networks, Bournemouth, UK,
112–114, 1991.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib68"><label>Schultz et al.(2021)</label><mixed-citation>
      
Schultz, M. G., Betancourt, C., Gong, B., Kleinert, F., Langguth, M., Leufen,
L. H., Mozaffari, A., and Stadtler, S.: Can deep learning beat numerical
weather prediction?, Philos. T. Roy. Soc. A, 379, 20200097,
<a href="https://doi.org/10.1098/rsta.2020.0097" target="_blank">https://doi.org/10.1098/rsta.2020.0097</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib69"><label>Silini et al.(2022)</label><mixed-citation>
      
Silini, R., Lerch, S., Mastrantonas, N., Kantz, H., Barreiro, M., and Masoller, C.: Improving the prediction of the Madden–Julian Oscillation of the ECMWF model by post-processing, Earth Syst. Dynam., 13, 1157–1165, <a href="https://doi.org/10.5194/esd-13-1157-2022" target="_blank">https://doi.org/10.5194/esd-13-1157-2022</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib70"><label>Taylor(2000)</label><mixed-citation>
      
Taylor, J. W.: A quantile regression neural network approach to estimating the
conditional density of multiperiod returns, J. Forecast., 19,
299–311, <a href="https://doi.org/10.1002/1099-131X(200007)19:4&lt;299::AID-FOR775&gt;3.0.CO;2-V" target="_blank">https://doi.org/10.1002/1099-131X(200007)19:4&lt;299::AID-FOR775&gt;3.0.CO;2-V</a>,
2000.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib71"><label>Thuemmel et al.(2023)</label><mixed-citation>
      
Thuemmel, J., Karlbauer, M., Otte, S., Zarfl, C., Martius, G., Ludwig, N.,
Scholten, T., Friedrich, U., Wulfmeyer, V., Goswami, B., and Butz, M. V.:
Inductive biases in deep learning models for weather prediction, arXiv,
<a href="https://doi.org/10.48550/arXiv.2304.04664" target="_blank">https://doi.org/10.48550/arXiv.2304.04664</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib72"><label>Vaswani et al.(2017)</label><mixed-citation>
      
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N.,
Kaiser, Å., and Polosukhin, I.: Attention is All you Need, in: Advances
in Neural Information Processing Systems, edited by Guyon, I.,
Luxburg, U. V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., and
Garnett, R., vol. 30, Curran Associates, Inc.,
<a href="https://proceedings.neurips.cc/paper_files/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf" target="_blank"/> (last access: 18 March 2024),
2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib73"><label>Watson(2022)</label><mixed-citation>
      
Watson, P. A. G.: Machine learning applications for weather and climate need
greater focus on extremes, Environ. Res. Lett., 17, 111004,
<a href="https://doi.org/10.1088/1748-9326/ac9d4e" target="_blank">https://doi.org/10.1088/1748-9326/ac9d4e</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib74"><label>Weyn et al.(2019)</label><mixed-citation>
      
Weyn, J. A., Durran, D. R., and Caruana, R.: Can Machines Learn to
Predict Weather? Using Deep Learning to Predict Gridded
500-hPa Geopotential Height From Historical Weather Data,
J. Adv. Model. Earth Sy., 11, 2680–2693,
<a href="https://doi.org/10.1029/2019MS001705" target="_blank">https://doi.org/10.1029/2019MS001705</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib75"><label>World Meteorological
Organization(2022)</label><mixed-citation>
      
World Meteorological Organization: Early warnings for all: Executive action
plan 2023-2027,
<a href="https://www.preventionweb.net/publication/early-warnings-all-executive-action-plan-2023-2027" target="_blank"/> (last access: 18 March 2024),
2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib76"><label>Zhang et al.(2023)</label><mixed-citation>
      
Zhang, Y., Long, M., Chen, K., Xing, L., Jin, R., Jordan, M. I., and Wang, J.:
Skilful nowcasting of extreme precipitation with NowcastNet, Nature, 619,
526–532, <a href="https://doi.org/10.1038/s41586-023-06184-4" target="_blank">https://doi.org/10.1038/s41586-023-06184-4</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib77"><label>Zhong et al.(2023)</label><mixed-citation>
      
Zhong, X., Chen, L., Liu, J., Lin, C., Qi, Y., and Li, H.: FuXi-Extreme:
Improving extreme rainfall and wind forecasts with diffusion model, arXiv,
<a href="https://doi.org/10.48550/arXiv.2310.19822" target="_blank">https://doi.org/10.48550/arXiv.2310.19822</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib78"><label>Zhu et al.(2017)</label><mixed-citation>
      
Zhu, A., Li, X., Mo, Z., and Wu, R.: Wind power prediction based on a
convolutional neural network, in: 2017 International Conference on
Circuits, Devices and Systems (ICCDS), Chengdu, China,  131–135,
<a href="https://doi.org/10.1109/ICCDS.2017.8120465" target="_blank">https://doi.org/10.1109/ICCDS.2017.8120465</a>, 2017.

    </mixed-citation></ref-html>--></article>
