<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing with OASIS Tables v3.0 20080202//EN" "https://jats.nlm.nih.gov/nlm-dtd/publishing/3.0/journalpub-oasis3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:oasis="http://docs.oasis-open.org/ns/oasis-exchange/table" xml:lang="en" dtd-version="3.0" article-type="research-article"><?xmltex \bartext{Model description paper}?>
  <front>
    <journal-meta><journal-id journal-id-type="publisher">GMD</journal-id><journal-title-group>
    <journal-title>Geoscientific Model Development</journal-title>
    <abbrev-journal-title abbrev-type="publisher">GMD</abbrev-journal-title><abbrev-journal-title abbrev-type="nlm-ta">Geosci. Model Dev.</abbrev-journal-title>
  </journal-title-group><issn pub-type="epub">1991-9603</issn><publisher>
    <publisher-name>Copernicus Publications</publisher-name>
    <publisher-loc>Göttingen, Germany</publisher-loc>
  </publisher></journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.5194/gmd-17-3839-2024</article-id><title-group><article-title>DEUCE v1.0: a neural network for probabilistic precipitation nowcasting with aleatoric and epistemic uncertainties</article-title><alt-title>DEUCE v1.0</alt-title>
      </title-group><?xmltex \runningtitle{DEUCE v1.0}?><?xmltex \runningauthor{B. Harnist et al.}?>
      <contrib-group>
        <contrib contrib-type="author" corresp="yes">
          <name><surname>Harnist</surname><given-names>Bent</given-names></name>
          <email>bent.harnist@fmi.fi</email>
        <ext-link>https://orcid.org/0000-0003-0079-7315</ext-link></contrib>
        <contrib contrib-type="author" corresp="no">
          <name><surname>Pulkkinen</surname><given-names>Seppo</given-names></name>
          
        <ext-link>https://orcid.org/0000-0002-1318-2814</ext-link></contrib>
        <contrib contrib-type="author" corresp="no">
          <name><surname>Mäkinen</surname><given-names>Terhi</given-names></name>
          
        </contrib>
        <aff id="aff1"><institution>Finnish Meteorological Institute, Erik Palménin aukio 1, 00560 Helsinki, Finland</institution>
        </aff>
      </contrib-group>
      <author-notes><corresp id="corr1">Bent Harnist (bent.harnist@fmi.fi)</corresp></author-notes><pub-date><day>14</day><month>May</month><year>2024</year></pub-date>
      
      <volume>17</volume>
      <issue>9</issue>
      <fpage>3839</fpage><lpage>3866</lpage>
      <history>
        <date date-type="received"><day>24</day><month>May</month><year>2023</year></date>
           <date date-type="rev-request"><day>5</day><month>June</month><year>2023</year></date>
           <date date-type="rev-recd"><day>8</day><month>March</month><year>2024</year></date>
           <date date-type="accepted"><day>24</day><month>March</month><year>2024</year></date>
      </history>
      <permissions>
        <copyright-statement>Copyright: © 2024 Bent Harnist et al.</copyright-statement>
        <copyright-year>2024</copyright-year>
      <license license-type="open-access"><license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p></license></permissions><self-uri xlink:href="https://gmd.copernicus.org/articles/17/3839/2024/gmd-17-3839-2024.html">This article is available from https://gmd.copernicus.org/articles/17/3839/2024/gmd-17-3839-2024.html</self-uri><self-uri xlink:href="https://gmd.copernicus.org/articles/17/3839/2024/gmd-17-3839-2024.pdf">The full text article is available as a PDF file from https://gmd.copernicus.org/articles/17/3839/2024/gmd-17-3839-2024.pdf</self-uri>
      <abstract><title>Abstract</title>

      <p id="d1e96">Precipitation nowcasting (forecasting locally for 0–6 h) serves both public security and industries, facilitating the mitigation of losses incurred due to, e.g., flash floods and is usually done by predicting weather radar echoes, which provide better performance than numerical weather prediction (NWP) at that scale. Probabilistic nowcasts are especially useful as they provide a desirable framework for operational decision-making. Many extrapolation-based statistical nowcasting methods exist, but they all suffer from a limited ability to capture the nonlinear growth and decay of precipitation, leading to a recent paradigm shift towards deep-learning methods which are more capable of representing these patterns.</p>

      <p id="d1e99">Despite its potential advantages, the application of deep learning in probabilistic nowcasting has only recently started to be explored. Here we develop a novel probabilistic precipitation nowcasting method, based on Bayesian neural networks with variational inference and the U-Net architecture, named DEUCE. The method estimates the total predictive uncertainty in the precipitation by combining estimates of the epistemic (knowledge-related and reducible) and heteroscedastic aleatoric (data-dependent and irreducible) uncertainties, using them to produce an ensemble of development scenarios for the following 60 min.</p>

      <p id="d1e102">DEUCE is trained and verified using Finnish Meteorological Institute radar composites compared to established classical models. Our model is found to produce both skillful and reliable probabilistic nowcasts based on various evaluation criteria. It improves the receiver operating characteristic (ROC) area under the curve scores 1 %–5 % over STEPS and LINDA-P baselines and comes close to the best-performer STEPS on a continuous ranked probability score (CRPS) metric. The reliability of DEUCE is demonstrated with, e.g., having the lowest expected calibration error at 20 and 25 dBZ reflectivity thresholds and coming second at 35 dBZ. On the other hand, the deterministic performance of ensemble means is found to be worse than that of extrapolation and LINDA-D baselines. Last, the composition of the predictive uncertainty is analyzed and described, with the conclusion that aleatoric uncertainty is more significant and informative than epistemic uncertainty in the DEUCE model.</p>
  </abstract>
    
<funding-group>
<award-group id="gs1">
<funding-source>Research Council of Finland</funding-source>
<award-id>341964</award-id>
</award-group>
</funding-group>
</article-meta>
  </front>
<body>
      

<sec id="Ch1.S1" sec-type="intro">
  <label>1</label><title>Introduction</title>
      <p id="d1e114">Predicting the amount and location of precipitation at local scales of a few kilometers for lead times ranging from minutes to hours, i.e., precipitation nowcasting, has recently grown into an important component of severe weather early-warning systems, particularly for those focused on predicting flash floods. Because of the intensification coupled with the increased frequency of extreme precipitation events brought by climate change, accurate estimates of future precipitation have increased in importance. However, the capacity of any nowcasting model to produce accurate estimates is limited and thus also having an idea of the reliability of the nowcast is operationally important. This can be addressed with ensemble nowcasts, which generate a set of possible scenarios, with which it is possible to estimate the probability of certain events.</p>
      <?pagebreak page3840?><p id="d1e117">Numerical weather prediction (NWP) is widely used for forecasts at longer timescales and with coarser grids <xref ref-type="bibr" rid="bib1.bibx5" id="paren.1"/>, with regional high-resolution models model generally having a grid resolution of a few kilometers and a refresh rate of typically 1 h. For example, the High-Resolution Rapid Refresh (HRRR) model developed by the United States' National Oceanic and Atmospheric Administration (NOAA) has a grid resolution of 3 km and a refresh rate of 1 h <xref ref-type="bibr" rid="bib1.bibx3" id="paren.2"/>. However, NWP does not achieve sufficient performance at the spatiotemporal scales typical of nowcasts, due to not yet having achieved numerical stability in these first few hours and due to the computational complexity of resolving atmospheric equations at sub-hour temporal resolutions and grid resolutions approaching the microscale (<inline-formula><mml:math id="M1" display="inline"><mml:mo lspace="0mm">≤</mml:mo></mml:math></inline-formula> 1 km) <xref ref-type="bibr" rid="bib1.bibx46 bib1.bibx36" id="paren.3"/>. Specialized nowcasting methods for precipitation have been developed in parallel with NWP and may be used in order to circumvent its problems in the domain. These mainly rely on forecasting the evolution of radar echo image sequences that act as a good proxy for ground-level precipitation and usually have a spatial resolution of <inline-formula><mml:math id="M2" display="inline"><mml:mo>∼</mml:mo></mml:math></inline-formula> 1 km and a temporal resolution of  <inline-formula><mml:math id="M3" display="inline"><mml:mo>∼</mml:mo></mml:math></inline-formula> 5 min, which are characteristic of weather radar observations.</p>
<sec id="Ch1.S1.SS1">
  <label>1.1</label><title>Extrapolation-based precipitation nowcasting</title>
      <p id="d1e158">The most important class of precipitation nowcasting models is based on the extrapolation of radar echoes along the background advection field. These models first estimate the advection field from a sequence of past radar images with methods such as variational echo tracking <xref ref-type="bibr" rid="bib1.bibx26" id="paren.4"/> or optical-flow-based methods like the Lucas–Kanade method <xref ref-type="bibr" rid="bib1.bibx29 bib1.bibx7" id="paren.5"/>. In the classical case of the pure extrapolation nowcast, the most recently observed frame is simply extrapolated along the estimated advection field, often using a semi-Lagrangian scheme <xref ref-type="bibr" rid="bib1.bibx45" id="paren.6"/>. Extrapolation nowcasting does not model the growth and decay of precipitation, so many extensions attempting to make up for that have been developed. One important method is Spectral Prognosis (S-PROG) by <xref ref-type="bibr" rid="bib1.bibx41" id="text.7"/>. S-PROG is based on the scale-dependence of the lifetime and evolution of features, decomposing the field into additive components corresponding to different spatial scales and evolving each of them separately using an autoregressive (AR2) model in Lagrangian (flow frame of reference) coordinates, enabling modeling the scale-dependent behavior of precipitation.</p>
      <p id="d1e173">STEPS (Short-Term Ensemble Prediction System) by <xref ref-type="bibr" rid="bib1.bibx8" id="text.8"/>, is an influential ensemble nowcasting model based on S-PROG. In STEPS, stochastic perturbations are added to the motion field in order to model its uncertainty. Just like with S-PROG, the growth and decay of the precipitation field is modeled by decomposing it using a cascade of scales with the autoregressive model applied to each of these scales separately in Lagrangian coordinates. Unlike in S-PROG, stochastic noise is injected at each scale, concurrent with the AR modeling. Over time, various models have expanded upon STEPS; one recent example is LINDA (Lagrangian INtegro-Difference equation model with autoregression) <xref ref-type="bibr" rid="bib1.bibx35" id="paren.9"/>, which uses an integro-difference equation model with rain cell detection and convolutions for modeling the loss of predictability at small scales. LINDA produces nowcasts that are particularly well-suited to the prediction of strong localized rainfall.</p>
</sec>
<sec id="Ch1.S1.SS2">
  <label>1.2</label><title>Deep-learning approaches to precipitation nowcasting</title>
      <p id="d1e190">With significant recent advances in deep learning, the interest in its use for precipitation nowcasting has increased. One of the first deep-learning models to have been used explicitly for precipitation nowcasting is the convolutional LSTM (ConvLSTM) model <xref ref-type="bibr" rid="bib1.bibx43" id="paren.10"/>, which combines the temporal prediction capacity of the long short-term memory (LSTM) neural networks with 3D convolutions modeling spatiotemporal features in one model for spatiotemporal nowcasting. ConvLSTM has later been improved by the TrajGRU model <xref ref-type="bibr" rid="bib1.bibx44" id="paren.11"/> that replaces the heavy LSTM structure with a lighter GRU (gated recurrent unit) structure and is capable of learning an active location variant structure for the recurrent connections.</p>
      <p id="d1e199">Apart from doing the temporal modeling using recurrent units, a popular approach has been to use fully convolutional neural networks, often two-dimensional, thus avoiding the modeling of explicit temporal dependencies. These networks have often been based on U-Net-type architectures, one early example of which is the model by <xref ref-type="bibr" rid="bib1.bibx2" id="text.12"/>, which predicts the exceedance of rainfall over three distinct intensity thresholds for a 1 h lead time. A more useful model is RainNet by <xref ref-type="bibr" rid="bib1.bibx4" id="text.13"/>. RainNet nowcasts rainfall continuously, one time step at a time, inserting the predicted frames back into the network in order to make multiple lead time predictions. Similarly, FureNET by <xref ref-type="bibr" rid="bib1.bibx32" id="text.14"/> nowcasts rainfall 1 h at a time using polarimetric input variables, in addition to observed rain rates, via multiple encoder branches and late fusion in the decoder of a residual U-Net architecture and brings improvement compared to using plain rain rates.</p>
      <p id="d1e211">The principal problem of using the above deep-learning models for deterministic precipitation nowcasting is that of the increasing blurring of nowcasts with lead time. This is the natural consequence of attempting to minimize the pixel-wise forecasting error in the presence of uncertainties inherent to the task of predicting precipitation. Such loss functions thus behave in the same fashion as S-PROG and STEPS by explicitly filtering out scales through their loss of predictability. One way to resolve the problem is to use generative modeling, which is the approach taken by <xref ref-type="bibr" rid="bib1.bibx37" id="text.15"/> with their Deep Generative Model of Radar (DGMR). DGMR is an adversarially trained convolutional gated recurrent unit (ConvGRU)-based generative model capable of generating realistic time series of future radar observations that outperform both classical and deep-learning baseline models. In addition to deterministic nowcasts, DGMR is also capable of making ensemble-based probabilistic nowcasts.</p>
      <?pagebreak page3841?><p id="d1e217">Making probabilistic precipitation nowcasts using deep learning has been explored less often than deterministic nowcasts, despite the clear benefit of the probabilistic approach in operational use. In addition to DGMR, other existing probabilistic models are MetNet <xref ref-type="bibr" rid="bib1.bibx47" id="paren.16"/> and its successor MetNet-2 <xref ref-type="bibr" rid="bib1.bibx11" id="paren.17"/>. MetNet aggregates weather radar, satellite, and orographic information over a large area to predict a probability distribution of rain rate per pixel in one forward pass for a single lead time, with an architecture consisting of a spatial aggregator of inputs, a ConvLSTM spatial encoder, and a spatial decoder with axial attention. The model is shown to outperform the HRRR NWP model on an F1 metric for lead times up to 8 h. MetNet-2 improves upon its predecessor by adding the data assimilation context as an input and aggregating data over a larger area. This enables it to outperform, or at worst rival, HRRR and HREF (High-Resolution Ensemble Forecast) models in continuous ranked probability score (CRPS) and critical success index (CSI) metrics for lead times up to 12 h.</p>
</sec>
<sec id="Ch1.S1.SS3">
  <label>1.3</label><title>Uncertainty quantification and Bayesian deep learning</title>
      <p id="d1e235">In addition to playing an important role in precipitation nowcasting, the importance of uncertainty quantification (UQ) has also been recognized in deep learning <xref ref-type="bibr" rid="bib1.bibx1" id="paren.18"/>. In the field of machine learning, uncertainty in the predictions can be divided into two separate components: epistemic and aleatoric uncertainty. Epistemic uncertainty represents the lack of knowledge in the model, and it is reducible through improving the model or bringing in more training data. Aleatoric uncertainty, on the other hand, is inherent to the input data, and no amount of additional training data or model improvement will reduce it. Aleatoric uncertainty that varies over the input data is said to be heteroscedastic; a constant uncertainty is called homoscedastic.</p>
      <p id="d1e241">Many approaches to the quantification of uncertainty have been developed on the deep-learning side. One particularly important theme driving the development in this realm has been operational safety and countering overconfident predictions made by black box models overfitting the training data. Bayesian neural networks (BNNs) have emerged as a candidate for addressing that issue. They work by placing probability distributions over the weights, which are estimated via the means of Bayesian inference and yield a predictive distribution for data through their marginalization.</p>
      <p id="d1e244">Although exact Bayesian inference is intractable for large neural networks, suitable approximations exist. These are commonly divided into Markov chain Monte Carlo (MCMC)- and variational inference (VI)-based methods <xref ref-type="bibr" rid="bib1.bibx21" id="paren.19"/>. MCMC methods predict better weight distributions but are more computationally expensive and thus often reserved for small-scale problems where performance is key. VI, on the other hand, is more scalable and has been applied to larger neural networks. The idea behind variational inference is to approximate the true posterior of weights with a simpler analytic one (the variational posterior) and to estimate the variational posterior which is the closest to the true one. Thanks to advances by <xref ref-type="bibr" rid="bib1.bibx14" id="text.20"/> and subsequently <xref ref-type="bibr" rid="bib1.bibx6" id="text.21"/> with the Bayes by Backprop (BBB) algorithm, it is now possible to use mini-batch optimization for mean field VI (i.e., assuming fully factorizable variational posteriors) on large networks, opening up possibilities for the use of VI in problems such as precipitation nowcasting that require large amounts of input data and numerous model parameters.</p>
      <p id="d1e256">Later, Monte Carlo Dropout <xref ref-type="bibr" rid="bib1.bibx13" id="paren.22"/> techniques, among other variants, have been identified as being equivalent to approximate Bayesian inference due to losing some model expressivity but gaining ease of implementation. Based on this, <xref ref-type="bibr" rid="bib1.bibx23" id="text.23"/> have developed a technique for estimating the epistemic and heteroscedastic aleatoric variance components separately in deep-learning regression tasks. They estimate the epistemic uncertainty with the variance of predictions made via Monte Carlo Dropout and add a separate component to their network for predicting the aleatoric component. The predictions are modeled as having Gaussian likelihoods, with means equal to the prediction point estimates and variances equal to the aleatoric term described. These terms are then learned by minimizing a Gaussian negative log-likelihood loss function and taking them and observations as inputs. This approach has recently started to be applied to problems such as the segmentation of satellite images <xref ref-type="bibr" rid="bib1.bibx10" id="paren.24"/>, remaining useful for life prognostics <xref ref-type="bibr" rid="bib1.bibx9" id="paren.25"/> and long-term synoptic-scale precipitation forecasts <xref ref-type="bibr" rid="bib1.bibx54" id="paren.26"/>.</p>
</sec>
<sec id="Ch1.S1.SS4">
  <label>1.4</label><title>Model idea and research questions</title>
      <p id="d1e282">We propose the Deep Ensemble-based Uncertainty Combining radar Echo nowcasting (DEUCE) model for probabilistic precipitation nowcasting. The idea of the model is to apply the aleatoric and epistemic decomposition of uncertainty by <xref ref-type="bibr" rid="bib1.bibx23" id="text.27"/> to a Bayesian convolutional neural network with mean field variational inference to produce a predictive distribution of radar reflectivity. This distribution is then sampled to generate ensemble nowcasts of weather radar echo images. The research questions that we will attempt to answer are the following: <list list-type="order"><list-item>
      <p id="d1e290"><italic>Can we produce both powerful and reliable ensemble precipitation nowcasts using Bayesian neural networks with uncertainty decomposition?</italic> Specifically, is such a model competitive when compared to classical baseline models when assessed with a variety of quantitative probabilistic prediction skill metrics and based on a qualitative assessment?</p></list-item><list-item>
      <p id="d1e296"><italic>What are the characteristics of the aleatoric/epistemic decomposition?</italic> We are interested in the evolution of uncertainties with prediction lead time and whether they<?pagebreak page3842?> capture different and complementary features of the total predictive uncertainty.</p></list-item><list-item>
      <p id="d1e302"><italic>Can the model additionally be useful in producing deterministic precipitation nowcasts by means of averaging multiple predictions and leveraging the regulatory effect of probability distributions placed on weights?</italic> Do such predictions perform competitively when assessed against classical baseline models using quantitative verification metrics? Also, what can those metrics tell us about the nature of the predictions?</p></list-item></list></p>
</sec>
</sec>
<sec id="Ch1.S2">
  <label>2</label><title>Model description</title>
      <p id="d1e316">DEUCE builds upon a U-Net-based convolutional neural network (CNN) model of deterministic precipitation nowcasting and turns it into a Bayesian neural network with variational inference for making the predictions stochastic, enabling us to model the uncertainty in this U-Net model. As mentioned, we build upon the work of <xref ref-type="bibr" rid="bib1.bibx23" id="text.28"/> for quantifying the uncertainty in the nowcasting task. Particularly, DEUCE attempts to decompose predictive uncertainty into aleatoric uncertainty (originating from data and irreducible) and epistemic uncertainty (induced by lacking knowledge and reducible) by predicting reflectivity fields along with the aleatoric uncertainty associated with them explicitly. Epistemic uncertainty, in turn, is estimated from the variance in the reflectivity fields sampled, and it is combined with aleatoric uncertainty at the inference time in order to yield an approximation of the total predictive uncertainty. From this point onward, symbols in bold will refer to quantities represented by multidimensional arrays which we interchangeably call tensors. Scalar quantities will be represented with non-bold symbols.</p>
<sec id="Ch1.S2.SS1">
  <label>2.1</label><title>Functional model</title>
      <p id="d1e329">A neural network <inline-formula><mml:math id="M4" display="inline"><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mi mathvariant="bold-italic">θ</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula> is a universal function approximator which can be used for regression tasks, mapping an input tensor <inline-formula><mml:math id="M5" display="inline"><mml:mi mathvariant="bold-italic">x</mml:mi></mml:math></inline-formula> to a predicted output tensor <inline-formula><mml:math id="M6" display="inline"><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover></mml:math></inline-formula> and approximating a ground-truth output tensor <inline-formula><mml:math id="M7" display="inline"><mml:mi mathvariant="bold-italic">y</mml:mi></mml:math></inline-formula>, using its learned network parameters <inline-formula><mml:math id="M8" display="inline"><mml:mi mathvariant="bold-italic">θ</mml:mi></mml:math></inline-formula> that are represented as a list of tensors. In the case of radar-based precipitation nowcasting with neural networks, we approximate a function mapping the spatiotemporal time series of past radar observation images <inline-formula><mml:math id="M9" display="inline"><mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mtext>in</mml:mtext></mml:msub></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>, where <inline-formula><mml:math id="M10" display="inline"><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mtext>in</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> corresponds to the number of input time steps and to future radar observation images <inline-formula><mml:math id="M11" display="inline"><mml:mrow><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mtext>out</mml:mtext></mml:msub></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>, where <inline-formula><mml:math id="M12" display="inline"><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mtext>out</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> corresponds to the number of output time steps. In the DEUCE model, both <inline-formula><mml:math id="M13" display="inline"><mml:mi mathvariant="bold-italic">x</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M14" display="inline"><mml:mi mathvariant="bold-italic">y</mml:mi></mml:math></inline-formula> represent processed radar reflectivity data, and the network <inline-formula><mml:math id="M15" display="inline"><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mi mathvariant="bold-italic">θ</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mo>,</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> outputs a tuple of predicted reflectivity field time series <inline-formula><mml:math id="M16" display="inline"><mml:mrow><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mn mathvariant="normal">2</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mtext>out</mml:mtext></mml:msub></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>, along with fields estimating the aleatoric uncertainties <inline-formula><mml:math id="M17" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>=</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mtext>out</mml:mtext></mml:msub></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:math></inline-formula> corresponding to each of the pixels of <inline-formula><mml:math id="M18" display="inline"><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover></mml:math></inline-formula>.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F1" specific-use="star"><?xmltex \currentcnt{1}?><?xmltex \def\figurename{Figure}?><label>Figure 1</label><caption><p id="d1e637">The DEUCE encoder and decoder architectural components depicted on the left, along with the architectural diagram making use of those components on the right. Feature maps at different scales are extracted in the encoder branch before being passed to decoder branches, providing the outputs of the network.</p></caption>
          <?xmltex \igopts{width=455.244094pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/17/3839/2024/gmd-17-3839-2024-f01.png"/>

        </fig>

      <p id="d1e646">For the task of precipitation nowcasting, the neural network has to be capable of outputting predictions for multiple lead times, i.e., discrete time steps in the future, corresponding to future radar observations. DEUCE achieves this using a variant of U-Net as its functional architecture, taking in a sequence of 12 radar reflectivity fields <inline-formula><mml:math id="M19" display="inline"><mml:mi mathvariant="bold-italic">x</mml:mi></mml:math></inline-formula>, predicting <inline-formula><mml:math id="M20" display="inline"><mml:mrow><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mo>,</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> that corresponds to the nowcast for the next 12 time steps in a single forward pass.</p>
      <p id="d1e675">A schematic representation of the main components of DEUCE and how they are connected is presented in Fig. <xref ref-type="fig" rid="Ch1.F1"/>. The architecture consists of a single encoder branch, extracting features from <inline-formula><mml:math id="M21" display="inline"><mml:mi mathvariant="bold-italic">x</mml:mi></mml:math></inline-formula> at different spatial scales and semantic levels. The feature maps from these different scales are preserved for later use through skip connections. The (largest-scale) latent state produced by the encoder and the intermediate feature maps mediated by skip connections are then fed to two independent decoders: one outputting <inline-formula><mml:math id="M22" display="inline"><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover></mml:math></inline-formula> and the other outputting <inline-formula><mml:math id="M23" display="inline"><mml:mrow><mml:mi>log⁡</mml:mi><mml:msubsup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mi mathvariant="normal">al</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:math></inline-formula>. Using separate decoders for the outputs is preferable over a single combined decoder to avoid the blending of adjacent features, which would be detrimental to the expressivity of the model.</p>
      <p id="d1e712">The network contains two-dimensional (spatial) convolutions. These are represented by <monospace>conv2d 3x3</monospace> and <monospace>conv2d 1x1</monospace> labels denoting layers with filter sizes 3 and 1, respectively. Using 2D convolutions, temporal dependencies are only present implicitly. This approximation casts the nowcasting task as a simple image sequence-to-sequence translation problem, which reduces the computational resources needed compared to explicit modeling of the temporal aspect. The convolutional layers use partial convolutions <xref ref-type="bibr" rid="bib1.bibx28" id="paren.29"/> in which missing values are masked, and only valid values are used to normalize the convolutions. Although we do not work with missing data, this design choice helps in providing better-quality predictions near image borders by reducing, e.g., various artifacts related to them. <monospace>ReLU</monospace> denotes the rectified linear unit activation function, <monospace>concat</monospace> is concatenation along the channel dimension, <monospace>upsample</monospace> is nearest-neighbor upsampling by a scale of 2, and <monospace>Max Pooling</monospace> is maximum pooling by a scale of 2.</p>
</sec>
<sec id="Ch1.S2.SS2">
  <label>2.2</label><title>Stochastic model</title>
      <p id="d1e745">Conventional neural networks are deterministic in their nature, meaning that they only ever yield the same output <inline-formula><mml:math id="M24" display="inline"><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover></mml:math></inline-formula> for a given input <inline-formula><mml:math id="M25" display="inline"><mml:mi mathvariant="bold-italic">x</mml:mi></mml:math></inline-formula> and parameters <inline-formula><mml:math id="M26" display="inline"><mml:mi mathvariant="bold-italic">θ</mml:mi></mml:math></inline-formula>. Our goal is to produce a reliable estimate of the uncertainty associated with the approximation produced by the neural network. Because this approximation is merely a function of the input data and the functional model including parameters, considering the uncertainty in these sources separately should allow the approximation of the total predictive uncertainty in nowcasts.</p>
      <p id="d1e772">Hence, epistemic uncertainty is modeled by placing probability distributions on functional model parameters <inline-formula><mml:math id="M27" display="inline"><mml:mi mathvariant="bold-italic">θ</mml:mi></mml:math></inline-formula>, effectively turning the model stochastic. A Bayesian approach is taken in this regard, placing a prior distribution upon the<?pagebreak page3843?> weights and estimating the most likely posterior distribution given that prior and the training data. The estimation of the true posterior is an intractable task for a large-scale neural network, which is why variational inference (VI) is used to learn approximate posterior estimates for weights. VI limits the space of acceptable posterior distributions to a parameterized family, whose learned parameters replace the point estimates of classical neural network weights. Here, we aim to minimize the Kullback–Leibler (KL) divergence <xref ref-type="bibr" rid="bib1.bibx25" id="paren.30"/> <inline-formula><mml:math id="M28" display="inline"><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mi mathvariant="normal">KL</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> between the true and variational posteriors, which is a measure of the similarity between two probability distributions. As such, the objective is stated as
            <disp-formula id="Ch1.E1" content-type="numbered"><label>1</label><mml:math id="M29" display="block"><mml:mtable rowspacing="0.2ex" class="split" displaystyle="true" columnalign="right left"><mml:mtr><mml:mtd><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>*</mml:mo></mml:msup></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mo>=</mml:mo><mml:mi>arg⁡</mml:mi><mml:munder><mml:mo movablelimits="false">min⁡</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi></mml:munder><mml:msub><mml:mi>D</mml:mi><mml:mi mathvariant="normal">KL</mml:mi></mml:msub><mml:mo>[</mml:mo><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mo>‖</mml:mo><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="script">D</mml:mi><mml:mo>)</mml:mo><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mo>=</mml:mo><mml:mi>arg⁡</mml:mi><mml:munder><mml:mo movablelimits="false">min⁡</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi></mml:munder><mml:mo movablelimits="false">∫</mml:mo><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mi>log⁡</mml:mi><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="script">D</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mfrac></mml:mstyle><mml:mi mathvariant="normal">d</mml:mi><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
          where <inline-formula><mml:math id="M30" display="inline"><mml:mi mathvariant="bold-italic">θ</mml:mi></mml:math></inline-formula> denotes the variational posterior parameters, <inline-formula><mml:math id="M31" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>*</mml:mo></mml:msup></mml:mrow></mml:math></inline-formula> the optimal parameters, <inline-formula><mml:math id="M32" display="inline"><mml:mi mathvariant="bold-italic">w</mml:mi></mml:math></inline-formula> the sampled network weights, <inline-formula><mml:math id="M33" display="inline"><mml:mrow><mml:mi mathvariant="script">D</mml:mi><mml:mo>=</mml:mo><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> the problem data, <inline-formula><mml:math id="M34" display="inline"><mml:mrow><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> the variational posterior, and <inline-formula><mml:math id="M35" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="script">D</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> the exact posterior of network weights. In practice, this is not directly solvable, so the optimization is accomplished through the maximization of an evidence lower-bound (ELBO) proxy objective. The objective is defined as

                <disp-formula specific-use="gather" content-type="numbered"><mml:math id="M36" display="block"><mml:mtable displaystyle="true"><mml:mlabeledtr id="Ch1.E2"><mml:mtd><mml:mtext>2</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mtable rowspacing="0.2ex" class="split" displaystyle="true" columnalign="right left"><mml:mtr><mml:mtd><mml:mrow><mml:mtext>ELBO</mml:mtext><mml:mo>(</mml:mo><mml:mi mathvariant="script">D</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="bold">E</mml:mi><mml:mrow><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msub><mml:mo>[</mml:mo><mml:mi>log⁡</mml:mi><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="script">D</mml:mi><mml:mo>)</mml:mo><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mo>-</mml:mo><mml:msub><mml:mi mathvariant="bold">E</mml:mi><mml:mrow><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msub><mml:mo>[</mml:mo><mml:mi>log⁡</mml:mi><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mo>]</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mtd></mml:mlabeledtr><mml:mlabeledtr id="Ch1.E3"><mml:mtd><mml:mtext>3</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mtable rowspacing="0.2ex" class="split" displaystyle="true" columnalign="right left"><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mo>=</mml:mo><mml:mover><mml:mover class="overbrace" accent="true"><mml:mrow><mml:msub><mml:mi mathvariant="bold">E</mml:mi><mml:mrow><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msub><mml:mo>[</mml:mo><mml:mi>log⁡</mml:mi><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="script">D</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mo>)</mml:mo><mml:mo>]</mml:mo></mml:mrow><mml:mo mathvariant="normal">︷</mml:mo></mml:mover><mml:mtext>likelihood</mml:mtext></mml:mover></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mspace width="1em" linebreak="nobreak"/><mml:mo>+</mml:mo><mml:munder><mml:munder class="underbrace"><mml:mrow><mml:mover><mml:mover accent="true" class="overbrace"><mml:mrow><mml:msub><mml:mi mathvariant="bold">E</mml:mi><mml:mrow><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msub><mml:mo>[</mml:mo><mml:mi>log⁡</mml:mi><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mo>)</mml:mo><mml:mo>]</mml:mo></mml:mrow><mml:mo mathvariant="normal">︷</mml:mo></mml:mover><mml:mtext>prior</mml:mtext></mml:mover><mml:mo>-</mml:mo><mml:mover><mml:mover class="overbrace" accent="true"><mml:mrow><mml:msub><mml:mi mathvariant="bold">E</mml:mi><mml:mrow><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msub><mml:mo>[</mml:mo><mml:mi>log⁡</mml:mi><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mo>]</mml:mo></mml:mrow><mml:mo mathvariant="normal">︷</mml:mo></mml:mover><mml:mtext>posterior</mml:mtext></mml:mover></mml:mrow><mml:mo mathvariant="normal">︸</mml:mo></mml:munder><mml:mrow><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:msub><mml:mi>D</mml:mi><mml:mtext>KL</mml:mtext></mml:msub><mml:mo>[</mml:mo><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mo>‖</mml:mo><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mo>)</mml:mo><mml:mo>]</mml:mo><mml:mo>,</mml:mo><mml:mtext>i.e., the complexity term</mml:mtext></mml:mrow></mml:munder><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula>

            consisting of the log-likelihood, log-prior, and log-posteriors, with the last two terms commonly grouped together as the complexity term. Here <inline-formula><mml:math id="M37" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">E</mml:mi><mml:mrow><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> denotes the expected value of the probability density of interest over the variational posteriors.</p>
      <p id="d1e1326">According to <xref ref-type="bibr" rid="bib1.bibx6" id="text.31"/>, in Bayesian neural networks and using mini-batch optimization, the ELBO objective, as stated in Eq. (<xref ref-type="disp-formula" rid="Ch1.E3"/>), can be approximated as
            <disp-formula id="Ch1.E4" content-type="numbered"><label>4</label><mml:math id="M38" display="block"><mml:mtable class="split" rowspacing="0.2ex" displaystyle="true" columnalign="right left"><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:msubsup><mml:mtext>ELBO</mml:mtext><mml:mi>i</mml:mi><mml:mi mathvariant="italic">π</mml:mi></mml:msubsup><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mo>≈</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mi>N</mml:mi></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:munderover><mml:mo mathsize="1.5em">(</mml:mo><mml:mi>log⁡</mml:mi><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>∣</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mspace linebreak="nobreak" width="1em"/><mml:mo>+</mml:mo><mml:msub><mml:mi mathvariant="italic">π</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mi>log⁡</mml:mi><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo><mml:mo>-</mml:mo><mml:msub><mml:mi mathvariant="italic">π</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mi>log⁡</mml:mi><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mo mathsize="1.5em">)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
          which acts as an unbiased Monte Carlo estimator of the ELBO and is our final loss function. Here, the cost is calculated for each <inline-formula><mml:math id="M39" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula>th of the <inline-formula><mml:math id="M40" display="inline"><mml:mi>M</mml:mi></mml:math></inline-formula> mini-batches in an epoch drawing <inline-formula><mml:math id="M41" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula> Monte Carlo samples of the variational posteriors of the weights each time. <inline-formula><mml:math id="M42" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">π</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> denotes an arbitrary weighting of the complexity term, using the same rule as in <xref ref-type="bibr" rid="bib1.bibx6" id="text.32"/> in this work, which is <inline-formula><mml:math id="M43" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">π</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msup><mml:mn mathvariant="normal">2</mml:mn><mml:mrow><mml:mi>M</mml:mi><mml:mo>-</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msup><mml:mo>/</mml:mo><mml:mo>(</mml:mo><mml:msup><mml:mn mathvariant="normal">2</mml:mn><mml:mi>M</mml:mi></mml:msup><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. This serves to make the regularization effect of the prior stronger earlier, allowing data to be more important later in the training. In DEUCE, the variational posterior distributions <inline-formula><mml:math id="M44" display="inline"><mml:mi>q</mml:mi></mml:math></inline-formula> are modeled as diagonal Gaussian distributions, and the Bayes By Backprop (BBB) algorithm, using the re-parameterization trick by <xref ref-type="bibr" rid="bib1.bibx6" id="text.33"/>, is employed for their optimization. The prior distribution <inline-formula><mml:math id="M45" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, on the contrary, is fixed as a hyperparameter and is identically and independently distributed for each parameter as a normal distribution with zero mean and a variance of 0.1. This allows us to potentially calculate the complexity cost in closed form <xref ref-type="bibr" rid="bib1.bibx19" id="paren.34"/>, rather than with the Monte Carlo estimate of Eq. (<xref ref-type="disp-formula" rid="Ch1.E4"/>), hence reducing the computational cost of training.</p>
      <?pagebreak page3844?><p id="d1e1580">The likelihood cost of Eq. (<xref ref-type="disp-formula" rid="Ch1.E4"/>), similar to <xref ref-type="bibr" rid="bib1.bibx23" id="text.35"/>, is modeled for the <inline-formula><mml:math id="M46" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula>th mini-batch and the <inline-formula><mml:math id="M47" display="inline"><mml:mi>n</mml:mi></mml:math></inline-formula>th Monte Carlo sample as the Gaussian log likelihood
            <disp-formula id="Ch1.E5" content-type="numbered"><label>5</label><mml:math id="M48" display="block"><mml:mrow><mml:mi>log⁡</mml:mi><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="script">D</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mi>P</mml:mi></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>P</mml:mi></mml:munderover><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:mfrac></mml:mstyle><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">s</mml:mi><mml:mi>p</mml:mi></mml:msub></mml:mrow></mml:msup><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>p</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mi>p</mml:mi></mml:msub><mml:msup><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>+</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:mfrac></mml:mstyle><mml:msub><mml:mi mathvariant="bold-italic">s</mml:mi><mml:mi>p</mml:mi></mml:msub><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
          where the cost is averaged over <inline-formula><mml:math id="M49" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mi mathvariant="normal">…</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:math></inline-formula> pixels of the <inline-formula><mml:math id="M50" display="inline"><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mi mathvariant="normal">out</mml:mi></mml:msub><mml:mo>×</mml:mo><mml:mi>W</mml:mi><mml:mo>×</mml:mo><mml:mi>H</mml:mi></mml:mrow></mml:math></inline-formula> spatiotemporal time series <inline-formula><mml:math id="M51" display="inline"><mml:mi mathvariant="bold-italic">s</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M52" display="inline"><mml:mi mathvariant="bold-italic">y</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M53" display="inline"><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover></mml:math></inline-formula>. <inline-formula><mml:math id="M54" display="inline"><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mi mathvariant="normal">out</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> refers to the length of the time series, <inline-formula><mml:math id="M55" display="inline"><mml:mi>W</mml:mi></mml:math></inline-formula> refers to the width of the images, and <inline-formula><mml:math id="M56" display="inline"><mml:mi>H</mml:mi></mml:math></inline-formula> refers to the height of the images. <inline-formula><mml:math id="M57" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M58" display="inline"><mml:mi>n</mml:mi></mml:math></inline-formula> indices of fields are omitted here for clarity. Here, <inline-formula><mml:math id="M59" display="inline"><mml:mi mathvariant="bold-italic">y</mml:mi></mml:math></inline-formula> denotes the observed reflectivity fields, <inline-formula><mml:math id="M60" display="inline"><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover></mml:math></inline-formula> denotes the predicted reflectivity fields using the <inline-formula><mml:math id="M61" display="inline"><mml:mi>n</mml:mi></mml:math></inline-formula>th weights sampled from the network, and <inline-formula><mml:math id="M62" display="inline"><mml:mrow><mml:mi mathvariant="bold-italic">s</mml:mi><mml:mo>:=</mml:mo><mml:mi>log⁡</mml:mi><mml:msup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> refers to the corresponding logarithm of the aleatoric variances predicted by the network with those weights. The logarithm of the aleatoric variances estimate is taken because optimizing using it is more computationally stable and was found to work better than simply using variance constrained to be positive with a ReLU output activation function, especially when dealing with variances approaching zero.</p>
</sec>
<sec id="Ch1.S2.SS3">
  <label>2.3</label><title>Generation of ensemble nowcasts</title>
      <p id="d1e1849">The procedure for producing the primary outputs of DEUCE to make probabilistic nowcasts is presented in Fig. <xref ref-type="fig" rid="Ch1.F2"/>. First, <inline-formula><mml:math id="M63" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula> raw network outputs are produced, which are stochastic pairs of reflectivity field sequences <inline-formula><mml:math id="M64" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mi>n</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and logarithmic aleatoric variance field sequences <inline-formula><mml:math id="M65" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mi>n</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. Each <inline-formula><mml:math id="M66" display="inline"><mml:mi>n</mml:mi></mml:math></inline-formula>th of those sampled outputs draws different weights from the learned variational posterior distributions, which is reflected in the output distribution. The <inline-formula><mml:math id="M67" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mi>n</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> are converted to their non-logarithmic version <inline-formula><mml:math id="M68" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mi>n</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:math></inline-formula>, and individual stochastic runs are stacked into a pair of raw ensembles <inline-formula><mml:math id="M69" display="inline"><mml:mrow><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mo>,</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>. At this point, the epistemic uncertainty is embedded in <inline-formula><mml:math id="M70" display="inline"><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover></mml:math></inline-formula>, but the aleatoric uncertainty is separate and only present in <inline-formula><mml:math id="M71" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mtext>al</mml:mtext><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:math></inline-formula>. Hence, in order to allow the combination of these uncertainties, <inline-formula><mml:math id="M72" display="inline"><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover></mml:math></inline-formula> is divided into the prediction mean <inline-formula><mml:math id="M73" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mtext>mean</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> and the epistemic variance <inline-formula><mml:math id="M74" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mtext>ep</mml:mtext><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:math></inline-formula> by taking the mean and variance over <inline-formula><mml:math id="M75" display="inline"><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover></mml:math></inline-formula>, respectively. Additionally, the aleatoric variance <inline-formula><mml:math id="M76" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> is summarized by taking its mean, denoted <inline-formula><mml:math id="M77" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mtext>al</mml:mtext><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:math></inline-formula>. These three outputs, <inline-formula><mml:math id="M78" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mtext>mean</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M79" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mtext>ep</mml:mtext><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:math></inline-formula>, and <inline-formula><mml:math id="M80" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mtext>al</mml:mtext><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:math></inline-formula>, form the base from which probabilistic nowcasts are computed.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F2" specific-use="star"><?xmltex \currentcnt{2}?><?xmltex \def\figurename{Figure}?><label>Figure 2</label><caption><p id="d1e2074">The prediction procedure for the primary outputs of the DEUCE is model illustrated. Each sampled output is computed separately with a forward pass through the network, yielding a time series of the predictions and the logarithmic aleatoric variances, which are converted back to variances. After agglomeration into a pair of raw ensembles, the prediction mean <inline-formula><mml:math id="M81" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mtext>mean</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> and the two types of uncertainties, the epistemic variances <inline-formula><mml:math id="M82" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mtext>ep</mml:mtext><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:math></inline-formula> and aleatoric variances <inline-formula><mml:math id="M83" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mtext>al</mml:mtext><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:math></inline-formula>, are computed from the pair. These three quantities are the ones used for producing the final prediction ensemble.</p></caption>
          <?xmltex \igopts{width=455.244094pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/17/3839/2024/gmd-17-3839-2024-f02.png"/>

        </fig>

      <p id="d1e2123">The total mean and uncertainty in the prediction can thus be estimated as
            <disp-formula id="Ch1.E6" content-type="numbered"><label>6</label><mml:math id="M84" display="block"><mml:mtable rowspacing="0.2ex" class="split" displaystyle="true" columnalign="right left"><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mtext>mean</mml:mtext></mml:msub></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mi>N</mml:mi></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:munderover><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mi>n</mml:mi></mml:msub><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:msubsup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mtext>pred</mml:mtext><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>≈</mml:mo><mml:mover><mml:mover class="overbrace" accent="true"><mml:mrow><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mi>N</mml:mi></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:munderover><mml:msubsup><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mi>n</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>-</mml:mo><mml:mo>(</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mi>N</mml:mi></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:munderover><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mi>n</mml:mi></mml:msub><mml:msup><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow><mml:mo mathvariant="normal">︷</mml:mo></mml:mover><mml:mrow><mml:msubsup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mtext>ep</mml:mtext><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:mover><mml:mo>+</mml:mo><mml:mover><mml:mover accent="true" class="overbrace"><mml:mrow><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mi>N</mml:mi></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:munderover><mml:msubsup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mi>n</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow><mml:mo mathvariant="normal">︷</mml:mo></mml:mover><mml:mrow><mml:msubsup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mtext>al</mml:mtext><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:mover><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
          where <inline-formula><mml:math id="M85" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mtext>pred</mml:mtext><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:math></inline-formula> denotes the predictive variance. This means that the predictive variance can be estimated as the sum of the variance of the predicted reflectivity fields, which is the epistemic variance, and of the mean of the predicted aleatoric variance fields. These quantities are sufficient for making probabilistic nowcasts such as calculating exceedance probabilities for future radar reflectivity values, allowing us to model the predictive distribution of this reflectivity as normally distributed with mean <inline-formula><mml:math id="M86" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mtext>mean</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> and variance <inline-formula><mml:math id="M87" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mtext>pred</mml:mtext><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:math></inline-formula>. This formulation is admissible, as reflectivity of precipitation (in dBZ units) is known to have a normal distribution which follows from the distribution of the precipitation rate being log-normal <xref ref-type="bibr" rid="bib1.bibx22" id="paren.36"/>. Additionally, it is interesting to note that the relationship between the ensemble and the predictive distribution here is opposite to that of NWP, where perturbations to initial conditions lead to ensembles that themselves define the predictive distribution <xref ref-type="bibr" rid="bib1.bibx5" id="paren.37"/>.</p>
      <p id="d1e2354">Nevertheless, some applications of probabilistic precipitation nowcasting – such as flood modeling – assume ensemble-based nowcasts, where each member of the ensemble represents a physically plausible precipitation scenario. One could of course randomly sample the predictive distribution to generate an ensemble, which would correctly approximate pixel-wise statistics, but the spatiotemporal structure of the fields would be lost. In an attempt to remedy this, we post-process outputs to generate ensemble members respecting the spatial covariance structure of the input field <inline-formula><mml:math id="M88" display="inline"><mml:mi mathvariant="bold-italic">x</mml:mi></mml:math></inline-formula> as
            <disp-formula id="Ch1.E7" content-type="numbered"><label>7</label><mml:math id="M89" display="block"><mml:mrow><mml:msubsup><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mi>n</mml:mi><mml:mtext>ens</mml:mtext></mml:msubsup><mml:mo>=</mml:mo><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mtext>mean</mml:mtext></mml:msub><mml:mo>+</mml:mo><mml:msqrt><mml:mrow><mml:msubsup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mtext>pred</mml:mtext><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:msqrt><mml:mo>⊗</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">ϵ</mml:mi><mml:mrow><mml:mtext>corr</mml:mtext><mml:mo>,</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
          where <inline-formula><mml:math id="M90" display="inline"><mml:mrow><mml:msubsup><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mi>n</mml:mi><mml:mtext>ens</mml:mtext></mml:msubsup><mml:mo>=</mml:mo><mml:msubsup><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mrow><mml:mi>n</mml:mi><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mtext>ens</mml:mtext></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mrow><mml:mi>n</mml:mi><mml:mo>,</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow><mml:mtext>ens</mml:mtext></mml:msubsup><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:msubsup><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mrow><mml:mi>n</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mtext>out</mml:mtext></mml:msub></mml:mrow><mml:mtext>ens</mml:mtext></mml:msubsup></mml:mrow></mml:math></inline-formula> denotes the newly generated ensemble member, <inline-formula><mml:math id="M91" display="inline"><mml:mo>⊗</mml:mo></mml:math></inline-formula> denotes an element-wise multiplication broadcast over <inline-formula><mml:math id="M92" display="inline"><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mtext>out</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> frames, and <inline-formula><mml:math id="M93" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">ϵ</mml:mi><mml:mrow><mml:mtext>corr</mml:mtext><mml:mo>,</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> is a correlated Gaussian random field of shape <inline-formula><mml:math id="M94" display="inline"><mml:mrow><mml:mi>W</mml:mi><mml:mo>×</mml:mo><mml:mi>H</mml:mi></mml:mrow></mml:math></inline-formula>. <inline-formula><mml:math id="M95" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">ϵ</mml:mi><mml:mrow><mml:mtext>corr</mml:mtext><mml:mo>,</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> is generated to match the average spatial correlation structure of <inline-formula><mml:math id="M96" display="inline"><mml:mi mathvariant="bold-italic">x</mml:mi></mml:math></inline-formula> using fast Fourier transform (FFT) filtering. The structure is obtained non-parametrically from the power spectrum of <inline-formula><mml:math id="M97" display="inline"><mml:mi mathvariant="bold-italic">x</mml:mi></mml:math></inline-formula> <xref ref-type="bibr" rid="bib1.bibx42" id="paren.38"/>. The technique is equivalent to that used to generate perturbation fields in STEPS <xref ref-type="bibr" rid="bib1.bibx34" id="paren.39"/>. Even though this method accounts for the spatial structure of the precipitation time series, it is not capable of modeling its temporal structure, which is assumed constant. The ensembles produced in this way shall be denoted <inline-formula><mml:math id="M98" display="inline"><mml:mrow><mml:msup><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mtext>ens</mml:mtext></mml:msup></mml:mrow></mml:math></inline-formula>, which is in contrast to the raw predicted reflectivity fields denoted <inline-formula><mml:math id="M99" display="inline"><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover></mml:math></inline-formula>.</p>
</sec>
</sec>
<sec id="Ch1.S3">
  <label>3</label><title>Experimental details</title>
      <?pagebreak page3845?><p id="d1e2606">This section presents the experiments performed. First, in Sect. <xref ref-type="sec" rid="Ch1.S3.SS1"/>, we present the dataset used, followed by the details related to the training of DEUCE in Sect. <xref ref-type="sec" rid="Ch1.S3.SS2"/>, and the verification experiments in Sect. <xref ref-type="sec" rid="Ch1.S3.SS3"/>. Additional technical details can, on the other hand, be found in Appendix <xref ref-type="sec" rid="App1.Ch1.S1"/>. <?xmltex \hack{\newpage}?></p>
<sec id="Ch1.S3.SS1">
  <label>3.1</label><title>Data</title>
      <p id="d1e2625">The dataset used for this work comes from the Finnish Meteorological Institute radar network. It consists of cropped lowest-altitude radar reflectivity composites chosen from rainy days during the summer period of the years 2019–2021. The dataset is identical to that used by <xref ref-type="bibr" rid="bib1.bibx38" id="text.40"/>, except for using longer time series. The composites are built from the two lowest-elevation angle scans interpolated into an <inline-formula><mml:math id="M100" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> km Cartesian grid. The chosen area covers southern Finland, with the bottom-left corner at coordinates (59.01° N, 20.55° E) and the top-right corner at coordinates (63.62° N, 30.27° E). The spatial extent of this crop is <inline-formula><mml:math id="M101" display="inline"><mml:mrow><mml:mn mathvariant="normal">512</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">512</mml:mn></mml:mrow></mml:math></inline-formula> km, corresponding to <inline-formula><mml:math id="M102" display="inline"><mml:mrow><mml:mn mathvariant="normal">512</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">512</mml:mn></mml:mrow></mml:math></inline-formula> pixel square images, suitable for training a neural network. The composites are available with a temporal resolution of 5 min. The extent of the bounding box is additionally illustrated in Fig. <xref ref-type="fig" rid="Ch1.F3"/>, along with the coverage of Finnish Meteorological Institute radars. From this we see that the advantage of the crop is that it has a higher-density radar cover than its surroundings.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F3" specific-use="star"><?xmltex \currentcnt{3}?><?xmltex \def\figurename{Figure}?><label>Figure 3</label><caption><p id="d1e2671">The Finnish Meteorological Institute radar network with its 11 radars and the bounding box used. Each radar is described by its three-letter code, with their 120 km coverage radii for snowfall in gray and the intersection of 250 km coverage radii for rainfall as the black outline. An example radar composite crop from a precipitation event (15 August 2019 at 15:00:00 UTC) is visualized in the enlarged version of the bounding box on the right.</p></caption>
          <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/17/3839/2024/gmd-17-3839-2024-f03.png"/>

        </fig>

      <p id="d1e2680">The data were selected on a day-by-day basis, selecting the 100 d with the most pixels having reflectivity values over 35 dBZ. The days were then divided into 6 h long blocks from which blocks with below 1 % of the pixels with reflectivity values over 20 dBZ were removed. These remaining blocks were then randomly split into training, validation, and verification datasets with a ratio of <inline-formula><mml:math id="M103" display="inline"><mml:mrow><mml:mn mathvariant="normal">6</mml:mn><mml:mo>:</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>:</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula>. The division into blocks was done in order to limit the number of successive time series present in different splits, as they exhibit high correlation, and not using any blocks would make the training, validation, and verification sets dependent, as the same events would be present in all of them. A time of 6 h was deemed a sufficient time for temporal correlations to mostly disappear. Last, 2 h long time series, corresponding to 24 images each, were then extracted from these blocks using a sliding window principle, with a stride of one, omitting those time series with missing data. The final training, validation, and verification datasets ended up containing 10 780, 1813, and 1666 time series, respectively.</p>
      <p id="d1e2700">The input time series were read from HDF5 files, stored there with an 8 bit scale-offset lossy compression scheme, ranging from <inline-formula><mml:math id="M104" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>32 to 96 dBZ at a resolution of 0.5 dBZ. The images were then converted to floating point values, and a threshold of 8 dBZ was applied, replacing values below the threshold with <inline-formula><mml:math id="M105" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>10 dBZ. This served as a simple way to remove non-meteorological targets and other clutter that could interfere with the training and prediction while maintaining most of the relevant precipitation echoes. Finally, the reflectivity values were normalized between zero and one. Computed predictions were converted back into reflectivity values by applying the inverse of the transformation before saving them, using the same scheme as with the input data.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F4"><?xmltex \currentcnt{4}?><?xmltex \def\figurename{Figure}?><label>Figure 4</label><caption><p id="d1e2719">Finnish Meteorological Institute composite crop dataset distribution of reflectivity. The threshold chosen below which the data are set to <inline-formula><mml:math id="M106" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>10 dBZ is shown with a dashed line. Additionally, a Gaussian probability density function (PDF) fit on the data above 20 dBZ is shown in red, which serves to illustrate the Gaussian distribution of precipitation reflectivity.</p></caption>
          <?xmltex \igopts{width=241.848425pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/17/3839/2024/gmd-17-3839-2024-f04.png"/>

        </fig>

      <p id="d1e2735">The dataset reflectivity distribution is depicted in Fig. <xref ref-type="fig" rid="Ch1.F4"/>, with the <inline-formula><mml:math id="M107" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>32 dBZ minimum value pixels left out of the histogram for clarity. The threshold is shown to divide the data into retained and discarded parts, and a Gaussian density is fitted to the part most likely to purely consist of precipitation, which exceeds 20 dBZ. Below the threshold, there seem to be multiple peaks in density, likely involving insects, birds, and miscellaneous clutter, as well as increasing noise the lower we go on the scale. While the highest-reflectivity density values seem to follow the Gaussian fit well, part of the density between 10 and 20 dBZ remains unexplained, and this range is likely to contain a mixture of precipitation and clutter. Still, this is not a major issue, since the most interesting<?pagebreak page3846?> precipitation to predict corresponds to reflectivity values well above 20 dBZ.</p>
</sec>
<sec id="Ch1.S3.SS2">
  <label>3.2</label><title>Training</title>
      <p id="d1e2755">For the training of the network, the Adam optimizer <xref ref-type="bibr" rid="bib1.bibx24" id="paren.41"/> was used with an initial learning rate of <inline-formula><mml:math id="M108" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">4</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> and other parameters set to their PyTorch default values. The network was trained with that learning rate for 20 epochs, after which the learning rate was lowered to <inline-formula><mml:math id="M109" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">5</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> for 8 more epochs, and finally further lowered to <inline-formula><mml:math id="M110" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">6</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> for 1 final epoch. A validation epoch was carried out after each epoch in which equitable threat score (ETS) <xref ref-type="bibr" rid="bib1.bibx20" id="paren.42"/> metrics were calculated for converted precipitation estimates (Sect. <xref ref-type="sec" rid="App1.Ch1.S1.SS1"/>) of predictions and summed over thresholds of 0.5, 1.0, 5.0, 10.0, 20.0, and 30.0 mm h<inline-formula><mml:math id="M111" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>, as well as each lead time. This validation score showed improvement over the whole training process.</p>
      <p id="d1e2833">The training procedure for DEUCE is presented in Fig. <xref ref-type="fig" rid="Ch1.F5"/> for a single epoch. Both input sequence lengths <inline-formula><mml:math id="M112" display="inline"><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mi mathvariant="normal">in</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and output sequence lengths <inline-formula><mml:math id="M113" display="inline"><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mi mathvariant="normal">out</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> were 12, corresponding to 1 h each. For the training and validation epochs, the batch size was set to two, and the number of produced Monte Carlo samples of posteriors <inline-formula><mml:math id="M114" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula> was set to two as well, which was the most that our GPU could fit during training. In order to increase the variance between the gradients of mini-batch members, Flipout re-parameterization <xref ref-type="bibr" rid="bib1.bibx51" id="paren.43"/> was applied to the sampled weights, multiplying the random sampling coefficient of weights with a random sign matrix and effectively adding randomness inside batches for a low computational cost. The closed-form of the KL divergence between two Gaussian distributions was used for the calculations of the ELBO complexity term instead of Monte Carlo estimates in the final model training.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F5" specific-use="star"><?xmltex \currentcnt{5}?><?xmltex \def\figurename{Figure}?><label>Figure 5</label><caption><p id="d1e2872">An illustrated training epoch for DEUCE. One loop corresponds to a single training sequence, which can be substituted for a single training mini-batch, taking multiple sequences in one batch. The DEUCE prediction process refers to that illustrated in Fig. <xref ref-type="fig" rid="Ch1.F2"/>, without the post-processing. The blue box labeled <inline-formula><mml:math id="M115" display="inline"><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mi mathvariant="normal">KL</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>q</mml:mi><mml:mo>‖</mml:mo><mml:mi>p</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> corresponds to the complexity term of the negative ELBO loss (to minimize), <inline-formula><mml:math id="M116" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">π</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> corresponds to its weighting coefficient, and the red box labeled <inline-formula><mml:math id="M117" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mi>log⁡</mml:mi><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> corresponds to the likelihood term of the negative ELBO loss. Monte Carlo estimates of the complexity term use sampled weights <inline-formula><mml:math id="M118" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, whereas the closed-form expression that we use is a function of parameters <inline-formula><mml:math id="M119" display="inline"><mml:mi mathvariant="bold-italic">θ</mml:mi></mml:math></inline-formula>.</p></caption>
          <?xmltex \igopts{width=483.69685pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/17/3839/2024/gmd-17-3839-2024-f05.png"/>

        </fig>

      <p id="d1e2960">The input time series <inline-formula><mml:math id="M120" display="inline"><mml:mrow><mml:msub><mml:mi>X</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> was pre-processed first, as described in Sect. <xref ref-type="sec" rid="Ch1.S3.SS1"/>, and, in the case of training data, was then augmented by applying in succession a random horizontal flip, a random vertical flip, and a rotation by an angle randomly chosen between 0, 90, 180, and 270°. This was done to improve the variety in the training dataset and consequently improve the generalization performance of the trained network.</p>
</sec>
<?pagebreak page3847?><sec id="Ch1.S3.SS3">
  <label>3.3</label><title>Verification</title>
      <p id="d1e2984">The performance of the DEUCE model is verified against the pySTEPS <xref ref-type="bibr" rid="bib1.bibx34" id="paren.44"/> implementation of multiple extrapolation-based precipitation methods. The verification is divided into the qualitative inspection of ensembles produced in two case studies, into an analysis of DEUCE uncertainty composition, into the verification of the (probabilistic) performance of the whole ensemble, and into the verification of the (deterministic) performance of the ensemble mean, i.e., its fidelity in representing the true variation in the radar images. The four types of verification performed, along with the relevant DEUCE product, the baseline models used, and the evaluation criteria are summarized in Table <xref ref-type="table" rid="Ch1.T1"/>.</p>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T1" specific-use="star"><?xmltex \currentcnt{1}?><label>Table 1</label><caption><p id="d1e2995">The four components of the verification process for DEUCE summarized.</p></caption><oasis:table frame="topbot"><?xmltex \begin{scaleboxenv}{.95}[.95]?><oasis:tgroup cols="4">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="left"/>
     <oasis:colspec colnum="3" colname="col3" align="left"/>
     <oasis:colspec colnum="4" colname="col4" align="left"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">DEUCE product</oasis:entry>
         <oasis:entry colname="col3">Baseline models (Sect. <xref ref-type="sec" rid="App1.Ch1.S1.SS2"/>)</oasis:entry>
         <oasis:entry colname="col4">Evaluation criteria</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">Case studies (Sect. <xref ref-type="sec" rid="Ch1.S3.SS3.SSS1"/>)</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M121" display="inline"><mml:mrow><mml:msup><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mi mathvariant="normal">ens</mml:mi></mml:msup></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">STEPS, LINDA-P</oasis:entry>
         <oasis:entry colname="col4">Ensemble mean/SD, exceedance probabilities</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Uncertainty composition (Sect. <xref ref-type="sec" rid="Ch1.S3.SS3.SSS2"/>)</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M122" display="inline"><mml:mrow><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mo>,</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">–</oasis:entry>
         <oasis:entry colname="col4">Case decomposed mean/SD and statistics</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Probabilistic perf. (Sect. <xref ref-type="sec" rid="Ch1.S3.SS3.SSS3"/>)</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M123" display="inline"><mml:mrow><mml:msup><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mi mathvariant="normal">ens</mml:mi></mml:msup></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">STEPS, LINDA-P</oasis:entry>
         <oasis:entry colname="col4">CRPS, reliability diagram, ROC AUC, rank hist.</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Deterministic perf. (Sect. <xref ref-type="sec" rid="Ch1.S3.SS3.SSS4"/>)</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M124" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mi mathvariant="normal">mean</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">Extrapolation, LINDA-D</oasis:entry>
         <oasis:entry colname="col4">ME, ETS, RAPSD</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup><?xmltex \end{scaleboxenv}?></oasis:table><?xmltex \gdef\@currentlabel{1}?></table-wrap>

      <p id="d1e3154">In probabilistic verification experiments, <inline-formula><mml:math id="M125" display="inline"><mml:mrow><mml:mi>N</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">48</mml:mn></mml:mrow></mml:math></inline-formula> ensemble members are used both for producing the raw outputs <inline-formula><mml:math id="M126" display="inline"><mml:mrow><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mo>,</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> and for drawing the post-processed ensemble <inline-formula><mml:math id="M127" display="inline"><mml:mrow><mml:msup><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mtext>ens</mml:mtext></mml:msup></mml:mrow></mml:math></inline-formula>, as well as for making the baseline ensemble model predictions. All of the predictions made for the verification of DEUCE are made until a 60 min lead time and thresholded at 8 dBZ, serving as an estimate for minimum observable precipitation. Four precipitation thresholds are considered where the verification involves evaluating the quality of a prediction exceeding a particular reflectivity value. Converted using the <inline-formula><mml:math id="M128" display="inline"><mml:mi>Z</mml:mi></mml:math></inline-formula>–<inline-formula><mml:math id="M129" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula> relationship presented in Sect. <xref ref-type="sec" rid="App1.Ch1.S1.SS1"/>, these are 20 dBZ (<inline-formula><mml:math id="M130" display="inline"><mml:mrow><mml:mo>≈</mml:mo><mml:mn mathvariant="normal">0.5</mml:mn></mml:mrow></mml:math></inline-formula> mm h<inline-formula><mml:math id="M131" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>), 25 dBZ (<inline-formula><mml:math id="M132" display="inline"><mml:mrow><mml:mo>≈</mml:mo><mml:mn mathvariant="normal">1.3</mml:mn></mml:mrow></mml:math></inline-formula> mm h<inline-formula><mml:math id="M133" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>), 35 dBZ (<inline-formula><mml:math id="M134" display="inline"><mml:mrow><mml:mo>≈</mml:mo><mml:mn mathvariant="normal">5.7</mml:mn></mml:mrow></mml:math></inline-formula> mm h<inline-formula><mml:math id="M135" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>), and 45 dBZ (<inline-formula><mml:math id="M136" display="inline"><mml:mrow><mml:mo>≈</mml:mo><mml:mn mathvariant="normal">25.5</mml:mn></mml:mrow></mml:math></inline-formula> mm h<inline-formula><mml:math id="M137" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>), which correspond to very light, light, moderate, and heavy rain, respectively. <?xmltex \hack{\newpage}?></p>
<sec id="Ch1.S3.SS3.SSS1">
  <label>3.3.1</label><title>Case studies</title>
      <p id="d1e3316">Two distinct rainfall events are chosen as case studies to provide a qualitative assessment, as well as a comparison of DEUCE nowcasts with the baseline probabilistic methods. The case studies each focus on an ensemble nowcast at a particular time step during the precipitation event that is chosen to include both large-scale weaker precipitation, which is characteristic of stratiform rainfall, or localized heavy precipitation, which is characteristic of convective rainfall. The latter has a shorter lifetime and has been traditionally harder to predict, but it is of interest to observe the performance of the model with both types. In addition, we include both instances of weakening and of intensification of echoes in the case studies. The cases are chosen from radar composite crops, with the area described in Sect. <xref ref-type="sec" rid="Ch1.S3.SS1"/> over the verification split, and the summer of the year 2022, which is separate from the dataset used for training, validation, and quantitative verification. The timestamp of the first case chosen is 9 July 2022 at 15:00:00 UTC. This case contains mostly convective rainfall, with some localized high rain rates. The second case chosen is 17 August 2021 at 16:50:00 UTC, which represents a very different scenario with large-scale, mostly stratiform rainfall. The radar images of the hour leading up to the timestamps are used as inputs and the following 1 h of radar images are predicted.</p>
      <p id="d1e3321">Three different visualizations of the cases are made at 5, 15, 30, and 60 min lead times, using the post-processed DEUCE ensembles <inline-formula><mml:math id="M138" display="inline"><mml:mrow><mml:msup><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mi mathvariant="normal">ens</mml:mi></mml:msup></mml:mrow></mml:math></inline-formula> and the probabilistic baseline models STEPS and LINDA-P described in Sect. <xref ref-type="sec" rid="App1.Ch1.S1.SS2"/> when appropriate. The first visualization is that of predictive means and standard deviations of the ensembles (in dBZ units). Here, DEUCE, STEPS, and LINDA-P are compared side by side.<?pagebreak page3848?> The second visualization is that of exceedance probabilities of DEUCE, STEPS, and LINDA-P ensemble nowcasts at a 25 dBZ reflectivity threshold. The third, and last, visualization depicts the exceedance probability of DEUCE in predicting reflectivity above 20, 25, 35, and 45 dBZ thresholds.</p>
</sec>
<sec id="Ch1.S3.SS3.SSS2">
  <label>3.3.2</label><title>Uncertainty composition analysis</title>
      <p id="d1e3348">The composition of the DEUCE predictive uncertainty is analyzed using the prediction of the first case study and using statistics aggregated over the verification dataset. For the case study prediction, the aleatoric and epistemic components of the predictive standard deviation are visualized next to the combined predictive uncertainty, the mean predictions, and the observations at lead times of 5, 15, 30, and 60 min. The statistics collected are the average magnitude of the aleatoric and epistemic standard deviation components under different conditions. These magnitudes are divided into bins corresponding to the prediction lead time and the observed reflectivity matching the pixel in question (5 dBZ bin width from 5 to 60 dBZ) and are collected for each prediction timestamp. The resulting statistics are visualized in the form of a histogram aggregated over the whole dataset and as bar plots showing the contribution of the uncertainties against lead time and observed reflectivity.</p>
</sec>
<sec id="Ch1.S3.SS3.SSS3">
  <label>3.3.3</label><title>Probabilistic performance verification</title>
      <p id="d1e3359">Probabilistic verification serves to assess the probabilistic predictive power of DEUCE ensembles, mostly in terms of prediction reliability and discrimination ability. In other words, it determines the quality and the variety of produced ensembles with regard to the true distribution of different future scenarios. Here, the DEUCE prediction is represented by the post-processed ensemble <inline-formula><mml:math id="M139" display="inline"><mml:mrow><mml:msup><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mtext>ens</mml:mtext></mml:msup></mml:mrow></mml:math></inline-formula>. The probabilistic baseline models used are STEPS <xref ref-type="bibr" rid="bib1.bibx8 bib1.bibx42" id="paren.45"/> and LINDA-P <xref ref-type="bibr" rid="bib1.bibx35" id="paren.46"/>. The description and configuration of those models are given in Sect. <xref ref-type="sec" rid="App1.Ch1.S1.SS2"/>.</p>
      <p id="d1e3384">The probabilistic performance metrics used are the continuous ranked probability score (CRPS) <xref ref-type="bibr" rid="bib1.bibx18 bib1.bibx52" id="paren.47"/>, which generalizes the mean absolute error in deterministic forecasts to probability distributions and is calculated for lead times up to 60 min. Next, the receiver operating characteristic (ROC) curve <xref ref-type="bibr" rid="bib1.bibx30 bib1.bibx52" id="paren.48"/>, along with the area under it (AUC), quantifies the discriminative power of the ensembles for predicting reflectivity values exceeding a certain threshold. ROC AUC is computed for reflectivity thresholds of 20, 25, 35, and 45 dBZ at lead times of 5, 15, 30, and 60 min. To measure forecast reliability and sharpness, we used the reliability diagram, along with its sharpness histogram <xref ref-type="bibr" rid="bib1.bibx52" id="paren.49"/>, and the expected calibration error (ECE) score <xref ref-type="bibr" rid="bib1.bibx31" id="paren.50"/>, which we computed for the same threshold and lead times as ROC curves. Finally, rank histograms <xref ref-type="bibr" rid="bib1.bibx52" id="paren.51"/> were calculated to measure the bias and spread of ensembles at lead times of 5, 15, 30, and 60 min. A detailed description of these metrics, along with the configurations used, is found in Sect. <xref ref-type="sec" rid="App1.Ch1.S1.SS3"/>.</p>
</sec>
<sec id="Ch1.S3.SS3.SSS4">
  <label>3.3.4</label><title>Deterministic performance verification</title>
      <p id="d1e3413">Deterministic verification serves to assess whether DEUCE ensemble means are useful themselves. It also gives insight into many interesting aspects of predictions, such as systematic biases and the possible loss of small-scale variability. Here, the DEUCE prediction is represented by the ensemble mean <inline-formula><mml:math id="M140" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mtext>mean</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula>. The deterministic baselines used are an extrapolation nowcast and LINDA-D <xref ref-type="bibr" rid="bib1.bibx35" id="paren.52"/>. The description and configuration of those models are again described in Sect. <xref ref-type="sec" rid="App1.Ch1.S1.SS2"/>.</p>
      <?pagebreak page3849?><p id="d1e3435">Three deterministic metrics are used to assess DEUCE ensemble means. The first is the mean error (ME) <xref ref-type="bibr" rid="bib1.bibx52" id="paren.53"/>, measuring the bias of nowcasts produced. The equitable threat score (ETS) <xref ref-type="bibr" rid="bib1.bibx20 bib1.bibx52" id="paren.54"/> then provides an estimate of the deterministic skill in forecasting reflectivity above a certain intensity threshold. It is calculated for lead times up to 60 min and thresholds of 20, 25, 35, and 45 dBZ. Finally, the radially averaged power spectral density (RAPSD) <xref ref-type="bibr" rid="bib1.bibx39 bib1.bibx49" id="paren.55"/> measures how well the power spectrum of reflectivity is maintained. It is summarized with a relative mean absolute error (MAE) score. We compute RAPSD for prediction lead times of 5, 15, 30, and 60 min. RAPSD is also calculated for individual <inline-formula><mml:math id="M141" display="inline"><mml:mrow><mml:msup><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mtext>ens</mml:mtext></mml:msup></mml:mrow></mml:math></inline-formula> members to analyze the possible contribution of the spatially correlated noise to maintaining the power spectrum. A detailed description of these metrics, along with the configurations used, is found in Sect. <xref ref-type="sec" rid="App1.Ch1.S1.SS4"/>. <?xmltex \hack{\newpage}?></p>
</sec>
</sec>
</sec>
<sec id="Ch1.S4">
  <label>4</label><title>Results</title>
      <p id="d1e3474">The results of the quantitative and qualitative analyses of model performance and fitness to the task indicate that DEUCE succeeds in its primary task of providing reasonably reliable probabilistic precipitation nowcasts but not in that of producing skillful deterministic nowcasts. This is illustrated by the summary of quantitative verification results is provided in Table <xref ref-type="table" rid="Ch1.T2"/>. These results will then be elaborated upon in detail and presented as figures in the following four subsections. Starting with the qualitative case study results in Sect. <xref ref-type="sec" rid="Ch1.S4.SS1"/>, we then present the composition of the uncertainty in Sect. <xref ref-type="sec" rid="Ch1.S4.SS2"/>, before continuing with the probabilistic performance metric results in Sect. <xref ref-type="sec" rid="Ch1.S4.SS3"/>, and finally presenting the deterministic performance metric results in Sect. <xref ref-type="sec" rid="Ch1.S4.SS4"/>.</p>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T2" specific-use="star"><?xmltex \currentcnt{2}?><label>Table 2</label><caption><p id="d1e3490">Quantitative verification metrics are summarized. The <inline-formula><mml:math id="M142" display="inline"><mml:mo>↑</mml:mo></mml:math></inline-formula> indicates that a higher score is better, while the <inline-formula><mml:math id="M143" display="inline"><mml:mo>↓</mml:mo></mml:math></inline-formula> indicates that a lower score is better. The best score amongst models is marked using bold font. ECE scores indicate the expected calibration error, which is an aggregate measure of reliability. AME stands for absolute mean error, and RAPSD rel. MAE score summary values indicate the relative mean absolute error between the power spectral density (PSD) of the observation and predictions. For AME, the sign of the mean error is reported in parentheses. Scores are averaged over the lead times for which they were calculated, except for RAPSD rel. MAE scores, as they are averaged over frequencies. Numerical values in the ECE, ROC AUC, and ETS score names indicate decibels relative to <inline-formula><mml:math id="M144" display="inline"><mml:mi>Z</mml:mi></mml:math></inline-formula> threshold values and in RAPSD lead time in minutes. ECE 45 results are omitted because the results are not comparable due to missing data in some of the bins of the DEUCE reliability diagrams.</p></caption><oasis:table frame="topbot"><?xmltex \begin{scaleboxenv}{.95}[.95]?><oasis:tgroup cols="8">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="left"/>
     <oasis:colspec colnum="3" colname="col3" align="left"/>
     <oasis:colspec colnum="4" colname="col4" align="left"/>
     <oasis:colspec colnum="5" colname="col5" align="left"/>
     <oasis:colspec colnum="6" colname="col6" align="left"/>
     <oasis:colspec colnum="7" colname="col7" align="left"/>
     <oasis:colspec colnum="8" colname="col8" align="left"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry namest="col2" nameend="col4">Probabilistic models </oasis:entry>
         <oasis:entry colname="col5"/>
         <oasis:entry namest="col6" nameend="col8">Deterministic models </oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">DEUCE (ours)</oasis:entry>
         <oasis:entry colname="col3">STEPS</oasis:entry>
         <oasis:entry colname="col4">LINDA-P</oasis:entry>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6">DEUCE mean (ours)</oasis:entry>
         <oasis:entry colname="col7">Extrapolation</oasis:entry>
         <oasis:entry colname="col8">LINDA-D</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">CRPS <inline-formula><mml:math id="M145" display="inline"><mml:mo>↓</mml:mo></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col2">1.29</oasis:entry>
         <oasis:entry colname="col3"><bold>1.27</bold></oasis:entry>
         <oasis:entry colname="col4">1.43</oasis:entry>
         <oasis:entry colname="col5">AME <inline-formula><mml:math id="M146" display="inline"><mml:mo>↓</mml:mo></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6">1.31 (–)</oasis:entry>
         <oasis:entry colname="col7"><bold>0.35</bold> (–)</oasis:entry>
         <oasis:entry colname="col8">0.53 (<inline-formula><mml:math id="M147" display="inline"><mml:mo lspace="0mm">+</mml:mo></mml:math></inline-formula>)</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">ECE 20 (<inline-formula><mml:math id="M148" display="inline"><mml:mrow><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mn mathvariant="normal">3</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>) <inline-formula><mml:math id="M149" display="inline"><mml:mo>↓</mml:mo></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col2"><bold>6.88</bold></oasis:entry>
         <oasis:entry colname="col3">9.45</oasis:entry>
         <oasis:entry colname="col4">13.36</oasis:entry>
         <oasis:entry colname="col5">ETS 20 <inline-formula><mml:math id="M150" display="inline"><mml:mo>↑</mml:mo></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6">0.442</oasis:entry>
         <oasis:entry colname="col7">0.435</oasis:entry>
         <oasis:entry colname="col8"><bold>0.454</bold></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">ECE 25 (<inline-formula><mml:math id="M151" display="inline"><mml:mrow><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mn mathvariant="normal">3</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>) <inline-formula><mml:math id="M152" display="inline"><mml:mo>↓</mml:mo></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col2"><bold>5.36</bold></oasis:entry>
         <oasis:entry colname="col3">6.64</oasis:entry>
         <oasis:entry colname="col4">8.44</oasis:entry>
         <oasis:entry colname="col5">ETS 25 <inline-formula><mml:math id="M153" display="inline"><mml:mo>↑</mml:mo></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6">0.299</oasis:entry>
         <oasis:entry colname="col7">0.341</oasis:entry>
         <oasis:entry colname="col8"><bold>0.371</bold></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">ECE 35 (<inline-formula><mml:math id="M154" display="inline"><mml:mrow><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mn mathvariant="normal">3</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>) <inline-formula><mml:math id="M155" display="inline"><mml:mo>↓</mml:mo></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col2">1.97</oasis:entry>
         <oasis:entry colname="col3"><bold>1.13</bold></oasis:entry>
         <oasis:entry colname="col4">2.04</oasis:entry>
         <oasis:entry colname="col5">ETS 35 <inline-formula><mml:math id="M156" display="inline"><mml:mo>↑</mml:mo></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6">0.047</oasis:entry>
         <oasis:entry colname="col7">0.134</oasis:entry>
         <oasis:entry colname="col8"><bold>0.162</bold></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">ECE 45 (<inline-formula><mml:math id="M157" display="inline"><mml:mrow><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mn mathvariant="normal">4</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>) <inline-formula><mml:math id="M158" display="inline"><mml:mo>↓</mml:mo></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col2">–</oasis:entry>
         <oasis:entry colname="col3">–</oasis:entry>
         <oasis:entry colname="col4">–</oasis:entry>
         <oasis:entry colname="col5">ETS 45  <inline-formula><mml:math id="M159" display="inline"><mml:mo>↑</mml:mo></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6">0.006</oasis:entry>
         <oasis:entry colname="col7">0.049</oasis:entry>
         <oasis:entry colname="col8"><bold>0.056</bold></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">ROC AUC 20 <inline-formula><mml:math id="M160" display="inline"><mml:mo>↑</mml:mo></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col2"><bold>0.968</bold></oasis:entry>
         <oasis:entry colname="col3">0.957</oasis:entry>
         <oasis:entry colname="col4">0.943</oasis:entry>
         <oasis:entry colname="col5">RAPSD rel. MAE 5  <inline-formula><mml:math id="M161" display="inline"><mml:mo>↓</mml:mo></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6">0.55</oasis:entry>
         <oasis:entry colname="col7"><bold>0.08</bold></oasis:entry>
         <oasis:entry colname="col8">0.39</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">ROC AUC 25 <inline-formula><mml:math id="M162" display="inline"><mml:mo>↑</mml:mo></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col2"><bold>0.960</bold></oasis:entry>
         <oasis:entry colname="col3">0.938</oasis:entry>
         <oasis:entry colname="col4">0.926</oasis:entry>
         <oasis:entry colname="col5">RAPSD rel. MAE 15 <inline-formula><mml:math id="M163" display="inline"><mml:mo>↓</mml:mo></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6">0.74</oasis:entry>
         <oasis:entry colname="col7"><bold>0.08</bold></oasis:entry>
         <oasis:entry colname="col8">0.52</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">ROC AUC 35 <inline-formula><mml:math id="M164" display="inline"><mml:mo>↑</mml:mo></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col2"><bold>0.885</bold></oasis:entry>
         <oasis:entry colname="col3">0.784</oasis:entry>
         <oasis:entry colname="col4">0.840</oasis:entry>
         <oasis:entry colname="col5">RAPSD rel. MAE 30 <inline-formula><mml:math id="M165" display="inline"><mml:mo>↓</mml:mo></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6">0.84</oasis:entry>
         <oasis:entry colname="col7"><bold>0.07</bold></oasis:entry>
         <oasis:entry colname="col8">0.58</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">ROC AUC 45 <inline-formula><mml:math id="M166" display="inline"><mml:mo>↑</mml:mo></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col2"><bold>0.706</bold></oasis:entry>
         <oasis:entry colname="col3">0.610</oasis:entry>
         <oasis:entry colname="col4">0.689</oasis:entry>
         <oasis:entry colname="col5">RAPSD rel. MAE 60 <inline-formula><mml:math id="M167" display="inline"><mml:mo>↓</mml:mo></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6">0.90</oasis:entry>
         <oasis:entry colname="col7"><bold>0.11</bold></oasis:entry>
         <oasis:entry colname="col8">0.65</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup><?xmltex \end{scaleboxenv}?></oasis:table><?xmltex \gdef\@currentlabel{2}?></table-wrap>

<sec id="Ch1.S4.SS1">
  <label>4.1</label><title>Case studies</title>
      <p id="d1e4026">The results of the first case study in Figs. <xref ref-type="fig" rid="Ch1.F6"/>, <xref ref-type="fig" rid="Ch1.F7"/>, and <xref ref-type="fig" rid="Ch1.F8"/> suggest that DEUCE ensemble nowcasts are able to give reasonable uncertainty and exceedance probability estimates at multiple thresholds and lead times and that DEUCE nowcasts look similar to those given by STEPS, despite being less grainy. The results of the second case study are detailed in Appendix <xref ref-type="sec" rid="App1.Ch1.S2"/>. Figure <xref ref-type="fig" rid="App1.Ch1.S3.F20"/> shows an example of what individual ensemble members look like at different prediction lead times for the first case. The predictions all start out being quite similar, but they eventually diverge, driven by the increasing predictive uncertainty and different patterns of correlated noise. The ensemble members exhibit variety while preserving a moderate amount of realism that is nevertheless limited by the increasing smoothing of the predictive mean and variance fields with lead time.</p>
<sec id="Ch1.S4.SS1.SSS1">
  <label>4.1.1</label><title>Ensemble mean and breadth</title>
      <p id="d1e4046">Ensemble mean and breadth as units of standard deviation are shown for the first case in Fig. <xref ref-type="fig" rid="Ch1.F6"/>. Here, we can see a trend in all models towards a loss of predicted reflectivity intensity and a disappearance of heavily localized echoes. However, these are in all models compensated by an increase in the spatial extent and the magnitude of the ensemble standard deviation. In LINDA-P, the effect of predicting in rain rate (mm h<inline-formula><mml:math id="M168" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>) is seen as the uncertainty in the cell borders emphasized. LINDA-P also generally exhibits a smaller and more uniform standard deviation than the other models. For a 1 h lead time, DEUCE seems to generally have an ensemble breadth a bit smaller than STEPS but higher than LINDA-P, with the most heterogeneity in the standard deviation values.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F6" specific-use="star"><?xmltex \currentcnt{6}?><?xmltex \def\figurename{Figure}?><label>Figure 6</label><caption><p id="d1e4065">The first case study ensemble means and breadths of DEUCE compared to STEPS and LINDA-P model predictions and observations for multiple lead times. The area covers southern Finland, starting at 15:00:00 UTC on 9 July 2022. The rows represent lead time and columns different instances of observations, model mean, and standard deviations. Missing values are indicated by a dark gray color.</p></caption>
            <?xmltex \igopts{width=497.923228pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/17/3839/2024/gmd-17-3839-2024-f06.png"/>

          </fig>

</sec>
<sec id="Ch1.S4.SS1.SSS2">
  <label>4.1.2</label><title>Reflectivity exceedance probabilities</title>
      <p id="d1e4082">Reflectivity probabilities of exceeding 25 dBZ predicted by the different models for the first case are shown in Fig. <xref ref-type="fig" rid="Ch1.F7"/>. Overall, DEUCE seems to provide balanced exceedance probabilities that are not missing any significant areas even after 1 h but are also not covering excessively large areas. Comparatively, STEPS tends to completely miss some significant portions, such as in the area highlighted in the southwest of Finland at 1 h, and generally seems to predict smaller probabilities for the evolution of smaller cells. LINDA-P, on the other hand, suffers from overconfidence and misplaces the evolution of multiple precipitation areas after 1 h. On a general level for all models compared, the advection field is well captured, while the growth and decay of echoes are often not very effectively forecast. The anisotropic structure of the uncertainty shown through exceedance probabilities is also much better captured by DEUCE and LINDA-P than STEPS. In addition, because it is not based on the extrapolation of radar echoes, there are no “dead zones” filled with NaN (not a number) values (dark gray color), and DEUCE is able to provide nowcasts with varying success in border regions where STEPS and LINDA-P predictions are not necessarily defined.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F7" specific-use="star"><?xmltex \currentcnt{7}?><?xmltex \def\figurename{Figure}?><label>Figure 7</label><caption><p id="d1e4089">The first case study reflectivity exceedance probabilities of 25 dBZ for DEUCE compared to STEPS and LINDA-P model predictions and observations for multiple lead times. The area covers southern Finland, starting at 15:00:00 UTC on 9 July 2022. The rows represent the lead time. In the leftmost column, actual observations are shown in light gray, with echo isotherms corresponding to the threshold marked in black. In other columns, the same isotherms are overlaid on the exceedance probabilities of the models that are depicted in shades of red. Missing values are indicated by a dark gray color. The blue circles labeled “<monospace>1</monospace>” highlight a case of DEUCE model improvement over baselines.</p></caption>
            <?xmltex \igopts{width=412.564961pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/17/3839/2024/gmd-17-3839-2024-f07.png"/>

          </fig>

      <p id="d1e4101">Last, the exceedance probabilities of DEUCE nowcasts for 15, 25, 35, and 45 dBZ reflectivity thresholds for the first case are shown in Fig. <xref ref-type="fig" rid="Ch1.F8"/>. We can see that DEUCE is able to nowcast an exceedance probability at all thresholds (which are indeed all exceeded at some place and point in the observations). Higher thresholds exhibit lower values and some misplacement of exceedance probabilities, as precipitation exceeding those is more difficult to predict and has smaller areas.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F8" specific-use="star"><?xmltex \currentcnt{8}?><?xmltex \def\figurename{Figure}?><label>Figure 8</label><caption><p id="d1e4109">The first case study reflectivity exceedance probabilities of DEUCE compared to observations for multiple lead times and reflectivity thresholds. The area covers southern Finland, starting at 15:00:00 UTC on 9 July 2022. The rows represent lead time. The leftmost column is observations, and the rest are exceedance probabilities at different thresholds. As in Fig. <xref ref-type="fig" rid="Ch1.F7"/>, the observation isotherms corresponding to the threshold in question overlay the exceedance probabilities depicted in shades of red.</p></caption>
            <?xmltex \igopts{width=441.017717pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/17/3839/2024/gmd-17-3839-2024-f08.png"/>

          </fig>

</sec>
</sec>
<sec id="Ch1.S4.SS2">
  <label>4.2</label><title>Analysis of the aleatoric and epistemic uncertainty dichotomy</title>
      <p id="d1e4130">The relative contribution of aleatoric and epistemic uncertainty for the first case study is presented in Fig. <xref ref-type="fig" rid="Ch1.F9"/>. We can see that most of the predictive uncertainty in fact comes from the aleatoric part. Epistemic uncertainty is of much smaller magnitude, and its contribution is further reduced when working in terms of variance in the calculation of predictive uncertainty. We can see that epistemic uncertainty does not extend as much away from the core of the predicted reflectivity as for aleatoric uncertainty, which reflects a small variance in the raw <inline-formula><mml:math id="M169" display="inline"><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover></mml:math></inline-formula> ensemble. This overlap can be seen in probabilistic nowcasts as aleatoric uncertainty overshadowing the contribution of epistemic uncertainty.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F9" specific-use="star"><?xmltex \currentcnt{9}?><?xmltex \def\figurename{Figure}?><label>Figure 9</label><caption><p id="d1e4147">Composition of the predictive uncertainty for the first case study. The area covers southern Finland, starting at 15:00:00 UTC on 9 July 2022. The rows represent lead time and columns observations (ground truth), the mean prediction, aleatoric and epistemic standard deviation components, and the combined predictive standard deviation.</p></caption>
          <?xmltex \igopts{width=412.564961pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/17/3839/2024/gmd-17-3839-2024-f09.png"/>

        </fig>

      <p id="d1e4156">A more detailed view of the contribution of aleatoric and epistemic components is provided in Fig. <xref ref-type="fig" rid="Ch1.F10"/>, with statistics over the whole verification dataset. A histogram of the uncertainties aggregated over all lead times and observed reflectivity values is shown to the left of Fig. <xref ref-type="fig" rid="Ch1.F10"/>. Epistemic uncertainty has a very narrow distribution, mostly between 0–5 dBZ, which means that its average value could not have varied much in different cases, lead times, and observed reflectivity values, pointing to a small model uncertainty response to these factors. Aleatoric uncertainty, on the other hand, has a long-tail distribution centered around 10 dBZ but<?pagebreak page3850?> going up to values over 30 dBZ, which means that we cannot exclude a dependence on these external factors.</p>
      <p id="d1e4164">Such dependencies are confirmed when inspecting the bar plots on the right of Fig. <xref ref-type="fig" rid="Ch1.F10"/>, where mean aleatoric uncertainty shows a clear dependence on prediction lead time and to some degree on observed reflectivity. Aleatoric uncertainty seems to clearly increase with lead time and to first slightly decrease before increasing again in relation to observed reflectivity. One possible explanation for this last observation is that reflectivity values below 20 dBZ often correspond to the edges of precipitation cells, which are difficult to predict, and that reflectivity values over 35 dBZ often correspond to heavy precipitation with a short lifetime and thus bad predictability. In between those, there are more predictable<?pagebreak page3851?> precipitation patterns, such as the interior of stratiform precipitation cells. Epistemic uncertainty, on the other hand, does not seem to show any particular dependence on prediction lead time, which might have something to do with the fact that the model predicts all lead times at once, making it possibly more difficult for the predictions to vary, depending on the lead time. There is, on the other hand, a slight increase in the epistemic uncertainty with observed reflectivity, which might be an accurate reflection of the relatively smaller amount of training data available for high observed reflectivity values.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F10" specific-use="star"><?xmltex \currentcnt{10}?><?xmltex \def\figurename{Figure}?><label>Figure 10</label><caption><p id="d1e4171">Visualization of the statistics on the composition of predictive uncertainty over the verification dataset. On the left, a histogram of aleatoric and epistemic standard deviation (SD) aggregated over all lead times and observed reflectivity values is shown. On the right, we arrange the same data into bar plots to show the relationship between the type of the uncertainty SD with prediction lead time <bold>(b)</bold> and observed ground truth radar reflectivity <bold>(c)</bold>.</p></caption>
          <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/17/3839/2024/gmd-17-3839-2024-f10.png"/>

        </fig>

</sec>
<sec id="Ch1.S4.SS3">
  <label>4.3</label><title>Probabilistic skill verification</title>
      <p id="d1e4194">The reliability diagrams and sharpness histograms for probabilistic nowcasts are depicted in Fig. <xref ref-type="fig" rid="Ch1.F11"/>. It can first be noted that, in general, DEUCE nowcasts are very close to the dashed black line, indicating a perfectly reliable forecast. Sometimes, this is similar to baseline models, but in some cases, such as a long lead time and a high threshold, DEUCE is closer to the diagonal than baselines. This is, however, not reflected in the ECE scores at 35 dBZ shown in Table <xref ref-type="table" rid="Ch1.T2"/>, as smaller forecast probabilities are weighted much higher due to their sample count here, making STEPS the most reliable model at 35 dBZ when using this metric. An important pattern is that compared to baseline models, DEUCE is prone to slight under-forecasting of the exceedance probabilities. This is particularly the case for a short lead time (5 min), where the effect is the most pronounced. As lead times grow longer and thresholds get higher, nowcasting gets harder, and there is an overall tendency in all models, but particularly LINDA-P, to over-forecast threshold exceedance.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F11" specific-use="star"><?xmltex \currentcnt{11}?><?xmltex \def\figurename{Figure}?><label>Figure 11</label><caption><p id="d1e4203">The reliability diagrams and sharpness histograms for DEUCE (yellow), STEPS (blue), and LINDA-P (red) model nowcasts at exceedance probability thresholds of 20, 25, 35, and 45 dBZ and at lead times of 5, 15, 30, and 60 min. Rows indicate the lead time and columns the exceedance probability threshold. The diagonal dashed black lines indicate perfect reliability.</p></caption>
          <?xmltex \igopts{width=497.923228pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/17/3839/2024/gmd-17-3839-2024-f11.png"/>

        </fig>

      <?pagebreak page3852?><p id="d1e4212">From the sharpness histogram, it is seen that the distribution of forecast probabilities is more or less uniform at low thresholds but more biased towards small exceedance probabilities at higher thresholds. These higher thresholds are where the difference between DEUCE and baselines is visible, as there is a considerably lower number of cases of high forecast probability in DEUCE than in baselines.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F12" specific-use="star"><?xmltex \currentcnt{12}?><?xmltex \def\figurename{Figure}?><label>Figure 12</label><caption><p id="d1e4218">Rank histograms of ensemble nowcasts, including DEUCE, STEPS, and LINDA-P, at lead times of 5, 15, 30, and 60 min over the verification set.</p></caption>
          <?xmltex \igopts{width=483.69685pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/17/3839/2024/gmd-17-3839-2024-f12.png"/>

        </fig>

      <p id="d1e4227">The rank histogram of nowcasts is shown in Fig. <xref ref-type="fig" rid="Ch1.F12"/>. It is apparent here that DEUCE is constantly slightly biased towards predicting reflectivity values that are too low and that the spread is large at short lead times but less significant later on. STEPS exhibits a very balanced flat histogram, but LINDA, on the other hand, has a U-shaped histogram characteristic of an ensemble breadth that is too small in general.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F13" specific-use="star"><?xmltex \currentcnt{13}?><?xmltex \def\figurename{Figure}?><label>Figure 13</label><caption><p id="d1e4234">The ROC area under the curve (AUC) values at lead times up to 60 min for DEUCE (yellow), STEPS (blue), and LINDA-P (red) model nowcasts at exceedance probability thresholds of 20, 25, 35, and 45 dBZ.</p></caption>
          <?xmltex \igopts{width=483.69685pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/17/3839/2024/gmd-17-3839-2024-f13.png"/>

        </fig>

      <p id="d1e4243">The results for the ROC area under the curve probabilistic nowcast metric are shown in Fig. <xref ref-type="fig" rid="Ch1.F13"/>. In this benchmark, DEUCE achieves the best results at all thresholds. We can notice that STEPS has good discriminative power at low thresholds but that it does not scale well to higher ones and that LINDA-P is not competitive at lower thresholds but excels as the threshold grows. Nevertheless, DEUCE manages to perform better than both in their skillful areas.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F14"><?xmltex \currentcnt{14}?><?xmltex \def\figurename{Figure}?><label>Figure 14</label><caption><p id="d1e4250">The CRPS score of DEUCE (yellow) compared to the ensemble baselines of STEPS (blue) and LINDA-P (red) is shown on the left. The mean error (ME) score for non-augmented ensemble mean predictions of DEUCE (yellow) compared to deterministic baseline extrapolation (green) and LINDA-D (red) nowcasts is shown on the right.</p></caption>
          <?xmltex \igopts{width=241.848425pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/17/3839/2024/gmd-17-3839-2024-f14.png"/>

        </fig>

      <p id="d1e4260">Last, the CRPS verification metric is depicted in Fig. <xref ref-type="fig" rid="Ch1.F14"/>. It can be seen that the lowest and best score is achieved by STEPS at all lead times. DEUCE comes second, slightly above STEPS, and LINDA-P lags far behind. Overall, with CRPS, it can be seen that DEUCE achieves adequate results of the order of baseline models.</p>
      <p id="d1e4265">From the quantitative probabilistic verification, it can be summarized that DEUCE achieves satisfactory and well-rounded performance. The model does not significantly lack in any category in particular and offers a good trade-off between forecast reliability and discriminatory power.</p>
</sec>
<sec id="Ch1.S4.SS4">
  <label>4.4</label><title>Deterministic skill verification</title>
      <p id="d1e4276">Here, we analyze the results of the comparison of the deterministic nowcast skill between DEUCE non-augmented mean predictions <inline-formula><mml:math id="M170" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mtext>mean</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> and baseline predictions. First, a depiction of the mean nowcasting error (ME) until a 60 min lead time is presented in Fig. <xref ref-type="fig" rid="Ch1.F14"/>. While extrapolation nowcasts have, on average, a ME slightly below zero, DEUCE is more strongly negatively biased, while LINDA-D is strongly positively biased.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F15" specific-use="star"><?xmltex \currentcnt{15}?><?xmltex \def\figurename{Figure}?><label>Figure 15</label><caption><p id="d1e4297">Equitable threat scores (ETSs) as a function of lead time for non-augmented DEUCE ensemble means (yellow) compared to those of extrapolation (green) and LINDA-D (red) deterministic baseline models for reflectivity thresholds of 20, 25, 35, and 45 dBZ at lead times up to 60 min. DEUCE ensemble means perform competitively for predicting reflectivity exceeding 20 dBZ but see their relative performance drop at higher reflectivity thresholds.</p></caption>
          <?xmltex \igopts{width=483.69685pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/17/3839/2024/gmd-17-3839-2024-f15.png"/>

        </fig>

      <p id="d1e4306"><?xmltex \hack{\newpage}?>Furthermore, the equitable threat score (ETS) results for reflectivity thresholds of 20, 25, and 35 dBZ are shown in Fig. <xref ref-type="fig" rid="Ch1.F15"/>. The ETS score of DEUCE is competitive for the 20 dBZ threshold, but at 25 dBZ, its progression at lead times longer than 30 min is already worse than the baseline. At 35 and 45 dBZ, DEUCE already performs worse than the baseline at any lead time examined. The reason for this weakness is the compound effect of intrinsic CNN prediction smoothing and the averaging of ensemble members. This smoothing effect is also visible in the radially averaged power spectral density (RAPSD) results for nowcasts presented in Fig. <xref ref-type="fig" rid="Ch1.F16"/>. Average RAPSD is computed for nowcasts a<?pagebreak page3854?>t lead times of 5, 15, 30, and 60 min. It can clearly be seen that, compared to baselines, the fields predicted by DEUCE lose more power at small spatial scales and that this effect is heavily amplified at longer lead times, which again illustrates the above compound effect. However, this effect seems to be damped by augmenting the mean prediction with uncertainty-weighted correlated noise following the structure of the input field, especially at longer lead times.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F16" specific-use="star"><?xmltex \currentcnt{16}?><?xmltex \def\figurename{Figure}?><label>Figure 16</label><caption><p id="d1e4317">Radially averaged power spectral density (RAPSD) for non-augmented DEUCE ensemble means (solid yellow) and individual augmented DEUCE ensemble members (dashed yellow) compared to those of the extrapolation (green) and LINDA-D (red) deterministic baseline models at lead times of 5, 15, 30, and 60 min. The observation RAPSD is shown as a dashed black line.</p></caption>
          <?xmltex \igopts{width=483.69685pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/17/3839/2024/gmd-17-3839-2024-f16.png"/>

        </fig>

</sec>
</sec>
<sec id="Ch1.S5">
  <label>5</label><title>Discussion</title>
      <p id="d1e4336">DEUCE probabilistic ensemble nowcasts proved to be both relatively reliable and skillful compared to STEPS and LINDA. In comparison, STEPS was often reliable but struggled to capture small-scale high reflectivity, and LINDA was better at this task but sometimes suffered from overconfidence. DEUCE seems to offer a good compromise as it does not suffer too much from either of those two defects. Nevertheless when analyzing the reliability diagram, it is apparent that DEUCE is under-confident, especially at short lead times. Further inspection of rank histograms (Fig. <xref ref-type="fig" rid="Ch1.F12"/>)<?pagebreak page3855?> – showing the distribution of the rank of observations among ensemble members – indicated that this under-confidence is expressed by (1) an ensemble spread that is too large and (2) a slight bias towards ensembles producing estimates that are too weak, which was also visible in the ensemble mean ME score. The ensemble spread that is too large may be a byproduct of attempting to predict the ensemble mean and aleatoric variances all at once, with the prior placed on weights that place a limit on the complexity of the model and in effect privileging the learning of longer lead times where the average errors are of a bigger magnitude.</p>
      <p id="d1e4341">We can also observe that most of the uncertainty is of an aleatoric nature and that the contribution of epistemic variance is universally low. In addition to having trained with a large amount of data, low epistemic variance is probably related to the variational inference mechanism, as the combination of Bayes by Backprop and Flipout re-parameterization has been shown by <xref ref-type="bibr" rid="bib1.bibx50" id="text.56"/> to yield epistemic uncertainty estimates that are too small when compared to Monte Carlo Dropout, deep ensembles, and Markov Chain Monte Carlo. Another factor that might have played a role in this is again predicting all lead times at once because adopting the iterative approach of RainNet <xref ref-type="bibr" rid="bib1.bibx4" id="paren.57"/> would have propagated previous epistemic uncertainties to subsequent lead times, possibly balancing out the contributions, and also allowing predictions to be made for an arbitrary number of time steps. On the other hand, abandoning the recursive prediction scheme of RainNet significantly reduces the time complexity of computing <inline-formula><mml:math id="M171" display="inline"><mml:mrow><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mo>,</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> from <inline-formula><mml:math id="M172" display="inline"><mml:mrow><mml:mi mathvariant="script">O</mml:mi><mml:mo>(</mml:mo><mml:mi>N</mml:mi><mml:mi>L</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> to <inline-formula><mml:math id="M173" display="inline"><mml:mrow><mml:mi mathvariant="script">O</mml:mi><mml:mo>(</mml:mo><mml:mi>N</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, where <inline-formula><mml:math id="M174" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula> is the sample size, and <inline-formula><mml:math id="M175" display="inline"><mml:mi>L</mml:mi></mml:math></inline-formula> is the number of prediction lead times.</p>
      <p id="d1e4413">Contrary to the preliminary iteration of the model focusing only on modeling epistemic uncertainty <xref ref-type="bibr" rid="bib1.bibx15" id="paren.58"/>, the current model better captures the increased spread of the predictive distribution with lead time through the aleatoric component. An alternative model – without a separate decoder branch for <inline-formula><mml:math id="M176" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> and only allocating a separate output channel to it – was not successful because the <inline-formula><mml:math id="M177" display="inline"><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover></mml:math></inline-formula> and <inline-formula><mml:math id="M178" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> that it learned were highly correlated, more blurry, and lacked expressivity. It is for this reason that the two-branch version was adopted.</p>
      <p id="d1e4451">In DEUCE, small-scale variability is steadily lost with increasing lead time, which is especially noticeable for lead times over 30 min. This is a problem for the production of realistic nowcasts, as the loss of small-scale variability is synonymous with the loss of information, limiting the expressivity of the model. However, the smoothing of radar image predictions can be justified in the case of a probabilistic model, assuming that the breadth of the ensemble is preserved despite the loss of high frequencies. This is because information will invariably be lost with time as we attempt to predict the evolution of a chaotic system through imperfect measurements. In the present case of DEUCE, the predictive distribution is modeled explicitly, giving us the ability to arbitrarily sample from it, so this loss of high-frequency components is not a major issue. For the preliminary version of the model <xref ref-type="bibr" rid="bib1.bibx15" id="paren.59"/>, however, losing small-scale details had adverse effects. This was because the model was trained with  homoscedastic (fixed as a hyperparameter) aleatoric uncertainty modeling, which was not taken into account when making predictions, resulting in an ensemble spread of smooth predictions that is too small and leading to vastly underestimated exceedance probabilities. The<?pagebreak page3856?> spatially correlated noise scheme for sampling the predictive uncertainty may help integrate DEUCE into applications where physically plausible ensemble members are necessary. Still, the lack of temporal correlation modeling inside the post-processed ensemble members and the smoothing of the predictive means and variances themselves limit realism, pointing to the limits of the taken approach. One grounded method for resolving this problem is to either implicitly or explicitly constrain the predictions to replicate the power spectrum of observations. Generative models such as GANs (Generative Adversarial Networks) fall into the category of implicit constraining. For example, the DGMR model by <xref ref-type="bibr" rid="bib1.bibx37" id="text.60"/> learns to model realistic spatial and temporal correlations by adversarially training the generator with two discriminators designed to discern those aspects.</p>
      <p id="d1e4461">We see that the issues of under-forecasting at short lead times and lacking small-scale variability could be linked and related to the model training settling for underperforming local optima. Avenues to mitigate this include longer training and different training strategies, using a bigger and more varied dataset, adapting the loss function, and improving the model itself. DEUCE was trained with only 29 epochs, which is a relatively small number. We attempted to train with a higher number of epochs, but this did not result in consistent improvement in model validation performance metrics. However, this behavior might have been related to our choice of learning rate scheduling. On the other hand, it might also be worthwhile to attempt to implement a curriculum learning strategy where easier to learn and earlier lead times would be learned first before allowing the network to learn to predict longer lead times. The dataset size itself may be increased by covering a variety of bounding boxes in the composite area and by including precipitation events from outside the summer period. The likelihood part of the loss function may be adapted so that higher weight is given to higher-reflectivity pixels. We, nevertheless, found this tricky to get right, as the experiments that we performed with weighting proportional to the inverse of the density of the reflectivity in the dataset distribution failed to produce reasonable nowcasts. From the perspective of the model, under-forecasting and the lack of small-scale details could be reduced by including spatial and channel (temporal for us) attention mechanisms, such as the Convolutional Block Attention Module (CBAM) <xref ref-type="bibr" rid="bib1.bibx53" id="paren.61"/>, which has been applied to improve (deterministic) precipitation nowcasting performance <xref ref-type="bibr" rid="bib1.bibx48" id="paren.62"/>. CBAM might enable, for example, sharper forecasts with smaller aleatoric uncertainties at short lead times without affecting the reliability at longer lead times.</p>
      <?pagebreak page3857?><p id="d1e4470">It is important to note that all models, DEUCE included, are generally not able to predict convective initiation, which is a notoriously hard problem to solve <xref ref-type="bibr" rid="bib1.bibx33" id="paren.63"/>. This is clearly illustrated with the new echoes appearing in the northwestern (continuously) and southern (between 30 and 60 min) parts of the second case (Fig. <xref ref-type="fig" rid="App1.Ch1.S2.F18"/>) for which very low exceedance probabilities are predicted. Adding polarimetric or vertical profile information as additional input channels and adopting a model less susceptible to blurring might improve this aspect of predictions.</p>
      <p id="d1e4478">With regard to the model development in general, some degree of hyperparameter optimization was performed. Those hyperparameters related to the functional model and the optimizer are mainly inherited from RainNet <xref ref-type="bibr" rid="bib1.bibx4" id="paren.64"/>, and those related to variational inference mostly originate from the preliminary version of the model <xref ref-type="bibr" rid="bib1.bibx15" id="paren.65"/>. There, the VI-related parameters specifically demanded non-trivial tuning for model convergence and acceptable result production, which might limit the immediate applicability of the model in its default state. Also, the local optimality of the current hyperparameters is not assured. Despite this, VI and epistemic uncertainty are not decisive factors in the model performance, and swapping out those components is a potential way forward. Moreover, 60 min (12 frames) of input data and the same length for predictions were picked without optimization in an attempt to preserve some symmetry between the network inputs and outputs. Although there is no consensus yet on how many frames are needed, as few as four input frames have be enough to saturate model performance in some conditions <xref ref-type="bibr" rid="bib1.bibx37" id="paren.66"/>, so tuning the ratio of input to output frames could be a viable thing to try.</p>
      <p id="d1e4490">Regarding the verification process, it is a pertinent question to ask whether the used baseline models were sufficient to validate the performance of DEUCE. In particular, the lack of deep-learning ensemble baselines is one weakness of the performed verification. It would have been particularly interesting to use the Deep Generative Model of Radar (DGMR) by <xref ref-type="bibr" rid="bib1.bibx37" id="text.67"/> as a baseline, as it represents the current state-of-the-art in deep-learning-based precipitation nowcasting and is capable of producing ensemble nowcasts. Unfortunately, we were not able to successfully train DGMR on our dataset using the resources that we had allocated for the task. Other models of interest that were not included in the verification are MetNet by <xref ref-type="bibr" rid="bib1.bibx47" id="text.68"/> and its successor MetNet-2 by <xref ref-type="bibr" rid="bib1.bibx11" id="text.69"/>, which use, e.g., orographic and satellite data in addition to radar data. We hope that further work will make the comparison of DEUCE probabilistic nowcasting performance to other deep-learning-based models possible.</p>
      <p id="d1e4502">One last point of concern regards the validity of the verification metrics used. The potential issues here mostly relate to the summarizing quantitative metrics of Table <xref ref-type="table" rid="Ch1.T2"/>. First, the relative RAPSD MAE metric for measuring the power spectrum fidelity of predictions uses a tighter sampling of points towards wavelengths representing small spatial scales, which biases it to give a higher weight to those scales. Although we are indeed mostly interested in small-scale variations, this property means that even big discrepancies in the power of large spatial scales will be under-represented. Next, the ECE metric used to summarize the reliability of ensemble models is very sensitive to variations of the order of magnitude of the number of samples per bin. This behavior is significant especially at higher exceedance thresholds, where almost all prediction probabilities are concentrated in the smallest probability bin, giving almost no weight to even the mildly successful nowcasting of rare but significant events of high heavy precipitation probability. This means that ECE does not necessarily provide a complete assessment of model reliability in the context of probabilistic precipitation nowcasting.</p>
</sec>
<sec id="Ch1.S6" sec-type="conclusions">
  <label>6</label><title>Conclusions</title>
      <p id="d1e4515">We developed a probabilistic precipitation nowcasting model named DEUCE, based on a Bayesian neural network with variational inference and featuring the combination of epistemic and aleatoric uncertainty estimates in an attempt to yield reliable yet powerful probabilistic predictions. The model succeeded at this primary task, performing competitively against the baseline STEPS and LINDA-P models that were judged using qualitative and quantitative evaluation.</p>
      <p id="d1e4518">It was found that DEUCE had issues with the representation of epistemic uncertainty, leading to most of the uncertainty appearing as aleatoric uncertainty, maybe due to the variational inference used. The aleatoric uncertainty exhibited a clear dependence on lead time and corresponding observed reflectivity, which are factors heavily influencing the predictability. The epistemic uncertainty, on the other hand, showed little dependence on these factors, with the exception of a slight increase with observed reflectivity, which might reflect the distribution of the training data. Based on this, aleatoric and epistemic uncertainties do indeed seem to capture complementary features of the predictive uncertainty.</p>
      <?pagebreak page3858?><p id="d1e4521">Deterministically, the ensemble means were found to perform worse compared to extrapolation and LINDA-D baselines, showing that the model in its current state is not useful in the deterministic case due to the excessive smoothing of predictions. This smoothing may also have affected the uncertainty composition, such that assuming the predictive mean to be fixed, i.e., with no improvement in skill, means that sharper reflectivity predictions would increase the epistemic uncertainty. As for the aleatoric uncertainty, the variation between individual draws would increase, but their average would not necessarily increase, except for cases where smoothing is the mechanism that hinders the prediction of large enough reflectivity values. In those cases, the average aleatoric uncertainty might decrease. <?xmltex \hack{\newpage}?> Looking into future research directions, DEUCE has a number of different facets upon which its performance could be improved. First, the underlying U-Net could potentially be replaced by a more powerful architecture capable of modeling explicit temporal dependencies. The spatiotemporal extent could be enlarged, and additional orographic, polarimetric, or satellite input channels could improve parts of the nowcasts. It is possible to additionally try to leverage other patterns for increasing predictability, such as operating in Lagrangian coordinates, as shown by <xref ref-type="bibr" rid="bib1.bibx38" id="text.70"/>, to increase prediction performance. From a probabilistic aspect, certain alternative inference methods, such as radial Bayesian neural networks <xref ref-type="bibr" rid="bib1.bibx12" id="paren.71"/> or deep ensembles, look promising as a potential way to ease the training and improve the representation of epistemic uncertainty. We could also think of directly appending the post-processing sampling with spatially correlated noise to the neural network or even learning context-dependent spatiotemporal correlation structures. The sampled outputs could then be, e.g., fed to a GAN-like discriminator module which would drive the processed outputs to be more realistic while retaining the uncertainty decomposition.</p>
      <p id="d1e4532">Regardless of its shortcomings, DEUCE is a first step in ensemble-based probabilistic precipitation nowcasting using Bayesian neural networks. The concurrent modeling of aleatoric and epistemic uncertainties has the potential to be useful for operational forecasters, and the model in its current state forms a strong yet relatively lightweight baseline for future developments in deep-learning-based probabilistic precipitation nowcasting.</p>
</sec>

      
      </body>
    <back><app-group>

<app id="App1.Ch1.S1">
  <?xmltex \currentcnt{A}?><label>Appendix A</label><title>Additional technical details</title>
<sec id="App1.Ch1.S1.SS1">
  <label>A1</label><title>Ground precipitation estimates from reflectivity</title>
      <p id="d1e4553">The formula <inline-formula><mml:math id="M179" display="inline"><mml:mrow><mml:mi>R</mml:mi><mml:mo>=</mml:mo><mml:mo>(</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mi>z</mml:mi><mml:mo>/</mml:mo><mml:mn mathvariant="normal">10</mml:mn></mml:mrow></mml:msup><mml:mo>/</mml:mo><mml:mn mathvariant="normal">223</mml:mn><mml:msup><mml:mo>)</mml:mo><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">1.53</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> was used in cases where an estimate of ground precipitation corresponding to lowest-level radar reflectivity composites was needed. Here <inline-formula><mml:math id="M180" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula> denotes precipitation estimates (in mm h<inline-formula><mml:math id="M181" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>), and <inline-formula><mml:math id="M182" display="inline"><mml:mi>z</mml:mi></mml:math></inline-formula> denotes radar reflectivity (in dBZ). The parameters of the <inline-formula><mml:math id="M183" display="inline"><mml:mi>Z</mml:mi></mml:math></inline-formula>–<inline-formula><mml:math id="M184" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula> relationship employed in the formula come from the work of <xref ref-type="bibr" rid="bib1.bibx27" id="text.72"/> and aim to estimate the amount of rainfall corresponding to radar reflectivity measurements from the Finnish Meteorological Institute polarimetric C-band radars in Finland.</p>
</sec>
<sec id="App1.Ch1.S1.SS2">
  <label>A2</label><title>Baseline models</title>
      <p id="d1e4644">There are two deterministic baseline models: a simple extrapolation nowcast and the deterministic variant of LINDA (LINDA-D). The extrapolation nowcast extrapolates the last input reflectivity field along a motion field calculated from the last four elements of the input time series. In the extrapolation nowcast and all other baseline methods, we use the dense Lucas–Kanade optical flow method with its default pySTEPS parameters for the computation of the motion field. In addition, all baseline nowcasting methods use the semi-Lagrangian integration scheme from pySTEPS for performing the extrapolation, with cubic interpolation and other parameters left to their default values.</p>
      <p id="d1e4647">LINDA, a more advanced extrapolation-based method capable of predicting high-intensity rainfall more accurately, serves as a natural benchmark in the deterministic and probabilistic cases for the ability of the model to capture convective rainfall evolution. LINDA predictions are made using reflectivity fields converted to rain rate, using the method described in Sect. <xref ref-type="sec" rid="App1.Ch1.S1.SS1"/>, as it is required for the model to work. LINDA models here use the last three input rain rate fields as input, in addition to the motion field. They do not use feature detection in order to reduce the prediction computation time over the verification set to more practical durations. The ensemble-producing version of LINDA, LINDA-P, is used as a probabilistic baseline model; while LINDA-D deterministic nowcasts do not add any perturbations, LINDA-P does add them, as well as velocity perturbations from <xref ref-type="bibr" rid="bib1.bibx8" id="text.73"/>, with <monospace>lucaskanade/fmi+mch</monospace> parameters <xref ref-type="bibr" rid="bib1.bibx34" id="paren.74"/>. Other parameters are set to be data-specific or to be their default values.</p>
      <p id="d1e4661">The STEPS model is used in addition to LINDA-P as a probabilistic baseline. While being a bit older and having lower discriminative power, it is a popular method for making reliable probabilistic precipitation nowcasts to this day. STEPS is applied to decibels relative to <inline-formula><mml:math id="M185" display="inline"><mml:mi>Z</mml:mi></mml:math></inline-formula> (dBZ) reflectivity fields and also takes in the last three input images, in addition to the motion field. Field perturbations and motion field perturbations are applied with the same parameters as with LINDA-P. Six cascade levels are used for the cascade decomposition, and the precipitation threshold of 8 dBZ is given as the lowest observable precipitation intensity.</p>
</sec>
<sec id="App1.Ch1.S1.SS3">
  <label>A3</label><title>Details on probabilistic verification metrics</title>
<sec id="App1.Ch1.S1.SS3.SSS1">
  <label>A3.1</label><title>Continuous ranked probability score (CRPS)</title>
      <p id="d1e4686">The CRPS generalizes the MAE to probability distributions by calculating the sum of the difference between the cumulative density function (CDF) of the nowcast and the empirical CDF of observations. It is defined as
              <disp-formula id="App1.Ch1.S1.E8" content-type="numbered"><label>A1</label><mml:math id="M186" display="block"><mml:mrow><mml:mi mathvariant="normal">CRPS</mml:mi><mml:mo>(</mml:mo><mml:mi>F</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:munderover><mml:mo movablelimits="false">∫</mml:mo><mml:mrow><mml:mo>-</mml:mo><mml:mi mathvariant="normal">∞</mml:mi></mml:mrow><mml:mi mathvariant="normal">∞</mml:mi></mml:munderover><mml:mo>(</mml:mo><mml:mi>F</mml:mi><mml:mo>(</mml:mo><mml:mover accent="true"><mml:mi>y</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mo>)</mml:mo><mml:mo>-</mml:mo><mml:mn mathvariant="double-struck">1</mml:mn><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mo>≥</mml:mo><mml:mover accent="true"><mml:mi>y</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mo>)</mml:mo><mml:msup><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mi mathvariant="normal">d</mml:mi><mml:mover accent="true"><mml:mi>y</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
            where <inline-formula><mml:math id="M187" display="inline"><mml:mover accent="true"><mml:mi>y</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover></mml:math></inline-formula> denote possible forecast values, <inline-formula><mml:math id="M188" display="inline"><mml:mrow><mml:mi>F</mml:mi><mml:mo>(</mml:mo><mml:mover accent="true"><mml:mi>y</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> denote the forecast CDF, and <inline-formula><mml:math id="M189" display="inline"><mml:mrow><mml:mn mathvariant="double-struck">1</mml:mn><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mo>≥</mml:mo><mml:mover accent="true"><mml:mi>y</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mo>)</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> denote the empirical CDF of observations <inline-formula><mml:math id="M190" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula>.</p>
</sec>
<sec id="App1.Ch1.S1.SS3.SSS2">
  <label>A3.2</label><title>Receiver operating characteristic (ROC) curve</title>
      <p id="d1e4830">The receiver operating characteristic (ROC) curve <xref ref-type="bibr" rid="bib1.bibx30 bib1.bibx52" id="paren.75"/> quantifies the discriminative power of an<?pagebreak page3859?> ensemble for predicting over a certain threshold by keeping track of the false alarm rate (FAR), i.e.,
              <disp-formula id="App1.Ch1.S1.E9" content-type="numbered"><label>A2</label><mml:math id="M191" display="block"><mml:mrow><mml:mtext>FAR</mml:mtext><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mtext>FP</mml:mtext><mml:mrow><mml:mtext>FP</mml:mtext><mml:mo>+</mml:mo><mml:mtext>TN</mml:mtext></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
            where the rate of false positives is indicated by FP, and the rate of true negatives is indicated by TN, and both are compared to the probability of detection (POD), i.e.,
              <disp-formula id="App1.Ch1.S1.E10" content-type="numbered"><label>A3</label><mml:math id="M192" display="block"><mml:mrow><mml:mtext>POD</mml:mtext><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mtext>TP</mml:mtext><mml:mrow><mml:mtext>TP</mml:mtext><mml:mo>+</mml:mo><mml:mtext>FN</mml:mtext></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
            where TP is the rate of true positives, and FN is the rate of false negatives. POD is regularly binned, and FAR is averaged over those bins, making a curve, the area under which (AUC) summarizes the overall discriminative power of the nowcasting method. A ROC AUC of 0.5 indicates zero skill, whereas a value of 1.0 indicates a perfect forecast. For ROC curve computations, we use 10 bins.</p>
</sec>
<sec id="App1.Ch1.S1.SS3.SSS3">
  <label>A3.3</label><title>Reliability diagram</title>
      <p id="d1e4890">The reliability diagram <xref ref-type="bibr" rid="bib1.bibx52" id="paren.76"/> measures the reliability of the forecast by presenting the observed relative frequencies of dBZ threshold exceedance events against the forecast probability of those events. Having these two values strongly correlate makes the forecast reliable. Reliability diagrams are built by dividing the forecast probabilities into bins (we choose 10) and incrementing them with associated binary indicators of whether the event happened. Sharpness histograms represent the number of events recorded in each forecast probability bin. They measure the relative “decisiveness” of the forecast, where a high decisiveness is associated with a convex histogram shape. A low decisiveness, on the other hand, can be discerned from a more uniform, or in the extreme case, a concave histogram shape.</p>
</sec>
<sec id="App1.Ch1.S1.SS3.SSS4">
  <label>A3.4</label><title>Expected calibration error (ECE)</title>
      <p id="d1e4904">The expected calibration error (ECE) <xref ref-type="bibr" rid="bib1.bibx31" id="paren.77"/> quantitatively summarizes the reliability of a model indicated by a reliability diagram. It is defined as
              <disp-formula id="App1.Ch1.S1.E11" content-type="numbered"><label>A4</label><mml:math id="M193" display="block"><mml:mrow><mml:mtext>ECE</mml:mtext><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mi>N</mml:mi></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>b</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>B</mml:mi></mml:munderover><mml:msub><mml:mi>n</mml:mi><mml:mi mathvariant="normal">b</mml:mi></mml:msub><mml:mo>∣</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi mathvariant="normal">b</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mi>o</mml:mi><mml:mi mathvariant="normal">b</mml:mi></mml:msub><mml:mo>∣</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
            with a total of <inline-formula><mml:math id="M194" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula> pairs of forecast probability and observation, forecast probabilities divided into <inline-formula><mml:math id="M195" display="inline"><mml:mi>B</mml:mi></mml:math></inline-formula> bins, with <inline-formula><mml:math id="M196" display="inline"><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi mathvariant="normal">b</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> observations per bin, <inline-formula><mml:math id="M197" display="inline"><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mi mathvariant="normal">b</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the mean bin forecast probability, and <inline-formula><mml:math id="M198" display="inline"><mml:mrow><mml:msub><mml:mi>o</mml:mi><mml:mi mathvariant="normal">b</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the corresponding observation frequency in the bin. ECE corresponds to the MAE of the reliability diagram to the diagonal and is weighted by the number of observations per bin.</p>
</sec>
<sec id="App1.Ch1.S1.SS3.SSS5">
  <label>A3.5</label><title>Rank histogram</title>
      <p id="d1e5018">Rank histograms <xref ref-type="bibr" rid="bib1.bibx52" id="paren.78"/> measure the bias and spread of ensemble nowcasts. They present a histogram of the rank of the true observed echo reflectivity among all ensemble members, where a convex histogram indicates a small spread and a concave histogram indicates a small spread. On the other hand, a higher frequency of low ranks for observations indicates a positive bias of predictions, and a higher frequency of high ranks for observations indicates a negative bias of predictions.</p>
</sec>
</sec>
<sec id="App1.Ch1.S1.SS4">
  <label>A4</label><title>Details on deterministic verification metrics</title>
<sec id="App1.Ch1.S1.SS4.SSS1">
  <label>A4.1</label><title>Mean error (ME)</title>
      <p id="d1e5040">The mean error (ME) <xref ref-type="bibr" rid="bib1.bibx52" id="paren.79"/> measures the bias of deterministic predictions. It is defined as
              <disp-formula id="App1.Ch1.S1.E12" content-type="numbered"><label>A5</label><mml:math id="M199" display="block"><mml:mrow><mml:mi mathvariant="normal">ME</mml:mi><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mi>P</mml:mi></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>P</mml:mi></mml:munderover><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>p</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mi>p</mml:mi></mml:msub></mml:mrow></mml:math></disp-formula>
            for images, or time series of them, with <inline-formula><mml:math id="M200" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula> pixels. This metric tells us about the mean bias of nowcasts produced. The absolute value of ME is used to give a quantitative summary of the bias of predictions made.</p>
</sec>
<sec id="App1.Ch1.S1.SS4.SSS2">
  <label>A4.2</label><title>Equitable threat score (ETS)</title>
      <p id="d1e5104">The equitable threat score (ETS) <xref ref-type="bibr" rid="bib1.bibx20 bib1.bibx52" id="paren.80"/> is an extension of the threat score, also known as the critical success index <xref ref-type="bibr" rid="bib1.bibx40" id="paren.81"/>. ETS aims to provide an estimate of deterministic skill in forecasting precipitation above a certain intensity threshold. This extension takes into account the effect of randomly occurring true positives. ETS is defined as
              <disp-formula id="App1.Ch1.S1.E13" content-type="numbered"><label>A6</label><mml:math id="M201" display="block"><mml:mtable class="split" rowspacing="0.2ex" displaystyle="true" columnalign="right"><mml:mtr><mml:mtd><mml:mrow><mml:mtext>ETS</mml:mtext><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mtext>TP</mml:mtext><mml:mo>-</mml:mo><mml:mtext>rnd</mml:mtext></mml:mrow><mml:mrow><mml:mtext>TP</mml:mtext><mml:mo>+</mml:mo><mml:mtext>FN</mml:mtext><mml:mo>+</mml:mo><mml:mtext>FP</mml:mtext><mml:mo>-</mml:mo><mml:mtext>rnd</mml:mtext></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mtext>where rnd</mml:mtext><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mo>(</mml:mo><mml:mtext>TP</mml:mtext><mml:mo>+</mml:mo><mml:mtext>FN</mml:mtext><mml:mo>)</mml:mo><mml:mo>(</mml:mo><mml:mtext>TP</mml:mtext><mml:mo>+</mml:mo><mml:mtext>FP</mml:mtext><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mtext>TP</mml:mtext><mml:mo>+</mml:mo><mml:mtext>FN</mml:mtext><mml:mo>+</mml:mo><mml:mtext>FP</mml:mtext><mml:mo>+</mml:mo><mml:mtext>TN</mml:mtext></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
            where the rnd term estimates the influence of random true positives.</p>
</sec>
<sec id="App1.Ch1.S1.SS4.SSS3">
  <label>A4.3</label><title>Radially averaged power spectrum density (RAPSD)</title>
      <?pagebreak page3860?><p id="d1e5207">The radially averaged power spectrum density (RAPSD) <xref ref-type="bibr" rid="bib1.bibx39 bib1.bibx49" id="paren.82"/> measures how well the power spectrum of precipitation is maintained when calculated for nowcasts at different lead times. RAPSD fidelity is summarized as
              <disp-formula id="App1.Ch1.S1.E14" content-type="numbered"><label>A7</label><mml:math id="M202" display="block"><mml:mrow><mml:mtext>RAPSD rel. MAE</mml:mtext><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mi>F</mml:mi></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>f</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>F</mml:mi></mml:munderover><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mo>∣</mml:mo><mml:msubsup><mml:mi>P</mml:mi><mml:mi>f</mml:mi><mml:mtext>obs</mml:mtext></mml:msubsup><mml:mo>-</mml:mo><mml:msubsup><mml:mi>P</mml:mi><mml:mi>f</mml:mi><mml:mtext>pred</mml:mtext></mml:msubsup><mml:mo>∣</mml:mo></mml:mrow><mml:mrow><mml:msubsup><mml:mi>P</mml:mi><mml:mi>f</mml:mi><mml:mtext>obs</mml:mtext></mml:msubsup></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
            which is the absolute error between the observed and predicted PSD, relative to observed PSD, averaged over frequencies. Here, <inline-formula><mml:math id="M203" display="inline"><mml:mi>F</mml:mi></mml:math></inline-formula> denotes the number of frequencies of the power spectrum, <inline-formula><mml:math id="M204" display="inline"><mml:mrow><mml:msubsup><mml:mi>P</mml:mi><mml:mi>f</mml:mi><mml:mtext>obs</mml:mtext></mml:msubsup></mml:mrow></mml:math></inline-formula> is the power of the observed field at the <inline-formula><mml:math id="M205" display="inline"><mml:mi>f</mml:mi></mml:math></inline-formula>th frequency, and <inline-formula><mml:math id="M206" display="inline"><mml:mrow><mml:msubsup><mml:mi>P</mml:mi><mml:mi>f</mml:mi><mml:mtext>pred</mml:mtext></mml:msubsup></mml:mrow></mml:math></inline-formula> is the power of the predicted field at the <inline-formula><mml:math id="M207" display="inline"><mml:mi>f</mml:mi></mml:math></inline-formula>th frequency. Taking the relative values allows the comparison of spectral densities on multiple scales. In the present case, PSD frequencies are sampled linearly, weighting corresponding wavelengths towards smaller scales, effectively biasing small-scale errors to be more important. This is, however, not necessarily a problem, as prediction fidelity at small scales is the most important question we seek to answer with RAPSD.</p>
</sec>
</sec>
<sec id="App1.Ch1.S1.SS5">
  <label>A5</label><title>Hardware and software packages used</title>
      <p id="d1e5332">The DEUCE model was built on PyTorch (version <monospace>1.12.1</monospace>). PyTorch Lightning (version <monospace>1.7.7</monospace>) was used to organize the neural network training and prediction workflow, and the TyXe library (version <monospace>0.0.1</monospace>) was used to turn DEUCE Bayesian, making use of the Pyro (version <monospace>1.4.0</monospace>) probabilistic programming language as its back-end for variational inference. The DEUCE training and prediction were performed using the Finnish IT Center for Science (CSC) supercomputer Puhti, using one NVIDIA V100 GPU with 32 GB of VRAM, 64 GB of RAM, and 10 cores from a 2.1 GHz Intel Xeon Gold 6230 CPU. For the evaluation of the model performance, we used the pySTEPS library (version <monospace>1.6.1</monospace>). It served to produce baseline extrapolation-based model nowcasts, to calculate verification metrics, and to help with their visualization. The pySTEPS-based verification pipeline was run on a computational server of the Finnish Meteorological Institute equipped with two Intel Xeon Gold 6138 2.0 GHz CPUs, each with 20 cores and 2 threads by core, as well as 192 GB of RAM. <?xmltex \hack{\newpage}?></p>
</sec>
</app>

<app id="App1.Ch1.S2">
  <?xmltex \currentcnt{B}?><label>Appendix B</label><title>Results of the second case study</title>
      <p id="d1e5360">The results of the second case study show a similar behavior to that of the first case study (Sect. <xref ref-type="sec" rid="Ch1.S4.SS1"/>) but generally lower uncertainty values, especially on the inside of the areas containing precipitation.</p>
<sec id="App1.Ch1.S2.SS1">
  <label>B1</label><title>Ensemble mean and breadth</title>
      <p id="d1e5372">Ensemble mean and breadth as units of standard deviation are shown for the second case in Fig. <xref ref-type="fig" rid="App1.Ch1.S2.F17"/>. Here, we observe a generally similar trend, as with the first case (Sect. <xref ref-type="sec" rid="Ch1.S4.SS1"/>), with the difference being that the ensemble breadth of STEPS is only higher than that of DEUCE towards the center of the rainfall areas, as it tends to be similar at the outskirts of those larger areas. Among the different models, we can also observe the most heterogeneity and anisotropy in the predictive distribution of DEUCE.</p>
</sec>
<sec id="App1.Ch1.S2.SS2">
  <label>B2</label><title>Reflectivity exceedance probabilities</title>
      <p id="d1e5387">Reflectivity probabilities of exceeding 25 dBZ predicted by the different models for the second case are depicted in Fig. <xref ref-type="fig" rid="App1.Ch1.S2.F18"/>, and many of the same comments can be made with respect to the first case (Sect. <xref ref-type="sec" rid="Ch1.S4.SS1"/>), with DEUCE seeming to offer the best balance between accuracy and lacking too many false positives.</p>
      <?pagebreak page3861?><p id="d1e5394">For this second case study, the exceedance probabilities of DEUCE at  thresholds of 15, 25, 35, and 45 dBZ are shown in Fig. <xref ref-type="fig" rid="App1.Ch1.S2.F19"/>. It is worth pointing out that DEUCE was not able to predict the growth of a substantial new rainfall area in the northwest of the composite, despite the predictive uncertainty being significant there, as shown in Fig. <xref ref-type="fig" rid="App1.Ch1.S2.F17"/>, which can be understood because the predictive means under the 8 dBZ threshold were generally closer to the minimum value of <inline-formula><mml:math id="M208" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>10 dBZ. <?xmltex \hack{\clearpage}?></p>

      <?xmltex \floatpos{h!}?><fig id="App1.Ch1.S2.F17"><?xmltex \currentcnt{B1}?><?xmltex \def\figurename{Figure}?><label>Figure B1</label><caption><p id="d1e5411">The second case study ensemble means and breadths of DEUCE compared to STEPS and LINDA-P model predictions and observations for multiple lead times. The area covers southern Finland, starting at 16:50:00 UTC on 17 August 2021. The rows represent lead time and columns different instances of observations, model mean, and standard deviations. Missing values are indicated by a dark gray color.</p></caption>
          <?xmltex \hack{\hsize\textwidth}?>
          <?xmltex \igopts{width=469.470472pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/17/3839/2024/gmd-17-3839-2024-f17.png"/>

        </fig>

      <?xmltex \floatpos{h!}?><fig id="App1.Ch1.S2.F18"><?xmltex \currentcnt{B2}?><?xmltex \def\figurename{Figure}?><label>Figure B2</label><caption><p id="d1e5425">The second case study reflectivity exceedance probabilities of 25 dBZ for DEUCE compared to STEPS and LINDA-P model predictions and observations for multiple lead times. The area covers southern Finland, starting at 16:50:00 UTC on 17 August 2021. The rows represent lead time. In the leftmost column, actual observations are shown in light gray, with echo isotherms corresponding to the threshold marked in black. In other columns, the same isotherms are overlaid on the exceedance probabilities of the models and depicted in shades of red. Missing values are indicated by a dark gray color.</p></caption>
          <?xmltex \hack{\hsize\textwidth}?>
          <?xmltex \igopts{width=284.527559pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/17/3839/2024/gmd-17-3839-2024-f18.png"/>

        </fig>

<?xmltex \hack{\clearpage}?><?xmltex \floatpos{h!}?><fig id="App1.Ch1.S2.F19"><?xmltex \currentcnt{B3}?><?xmltex \def\figurename{Figure}?><label>Figure B3</label><caption><p id="d1e5439">The second case study reflectivity exceedance probabilities of DEUCE compared to observations for multiple lead times and reflectivity thresholds. The area covers southern Finland, starting at 16:50:00 UTC on 17 August 2021. The rows represent lead time. The leftmost column is observations, and the rest are exceedance probabilities at different thresholds. As in Fig. <xref ref-type="fig" rid="App1.Ch1.S2.F18"/>, the observation isotherms corresponding to the threshold in question overlay the exceedance probabilities depicted in shades of red.</p></caption>
          <?xmltex \hack{\hsize\textwidth}?>
          <?xmltex \igopts{width=355.659449pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/17/3839/2024/gmd-17-3839-2024-f19.png"/>

        </fig>

<?xmltex \hack{\clearpage}?>
</sec>
</app>

<?pagebreak page3863?><app id="App1.Ch1.S3">
  <?xmltex \currentcnt{C}?><label>Appendix C</label><title>Additional figures</title>

      <?xmltex \floatpos{h!}?><fig id="App1.Ch1.S3.F20"><?xmltex \currentcnt{C1}?><?xmltex \def\figurename{Figure}?><label>Figure C1</label><caption><p id="d1e5465">Five randomly selected examples of post-processed DEUCE ensemble members on the first case studied, whose area covers southern Finland, starting at 15:00:00 UTC on 9 July 2022. The prediction lead times illustrated are 5, 15, 30, and 60 min.</p></caption>
        <?xmltex \hack{\hsize\textwidth}?>
        <?xmltex \igopts{width=327.206693pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/17/3839/2024/gmd-17-3839-2024-f20.png"/>

      </fig>

<?xmltex \hack{\clearpage}?>
</app>
  </app-group><notes notes-type="codedataavailability"><title>Code and data availability</title>

      <p id="d1e5482">The data used for the production of the results are available online <xref ref-type="bibr" rid="bib1.bibx17" id="paren.83"/> at  <ext-link xlink:href="https://doi.org/10.23728/fmi-b2share.3efcfc9080fe4871bd756c45373e7c11" ext-link-type="DOI">10.23728/fmi-b2share.3efcfc9080fe4871bd756c45373e7c11</ext-link>. These data include the input data used for the training of DEUCE, prediction generation, and observations for the verification. Pre-trained model checkpoints, the script used to gather neural network inputs into an HDF5 file, and computed metric data are also included.</p>

      <p id="d1e5491">The source code with instructions for the reproduction of results is available online <xref ref-type="bibr" rid="bib1.bibx16" id="paren.84"/> at <ext-link xlink:href="https://doi.org/10.5281/zenodo.7961954" ext-link-type="DOI">10.5281/zenodo.7961954</ext-link> and from GitHub at <uri>https://github.com/fmidev/deuce-nowcasting</uri> (last access: 2 May 2024). This code is used for the training and nowcast generation of DEUCE, the production of baseline nowcasts, the computation of metrics, and the creation of figures presenting these metrics.</p>
  </notes><notes notes-type="authorcontribution"><title>Author contributions</title>

      <p id="d1e5506">BH performed the experiments and analysis, developed the methodology, wrote and maintained software, performed the validation, as well as the visualization work, and wrote the original draft. SP acquired funding and resources and also administrated the project. SP and TM jointly supervised BH in this work. BH, SP, and TM took part in the conceptualization and in the review and editing of the paper.</p>
  </notes><notes notes-type="competinginterests"><title>Competing interests</title>

      <p id="d1e5512">The contact author has declared that none of the authors has any competing interests.</p>
  </notes><notes notes-type="disclaimer"><title>Disclaimer</title>

      <p id="d1e5518">Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. While Copernicus Publications makes every effort to include appropriate place names, the final responsibility lies with the authors.</p>
  </notes><ack><title>Acknowledgements</title><p id="d1e5524">We gratefully acknowledge the Finnish CSC IT Center For Science, for providing us with the computational resources used for the development of the model, and Venkatachalam Chandrasekaran from Colorado State University, for his advice and critical comments regarding the work.</p></ack><notes notes-type="financialsupport"><title>Financial support</title>

      <p id="d1e5529">This research has been supported by the Research Council of Finland (grant no. 341964).</p>
  </notes><notes notes-type="reviewstatement"><title>Review statement</title>

      <p id="d1e5536">This paper was edited by Shu-Chih Yang and reviewed by Jatan Buch and one anonymous referee.</p>
  </notes><?xmltex \hack{\newpage}?><ref-list>
    <title>References</title>

      <ref id="bib1.bibx1"><?xmltex \def\ref@label{{Abdar et~al.(2021)Abdar, Pourpanah, Hussain, Rezazadegan, Liu,
Ghavamzadeh, Fieguth, Cao, Khosravi, Acharya, Makarenkov, and
Nahavandi}}?><label>Abdar et al.(2021)Abdar, Pourpanah, Hussain, Rezazadegan, Liu, Ghavamzadeh, Fieguth, Cao, Khosravi, Acharya, Makarenkov, and Nahavandi</label><?label abdar_review_2021?><mixed-citation>Abdar, M., Pourpanah, F., Hussain, S., Rezazadegan, D., Liu, L., Ghavamzadeh, M., Fieguth, P., Cao, X., Khosravi, A., Acharya, U. R., Makarenkov, V., and Nahavandi, S.: A review of uncertainty quantification in deep learning: Techniques, applications and challenges, Inform. Fusion, 76, 243–297, <ext-link xlink:href="https://doi.org/10.1016/j.inffus.2021.05.008" ext-link-type="DOI">10.1016/j.inffus.2021.05.008</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx2"><?xmltex \def\ref@label{{Agrawal et~al.(2019)Agrawal, Barrington, Bromberg, Burge, Gazen, and
Hickey}}?><label>Agrawal et al.(2019)Agrawal, Barrington, Bromberg, Burge, Gazen, and Hickey</label><?label agrawal_machine_2019?><mixed-citation>Agrawal, S., Barrington, L., Bromberg, C., Burge, J., Gazen, C., and Hickey, J.: Machine Learning for Precipitation Nowcasting from Radar Images, arXiv [preprint], <ext-link xlink:href="https://doi.org/10.48550/arXiv.1912.12132" ext-link-type="DOI">10.48550/arXiv.1912.12132</ext-link>,  2019.</mixed-citation></ref>
      <ref id="bib1.bibx3"><?xmltex \def\ref@label{{Alexander et~al.(2020)Alexander, Dowell, Hu, Olson, Smirnova, Ladwig,
Weygandt, Kenyon, James, Lin, Grell, Ge, Alcott, Benjamin, Brown, Toy,
Ahmadov, Back, Duda, Smith, Hamilton, Jamison, Jankov, and
Turner}}?><label>Alexander et al.(2020)Alexander, Dowell, Hu, Olson, Smirnova, Ladwig, Weygandt, Kenyon, James, Lin, Grell, Ge, Alcott, Benjamin, Brown, Toy, Ahmadov, Back, Duda, Smith, Hamilton, Jamison, Jankov, and Turner</label><?label alexander_rapid_2020?><mixed-citation>Alexander, C., Dowell, D. C., Hu, M., Olson, J., Smirnova, T., Ladwig, T., Weygandt, S., Kenyon, J. S., James, E., Lin, H., Grell, G., Ge, G., Alcott, T., Benjamin, S., Brown, J. M., Toy, M. D., Ahmadov, R., Back, A., Duda, J. D., Smith, M. B., Hamilton, J. A., Jamison, B. D., Jankov, I., and Turner, D. D.: Rapid Refresh (RAP) and High Resolution Rapid Refresh (HRRR) Model Development, 100th Annual AMS Meeting, Boston Convention and Exhibition Center 415 Summer St. Boston, MA, <uri>https://rapidrefresh.noaa.gov/pdf/Alexander_AMS_NWP_2020.pdf</uri> (last access: 2 May 2024), 2020.</mixed-citation></ref>
      <ref id="bib1.bibx4"><?xmltex \def\ref@label{{Ayzel et~al.(2020)Ayzel, Scheffer, and
Heistermann}}?><label>Ayzel et al.(2020)Ayzel, Scheffer, and Heistermann</label><?label ayzel_rainnet_2020?><mixed-citation>Ayzel, G., Scheffer, T., and Heistermann, M.: RainNet v1.0: a convolutional neural network for radar-based precipitation nowcasting, Geosci. Model Dev., 13, 2631–2644, <ext-link xlink:href="https://doi.org/10.5194/gmd-13-2631-2020" ext-link-type="DOI">10.5194/gmd-13-2631-2020</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx5"><?xmltex \def\ref@label{{Bauer et~al.(2015)Bauer, Thorpe, and Brunet}}?><label>Bauer et al.(2015)Bauer, Thorpe, and Brunet</label><?label bauer_quiet_2015?><mixed-citation>Bauer, P., Thorpe, A., and Brunet, G.: The quiet revolution of numerical weather prediction, Nature, 525, 47–55, <ext-link xlink:href="https://doi.org/10.1038/nature14956" ext-link-type="DOI">10.1038/nature14956</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx6"><?xmltex \def\ref@label{{Blundell et~al.(2015)Blundell, Cornebise, Kavukcuoglu, and
Wierstra}}?><label>Blundell et al.(2015)Blundell, Cornebise, Kavukcuoglu, and Wierstra</label><?label blundell_weight_2015?><mixed-citation>Blundell, C., Cornebise, J., Kavukcuoglu, K., and Wierstra, D.: Weight Uncertainty in Neural Network, in: Proceedings of the 32nd International Conference on Machine Learning,   1613–1622, PMLR, Lille, France, <uri>https://proceedings.mlr.press/v37/blundell15.html</uri> (last access: 2 May 2024), 2015.</mixed-citation></ref>
      <ref id="bib1.bibx7"><?xmltex \def\ref@label{{Bouguet(2001)}}?><label>Bouguet(2001)</label><?label bouguet2001pyramidal?><mixed-citation>Bouguet, J.-Y.: Pyramidal implementation of the affine lucas kanade feature tracker description of the algorithm, Intel corporation, 5, <uri>http://robots.stanford.edu/cs223b04/algo_tracking.pdf</uri> (last access: 2 May 2024), 2001.</mixed-citation></ref>
      <ref id="bib1.bibx8"><?xmltex \def\ref@label{{Bowler et~al.(2006)Bowler, Pierce, and Seed}}?><label>Bowler et al.(2006)Bowler, Pierce, and Seed</label><?label bowler_steps_2006?><mixed-citation>Bowler, N. E., Pierce, C. E., and Seed, A. W.: STEPS: A probabilistic precipitation forecasting scheme which merges an extrapolation nowcast with downscaled NWP, Q. J. Roy. Meteor. Soc., 132, 2127–2155, <ext-link xlink:href="https://doi.org/10.1256/qj.04.100" ext-link-type="DOI">10.1256/qj.04.100</ext-link>, 2006.</mixed-citation></ref>
      <ref id="bib1.bibx9"><?xmltex \def\ref@label{{Caceres et~al.(2021)Caceres, Gonzalez, Zhou, and
Droguett}}?><label>Caceres et al.(2021)Caceres, Gonzalez, Zhou, and Droguett</label><?label caceres_probabilistic_2021?><mixed-citation>Caceres, J., Gonzalez, D., Zhou, T., and Droguett, E. L.: A probabilistic Bayesian recurrent neural network for remaining useful life prognostics considering epistemic and aleatory uncertainties, Struct. Contr. Health Monit., 28, e2811, <ext-link xlink:href="https://doi.org/10.1002/stc.2811" ext-link-type="DOI">10.1002/stc.2811</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx10"><?xmltex \def\ref@label{{Dechesne et~al.(2021)Dechesne, Lassalle, and
Lefèvre}}?><label>Dechesne et al.(2021)Dechesne, Lassalle, and Lefèvre</label><?label dechesne_bayesian_2021?><mixed-citation>Dechesne, C., Lassalle, P., and Lefèvre, S.: Bayesian U-Net: Estimating Uncertainty in Semantic Segmentation of Earth Observation Images, Remote Sens., 13, 3836, <ext-link xlink:href="https://doi.org/10.3390/rs13193836" ext-link-type="DOI">10.3390/rs13193836</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx11"><?xmltex \def\ref@label{{Espeholt et~al.(2022)Espeholt, Agrawal, Sønderby, Kumar, Heek,
Bromberg, Gazen, Carver, Andrychowicz, Hickey, Bell, and
Kalchbrenner}}?><label>Espeholt et al.(2022)Espeholt, Agrawal, Sønderby, Kumar, Heek, Bromberg, Gazen, Carver, Andrychowicz, Hickey, Bell, and Kalchbrenner</label><?label espeholt_deep_2022?><mixed-citation>Espeholt, L., Agrawal, S., Sønderby, C., Kumar, M., Heek, J., Bromberg, C., Gazen, C., Carver, R., Andrychowicz, M., Hickey, J., Bell, A., and Kalchbrenner, N.: Deep learning for twelve hour precipitation forecasts, Nat. Commun., 13, 5145, <ext-link xlink:href="https://doi.org/10.1038/s41467-022-32483-x" ext-link-type="DOI">10.1038/s41467-022-32483-x</ext-link>, 2022.</mixed-citation></ref>
      <?pagebreak page3865?><ref id="bib1.bibx12"><?xmltex \def\ref@label{{Farquhar et~al.(2020)Farquhar, Osborne, and
Gal}}?><label>Farquhar et al.(2020)Farquhar, Osborne, and Gal</label><?label farquhar_radial_2020?><mixed-citation>Farquhar, S., Osborne, M. A., and Gal, Y.: Radial Bayesian Neural Networks: Beyond Discrete Support In Large-Scale Bayesian Deep Learning, in: Proceedings of the Twenty Third International Conference on Artificial Intelligence and Statistics, PMLR,  1352–1362, <uri>https://proceedings.mlr.press/v108/farquhar20a.html</uri> (last access: 2 May 2024), 2020.</mixed-citation></ref>
      <ref id="bib1.bibx13"><?xmltex \def\ref@label{{Gal and Ghahramani(2016)}}?><label>Gal and Ghahramani(2016)</label><?label gal_dropout_2016?><mixed-citation>Gal, Y. and Ghahramani, Z.: Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning, in: Proceedings of The 33rd International Conference on Machine Learning, PMLR, New York, NY, USA,  1050–1059, <uri>https://proceedings.mlr.press/v48/gal16.html</uri> (last access: 2 May 2024), 2016.</mixed-citation></ref>
      <ref id="bib1.bibx14"><?xmltex \def\ref@label{{Graves(2011)}}?><label>Graves(2011)</label><?label graves_practical_2011?><mixed-citation>Graves, A.: Practical Variational Inference for Neural Networks, in: Advances in Neural Information Processing Systems, vol. 24, Curran Associates, Inc., Granada, Spain,   2348–2356,  ISBN 978-1-61839-599-3, <uri>https://papers.nips.cc/paper_files/paper/2011/hash/7eb3c8be3d411e8ebfab08eba5f49632-Abstract.html</uri> (last access: 2 May 2024), 2011.</mixed-citation></ref>
      <ref id="bib1.bibx15"><?xmltex \def\ref@label{{Harnist(2022)}}?><label>Harnist(2022)</label><?label Harnist2022?><mixed-citation>Harnist, B.: Probabilistic Precipitation Nowcasting using Bayesian Convolutional Neural Networks, Master's thesis, Aalto University, School of Science, <uri>http://urn.fi/URN:NBN:fi:aalto-202208285227</uri> (last access: 2 May 2024), 2022.</mixed-citation></ref>
      <ref id="bib1.bibx16"><?xmltex \def\ref@label{{Harnist(2023)}}?><label>Harnist(2023)</label><?label bent_harnist_2023_7961955?><mixed-citation>Harnist, B.: fmidev/deuce-nowcasting: Initial release of the source code for the manuscript,  Zenodo [code], <ext-link xlink:href="https://doi.org/10.5281/zenodo.7961955" ext-link-type="DOI">10.5281/zenodo.7961955</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx17"><?xmltex \def\ref@label{{Harnist et~al.(2023)Harnist, Pulkkinen, and
Mäkinen}}?><label>Harnist et al.(2023)Harnist, Pulkkinen, and Mäkinen</label><?label Harnist?><mixed-citation>Harnist, B., Pulkkinen, S., and Mäkinen, T.: Data for the manuscript   “DEUCE v1.0: A neural network for probabilistic precipitation nowcasting with aleatoric and epistemic uncertainties” by Harnist et al. (2023), Finnish Meteorological Institute [data set], <ext-link xlink:href="https://doi.org/10.23728/FMI-B2SHARE.3EFCFC9080FE4871BD756C45373E7C11" ext-link-type="DOI">10.23728/FMI-B2SHARE.3EFCFC9080FE4871BD756C45373E7C11</ext-link>,  2023.</mixed-citation></ref>
      <ref id="bib1.bibx18"><?xmltex \def\ref@label{{Hersbach(2000)}}?><label>Hersbach(2000)</label><?label hersbach_decomposition_2000?><mixed-citation>Hersbach, H.: Decomposition of the Continuous Ranked Probability Score for Ensemble Prediction Systems, Weather   Forecast., 15, 559–570, <ext-link xlink:href="https://doi.org/10.1175/1520-0434(2000)015&lt;0559:DOTCRP&gt;2.0.CO;2" ext-link-type="DOI">10.1175/1520-0434(2000)015&lt;0559:DOTCRP&gt;2.0.CO;2</ext-link>, 2000.</mixed-citation></ref>
      <ref id="bib1.bibx19"><?xmltex \def\ref@label{{Hershey and Olsen(2007)}}?><label>Hershey and Olsen(2007)</label><?label hershey_approximating_2007?><mixed-citation>Hershey, J. R. and Olsen, P. A.: Approximating the Kullback Leibler Divergence Between Gaussian Mixture Models, in: 2007 IEEE International Conference on Acoustics, Speech and Signal Processing – ICASSP '07, vol. 4, Honolulu, HI, USA, IV–317–IV–320, <ext-link xlink:href="https://doi.org/10.1109/ICASSP.2007.366913" ext-link-type="DOI">10.1109/ICASSP.2007.366913</ext-link>, 2007.</mixed-citation></ref>
      <ref id="bib1.bibx20"><?xmltex \def\ref@label{{Hogan et~al.(2010)Hogan, Ferro, Jolliffe, and
Stephenson}}?><label>Hogan et al.(2010)Hogan, Ferro, Jolliffe, and Stephenson</label><?label hogan_equitability_2010?><mixed-citation>Hogan, R. J., Ferro, C. A. T., Jolliffe, I. T., and Stephenson, D. B.: Equitability Revisited: Why the “Equitable Threat Score” Is Not Equitable, Weather  Forecast., 25, 710–726, <ext-link xlink:href="https://doi.org/10.1175/2009WAF2222350.1" ext-link-type="DOI">10.1175/2009WAF2222350.1</ext-link>, 2010.</mixed-citation></ref>
      <ref id="bib1.bibx21"><?xmltex \def\ref@label{{Jospin et~al.(2022)Jospin, Laga, Boussaid, Buntine, and
Bennamoun}}?><label>Jospin et al.(2022)Jospin, Laga, Boussaid, Buntine, and Bennamoun</label><?label jospin_hands-bayesian_2022?><mixed-citation>Jospin, L. V., Laga, H., Boussaid, F., Buntine, W., and Bennamoun, M.: Hands-On Bayesian Neural Networks – A Tutorial for Deep Learning Users, IEEE Comput. Intell. Mag., 17, 29–48, <ext-link xlink:href="https://doi.org/10.1109/MCI.2022.3155327" ext-link-type="DOI">10.1109/MCI.2022.3155327</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx22"><?xmltex \def\ref@label{{Kedem and Chiu(1987)}}?><label>Kedem and Chiu(1987)</label><?label kedem_lognormality_1987?><mixed-citation>Kedem, B. and Chiu, L. S.: On the lognormality of rain rate, P. Natl. Acad. Sci. USA, 84, 901–905, <ext-link xlink:href="https://doi.org/10.1073/pnas.84.4.901" ext-link-type="DOI">10.1073/pnas.84.4.901</ext-link>, 1987.</mixed-citation></ref>
      <ref id="bib1.bibx23"><?xmltex \def\ref@label{{Kendall and Gal(2017)}}?><label>Kendall and Gal(2017)</label><?label kendall_what_2017?><mixed-citation> Kendall, A. and Gal, Y.: What uncertainties do we need in Bayesian deep learning for computer vision?, in: Proceedings of the 31st International Conference on Neural Information Processing Systems, NIPS'17, Curran Associates Inc., Red Hook, NY, USA,  5580–5590, ISBN 978-1-5108-6096-4, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx24"><?xmltex \def\ref@label{{Kingma and Ba(2015)}}?><label>Kingma and Ba(2015)</label><?label kingma_adam_2015?><mixed-citation>Kingma, D. P. and Ba, J.: Adam: A Method for Stochastic Optimization, in: Proceedings of the 3rd International Conference on Learning Representations (ICLR), San Diego, CA, USA, <ext-link xlink:href="https://doi.org/10.48550/arXiv.1412.6980" ext-link-type="DOI">10.48550/arXiv.1412.6980</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx25"><?xmltex \def\ref@label{{Kullback and Leibler(1951)}}?><label>Kullback and Leibler(1951)</label><?label kullback_information_1951?><mixed-citation>Kullback, S. and Leibler, R. A.: On Information and Sufficiency,   Ann. Math. Stat., 22, 79–86, <ext-link xlink:href="https://doi.org/10.1214/aoms/1177729694" ext-link-type="DOI">10.1214/aoms/1177729694</ext-link>, 1951.</mixed-citation></ref>
      <ref id="bib1.bibx26"><?xmltex \def\ref@label{{Laroche and Zawadzki(1995)}}?><label>Laroche and Zawadzki(1995)</label><?label laroche_retrievals_1995?><mixed-citation>Laroche, S. and Zawadzki, I.: Retrievals of Horizontal Winds from Single-Doppler Clear-Air Data by Methods of Cross Correlation and Variational Analysis, J. Atmos. Ocean. Tech., 12, 721–738, <ext-link xlink:href="https://doi.org/10.1175/1520-0426(1995)012&lt;0721:ROHWFS&gt;2.0.CO;2" ext-link-type="DOI">10.1175/1520-0426(1995)012&lt;0721:ROHWFS&gt;2.0.CO;2</ext-link>, 1995.</mixed-citation></ref>
      <ref id="bib1.bibx27"><?xmltex \def\ref@label{{Leinonen et~al.(2012)Leinonen, Moisseev, Leskinen, and
Petersen}}?><label>Leinonen et al.(2012)Leinonen, Moisseev, Leskinen, and Petersen</label><?label leinonen_climatology_2012?><mixed-citation>Leinonen, J., Moisseev, D., Leskinen, M., and Petersen, W. A.: A Climatology of Disdrometer Measurements of Rainfall in Finland over Five Years with Implications for Global Radar Observations, J. Appl. Meteorol. Clim., 51, 392–404, <ext-link xlink:href="https://doi.org/10.1175/JAMC-D-11-056.1" ext-link-type="DOI">10.1175/JAMC-D-11-056.1</ext-link>, 2012.</mixed-citation></ref>
      <ref id="bib1.bibx28"><?xmltex \def\ref@label{{Liu et~al.(2018)Liu, Reda, Shih, Wang, Tao, and
Catanzaro}}?><label>Liu et al.(2018)Liu, Reda, Shih, Wang, Tao, and Catanzaro</label><?label liu_image_2018?><mixed-citation>Liu, G., Reda, F. A., Shih, K. J., Wang, T.-C., Tao, A., and Catanzaro, B.: Image Inpainting for Irregular Holes Using Partial Convolutions, in: Computer Vision – ECCV 2018 Proceedings, Part XI, edited by Ferrari, V., Hebert, M., Sminchisescu, C., and Weiss, Y., Lecture Notes in Computer Science,   Springer International Publishing, Munich, Germany, 89–105,  ISBN 978-3-030-01252-6, <ext-link xlink:href="https://doi.org/10.1007/978-3-030-01252-6_6" ext-link-type="DOI">10.1007/978-3-030-01252-6_6</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx29"><?xmltex \def\ref@label{{Lucas and Kanade(1981)}}?><label>Lucas and Kanade(1981)</label><?label lucas_iterative_1981?><mixed-citation>Lucas, B. D. and Kanade, T.: An Iterative Image Registration Technique with an Application to Stereo Vision, in: IJCAI'81: Proceedings of the 7th international joint conference on Artificial intelligence, vol. 2, University of British Columbia Vancouver, B.C., Canada,   p. 674, <uri>https://hal.science/hal-03697340</uri> (last access: 2 May 2024), 1981.</mixed-citation></ref>
      <ref id="bib1.bibx30"><?xmltex \def\ref@label{{Mason(1982)}}?><label>Mason(1982)</label><?label mason1982model?><mixed-citation> Mason, I.: A model for assessment of weather forecasts, Austr. Meteorol. Mag., 30, 291–303, 1982.</mixed-citation></ref>
      <ref id="bib1.bibx31"><?xmltex \def\ref@label{{Naeini et~al.(2015)Naeini, Cooper, and
Hauskrecht}}?><label>Naeini et al.(2015)Naeini, Cooper, and Hauskrecht</label><?label naeini_obtaining_2015?><mixed-citation>Naeini, M. P., Cooper, G., and Hauskrecht, M.: Obtaining Well Calibrated Probabilities Using Bayesian Binning, in: Proceedings of the Twenty-Ninth AAAI Conference on Artificial Intelligence, vol. 29,  <ext-link xlink:href="https://doi.org/10.1609/aaai.v29i1.9602" ext-link-type="DOI">10.1609/aaai.v29i1.9602</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx32"><?xmltex \def\ref@label{{Pan et~al.(2021)Pan, Lu, Zhao, Huang, Wang, and
Chen}}?><label>Pan et al.(2021)Pan, Lu, Zhao, Huang, Wang, and Chen</label><?label pan_improving_2021?><mixed-citation>Pan, X., Lu, Y., Zhao, K., Huang, H., Wang, M., and Chen, H.: Improving Nowcasting of Convective Development by Incorporating Polarimetric Radar Variables Into a Deep-Learning Model, Geophys. Res. Lett., 48, e2021GL095 302, <ext-link xlink:href="https://doi.org/10.1029/2021GL095302" ext-link-type="DOI">10.1029/2021GL095302</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx33"><?xmltex \def\ref@label{{Prudden et~al.(2020)Prudden, Adams, Kangin, Robinson, Ravuri,
Mohamed, and Arribas}}?><label>Prudden et al.(2020)Prudden, Adams, Kangin, Robinson, Ravuri, Mohamed, and Arribas</label><?label prudden_review_2020?><mixed-citation>Prudden, R., Adams, S., Kangin, D., Robinson, N., Ravuri, S., Mohamed, S., and Arribas, A.: A review of radar-based nowcasting of precipitation and applicable machine learning techniques,  arXiv [preprint], <ext-link xlink:href="https://doi.org/10.48550/arXiv.2005.04988" ext-link-type="DOI">10.48550/arXiv.2005.04988</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx34"><?xmltex \def\ref@label{{Pulkkinen et~al.(2019)Pulkkinen, Nerini, Pérez~Hortal,
Velasco-Forero, Seed, Germann, and Foresti}}?><label>Pulkkinen et al.(2019)Pulkkinen, Nerini, Pérez Hortal, Velasco-Forero, Seed, Germann, and Foresti</label><?label pulkkinen_pysteps_2019?><mixed-citation>Pulkkinen, S., Nerini, D., Pérez Hortal, A. A., Velasco-Forero, C., Seed, A., Germann, U., and Foresti, L.: Pysteps: an open-source Python library for probabilistic precipitation nowcasting (v1.0), Geosci. Model Dev., 12, 4185–4219, <ext-link xlink:href="https://doi.org/10.5194/gmd-12-4185-2019" ext-link-type="DOI">10.5194/gmd-12-4185-2019</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx35"><?xmltex \def\ref@label{{Pulkkinen et~al.(2021)Pulkkinen, Chandrasekar, and
Niemi}}?><label>Pulkkinen et al.(2021)Pulkkinen, Chandrasekar, and Niemi</label><?label pulkkinen_lagrangian_2021?><mixed-citation>Pulkkinen, S., Chandrasekar, V., and Niemi, T.: Lagrangian Integro-Difference Equation Model for Precipitation Nowcasting, J.  Atmos. Ocean. Tech., 38, 2125–2145, <ext-link xlink:href="https://doi.org/10.1175/JTECH-D-21-0013.1" ext-link-type="DOI">10.1175/JTECH-D-21-0013.1</ext-link>,  2021.</mixed-citation></ref>
      <ref id="bib1.bibx36"><?xmltex \def\ref@label{{Radhakrishnan and Chandrasekar(2020)}}?><label>Radhakrishnan and Chandrasekar(2020)</label><?label radhakrishnan_casa_2020?><mixed-citation>Radhakrishnan, C. and Chandrasekar, V.: CASA Prediction System over Dallas–Fort Worth Urban Network: Blending of Nowcasting and High-Resolution Numerical Weather Prediction Model, J. Atmos. Ocean. Tech., 37, 211–228, <ext-link xlink:href="https://doi.org/10.1175/JTECH-D-18-0192.1" ext-link-type="DOI">10.1175/JTECH-D-18-0192.1</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx37"><?xmltex \def\ref@label{{Ravuri et~al.(2021)Ravuri, Lenc, Willson, Kangin, Lam, Mirowski,
Fitzsimons, Athanassiadou, Kashem, Madge, Prudden, Mandhane, Clark, Brock,
Simonyan, Hadsell, Robinson, Clancy, Arribas, and
Mohamed}}?><label>Ravuri et al.(2021)Ravuri, Lenc, Willson, Kangin, Lam, Mirowski, Fitzsimons, Athanassiadou, Kashem, Madge, Prudden, Mandhane, Clark, Brock, Simonyan, Hadsell, Robinson, Clancy, Arribas, and Mohamed</label><?label ravuri_skilful_2021?><mixed-citation>Ravuri, S., Lenc, K., Willson, M., Kangin, D., Lam, R., Mi<?pagebreak page3866?>rowski, P., Fitzsimons, M., Athanassiadou, M., Kashem, S., Madge, S., Prudden, R., Mandhane, A., Clark, A., Brock, A., Simonyan, K., Hadsell, R., Robinson, N., Clancy, E., Arribas, A., and Mohamed, S.: Skilful precipitation nowcasting using deep generative models of radar, Nature, 597, 672–677, <ext-link xlink:href="https://doi.org/10.1038/s41586-021-03854-z" ext-link-type="DOI">10.1038/s41586-021-03854-z</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx38"><?xmltex \def\ref@label{{Ritvanen et~al.(2023)Ritvanen, Harnist, Aldana, Mäkinen, and
Pulkkinen}}?><label>Ritvanen et al.(2023)Ritvanen, Harnist, Aldana, Mäkinen, and Pulkkinen</label><?label ritvanen_advection-free_2023?><mixed-citation>Ritvanen, J., Harnist, B., Aldana, M., Mäkinen, T., and Pulkkinen, S.: Advection-Free Convolutional Neural Network for Convective Rainfall Nowcasting, IEEE J. Sel. Top. Appl.,   1–16, <ext-link xlink:href="https://doi.org/10.1109/JSTARS.2023.3238016" ext-link-type="DOI">10.1109/JSTARS.2023.3238016</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx39"><?xmltex \def\ref@label{{Ruzanski and Chandrasekar(2011)}}?><label>Ruzanski and Chandrasekar(2011)</label><?label ruzanski_scale_2011?><mixed-citation>Ruzanski, E. and Chandrasekar, V.: Scale Filtering for Improved Nowcasting Performance in a High-Resolution X-Band Radar Network, IEEE T. Geosci. Remote, 49, 2296–2307, <ext-link xlink:href="https://doi.org/10.1109/TGRS.2010.2103946" ext-link-type="DOI">10.1109/TGRS.2010.2103946</ext-link>, 2011.</mixed-citation></ref>
      <ref id="bib1.bibx40"><?xmltex \def\ref@label{{Schaefer(1990)}}?><label>Schaefer(1990)</label><?label schaefer_critical_1990?><mixed-citation>Schaefer, J. T.: The Critical Success Index as an Indicator of Warning Skill, Weather  Forecast., 5, 570–575, <ext-link xlink:href="https://doi.org/10.1175/1520-0434(1990)005&lt;0570:TCSIAA&gt;2.0.CO;2" ext-link-type="DOI">10.1175/1520-0434(1990)005&lt;0570:TCSIAA&gt;2.0.CO;2</ext-link>, 1990.</mixed-citation></ref>
      <ref id="bib1.bibx41"><?xmltex \def\ref@label{{Seed(2003)}}?><label>Seed(2003)</label><?label seed_dynamic_2003?><mixed-citation>Seed, A. W.: A Dynamic and Spatial Scaling Approach to Advection Forecasting, J. Appl. Meteorol. Clim., 42, 381–388, <ext-link xlink:href="https://doi.org/10.1175/1520-0450(2003)042&lt;0381:ADASSA&gt;2.0.CO;2" ext-link-type="DOI">10.1175/1520-0450(2003)042&lt;0381:ADASSA&gt;2.0.CO;2</ext-link>, 2003.</mixed-citation></ref>
      <ref id="bib1.bibx42"><?xmltex \def\ref@label{{Seed et~al.(2013)Seed, Pierce, and Norman}}?><label>Seed et al.(2013)Seed, Pierce, and Norman</label><?label seed_formulation_2013?><mixed-citation>Seed, A. W., Pierce, C. E., and Norman, K.: Formulation and evaluation of a scale decomposition-based stochastic precipitation nowcast scheme, Water Resour. Res., 49, 6624–6641, <ext-link xlink:href="https://doi.org/10.1002/wrcr.20536" ext-link-type="DOI">10.1002/wrcr.20536</ext-link>, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx43"><?xmltex \def\ref@label{{Shi et~al.(2015)Shi, Chen, Wang, Yeung, Wong, and
Woo}}?><label>Shi et al.(2015)Shi, Chen, Wang, Yeung, Wong, and Woo</label><?label shi_convolutional_2015?><mixed-citation> Shi, X., Chen, Z., Wang, H., Yeung, D.-Y., Wong, W.-K., and Woo, W.-C.: Convolutional LSTM Network: a machine learning approach for precipitation nowcasting, in: Proceedings of the 28th International Conference on Neural Information Processing Systems – Volume 1, NIPS'15, MIT Press, Cambridge, MA, USA, 802–810, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx44"><?xmltex \def\ref@label{{Shi et~al.(2017)Shi, Gao, Lausen, Wang, Yeung, Wong, and
Woo}}?><label>Shi et al.(2017)Shi, Gao, Lausen, Wang, Yeung, Wong, and Woo</label><?label shi_deep_2017?><mixed-citation> Shi, X., Gao, Z., Lausen, L., Wang, H., Yeung, D.-Y., Wong, W.-K., and Woo, W.-C.: Deep learning for precipitation nowcasting: a benchmark and a new model, in: Proceedings of the 31st International Conference on Neural Information Processing Systems, NIPS'17,   Curran Associates Inc., Red Hook, NY, USA, 5622–5632, ISBN 978-1-5108-6096-4, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx45"><?xmltex \def\ref@label{{Staniforth and Côté(1991)}}?><label>Staniforth and Côté(1991)</label><?label staniforth_semi-lagrangian_1991?><mixed-citation>Staniforth, A. and Côté, J.: Semi-Lagrangian Integration Schemes for Atmospheric Models – A Review, Mon. Weather Rev., 119, 2206–2223, <ext-link xlink:href="https://doi.org/10.1175/1520-0493(1991)119&lt;2206:SLISFA&gt;2.0.CO;2" ext-link-type="DOI">10.1175/1520-0493(1991)119&lt;2206:SLISFA&gt;2.0.CO;2</ext-link>, 1991. </mixed-citation></ref><?xmltex \hack{\newpage}?>
      <ref id="bib1.bibx46"><?xmltex \def\ref@label{{Sun et~al.(2014)Sun, Xue, Wilson, Zawadzki, Ballard,
Onvlee-Hooimeyer, Joe, Barker, Li, Golding, Xu, and Pinto}}?><label>Sun et al.(2014)Sun, Xue, Wilson, Zawadzki, Ballard, Onvlee-Hooimeyer, Joe, Barker, Li, Golding, Xu, and Pinto</label><?label sun_use_2014?><mixed-citation>Sun, J., Xue, M., Wilson, J. W., Zawadzki, I., Ballard, S. P., Onvlee-Hooimeyer, J., Joe, P., Barker, D. M., Li, P.-W., Golding, B., Xu, M., and Pinto, J.: Use of NWP for Nowcasting Convective Precipitation: Recent Progress and Challenges, B. Am. Meteorol. Soc., 95, 409–426, <ext-link xlink:href="https://doi.org/10.1175/BAMS-D-11-00263.1" ext-link-type="DOI">10.1175/BAMS-D-11-00263.1</ext-link>, 2014.</mixed-citation></ref>
      <ref id="bib1.bibx47"><?xmltex \def\ref@label{{Sønderby et~al.(2020)Sønderby, Espeholt, Heek, Dehghani, Oliver,
Salimans, Agrawal, Hickey, and Kalchbrenner}}?><label>Sønderby et al.(2020)Sønderby, Espeholt, Heek, Dehghani, Oliver, Salimans, Agrawal, Hickey, and Kalchbrenner</label><?label sonderby_metnet_2020?><mixed-citation>Sønderby, C. K., Espeholt, L., Heek, J., Dehghani, M., Oliver, A., Salimans, T., Agrawal, S., Hickey, J., and Kalchbrenner, N.: MetNet: A Neural Weather Model for Precipitation Forecasting, arXiv [preprint], <ext-link xlink:href="https://doi.org/10.48550/arXiv.2003.12140" ext-link-type="DOI">10.48550/arXiv.2003.12140</ext-link>,  2020.</mixed-citation></ref>
      <ref id="bib1.bibx48"><?xmltex \def\ref@label{{Trebing et~al.(2021)Trebing, Stanczyk, and
Mehrkanoon}}?><label>Trebing et al.(2021)Trebing, Stanczyk, and Mehrkanoon</label><?label trebing2021smaat?><mixed-citation> Trebing, K., Stanczyk, T., and Mehrkanoon, S.: SmaAt-UNet: Precipitation nowcasting using a small attention-UNet architecture, Pattern Recogn. Lett., 145, 178–186, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx49"><?xmltex \def\ref@label{{Ulichney(1988)}}?><label>Ulichney(1988)</label><?label ulichney_dithering_1988?><mixed-citation>Ulichney, R.: Dithering with blue noise, P. IEEE, 76, 56–79, <ext-link xlink:href="https://doi.org/10.1109/5.3288" ext-link-type="DOI">10.1109/5.3288</ext-link>, 1988.</mixed-citation></ref>
      <ref id="bib1.bibx50"><?xmltex \def\ref@label{{Valdenegro-Toro and Mori(2022)}}?><label>Valdenegro-Toro and Mori(2022)</label><?label valdenegro-toro_deeper_2022?><mixed-citation>Valdenegro-Toro, M. and Mori, D. S.: A Deeper Look into Aleatoric and Epistemic Uncertainty Disentanglement, in: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW),   IEEE, New Orleans, LA, USA, ISBN 978-1-66548-739-9, 1508–1516, <ext-link xlink:href="https://doi.org/10.1109/CVPRW56347.2022.00157" ext-link-type="DOI">10.1109/CVPRW56347.2022.00157</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx51"><?xmltex \def\ref@label{{Wen et~al.(2018)Wen, Vicol, Ba, Tran, and Grosse}}?><label>Wen et al.(2018)Wen, Vicol, Ba, Tran, and Grosse</label><?label wen_flipout_2018?><mixed-citation>Wen, Y., Vicol, P., Ba, J., Tran, D., and Grosse, R. B.: Flipout: Efficient Pseudo-Independent Weight Perturbations on Mini-Batches, in: 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada,  30 April–3 May   2018, Conference Track Proceedings, OpenReview.net, <uri>https://openreview.net/forum?id=rJNpifWAb</uri> (last access: 2 May 2024), 2018.</mixed-citation></ref>
      <ref id="bib1.bibx52"><?xmltex \def\ref@label{{Wilks(2011)}}?><label>Wilks(2011)</label><?label wilks_statistical_2011?><mixed-citation> Wilks, D. S.: Statistical Methods in the Atmospheric Sciences, Academic Press, 3rd Edn., ISBN 978-0-12-385022-5, 2011.</mixed-citation></ref>
      <ref id="bib1.bibx53"><?xmltex \def\ref@label{{Woo et~al.(2018)Woo, Park, Lee, and Kweon}}?><label>Woo et al.(2018)Woo, Park, Lee, and Kweon</label><?label woo2018cbam?><mixed-citation>Woo, S., Park, J., Lee, J.-Y., and Kweon, I. S.: Cbam: Convolutional block attention module, in: Proceedings of the European conference on computer vision (ECCV), Munich, Germany, 3–19, <ext-link xlink:href="https://doi.org/10.1007/978-3-030-01234-2_1" ext-link-type="DOI">10.1007/978-3-030-01234-2_1</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx54"><?xmltex \def\ref@label{{Xu et~al.(2022)Xu, Chen, Yang, Yu, and Chen}}?><label>Xu et al.(2022)Xu, Chen, Yang, Yu, and Chen</label><?label xu_quantifying_2022?><mixed-citation>Xu, L., Chen, N., Yang, C., Yu, H., and Chen, Z.: Quantifying the uncertainty of precipitation forecasting using probabilistic deep learning, Hydrol. Earth Syst. Sci., 26, 2923–2938, <ext-link xlink:href="https://doi.org/10.5194/hess-26-2923-2022" ext-link-type="DOI">10.5194/hess-26-2923-2022</ext-link>, 2022.</mixed-citation></ref>

  </ref-list></back>
    <!--<article-title-html>DEUCE v1.0: a neural network for probabilistic precipitation nowcasting with aleatoric and epistemic uncertainties</article-title-html>
<abstract-html/>
<ref-html id="bib1.bib1"><label>Abdar et al.(2021)Abdar, Pourpanah, Hussain, Rezazadegan, Liu,
Ghavamzadeh, Fieguth, Cao, Khosravi, Acharya, Makarenkov, and
Nahavandi</label><mixed-citation>
      
Abdar, M., Pourpanah, F., Hussain, S., Rezazadegan, D., Liu, L., Ghavamzadeh,
M., Fieguth, P., Cao, X., Khosravi, A., Acharya, U. R., Makarenkov, V., and
Nahavandi, S.: A review of uncertainty quantification in deep learning:
Techniques, applications and challenges, Inform. Fusion, 76, 243–297,
<a href="https://doi.org/10.1016/j.inffus.2021.05.008" target="_blank">https://doi.org/10.1016/j.inffus.2021.05.008</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib2"><label>Agrawal et al.(2019)Agrawal, Barrington, Bromberg, Burge, Gazen, and
Hickey</label><mixed-citation>
      
Agrawal, S., Barrington, L., Bromberg, C., Burge, J., Gazen, C., and Hickey,
J.: Machine Learning for Precipitation Nowcasting from Radar
Images, arXiv [preprint], <a href="https://doi.org/10.48550/arXiv.1912.12132" target="_blank">https://doi.org/10.48550/arXiv.1912.12132</a>,  2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib3"><label>Alexander et al.(2020)Alexander, Dowell, Hu, Olson, Smirnova, Ladwig,
Weygandt, Kenyon, James, Lin, Grell, Ge, Alcott, Benjamin, Brown, Toy,
Ahmadov, Back, Duda, Smith, Hamilton, Jamison, Jankov, and
Turner</label><mixed-citation>
      
Alexander, C., Dowell, D. C., Hu, M., Olson, J., Smirnova, T., Ladwig, T.,
Weygandt, S., Kenyon, J. S., James, E., Lin, H., Grell, G., Ge, G., Alcott,
T., Benjamin, S., Brown, J. M., Toy, M. D., Ahmadov, R., Back, A., Duda,
J. D., Smith, M. B., Hamilton, J. A., Jamison, B. D., Jankov, I., and Turner,
D. D.: Rapid Refresh (RAP) and High Resolution Rapid Refresh
(HRRR) Model Development, 100th Annual AMS Meeting, Boston Convention and Exhibition Center 415
Summer St. Boston, MA,
<a href="https://rapidrefresh.noaa.gov/pdf/Alexander_AMS_NWP_2020.pdf" target="_blank"/> (last access: 2 May 2024),
2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib4"><label>Ayzel et al.(2020)Ayzel, Scheffer, and
Heistermann</label><mixed-citation>
      
Ayzel, G., Scheffer, T., and Heistermann, M.: RainNet v1.0: a convolutional neural network for radar-based precipitation nowcasting, Geosci. Model Dev., 13, 2631–2644, <a href="https://doi.org/10.5194/gmd-13-2631-2020" target="_blank">https://doi.org/10.5194/gmd-13-2631-2020</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib5"><label>Bauer et al.(2015)Bauer, Thorpe, and Brunet</label><mixed-citation>
      
Bauer, P., Thorpe, A., and Brunet, G.: The quiet revolution of numerical
weather prediction, Nature, 525, 47–55, <a href="https://doi.org/10.1038/nature14956" target="_blank">https://doi.org/10.1038/nature14956</a>, 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib6"><label>Blundell et al.(2015)Blundell, Cornebise, Kavukcuoglu, and
Wierstra</label><mixed-citation>
      
Blundell, C., Cornebise, J., Kavukcuoglu, K., and Wierstra, D.: Weight
Uncertainty in Neural Network, in: Proceedings of the 32nd
International Conference on Machine Learning,   1613–1622, PMLR,
Lille, France,
<a href="https://proceedings.mlr.press/v37/blundell15.html" target="_blank"/> (last access: 2 May 2024), 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib7"><label>Bouguet(2001)</label><mixed-citation>
      
Bouguet, J.-Y.: Pyramidal implementation of the affine lucas kanade
feature tracker description of the algorithm, Intel corporation, 5, <a href="http://robots.stanford.edu/cs223b04/algo_tracking.pdf" target="_blank"/> (last access: 2 May 2024), 2001.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib8"><label>Bowler et al.(2006)Bowler, Pierce, and Seed</label><mixed-citation>
      
Bowler, N. E., Pierce, C. E., and Seed, A. W.: STEPS: A probabilistic
precipitation forecasting scheme which merges an extrapolation nowcast with
downscaled NWP, Q. J. Roy. Meteor. Soc., 132,
2127–2155, <a href="https://doi.org/10.1256/qj.04.100" target="_blank">https://doi.org/10.1256/qj.04.100</a>, 2006.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib9"><label>Caceres et al.(2021)Caceres, Gonzalez, Zhou, and
Droguett</label><mixed-citation>
      
Caceres, J., Gonzalez, D., Zhou, T., and Droguett, E. L.: A probabilistic
Bayesian recurrent neural network for remaining useful life prognostics
considering epistemic and aleatory uncertainties, Struct. Contr.
Health Monit., 28, e2811, <a href="https://doi.org/10.1002/stc.2811" target="_blank">https://doi.org/10.1002/stc.2811</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib10"><label>Dechesne et al.(2021)Dechesne, Lassalle, and
Lefèvre</label><mixed-citation>
      
Dechesne, C., Lassalle, P., and Lefèvre, S.: Bayesian U-Net: Estimating
Uncertainty in Semantic Segmentation of Earth Observation Images,
Remote Sens., 13, 3836, <a href="https://doi.org/10.3390/rs13193836" target="_blank">https://doi.org/10.3390/rs13193836</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib11"><label>Espeholt et al.(2022)Espeholt, Agrawal, Sønderby, Kumar, Heek,
Bromberg, Gazen, Carver, Andrychowicz, Hickey, Bell, and
Kalchbrenner</label><mixed-citation>
      
Espeholt, L., Agrawal, S., Sønderby, C., Kumar, M., Heek, J., Bromberg, C.,
Gazen, C., Carver, R., Andrychowicz, M., Hickey, J., Bell, A., and
Kalchbrenner, N.: Deep learning for twelve hour precipitation forecasts,
Nat. Commun., 13, 5145, <a href="https://doi.org/10.1038/s41467-022-32483-x" target="_blank">https://doi.org/10.1038/s41467-022-32483-x</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib12"><label>Farquhar et al.(2020)Farquhar, Osborne, and
Gal</label><mixed-citation>
      
Farquhar, S., Osborne, M. A., and Gal, Y.: Radial Bayesian Neural
Networks: Beyond Discrete Support In Large-Scale Bayesian
Deep Learning, in: Proceedings of the Twenty Third International
Conference on Artificial Intelligence and Statistics,
PMLR,  1352–1362,
<a href="https://proceedings.mlr.press/v108/farquhar20a.html" target="_blank"/> (last access: 2 May 2024), 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib13"><label>Gal and Ghahramani(2016)</label><mixed-citation>
      
Gal, Y. and Ghahramani, Z.: Dropout as a Bayesian Approximation:
Representing Model Uncertainty in Deep Learning, in: Proceedings of
The 33rd International Conference on Machine Learning,
PMLR, New York, NY, USA,  1050–1059,
<a href="https://proceedings.mlr.press/v48/gal16.html" target="_blank"/> (last access: 2 May 2024), 2016.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib14"><label>Graves(2011)</label><mixed-citation>
      
Graves, A.: Practical Variational Inference for Neural Networks, in:
Advances in Neural Information Processing Systems, vol. 24,
Curran Associates, Inc., Granada, Spain,   2348–2356,  ISBN 978-1-61839-599-3,
<a href="https://papers.nips.cc/paper_files/paper/2011/hash/7eb3c8be3d411e8ebfab08eba5f49632-Abstract.html" target="_blank"/> (last access: 2 May 2024),
2011.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib15"><label>Harnist(2022)</label><mixed-citation>
      
Harnist, B.: Probabilistic Precipitation Nowcasting using Bayesian
Convolutional Neural Networks, Master's thesis, Aalto University, School of
Science, <a href="http://urn.fi/URN:NBN:fi:aalto-202208285227" target="_blank"/> (last access: 2 May 2024), 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib16"><label>Harnist(2023)</label><mixed-citation>
      
Harnist, B.: fmidev/deuce-nowcasting: Initial release of the source code for
the manuscript,  Zenodo [code], <a href="https://doi.org/10.5281/zenodo.7961955" target="_blank">https://doi.org/10.5281/zenodo.7961955</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib17"><label>Harnist et al.(2023)Harnist, Pulkkinen, and
Mäkinen</label><mixed-citation>
      
Harnist, B., Pulkkinen, S., and Mäkinen, T.: Data for the manuscript   “DEUCE
v1.0: A neural network for probabilistic precipitation nowcasting with
aleatoric and epistemic uncertainties” by Harnist et al. (2023), Finnish
Meteorological Institute [data set],
<a href="https://doi.org/10.23728/FMI-B2SHARE.3EFCFC9080FE4871BD756C45373E7C11" target="_blank">https://doi.org/10.23728/FMI-B2SHARE.3EFCFC9080FE4871BD756C45373E7C11</a>,  2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib18"><label>Hersbach(2000)</label><mixed-citation>
      
Hersbach, H.: Decomposition of the Continuous Ranked Probability Score
for Ensemble Prediction Systems, Weather   Forecast., 15, 559–570,
<a href="https://doi.org/10.1175/1520-0434(2000)015&lt;0559:DOTCRP&gt;2.0.CO;2" target="_blank">https://doi.org/10.1175/1520-0434(2000)015&lt;0559:DOTCRP&gt;2.0.CO;2</a>, 2000.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib19"><label>Hershey and Olsen(2007)</label><mixed-citation>
      
Hershey, J. R. and Olsen, P. A.: Approximating the Kullback Leibler
Divergence Between Gaussian Mixture Models, in: 2007 IEEE
International Conference on Acoustics, Speech and Signal
Processing – ICASSP '07, vol. 4, Honolulu, HI, USA,
IV–317–IV–320, <a href="https://doi.org/10.1109/ICASSP.2007.366913" target="_blank">https://doi.org/10.1109/ICASSP.2007.366913</a>, 2007.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib20"><label>Hogan et al.(2010)Hogan, Ferro, Jolliffe, and
Stephenson</label><mixed-citation>
      
Hogan, R. J., Ferro, C. A. T., Jolliffe, I. T., and Stephenson, D. B.:
Equitability Revisited: Why the “Equitable Threat Score” Is Not
Equitable, Weather  Forecast., 25, 710–726,
<a href="https://doi.org/10.1175/2009WAF2222350.1" target="_blank">https://doi.org/10.1175/2009WAF2222350.1</a>, 2010.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib21"><label>Jospin et al.(2022)Jospin, Laga, Boussaid, Buntine, and
Bennamoun</label><mixed-citation>
      
Jospin, L. V., Laga, H., Boussaid, F., Buntine, W., and Bennamoun, M.:
Hands-On Bayesian Neural Networks – A Tutorial for Deep
Learning Users, IEEE Comput. Intell. Mag., 17, 29–48,
<a href="https://doi.org/10.1109/MCI.2022.3155327" target="_blank">https://doi.org/10.1109/MCI.2022.3155327</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib22"><label>Kedem and Chiu(1987)</label><mixed-citation>
      
Kedem, B. and Chiu, L. S.: On the lognormality of rain rate, P.
Natl. Acad. Sci. USA, 84, 901–905, <a href="https://doi.org/10.1073/pnas.84.4.901" target="_blank">https://doi.org/10.1073/pnas.84.4.901</a>,
1987.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib23"><label>Kendall and Gal(2017)</label><mixed-citation>
      
Kendall, A. and Gal, Y.: What uncertainties do we need in Bayesian deep
learning for computer vision?, in: Proceedings of the 31st International
Conference on Neural Information Processing Systems, NIPS'17,
Curran Associates Inc., Red Hook, NY, USA,  5580–5590, ISBN
978-1-5108-6096-4, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib24"><label>Kingma and Ba(2015)</label><mixed-citation>
      
Kingma, D. P. and Ba, J.: Adam: A Method for Stochastic Optimization,
in: Proceedings of the 3rd International Conference on Learning
Representations (ICLR), San Diego, CA, USA,
<a href="https://doi.org/10.48550/arXiv.1412.6980" target="_blank">https://doi.org/10.48550/arXiv.1412.6980</a>, 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib25"><label>Kullback and Leibler(1951)</label><mixed-citation>
      
Kullback, S. and Leibler, R. A.: On Information and Sufficiency,   Ann.
Math. Stat., 22, 79–86, <a href="https://doi.org/10.1214/aoms/1177729694" target="_blank">https://doi.org/10.1214/aoms/1177729694</a>, 1951.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib26"><label>Laroche and Zawadzki(1995)</label><mixed-citation>
      
Laroche, S. and Zawadzki, I.: Retrievals of Horizontal Winds from
Single-Doppler Clear-Air Data by Methods of Cross Correlation
and Variational Analysis, J. Atmos. Ocean. Tech.,
12, 721–738, <a href="https://doi.org/10.1175/1520-0426(1995)012&lt;0721:ROHWFS&gt;2.0.CO;2" target="_blank">https://doi.org/10.1175/1520-0426(1995)012&lt;0721:ROHWFS&gt;2.0.CO;2</a>, 1995.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib27"><label>Leinonen et al.(2012)Leinonen, Moisseev, Leskinen, and
Petersen</label><mixed-citation>
      
Leinonen, J., Moisseev, D., Leskinen, M., and Petersen, W. A.: A Climatology
of Disdrometer Measurements of Rainfall in Finland over Five
Years with Implications for Global Radar Observations, J.
Appl. Meteorol. Clim., 51, 392–404,
<a href="https://doi.org/10.1175/JAMC-D-11-056.1" target="_blank">https://doi.org/10.1175/JAMC-D-11-056.1</a>, 2012.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib28"><label>Liu et al.(2018)Liu, Reda, Shih, Wang, Tao, and
Catanzaro</label><mixed-citation>
      
Liu, G., Reda, F. A., Shih, K. J., Wang, T.-C., Tao, A., and Catanzaro, B.:
Image Inpainting for Irregular Holes Using Partial Convolutions,
in: Computer Vision – ECCV 2018 Proceedings, Part XI, edited by
Ferrari, V., Hebert, M., Sminchisescu, C., and Weiss, Y., Lecture Notes in
Computer Science,   Springer International Publishing, Munich,
Germany, 89–105,  ISBN 978-3-030-01252-6, <a href="https://doi.org/10.1007/978-3-030-01252-6_6" target="_blank">https://doi.org/10.1007/978-3-030-01252-6_6</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib29"><label>Lucas and Kanade(1981)</label><mixed-citation>
      
Lucas, B. D. and Kanade, T.: An Iterative Image Registration Technique
with an Application to Stereo Vision, in: IJCAI'81: Proceedings of
the 7th international joint conference on Artificial intelligence, vol. 2,
University of British Columbia Vancouver, B.C., Canada,   p. 674,
<a href="https://hal.science/hal-03697340" target="_blank"/> (last access: 2 May 2024), 1981.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib30"><label>Mason(1982)</label><mixed-citation>
      
Mason, I.: A model for assessment of weather forecasts, Austr.
Meteorol. Mag., 30, 291–303, 1982.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib31"><label>Naeini et al.(2015)Naeini, Cooper, and
Hauskrecht</label><mixed-citation>
      
Naeini, M. P., Cooper, G., and Hauskrecht, M.: Obtaining Well Calibrated
Probabilities Using Bayesian Binning, in: Proceedings of the
Twenty-Ninth AAAI Conference on Artificial Intelligence, vol. 29,  <a href="https://doi.org/10.1609/aaai.v29i1.9602" target="_blank">https://doi.org/10.1609/aaai.v29i1.9602</a>, 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib32"><label>Pan et al.(2021)Pan, Lu, Zhao, Huang, Wang, and
Chen</label><mixed-citation>
      
Pan, X., Lu, Y., Zhao, K., Huang, H., Wang, M., and Chen, H.: Improving
Nowcasting of Convective Development by Incorporating Polarimetric
Radar Variables Into a Deep-Learning Model, Geophys. Res.
Lett., 48, e2021GL095&thinsp;302, <a href="https://doi.org/10.1029/2021GL095302" target="_blank">https://doi.org/10.1029/2021GL095302</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib33"><label>Prudden et al.(2020)Prudden, Adams, Kangin, Robinson, Ravuri,
Mohamed, and Arribas</label><mixed-citation>
      
Prudden, R., Adams, S., Kangin, D., Robinson, N., Ravuri, S., Mohamed, S., and
Arribas, A.: A review of radar-based nowcasting of precipitation and
applicable machine learning techniques,  arXiv [preprint], <a href="https://doi.org/10.48550/arXiv.2005.04988" target="_blank">https://doi.org/10.48550/arXiv.2005.04988</a>,
2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib34"><label>Pulkkinen et al.(2019)Pulkkinen, Nerini, Pérez Hortal,
Velasco-Forero, Seed, Germann, and Foresti</label><mixed-citation>
      
Pulkkinen, S., Nerini, D., Pérez Hortal, A. A., Velasco-Forero, C., Seed, A., Germann, U., and Foresti, L.: Pysteps: an open-source Python library for probabilistic precipitation nowcasting (v1.0), Geosci. Model Dev., 12, 4185–4219, <a href="https://doi.org/10.5194/gmd-12-4185-2019" target="_blank">https://doi.org/10.5194/gmd-12-4185-2019</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib35"><label>Pulkkinen et al.(2021)Pulkkinen, Chandrasekar, and
Niemi</label><mixed-citation>
      
Pulkkinen, S., Chandrasekar, V., and Niemi, T.: Lagrangian
Integro-Difference Equation Model for Precipitation Nowcasting,
J.  Atmos. Ocean. Tech., 38, 2125–2145,
<a href="https://doi.org/10.1175/JTECH-D-21-0013.1" target="_blank">https://doi.org/10.1175/JTECH-D-21-0013.1</a>,  2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib36"><label>Radhakrishnan and Chandrasekar(2020)</label><mixed-citation>
      
Radhakrishnan, C. and Chandrasekar, V.: CASA Prediction System over
Dallas–Fort Worth Urban Network: Blending of Nowcasting and
High-Resolution Numerical Weather Prediction Model, J.
Atmos. Ocean. Tech., 37, 211–228,
<a href="https://doi.org/10.1175/JTECH-D-18-0192.1" target="_blank">https://doi.org/10.1175/JTECH-D-18-0192.1</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib37"><label>Ravuri et al.(2021)Ravuri, Lenc, Willson, Kangin, Lam, Mirowski,
Fitzsimons, Athanassiadou, Kashem, Madge, Prudden, Mandhane, Clark, Brock,
Simonyan, Hadsell, Robinson, Clancy, Arribas, and
Mohamed</label><mixed-citation>
      
Ravuri, S., Lenc, K., Willson, M., Kangin, D., Lam, R., Mirowski, P.,
Fitzsimons, M., Athanassiadou, M., Kashem, S., Madge, S., Prudden, R.,
Mandhane, A., Clark, A., Brock, A., Simonyan, K., Hadsell, R., Robinson, N.,
Clancy, E., Arribas, A., and Mohamed, S.: Skilful precipitation nowcasting
using deep generative models of radar, Nature, 597, 672–677,
<a href="https://doi.org/10.1038/s41586-021-03854-z" target="_blank">https://doi.org/10.1038/s41586-021-03854-z</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib38"><label>Ritvanen et al.(2023)Ritvanen, Harnist, Aldana, Mäkinen, and
Pulkkinen</label><mixed-citation>
      
Ritvanen, J., Harnist, B., Aldana, M., Mäkinen, T., and Pulkkinen, S.:
Advection-Free Convolutional Neural Network for Convective
Rainfall Nowcasting, IEEE J. Sel. Top. Appl.,   1–16,
<a href="https://doi.org/10.1109/JSTARS.2023.3238016" target="_blank">https://doi.org/10.1109/JSTARS.2023.3238016</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib39"><label>Ruzanski and Chandrasekar(2011)</label><mixed-citation>
      
Ruzanski, E. and Chandrasekar, V.: Scale Filtering for Improved
Nowcasting Performance in a High-Resolution X-Band Radar
Network, IEEE T. Geosci. Remote, 49,
2296–2307, <a href="https://doi.org/10.1109/TGRS.2010.2103946" target="_blank">https://doi.org/10.1109/TGRS.2010.2103946</a>, 2011.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib40"><label>Schaefer(1990)</label><mixed-citation>
      
Schaefer, J. T.: The Critical Success Index as an Indicator of
Warning Skill, Weather  Forecast., 5, 570–575,
<a href="https://doi.org/10.1175/1520-0434(1990)005&lt;0570:TCSIAA&gt;2.0.CO;2" target="_blank">https://doi.org/10.1175/1520-0434(1990)005&lt;0570:TCSIAA&gt;2.0.CO;2</a>, 1990.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib41"><label>Seed(2003)</label><mixed-citation>
      
Seed, A. W.: A Dynamic and Spatial Scaling Approach to Advection
Forecasting, J. Appl. Meteorol. Clim., 42, 381–388,
<a href="https://doi.org/10.1175/1520-0450(2003)042&lt;0381:ADASSA&gt;2.0.CO;2" target="_blank">https://doi.org/10.1175/1520-0450(2003)042&lt;0381:ADASSA&gt;2.0.CO;2</a>, 2003.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib42"><label>Seed et al.(2013)Seed, Pierce, and Norman</label><mixed-citation>
      
Seed, A. W., Pierce, C. E., and Norman, K.: Formulation and evaluation of a
scale decomposition-based stochastic precipitation nowcast scheme, Water
Resour. Res., 49, 6624–6641, <a href="https://doi.org/10.1002/wrcr.20536" target="_blank">https://doi.org/10.1002/wrcr.20536</a>, 2013.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib43"><label>Shi et al.(2015)Shi, Chen, Wang, Yeung, Wong, and
Woo</label><mixed-citation>
      
Shi, X., Chen, Z., Wang, H., Yeung, D.-Y., Wong, W.-K., and Woo, W.-C.:
Convolutional LSTM Network: a machine learning approach for precipitation
nowcasting, in: Proceedings of the 28th International Conference on
Neural Information Processing Systems – Volume 1, NIPS'15,
MIT Press, Cambridge, MA, USA, 802–810, 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib44"><label>Shi et al.(2017)Shi, Gao, Lausen, Wang, Yeung, Wong, and
Woo</label><mixed-citation>
      
Shi, X., Gao, Z., Lausen, L., Wang, H., Yeung, D.-Y., Wong, W.-K., and Woo,
W.-C.: Deep learning for precipitation nowcasting: a benchmark and a new
model, in: Proceedings of the 31st International Conference on Neural
Information Processing Systems, NIPS'17,   Curran
Associates Inc., Red Hook, NY, USA, 5622–5632, ISBN 978-1-5108-6096-4, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib45"><label>Staniforth and Côté(1991)</label><mixed-citation>
      
Staniforth, A. and Côté, J.: Semi-Lagrangian Integration Schemes for
Atmospheric Models – A Review, Mon. Weather Rev., 119,
2206–2223, <a href="https://doi.org/10.1175/1520-0493(1991)119&lt;2206:SLISFA&gt;2.0.CO;2" target="_blank">https://doi.org/10.1175/1520-0493(1991)119&lt;2206:SLISFA&gt;2.0.CO;2</a>, 1991.


    </mixed-citation></ref-html>
<ref-html id="bib1.bib46"><label>Sun et al.(2014)Sun, Xue, Wilson, Zawadzki, Ballard,
Onvlee-Hooimeyer, Joe, Barker, Li, Golding, Xu, and Pinto</label><mixed-citation>
      
Sun, J., Xue, M., Wilson, J. W., Zawadzki, I., Ballard, S. P.,
Onvlee-Hooimeyer, J., Joe, P., Barker, D. M., Li, P.-W., Golding, B., Xu, M.,
and Pinto, J.: Use of NWP for Nowcasting Convective Precipitation:
Recent Progress and Challenges, B. Am. Meteorol.
Soc., 95, 409–426, <a href="https://doi.org/10.1175/BAMS-D-11-00263.1" target="_blank">https://doi.org/10.1175/BAMS-D-11-00263.1</a>, 2014.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib47"><label>Sønderby et al.(2020)Sønderby, Espeholt, Heek, Dehghani, Oliver,
Salimans, Agrawal, Hickey, and Kalchbrenner</label><mixed-citation>
      
Sønderby, C. K., Espeholt, L., Heek, J., Dehghani, M., Oliver, A., Salimans,
T., Agrawal, S., Hickey, J., and Kalchbrenner, N.: MetNet: A Neural
Weather Model for Precipitation Forecasting, arXiv [preprint],
<a href="https://doi.org/10.48550/arXiv.2003.12140" target="_blank">https://doi.org/10.48550/arXiv.2003.12140</a>,  2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib48"><label>Trebing et al.(2021)Trebing, Stanczyk, and
Mehrkanoon</label><mixed-citation>
      
Trebing, K., Stanczyk, T., and Mehrkanoon, S.: SmaAt-UNet: Precipitation
nowcasting using a small attention-UNet architecture, Pattern Recogn.
Lett., 145, 178–186, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib49"><label>Ulichney(1988)</label><mixed-citation>
      
Ulichney, R.: Dithering with blue noise, P. IEEE, 76, 56–79,
<a href="https://doi.org/10.1109/5.3288" target="_blank">https://doi.org/10.1109/5.3288</a>, 1988.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib50"><label>Valdenegro-Toro and Mori(2022)</label><mixed-citation>
      
Valdenegro-Toro, M. and Mori, D. S.: A Deeper Look into Aleatoric and
Epistemic Uncertainty Disentanglement, in: 2022 IEEE/CVF
Conference on Computer Vision and Pattern Recognition Workshops
(CVPRW),   IEEE, New Orleans, LA, USA, ISBN
978-1-66548-739-9, 1508–1516, <a href="https://doi.org/10.1109/CVPRW56347.2022.00157" target="_blank">https://doi.org/10.1109/CVPRW56347.2022.00157</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib51"><label>Wen et al.(2018)Wen, Vicol, Ba, Tran, and Grosse</label><mixed-citation>
      
Wen, Y., Vicol, P., Ba, J., Tran, D., and Grosse, R. B.: Flipout: Efficient
Pseudo-Independent Weight Perturbations on Mini-Batches, in: 6th
International Conference on Learning Representations, ICLR 2018, Vancouver,
BC, Canada,  30 April–3 May   2018, Conference Track Proceedings,
OpenReview.net, <a href="https://openreview.net/forum?id=rJNpifWAb" target="_blank"/> (last access: 2 May 2024),
2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib52"><label>Wilks(2011)</label><mixed-citation>
      
Wilks, D. S.: Statistical Methods in the Atmospheric Sciences, Academic
Press, 3rd Edn., ISBN 978-0-12-385022-5, 2011.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib53"><label>Woo et al.(2018)Woo, Park, Lee, and Kweon</label><mixed-citation>
      
Woo, S., Park, J., Lee, J.-Y., and Kweon, I. S.: Cbam: Convolutional block
attention module, in: Proceedings of the European conference on computer
vision (ECCV), Munich, Germany, 3–19, <a href="https://doi.org/10.1007/978-3-030-01234-2_1" target="_blank">https://doi.org/10.1007/978-3-030-01234-2_1</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib54"><label>Xu et al.(2022)Xu, Chen, Yang, Yu, and Chen</label><mixed-citation>
      
Xu, L., Chen, N., Yang, C., Yu, H., and Chen, Z.: Quantifying the uncertainty of precipitation forecasting using probabilistic deep learning, Hydrol. Earth Syst. Sci., 26, 2923–2938, <a href="https://doi.org/10.5194/hess-26-2923-2022" target="_blank">https://doi.org/10.5194/hess-26-2923-2022</a>, 2022.

    </mixed-citation></ref-html>--></article>
