<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing with OASIS Tables v3.0 20080202//EN" "https://jats.nlm.nih.gov/nlm-dtd/publishing/3.0/journalpub-oasis3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:oasis="http://docs.oasis-open.org/ns/oasis-exchange/table" xml:lang="en" dtd-version="3.0" article-type="research-article">
  <front>
    <journal-meta><journal-id journal-id-type="publisher">GMD</journal-id><journal-title-group>
    <journal-title>Geoscientific Model Development</journal-title>
    <abbrev-journal-title abbrev-type="publisher">GMD</abbrev-journal-title><abbrev-journal-title abbrev-type="nlm-ta">Geosci. Model Dev.</abbrev-journal-title>
  </journal-title-group><issn pub-type="epub">1991-9603</issn><publisher>
    <publisher-name>Copernicus Publications</publisher-name>
    <publisher-loc>Göttingen, Germany</publisher-loc>
  </publisher></journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.5194/gmd-19-8565-2026</article-id><title-group><article-title>Evolving beyond collapse: an adaptive particle batch smoother for cryospheric data assimilation</article-title><alt-title>AdaPBS for cryospheric data assimilation</alt-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author" equal-contrib="yes" corresp="no" rid="aff1">
          <name><surname>Aalstad</surname><given-names>Kristoffer</given-names></name>
          
        <ext-link>https://orcid.org/0000-0002-2475-3731</ext-link></contrib>
        <contrib contrib-type="author" equal-contrib="yes" corresp="yes" rid="aff2">
          <name><surname>Alonso-González</surname><given-names>Esteban</given-names></name>
          <email>alonsoe@ipe.csic.es</email>
        <ext-link>https://orcid.org/0000-0002-1883-3823</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1">
          <name><surname>Pirk</surname><given-names>Norbert</given-names></name>
          
        <ext-link>https://orcid.org/0000-0002-8137-2329</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1">
          <name><surname>Westermann</surname><given-names>Sebastian</given-names></name>
          
        </contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1">
          <name><surname>Willmes</surname><given-names>Clarissa</given-names></name>
          
        <ext-link>https://orcid.org/0000-0001-9424-1996</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1">
          <name><surname>Yang</surname><given-names>Ruitang</given-names></name>
          
        <ext-link>https://orcid.org/0000-0001-5145-940X</ext-link></contrib>
        <aff id="aff1"><label>1</label><institution>Department of Geosciences, University of Oslo (UiO), Oslo, Norway</institution>
        </aff>
        <aff id="aff2"><label>2</label><institution>Instituto Pirenaico de Ecología, Consejo Superior de Investigaciones Científicas (IPE-CSIC), Jaca, Spain</institution>
        </aff><author-comment content-type="econtrib"><p>These authors contributed equally to this work.</p></author-comment>
      </contrib-group>
      <author-notes><corresp id="corr1">Esteban Alonso-González (alonsoe@ipe.csic.es)</corresp></author-notes><pub-date><day>16</day><month>September</month><year>2026</year></pub-date>
      
      <volume>19</volume>
      <issue>18</issue>
      <fpage>8565</fpage><lpage>8596</lpage>
      <history>
        <date date-type="received"><day>11</day><month>February</month><year>2026</year></date>
           <date date-type="rev-request"><day>23</day><month>April</month><year>2026</year></date>
           <date date-type="rev-recd"><day>20</day><month>July</month><year>2026</year></date>
           <date date-type="accepted"><day>23</day><month>August</month><year>2026</year></date>
      </history>
      <permissions>
        <copyright-statement>Copyright: © 2026 Kristoffer Aalstad et al.</copyright-statement>
        <copyright-year>2026</copyright-year>
      <license license-type="open-access"><license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p></license></permissions><self-uri xlink:href="https://gmd.copernicus.org/articles/19/8565/2026/gmd-19-8565-2026.html">This article is available from https://gmd.copernicus.org/articles/19/8565/2026/gmd-19-8565-2026.html</self-uri><self-uri xlink:href="https://gmd.copernicus.org/articles/19/8565/2026/gmd-19-8565-2026.pdf">The full text article is available as a PDF file from https://gmd.copernicus.org/articles/19/8565/2026/gmd-19-8565-2026.pdf</self-uri>
      <abstract><title>Abstract</title>

      <p id="d2e137">We present a new adaptive particle-based data assimilation scheme for cryospheric applications that leverages promising developments in importance sampling. The proposed approach seeks to combine some of the advantages of two widely used classes of schemes: particle methods and iterative ensemble Kalman methods. Specifically, it extends the Particle Batch Smoother (PBS) that is commonly used in cryospheric data assimilation, with the Adaptive Multiple Importance Sampling algorithm. This adaptive formulation transforms the PBS into an iterative scheme with improved resilience against ensemble collapse and the ability to implement early-stopping strategies. As such, computational cost is automatically adapted to the complexity of the problem at hand, even down to the grid-cell and water year level in distributed multiyear simulations.</p>

      <p id="d2e140">In homage to the schemes that it builds on, we coin this new algorithm the Adaptive Particle Batch Smoother (AdaPBS) and we test it across a range of scenarios. First, we conducted an intercomparison of some of the most commonly used cryospheric data assimilation algorithms using Markov Chain Monte Carlo (MCMC) simulation as a costly gold-standard benchmark in a simplified temperature index model assimilating snow depth observations. We further evaluated AdaPBS by assimilating snow depth observations from the ESMSnowMIP project at <inline-formula><mml:math id="M1" display="inline"><mml:mn mathvariant="normal">6</mml:mn></mml:math></inline-formula> different sites spanning <inline-formula><mml:math id="M2" display="inline"><mml:mn mathvariant="normal">3</mml:mn></mml:math></inline-formula> continents, using an ensemble of simulations generated with the more complex Flexible Snow Model (FSM2). Our results demonstrate that AdaPBS is a robust and reliable tool, outperforming or at least matching the performance of other commonly used algorithms and successfully handling complex cases with dense observational datasets. All experiments were carried out using the open-source Multiple Snow Data Assimilation System (MuSA) toolbox, which now includes AdaPBS and MCMC among the growing list of available cryospheric data assimilation methods. Beyond our cryospheric focus, the scheme has the potential to be applied directly to the closely related fields of land surface and hydrological data assimilation as well as more general geoscientific Bayesian inference problems.</p>
  </abstract>
    
<funding-group>
<award-group id="gs1">
<funding-source>European Space Agency</funding-source>
<award-id>ESA-CCI Research Fellowship (PATCHES project)</award-id>
<award-id>ESA-CCI Research Fellowship (SnowHotspots project)</award-id>
</award-group>
<award-group id="gs2">
<funding-source>Agencia Estatal de Investigación</funding-source>
<award-id>YC2023-044416-I</award-id>
</award-group>
<award-group id="gs3">
<funding-source>European Research Council</funding-source>
<award-id>GLACMASS project #101096057</award-id>
</award-group>
<award-group id="gs4">
<funding-source>Universitetet i Oslo</funding-source>
<award-id>LATICE</award-id>
</award-group>
</funding-group>
</article-meta>
  </front>
<body>
      

<sec id="Ch1.S1" sec-type="intro">
  <label>1</label><title>Introduction</title>
      <p id="d2e166">Billions of people dwell downstream of cryosphere-dominated basins that provide seasonal snowmelt and glacier meltwater as vital freshwater resources <xref ref-type="bibr" rid="bib1.bibx11 bib1.bibx68" id="paren.1"/>. Seasonal snow and glaciers in these cold regions, i.e. mountains and/or high latitudes,  constitute a key natural water storage system <xref ref-type="bibr" rid="bib1.bibx48" id="paren.2"/>, providing vital water resources during spring and summer for drinking water, agriculture, hydropower, and ecosystems. The cryosphere provides additional climate services as our planet's air conditioner <xref ref-type="bibr" rid="bib1.bibx42" id="paren.3"/>, for example by modulating the global energy cycle <xref ref-type="bibr" rid="bib1.bibx108" id="paren.4"/>, preserving considerable amounts of organic carbon in permafrost as opposed to the atmosphere <xref ref-type="bibr" rid="bib1.bibx103" id="paren.5"/>, and storing frozen water in glaciers rather than raising sea levels <xref ref-type="bibr" rid="bib1.bibx111" id="paren.6"/>.  These services are threatened by ongoing anthropogenic global warming <xref ref-type="bibr" rid="bib1.bibx58" id="paren.7"/>, which is amplified in cold regions leading to strong perturbations of the terrestrial cryosphere according to the vast literature summarized in reports by the Intergovernmental Panel on Climate Change <xref ref-type="bibr" rid="bib1.bibx66 bib1.bibx92" id="paren.8"><named-content content-type="pre">e.g.</named-content></xref>. Thus, the development and implementation of improved cryospheric monitoring systems constitute a priority for humanity with numerous downstream scientific and operational applications.</p>
      <p id="d2e196">Due to the harsh conditions and heterogeneity of the cold regions where the cryosphere manifests, it is usually challenging to deploy representative ground-based monitoring networks. In this context, remotely sensed retrievals of surface properties have emerged as a partial solution to help monitor the state of the cryosphere at regional to global scales <xref ref-type="bibr" rid="bib1.bibx49" id="paren.9"/>. Unfortunately, the quest for direct, accurate, and gap-free estimates of key cryospheric states, such as snow mass <xref ref-type="bibr" rid="bib1.bibx32" id="paren.10"/>, is a recalcitrant problem. Remotely sensed information that is retrievable from space is often limited to surface processes that are only indirectly related to the full internal state of the cryosphere, further complicating its study as a dynamic fresh water reservoir. As mentioned, space-borne estimates of snow mass, also known as snow water equivalent (SWE), in complex terrain remain elusive leading to a big gap between what satellites observe and what hydrologists need <xref ref-type="bibr" rid="bib1.bibx32" id="paren.11"/>. Global passive microwave remote sensing products are able to directly retrieve SWE-related information exhibit a pixel resolution on the order of 10 km that is too coarse for many applications <xref ref-type="bibr" rid="bib1.bibx135" id="paren.12"/>. This gap obfuscates the important task of connecting snowpack storage and fluxes in complex terrain with downstream hydrology and land surface models <xref ref-type="bibr" rid="bib1.bibx29" id="paren.13"/>. At the same time, mechanistic numerical modeling has become a powerful tool to simulate the evolution of the snowpack <xref ref-type="bibr" rid="bib1.bibx41" id="paren.14"/> and the other main components of the terrestrial cryosphere in the form of glaciers <xref ref-type="bibr" rid="bib1.bibx111" id="paren.15"/> and permafrost <xref ref-type="bibr" rid="bib1.bibx131" id="paren.16"/>. Nonetheless, numerical models rely on the availability of accurate high resolution meteorological forcings <xref ref-type="bibr" rid="bib1.bibx60" id="paren.17"/>. This is currently not available on the global scale where state-of-the-art products <xref ref-type="bibr" rid="bib1.bibx64" id="paren.18"/> do not resolve key processes or even major terrain features. In addition, all physically-based numerical models rely at least in part on empirical parameterizations with parameters that are generally uncertain and not transferable <xref ref-type="bibr" rid="bib1.bibx74" id="paren.19"/>.</p>
      <p id="d2e233">The shortcomings of observations and models can be greatly minimized by leveraging data assimilation (DA) techniques, with satellite-based cryospheric DA emerging as a particularly promising approach <xref ref-type="bibr" rid="bib1.bibx76 bib1.bibx53 bib1.bibx7 bib1.bibx133 bib1.bibx134" id="paren.20"/>. Using cryospheric DA techniques it is possible to perform uncertainty-aware monitoring of ungauged areas by simultaneously constraining the uncertainty in satellite observations and simulations, to get the best of both worlds. Data assimilation is a promising way to infer uncertain parameters and update model states <xref ref-type="bibr" rid="bib1.bibx105 bib1.bibx44" id="paren.21"/>, with many applications in the development of reanalysis <xref ref-type="bibr" rid="bib1.bibx86 bib1.bibx64 bib1.bibx120" id="paren.22"/> and the implementation of operational forecasting systems <xref ref-type="bibr" rid="bib1.bibx84 bib1.bibx18 bib1.bibx126" id="paren.23"/>. Moreover, data assimilation allows us to leverage global satellite data provided by space agencies <xref ref-type="bibr" rid="bib1.bibx1 bib1.bibx5" id="paren.24"/>, making the most of the long-term information from the existing climate data record while also offering the potential to ingest emerging satellite retrievals <xref ref-type="bibr" rid="bib1.bibx25 bib1.bibx89" id="paren.25"/> in a physically consistent way.</p>
      <p id="d2e255">Various ensemble-based (also known as Monte Carlo) cryospheric DA algorithms have been proposed for this purpose <xref ref-type="bibr" rid="bib1.bibx7" id="paren.26"><named-content content-type="pre">see e.g.</named-content></xref>. Two main families of ensemble-based schemes have been deployed to date in cryospheric DA: ensemble Kalman methods and particle methods. Of late, the latter particle-based methods have become especially popular in cryospheric applications <xref ref-type="bibr" rid="bib1.bibx78 bib1.bibx85 bib1.bibx19 bib1.bibx96 bib1.bibx84 bib1.bibx27 bib1.bibx101 bib1.bibx46 bib1.bibx116 bib1.bibx81 bib1.bibx6 bib1.bibx24 bib1.bibx75 bib1.bibx54 bib1.bibx99 bib1.bibx120 bib1.bibx17" id="paren.27"/>. This is due to their relatively simple implementation when used in their most basic `bootstrap' form <xref ref-type="bibr" rid="bib1.bibx57 bib1.bibx113" id="paren.28"/>, their ease of interpretation, and the few assumptions they require <xref ref-type="bibr" rid="bib1.bibx76" id="paren.29"/>. However, issues with their practical implementation can make these methods problematic in certain settings. It is well known that if the proposal distribution, typically the prior, differs significantly from the target posterior distribution <xref ref-type="bibr" rid="bib1.bibx83 bib1.bibx104" id="paren.30"/>, then all probability mass risks collapsing to a few or even a single particle, a problem known as particle degeneracy or ensemble collapse <xref ref-type="bibr" rid="bib1.bibx117 bib1.bibx94 bib1.bibx95" id="paren.31"/>. Degeneracy results in suboptimal approximate Bayesian inference by greatly degrading uncertainty quantification in the posterior simulations. This issue is aggravated by incorporating highly informative (i.e., numerous and/or very precise) observations, as well as when working in higher dimensional state and parameter spaces <xref ref-type="bibr" rid="bib1.bibx126" id="paren.32"/>. Drawing on <xref ref-type="bibr" rid="bib1.bibx83" id="text.33"/>, a useful analogy is looking for a needle (target posterior) in a haystack (prior proposal): a larger haystack (broader, higher-dimensionality) or a smaller needle (bigger data, more informative observations) exponentially reduces the chance of hitting the target needle when probing the haystack. While this needle in a haystack effect is a general problem, particle methods are especially vulnerable to it due to high Monte Carlo variance induced by the mismatch between the proposal and target <xref ref-type="bibr" rid="bib1.bibx83 bib1.bibx104" id="paren.34"/>, challenges with localization <xref ref-type="bibr" rid="bib1.bibx45" id="paren.35"/>, and generally poor scaling with high effective dimensions <xref ref-type="bibr" rid="bib1.bibx117 bib1.bibx94" id="paren.36"/>.</p>
      <p id="d2e295">Ensemble Kalman methods have demonstrated high resistance to collapse, even in high-dimensional systems. However, this is achieved by relying on the strong underlying assumption (although not a strict requirement) that all distributions involved in the analysis are Gaussian and that both the dynamical and observational models are linear for optimal inference <xref ref-type="bibr" rid="bib1.bibx44" id="paren.37"/>. These conditions are typically not met in cryospheric DA problems in particular <xref ref-type="bibr" rid="bib1.bibx76" id="paren.38"/> or more generally in geosciences <xref ref-type="bibr" rid="bib1.bibx18" id="paren.39"/>, resulting in these methods often being (arguably prematurely) discarded out of hand in many settings. To mitigate the linear assumption, the use of multiple data assimilation (MDA) iterations that allows a more progressive transition from prior to posterior has emerged as a promising enhancement of ensemble Kalman methods <xref ref-type="bibr" rid="bib1.bibx39 bib1.bibx1 bib1.bibx7 bib1.bibx59" id="paren.40"/>. Moreover, this relaxation of the linear assumption via iteration also collaterally permits workarounds to soften the Gaussian assumption. In particular, transformation techniques <xref ref-type="bibr" rid="bib1.bibx50" id="paren.41"/>, such as Gaussian anamorphosis <xref ref-type="bibr" rid="bib1.bibx12 bib1.bibx18 bib1.bibx1" id="paren.42"/>, allow the ensemble Kalman update to occur in a transformed Gaussian space rather than in the possibly bounded model space at the cost of an additional non-linear transform.</p>
      <p id="d2e317">Previous work has demonstrated the potential of MDA-based iterative ensemble Kalman methods by showing that they can outperform or at least match other computationally tractable Monte Carlo algorithms in various comparisons <xref ref-type="bibr" rid="bib1.bibx1 bib1.bibx7 bib1.bibx102 bib1.bibx71" id="paren.43"/>. The efficacy of this iterative ensemble Kalman approach has also been demonstrated in higher dimensional spatio-temporal cryospheric DA problems <xref ref-type="bibr" rid="bib1.bibx8 bib1.bibx89 bib1.bibx9" id="paren.44"/> and in other complex non-linear and high dimensional inference problems such as gradient-free training of deep neural networks <xref ref-type="bibr" rid="bib1.bibx103" id="paren.45"/>. Despite the clear benefits of iterative ensemble Kalman methods, there are some issues that need to be considered. Although the underlying linear and Gaussian assumptions can be strongly relaxed particularly in the limit of a large number of MDA iterations, this does not mean that they do not hold the potential of affecting the results without incurring a considerable computational cost. This makes the number of iterations an important hyperparameter, which has to be chosen with caution depending on the complexity of the problem. In addition, in the seminal MDA approach with fixed observation error inflation <xref ref-type="bibr" rid="bib1.bibx39" id="paren.46"/>, the number of MDA iterations must be chosen a priori and it is necessary to perform them all to avoid violating the consistency of Bayesian inference <xref ref-type="bibr" rid="bib1.bibx119 bib1.bibx7 bib1.bibx95" id="paren.47"/>, which complicates the implementation of early stopping strategies. Recent iterative ensemble Kalman schemes can circumvent the need to fix the number of iterations <xref ref-type="bibr" rid="bib1.bibx47 bib1.bibx59" id="paren.48"/>, although the adaptive pseudo-timestep may still require a relatively large number of iterations to converge. Furthermore, with a large number of observations (in the order of thousands), owing to high spatio-temporal density and/or large DA windows, ensemble Kalman-based methods can be  more costly than particle-based methods due to the large linear algebra operations involved in the computation of the Kalman gain in the analysis <xref ref-type="bibr" rid="bib1.bibx44" id="paren.49"/>.</p>
      <p id="d2e342">In this paper, we explore the potential to overcome the shortcomings of particle-based methods through the use of iterations. In doing so, we have been inspired by combining ideas from several established algorithms, namely the aforementioned PBS <xref ref-type="bibr" rid="bib1.bibx85" id="paren.50"/> and iterative ensemble Kalman methods <xref ref-type="bibr" rid="bib1.bibx39 bib1.bibx119" id="paren.51"/>, to build a new cryospheric data assimilation method based on developments in adaptive importance sampling <xref ref-type="bibr" rid="bib1.bibx26 bib1.bibx15" id="paren.52"/>. This new iterative and adaptive particle-based method that we coin the adaptive PBS (AdaPBS) has the potential to evolve beyond collapse unlike traditional particle-based methods, while making use of fewer assumptions than ensemble Kalman-based methods and allowing the implementation of early stopping strategies that save substantial computational cost. In concurrent studies, we also demonstrate successful applications of AdaPBS to challenging cryospheric data assimilation problems related to glacier <xref ref-type="bibr" rid="bib1.bibx134" id="paren.53"/> and permafrost <xref ref-type="bibr" rid="bib1.bibx133" id="paren.54"/> modeling. Our contribution here is devoted to describing this new scheme in detail and benchmarking it against existing schemes in several snow data assimilation experiments. By performing these experiments in the open source MuSA snow data assimilation toolbox <xref ref-type="bibr" rid="bib1.bibx7" id="paren.55"/>, a working AdaPBS Python code implementation is made available to the cryospheric community to freely use and remix <xref ref-type="bibr" rid="bib1.bibx3" id="paren.56"/>. In the following, we will outline the relevant theory of Bayesian cryospheric data assimilation in the context of this new adaptive particle method. Subsequently, we perform algorithm benchmarks in different scenarios of varying difficulty with different models to demonstrate the potential of the algorithm.</p>
</sec>
<sec id="Ch1.S2">
  <label>2</label><title>Cryospheric data assimilation</title>
<sec id="Ch1.S2.SS1">
  <label>2.1</label><title>Bayesian inference</title>
      <p id="d2e382">Data assimilation, loosely the fusion of data and models, can be formalized as the application of Bayesian inference <xref ref-type="bibr" rid="bib1.bibx132" id="paren.57"/>. As such, we begin the section by briefly reviewing the key ideas behind Bayesian inference that is at the core of most modern DA schemes, including those explored herein. We refer the reader to the comprehensive texts of <xref ref-type="bibr" rid="bib1.bibx50" id="text.58"/>, <xref ref-type="bibr" rid="bib1.bibx113" id="text.59"/>, <xref ref-type="bibr" rid="bib1.bibx95" id="text.60"/> and <xref ref-type="bibr" rid="bib1.bibx44" id="text.61"/> for a thorough Bayesian treatment from the perspectives of statistics, applied mathematics, machine learning and geophysical DA, respectively. A more thorough analysis of the connection between formal Bayesian inference and practical cryospheric DA is provided in <xref ref-type="bibr" rid="bib1.bibx7" id="text.62"/>.</p>
      <p id="d2e404">In essence, Bayesian inference can be viewed as updating beliefs about some uncertain (also known as random) variables <inline-formula><mml:math id="M3" display="inline"><mml:mi mathvariant="bold-italic">θ</mml:mi></mml:math></inline-formula> given some data <inline-formula><mml:math id="M4" display="inline"><mml:mrow><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>∈</mml:mo><mml:msup><mml:mi mathvariant="double-struck">R</mml:mi><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">o</mml:mi></mml:msub></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>. The uncertain variables in the vector <inline-formula><mml:math id="M5" display="inline"><mml:mi mathvariant="bold-italic">θ</mml:mi></mml:math></inline-formula> can generally be of any form. Here, without loss of generality, we will restrict our attention to continuous model parameters <inline-formula><mml:math id="M6" display="inline"><mml:mrow><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>∈</mml:mo><mml:msup><mml:mi mathvariant="double-struck">R</mml:mi><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">p</mml:mi></mml:msub></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>. Belief updating can then be achieved by combining the basic sum and product rules of probability <xref ref-type="bibr" rid="bib1.bibx69 bib1.bibx83" id="paren.63"/> to obtain Bayes' rule

            <disp-formula id="Ch1.E1" content-type="numbered"><label>1</label><mml:math id="M7" display="block"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mfrac></mml:mstyle><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

          which states that the posterior belief after conditioning on the data, <inline-formula><mml:math id="M8" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, is proportional to the product of the likelihood, <inline-formula><mml:math id="M9" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, which is loosely speaking what the data tell us, and the prior, <inline-formula><mml:math id="M10" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, which encodes what we believed about <inline-formula><mml:math id="M11" display="inline"><mml:mi mathvariant="bold-italic">θ</mml:mi></mml:math></inline-formula> before considering the data. For the model evidence term <inline-formula><mml:math id="M12" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> in the denominator of Eq. (<xref ref-type="disp-formula" rid="Ch1.E1"/>) we again combine the product and sum rules to see that

            <disp-formula id="Ch1.E2" content-type="numbered"><label>2</label><mml:math id="M13" display="block"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mo movablelimits="false">∫</mml:mo><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mi mathvariant="normal">d</mml:mi><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>=</mml:mo><mml:mo movablelimits="false">∫</mml:mo><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mi mathvariant="normal">d</mml:mi><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

          where here and throughout this study the limits of integration are implicitly over the entire support of the integrand. From Eq. (<xref ref-type="disp-formula" rid="Ch1.E2"/>) the evidence is independent of the parameters <inline-formula><mml:math id="M14" display="inline"><mml:mi mathvariant="bold-italic">θ</mml:mi></mml:math></inline-formula> and simply the integral of the numerator of Eq. (<xref ref-type="disp-formula" rid="Ch1.E1"/>) which ensures that the posterior integrates to one over its support. Thus, the evidence can be seen as just a normalizing constant, although it plays an important role as a marginal likelihood for higher levels of inference, since it is implicitly conditioned on the model <xref ref-type="bibr" rid="bib1.bibx83 bib1.bibx95" id="paren.64"/>. In this Bayesian framework, the probability densities <inline-formula><mml:math id="M15" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mo>⋅</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> encode beliefs (i.e. epistemic uncertainties) about their arguments. If the subjective nature of beliefs is unpalatable, it may help to imagine that these are the beliefs held by an abstract numerical agent with a probabilistic model of the world <xref ref-type="bibr" rid="bib1.bibx63" id="paren.65"/>. This probabilistic numerics perspective is instructive for the adaptive methods that we will present here.</p>
      <p id="d2e693">In theory, by inspection of Eq. (<xref ref-type="disp-formula" rid="Ch1.E1"/>), performing Bayesian inference is simply (up to a normalizing constant) multiplying the prior and the likelihood. Naively then, we could just enumerate the posterior through an exhaustive grid approximation provided that we select a sufficiently fine discretization of the variables <inline-formula><mml:math id="M16" display="inline"><mml:mi mathvariant="bold-italic">θ</mml:mi></mml:math></inline-formula> <xref ref-type="bibr" rid="bib1.bibx83" id="paren.66"/>. However, as soon as we have more than a couple of uncertain variables in <inline-formula><mml:math id="M17" display="inline"><mml:mi mathvariant="bold-italic">θ</mml:mi></mml:math></inline-formula> and/or very informative data in <inline-formula><mml:math id="M18" display="inline"><mml:mi mathvariant="bold-italic">y</mml:mi></mml:math></inline-formula>, this grid approach becomes intractable due to the so-called curse of dimensionality <xref ref-type="bibr" rid="bib1.bibx95" id="paren.67"/>. Recalling and extending the earlier analogy, the task of inferring a possibly multi-modal posterior distribution <inline-formula><mml:math id="M19" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is akin to looking for an unknown number of needles in a multi-dimensional haystack: the larger the support (e.g. dimensionality) of the prior, the larger the haystack, and the more informative (numerous and/or accurate) the observations, the smaller the needles. It is this computational challenge that has spurred the development of sophisticated Bayesian inference algorithms, including ensemble-based (also known as Monte Carlo) algorithms that emerged from statistical physics <xref ref-type="bibr" rid="bib1.bibx83 bib1.bibx110" id="paren.68"/> and have, along with variational methods, enjoyed widespread adoption in geophysical DA <xref ref-type="bibr" rid="bib1.bibx105 bib1.bibx44" id="paren.69"/>. In this study, we restrict our attention to ensemble-based methods since these gradient-free methods are the most widely used in cryospheric DA due to their relative ease of use and comparatively robust uncertainty quantification <xref ref-type="bibr" rid="bib1.bibx7" id="paren.70"><named-content content-type="pre">see</named-content><named-content content-type="post">and references therein</named-content></xref>. At the same time, more research is warranted to explore the ever-growing plethora of inference algorithms <xref ref-type="bibr" rid="bib1.bibx44 bib1.bibx63 bib1.bibx95 bib1.bibx113" id="paren.71"/>, many of which remain relatively untested in cryospheric science and geoscience more generally.</p>
</sec>
<sec id="Ch1.S2.SS2">
  <label>2.2</label><title>Cryospheric inverse problems</title>
      <p id="d2e768">To help concretize the above formalization of cryospheric DA, it is instructive to consider the perspective of inverse modeling <xref ref-type="bibr" rid="bib1.bibx112" id="paren.72"/>. Let <inline-formula><mml:math id="M20" display="inline"><mml:mrow><mml:mi mathvariant="script">G</mml:mi><mml:mo>(</mml:mo><mml:mo>⋅</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> denote our forward (or data generating) model that maps from the uncertain snow model parameters to predicted (modeled) snow observations <inline-formula><mml:math id="M21" display="inline"><mml:mrow><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo stretchy="true" mathvariant="normal">^</mml:mo></mml:mover><mml:mo>∈</mml:mo><mml:msup><mml:mi mathvariant="double-struck">R</mml:mi><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">o</mml:mi></mml:msub></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>, i.e., <inline-formula><mml:math id="M22" display="inline"><mml:mrow><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo stretchy="true" mathvariant="normal">^</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:mi mathvariant="script">G</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. In practice, the forward model combines an observation model for <inline-formula><mml:math id="M23" display="inline"><mml:mrow><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo mathvariant="normal" stretchy="true">^</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:mi mathvariant="script">H</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> with a dynamical model for the full model state trajectory <inline-formula><mml:math id="M24" display="inline"><mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>=</mml:mo><mml:mi mathvariant="script">M</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">ϕ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and a (bounding) parameter transformation step <inline-formula><mml:math id="M25" display="inline"><mml:mrow><mml:mi mathvariant="bold-italic">ϕ</mml:mi><mml:mo>=</mml:mo><mml:mi mathvariant="script">T</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> <xref ref-type="bibr" rid="bib1.bibx50 bib1.bibx7 bib1.bibx102" id="paren.73"/>.  Next, consider the typical case where we are given a set of noisy snow observations <inline-formula><mml:math id="M26" display="inline"><mml:mi mathvariant="bold-italic">y</mml:mi></mml:math></inline-formula> which we assume to be related to some true snow observable <inline-formula><mml:math id="M27" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>⋆</mml:mo></mml:msup></mml:mrow></mml:math></inline-formula> through <inline-formula><mml:math id="M28" display="inline"><mml:mrow><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>=</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>⋆</mml:mo></mml:msup><mml:mo>+</mml:mo><mml:mi mathvariant="bold-italic">ϵ</mml:mi></mml:mrow></mml:math></inline-formula>, where <inline-formula><mml:math id="M29" display="inline"><mml:mi mathvariant="bold-italic">ϵ</mml:mi></mml:math></inline-formula> is the observation error. By making the usual strong constraint (or perfect model) assumption <xref ref-type="bibr" rid="bib1.bibx44" id="paren.74"/> that the forward model maps perfectly (without error) from the parameters to the observed variables, we can define some true parameter set  <inline-formula><mml:math id="M30" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>⋆</mml:mo></mml:msup></mml:mrow></mml:math></inline-formula> to exist such that <inline-formula><mml:math id="M31" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>⋆</mml:mo></mml:msup><mml:mo>=</mml:mo><mml:mi mathvariant="script">G</mml:mi><mml:mo>(</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>⋆</mml:mo></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. In general, these parameters can include internal model parameters, initial conditions, and/or forcing terms. This corresponds to a widely used strong constraint version of the forcing formulation of cryospheric DA, wherein model states are completely determined by this parameter vector and the forward model <xref ref-type="bibr" rid="bib1.bibx7" id="paren.75"/>. We adopt this approach throughout without loss of generality, as the forcing formulation can accommodate a weak constraint with model error <xref ref-type="bibr" rid="bib1.bibx44" id="paren.76"/>.</p>
      <p id="d2e981">Using the definition of the observation error, we can establish the following forward relationship between the noisy observations we are given and some presumed true parameters of interest

            <disp-formula id="Ch1.E3" content-type="numbered"><label>3</label><mml:math id="M32" display="block"><mml:mrow><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>=</mml:mo><mml:mi mathvariant="script">G</mml:mi><mml:mfenced open="(" close=")"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>⋆</mml:mo></mml:msup></mml:mrow></mml:mfenced><mml:mo>+</mml:mo><mml:mi mathvariant="bold-italic">ϵ</mml:mi><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>

          At least conceptually, the task in cryospheric data assimilation can now be cast as somehow <italic>inverting</italic> <inline-formula><mml:math id="M33" display="inline"><mml:mrow><mml:mi mathvariant="script">G</mml:mi><mml:mfenced open="(" close=")"><mml:mo>⋅</mml:mo></mml:mfenced></mml:mrow></mml:math></inline-formula> in Eq. (<xref ref-type="disp-formula" rid="Ch1.E3"/>) to solve for <inline-formula><mml:math id="M34" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>⋆</mml:mo></mml:msup></mml:mrow></mml:math></inline-formula>. Unfortunately, even in an idealized linear and noise-free (<inline-formula><mml:math id="M35" display="inline"><mml:mrow><mml:mi mathvariant="bold-italic">ϵ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula>) case, this is typically an ill-posed problem in that an exact solution <inline-formula><mml:math id="M36" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>⋆</mml:mo></mml:msup></mml:mrow></mml:math></inline-formula> may not exist or be unique <xref ref-type="bibr" rid="bib1.bibx112" id="paren.77"/>. In the more challenging noisy and possibly non-linear practical settings the problem is always ill-posed since we invariably assimilate noisy observations where the observation error is <italic>uncertain</italic>.</p>
      <p id="d2e1067">Due to ill-posedness, the quest for a universally optimal (let alone exact) solution  <inline-formula><mml:math id="M37" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>⋆</mml:mo></mml:msup></mml:mrow></mml:math></inline-formula> is nonsensical since infinitely many solutions <inline-formula><mml:math id="M38" display="inline"><mml:mi mathvariant="bold-italic">θ</mml:mi></mml:math></inline-formula> are typically admissible. A way forwards is to instead seek a probabilistic (i.e. Bayesian) solution to the inverse problem in Eq. (<xref ref-type="disp-formula" rid="Ch1.E3"/>) where we construct an observation error model by treating <inline-formula><mml:math id="M39" display="inline"><mml:mi mathvariant="bold-italic">ϵ</mml:mi></mml:math></inline-formula> as an uncertain variable. For convenience, as is common practice in DA <xref ref-type="bibr" rid="bib1.bibx18" id="paren.78"/>, we assume independent additive zero-mean Gaussian observation errors <inline-formula><mml:math id="M40" display="inline"><mml:mrow><mml:mi mathvariant="bold-italic">ϵ</mml:mi><mml:mo>∼</mml:mo><mml:mi>N</mml:mi><mml:mo>(</mml:mo><mml:mn mathvariant="bold">0</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="bold">R</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> where <inline-formula><mml:math id="M41" display="inline"><mml:mi mathvariant="bold">R</mml:mi></mml:math></inline-formula> is an <inline-formula><mml:math id="M42" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">o</mml:mi></mml:msub><mml:mo>×</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">o</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> diagonal observation error covariance matrix. This assumption helps formulate the likelihood <inline-formula><mml:math id="M43" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, the probability of obtaining the fixed observations <inline-formula><mml:math id="M44" display="inline"><mml:mi mathvariant="bold-italic">y</mml:mi></mml:math></inline-formula> given that a parameter set <inline-formula><mml:math id="M45" display="inline"><mml:mi mathvariant="bold-italic">θ</mml:mi></mml:math></inline-formula> is true, since if <inline-formula><mml:math id="M46" display="inline"><mml:mrow><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>=</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>⋆</mml:mo></mml:msup></mml:mrow></mml:math></inline-formula> (due to the conditional <inline-formula><mml:math id="M47" display="inline"><mml:mrow><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi></mml:mrow></mml:math></inline-formula>) then from Eq. (<xref ref-type="disp-formula" rid="Ch1.E3"/>) <inline-formula><mml:math id="M48" display="inline"><mml:mrow><mml:mi mathvariant="bold-italic">ϵ</mml:mi><mml:mo>=</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>-</mml:mo><mml:mi mathvariant="script">G</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the residual so we apply the assumed observation error model to obtain a Gaussian likelihood <xref ref-type="bibr" rid="bib1.bibx44" id="paren.79"/>

            <disp-formula id="Ch1.E4" content-type="numbered"><label>4</label><mml:math id="M49" display="block"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mi>y</mml:mi></mml:msub><mml:mtext>exp</mml:mtext><mml:mfenced open="(" close=")"><mml:mrow><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:mfrac></mml:mstyle><mml:msup><mml:mfenced close="]" open="["><mml:mrow><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>-</mml:mo><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo mathvariant="normal" stretchy="true">^</mml:mo></mml:mover></mml:mrow></mml:mfenced><mml:mi mathvariant="normal">T</mml:mi></mml:msup><mml:msup><mml:mi mathvariant="bold">R</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup><mml:mfenced open="[" close="]"><mml:mrow><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>-</mml:mo><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo mathvariant="normal" stretchy="true">^</mml:mo></mml:mover></mml:mrow></mml:mfenced></mml:mrow></mml:mfenced><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

          where <inline-formula><mml:math id="M50" display="inline"><mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mi>y</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mi mathvariant="normal">det</mml:mi><mml:mo>(</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:mi mathvariant="italic">π</mml:mi><mml:mi mathvariant="bold">R</mml:mi><mml:msup><mml:mo>)</mml:mo><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M51" display="inline"><mml:mrow><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo stretchy="true" mathvariant="normal">^</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:mi mathvariant="script">G</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> denotes the predicted (i.e. modeled) observations given a particular parameter set <inline-formula><mml:math id="M52" display="inline"><mml:mi mathvariant="bold-italic">θ</mml:mi></mml:math></inline-formula> whereby diagnosing the likelihood in Eq. (<xref ref-type="disp-formula" rid="Ch1.E4"/>) requires point-wise evaluations of the forward model to evaluate the residual. In accordance with the likelihood principle, <inline-formula><mml:math id="M53" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> should be viewed as a function of the uncertain parameters <inline-formula><mml:math id="M54" display="inline"><mml:mi mathvariant="bold-italic">θ</mml:mi></mml:math></inline-formula> rather than the fixed (albeit noisy) observations <inline-formula><mml:math id="M55" display="inline"><mml:mi mathvariant="bold-italic">y</mml:mi></mml:math></inline-formula> that we are assimilating <xref ref-type="bibr" rid="bib1.bibx83" id="paren.80"/>. If we now combine this likelihood with a regularizing prior <inline-formula><mml:math id="M56" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> that encodes initial beliefs concerning the parameters, then in principle the full probabilistic solution to the inverse problem Eq. (<xref ref-type="disp-formula" rid="Ch1.E3"/>) is obtained by inferring the posterior through Bayes' rule in Eq. (<xref ref-type="disp-formula" rid="Ch1.E1"/>). As noted by <xref ref-type="bibr" rid="bib1.bibx44" id="text.81"/> this process of Bayesian inference is just point-wise multiplication that does not in itself involve any explicit (matrix or function) inversion, but it nonetheless offers a probabilistic framework for solving general geophysical inverse problems <xref ref-type="bibr" rid="bib1.bibx112" id="paren.82"/>. Although distinctions are sometimes made between DA and inverse modeling, the two fields are highly complementary and unified under the umbrella of Bayesian inference <xref ref-type="bibr" rid="bib1.bibx105 bib1.bibx44 bib1.bibx112" id="paren.83"/>. Thereby, the process of cryospheric DA can be viewed as the solution to dynamical (time-varying) cryospheric inverse problems. The Bayesian dynamics then determines if one is solving a filtering or more general smoothing problem <xref ref-type="bibr" rid="bib1.bibx7 bib1.bibx113" id="paren.84"/>, either way the solutions are typically obtained via numerical approximation of Bayesian inference which is our focus.</p>
</sec>
<sec id="Ch1.S2.SS3">
  <label>2.3</label><title>Markov chain Monte Carlo</title>
      <p id="d2e1448">Markov Chain Monte Carlo (MCMC) methods, which are asymptotically exact, are widely considered the reference computational tool for approximating Bayesian posteriors in practice <xref ref-type="bibr" rid="bib1.bibx50 bib1.bibx95" id="paren.85"/>. Here we provide a brief overview of MCMC since we will use it as a gold-standard computational benchmark against which to gauge the performance of other cryospheric DA methods <xref ref-type="bibr" rid="bib1.bibx77" id="paren.86"/>. As expounded in <xref ref-type="bibr" rid="bib1.bibx110" id="text.87"/>, MCMC originated with the seminal physics paper of <xref ref-type="bibr" rid="bib1.bibx93" id="text.88"/>, whose results were later generalized to statistics by <xref ref-type="bibr" rid="bib1.bibx62" id="text.89"/>, introducing the Random Walk Metropolis (RWM) algorithm that is the ancestor of modern MCMC methods, among which gradient-based methods such as Hamiltonian Monte Carlo <xref ref-type="bibr" rid="bib1.bibx83 bib1.bibx98" id="paren.90"/> are arguably the state-of-the-art. All MCMC schemes construct Markov chains so as to asymptotically sample from a target distribution of interest, where the natural choice for Bayesian DA is the posterior <inline-formula><mml:math id="M57" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. More specifically, MCMC algorithms take sequential Markovian (memoryless) steps through the parameter space and probabilistically either accept or reject the newly proposed step by comparing its posterior density to that of the current step, eventually obtaining samples from the posterior distribution in Eq. (<xref ref-type="disp-formula" rid="Ch1.E1"/>).</p>
      <p id="d2e1490">We base our implementation here on the aforementioned RWM algorithm which remains arguably the most archetypal MCMC method. In this approach, a new step <inline-formula><mml:math id="M58" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mo>*</mml:mo></mml:msubsup></mml:mrow></mml:math></inline-formula> in the chain is proposed by randomly drawing from a proposal distribution <inline-formula><mml:math id="M59" display="inline"><mml:mrow><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mo>*</mml:mo></mml:msubsup><mml:mo>∣</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> that is conditioned (usually by centering) on the current step <inline-formula><mml:math id="M60" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. This proposed step is probabilistically accepted based on the value of the acceptance ratio

            <disp-formula id="Ch1.E5" content-type="numbered"><label>5</label><mml:math id="M61" display="block"><mml:mrow><mml:msub><mml:mi mathvariant="italic">α</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>∣</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mo>*</mml:mo></mml:msubsup><mml:mo>)</mml:mo><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mo>*</mml:mo></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>∣</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mfrac></mml:mstyle><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

          on the condition that the acceptance rule <inline-formula><mml:math id="M62" display="inline"><mml:mrow><mml:msub><mml:mi>u</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>≤</mml:mo><mml:mi mathvariant="normal">min</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="italic">α</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> holds, where <inline-formula><mml:math id="M63" display="inline"><mml:mrow><mml:msub><mml:mi>u</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>∼</mml:mo><mml:mi>U</mml:mi><mml:mo>(</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is a realization of a random variable that is uniformly distributed between <inline-formula><mml:math id="M64" display="inline"><mml:mn mathvariant="normal">0</mml:mn></mml:math></inline-formula> and <inline-formula><mml:math id="M65" display="inline"><mml:mn mathvariant="normal">1</mml:mn></mml:math></inline-formula>, otherwise it is rejected. The evidence <inline-formula><mml:math id="M66" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, which appears as a constant in the posterior density Eq. (<xref ref-type="disp-formula" rid="Ch1.E1"/>) for both the current and the proposed step, cancels out in the acceptance ratio, so we do not need to estimate this intractable quantity. This, together with asymptotic guarantees, helps explain the historical popularity of MCMC sampling.  The more general form of Eq. (<xref ref-type="disp-formula" rid="Ch1.E5"/>) introduced by <xref ref-type="bibr" rid="bib1.bibx62" id="text.91"/> also involves proposal densities, but these cancel out in RWM and were hence omitted here. Upon acceptance, the chain moves to the proposed point in parameter space such that <inline-formula><mml:math id="M67" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mo>*</mml:mo></mml:msubsup></mml:mrow></mml:math></inline-formula>. Upon rejection, the chain stays at the current point  <inline-formula><mml:math id="M68" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. Note that a proposed point will always be accepted if it has a higher posterior density since then <inline-formula><mml:math id="M69" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">α</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> so the acceptance rule always holds. Moreover, the proposed point <inline-formula><mml:math id="M70" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mo>*</mml:mo></mml:msubsup></mml:mrow></mml:math></inline-formula> will always have a non-zero probability of being accepted. This probabilistic acceptance rule helps ensure that the chain will asymptotically (<inline-formula><mml:math id="M71" display="inline"><mml:mrow><mml:mi>i</mml:mi><mml:mo>→</mml:mo><mml:mi mathvariant="normal">∞</mml:mi></mml:mrow></mml:math></inline-formula>) sample from the posterior. To sample from the posterior with MCMC, one needs to run the Markov chain for enough iterations to ensure that it mixes properly and converges to sampling from the posterior. In practice, although there are some diagnostics that can be used as a guide <xref ref-type="bibr" rid="bib1.bibx50" id="paren.92"/>, what constitutes enough iterations is uncertain a priori and can remain challenging to verify post hoc. As such, MCMC is typically run for a very large number (i.e., tens of thousands) of iterations to ensure convergence <xref ref-type="bibr" rid="bib1.bibx23" id="paren.93"/>. An initial part of the chain is discarded as a burn-in period to avoid the biasing effects of the chain initialization.  Due to the Markov property, where the next step only depends on the current step, it is also clear that the subsequent iterations in the chain will be auto-correlated rather than independent. As such, the remaining part of the chain can be subsampled uniformly at random to obtain more independent samples from the posterior distribution.</p>
      <p id="d2e1859">The most standard RWM approach employs a multivariate normal (Gaussian) proposal of the form <inline-formula><mml:math id="M72" display="inline"><mml:mrow><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">ϕ</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mo>*</mml:mo></mml:msubsup><mml:mo>∣</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">ϕ</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mi>N</mml:mi><mml:mo>(</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">ϕ</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mo>*</mml:mo></mml:msubsup><mml:mo>∣</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">ϕ</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold">Σ</mml:mi><mml:mi>q</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> where the mean is the current point <inline-formula><mml:math id="M73" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">ϕ</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M74" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">Σ</mml:mi><mml:mi>q</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the proposal covariance matrix. To reduce the number of tuning parameters in the proposal, the latter can be made isotropic by using a scalar (constant diagonal) matrix of the form <inline-formula><mml:math id="M75" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">Σ</mml:mi><mml:mi>q</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msubsup><mml:mi mathvariant="italic">σ</mml:mi><mml:mi>q</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mi mathvariant="bold">I</mml:mi></mml:mrow></mml:math></inline-formula> where <inline-formula><mml:math id="M76" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="italic">σ</mml:mi><mml:mi>q</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:math></inline-formula> is the proposal variance and <inline-formula><mml:math id="M77" display="inline"><mml:mi mathvariant="bold">I</mml:mi></mml:math></inline-formula> is an identity matrix. In practice, good mixing of the RWM method requires judicious hand-tuning of <inline-formula><mml:math id="M78" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi>q</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, and even then the use of an isotropic covariance matrix remains wasteful to efficiently explore higher dimensional parameter spaces given that the posterior is often anisotropic <xref ref-type="bibr" rid="bib1.bibx83" id="paren.94"/>. A relatively simple workaround is to employ an adaptive RWM method that automatically modifies a generally anisotropic proposal covariance on the fly using the history of the evolving Markov chain to obtain better mixing properties. Among the existing adaptive MCMC algorithms, here we chose to employ the robust adaptive Metropolis (RAM) method proposed by <xref ref-type="bibr" rid="bib1.bibx127" id="text.95"/> using the hyperparameters suggested therein with <inline-formula><mml:math id="M79" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">s</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">20</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mn mathvariant="normal">3</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> steps and discarding the first <inline-formula><mml:math id="M80" display="inline"><mml:mn mathvariant="normal">10</mml:mn></mml:math></inline-formula>  % as a burn-in phase <xref ref-type="bibr" rid="bib1.bibx95" id="paren.96"/>. The RAM method was chosen in-lieu of even more sophisticated MCMC methods such as Hamiltonian Monte Carlo <xref ref-type="bibr" rid="bib1.bibx98" id="paren.97"/> that may have mixed even faster since the RAM method is considerably easier to implement. Moreover, by running RAM for tens of thousands of iterations we can be fairly confident that it converges and represents a strong benchmark. A crucial point here is that it is likely that most if not all MCMC methods are too computationally expensive to deploy at scale (i.e., for large areas) in practical cryospheric DA, although further research into applying more sophisticated MCMC schemes <xref ref-type="bibr" rid="bib1.bibx95" id="paren.98"/> is needed. Herein, MCMC in the form om RAM is used as a reference standard against which we gauge the performance of the more tractable ensemble-based cryospheric DA schemes <xref ref-type="bibr" rid="bib1.bibx77" id="paren.99"/>, especially the adaptive particle smoother that is the focus of this study.</p>
</sec>
<sec id="Ch1.S2.SS4">
  <label>2.4</label><title>Ensemble Kalman methods</title>
      <p id="d2e2060">Ensemble Kalman methods, introduced by <xref ref-type="bibr" rid="bib1.bibx43" id="text.100"/> and described in detail in <xref ref-type="bibr" rid="bib1.bibx44" id="text.101"/>, are ensemble-based extensions of the classic Kalman methods <xref ref-type="bibr" rid="bib1.bibx70" id="paren.102"/> which provide exact Bayesian inference methods for linear Gaussian models <xref ref-type="bibr" rid="bib1.bibx113" id="paren.103"/>. In such models the mapping from hidden states and/or parameters to observations is linear while both the prior and the likelihood are Gaussian. Ensemble Kalman methods help to relax these assumptions in the sense that approximate yet efficient inference is still possible when they are violated, making them applicable also to non-linear geophysical problems <xref ref-type="bibr" rid="bib1.bibx44" id="paren.104"/>. Although other non-linear variations on classic Kalman methods also exist <xref ref-type="bibr" rid="bib1.bibx113" id="paren.105"/>, the ease of implementation and the robust performance of ensemble Kalman methods have made them highly applicable for geophysical DA in general <xref ref-type="bibr" rid="bib1.bibx18" id="paren.106"/> as well as cryospheric DA in particular <xref ref-type="bibr" rid="bib1.bibx7" id="paren.107"><named-content content-type="pre">see</named-content><named-content content-type="post">and references therein</named-content></xref>. Crucially, the more recent development of iterative ensemble Kalman methods <xref ref-type="bibr" rid="bib1.bibx39 bib1.bibx47" id="paren.108"/> have helped to enhance the capability of this class of DA schemes for highly non-linear and/or complex problems <xref ref-type="bibr" rid="bib1.bibx102 bib1.bibx103 bib1.bibx71" id="paren.109"><named-content content-type="pre">e.g.</named-content></xref>.</p>
      <p id="d2e2100">In practical cryospheric DA, several numerical experiments have previously shown that iterative ensemble Kalman methods can outperform basic particle methods while maintaining more robust posterior uncertainty quantification <xref ref-type="bibr" rid="bib1.bibx1 bib1.bibx7" id="paren.110"/>. It is thus instructive to also include such experiments here to provide an additional benchmark for the performance of the proposed adaptive particle method. Rather than providing a high-fidelity benchmark like the MCMC experiments, the more computationally tractable iterative ensemble Kalman experiments should be seen as a more practical operational benchmark for cryospheric DA. For the sake of completeness, the rest of this section provides a brief overview on the implementation of iterative ensemble Kalman methods. We refer to <xref ref-type="bibr" rid="bib1.bibx44" id="text.111"/> and <xref ref-type="bibr" rid="bib1.bibx7" id="text.112"/> for more details on the theory of ensemble Kalman methods and their implementation for cryospheric DA, respectively.</p>
      <p id="d2e2112">Here we adopt a specific iterative ensemble Kalman method known as the ensemble smoother with multiple data assimilation <xref ref-type="bibr" rid="bib1.bibx39" id="paren.113"><named-content content-type="pre">ES-MDA;</named-content></xref> due to its relative ease of implementation and robust performance <xref ref-type="bibr" rid="bib1.bibx1 bib1.bibx7 bib1.bibx8" id="paren.114"/>. We note in passing that other promising variations on the iterative ensemble Kalman method exist <xref ref-type="bibr" rid="bib1.bibx47 bib1.bibx44" id="paren.115"/> and are worthy of further investigation in cryospheric DA, but we do not expect their performance to differ markedly from the ES-MDA. The ES-MDA scheme is initialized by sampling an initial ensemble of <inline-formula><mml:math id="M81" display="inline"><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> parameter vectors from the prior <inline-formula><mml:math id="M82" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>∼</mml:mo><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and it then proceeds by cycling between a prediction and update step for <inline-formula><mml:math id="M83" display="inline"><mml:mrow><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">a</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> iterations:</p>
      <p id="d2e2201"><list list-type="order">
            <list-item>

      <p id="d2e2206">Run the forward model to obtain the predicted observations <inline-formula><mml:math id="M84" display="inline"><mml:mrow><mml:msubsup><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo stretchy="true" mathvariant="normal">^</mml:mo></mml:mover><mml:mi>i</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mi mathvariant="script">G</mml:mi><mml:mfenced close=")" open="("><mml:mrow><mml:msubsup><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:mfenced></mml:mrow></mml:math></inline-formula> from the state <inline-formula><mml:math id="M85" display="inline"><mml:mi mathvariant="bold-italic">x</mml:mi></mml:math></inline-formula> (see Sect. <xref ref-type="sec" rid="Ch1.S2.SS2"/>).</p>
            </list-item>
            <list-item>

      <p id="d2e2260">If <inline-formula><mml:math id="M86" display="inline"><mml:mrow><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>&lt;</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">a</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, perform a tempered ensemble Kalman update step

                  <disp-formula id="Ch1.E6" content-type="numbered"><label>6</label><mml:math id="M87" display="block"><mml:mrow><mml:msubsup><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:msup><mml:mi mathvariant="bold">K</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mfenced close=")" open="("><mml:mrow><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>-</mml:mo><mml:msubsup><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo stretchy="true" mathvariant="normal">^</mml:mo></mml:mover><mml:mi>i</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:mfenced><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

                where <inline-formula><mml:math id="M88" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold">K</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> is the tempered ensemble Kalman gain for iteration <inline-formula><mml:math id="M89" display="inline"><mml:mi mathvariant="normal">ℓ</mml:mi></mml:math></inline-formula> obtained from (ensemble) covariance matrices <xref ref-type="bibr" rid="bib1.bibx44" id="paren.116"/>.</p>
            </list-item>
          </list>The steps above are implicitly carried out for all <inline-formula><mml:math id="M90" display="inline"><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> ensemble members, with <inline-formula><mml:math id="M91" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">a</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> iterations of the prediction step and <inline-formula><mml:math id="M92" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">a</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> iterations of the update step. The update step itself is essentially free, so the total cost of the ES-MDA is <inline-formula><mml:math id="M93" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">a</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>)</mml:mo><mml:mo>×</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> forward model simulations where it is possible to parallelize across the ensemble dimension in each iteration <inline-formula><mml:math id="M94" display="inline"><mml:mi mathvariant="normal">ℓ</mml:mi></mml:math></inline-formula>. Recall from Sect. <xref ref-type="sec" rid="Ch1.S2.SS2"/> that several steps (observation, dynamics, transformation) are baked into the forward model <inline-formula><mml:math id="M95" display="inline"><mml:mrow><mml:mi mathvariant="script">G</mml:mi><mml:mo>(</mml:mo><mml:mo>⋅</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. Therein, it is the cost of running the dynamical model <inline-formula><mml:math id="M96" display="inline"><mml:mrow><mml:mi mathvariant="script">M</mml:mi><mml:mo>(</mml:mo><mml:mo>⋅</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> to obtain the ensemble of hidden model states <inline-formula><mml:math id="M97" display="inline"><mml:mi mathvariant="bold-italic">x</mml:mi></mml:math></inline-formula> of interest that completely dominates the computational burden of <inline-formula><mml:math id="M98" display="inline"><mml:mrow><mml:mi mathvariant="script">G</mml:mi><mml:mo>(</mml:mo><mml:mo>⋅</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> in step 1. So, while (single chain) MCMC costs tens of thousands of strictly sequential forward model runs, with a typical setting of <inline-formula><mml:math id="M99" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">a</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">4</mml:mn></mml:mrow></mml:math></inline-formula> <xref ref-type="bibr" rid="bib1.bibx1 bib1.bibx102 bib1.bibx7" id="paren.117"/> ES-MDA only incurs a computational cost of <inline-formula><mml:math id="M100" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">a</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>=</mml:mo><mml:mn mathvariant="normal">5</mml:mn></mml:mrow></mml:math></inline-formula> iterations of an ensemble of <inline-formula><mml:math id="M101" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">100</mml:mn></mml:mrow></mml:math></inline-formula> parallelizable forward model runs. With <inline-formula><mml:math id="M102" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">a</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> the ES-MDA reverts to the original (non-iterative) ensemble smoother (ES) scheme proposed by <xref ref-type="bibr" rid="bib1.bibx125" id="text.118"/> which we also include here in the benchmarking of the new AdaPBS method. We emphasize that even this non-iterative ES with <inline-formula><mml:math id="M103" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">a</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> has a cost of <inline-formula><mml:math id="M104" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">a</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>)</mml:mo><mml:mo>×</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">200</mml:mn></mml:mrow></mml:math></inline-formula> model runs as it requires rerunning an ensemble of model simulations with the updated parameters to obtain posterior state predictions.</p>
</sec>
<sec id="Ch1.S2.SS5">
  <label>2.5</label><title>Particle methods</title>
      <p id="d2e2639">Particle methods <xref ref-type="bibr" rid="bib1.bibx124 bib1.bibx126" id="paren.119"/>, also known as sequential Monte Carlo (SMC) <xref ref-type="bibr" rid="bib1.bibx21" id="paren.120"/>, rose to prominence with the work of <xref ref-type="bibr" rid="bib1.bibx57" id="text.121"/> and <xref ref-type="bibr" rid="bib1.bibx72" id="text.122"/> around the same time as ensemble Kalman methods <xref ref-type="bibr" rid="bib1.bibx43" id="paren.123"/>. Moreover, similar inference methods have arguably independently been (re)discovered in geoscience <xref ref-type="bibr" rid="bib1.bibx14 bib1.bibx125 bib1.bibx85" id="paren.124"/>. These particle methods also have roots back to the dawn of Monte Carlo methods in physics <xref ref-type="bibr" rid="bib1.bibx61 bib1.bibx110" id="paren.125"/>. The key inference mechanism that powers particle methods is importance sampling <xref ref-type="bibr" rid="bib1.bibx83" id="paren.126"/>, which weighs the importance of an ensemble of particles (ensemble members) <inline-formula><mml:math id="M105" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> according to their posterior probability density and the density of the proposal. By performing this sequentially in time and resampling the particles based on their weights (i.e., fitness) after each importance sampling step, we recover the algorithm known as sequential importance resampling <xref ref-type="bibr" rid="bib1.bibx114 bib1.bibx31" id="paren.127"><named-content content-type="pre">SIR;</named-content></xref> at the core of particle methods used for Bayesian filtering and smoothing <xref ref-type="bibr" rid="bib1.bibx72" id="paren.128"/> and static inference problems <xref ref-type="bibr" rid="bib1.bibx20" id="paren.129"/>. The cycling of prediction (mutation) and updating (selection) has clear connections with metaheuristic evolutionary methods <xref ref-type="bibr" rid="bib1.bibx16" id="paren.130"/> such as genetic algorithms <xref ref-type="bibr" rid="bib1.bibx67" id="paren.131"/> that can be mathematically formalized as particle methods <xref ref-type="bibr" rid="bib1.bibx30" id="paren.132"/>. This perspective also provides a connection to other evolutionary-inspired Bayesian such as Shuffled Complex Evolution Metropolis <xref ref-type="bibr" rid="bib1.bibx129" id="paren.133"><named-content content-type="pre">SCEM-UA;</named-content></xref> and Differential Evolution Adaptive Metropolis <xref ref-type="bibr" rid="bib1.bibx130" id="paren.134"><named-content content-type="pre">DREAM;</named-content></xref>, that remain state of the art optimizers and samplers, respectively, in hydrology. However, following <xref ref-type="bibr" rid="bib1.bibx118" id="text.135"/> we emphasize that the adaptive particle method proposed herein and our playful title use evolutionary principles as a conceptual model for inspiration and not as a justification. Justification can be found in the literature on the convergence of general SMC methods <xref ref-type="bibr" rid="bib1.bibx21" id="paren.136"/> and more specifically adaptive multiple importance sampling <xref ref-type="bibr" rid="bib1.bibx87" id="paren.137"/>. Having introduced the notion of particle methods, the rest of this section will outline the principle of importance sampling and how this is incorporated into SIR. Together, this basic SIR theory suffices to grasp vanilla particle methods related to the seminal “bootstrap” particle filter of <xref ref-type="bibr" rid="bib1.bibx57" id="text.138"/>. These basic particle methods have become popular approaches to DA in snow science <xref ref-type="bibr" rid="bib1.bibx78 bib1.bibx85 bib1.bibx19 bib1.bibx84 bib1.bibx101 bib1.bibx46 bib1.bibx116 bib1.bibx6 bib1.bibx24 bib1.bibx99 bib1.bibx54 bib1.bibx120" id="paren.139"><named-content content-type="pre">e.g.</named-content></xref> and glaciology <xref ref-type="bibr" rid="bib1.bibx96 bib1.bibx75 bib1.bibx17" id="paren.140"/>.</p>
<sec id="Ch1.S2.SS5.SSS1">
  <label>2.5.1</label><title>Monte Carlo integration</title>
      <p id="d2e2737">Importance sampling is a generalized form of indirect Monte Carlo integration that can be used to estimate expectations with respect to a complex target distribution, in our case the posterior <inline-formula><mml:math id="M106" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, by sampling from a simpler proposal distribution <inline-formula><mml:math id="M107" display="inline"><mml:mrow><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> <xref ref-type="bibr" rid="bib1.bibx83" id="paren.141"/>. To unpack this definition, it can be helpful to recall how basic direct Monte Carlo integration works in the context of estimating expectations. The posterior expectations we wish to estimate are of the form <xref ref-type="bibr" rid="bib1.bibx113" id="paren.142"/>

              <disp-formula id="Ch1.E7" content-type="numbered"><label>7</label><mml:math id="M108" display="block"><mml:mrow><mml:mi>E</mml:mi><mml:mfenced open="[" close="]"><mml:mrow><mml:mi>g</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi></mml:mrow></mml:mfenced><mml:mo>=</mml:mo><mml:mo movablelimits="false">∫</mml:mo><mml:mi>g</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mi mathvariant="normal">d</mml:mi><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

            where we recall that the integral is implicitly over the support of the integrand. Here <inline-formula><mml:math id="M109" display="inline"><mml:mrow><mml:mi>g</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is an arbitrary (possibly vector-valued) function of the parameters <inline-formula><mml:math id="M110" display="inline"><mml:mi mathvariant="bold-italic">θ</mml:mi></mml:math></inline-formula> that we may wish to take the posterior expectation of. For example, using <inline-formula><mml:math id="M111" display="inline"><mml:mrow><mml:mi>g</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi></mml:mrow></mml:math></inline-formula> in Eq. (<xref ref-type="disp-formula" rid="Ch1.E7"/>) yields the posterior mean <inline-formula><mml:math id="M112" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">μ</mml:mi><mml:mo stretchy="true" mathvariant="normal">^</mml:mo></mml:mover><mml:mi mathvariant="bold-italic">θ</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> while using <inline-formula><mml:math id="M113" display="inline"><mml:mrow><mml:mi>g</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>-</mml:mo><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">μ</mml:mi><mml:mo stretchy="true" mathvariant="normal">^</mml:mo></mml:mover><mml:mi mathvariant="bold-italic">θ</mml:mi></mml:msub><mml:msup><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> in Eq. (<xref ref-type="disp-formula" rid="Ch1.E7"/>) yields the posterior variance <inline-formula><mml:math id="M114" display="inline"><mml:mrow><mml:msubsup><mml:mover accent="true"><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mo mathvariant="normal" stretchy="true">^</mml:mo></mml:mover><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:math></inline-formula>. Now suppose we could generate <inline-formula><mml:math id="M115" display="inline"><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> independent samples from the posterior <inline-formula><mml:math id="M116" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>∼</mml:mo><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> then we could obtain a direct Monte Carlo estimate of Eq. (<xref ref-type="disp-formula" rid="Ch1.E7"/>) through the sample mean

              <disp-formula id="Ch1.E8" content-type="numbered"><label>8</label><mml:math id="M117" display="block"><mml:mrow><mml:mi>E</mml:mi><mml:mfenced open="[" close="]"><mml:mrow><mml:mi>g</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi></mml:mrow></mml:mfenced><mml:mo>≃</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:munderover><mml:mi>g</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

            which, thanks to the law of large numbers and the central limit theorem, will be an unbiased estimate that asymptotically (<inline-formula><mml:math id="M118" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub><mml:mo>→</mml:mo><mml:mi mathvariant="normal">∞</mml:mi></mml:mrow></mml:math></inline-formula>) converges to the true posterior expectation in Eq. (<xref ref-type="disp-formula" rid="Ch1.E7"/>) with a Monte Carlo error that decays at a rate <inline-formula><mml:math id="M119" display="inline"><mml:mrow><mml:mi mathvariant="script">O</mml:mi><mml:mo>(</mml:mo><mml:msubsup><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> <xref ref-type="bibr" rid="bib1.bibx21" id="paren.143"/>. Although this error reduction rate is relatively slow <xref ref-type="bibr" rid="bib1.bibx63" id="paren.144"/>, e.g. increasing the ensemble size by a factor <inline-formula><mml:math id="M120" display="inline"><mml:mn mathvariant="normal">100</mml:mn></mml:math></inline-formula> only reduces the error by a factor <inline-formula><mml:math id="M121" display="inline"><mml:mn mathvariant="normal">10</mml:mn></mml:math></inline-formula>, it is independent of dimension which makes these Monte Carlo (also known as ensemble-based) methods a viable (and sometimes the sole) option for challenging inference problems that arise in geophysical DA <xref ref-type="bibr" rid="bib1.bibx18" id="paren.145"/>. Nonetheless, following <xref ref-type="bibr" rid="bib1.bibx63" id="text.146"/>, it is worth remembering that Monte Carlo methods are a last resort in line with the principle of <xref ref-type="bibr" rid="bib1.bibx69" id="text.147"/> that for every randomized method there is usually a better performing deterministic method that requires more thought.</p>
</sec>
<sec id="Ch1.S2.SS5.SSS2">
  <label>2.5.2</label><title>Importance sampling</title>
      <p id="d2e3139">The obvious problem with this method is that we are not able to directly sample independently (let alone efficiently) from the exact posterior distribution. In fact, posterior sampling is often the very problem that we need to solve. Nonetheless, once we have obtained posterior samples we can use Monte Carlo integration to approximate the desired posterior expectations of interest. Slow MCMC methods (Sect. <xref ref-type="sec" rid="Ch1.S2.SS3"/>) are only asymptotically exact samplers that do not provide independent samples. The efficient yet approximate ensemble Kalman methods (Sect. <xref ref-type="sec" rid="Ch1.S2.SS4"/>) are only exact samplers for Gaussian linear models. This is where the more general and indirect Monte Carlo technique of importance sampling shines.</p>
      <p id="d2e3146">Defining a proposal distribution <inline-formula><mml:math id="M122" display="inline"><mml:mrow><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> that we can easily generate independent samples from and multiplying the integrand in Eq. (<xref ref-type="disp-formula" rid="Ch1.E7"/>) by <inline-formula><mml:math id="M123" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>=</mml:mo><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mo>/</mml:mo><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> then clearly

              <disp-formula id="Ch1.E9" content-type="numbered"><label>9</label><mml:math id="M124" display="block"><mml:mrow><mml:mi>E</mml:mi><mml:mfenced close="]" open="["><mml:mrow><mml:mi>g</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi></mml:mrow></mml:mfenced><mml:mo>=</mml:mo><mml:mo movablelimits="false">∫</mml:mo><mml:mi>g</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mfrac></mml:mstyle><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mi mathvariant="normal">d</mml:mi><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

            under the sufficient condition that the support of the proposal encompasses that of the posterior, i.e. <inline-formula><mml:math id="M125" display="inline"><mml:mrow><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula> wherever <inline-formula><mml:math id="M126" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula> <xref ref-type="bibr" rid="bib1.bibx21" id="paren.148"/>. If we now generate independent samples from the proposal  <inline-formula><mml:math id="M127" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>∼</mml:mo><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> we can obtain the following importance sampling estimate of Eq. (<xref ref-type="disp-formula" rid="Ch1.E7"/>) from Eq. (<xref ref-type="disp-formula" rid="Ch1.E9"/>)

              <disp-formula id="Ch1.E10" content-type="numbered"><label>10</label><mml:math id="M128" display="block"><mml:mrow><mml:mi>E</mml:mi><mml:mfenced open="[" close="]"><mml:mrow><mml:mi>g</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi></mml:mrow></mml:mfenced><mml:mo>≃</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:munderover><mml:mi>g</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mfrac></mml:mstyle><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

            where we have great freedom in the choice of the proposal <inline-formula><mml:math id="M129" display="inline"><mml:mrow><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> <xref ref-type="bibr" rid="bib1.bibx124" id="paren.149"/>.</p>
</sec>
<sec id="Ch1.S2.SS5.SSS3">
  <label>2.5.3</label><title>Self-normalized importance sampling</title>
      <p id="d2e3458">Unfortunately we still can not evaluate Eq. (<xref ref-type="disp-formula" rid="Ch1.E10"/>) since the posterior density <inline-formula><mml:math id="M130" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> term defined in Eq. (<xref ref-type="disp-formula" rid="Ch1.E1"/>) is only known up to an <italic>unknown</italic> normalizing constant, namely the evidence. The evidence term <inline-formula><mml:math id="M131" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> in Eq. (<xref ref-type="disp-formula" rid="Ch1.E2"/>) is a typically intractable integral over the unnormalized posterior <inline-formula><mml:math id="M132" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. Nonetheless, comparing Eqs. (<xref ref-type="disp-formula" rid="Ch1.E2"/>) and (<xref ref-type="disp-formula" rid="Ch1.E7"/>) the evidence can be seen as the prior (rather than posterior) expectation of the likelihood by setting <inline-formula><mml:math id="M133" display="inline"><mml:mrow><mml:mi>g</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and replacing the posterior with the prior in Eq. (<xref ref-type="disp-formula" rid="Ch1.E7"/>). Analogously to Eq. (<xref ref-type="disp-formula" rid="Ch1.E9"/>), we can then recast the evidence as an expectation involving the proposal density

              <disp-formula id="Ch1.E11" content-type="numbered"><label>11</label><mml:math id="M134" display="block"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mo movablelimits="false">∫</mml:mo><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mfrac></mml:mstyle><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mi mathvariant="normal">d</mml:mi><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

            so that we can use proposal samples <inline-formula><mml:math id="M135" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>∼</mml:mo><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> to obtain the estimate

              <disp-formula id="Ch1.E12" content-type="numbered"><label>12</label><mml:math id="M136" display="block"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo><mml:mo>≃</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:munderover><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>∣</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mfrac></mml:mstyle><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:munderover><mml:msub><mml:mover accent="true"><mml:mi>w</mml:mi><mml:mo mathvariant="normal" stretchy="true">̃</mml:mo></mml:mover><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

            where we have defined the unnormalized weights <inline-formula><mml:math id="M137" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi>w</mml:mi><mml:mo stretchy="true" mathvariant="normal">̃</mml:mo></mml:mover><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mover accent="true"><mml:mi>w</mml:mi><mml:mo mathvariant="normal" stretchy="true">̃</mml:mo></mml:mover><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>∣</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>/</mml:mo><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> as a useful shorthand. Now we are in a position to revisit the importance sampling estimate of the posterior expectation Eq. (<xref ref-type="disp-formula" rid="Ch1.E10"/>) which, using the definition of the posterior in Eq. (<xref ref-type="disp-formula" rid="Ch1.E1"/>), can be re-written as

              <disp-formula id="Ch1.E13" content-type="numbered"><label>13</label><mml:math id="M138" display="block"><mml:mrow><mml:mi>E</mml:mi><mml:mfenced close="]" open="["><mml:mrow><mml:mi>g</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi></mml:mrow></mml:mfenced><mml:mo>≃</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:munderover><mml:mi>g</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>∣</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mfrac></mml:mstyle><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>

            If we now use the definition of the unnormalized weights <inline-formula><mml:math id="M139" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi>w</mml:mi><mml:mo stretchy="true" mathvariant="normal">̃</mml:mo></mml:mover><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and insert for the approximation of <inline-formula><mml:math id="M140" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> in Eq. (<xref ref-type="disp-formula" rid="Ch1.E12"/>), we obtain the so-called self-normalized importance sampling (SNIS) estimate <xref ref-type="bibr" rid="bib1.bibx104 bib1.bibx95" id="paren.150"/>

              <disp-formula id="Ch1.E14" content-type="numbered"><label>14</label><mml:math id="M141" display="block"><mml:mtable rowspacing="0.2ex" class="split" displaystyle="true" columnalign="right left"><mml:mtr><mml:mtd><mml:mrow><mml:mi>E</mml:mi><mml:mfenced open="[" close="]"><mml:mrow><mml:mi>g</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi></mml:mrow></mml:mfenced></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mo>≃</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:munderover><mml:mi>g</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi>w</mml:mi><mml:mo stretchy="true" mathvariant="normal">̃</mml:mo></mml:mover><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:msubsup><mml:mo>∑</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:msubsup><mml:msub><mml:mover accent="true"><mml:mi>w</mml:mi><mml:mo stretchy="true" mathvariant="normal">̃</mml:mo></mml:mover><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mo>=</mml:mo><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:munderover><mml:mi>g</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>

            where the normalized weights <inline-formula><mml:math id="M142" display="inline"><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> are defined by <inline-formula><mml:math id="M143" display="inline"><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>w</mml:mi><mml:mo mathvariant="normal" stretchy="true">̃</mml:mo></mml:mover><mml:mi>i</mml:mi></mml:msub><mml:mo>/</mml:mo><mml:msubsup><mml:mo>∑</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:msubsup><mml:msub><mml:mover accent="true"><mml:mi>w</mml:mi><mml:mo mathvariant="normal" stretchy="true">̃</mml:mo></mml:mover><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> with the property that <inline-formula><mml:math id="M144" display="inline"><mml:mrow><mml:msubsup><mml:mo>∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:msubsup><mml:msub><mml:mi>w</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula>. Although the additional approximation has downsides <xref ref-type="bibr" rid="bib1.bibx104" id="paren.151"/>, this self-normalization step means that we can ignore normalizing constant terms in the proposal, prior, and likelihood since these will cancel upon normalization.</p>
      <p id="d2e4247">The SNIS estimate of the posterior expectation in Eq. (<xref ref-type="disp-formula" rid="Ch1.E14"/>) forms the basis of the majority of SIR-based particle methods applied to the geosciences  <xref ref-type="bibr" rid="bib1.bibx126" id="paren.152"/>. Moreover, SNIS is formally mathematically equivalent to using a particle approximation <xref ref-type="bibr" rid="bib1.bibx113" id="paren.153"/> of the posterior in Eq. (<xref ref-type="disp-formula" rid="Ch1.E7"/>) of the form

              <disp-formula id="Ch1.E15" content-type="numbered"><label>15</label><mml:math id="M145" display="block"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo><mml:mo>≃</mml:mo><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:munderover><mml:msub><mml:mi>w</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mi mathvariant="italic">δ</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>-</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

            with weights <inline-formula><mml:math id="M146" display="inline"><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> given by Eq. (<xref ref-type="disp-formula" rid="Ch1.E14"/>) and <inline-formula><mml:math id="M147" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>∼</mml:mo><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. This is relatively trivial to verify by recalling the sifting property of the Dirac delta function that <inline-formula><mml:math id="M148" display="inline"><mml:mrow><mml:mo>∫</mml:mo><mml:mi>g</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mi mathvariant="italic">δ</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>-</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mi mathvariant="normal">d</mml:mi><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>=</mml:mo><mml:mi>g</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> <xref ref-type="bibr" rid="bib1.bibx95" id="paren.154"/>. On the one hand, this particle approximation helps conceptualize the posterior estimate in Eq. (<xref ref-type="disp-formula" rid="Ch1.E15"/>) as a sum of point particles whose relative importance is given by their weights. On the other hand, the SNIS formalism that we have outlined clarifies the origin of these weights.</p>
</sec>
<sec id="Ch1.S2.SS5.SSS4">
  <label>2.5.4</label><title>Resampling</title>
      <p id="d2e4419">In settings such as Bayesian filtering in state space models <xref ref-type="bibr" rid="bib1.bibx57 bib1.bibx72 bib1.bibx51 bib1.bibx113" id="paren.155"/> or applying data tempering for static parameter inference with large amounts of data <xref ref-type="bibr" rid="bib1.bibx20 bib1.bibx95 bib1.bibx122" id="paren.156"/> it can be helpful to embed importance sampling in Sequential Importance Resampling (SIR) that is the basis of particle filtering <xref ref-type="bibr" rid="bib1.bibx21" id="paren.157"/>. The workflow in SIR is to sequentially assimilate data as it becomes available in time or through minibatches and then to perform a resampling step on the dynamic weights after each SNIS step. This exploits the computational benefits of the sequential nature of Bayesian inference, where the posterior of the current step can become the prior for the next step <xref ref-type="bibr" rid="bib1.bibx113" id="paren.158"/>, while using resampling to avoid the inevitable weight degeneracy that would otherwise occur. Herein, our focus is on inferring static parameters via batch smoothing within a given water year so the sequential aspect of SIR is not as important. At the same time, the adaptive particle method presented herein could also be embedded within a sequential particle filtering framework. Moreover, the adaptation step that we use does make use of particle resampling so we also briefly outline what resampling entails.</p>
      <p id="d2e4434">In the resampling step, particles are resampled with replacement according to the probability mass given by their weights. As such, more fit high weight particles are copied whereas less fit low weight particles are removed.  Many particle resampling methods exist <xref ref-type="bibr" rid="bib1.bibx79" id="paren.159"/> and the particular method used is often of secondary importance as long as it is valid. After resampling, all particles are assigned an equal weight of <inline-formula><mml:math id="M149" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and can be treated as independent samples from the target allowing  for straightforward Monte Carlo integration to approximate expectations. The fact that resampling results in equal weights means that it avoids weight degeneracy since there will no longer be just a few particles carrying all the weight. Although resampling trivially solves the weight degeneracy problem, it does not solve the more general problem of path degeneracy wherein many identical resampled particles results in a poor posterior approximation <xref ref-type="bibr" rid="bib1.bibx95" id="paren.160"/>. A promising way to alleviate the path degeneracy problem is to rejuvenate the ensemble of particles through the resample-move algorithm by taking a few MCMC steps targeting the posterior to diversify the ensemble after resampling <xref ref-type="bibr" rid="bib1.bibx51 bib1.bibx122" id="paren.161"/>. Since we do not pursue filtering herein, we only mention this resample-move strategy in passing to motivate further work on particle filtering in cryospheric DA using the adaptive method we will present.</p>
</sec>
</sec>
<sec id="Ch1.S2.SS6">
  <label>2.6</label><title>Adaptive particle methods</title>
      <p id="d2e4470">The typical particle methods used in cryospheric DA <xref ref-type="bibr" rid="bib1.bibx7" id="paren.162"><named-content content-type="pre">see</named-content><named-content content-type="post">and references therein</named-content></xref> can be considered to be basic particle methods in the spirit of the seminal bootstrap particle filter <xref ref-type="bibr" rid="bib1.bibx57" id="paren.163"/> in that they do not leverage recent developments <xref ref-type="bibr" rid="bib1.bibx126 bib1.bibx21" id="paren.164"/>. They are basic in the sense that they use the simplest and most convenient choice of proposal distribution <xref ref-type="bibr" rid="bib1.bibx124" id="paren.165"/>, namely using the prior as the proposal. In this case, i.e. inserting <inline-formula><mml:math id="M150" display="inline"><mml:mrow><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> in Eq. (<xref ref-type="disp-formula" rid="Ch1.E12"/>), the unnormalized weights simply correspond to likelihood evaluations <inline-formula><mml:math id="M151" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi>w</mml:mi><mml:mo stretchy="true" mathvariant="normal">̃</mml:mo></mml:mover><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>∣</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. Thereby, the normalized weights in the weighted mean approximation of the posterior expectation in Eq. (<xref ref-type="disp-formula" rid="Ch1.E14"/>) just become normalized likelihoods of the form <inline-formula><mml:math id="M152" display="inline"><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>∣</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>/</mml:mo><mml:msubsup><mml:mo>∑</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:msubsup><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>∣</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. This choice of proposal also makes it easy to implement particle filtering via SIR since one does not need direct access to the dynamic prior which is generally no longer available in closed form. Nonetheless, in terms of the importance sampling step itself the use of the prior as the proposal is generally highly suboptimal. This has motivated more sophisticated particle methods such as those involving tempering <xref ref-type="bibr" rid="bib1.bibx21 bib1.bibx95" id="paren.166"/>, hybrid methods <xref ref-type="bibr" rid="bib1.bibx102 bib1.bibx113" id="paren.167"/>, and adaptive importance sampling methods <xref ref-type="bibr" rid="bib1.bibx15" id="paren.168"/>. Our focus will be on the latter, although we note in passing that it is entirely possible to combine these approaches <xref ref-type="bibr" rid="bib1.bibx73" id="paren.169"><named-content content-type="pre">e.g.</named-content></xref>.</p>
      <p id="d2e4632">The idea behind adaptive importance sampling is fairly simple: apply importance sampling iteratively and use the results of the previous iteration to adapt the proposal for the next. The reason that this approach tends to work is also clear: after each iteration the weighted particle ensemble tends to be a better approximation of the posterior than the proposal and so we can use these particles to allow the proposals to evolve by adapting toward the posterior.  As such, adaptive importance sampling leverages the power of iterations that iterative ensemble Kalman methods also exploit. At the same time, adaptive importance sampling adapts the proposal to the complexity of the target posterior at hand so that the cost of the algorithm is not set in stone. Thereby, these adaptive algorithms can converge rapidly via early stopping and only in the worst case will they run for a predefined maximum number of iterations. A plethora of adaptive importance sampling methods exist <xref ref-type="bibr" rid="bib1.bibx15" id="paren.170"/> and many of these are worthy of further investigation in cryospheric DA. The AMIS method that we chose to adopt was based on preliminary non-exhaustive testing of a subset of these methods and came out on top in terms of efficiency and performance.</p>
<sec id="Ch1.S2.SS6.SSS1">
  <label>2.6.1</label><title>Adaptive multiple importance sampling</title>
      <p id="d2e4645">Here we introduce the general adaptive multiple importance sampling (AMIS) algorithm proposed by <xref ref-type="bibr" rid="bib1.bibx26" id="text.171"/>. Our particular adaptation of this method to cryospheric DA with the AdaPBS is described in the subsequent section. This AMIS approach is adaptive in a couple of ways. Firstly, the proposal distribution iteratively adapts to better approximate the target posterior distribution. Secondly, the algorithm has an (optional) convergence criterion that is reached if an adequately high effective sample size is obtained. This means that AMIS will revert to the behaviour of computationally cheaper non-iterative basic importance sampling as typically used in cryospheric DA in cases where the basic method already provides a good enough approximation of the posterior. Note that these two adaptive features are common among many adaptive importance sampling methods as outlined in <xref ref-type="bibr" rid="bib1.bibx15" id="text.172"/>. What makes the AMIS method unique is that it employs a so-called deterministic mixture proposal <xref ref-type="bibr" rid="bib1.bibx100" id="paren.173"/> which makes it stable, relatively quick to converge, and waste-free in that the entire history of particles is used at each iteration. Due to this latter waste-free property, the AMIS algorithm is somewhat involved and requires tracking this evolving particle history through several generations. As such, it is easiest to present AMIS through a series of 7 sequential steps that we cycle through iteratively. That is, we initialize the iteration counter <inline-formula><mml:math id="M153" display="inline"><mml:mrow><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> and set the current sampling proposal to the prior <inline-formula><mml:math id="M154" display="inline"><mml:mrow><mml:msup><mml:mi>q</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and then proceed as follows:</p>
      <p id="d2e4701"><list list-type="order">
              <list-item>

      <p id="d2e4706">Generate an ensemble of particles indexed by <inline-formula><mml:math id="M155" display="inline"><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:msubsup><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula> from the current sampling proposal

                    <disp-formula id="Ch1.E16" content-type="numbered"><label>16</label><mml:math id="M156" display="block"><mml:mrow><mml:msubsup><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>∼</mml:mo><mml:msup><mml:mi>q</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
              </list-item>
              <list-item>

      <p id="d2e4776">For the current particle history, that is for all <inline-formula><mml:math id="M157" display="inline"><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi></mml:mrow></mml:math></inline-formula> historical iterations and the ensembles of <inline-formula><mml:math id="M158" display="inline"><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:msubsup><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula> particles in these iterations, evaluate the <italic>deterministic mixture</italic> (DM) proposal density <inline-formula><mml:math id="M159" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="italic">υ</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> for each particle <inline-formula><mml:math id="M160" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula>

                    <disp-formula id="Ch1.E17" content-type="numbered"><label>17</label><mml:math id="M161" display="block"><mml:mrow><mml:msup><mml:mi mathvariant="italic">υ</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>(</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:msubsup><mml:mi>N</mml:mi><mml:mi mathvariant="normal">h</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi mathvariant="normal">ℓ</mml:mi></mml:munderover><mml:msubsup><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>j</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:msup><mml:mi>q</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>j</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>(</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>)</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

                  where <inline-formula><mml:math id="M162" display="inline"><mml:mrow><mml:msubsup><mml:mi>N</mml:mi><mml:mi mathvariant="normal">h</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msubsup><mml:mo>∑</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi mathvariant="normal">ℓ</mml:mi></mml:msubsup><mml:msubsup><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula> is the total number of particles in the current particle history. The ratio <inline-formula><mml:math id="M163" display="inline"><mml:mrow><mml:msubsup><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>j</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>/</mml:mo><mml:msubsup><mml:mi>N</mml:mi><mml:mi mathvariant="normal">h</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula> corresponds to the weight of each iteration's proposal density <inline-formula><mml:math id="M164" display="inline"><mml:mrow><mml:msup><mml:mi>q</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>j</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> in the current DM <inline-formula><mml:math id="M165" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="italic">υ</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>. The particles <inline-formula><mml:math id="M166" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula> for each iteration <inline-formula><mml:math id="M167" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula> have been drawn independently from their respective sampling proposal distributions <inline-formula><mml:math id="M168" display="inline"><mml:mrow><mml:msup><mml:mi>q</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>. Nonetheless, using this DM concept from <xref ref-type="bibr" rid="bib1.bibx100" id="text.174"/> we can treat all the particles <italic>as if</italic> they were collectively drawn from the mixture <inline-formula><mml:math id="M169" display="inline"><mml:mrow><mml:mi mathvariant="italic">υ</mml:mi><mml:mo>(</mml:mo><mml:mo>⋅</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> where we happened to end up with exactly <inline-formula><mml:math id="M170" display="inline"><mml:mrow><mml:msubsup><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula> particles from each proposal.</p>
              </list-item>
              <list-item>

      <p id="d2e5163">Using the DM in Eq. (<xref ref-type="disp-formula" rid="Ch1.E17"/>) as our effective proposal for importance sampling compute unnormalized weights for the current particle history through

                    <disp-formula id="Ch1.E18" content-type="numbered"><label>18</label><mml:math id="M171" display="block"><mml:mrow><mml:msubsup><mml:mover accent="true"><mml:mi>w</mml:mi><mml:mo stretchy="true" mathvariant="normal">̃</mml:mo></mml:mover><mml:mi>i</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>∣</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>)</mml:mo><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:msup><mml:mi mathvariant="italic">υ</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>(</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:mfrac></mml:mstyle><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula></p>
              </list-item>
              <list-item>

      <p id="d2e5266">Using the entire particle history, compute <inline-formula><mml:math id="M172" display="inline"><mml:mrow><mml:msubsup><mml:mi>N</mml:mi><mml:mi mathvariant="normal">h</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula> self-normalized particle weights

                    <disp-formula id="Ch1.E19" content-type="numbered"><label>19</label><mml:math id="M173" display="block"><mml:mrow><mml:msubsup><mml:mi>w</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msubsup><mml:mover accent="true"><mml:mi>w</mml:mi><mml:mo stretchy="true" mathvariant="normal">̃</mml:mo></mml:mover><mml:mi>i</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup></mml:mrow><mml:mrow><mml:msubsup><mml:mi>N</mml:mi><mml:mi mathvariant="normal">h</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:msubsup><mml:mi>Z</mml:mi><mml:mi mathvariant="italic">υ</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:mfrac></mml:mstyle><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

                  where the normalizing constant in the denominator

                    <disp-formula id="Ch1.E20" content-type="numbered"><label>20</label><mml:math id="M174" display="block"><mml:mrow><mml:msubsup><mml:mi>Z</mml:mi><mml:mi mathvariant="italic">υ</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:msubsup><mml:mi>N</mml:mi><mml:mi mathvariant="normal">h</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi mathvariant="normal">ℓ</mml:mi></mml:munderover><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mrow><mml:msubsup><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:munderover><mml:msubsup><mml:mover accent="true"><mml:mi>w</mml:mi><mml:mo stretchy="true" mathvariant="normal">̃</mml:mo></mml:mover><mml:mi>i</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

                  also provides estimate for the model evidence using the DM proposal since

                    <disp-formula id="Ch1.E21" content-type="numbered"><label>21</label><mml:math id="M175" display="block"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mo movablelimits="false">∫</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:msup><mml:mi mathvariant="italic">υ</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mfrac></mml:mstyle><mml:msup><mml:mi mathvariant="italic">υ</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mi mathvariant="normal">d</mml:mi><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>≃</mml:mo><mml:msubsup><mml:mi>Z</mml:mi><mml:mi mathvariant="italic">υ</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula></p>
              </list-item>
              <list-item>

      <p id="d2e5542">Compute the effective sample size <xref ref-type="bibr" rid="bib1.bibx37" id="paren.175"><named-content content-type="pre">cf.</named-content></xref> using <italic>all</italic> the <inline-formula><mml:math id="M176" display="inline"><mml:mrow><mml:msubsup><mml:mi>N</mml:mi><mml:mi mathvariant="normal">h</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula> historical particles

                    <disp-formula id="Ch1.E22" content-type="numbered"><label>22</label><mml:math id="M177" display="block"><mml:mrow><mml:msubsup><mml:mi>N</mml:mi><mml:mi mathvariant="normal">eff</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi mathvariant="normal">ℓ</mml:mi></mml:munderover><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mrow><mml:msubsup><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:munderover><mml:mo>(</mml:mo><mml:msubsup><mml:mi>w</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:msup><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

                  if the predefined maximum number of iterations has been reached (<inline-formula><mml:math id="M178" display="inline"><mml:mrow><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) or a desired effective sample size is obtained <inline-formula><mml:math id="M179" display="inline"><mml:mrow><mml:msubsup><mml:mi>N</mml:mi><mml:mi mathvariant="normal">eff</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>≥</mml:mo><mml:mi mathvariant="italic">τ</mml:mi><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> given a predefined threshold <inline-formula><mml:math id="M180" display="inline"><mml:mrow><mml:mn mathvariant="normal">0</mml:mn><mml:mo>&lt;</mml:mo><mml:mi mathvariant="italic">τ</mml:mi><mml:mo>≤</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> then this <inline-formula><mml:math id="M181" display="inline"><mml:mi mathvariant="normal">ℓ</mml:mi></mml:math></inline-formula> becomes the final AMIS iteration which we denote as <inline-formula><mml:math id="M182" display="inline"><mml:mi>L</mml:mi></mml:math></inline-formula>. If this stopping criterion is not satisfied, apply the clipping approach of <xref ref-type="bibr" rid="bib1.bibx73" id="text.176"/> to improve proposal adaptation. First, identify the <inline-formula><mml:math id="M183" display="inline"><mml:mi mathvariant="script">T</mml:mi></mml:math></inline-formula>th largest unnormalized weight denoted <inline-formula><mml:math id="M184" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi>w</mml:mi><mml:mo mathvariant="normal" stretchy="true">̃</mml:mo></mml:mover><mml:mi mathvariant="script">T</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> where <inline-formula><mml:math id="M185" display="inline"><mml:mrow><mml:mi mathvariant="script">T</mml:mi><mml:mo>=</mml:mo><mml:mi mathvariant="normal">round</mml:mi><mml:mfenced open="(" close=")"><mml:mrow><mml:mi mathvariant="italic">τ</mml:mi><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:mfenced></mml:mrow></mml:math></inline-formula>. Subsequently, as long as <inline-formula><mml:math id="M186" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi>w</mml:mi><mml:mo mathvariant="normal" stretchy="true">̃</mml:mo></mml:mover><mml:mi mathvariant="script">T</mml:mi></mml:msub><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula>, apply weight clipping

                    <disp-formula id="Ch1.E23" content-type="numbered"><label>23</label><mml:math id="M187" display="block"><mml:mrow><mml:msubsup><mml:mover accent="true"><mml:mi>w</mml:mi><mml:mo mathvariant="normal" stretchy="true">̃</mml:mo></mml:mover><mml:mi>i</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>←</mml:mo><mml:mo movablelimits="false">min⁡</mml:mo><mml:mfenced open="(" close=")"><mml:mrow><mml:msubsup><mml:mover accent="true"><mml:mi>w</mml:mi><mml:mo stretchy="true" mathvariant="normal">̃</mml:mo></mml:mover><mml:mi>i</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>w</mml:mi><mml:mo mathvariant="normal" stretchy="true">̃</mml:mo></mml:mover><mml:mi mathvariant="script">T</mml:mi></mml:msub></mml:mrow></mml:mfenced><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

                  with subsequent renormalization via Eq. (<xref ref-type="disp-formula" rid="Ch1.E19"/>) . Clipping ensures equality of the <inline-formula><mml:math id="M188" display="inline"><mml:mi mathvariant="script">T</mml:mi></mml:math></inline-formula> largest clipped weights which guarantees a clipped <inline-formula><mml:math id="M189" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">eff</mml:mi></mml:msub><mml:mo>≥</mml:mo><mml:mi mathvariant="script">T</mml:mi></mml:mrow></mml:math></inline-formula> that in turn leads to a more robust and less degenerate sampling proposal adaptation.</p>
              </list-item>
              <list-item>

      <p id="d2e5873">Resample an ensemble of <inline-formula><mml:math id="M190" display="inline"><mml:mrow><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> equally weighted particles <inline-formula><mml:math id="M191" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>r</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula> using the <inline-formula><mml:math id="M192" display="inline"><mml:mrow><mml:msubsup><mml:mi>N</mml:mi><mml:mi mathvariant="normal">h</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula> weights <inline-formula><mml:math id="M193" display="inline"><mml:mrow><mml:msubsup><mml:mi>w</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula> of the current particle history obtained from Eq. (<xref ref-type="disp-formula" rid="Ch1.E19"/>) via Eq. (<xref ref-type="disp-formula" rid="Ch1.E23"/>) if needed. Use of the index <inline-formula><mml:math id="M194" display="inline"><mml:mi>r</mml:mi></mml:math></inline-formula> emphasizes this is an ensemble of <italic>resampled</italic> particles which approximate the posterior rather than draws from the sampling proposals Eq. (<xref ref-type="disp-formula" rid="Ch1.E16"/>).</p>
              </list-item>
              <list-item>

      <p id="d2e5973">If the final AMIS iteration has been reached <inline-formula><mml:math id="M195" display="inline"><mml:mrow><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>=</mml:mo><mml:mi>L</mml:mi></mml:mrow></mml:math></inline-formula> according to the stopping criterion in step 5, stop iterating and use the resampled particles <inline-formula><mml:math id="M196" display="inline"><mml:mrow><mml:mo mathvariant="italic">{</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>r</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:msubsup><mml:mo mathvariant="italic">}</mml:mo><mml:mrow><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula> as a particle approximation of the posterior distribution. Otherwise, use the resampled particles to construct a new sampling proposal distribution <inline-formula><mml:math id="M197" display="inline"><mml:mrow><mml:msup><mml:mi>q</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> for the next iteration, then update the iteration counter <inline-formula><mml:math id="M198" display="inline"><mml:mrow><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>←</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> and return to step 1.</p>
              </list-item>
            </list></p>
      <p id="d2e6062">Generating an adaptive particle history by iterating the sequence of steps 1–7 outlined above until convergence is the general workflow for AMIS. Nonetheless, to be able to implement this algorithm for particle smoothing in data assimilation we need to specify the sampling proposal distributions <inline-formula><mml:math id="M199" display="inline"><mml:mrow><mml:msup><mml:mi>q</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> and how the resampled particles can be used to update these proposals from one iteration to the next in step 7.</p>
</sec>
<sec id="Ch1.S2.SS6.SSS2">
  <label>2.6.2</label><title>Adaptive particle batch smoother</title>
      <p id="d2e6089">Recall that we seek to enhance a popular cryospheric DA algorithm known as the PBS that was introduced by <xref ref-type="bibr" rid="bib1.bibx85" id="text.177"/> and has since been widely adopted for cryospheric reanalysis <xref ref-type="bibr" rid="bib1.bibx86 bib1.bibx96 bib1.bibx27 bib1.bibx46 bib1.bibx6 bib1.bibx81 bib1.bibx54 bib1.bibx17 bib1.bibx120" id="paren.178"><named-content content-type="pre">e.g.</named-content></xref> by embedding it within the powerful and more general adaptive framework of the AMIS algorithm <xref ref-type="bibr" rid="bib1.bibx26" id="paren.179"/>. We will refer to this proposed new DA method as the adaptive PBS, or AdaPBS for short. In the current AdaPBS implementation, we immediately make one simplification compared to the general AMIS algorithm laid out above which is to fix the number of particles <inline-formula><mml:math id="M200" display="inline"><mml:mrow><mml:msubsup><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula> sampled in each iteration to a constant ensemble size, i.e. <inline-formula><mml:math id="M201" display="inline"><mml:mrow><mml:msubsup><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> for all iterations <inline-formula><mml:math id="M202" display="inline"><mml:mi mathvariant="normal">ℓ</mml:mi></mml:math></inline-formula>. The parsimonious choice of a constant ensemble size for all iterations reduces the number of hyperparameters that the user has to specify and facilitates comparisons with the conventional PBS with a given ensemble size. We still suspect that varying the size of the ensemble during the iterations may lead to performance improvements, which is a topic worthy of future research. In our case of a fixed ensemble size at each iteration, the ratio in the DM simplifies to <inline-formula><mml:math id="M203" display="inline"><mml:mrow><mml:msubsup><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>j</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>/</mml:mo><mml:msubsup><mml:mi>N</mml:mi><mml:mi mathvariant="normal">h</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi></mml:mrow></mml:math></inline-formula> since <inline-formula><mml:math id="M204" display="inline"><mml:mrow><mml:msubsup><mml:mi>N</mml:mi><mml:mi mathvariant="normal">h</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> so each component of the DM proposal is equally weighted. As such, the typically suboptimal choice of the prior as the initial proposal, which is implicit in the PBS, with a mixture weight of <inline-formula><mml:math id="M205" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi></mml:mrow></mml:math></inline-formula> becomes less influential as <inline-formula><mml:math id="M206" display="inline"><mml:mi mathvariant="normal">ℓ</mml:mi></mml:math></inline-formula> increases and the latter adaptive iterations together occupy a continuously increasing total mixture weight of <inline-formula><mml:math id="M207" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>)</mml:mo><mml:mo>/</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>→</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> as <inline-formula><mml:math id="M208" display="inline"><mml:mrow><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>→</mml:mo><mml:mi mathvariant="normal">∞</mml:mi></mml:mrow></mml:math></inline-formula>.</p>
      <p id="d2e6277">For simplicity, but without loss of generality, in the current version of AdaPBS we employ multivariate normal (Gaussian) proposal distributions <inline-formula><mml:math id="M209" display="inline"><mml:mrow><mml:msup><mml:mi>q</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mi>N</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>∣</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">μ</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mi mathvariant="bold">C</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> that are parametrized by hyperparameters in the form of a mean vector <inline-formula><mml:math id="M210" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">μ</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> and covariance matrix <inline-formula><mml:math id="M211" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold">C</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>. In principle, we could set these proposals to anything with at least the same support as the target posterior so we have great freedom in picking our proposal distribution. In practice, however, importance sampling works better if the proposal is similar to the target posterior distribution. As such, it is advantageous to update the hyperparameters of the successive proposal distributions using the current particle approximation of the posterior in the form of the ensemble <inline-formula><mml:math id="M212" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>r</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula> obtained in step 6 of AMIS. Recalling that this is an equally weighted ensemble, an obvious choice is to estimate these hyperparameters using ensemble statistics from the successive posterior approximations. Thus, in step 7 of AMIS the mean vector and covariance matrix for the Gaussian proposal in the next iteration <inline-formula><mml:math id="M213" display="inline"><mml:mrow><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> are simply set to the ensemble mean vector

              <disp-formula id="Ch1.E24" content-type="numbered"><label>24</label><mml:math id="M214" display="block"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">μ</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:msup><mml:mi mathvariant="bold">Θ</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mn mathvariant="bold">1</mml:mn><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

            and ensemble covariance matrix

              <disp-formula id="Ch1.E25" content-type="numbered"><label>25</label><mml:math id="M215" display="block"><mml:mrow><mml:msup><mml:mi mathvariant="bold">C</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:mfenced open="(" close=")"><mml:mrow><mml:msup><mml:mi mathvariant="bold">Θ</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">μ</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:msup><mml:mn mathvariant="bold">1</mml:mn><mml:mi mathvariant="normal">T</mml:mi></mml:msup></mml:mrow></mml:mfenced><mml:msup><mml:mfenced open="(" close=")"><mml:mrow><mml:msup><mml:mi mathvariant="bold">Θ</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">μ</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:msup><mml:mn mathvariant="bold">1</mml:mn><mml:mi mathvariant="normal">T</mml:mi></mml:msup></mml:mrow></mml:mfenced><mml:mi mathvariant="normal">T</mml:mi></mml:msup><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

            where the particle ensemble <inline-formula><mml:math id="M216" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>r</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula> is stored as column vectors in the <inline-formula><mml:math id="M217" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">p</mml:mi></mml:msub><mml:mo>×</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> matrix <inline-formula><mml:math id="M218" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold">Θ</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M219" display="inline"><mml:mn mathvariant="bold">1</mml:mn></mml:math></inline-formula> is a <inline-formula><mml:math id="M220" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub><mml:mo>×</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> column vector of ones.</p>
      <p id="d2e6626">If, using suitable transformations, we further require a prior <inline-formula><mml:math id="M221" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mi>N</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">μ</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold">C</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> that is Gaussian and employing the usual Gaussian likelihood Eq. (<xref ref-type="disp-formula" rid="Ch1.E4"/>) of the PBS <inline-formula><mml:math id="M222" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mi>N</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>∣</mml:mo><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo stretchy="true" mathvariant="normal">^</mml:mo></mml:mover><mml:mo>,</mml:mo><mml:mi mathvariant="bold">R</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> where the mean <inline-formula><mml:math id="M223" display="inline"><mml:mrow><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo stretchy="true" mathvariant="normal">^</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo stretchy="true" mathvariant="normal">^</mml:mo></mml:mover><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the predicted observation vector and <inline-formula><mml:math id="M224" display="inline"><mml:mi mathvariant="bold">R</mml:mi></mml:math></inline-formula> is the observation error covariance then the (unnormalized) target posterior is

              <disp-formula id="Ch1.E26" content-type="numbered"><label>26</label><mml:math id="M225" display="block"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mi>A</mml:mi><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mi mathvariant="normal">exp</mml:mi><mml:mfenced open="(" close=")"><mml:mrow><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:mfrac></mml:mstyle><mml:mi mathvariant="italic">ϕ</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mfenced><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

            where <inline-formula><mml:math id="M226" display="inline"><mml:mrow><mml:mi>A</mml:mi><mml:mo>=</mml:mo><mml:mi mathvariant="normal">det</mml:mi><mml:mo>(</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:mi mathvariant="italic">π</mml:mi><mml:mi mathvariant="bold">R</mml:mi><mml:msup><mml:mo>)</mml:mo><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:msup><mml:mi mathvariant="normal">det</mml:mi><mml:mo>(</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:mi mathvariant="italic">π</mml:mi><mml:mi mathvariant="bold">C</mml:mi><mml:msup><mml:mo>)</mml:mo><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> is a constant that does not depend on <inline-formula><mml:math id="M227" display="inline"><mml:mi mathvariant="bold-italic">θ</mml:mi></mml:math></inline-formula> and we have introduced the unnormalized and scaled negative log posterior <inline-formula><mml:math id="M228" display="inline"><mml:mi mathvariant="italic">ϕ</mml:mi></mml:math></inline-formula> (not to be confused with the uncertain bounded parameters <inline-formula><mml:math id="M229" display="inline"><mml:mrow><mml:mi mathvariant="bold-italic">φ</mml:mi><mml:mo>=</mml:mo><mml:mi mathvariant="script">T</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>) defined as

              <disp-formula id="Ch1.E27" content-type="numbered"><label>27</label><mml:math id="M230" display="block"><mml:mrow><mml:mi mathvariant="italic">ϕ</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:msup><mml:mfenced close="]" open="["><mml:mrow><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>-</mml:mo><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo mathvariant="normal" stretchy="true">^</mml:mo></mml:mover></mml:mrow></mml:mfenced><mml:mi mathvariant="normal">T</mml:mi></mml:msup><mml:msup><mml:mi mathvariant="bold">R</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup><mml:mfenced close="]" open="["><mml:mrow><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>-</mml:mo><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo stretchy="true" mathvariant="normal">^</mml:mo></mml:mover></mml:mrow></mml:mfenced><mml:mo>+</mml:mo><mml:msup><mml:mfenced close="]" open="["><mml:mrow><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>-</mml:mo><mml:mi mathvariant="bold-italic">μ</mml:mi></mml:mrow></mml:mfenced><mml:mi mathvariant="normal">T</mml:mi></mml:msup><mml:msup><mml:mi mathvariant="bold">C</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup><mml:mfenced close="]" open="["><mml:mrow><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>-</mml:mo><mml:mi mathvariant="bold-italic">μ</mml:mi></mml:mrow></mml:mfenced><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

            for economy. We may construct a similar expression for the Gaussian proposal distributions <inline-formula><mml:math id="M231" display="inline"><mml:mrow><mml:msup><mml:mi>q</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mi>N</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>∣</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">μ</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mi mathvariant="bold">C</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> of the form

              <disp-formula id="Ch1.E28" content-type="numbered"><label>28</label><mml:math id="M232" display="block"><mml:mrow><mml:msup><mml:mi>q</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:msup><mml:mi>c</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mi mathvariant="normal">exp</mml:mi><mml:mfenced open="(" close=")"><mml:mrow><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:mfrac></mml:mstyle><mml:msup><mml:mi mathvariant="italic">ψ</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mfenced><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

            where we have defined the shorthand function

              <disp-formula id="Ch1.E29" content-type="numbered"><label>29</label><mml:math id="M233" display="block"><mml:mrow><mml:msup><mml:mi mathvariant="italic">ψ</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:msup><mml:mfenced close="]" open="["><mml:mrow><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">μ</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:mfenced><mml:mi mathvariant="normal">T</mml:mi></mml:msup><mml:msup><mml:msup><mml:mi mathvariant="bold">C</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup><mml:mfenced open="[" close="]"><mml:mrow><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">μ</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:mfenced><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

            where <inline-formula><mml:math id="M234" display="inline"><mml:mrow><mml:msup><mml:mi>c</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mi mathvariant="normal">det</mml:mi><mml:mo>(</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:mi mathvariant="italic">π</mml:mi><mml:msup><mml:mi mathvariant="bold">C</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:msup><mml:mo>)</mml:mo><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> is a normalizing constant that only depends on the sampling proposal covariance <inline-formula><mml:math id="M235" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold">C</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>. Inserting Eq. (<xref ref-type="disp-formula" rid="Ch1.E29"/>) in Eq. (<xref ref-type="disp-formula" rid="Ch1.E17"/>) with fixed <inline-formula><mml:math id="M236" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> the Gaussian DM proposal at iteration <inline-formula><mml:math id="M237" display="inline"><mml:mi mathvariant="normal">ℓ</mml:mi></mml:math></inline-formula> has the following simple analytical form

              <disp-formula id="Ch1.E30" content-type="numbered"><label>30</label><mml:math id="M238" display="block"><mml:mrow><mml:msup><mml:mi mathvariant="italic">υ</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mi mathvariant="normal">ℓ</mml:mi></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi mathvariant="normal">ℓ</mml:mi></mml:munderover><mml:msup><mml:mi>c</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>j</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mi mathvariant="normal">exp</mml:mi><mml:mfenced open="(" close=")"><mml:mrow><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:mfrac></mml:mstyle><mml:msup><mml:mi mathvariant="italic">ψ</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>j</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mfenced><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

            whereby it is readily verified that a numerically stable estimate of the logarithm of the unnormalized weights in Eq. (<xref ref-type="disp-formula" rid="Ch1.E18"/>) for the particle history (<inline-formula><mml:math id="M239" display="inline"><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M240" display="inline"><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) at iteration <inline-formula><mml:math id="M241" display="inline"><mml:mi mathvariant="normal">ℓ</mml:mi></mml:math></inline-formula> is given by

              <disp-formula id="Ch1.E31" content-type="numbered"><label>31</label><mml:math id="M242" display="block"><mml:mrow><mml:mi>log⁡</mml:mi><mml:mo>(</mml:mo><mml:msubsup><mml:mover accent="true"><mml:mi>w</mml:mi><mml:mo mathvariant="normal" stretchy="true">̃</mml:mo></mml:mover><mml:mi>i</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mi>log⁡</mml:mi><mml:mo>(</mml:mo><mml:mi>A</mml:mi><mml:mo>)</mml:mo><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.5</mml:mn><mml:mi mathvariant="italic">ϕ</mml:mi><mml:mo>(</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>)</mml:mo><mml:mo>-</mml:mo><mml:msub><mml:mi mathvariant="normal">LSE</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mfenced close=")" open="("><mml:mrow><mml:msup><mml:mi>a</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>j</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>(</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:mfenced><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

            where with some abuse of notation the log-sum-exp term <inline-formula><mml:math id="M243" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">LSE</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mo>⋅</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> related to the DM is

              <disp-formula id="Ch1.E32" content-type="numbered"><label>32</label><mml:math id="M244" display="block"><mml:mtable class="split" rowspacing="0.2ex" displaystyle="true" columnalign="right left"><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:msub><mml:mi mathvariant="normal">LSE</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:msup><mml:mi>a</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>j</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>(</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>)</mml:mo><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:msup><mml:mi>a</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mo>⋆</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>(</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mspace linebreak="nobreak" width="1em"/><mml:mo>+</mml:mo><mml:mi>log⁡</mml:mi><mml:mfenced open="(" close=")"><mml:mrow><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi mathvariant="normal">ℓ</mml:mi></mml:munderover><mml:mi>exp⁡</mml:mi><mml:mfenced close=")" open="("><mml:mrow><mml:msup><mml:mi>a</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>j</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>(</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>)</mml:mo><mml:mo>-</mml:mo><mml:msup><mml:mi>a</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mo>⋆</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>(</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:mfenced></mml:mrow></mml:mfenced><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>

            which is implicitly a function with <inline-formula><mml:math id="M245" display="inline"><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi></mml:mrow></mml:math></inline-formula> arguments

              <disp-formula id="Ch1.E33" content-type="numbered"><label>33</label><mml:math id="M246" display="block"><mml:mrow><mml:msup><mml:mi>a</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>j</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>(</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mi>log⁡</mml:mi><mml:mo>(</mml:mo><mml:msup><mml:mi>c</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>j</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>)</mml:mo><mml:mo>-</mml:mo><mml:mi>log⁡</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.5</mml:mn><mml:msup><mml:mi mathvariant="italic">ψ</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>j</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>(</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>)</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

            whose maximum across <inline-formula><mml:math id="M247" display="inline"><mml:mi mathvariant="normal">ℓ</mml:mi></mml:math></inline-formula> iterations is <inline-formula><mml:math id="M248" display="inline"><mml:mrow><mml:msup><mml:mi>a</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mo>⋆</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>(</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mo>max⁡</mml:mo><mml:mi>j</mml:mi></mml:msub><mml:mfenced open="(" close=")"><mml:mrow><mml:msup><mml:mi>a</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>j</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>(</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:mfenced></mml:mrow></mml:math></inline-formula>. We emphasize that to obtain the logarithm of the unnormalized weights Eq. (<xref ref-type="disp-formula" rid="Ch1.E31"/>) for the particle history at each iteration <inline-formula><mml:math id="M249" display="inline"><mml:mi mathvariant="normal">ℓ</mml:mi></mml:math></inline-formula> both the negative log posterior in Eq. (<xref ref-type="disp-formula" rid="Ch1.E27"/>) and the LSE form of the DM in Eq. (<xref ref-type="disp-formula" rid="Ch1.E32"/>) have to be evaluated for each of the <inline-formula><mml:math id="M250" display="inline"><mml:mrow><mml:mi mathvariant="normal">ℓ</mml:mi><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> particles. To obtain the logarithm of the self-normalized weights Eq. (<xref ref-type="disp-formula" rid="Ch1.E19"/>) we also need to use the logarithm of all the <inline-formula><mml:math id="M251" display="inline"><mml:mrow><mml:mi mathvariant="normal">ℓ</mml:mi><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> historical unnormalized weights from Eq. (<xref ref-type="disp-formula" rid="Ch1.E31"/>) to evaluate the logarithm of the evidence approximation Eq. (<xref ref-type="disp-formula" rid="Ch1.E20"/>) which is given by

              <disp-formula id="Ch1.E34" content-type="numbered"><label>34</label><mml:math id="M252" display="block"><mml:mrow><mml:mi>log⁡</mml:mi><mml:mo>(</mml:mo><mml:msubsup><mml:mi>Z</mml:mi><mml:mi mathvariant="italic">υ</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:mi>log⁡</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>+</mml:mo><mml:msub><mml:mi mathvariant="normal">LSE</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mfenced close=")" open="("><mml:mrow><mml:mi>log⁡</mml:mi><mml:mo>(</mml:mo><mml:msubsup><mml:mover accent="true"><mml:mi>w</mml:mi><mml:mo mathvariant="normal" stretchy="true">̃</mml:mo></mml:mover><mml:mi>i</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:mfenced><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

            where the new log-sum-exp term <inline-formula><mml:math id="M253" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">LSE</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mo>⋅</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is implicitly a function with <inline-formula><mml:math id="M254" display="inline"><mml:mrow><mml:mi mathvariant="normal">ℓ</mml:mi><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> arguments

              <disp-formula id="Ch1.E35" content-type="numbered"><label>35</label><mml:math id="M255" display="block"><mml:mtable rowspacing="0.2ex" class="split" displaystyle="true" columnalign="right left"><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:msub><mml:mi mathvariant="normal">LSE</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mfenced open="(" close=")"><mml:mrow><mml:mi>log⁡</mml:mi><mml:mo>(</mml:mo><mml:msubsup><mml:mover accent="true"><mml:mi>w</mml:mi><mml:mo stretchy="true" mathvariant="normal">̃</mml:mo></mml:mover><mml:mi>i</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:mfenced><mml:mo>=</mml:mo><mml:mi mathvariant="normal">Ω</mml:mi></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mspace linebreak="nobreak" width="1em"/><mml:mo>+</mml:mo><mml:mi>log⁡</mml:mi><mml:mfenced close=")" open="("><mml:mrow><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi mathvariant="normal">ℓ</mml:mi></mml:munderover><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:munderover><mml:mi>exp⁡</mml:mi><mml:mfenced open="(" close=")"><mml:mrow><mml:mo>(</mml:mo><mml:mi>log⁡</mml:mi><mml:mo>(</mml:mo><mml:msubsup><mml:mover accent="true"><mml:mi>w</mml:mi><mml:mo mathvariant="normal" stretchy="true">̃</mml:mo></mml:mover><mml:mi>i</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>)</mml:mo><mml:mo>-</mml:mo><mml:mi mathvariant="normal">Ω</mml:mi></mml:mrow></mml:mfenced></mml:mrow></mml:mfenced><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>

            where <inline-formula><mml:math id="M256" display="inline"><mml:mrow><mml:mi mathvariant="normal">Ω</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mo>max⁡</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mfenced open="(" close=")"><mml:mrow><mml:mi>log⁡</mml:mi><mml:mo>(</mml:mo><mml:msubsup><mml:mover accent="true"><mml:mi>w</mml:mi><mml:mo stretchy="true" mathvariant="normal">̃</mml:mo></mml:mover><mml:mi>i</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:mfenced></mml:mrow></mml:math></inline-formula>. Note that Eq. (<xref ref-type="disp-formula" rid="Ch1.E34"/>) only needs to be evaluated once per iteration <inline-formula><mml:math id="M257" display="inline"><mml:mi mathvariant="normal">ℓ</mml:mi></mml:math></inline-formula>. Finally, combining Eqs. (<xref ref-type="disp-formula" rid="Ch1.E31"/>) and  (<xref ref-type="disp-formula" rid="Ch1.E34"/>), the stable logarithm of the self-normalized weights Eq. (<xref ref-type="disp-formula" rid="Ch1.E19"/>) that we seek can now be computed as

              <disp-formula id="Ch1.E36" content-type="numbered"><label>36</label><mml:math id="M258" display="block"><mml:mrow><mml:mi>log⁡</mml:mi><mml:mo>(</mml:mo><mml:msubsup><mml:mi>w</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mi>log⁡</mml:mi><mml:mo>(</mml:mo><mml:msubsup><mml:mover accent="true"><mml:mi>w</mml:mi><mml:mo mathvariant="normal" stretchy="true">̃</mml:mo></mml:mover><mml:mi>i</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>)</mml:mo><mml:mo>-</mml:mo><mml:mi>log⁡</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>-</mml:mo><mml:mi>log⁡</mml:mi><mml:mo>(</mml:mo><mml:msubsup><mml:mi>Z</mml:mi><mml:mi mathvariant="italic">υ</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>)</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

            which we have expressed in terms of the log evidence approximation Eq. (<xref ref-type="disp-formula" rid="Ch1.E34"/>) since this is a useful additional output from AdaPBS for hierarchical Bayesian inference <xref ref-type="bibr" rid="bib1.bibx109 bib1.bibx95" id="paren.180"/>. The self-normalized weights for AdaPBS are now obtained trivially by taking the exponential of Eq. (<xref ref-type="disp-formula" rid="Ch1.E36"/>). These weights can then be resampled <inline-formula><mml:math id="M259" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> times to obtain an equally weighted ensemble of particles as described in step 6 of the AMIS algorihtm. Next, as described in step 7, upon convergence these <inline-formula><mml:math id="M260" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> resampled particles are output as a particle approximation of the posterior. Otherwise we use these particles to design the Gaussian sampling proposal through Eqs. (<xref ref-type="disp-formula" rid="Ch1.E24"/>) and (<xref ref-type="disp-formula" rid="Ch1.E25"/>) for the next iteration <inline-formula><mml:math id="M261" display="inline"><mml:mrow><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula>.</p>
      <p id="d2e8321">The overall workflow of the AdaPBS algorithm is visualized in Fig. <xref ref-type="fig" rid="F1"/>. As with the standard PBS, AdaPBS is initialized by sampling a particle ensemble from the prior, running this ensemble through the forward model, and computing self-normalized importance weights via Eq. (<xref ref-type="disp-formula" rid="Ch1.E36"/>). Ensemble diversity is then diagnosed using the effective sample size <inline-formula><mml:math id="M262" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">eff</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> in Eq. (<xref ref-type="disp-formula" rid="Ch1.E22"/>). If <inline-formula><mml:math id="M263" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">eff</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> exceeds a predefined threshold, the algorithm stops successfully, outputting this diverse ensemble as the posterior approximation. Otherwise, it adapts a new sampling proposal using clipping Eq. (<xref ref-type="disp-formula" rid="Ch1.E23"/>) and ensemble statistics via Eqs. (<xref ref-type="disp-formula" rid="Ch1.E24"/>) and  (<xref ref-type="disp-formula" rid="Ch1.E25"/>). A new particle ensemble is drawn from this adapted sampling proposal, run through the forward model, and weights are obtained using Eq. (<xref ref-type="disp-formula" rid="Ch1.E36"/>), now incorporating the entire particle history via a mixture proposal Eq. (<xref ref-type="disp-formula" rid="Ch1.E17"/>). Adaptation continues until the ESS-based diversity threshold is met or a predefined maximum number of iterations is reached. Workflow of the adaptive particle batch smoother (AdaPBS), illustrating its extension (green) from the non-iterative PBS method (yellow). The algorithm iterates by adapting the proposal until the effective sample size (<inline-formula><mml:math id="M264" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">eff</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) reaches a predefined fraction of the ensemble size (<inline-formula><mml:math id="M265" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>), determined by the diversity threshold (<inline-formula><mml:math id="M266" display="inline"><mml:mi mathvariant="italic">τ</mml:mi></mml:math></inline-formula>), or until a maximum number of iterations (<inline-formula><mml:math id="M267" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) is reached. AdaPBS is mathematically equivalent to PBS if the diversity threshold is met in the first iteration (<inline-formula><mml:math id="M268" display="inline"><mml:mrow><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula>) or if the algorithm is constrained to a single iteration (<inline-formula><mml:math id="M269" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula>).</p>

      <fig id="F1" specific-use="star"><label>Figure 1</label><caption><p id="d2e8434">Flowchart showing the workflow in the adaptive particle batch smoother (AdaPBS) as an extension (green) of the non-iterative PBS method (yellow). The AdaPBS algorithm continues iterating by adapting the proposal until convergence, achieved when the effective sample size (<inline-formula><mml:math id="M270" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">eff</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) reaches at least a fraction of the ensemble size (<inline-formula><mml:math id="M271" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) determined by the desired diversity threshold (<inline-formula><mml:math id="M272" display="inline"><mml:mi mathvariant="italic">τ</mml:mi></mml:math></inline-formula>), or until a pre-defined maximum number of iterations (<inline-formula><mml:math id="M273" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) is reached. Both in the special case of convergence in the first iteration (<inline-formula><mml:math id="M274" display="inline"><mml:mrow><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula>) and in the non-iterative case (<inline-formula><mml:math id="M275" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula>), AdaPBS reduces to PBS.</p></caption>
            <graphic xlink:href="https://gmd.copernicus.org/articles/19/8565/2026/gmd-19-8565-2026-f01.png"/>

          </fig>

</sec>
</sec>
<sec id="Ch1.S2.SS7">
  <label>2.7</label><title>Probabilistic evaluation diagnostics</title>
      <p id="d2e8520">To evaluate the performance of the respective data assimilation schemes we employ two probabilistic evaluation diagnostics, namely the Continuous Ranked Probability Score <xref ref-type="bibr" rid="bib1.bibx56" id="paren.181"><named-content content-type="pre">CRPS;</named-content></xref> and the Kullback–Leibler Divergence <xref ref-type="bibr" rid="bib1.bibx95" id="paren.182"><named-content content-type="pre">KLD;</named-content></xref>. These probabilistic verification measures allow us to evaluate the performance of the entire approximate posterior distribution as represented by an ensemble of particles rather than just point estimates such as the posterior mean.</p>
      <p id="d2e8533">The CRPS is a strictly proper scoring rule defined by <xref ref-type="bibr" rid="bib1.bibx55" id="paren.183"/>

            <disp-formula id="Ch1.E37" content-type="numbered"><label>37</label><mml:math id="M276" display="block"><mml:mrow><mml:mi mathvariant="normal">CRPS</mml:mi><mml:mfenced open="(" close=")"><mml:mrow><mml:mi>P</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mo>⋆</mml:mo></mml:msup></mml:mrow></mml:mfenced><mml:mo>=</mml:mo><mml:mo movablelimits="false">∫</mml:mo><mml:msup><mml:mfenced open="[" close="]"><mml:mrow><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo><mml:mo>-</mml:mo><mml:mi>H</mml:mi><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>-</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mo>⋆</mml:mo></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mi mathvariant="normal">d</mml:mi><mml:mi>x</mml:mi><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

          where <inline-formula><mml:math id="M277" display="inline"><mml:mrow><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:msubsup><mml:mo>∫</mml:mo><mml:mi mathvariant="normal">∞</mml:mi><mml:mi>x</mml:mi></mml:msubsup><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">χ</mml:mi><mml:mo>)</mml:mo><mml:mi mathvariant="normal">d</mml:mi><mml:mi mathvariant="italic">χ</mml:mi></mml:mrow></mml:math></inline-formula> is the cumulative (posterior or prior) probability distribution of the uncertain variable or parameter <inline-formula><mml:math id="M278" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> of interest with density <inline-formula><mml:math id="M279" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> that we are inferring, <inline-formula><mml:math id="M280" display="inline"><mml:mrow><mml:msup><mml:mi>x</mml:mi><mml:mo>⋆</mml:mo></mml:msup></mml:mrow></mml:math></inline-formula> is the reference truth value, and <inline-formula><mml:math id="M281" display="inline"><mml:mrow><mml:mi>H</mml:mi><mml:mo>(</mml:mo><mml:mo>⋅</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the Heaviside function which is <inline-formula><mml:math id="M282" display="inline"><mml:mn mathvariant="normal">1</mml:mn></mml:math></inline-formula> for positive arguments (here <inline-formula><mml:math id="M283" display="inline"><mml:mrow><mml:mi>x</mml:mi><mml:mo>&gt;</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mo>⋆</mml:mo></mml:msup></mml:mrow></mml:math></inline-formula>) and <inline-formula><mml:math id="M284" display="inline"><mml:mn mathvariant="normal">0</mml:mn></mml:math></inline-formula>  otherwise. The CRPS is non-negative and inherits the same units as the variable <inline-formula><mml:math id="M285" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> of interest. In the special case of a deterministic forecast, where <inline-formula><mml:math id="M286" display="inline"><mml:mrow><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is also a Heaviside function, the CRPS reduces to the absolute error. As such, the CRPS is negatively oriented with the best possible value being <inline-formula><mml:math id="M287" display="inline"><mml:mn mathvariant="normal">0</mml:mn></mml:math></inline-formula> indicating perfect agreement with the reference. In the general non-degenerate probabilistic setting the CRPS measures both the accuracy (goodness of fit) and the precision (sharpness) of the distribution <inline-formula><mml:math id="M288" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> relative to the reference <inline-formula><mml:math id="M289" display="inline"><mml:mrow><mml:msup><mml:mi>x</mml:mi><mml:mo>⋆</mml:mo></mml:msup></mml:mrow></mml:math></inline-formula>, making it an apt choice for evaluating ensemble-based data assimilation experiments in cryospheric science <xref ref-type="bibr" rid="bib1.bibx101 bib1.bibx17" id="paren.184"/>. In practice, in line with previous studies <xref ref-type="bibr" rid="bib1.bibx8 bib1.bibx89 bib1.bibx9" id="paren.185"/>, we estimate the CRPS for probabilistic predictions obtained via ensemble-based DA schemes in MuSA by assuming that the marginal predictive prior and posterior snow depth distributions follow a normal distribution. Thereby, we only need to store the mean and standard deviation of the ensemble predictions of each scheme which can be plugged into the simple analytical expression for the Gaussian CRPS given by <xref ref-type="bibr" rid="bib1.bibx56" id="text.186"/>.</p>
      <p id="d2e8776">The CRPS is used to evaluate the performance of the ensemble of snowpack states obtained in MuSA by comparing them to assimilated or independent observations. In a similar vein, we employ the  KLD as a probabilistic evaluation diagnostic that quantifies how close the approximate posterior distributions over parameters are to the reference posterior obtained through MCMC simulation using the RAM method. In particular, we use the so-called reverse KLD defined by <xref ref-type="bibr" rid="bib1.bibx95" id="paren.187"/>

            <disp-formula id="Ch1.E38" content-type="numbered"><label>38</label><mml:math id="M290" display="block"><mml:mtable rowspacing="0.2ex" class="split" displaystyle="true" columnalign="right left"><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mi mathvariant="normal">KL</mml:mi></mml:msub></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mo>(</mml:mo><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo><mml:mo>∣</mml:mo><mml:mo>∣</mml:mo><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mo movablelimits="false">∫</mml:mo><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mi>log⁡</mml:mi><mml:mfenced close=")" open="("><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mfrac></mml:mstyle></mml:mfenced><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mi mathvariant="normal">d</mml:mi><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>

          where <inline-formula><mml:math id="M291" display="inline"><mml:mrow><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>|</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is an approximating distribution of the target posterior distribution <inline-formula><mml:math id="M292" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>|</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. The KLD, also known as relative entropy <xref ref-type="bibr" rid="bib1.bibx83" id="paren.188"/>, is a dimensionless measure of the distance between the two distributions that it takes as arguments. As with the CRPS, it is negatively oriented and non-negative with the best possible value being <inline-formula><mml:math id="M293" display="inline"><mml:mn mathvariant="normal">0</mml:mn></mml:math></inline-formula> in the case that the input distributions used as arguments are equal. It happens to be asymmetric in these arguments, and this so-called reverse form of the KLD is widely used as an objective function in variational Bayesian inference which recasts inference as an optimization problem over tractable variational distributions <inline-formula><mml:math id="M294" display="inline"><mml:mrow><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>|</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> used to approximate the intractable posterior <inline-formula><mml:math id="M295" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>|</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> <xref ref-type="bibr" rid="bib1.bibx95" id="paren.189"/>. While this use of reverse KLD as an objective function motivated our use of this metric, we use it in the slightly different context of evaluation rather than optimization. In particular, we use the reference posterior samples simulated via MCMC using RAM to define the target posterior distribution <inline-formula><mml:math id="M296" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>|</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and we use the reverse KLD to measure how close the more tractable approximations <inline-formula><mml:math id="M297" display="inline"><mml:mrow><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>|</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> from the ensemble-based DA schemes, including the AdaPBS, are to this target. To simplify the calculation of the KLD, especially given that we only have samples from the distributions <inline-formula><mml:math id="M298" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M299" display="inline"><mml:mi>q</mml:mi></mml:math></inline-formula>, we focus on the reverse KLD of the marginal distributions of the perturbation parameters <inline-formula><mml:math id="M300" display="inline"><mml:mi mathvariant="italic">θ</mml:mi></mml:math></inline-formula> in the vector <inline-formula><mml:math id="M301" display="inline"><mml:mi mathvariant="bold-italic">θ</mml:mi></mml:math></inline-formula> in transformed space as introduced in Sect. <xref ref-type="sec" rid="Ch1.S3.SS2"/>. Note that the KLD is invariant to transformations <xref ref-type="bibr" rid="bib1.bibx95" id="paren.190"/>, and by sticking to the transformed unbounded space we can better approximate these marginals as Gaussian. Assuming Gaussianity, the marginal reverse KLD has a simple analytical solution of the form <xref ref-type="bibr" rid="bib1.bibx95" id="paren.191"><named-content content-type="pre">see supplementary material 5.1.2 of</named-content></xref>

            <disp-formula id="Ch1.E39" content-type="numbered"><label>39</label><mml:math id="M302" display="block"><mml:mtable class="split" rowspacing="0.2ex" displaystyle="true" columnalign="right left"><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mi mathvariant="normal">KL</mml:mi></mml:msub></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mo>(</mml:mo><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo><mml:mo>∣</mml:mo><mml:mo>∣</mml:mo><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo><mml:mo>)</mml:mo><mml:mo>≃</mml:mo><mml:mi>log⁡</mml:mi><mml:mfenced close=")" open="("><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi>p</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi>q</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mstyle></mml:mfenced><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:mfrac></mml:mstyle></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mo>+</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi>p</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi>q</mml:mi></mml:msub><mml:msup><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>+</mml:mo><mml:msubsup><mml:mi mathvariant="italic">σ</mml:mi><mml:mi>q</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:msubsup><mml:mi mathvariant="italic">σ</mml:mi><mml:mi>p</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>

          where <inline-formula><mml:math id="M303" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi>q</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M304" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi>q</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> are the mean and standard deviation, respectively, of the ensemble-based approximation <inline-formula><mml:math id="M305" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>∼</mml:mo><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>|</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. In addition to computing reverse KLD of the various ensemble-based DA schemes in MuSA we also compute the reverse KLD of the marginal prior, i.e. setting <inline-formula><mml:math id="M306" display="inline"><mml:mrow><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>|</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, as an additional benchmark to help contextualize this distance measure. For all these reverse KLD calculations, we use the MCMC samples obtained via RAM (Sect. <xref ref-type="sec" rid="Ch1.S2.SS3"/>) to define the reference marginal posterior distribution <inline-formula><mml:math id="M307" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>∼</mml:mo><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>|</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> with reference sample mean <inline-formula><mml:math id="M308" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi>p</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and standard deviation <inline-formula><mml:math id="M309" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi>p</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>.</p>
</sec>
</sec>
<sec id="Ch1.S3">
  <label>3</label><title>Data and methods</title>
<sec id="Ch1.S3.SS1">
  <label>3.1</label><title>Data assimilation framework and experimental design</title>
      <p id="d2e9319">All experiments presented in this study were developed using the Multiple Snow Data Assimilation System (MuSA) <xref ref-type="bibr" rid="bib1.bibx7" id="paren.192"/>. MuSA is a versatile, open-source data assimilation tool designed to integrate diverse snowpack observations with an ensemble of numerical snowpack simulations. Its modular architecture allows users to implement different algorithms and snowpack models. In addition to the set of algorithms already available in MuSA, we have implemented three new algorithms to develop the experiments proposed in this work. This includes the aforementioned AdaPBS algorithm and two MCMC algorithms, namely RWM and RAM, where RAM is used as a sta. The updated MuSA code is released as a new version of the tool <xref ref-type="bibr" rid="bib1.bibx3" id="paren.193"/>. In order to provide an intuitive understanding of the behavior of the AdaPBS, we have performed several experiments that compare its performance to that of previously proposed ensemble-based cryospheric DA algorithms.</p>
</sec>
<sec id="Ch1.S3.SS2">
  <label>3.2</label><title>Experiment 1: Assimilating drone-based snow depth in a temperature index model</title>
      <p id="d2e9336">First, we assimilated snow depth observations in an ensemble of simulations generated by a simple temperature index model <xref ref-type="bibr" rid="bib1.bibx65" id="paren.194"/> implemented in MuSA. We emphasize that the primary objective of these experiments is to assess the performance of the AdaPBS algorithm by exploiting a simpler yet computationally efficient model and not to achieve the most physically complex simulations possible. Moreover, in part due to their simplicity in terms of input data and few calibration parameters, temperature index models are being used effectively in operational settings <xref ref-type="bibr" rid="bib1.bibx82" id="paren.195"/> as well as for hemispheric-scale snow reanalysis <xref ref-type="bibr" rid="bib1.bibx35" id="paren.196"/> and global glacier modeling <xref ref-type="bibr" rid="bib1.bibx111" id="paren.197"/>. Our temperature index forward modeling implementation relies on the strong assumption of constant density, fixed at a  climatological value of 300 kg m<sup>−3</sup>, along with a seasonally constant hourly temperature index factor <inline-formula><mml:math id="M311" display="inline"><mml:mi>a</mml:mi></mml:math></inline-formula> set to <inline-formula><mml:math id="M312" display="inline"><mml:mn mathvariant="normal">0.1375</mml:mn></mml:math></inline-formula> mm h<sup>−1</sup> K<sup>−1</sup>. Typically <inline-formula><mml:math id="M315" display="inline"><mml:mi>a</mml:mi></mml:math></inline-formula> would also be treated as an uncertain parameter <xref ref-type="bibr" rid="bib1.bibx111" id="paren.198"><named-content content-type="pre">e.g.</named-content></xref>, but to facilitate the comparison of schemes and visualizing the results we chose to fix it here. Thus, the near surface air temperature <inline-formula><mml:math id="M316" display="inline"><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mi>n</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> [K] dependent snowmelt rate <inline-formula><mml:math id="M317" display="inline"><mml:mrow><mml:msub><mml:mi>M</mml:mi><mml:mi>n</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> [mm h<sup>−1</sup>] at each hourly timestep <inline-formula><mml:math id="M319" display="inline"><mml:mi>n</mml:mi></mml:math></inline-formula> is computed as follows:

            <disp-formula id="Ch1.E40" content-type="numbered"><label>40</label><mml:math id="M320" display="block"><mml:mrow><mml:msub><mml:mi>M</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mo movablelimits="false">max⁡</mml:mo><mml:mfenced close=")" open="("><mml:mrow><mml:mi>a</mml:mi><mml:mfenced close="]" open="["><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:mi>b</mml:mi><mml:mo>-</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:mfenced><mml:mo>,</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:mfenced><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

          where the melting temperature <inline-formula><mml:math id="M321" display="inline"><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">273.15</mml:mn></mml:mrow></mml:math></inline-formula> [K] is set to the freezing point of water. Often this is treated as a calibration parameter to account for errors in the air temperature forcing, but we do this more explicitly by perturbing the air temperature field itself with an additive bias parameter <inline-formula><mml:math id="M322" display="inline"><mml:mi>b</mml:mi></mml:math></inline-formula> that we seek to infer. By combining the snowmelt rate with the snowfall rate <inline-formula><mml:math id="M323" display="inline"><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mi>n</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> [mm h<sup>−1</sup>], the hourly SWE <inline-formula><mml:math id="M325" display="inline"><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mi>n</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> [mm] is updated via

            <disp-formula id="Ch1.E41" content-type="numbered"><label>41</label><mml:math id="M326" display="block"><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo movablelimits="false">max⁡</mml:mo><mml:mfenced open="(" close=")"><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:mi>c</mml:mi><mml:msub><mml:mi>S</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mi>M</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:mfenced><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

          where <inline-formula><mml:math id="M327" display="inline"><mml:mi>c</mml:mi></mml:math></inline-formula> is the precipitation perturbation parameter that we also seek to infer. Note that the air temperature bias <inline-formula><mml:math id="M328" display="inline"><mml:mi>b</mml:mi></mml:math></inline-formula> also affects the snowfall rate <inline-formula><mml:math id="M329" display="inline"><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mi>n</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> directly because the precipitation phase is diagnosed using the perturbed air temperature <inline-formula><mml:math id="M330" display="inline"><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:mi>b</mml:mi></mml:mrow></mml:math></inline-formula>.</p>
      <p id="d2e9648">The temperature index model DA experiments were conducted in the Izas experimental catchment, in the Central Spanish Pyrenees <xref ref-type="bibr" rid="bib1.bibx106" id="paren.199"/>. The snow depth data were acquired through drone surveys utilizing structure-from-motion techniques <xref ref-type="bibr" rid="bib1.bibx107" id="paren.200"/>. Meteorological forcing data in the form of air temperature and total precipitation (<inline-formula><mml:math id="M331" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula>), were generated using the Micromet downscaling tool <xref ref-type="bibr" rid="bib1.bibx80" id="paren.201"/>, driven by the ERA5 reanalysis <xref ref-type="bibr" rid="bib1.bibx64" id="paren.202"/>. These experiments were conducted for a single grid cell from which we retrieve both the forcing and observations using a grid spacing of 5 m spatial resolution. The cell was located in a concavity within the basin and therefore exhibited significant snow accumulations resulting from snow redistribution that makes it particularly challenging to simulate for the temperature index model with standard forcing. These drone-acquired data and forcing datasets have previously been used in data assimilation studies <xref ref-type="bibr" rid="bib1.bibx7 bib1.bibx8" id="paren.203"/> Following these studies and the findings of <xref ref-type="bibr" rid="bib1.bibx107" id="text.204"/>, we assume an observation error standard deviation of <inline-formula><mml:math id="M332" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi>y</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.2</mml:mn></mml:mrow></mml:math></inline-formula> m for the drone-based snow depth retrievals. The corresponding observation error variance <inline-formula><mml:math id="M333" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="italic">σ</mml:mi><mml:mi>y</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:math></inline-formula> is used to construct the diagonal observation error covariance matrix <inline-formula><mml:math id="M334" display="inline"><mml:mi mathvariant="bold">R</mml:mi></mml:math></inline-formula> in the likelihood Eq. (<xref ref-type="disp-formula" rid="Ch1.E4"/>). The ensemble of simulations was generated by perturbing the precipitation and temperature timeseries with the pseudo-randomly sampled constant in time prior perturbation parameters <inline-formula><mml:math id="M335" display="inline"><mml:mi>b</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M336" display="inline"><mml:mi>c</mml:mi></mml:math></inline-formula> covering the 2018/2019 snow season. The data assimilation window covered the whole snow season and, as such, our experiments were developed using batch smoothers, but we would still expect similar relative performance (i.e. ranking) of these cryospheric DA schemes when applied as filters.</p>
      <p id="d2e9729">For simplicity, and without loss of generality for the DA, we assume a priori parameter independence, such that the joint prior factorizes into the product of its marginal prior distributions. This assumption can be relaxed using additional background knowledge on parameter dependence <xref ref-type="bibr" rid="bib1.bibx102 bib1.bibx8" id="paren.205"><named-content content-type="pre">e.g.,</named-content></xref>. Yet, even with prior independence, correlation between parameters can be inferred in the posterior via the likelihood <xref ref-type="bibr" rid="bib1.bibx17" id="paren.206"/>. The prior probability distribution for the additive air temperature bias is a normal distribution with mean <inline-formula><mml:math id="M337" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi>b</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula> and standard deviation <inline-formula><mml:math id="M338" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi>b</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula>. For the prior on the multiplicative precipitation scaling, we use a strictly positive lognormal distribution with an associated normal mean <inline-formula><mml:math id="M339" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.1</mml:mn></mml:mrow></mml:math></inline-formula> (so the actual median perturbation in model space is <inline-formula><mml:math id="M340" display="inline"><mml:mrow><mml:mi>exp⁡</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1.11</mml:mn></mml:mrow></mml:math></inline-formula>) and standard deviation <inline-formula><mml:math id="M341" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.5</mml:mn></mml:mrow></mml:math></inline-formula>. Note that we carry out inference directly on the unbounded space of the log-transformed precipitation perturbation parameter and then apply an exponential transform to map this back to the model parameter <inline-formula><mml:math id="M342" display="inline"><mml:mi>c</mml:mi></mml:math></inline-formula> following common practice <xref ref-type="bibr" rid="bib1.bibx50 bib1.bibx7" id="paren.207"/>. The location parameters of these prior distributions were chosen in a conservative manner, i.e., near unbiased, with a mean temperature bias of <inline-formula><mml:math id="M343" display="inline"><mml:mn mathvariant="normal">0</mml:mn></mml:math></inline-formula> and a precipitation scaling centered close to <inline-formula><mml:math id="M344" display="inline"><mml:mn mathvariant="normal">1</mml:mn></mml:math></inline-formula>. This corresponds to quite a general case where, without additional background knowledge or data, the orientation of the forcing bias is not known a priori. The relatively large prior spread parameters (<inline-formula><mml:math id="M345" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M346" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi>b</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) encode considerable initial epistemic uncertainty in the two forcing perturbation parameters.</p>
      <p id="d2e9869">Here we have compared the results of the approximate Bayesian inference using different DA algorithms, to evaluate and benchmark the performance of the novel AdaPBS algorithm. First, we have solved the data assimilation problem using adaptive MCMC using RAM <xref ref-type="bibr" rid="bib1.bibx127" id="paren.208"/>, which we consider a gold-standard albeit computationally costly reference. Taking inspiration from <xref ref-type="bibr" rid="bib1.bibx38" id="text.209"/>, we initialized the Markov chain used in RAM at the approximate posterior mean obtained from ES-MDA with <inline-formula><mml:math id="M347" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">a</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">4</mml:mn></mml:mrow></mml:math></inline-formula> iterations and <inline-formula><mml:math id="M348" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">100</mml:mn></mml:mrow></mml:math></inline-formula> ensemble members. The idea is to accelerate MCMC by starting the chain within or close to the typical set of the target posterior distribution. The length of the Markov chain in the RAM algorithm was set to 20 000 and the initial <inline-formula><mml:math id="M349" display="inline"><mml:mn mathvariant="normal">10</mml:mn></mml:math></inline-formula>  % of the samples were discarded as burn in. Discarding the burn in period avoids initialization artifacts whereas the relatively long chain gives the Markov Chain time to reach its target posterior distribution.</p>
      <p id="d2e9916">We then performed the same exercise using a Particle Batch Smoother (PBS), an ES, an ES-MDA and finally we ran the AdaPBS. The hyperparameters of the analysis (i.e. the <inline-formula><mml:math id="M350" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">100</mml:mn></mml:mrow></mml:math></inline-formula> number of ensemble members and the prior distributions) were the same in all cases. In the case of the AdaPBS, the <inline-formula><mml:math id="M351" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">eff</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> limit was set to <inline-formula><mml:math id="M352" display="inline"><mml:mn mathvariant="normal">30</mml:mn></mml:math></inline-formula>  %, which means that for our <inline-formula><mml:math id="M353" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">100</mml:mn></mml:mrow></mml:math></inline-formula> member ensemble at least 30 particles should show non-negligible weights before stopping the iterations. The results of each algorithm were compared with the samples from MCMC-RAM which we used as a reference posterior. The agreement between the approximate posteriors obtained via ensemble-based DA  and the MCMC reference was quantified using the KLD as a probabilistic evaluation diagnostic as described in Sect. <xref ref-type="sec" rid="Ch1.S2.SS7"/>.</p>
</sec>
<sec id="Ch1.S3.SS3">
  <label>3.3</label><title>Experiment 2: Assimilating hourly ESM-SnowMIP snow depth data in FSM2</title>
      <p id="d2e9977">In the second set of experiments, we conducted several site-level simulations using a physics-based snow model of intermediate complexity, the Flexible Snow Model <xref ref-type="bibr" rid="bib1.bibx41" id="paren.210"><named-content content-type="pre">FSM2;</named-content></xref>. In these experiments, we assimilated hourly observations of snow depth at six different geographical locations using both AdaPBS and ES-MDA for comparison. For simplicity, we assume the same observation error standard deviation of <inline-formula><mml:math id="M354" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi>y</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.2</mml:mn></mml:mrow></mml:math></inline-formula> m as in Experiment 1. While this is likely to be an overestimate of individual measurement errors at these sites, inflating independent observation errors in our model helps implicitly account for the fact that the actual errors in these data are correlated in time, which reduces their information content <xref ref-type="bibr" rid="bib1.bibx44" id="paren.211"/>. Assimilating snow depth observations at different locations over different periods of time allowed us to obtain a general overview of the behavior of the algorithms in different climates and weather regimes. Observations were obtained from a database generated under the framework of the Earth System Model-Snow Model Intercomparison Project <xref ref-type="bibr" rid="bib1.bibx91" id="paren.212"><named-content content-type="pre">ESM-SnowMIP;</named-content></xref>. Specifically, we conducted simulations at Reynolds Mountain East, Idaho, USA (RME); Sapporo, Japan (SAP); Senator Beck, Colorado, USA (SNB); Sodankylä, Finland (SOD); Swamp Angel, Colorado, USA (SWE) and Weissfluhjoch, Switzerland (WFJ). The meteorological forcing was obtained directly from the nearest grid cell of the ERA5 reanalysis <xref ref-type="bibr" rid="bib1.bibx64" id="paren.213"/> without applying topographic corrections. While site-level meteorological forcing was available from <xref ref-type="bibr" rid="bib1.bibx91" id="text.214"/>, we used the more uncertain, unadjusted global reanalysis data to provide a more challenging experimental setup that is generally applicable to any site with snow depth data. Specifically, each water year, we assimilated a large batch of <inline-formula><mml:math id="M355" display="inline"><mml:mrow><mml:mn mathvariant="normal">24</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">365</mml:mn><mml:mo>=</mml:mo><mml:mn mathvariant="normal">8760</mml:mn></mml:mrow></mml:math></inline-formula> hourly snow depth observations.</p>
      <p id="d2e10031">To explore the performance of the algorithms in the case of a higher dimensional parameter space, we have generated 7 perturbation parameters to correct the forcing (air temperature, precipitation, relative humidity, surface pressure, wind speed, and incoming long- and short-wave radiations). On top of this, we have also included the 12 main tunable internal model parameters related to snow (not vegetation) in FSM2 <xref ref-type="bibr" rid="bib1.bibx41" id="paren.215"/> as listed in Table <xref ref-type="table" rid="T1"/>. The prior distributions for the temperature and precipitation parameters are the same as those for the previous experiment described in Sect. <xref ref-type="sec" rid="Ch1.S3.SS2"/>. For the new model parameters, we used a multiplicative logit-normal distribution bounded in the physical space between 0.8 and 1.2, defined by the moments of its underlying normal distribution <inline-formula><mml:math id="M356" display="inline"><mml:mrow><mml:mi mathvariant="italic">μ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M357" display="inline"><mml:mrow><mml:mi mathvariant="italic">σ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula>. As such, we perturb these new parameters in the <inline-formula><mml:math id="M358" display="inline"><mml:mrow><mml:mo>±</mml:mo><mml:mn mathvariant="normal">20</mml:mn></mml:mrow></mml:math></inline-formula>  % range from their reference values in the physical space, ensuring that the combinations parameters remain in their physical bounds, maintaining the numerical stability of FSM2. We performed a dependent validation to compare the performance of AdaPBS with ES-MDA and PBS, by comparing the posterior snow simulations with the assimilated observations, using bias and root mean squared error (RMSE) as well as an uncertainty-aware metric in the form of the continuous ranked probability score (CRPS) described in Sect. <xref ref-type="sec" rid="Ch1.S2.SS7"/>. We removed the timesteps where <italic>both</italic> the simulation and observations had snow depths of zero in the computation of the error and CRPS to avoid artificially deflating the validation metrics by not considering the seasonality of the snow depth values. That is to say, we are primarily concerned with model performance in the periods in which the model or observations indicate the presence of a seasonal snowpack. It is only in these periods that the model is susceptible to snow commission or omission errors rather than the mostly trivially correct no-snow prediction outside the snow season.</p>
      <p id="d2e10081">Unlike Experiment 1 with the temperature index model, we did not run MCMC in this more complex experiment due to the high computational cost and marginal additional benefits to the analysis. As such, the KLD metric could not be used without an MCMC-based reference distribution. Instead, this experiment focused on comparing the predictive performance in state space of AdaPBS to that of ES-MDA, an advanced and competitive baseline for DA with more complex snow models <xref ref-type="bibr" rid="bib1.bibx7 bib1.bibx9" id="paren.216"/>. In this higher-dimensional experiment, the entire 19-dimensional parameter space is too large for comprehensive and instructive visualization. An alternative could be to investigate the importance of each parameter by estimating the information gain <xref ref-type="bibr" rid="bib1.bibx123" id="paren.217"/>, but such parameter analysis strays beyond the focus of this experiment. Except for a handful of forcing parameters partly analyzed in Experiment 1, the majority of the 19 parameters are so-called nuisance parameters <xref ref-type="bibr" rid="bib1.bibx69" id="paren.218"/>. They are not of primary interest beyond their effect on the state and their role in making the DA problem more challenging. We thus focus on the performance of AdaPBS and ES-MDA for estimating the snow depth state observed at the selected ESM-SnowMIP sites. The non-iterative PBS is also tested in this experiment to directly assess the impact of the adaptations within the iterative AdaPBS.</p>

<table-wrap id="T1"><label>Table 1</label><caption><p id="d2e10097">Tabular overview of the internal FSM2 parameters that we perturbed at ESM-SnowMIP sites. Columns from left to right: the parameter name in the FSM2 code, its default value, physical units (– is dimensionless), and a description adapted from <xref ref-type="bibr" rid="bib1.bibx40" id="text.219"/>.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="4">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="left"/>
     <oasis:colspec colnum="3" colname="col3" align="left"/>
     <oasis:colspec colnum="4" colname="col4" align="left"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Name</oasis:entry>
         <oasis:entry colname="col2">Default</oasis:entry>
         <oasis:entry colname="col3">Units</oasis:entry>
         <oasis:entry colname="col4">Description</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1"><monospace>asmn</monospace></oasis:entry>
         <oasis:entry colname="col2">0.5</oasis:entry>
         <oasis:entry colname="col3">–</oasis:entry>
         <oasis:entry colname="col4">Min. albedo for melting snow</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"><monospace>asmx</monospace></oasis:entry>
         <oasis:entry colname="col2">0.85</oasis:entry>
         <oasis:entry colname="col3">–</oasis:entry>
         <oasis:entry colname="col4">Max. albedo for fresh snow</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"><monospace>eta0</monospace></oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M359" display="inline"><mml:mrow><mml:mn mathvariant="normal">3.7</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mn mathvariant="normal">7</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">Pa s</oasis:entry>
         <oasis:entry colname="col4">Reference snow viscosity</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"><monospace>hfsn</monospace></oasis:entry>
         <oasis:entry colname="col2">0.1</oasis:entry>
         <oasis:entry colname="col3">m</oasis:entry>
         <oasis:entry colname="col4">Snow cover fraction depth scale</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"><monospace>rgr0</monospace></oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M360" display="inline"><mml:mrow><mml:mn mathvariant="normal">5</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">5</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">m</oasis:entry>
         <oasis:entry colname="col4">Fresh snow grain radius</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"><monospace>rhow</monospace></oasis:entry>
         <oasis:entry colname="col2">300</oasis:entry>
         <oasis:entry colname="col3">kg m<sup>−3</sup></oasis:entry>
         <oasis:entry colname="col4">Wind-packed snow density</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"><monospace>Salb</monospace></oasis:entry>
         <oasis:entry colname="col2">10</oasis:entry>
         <oasis:entry colname="col3">kg m<sup>−2</sup></oasis:entry>
         <oasis:entry colname="col4">Snowfall to refresh albedo</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"><monospace>snda</monospace></oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M363" display="inline"><mml:mrow><mml:mn mathvariant="normal">2.8</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">6</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">s<sup>−1</sup></oasis:entry>
         <oasis:entry colname="col4">Thermal metamorphism parameter</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"><monospace>tcld</monospace></oasis:entry>
         <oasis:entry colname="col2">1000</oasis:entry>
         <oasis:entry colname="col3">h</oasis:entry>
         <oasis:entry colname="col4">Cold snow albedo decay time scale</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"><monospace>tmlt</monospace></oasis:entry>
         <oasis:entry colname="col2">100</oasis:entry>
         <oasis:entry colname="col3">h</oasis:entry>
         <oasis:entry colname="col4">Melting snow albedo decay time scale</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"><monospace>Wirr</monospace></oasis:entry>
         <oasis:entry colname="col2">0.03</oasis:entry>
         <oasis:entry colname="col3">–</oasis:entry>
         <oasis:entry colname="col4">Irreducible liquid water content of snow</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"><monospace>z0sn</monospace></oasis:entry>
         <oasis:entry colname="col2">0.001</oasis:entry>
         <oasis:entry colname="col3">m</oasis:entry>
         <oasis:entry colname="col4">Snow surface roughness length</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

</sec>
</sec>
<sec id="Ch1.S4">
  <label>4</label><title>Results and discussion</title>
<sec id="Ch1.S4.SS1">
  <label>4.1</label><title>Inter-comparison of algorithms using a temperature index model</title>
      <p id="d2e10426">Despite relying on identical priors, observations, observation error model, and forward model in the form of a temperature index snow model, the performance of the DA algorithms differed substantially. This conclusion aligns with the few previous publications that developed inter-comparison experiments on ensemble-based cryospheric data assimilation algorithms <xref ref-type="bibr" rid="bib1.bibx78 bib1.bibx85 bib1.bibx1 bib1.bibx7 bib1.bibx17" id="paren.220"/> and highlights the importance of selecting the algorithm with careful consideration of the problem at hand by balancing accuracy and computational cost.</p>
      <p id="d2e10432">Given the highly informative nature of the drone-based snow depth observations, the posterior approximation obtained through MCMC samples via RAM showed a challenging, narrow, banana-shaped distribution (Fig. <xref ref-type="fig" rid="F2"/>). This demonstrates that even in this relatively simple case – where in practice we are only updating <inline-formula><mml:math id="M365" display="inline"><mml:mn mathvariant="normal">2</mml:mn></mml:math></inline-formula> parameters (<inline-formula><mml:math id="M366" display="inline"><mml:mi>b</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M367" display="inline"><mml:mi>c</mml:mi></mml:math></inline-formula>) by assimilating just a handful (<inline-formula><mml:math id="M368" display="inline"><mml:mn mathvariant="normal">5</mml:mn></mml:math></inline-formula>) of snow depth observations – efficiently sampling from the posterior is nonetheless a challenging task. This is probably exacerbated in this case by the fact that we are using batch smoothers. On the one hand, such smoothers can provide more information than filters by assimilating the complete trajectory of observations in the water year all at once and propagating information backwards in time. On the other hand, the larger immediate information gain from batch smoothers can make posterior sampling more challenging than with sequential filters.</p>
      <p id="d2e10465">The original PBS algorithm of <xref ref-type="bibr" rid="bib1.bibx85" id="text.221"/> has several benefits. It is relatively easy to implement, relies on few assumptions, and its computational cost is lower than that of other schemes, even when considering the same number of members in the ensemble since (unlike the ES) it does not require any reruns. These benefits do come at a cost, as PBS is prone to collapse in more challenging settings. This means that sometimes, especially with very informative (i.e. numerous and/or accurate) observations, the ensemble collapses and degeneracy ensues <xref ref-type="bibr" rid="bib1.bibx117 bib1.bibx94" id="paren.222"/>. Using a familiar analogy, this is like looking for smaller needles in a haystack. A similar result arises if the priors are relatively far (as measured, e.g., by a modified variant of the KLD in Sect. <xref ref-type="sec" rid="Ch1.S2.SS7"/>) from the posterior in the parameter space, such as in the case of overconfident and/or strongly biased priors. More generally, “far” is not just the distance between the prior and posterior mean, but whenever the overlap between the prior and the posterior is small, including when a broad and/or high-dimensional prior encompasses a narrow posterior. Following a similar analogy, this is akin to having a larger haystack. These needle and haystack effects can yield sub-optimal solutions with high Monte Carlo variance <xref ref-type="bibr" rid="bib1.bibx110" id="paren.223"/> that make it difficult, if not impossible, to accurately estimate uncertainty. So even if SNIS, which powers the PBS, is asymptotically unbiased <xref ref-type="bibr" rid="bib1.bibx104" id="paren.224"/>, the PBS in particular and basic importance sampling in general may exhibit a large variance since the prior is implicitly used as the proposal, which is generally far from the posterior <xref ref-type="bibr" rid="bib1.bibx83" id="paren.225"/>. In the worst case, this will result in complete particle degeneracy with <inline-formula><mml:math id="M369" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">eff</mml:mi></mml:msub><mml:mo>≃</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula>. This was clearly the case in our particular experiment, as shown in Fig. <xref ref-type="fig" rid="F2"/>, where the majority of the weight collapsed onto a single particle despite the fact that the range of the prior parameter distribution clearly encompassed the reference posterior MCMC samples obtained via the RAM algorithm. In the model state space, the fact that a single particle carries all the probability leads to a degenerate posterior snow depth ensemble that is not close to the assimilated observations and completely lacks well-calibrated uncertainty quantification.</p>
      <p id="d2e10503">In this case, the skewed prior on the precipitation perturbation combined with the challenging banana shape of the posterior are key challenges leading to degeneracy since very few particles were sampled in the range of the typical set of the posterior. The situation would improve if a more informative prior were used with a higher median precipitation perturbation parameter and even a positive correlation between the perturbations to better sample the banana shape. The problem, of course, is that these prior hyperparameter settings are rarely, if ever, known a priori, even from expert knowledge, and setting them directly based on the data leads to incoherent “double-dipping” via so-called empirical Bayes <xref ref-type="bibr" rid="bib1.bibx109 bib1.bibx50 bib1.bibx95" id="paren.226"/>. Adaptive particle methods present a solution that, instead of incoherently changing the prior to avoid degeneracy, simply adapts the proposal towards the target posterior so as to perform more sample efficient approximate Bayesian inference that remains coherent.</p>
      <p id="d2e10510">It may seem surprising that importance sampling via PBS struggles to accurately capture the reference posterior in this apparently simple experiment with a two-dimensional parameter space. An alternative would be gradient-based optimization techniques, such as gradient descent or Newton's method, used in both machine learning <xref ref-type="bibr" rid="bib1.bibx95" id="paren.227"/> and variational DA <xref ref-type="bibr" rid="bib1.bibx10" id="paren.228"/>, to identify the mode with relatively few iterations and find a Laplace approximation of the posterior <xref ref-type="bibr" rid="bib1.bibx83" id="paren.229"/>. However, as is often the case in practical DA in the absence of automatic differentiation <xref ref-type="bibr" rid="bib1.bibx95" id="paren.230"/>, we do not have access to gradients and are focused on the harder problem of “black-box” (i.e., gradient-free) inference that is naturally handled by ensemble-based methods <xref ref-type="bibr" rid="bib1.bibx112" id="paren.231"/>. While similar low-dimensional calibration problems are routinely encountered in hydrology and can be tackled via methods such as generalized likelihood uncertainty estimation <xref ref-type="bibr" rid="bib1.bibx14" id="paren.232"><named-content content-type="pre">GLUE;</named-content></xref>, SCEM-UA optimization <xref ref-type="bibr" rid="bib1.bibx129" id="paren.233"/>, or DREAM MCMC sampling <xref ref-type="bibr" rid="bib1.bibx130" id="paren.234"/>, Nonetheless, these methods usually involve orders of magnitude more forward model evaluations than the <inline-formula><mml:math id="M370" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">100</mml:mn></mml:mrow></mml:math></inline-formula> particles used by PBS in this experiment. For example, <xref ref-type="bibr" rid="bib1.bibx22" id="text.235"/> demonstrated the need for a computationally intensive SODA approach, using SCEM-UA in an outer loop to optimize parameters, together with an EnKF inner loop for state estimation, to rigorously calibrate a simple two-parameter snow model. Furthermore, the widely used hydrological calibration technique GLUE <xref ref-type="bibr" rid="bib1.bibx14" id="paren.236"/>, tailored to identify equifinality <xref ref-type="bibr" rid="bib1.bibx13" id="paren.237"/>, is equivalent to PBS when using a Gaussian likelihood. As such, GLUE could only outperform PBS by increasing the ensemble size or modifying the likelihood, which are equally applicable to PBS. Both black-box optimization and inference are challenging problems where particle methods have proven to be highly competitive <xref ref-type="bibr" rid="bib1.bibx33 bib1.bibx21" id="paren.238"/>. In fact, these particle methods can be seen as formalizing highly successful meta-heuristic evolutionary algorithms that are widely used for optimization <xref ref-type="bibr" rid="bib1.bibx30" id="paren.239"/>. The takeaway here is that particle methods can be improved either by increasing the ensemble size, at considerable computational expense given typically poor scaling <xref ref-type="bibr" rid="bib1.bibx117" id="paren.240"/>, or by adapting them for example using various iterative methods <xref ref-type="bibr" rid="bib1.bibx15 bib1.bibx21" id="paren.241"/>.</p>

      <fig id="F2" specific-use="star"><label>Figure 2</label><caption><p id="d2e10579">Results from the particle batch smoother (PBS) assimilating drone-based snow depth observations in a temperature index model for a single cell at Izas in water year 2019: The model parameter space visualized as the temperature bias parameter <inline-formula><mml:math id="M371" display="inline"><mml:mi>b</mml:mi></mml:math></inline-formula> on the <inline-formula><mml:math id="M372" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula>-axis and the precipitation correction <inline-formula><mml:math id="M373" display="inline"><mml:mi>c</mml:mi></mml:math></inline-formula> on the <inline-formula><mml:math id="M374" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula>-axis with the prior ensemble of particles shown with red dots while <inline-formula><mml:math id="M375" display="inline"><mml:mn mathvariant="normal">2</mml:mn></mml:math></inline-formula> green stars with size proportional to weight indicate particles with non-negligible PBS weights (<inline-formula><mml:math id="M376" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">eff</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1.26</mml:mn></mml:mrow></mml:math></inline-formula>) and the blue banana-shaped distribution shows the reference posterior MCMC samples obtained via RAM (left panel); The model state space (right panel) showing the trajectory of snow depth with mean (solid line) <inline-formula><mml:math id="M377" display="inline"><mml:mrow><mml:mo>±</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> standard deviation (shading)  from the prior (orange), PBS posterior (green), and reference MCMC posterior (blue), along with the assimilated observations (yellow dots).</p></caption>
          <graphic xlink:href="https://gmd.copernicus.org/articles/19/8565/2026/gmd-19-8565-2026-f02.png"/>

        </fig>

      <p id="d2e10649">The ES <xref ref-type="bibr" rid="bib1.bibx125" id="paren.242"/>, i.e. the batch smoother version of the EnKF <xref ref-type="bibr" rid="bib1.bibx43" id="paren.243"/>, has also been widely used in cryospheric DA particularly in reanalysis settings <xref ref-type="bibr" rid="bib1.bibx34 bib1.bibx52" id="paren.244"/>. Although the EnKF has mainly been used to directly update snow model states <xref ref-type="bibr" rid="bib1.bibx28" id="paren.245"><named-content content-type="pre">e.g.</named-content></xref>, the ES is better suited to update the forcing and internal parameters in a forcing formulation of the data assimilation problem <xref ref-type="bibr" rid="bib1.bibx44" id="paren.246"/>. Since the update moves the parameters in parameter space through the ensemble Kalman analysis step, with the ES it is necessary to generate an ensemble of snow model simulations twice, once with the prior parameters and once with the posterior parameters. This has the obvious advantage of maintaining model stability and physical consistency as long as the parameters remain within their physical bounds <xref ref-type="bibr" rid="bib1.bibx7" id="paren.247"/>. The latter boundedness can be easily controlled using transformation techniques such as Gaussian anamorphosis that also help to accommodate the Gaussian assumption <xref ref-type="bibr" rid="bib1.bibx12 bib1.bibx1" id="paren.248"/>. However, with the ES this forcing formulation approach entails a non-parallelizable doubling of the computational cost (<inline-formula><mml:math id="M378" display="inline"><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> simulations) compared to solving the same problem with the same ensemble size using the PBS (<inline-formula><mml:math id="M379" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> simulations). In practice, the fact that the ES is less prone to ensemble collapse allows one to reduce the ensemble size compared to PBS such that in some problems the ES may actually end up being less costly. This gives a lot of flexibility when configuring the assimilation system according to the available computational resources especially in spatio-temporal settings where localization methods are more easily implemented with ensemble Kalman schemes <xref ref-type="bibr" rid="bib1.bibx44" id="paren.249"/>. Despite these advantages, the linear assumption in combination with models used in cryospheric sciences, which are typically non-linear, can lead to highly suboptimal posterior approximations with poorly calibrated and even underconfident uncertainty estimates. In fact, this is what we see in Fig. <xref ref-type="fig" rid="F3"/> with slight and underconfident updates both in parameter and state space, in line with previous studies <xref ref-type="bibr" rid="bib1.bibx1 bib1.bibx7" id="paren.250"/>.</p>

      <fig id="F3" specific-use="star"><label>Figure 3</label><caption><p id="d2e10711">Analogous to Fig. <xref ref-type="fig" rid="F2"/> but for the ensemble smoother (ES) with the prior (top left) and posterior (top right) in the model parameter space and the trajectory of snow depth (bottom) in the model state space.</p></caption>
          <graphic xlink:href="https://gmd.copernicus.org/articles/19/8565/2026/gmd-19-8565-2026-f03.png"/>

        </fig>

      <p id="d2e10722">The typical underlying linear assumption of ES, which is the main cause of suboptimal results, can be strongly violated in practical cryospheric DA problems. The iterative ES-MDA <xref ref-type="bibr" rid="bib1.bibx39" id="paren.251"/> has proven to be a powerful cryospheric DA algorithm in the few previous inter-comparisons that have been carried out to date <xref ref-type="bibr" rid="bib1.bibx1 bib1.bibx7" id="paren.252"/>. Iterations allow for a progressive movement of parameters towards the posterior via likelihood tempering <xref ref-type="bibr" rid="bib1.bibx21 bib1.bibx95" id="paren.253"/>, which in practice limits the effects of model non-linearity as shown in Fig. <xref ref-type="fig" rid="F4"/>. One limitation is that the number of iterations must be pre-set in the ES-MDA of <xref ref-type="bibr" rid="bib1.bibx39" id="text.254"/>, making it an important hyperparameter that must be carefully adjusted balancing between computational cost and solution robustness. When implemented in a distributed cell-by-cell manner, this is problematic, as it is difficult to adjust the number of iterations cell by cell, opting in practice for a fixed number of iterations for the whole domain. Such a global “one size fits all” approach is very likely to result in a waste of computational resources through an unnecessary number of model realizations compared to a locally adaptive approach. In addition, the computational cost of the linear algebra operations needed for ES-MDA may become non-negligible in the case of assimilating a large number of observations.</p>

      <fig id="F4" specific-use="star"><label>Figure 4</label><caption><p id="d2e10742">Analogous to Fig. <xref ref-type="fig" rid="F3"/> but for an iterative ensemble smoother (IES) in the form of the ensemble smoother with multiple data assimilation (ES-MDA). The top panels show the evolution from the prior to the posterior ensemble members (red) across the MDA iterations along with the reference banana-shaped MCMC posterior samples (blue). As before, the bottom panel shows the predicted snow depth in model state space, but with a better calibrated IES posterior (green). The zero-based iteration counter from <inline-formula><mml:math id="M380" display="inline"><mml:mn mathvariant="normal">0</mml:mn></mml:math></inline-formula> to <inline-formula><mml:math id="M381" display="inline"><mml:mn mathvariant="normal">4</mml:mn></mml:math></inline-formula> corresponds to <inline-formula><mml:math id="M382" display="inline"><mml:mi mathvariant="normal">ℓ</mml:mi></mml:math></inline-formula> in the algorithm description in Sect. <xref ref-type="sec" rid="Ch1.S2.SS4"/> with <inline-formula><mml:math id="M383" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">a</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">4</mml:mn></mml:mrow></mml:math></inline-formula> iterations of the update step and <inline-formula><mml:math id="M384" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">a</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>=</mml:mo><mml:mn mathvariant="normal">5</mml:mn></mml:mrow></mml:math></inline-formula> iterations of the prediction step.</p></caption>
          <graphic xlink:href="https://gmd.copernicus.org/articles/19/8565/2026/gmd-19-8565-2026-f04.png"/>

        </fig>

      <p id="d2e10811">Here we introduce a new algorithm, namely the AdaPBS, as an effort to combine the advantages of the different algorithms presented above into a single tool. The AdaPBS has the potential to be a powerful iterative DA method, aspiring to sample from the posterior with performance similar to that of ES-MDA. Unlike the ES-MDA, AdaPBS does not rely on assumptions of linearity or Gaussianity, which facilitates its implementation, especially for users with limited experience in data assimilation. This also opens up for the possibility of using tailored likelihood functions such as those involving zero-inflation <xref ref-type="bibr" rid="bib1.bibx115 bib1.bibx121" id="paren.255"/> that may be better suited to double-bounded cryospheric variables.</p>

      <fig id="F5" specific-use="star"><label>Figure 5</label><caption><p id="d2e10819">Analogous to Fig. <xref ref-type="fig" rid="F4"/> but for the adaptive particle batch smoother (AdaPBS) showing the particle and effective sample size (ESS, i.e., <inline-formula><mml:math id="M385" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">eff</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) evolution across iterations. Note that the results after the initial iteration (top left) correspond exactly (ignoring Monte Carlo variance) to those that would be obtained from the PBS in Fig. <xref ref-type="fig" rid="F2"/>. The zero-based iteration counter from <inline-formula><mml:math id="M386" display="inline"><mml:mn mathvariant="normal">0</mml:mn></mml:math></inline-formula> to <inline-formula><mml:math id="M387" display="inline"><mml:mn mathvariant="normal">4</mml:mn></mml:math></inline-formula> corresponds to <inline-formula><mml:math id="M388" display="inline"><mml:mrow><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> in the algorithm description in Sect. <xref ref-type="sec" rid="Ch1.S2.SS6"/>, so here the AdaPBS converged (here <inline-formula><mml:math id="M389" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">eff</mml:mi></mml:msub><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">30</mml:mn></mml:mrow></mml:math></inline-formula>) in <inline-formula><mml:math id="M390" display="inline"><mml:mrow><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">5</mml:mn></mml:mrow></mml:math></inline-formula> iterations with exactly the same computational cost as the ES-MDA with <inline-formula><mml:math id="M391" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">a</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>=</mml:mo><mml:mn mathvariant="normal">5</mml:mn></mml:mrow></mml:math></inline-formula> iterations of the prediction step.</p></caption>
          <graphic xlink:href="https://gmd.copernicus.org/articles/19/8565/2026/gmd-19-8565-2026-f05.png"/>

        </fig>

      <p id="d2e10918">A clear computational advantage of the AdaPBS over ES-MDA is that it does not require pre-selecting a fixed number of iterations, enabling the use of early stopping strategies throughout the simulation domain. This allows the number of iterations to be automatically adapted for each cell in a spatio-temporally distributed model domain, depending on the difficulty of the local (or localized) inference problem being solved. We have used an early stopping criterion that stops the iterations when a threshold on the classical estimate of the ESS (<inline-formula><mml:math id="M392" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">eff</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> in Eq. <xref ref-type="disp-formula" rid="Ch1.E22"/>), i.e., a minimum number of ensemble members with considerable weight is reached, making it more robust to collapse than PBS. In this case, we used a relatively high diversity threshold of <inline-formula><mml:math id="M393" display="inline"><mml:mrow><mml:mi mathvariant="italic">τ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.3</mml:mn></mml:mrow></mml:math></inline-formula> for <inline-formula><mml:math id="M394" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">eff</mml:mi></mml:msub><mml:mo>/</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, which in practice resulted in the same number of iterations as ES-MDA. However, it should be noted that in this example, a respectable <inline-formula><mml:math id="M395" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">eff</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">22</mml:mn></mml:mrow></mml:math></inline-formula> was already achieved by iteration 3 as shown in Fig. <xref ref-type="fig" rid="F5"/>, which may be sufficient for many applications.</p>
      <p id="d2e10982">Thus, a key hyperparameter to set in AdaPBS is the diversity threshold <inline-formula><mml:math id="M396" display="inline"><mml:mi mathvariant="italic">τ</mml:mi></mml:math></inline-formula>, which controls the target minimum effective sample size <inline-formula><mml:math id="M397" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">eff</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. Since <inline-formula><mml:math id="M398" display="inline"><mml:mi mathvariant="italic">τ</mml:mi></mml:math></inline-formula> is directly related to the desired performance of the algorithm, it is more interpretable for users than the fixed number of iterations <inline-formula><mml:math id="M399" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">a</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> required in ES-MDA <xref ref-type="bibr" rid="bib1.bibx39" id="paren.256"/>. This makes AdaPBS an easy-to-implement algorithm, resistant to collapse, and potentially able to significantly reduce computational cost compared to ES-MDA when applied to large domains with many cells. Furthermore, even for very challenging cases with AdaPBS struggling to reach the threshold diversity <inline-formula><mml:math id="M400" display="inline"><mml:mi mathvariant="italic">τ</mml:mi></mml:math></inline-formula>, the algorithm will still terminate after a predefined maximum number of iterations <inline-formula><mml:math id="M401" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> which can be allocated depending on the computational budget at hand. Even in these challenging edge cases we have found that although the particle diversity and resulting uncertainty quantification may be worse than desired, the AdaPBS method will by construction still typically provide more accurate and reliable inference than the PBS.</p>
      <p id="d2e11044">In Table <xref ref-type="table" rid="T2"/> we compare the KLD distance measures for the marginal posterior distributions obtained for temperature and precipitation parameters using the respective algorithms. This shows that although AdaPBS is not capable of reaching precisely the same low KLD values as the ES-MDA, it shows quite a comparable performance to this state-of-the art iterative ensemble Kalman method. In fact, the quantitative differences appear minimal when the distributions are compared graphically in Fig. <xref ref-type="fig" rid="F6"/>. More generally, it is clear that the iterative methods, namely the AdaPBS and ES-MDA, are much closer to being able to match the performance of MCMC than their non-iterative counterparts, namely the PBS and ES. Crucially, the AdaPBS can adapt to the complexity of the problem at hand and can thus be guaranteed to incur a lower computational or in the worst case equal cost to the ES-MDA by setting <inline-formula><mml:math id="M402" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>max⁡</mml:mo></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">a</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula>. This cost difference is highlighted in the cost column of Table <xref ref-type="table" rid="T2"/>, where in the best case with a less complex target posterior the AdaPBS has the chance to trigger early stopping already after a single iteration (<inline-formula><mml:math id="M403" display="inline"><mml:mrow><mml:mi>L</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula>) and computationally affordable as the PBS on which it is based. Moreover, like the PBS, the AdaPBS can handle more general probabilistic models than the Gaussian models assumed by the ES-MDA and other ensemble Kalman methods.</p>

<table-wrap id="T2" specific-use="star"><label>Table 2</label><caption><p id="d2e11090">Performance of cryospheric DA algorithms in terms of inferential accuracy and cost. Accuracy is gauged in terms of how well the posterior approximations <inline-formula><mml:math id="M404" display="inline"><mml:mi>q</mml:mi></mml:math></inline-formula> from the algorithms match the reference posterior estimate <inline-formula><mml:math id="M405" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula> from RAM as measured by the Gaussian approximation of the marginal reverse KLD <inline-formula><mml:math id="M406" display="inline"><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mi mathvariant="normal">KL</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>q</mml:mi><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mi>p</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> in Eq. (<xref ref-type="disp-formula" rid="Ch1.E39"/>) for air temperature bias <inline-formula><mml:math id="M407" display="inline"><mml:mi>b</mml:mi></mml:math></inline-formula> and the precipitation correction <inline-formula><mml:math id="M408" display="inline"><mml:mi>c</mml:mi></mml:math></inline-formula> in the temperature index model experiments. Computational cost is measured in terms of the number of forward model runs as multiples of the ensemble size <inline-formula><mml:math id="M409" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">100</mml:mn></mml:mrow></mml:math></inline-formula>, ES-MDA iterations <inline-formula><mml:math id="M410" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">a</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">4</mml:mn></mml:mrow></mml:math></inline-formula>, AdaPBS iterations <inline-formula><mml:math id="M411" display="inline"><mml:mrow><mml:mi>L</mml:mi><mml:mo>≤</mml:mo><mml:msub><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> where typically <inline-formula><mml:math id="M412" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">a</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula>, and the number of Markov chain steps in RAM <inline-formula><mml:math id="M413" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">s</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">20</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mn mathvariant="normal">3</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>. </p></caption><oasis:table frame="topbot"><oasis:tgroup cols="4">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="left"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Algorithm</oasis:entry>
         <oasis:entry colname="col2">Temperature <inline-formula><mml:math id="M414" display="inline"><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mi mathvariant="normal">KL</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>q</mml:mi><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mi>p</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">Precipitation  <inline-formula><mml:math id="M415" display="inline"><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mi mathvariant="normal">KL</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>q</mml:mi><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mi>p</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">Cost</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">Prior</oasis:entry>
         <oasis:entry colname="col2">625.80</oasis:entry>
         <oasis:entry colname="col3">52 893.24</oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M416" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">PBS</oasis:entry>
         <oasis:entry colname="col2">85.20</oasis:entry>
         <oasis:entry colname="col3">186.69</oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M417" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">ES</oasis:entry>
         <oasis:entry colname="col2">878.32</oasis:entry>
         <oasis:entry colname="col3">8418.42</oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M418" display="inline"><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">IES (ES-MDA)</oasis:entry>
         <oasis:entry colname="col2">3.60</oasis:entry>
         <oasis:entry colname="col3">27.66</oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M419" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">a</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>)</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">AdaPBS</oasis:entry>
         <oasis:entry colname="col2">5.59</oasis:entry>
         <oasis:entry colname="col3">47.31</oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M420" display="inline"><mml:mrow><mml:mi>L</mml:mi><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub><mml:mo>≤</mml:mo><mml:msub><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">RAM (<inline-formula><mml:math id="M421" display="inline"><mml:mrow><mml:mi>q</mml:mi><mml:mo>=</mml:mo><mml:mi>p</mml:mi></mml:mrow></mml:math></inline-formula>)</oasis:entry>
         <oasis:entry colname="col2">0</oasis:entry>
         <oasis:entry colname="col3">0</oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M422" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">s</mml:mi></mml:msub><mml:mo>≫</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <fig id="F6" specific-use="star"><label>Figure 6</label><caption><p id="d2e11510">Comparison of Gaussian kernel density estimates (KDEs) obtained via <monospace>scipy.stats.gaussian_kde</monospace> <xref ref-type="bibr" rid="bib1.bibx128" id="paren.257"><named-content content-type="pre">Scott's rule bandwidth;</named-content></xref> of the marginal prior and marginal approximate posterior distributions of the parameters from the respective algorithms in the transformed (unbounded) space. The MCMC posterior (orange) from RAM is considered the gold-standard.</p></caption>
          <graphic xlink:href="https://gmd.copernicus.org/articles/19/8565/2026/gmd-19-8565-2026-f06.png"/>

        </fig>

      <p id="d2e11527">It should also be noted that the estimated posterior uncertainty differs between ES-MDA and AdaPBS. The posterior standard deviation in ES-MDA is slightly higher compared to that of AdaPBS. At least in the present example, these differences are nonetheless relatively minor, especially in the model state space visualized here in terms of snow depth. While quantitative probabilistic evaluation using diagnostics such as KLD is helpful, qualitative visual comparisons to an MCMC reference posterior ensemble trajectory in state space can be equally instructive. Inspecting the snow depth trajectories in Figs. <xref ref-type="fig" rid="F2"/>–<xref ref-type="fig" rid="F5"/> shows that the patterns identified in parameter space carry over to snow depth in state space. Using the RAM-based MCMC posterior trajectory as a reference: PBS is biased and highly degenerate, ES is biased and highly under-confident, ES-MDA is unbiased yet slightly under-confident, whereas AdaPBS overlaps almost perfectly with the reference ensemble. To date, most snow DA studies focus primarily on the predictive performance of point estimates such as the posterior mean, median, or mode <xref ref-type="bibr" rid="bib1.bibx86 bib1.bibx1 bib1.bibx6 bib1.bibx99" id="paren.258"><named-content content-type="pre">e.g.</named-content></xref>. Nonetheless, as snow DA continues to transition towards fully Bayesian methods, an emphasis on rigorous uncertainty quantification will become increasingly important for probabilistic prediction <xref ref-type="bibr" rid="bib1.bibx105" id="paren.259"/> and decision-making <xref ref-type="bibr" rid="bib1.bibx109" id="paren.260"/>. As we have shown, benchmarking via MCMC reference posteriors <xref ref-type="bibr" rid="bib1.bibx77" id="paren.261"/> allows for direct qualitative and quantitative evaluation of more approximate posterior estimates obtained from tractable DA algorithms such as AdaPBS.</p>
      <p id="d2e11549">Despite the many advantages of the AdaPBS scheme that we have highlighted, the ES-MDA still retains a key advantage: There are more localization methods developed for ensemble Kalman-based algorithms <xref ref-type="bibr" rid="bib1.bibx44" id="paren.262"/> than for those based on particle methods <xref ref-type="bibr" rid="bib1.bibx45" id="paren.263"/>. This facilitates the development of spatiotemporal assimilation initiatives <xref ref-type="bibr" rid="bib1.bibx8 bib1.bibx89 bib1.bibx9" id="paren.264"/>, as opposed to the purely temporal example presented here, which allows non-local observations in remote cells of a distributed model domain to be considered to update the local cell in question. In this way, information can be propagated in space, correcting areas of the domain even when they have no local observations, or integrating point-scale observations into distributed simulations. A future line of research with great potential will be the development of these spatio-temporal methods for AdaPBS and other particle-based methods, which remains at the frontier of current DA research. Moreover, in conjunction with the localization problem, it remains to be seen if adaptive particle methods can be scaled up to as high-dimensional problems as ensemble Kalman methods <xref ref-type="bibr" rid="bib1.bibx18 bib1.bibx103" id="paren.265"/>. Nonetheless, the results from our inter-comparison of cryospheric DA schemes with a temperature index model show that these adaptive particle methods are likely to perform favorably in local cryospheric DA problems with low dimensional parameter spaces. In the ensuing experiments with FSM2, we push the model complexity, number of observations, and the dimensionality of the parameter space considerably to further test the AdaPBS.</p>
</sec>
<sec id="Ch1.S4.SS2">
  <label>4.2</label><title>Applying AdaPBS in FSM2 at SnowMIP sites</title>
      <p id="d2e11572">The large number of assimilated observations in these experiments with thousands of observations per annual DA window represents a challenge for any batch smoothing assimilation algorithm. On average, AdaPBS required 8 iterations per season to address this problem. In principle, such a high number of iterations in AdaPBS would be expected to result in a considerable increase in computational cost compared to the typical number of iterations used by ES-MDA in most cryospheric applications (which in our experience is around four) since a higher number of iterations involve a large number of sequential model realizations. However in this case, the computational cost of the linear algebra in ES-MDA scaled with the number of observations, resulting in a much higher execution time compared with AdaPBS. The site with the largest run times due to the longest simulation, namely WSJ with 6 water years, took 16 min on a single core with AdaPBS compared with 5 h using two cores for ES-MDA. Nonetheless, this conclusion should be interpreted with some caution, as both linear algebra operations and ensemble generation can be parallelized. There may be solutions to circumvent these issues depending on the problem at a hand related to the the parallelization scheme of the computing infrastructure, the hyperparamters (i.e. number of iterations for ES-MDA and diversity threshold for AdaPBS), and further optimizing linear algebra operations <xref ref-type="bibr" rid="bib1.bibx63" id="paren.266"/>. Moreover, previous experience suggests that a larger number of iterations of ES-MDA, potentially exceeding <inline-formula><mml:math id="M423" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">a</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">10</mml:mn></mml:mrow></mml:math></inline-formula>, may also be necessary to ensure convergence based on earlier work applying the ES-MDA to higher-dimensional problems <xref ref-type="bibr" rid="bib1.bibx103" id="paren.267"/>. In a similar vein, the validation metrics of AdaPBS may improve if a larger <inline-formula><mml:math id="M424" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">eff</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> threshold is selected, at the cost of increasing the computational expense. More generally, the large number of observations can be assimilated in a more tractable manner via Bayesian filtering <xref ref-type="bibr" rid="bib1.bibx113" id="paren.268"/> and tempering <xref ref-type="bibr" rid="bib1.bibx21" id="paren.269"/> techniques. Combining adaptive particle methods with these techniques is likely to be a promising future research direction in cryospheric DA.</p>

      <fig id="F7" specific-use="star"><label>Figure 7</label><caption><p id="d2e11616">Comparison of the performance of the ESA-MDA (top) and the AdaPBS (bottom) at the Weissfluhjoch (WFJ) ESM-SnowMIP site in the eastern Swiss Alps. The red and green lines with corresponding shading show the prior (“open-loop” in red) and posterior (“updated” in green) mean <inline-formula><mml:math id="M425" display="inline"><mml:mrow><mml:mo>±</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> standard deviation along with the hourly assimilated snow depth observations (purple dots). </p></caption>
          <graphic xlink:href="https://gmd.copernicus.org/articles/19/8565/2026/gmd-19-8565-2026-f07.png"/>

        </fig>

      <p id="d2e11635">The results of AdaPBS in this more challenging setting with higher-dimensional parameter and observation spaces are similar to those obtained by ES-MDA, as shown in Table <xref ref-type="table" rid="T3"/>. It is difficult to conclude confidently which of the two schemes works best, so we conclude tentatively that their performance is similar even in scenarios as complex as the one proposed. The biggest differences are found at SNB, both in terms of RMSE (ES-MDA: 0.21, AdaPBS: 0.44) and CRPS (ES-MDA: 0.15, AdaPBS: 0.29), despite both methods presenting a very similar bias for this site. While the AdaPBS has a tendency to either match or perform slightly worse than ES-MDA in terms of RMSE and CRPS, it generally exhibits a lower bias than ES-MDA particularly for SWA (ES-MDA: <inline-formula><mml:math id="M426" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.16, AdaPBS: <inline-formula><mml:math id="M427" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.06). However, the fact that there are notable improvements with AdaPBS compared to ES-MDA in only one of the sites (SWA), worse performance at one site (SNB), and nearly identical performance at the remaining sites, is clearly not sufficient grounds to claim that there is considerable differences in the performance of these schemes. The results in Table <xref ref-type="table" rid="T3"/> also show that AdaPBS consistently outperforms PBS in this experiment for all metrics across all sites, with the exception of SOD where the low bias differs by only 0.01. Overall, when considering metrics averaged over all sites, use of AdaPBS leads to a 25 %, 41 %, and 60 % reduction in RMSE, CRPS, and bias, respectively, compared to PBS. A likely reason for the large gains from adaptation is that the non-iterative PBS struggles to capture the dense observations, as demonstrated by the average of eight iterations needed in AdaPBS. As a result, PBS often degenerates completely to a suboptimal solution. Although not explicitly evaluated here, a single ES update is expected to also lead to a suboptimal solution due to the combined effects of nonlinearity and the highly informative nature of the observations. This expectation is supported by Experiment 1 as well as previous experiments comparing the ES-MDA to the ES in snow data assimilation <xref ref-type="bibr" rid="bib1.bibx1 bib1.bibx7" id="paren.270"/> and other applications <xref ref-type="bibr" rid="bib1.bibx102 bib1.bibx103 bib1.bibx71" id="paren.271"/>. It is thus encouraging to see the AdaPBS, a purely particle-based iterative batch smoother, be able to match the performance of the ES-MDA given that AdaPBS provides added flexibility in terms of adapting to the problem at hand through early stopping and is readily extended to non-Gaussian likelihoods <xref ref-type="bibr" rid="bib1.bibx115" id="paren.272"><named-content content-type="pre">e.g.</named-content></xref>. Figure <xref ref-type="fig" rid="F7"/> demonstrates the generally very good performance of both the ES-MDA and AdaPBS algorithms at WFJ in terms of matching the temporally dense snow depth observations. The same holds for most of the other ESM-SnowMIP sites visualized in the Supplement (see Fig. S1). However, suboptimal posterior predictions are obtained in several water years, especially with AdaPBS but also ES-MDA, at the SNB which turns out to be the most challenging among all the sites (see Fig. S3). Nonetheless, it should be noted that no explicit ERA5 forcing downscaling has been performed other than implicitly by the assimilation itself by inferring the forcing perturbation parameters. The resolution of ERA5 may not be sufficient to capture all the particularities of local observations even after the implicit downscaling at least with the current probabilistic model configuration.</p>

<table-wrap id="T3" specific-use="star"><label>Table 3</label><caption><p id="d2e11674">Snow depth evaluation metrics in the form of the root mean square error (RMSE), mean continuous ranked probability score (CRPS), and mean error (bias), for the Prior, PBS, IES (ES-MDA), and AdaPBS across 6 ESM-SnowMIP sites.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="13">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:colspec colnum="5" colname="col5" align="right" colsep="1"/>
     <oasis:colspec colnum="6" colname="col6" align="right"/>
     <oasis:colspec colnum="7" colname="col7" align="right"/>
     <oasis:colspec colnum="8" colname="col8" align="right"/>
     <oasis:colspec colnum="9" colname="col9" align="right" colsep="1"/>
     <oasis:colspec colnum="10" colname="col10" align="right"/>
     <oasis:colspec colnum="11" colname="col11" align="right"/>
     <oasis:colspec colnum="12" colname="col12" align="right"/>
     <oasis:colspec colnum="13" colname="col13" align="right"/>
     <oasis:thead>
       <oasis:row>
         <oasis:entry colname="col1">ESM-SnowMIP</oasis:entry>
         <oasis:entry rowsep="1" namest="col2" nameend="col5" align="center" colsep="1">RMSE </oasis:entry>
         <oasis:entry rowsep="1" namest="col6" nameend="col9" align="center" colsep="1">Mean CRPS </oasis:entry>
         <oasis:entry rowsep="1" namest="col10" nameend="col13" align="center">Bias </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Site (Location)</oasis:entry>
         <oasis:entry colname="col2">Prior</oasis:entry>
         <oasis:entry colname="col3">PBS</oasis:entry>
         <oasis:entry colname="col4">IES</oasis:entry>
         <oasis:entry colname="col5">AdaPBS</oasis:entry>
         <oasis:entry colname="col6">Prior</oasis:entry>
         <oasis:entry colname="col7">PBS</oasis:entry>
         <oasis:entry colname="col8">IES</oasis:entry>
         <oasis:entry colname="col9">AdaPBS</oasis:entry>
         <oasis:entry colname="col10">Prior</oasis:entry>
         <oasis:entry colname="col11">PBS</oasis:entry>
         <oasis:entry colname="col12">IES</oasis:entry>
         <oasis:entry colname="col13">AdaPBS</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">RME (Idaho, USA)</oasis:entry>
         <oasis:entry colname="col2">0.31</oasis:entry>
         <oasis:entry colname="col3">0.22</oasis:entry>
         <oasis:entry colname="col4">0.10</oasis:entry>
         <oasis:entry colname="col5">0.11</oasis:entry>
         <oasis:entry colname="col6">0.29</oasis:entry>
         <oasis:entry colname="col7">0.18</oasis:entry>
         <oasis:entry colname="col8">0.08</oasis:entry>
         <oasis:entry colname="col9">0.10</oasis:entry>
         <oasis:entry colname="col10">0.01</oasis:entry>
         <oasis:entry colname="col11"><inline-formula><mml:math id="M428" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.06</oasis:entry>
         <oasis:entry colname="col12">0.01</oasis:entry>
         <oasis:entry colname="col13">0.01</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">SAP (Hokkaido, Japan)</oasis:entry>
         <oasis:entry colname="col2">0.22</oasis:entry>
         <oasis:entry colname="col3">0.11</oasis:entry>
         <oasis:entry colname="col4">0.07</oasis:entry>
         <oasis:entry colname="col5">0.07</oasis:entry>
         <oasis:entry colname="col6">0.14</oasis:entry>
         <oasis:entry colname="col7">0.14</oasis:entry>
         <oasis:entry colname="col8">0.04</oasis:entry>
         <oasis:entry colname="col9">0.05</oasis:entry>
         <oasis:entry colname="col10"><inline-formula><mml:math id="M429" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.16</oasis:entry>
         <oasis:entry colname="col11"><inline-formula><mml:math id="M430" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.03</oasis:entry>
         <oasis:entry colname="col12"><inline-formula><mml:math id="M431" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.01</oasis:entry>
         <oasis:entry colname="col13">0.00</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">SNB (Colorado, USA)</oasis:entry>
         <oasis:entry colname="col2">0.65</oasis:entry>
         <oasis:entry colname="col3">0.48</oasis:entry>
         <oasis:entry colname="col4">0.21</oasis:entry>
         <oasis:entry colname="col5">0.44</oasis:entry>
         <oasis:entry colname="col6">0.36</oasis:entry>
         <oasis:entry colname="col7">0.34</oasis:entry>
         <oasis:entry colname="col8">0.15</oasis:entry>
         <oasis:entry colname="col9">0.29</oasis:entry>
         <oasis:entry colname="col10"><inline-formula><mml:math id="M432" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.10</oasis:entry>
         <oasis:entry colname="col11"><inline-formula><mml:math id="M433" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.12</oasis:entry>
         <oasis:entry colname="col12"><inline-formula><mml:math id="M434" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.07</oasis:entry>
         <oasis:entry colname="col13"><inline-formula><mml:math id="M435" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.09</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">SOD (Lapland, Finland)</oasis:entry>
         <oasis:entry colname="col2">0.11</oasis:entry>
         <oasis:entry colname="col3">0.07</oasis:entry>
         <oasis:entry colname="col4">0.06</oasis:entry>
         <oasis:entry colname="col5">0.06</oasis:entry>
         <oasis:entry colname="col6">0.08</oasis:entry>
         <oasis:entry colname="col7">0.08</oasis:entry>
         <oasis:entry colname="col8">0.04</oasis:entry>
         <oasis:entry colname="col9">0.05</oasis:entry>
         <oasis:entry colname="col10"><inline-formula><mml:math id="M436" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.07</oasis:entry>
         <oasis:entry colname="col11">0.00</oasis:entry>
         <oasis:entry colname="col12">0.00</oasis:entry>
         <oasis:entry colname="col13">0.01</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">SWA (Colorado, USA)</oasis:entry>
         <oasis:entry colname="col2">0.48</oasis:entry>
         <oasis:entry colname="col3">0.28</oasis:entry>
         <oasis:entry colname="col4">0.28</oasis:entry>
         <oasis:entry colname="col5">0.27</oasis:entry>
         <oasis:entry colname="col6">0.35</oasis:entry>
         <oasis:entry colname="col7">0.33</oasis:entry>
         <oasis:entry colname="col8">0.20</oasis:entry>
         <oasis:entry colname="col9">0.21</oasis:entry>
         <oasis:entry colname="col10"><inline-formula><mml:math id="M437" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.27</oasis:entry>
         <oasis:entry colname="col11"><inline-formula><mml:math id="M438" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.09</oasis:entry>
         <oasis:entry colname="col12"><inline-formula><mml:math id="M439" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.16</oasis:entry>
         <oasis:entry colname="col13"><inline-formula><mml:math id="M440" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.06</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">WFJ (Swiss Alps)</oasis:entry>
         <oasis:entry colname="col2">0.33</oasis:entry>
         <oasis:entry colname="col3">0.25</oasis:entry>
         <oasis:entry colname="col4">0.10</oasis:entry>
         <oasis:entry colname="col5">0.11</oasis:entry>
         <oasis:entry colname="col6">0.28</oasis:entry>
         <oasis:entry colname="col7">0.27</oasis:entry>
         <oasis:entry colname="col8">0.08</oasis:entry>
         <oasis:entry colname="col9">0.09</oasis:entry>
         <oasis:entry colname="col10">0.15</oasis:entry>
         <oasis:entry colname="col11"><inline-formula><mml:math id="M441" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.01</oasis:entry>
         <oasis:entry colname="col12"><inline-formula><mml:math id="M442" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.01</oasis:entry>
         <oasis:entry colname="col13"><inline-formula><mml:math id="M443" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.01</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Mean over sites</oasis:entry>
         <oasis:entry colname="col2">0.35</oasis:entry>
         <oasis:entry colname="col3">0.24</oasis:entry>
         <oasis:entry colname="col4">0.14</oasis:entry>
         <oasis:entry colname="col5">0.18</oasis:entry>
         <oasis:entry colname="col6">0.25</oasis:entry>
         <oasis:entry colname="col7">0.22</oasis:entry>
         <oasis:entry colname="col8">0.10</oasis:entry>
         <oasis:entry colname="col9">0.13</oasis:entry>
         <oasis:entry colname="col10"><inline-formula><mml:math id="M444" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.07</oasis:entry>
         <oasis:entry colname="col11"><inline-formula><mml:math id="M445" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.05</oasis:entry>
         <oasis:entry colname="col12"><inline-formula><mml:math id="M446" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.04</oasis:entry>
         <oasis:entry colname="col13"><inline-formula><mml:math id="M447" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.02</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

</sec>
<sec id="Ch1.S4.SS3">
  <label>4.3</label><title>Recommendations for applying and extending the AdaPBS</title>
      <p id="d2e12212">Broadening the discussion, we provide recommendations informed by the above snow DA experiments and recent applications of AdaPBS to permafrost <xref ref-type="bibr" rid="bib1.bibx133" id="paren.273"/> and glaciers <xref ref-type="bibr" rid="bib1.bibx134" id="paren.274"/>. In addition to ensemble size <inline-formula><mml:math id="M448" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, which is foundational for all ensemble-based DA, two key algorithm-specific hyperparameters control AdaPBS: the diversity threshold <inline-formula><mml:math id="M449" display="inline"><mml:mrow><mml:mn mathvariant="normal">0</mml:mn><mml:mo>&lt;</mml:mo><mml:mi mathvariant="italic">τ</mml:mi><mml:mo>&lt;</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> and the maximum number of iterations <inline-formula><mml:math id="M450" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub><mml:mo>≥</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> (Sect. <xref ref-type="sec" rid="Ch1.S2.SS6"/>). On the one hand, <inline-formula><mml:math id="M451" display="inline"><mml:mi mathvariant="italic">τ</mml:mi></mml:math></inline-formula> serves as a lower bound for the desired quality of the  posterior approximation: an ensemble run that converges successfully will achieve an <inline-formula><mml:math id="M452" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">eff</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> of <italic>at least</italic> <inline-formula><mml:math id="M453" display="inline"><mml:mrow><mml:mi mathvariant="italic">τ</mml:mi><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. On the other hand, <inline-formula><mml:math id="M454" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> serves as an upper bound on the computational cost: AdaPBS will run for <italic>at most</italic> <inline-formula><mml:math id="M455" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> iterations with a total of <inline-formula><mml:math id="M456" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> model runs, but it may converge and terminate much earlier. This upper bound is only reached if AdaPBS converges in the final possible iteration or, in the worst case, if it fails to converge within the prescribed <inline-formula><mml:math id="M457" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> iterations. There is thus an inherent tradeoff in these hyperparameters: a higher threshold <inline-formula><mml:math id="M458" display="inline"><mml:mrow><mml:mi mathvariant="italic">τ</mml:mi><mml:mo>≃</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> helps improve the particle approximation of the posterior while often requiring more iterations for convergence, demanding a larger <inline-formula><mml:math id="M459" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> with increased computational cost. At the opposite extreme, one could simply set <inline-formula><mml:math id="M460" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> to minimize computational cost, but then AdaPBS is identical to the PBS and has no way to adapt and evolve beyond degeneracy when it occurs. A useful default is <inline-formula><mml:math id="M461" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">5</mml:mn></mml:mrow></mml:math></inline-formula>, mirroring the typical cost of the ES-MDA that uses <inline-formula><mml:math id="M462" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">a</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> ensemble iterations with <inline-formula><mml:math id="M463" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">a</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">4</mml:mn></mml:mrow></mml:math></inline-formula> assimilation cycles <xref ref-type="bibr" rid="bib1.bibx9" id="paren.275"/>, but where AdaPBS imposes fewer assumptions and can stop early upon convergence.</p>
      <p id="d2e12440">The optimal balance between these two hyperparameters depends on the problem at hand and the available computational budget. Setting the threshold in the range <inline-formula><mml:math id="M464" display="inline"><mml:mrow><mml:mn mathvariant="normal">0.1</mml:mn><mml:mo>&lt;</mml:mo><mml:mi mathvariant="italic">τ</mml:mi><mml:mo>&lt;</mml:mo><mml:mn mathvariant="normal">0.5</mml:mn></mml:mrow></mml:math></inline-formula> is generally enough to obtain a good posterior approximation, while the number of iterations depends on the problem. This number may vary considerably with the complexity of the problem and even within a given distributed multi-year simulation. Ensuring convergence across the board may require setting <inline-formula><mml:math id="M465" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> to 10 or more. In practice, it is often entirely acceptable to end up with a low percentage (say <inline-formula><mml:math id="M466" display="inline"><mml:mrow><mml:mo>≤</mml:mo><mml:mn mathvariant="normal">10</mml:mn></mml:mrow></mml:math></inline-formula>  %) of cases in which the algorithm failed to converge. We emphasize that failure to converge does not imply a complete failure of the algorithm, but rather that the particle approximation is not as diverse as desired for uncertainty quantification via (e.g.) the posterior variance. The ensemble may still provide a reasonable point estimate for the posterior mean and avoid complete degeneracy (<inline-formula><mml:math id="M467" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">eff</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula>). Since the AdaPBS is designed as a direct extension of the PBS, instances of failed convergence will correspond to cases when the PBS would be at least as degenerate.</p>
      <p id="d2e12495">Experiments 1 and 2 demonstrate the role of problem complexity. For the simpler temperature index model with two parameters, AdaPBS exceeded the desired threshold of <inline-formula><mml:math id="M468" display="inline"><mml:mrow><mml:mi mathvariant="italic">τ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.3</mml:mn></mml:mrow></mml:math></inline-formula> (i.e., <inline-formula><mml:math id="M469" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">eff</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">30</mml:mn></mml:mrow></mml:math></inline-formula> with <inline-formula><mml:math id="M470" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">100</mml:mn></mml:mrow></mml:math></inline-formula>) with <inline-formula><mml:math id="M471" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">eff</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">84</mml:mn></mml:mrow></mml:math></inline-formula> in 5 iterations, providing a strong approximation that nearly converged already at 4 iterations (Fig. <xref ref-type="fig" rid="F5"/>). In Experiment 2, the intermediate complexity snow model FSM2, a higher dimensional parameter space, and numerous observations conspired to make inference challenging <xref ref-type="bibr" rid="bib1.bibx117 bib1.bibx94" id="paren.276"/>. Following the “needle in a haystack” analogy for inferring the posterior: increased model complexity may result in an oddly shaped needle, more informative data makes the needle smaller, and higher dimensionality makes the haystack bigger. Thus, AdaPBS required, on average, nearly twice as many iterations (<inline-formula><mml:math id="M472" display="inline"><mml:mn mathvariant="normal">8</mml:mn></mml:math></inline-formula>) to converge with the same <inline-formula><mml:math id="M473" display="inline"><mml:mrow><mml:mi mathvariant="italic">τ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.3</mml:mn></mml:mrow></mml:math></inline-formula> as Experiment 1. For both experiments, the number of iterations could have been reduced by lowering the bar on diversity to say <inline-formula><mml:math id="M474" display="inline"><mml:mrow><mml:mi mathvariant="italic">τ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.1</mml:mn></mml:mrow></mml:math></inline-formula>, which can suffice in operational settings focused on prediction using point estimates rather than uncertainty quantification <xref ref-type="bibr" rid="bib1.bibx99" id="paren.277"><named-content content-type="pre">e.g.</named-content></xref>.</p>
      <p id="d2e12597">Valuable lessons can also be learned from the two other contemporaneous studies using AdaPBS for cryospheric DA. <xref ref-type="bibr" rid="bib1.bibx133" id="text.278"/> used AdaPBS to assimilate Sentinel-2 and in-situ FSCA data into the detailed multiphysics cryospheric model CryoGrid <xref ref-type="bibr" rid="bib1.bibx131" id="paren.279"/>, inferring two parameters to constrain ground surface temperatures around Bayelva near Ny-Ålesund, Svalbard. To configure AdaPBS, the maximum number of iterations was set to <inline-formula><mml:math id="M475" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">8</mml:mn></mml:mrow></mml:math></inline-formula>, and the diversity threshold to <inline-formula><mml:math id="M476" display="inline"><mml:mrow><mml:mi mathvariant="italic">τ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.3</mml:mn></mml:mrow></mml:math></inline-formula>. In this two-dimensional parameter space, across all grid cells, convergence was achieved within an average of <inline-formula><mml:math id="M477" display="inline"><mml:mrow><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:math></inline-formula> iterations, and only rarely did AdaPBS fail to converge. To contend with the relatively high computational cost of this CryoGrid configuration, parallel simulations were run with the ensemble size <inline-formula><mml:math id="M478" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">20</mml:mn></mml:mrow></mml:math></inline-formula> equal to the number of locally available cores.  As such, AdaPBS could optimally leverage parallelization to avoid waiting for new particles to complete while benefiting from adaptation across iterations to avoid degeneracy. For example, the total computational cost of <inline-formula><mml:math id="M479" display="inline"><mml:mn mathvariant="normal">5</mml:mn></mml:math></inline-formula> iterations with this configuration is identical to running a non-adaptive PBS with a larger ensemble of <inline-formula><mml:math id="M480" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">100</mml:mn></mml:mrow></mml:math></inline-formula>, as both would require <inline-formula><mml:math id="M481" display="inline"><mml:mn mathvariant="normal">5</mml:mn></mml:math></inline-formula> sequential batches of <inline-formula><mml:math id="M482" display="inline"><mml:mn mathvariant="normal">20</mml:mn></mml:math></inline-formula> parallel simulations to finish. This approach can be extended to spatially distributed simulations across large domains, where AdaPBS is, like all ensemble-based methods, highly parallelizable in the ensemble dimension. Furthermore, later adaptive iterations will be cheaper to run as some grid cells will have converged and can be skipped. It may also be possible to fold in new cells to replace converged cells, helping to maintain an even computational load. For very large scale applications, it is advantageous to jointly chunk simulations in both the ensemble and spatial dimensions, vectorizing within chunks and parallelizing across. Generally, multiplying the <inline-formula><mml:math id="M483" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> hyperparameter with typical model run times provides a useful worst-case estimate for wall-clock time for allocating computing resources.</p>
      <p id="d2e12709">Further lessons can be learned from <xref ref-type="bibr" rid="bib1.bibx134" id="text.280"/>, who used AdaPBS to jointly calibrate 4 parameters controlling surface mass balance and frontal ablation in the coupled PyGEM-OGGM glacier models <xref ref-type="bibr" rid="bib1.bibx111 bib1.bibx88" id="paren.281"/>. Therein, the evolution of <inline-formula><mml:math id="M484" display="inline"><mml:mn mathvariant="normal">71</mml:mn></mml:math></inline-formula> marine-terminating glaciers on Svalbard was inferred for the period 2000–2019 by jointly assimilating glacier-wide mass balance and terminus-positions. Using a configuration with <inline-formula><mml:math id="M485" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">200</mml:mn></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M486" display="inline"><mml:mrow><mml:mi mathvariant="italic">τ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.1</mml:mn></mml:mrow></mml:math></inline-formula>, and <inline-formula><mml:math id="M487" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">10</mml:mn></mml:mrow></mml:math></inline-formula>, AdaPBS converged successfully within 10 iterations for 87 % of the glaciers considered. In terms of performance, the posterior mean from AdaPBS considerably reduced root mean squared errors compared to the prior mean for both the converged and non-converged glaciers. By extending the maximum number of iterations to <inline-formula><mml:math id="M488" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">100</mml:mn></mml:mrow></mml:math></inline-formula>, AdaPBS also managed to converge for most of the remaining glaciers. This showcases both the need for selecting an appropriate AdaPBS configuration and that infrequent non-convergence need not adversely affect overall performance. Moreover, as suggested in <xref ref-type="bibr" rid="bib1.bibx134" id="text.282"/>, methods such as Bayesian optimization <xref ref-type="bibr" rid="bib1.bibx63" id="paren.283"/> can be used to automatically tune these hyperparameters.</p>
      <p id="d2e12789">The AdaPBS algorithm that we have presented relies on adaptive importance sampling <xref ref-type="bibr" rid="bib1.bibx15" id="paren.284"/>, specifically AMIS <xref ref-type="bibr" rid="bib1.bibx26" id="paren.285"/>, to extend the PBS <xref ref-type="bibr" rid="bib1.bibx85" id="paren.286"/> by allowing it to evolve beyond collapse when degeneracy strikes. At the same time, adaptation is but one direction for improvement to make particle methods less prone to degeneracy. Likelihood tempering <xref ref-type="bibr" rid="bib1.bibx97 bib1.bibx95" id="paren.287"/>, which also powers the ES-MDA <xref ref-type="bibr" rid="bib1.bibx39 bib1.bibx119" id="paren.288"/>, inflates the observation error covariance and iteratively ensures a smoother transition from prior to posterior. Data tempering <xref ref-type="bibr" rid="bib1.bibx20 bib1.bibx95" id="paren.289"/> randomly divides the observation vector into mini-batches and assimilates these sequentially, allowing particle methods to assimilate more data without degeneracy <xref ref-type="bibr" rid="bib1.bibx122 bib1.bibx123" id="paren.290"><named-content content-type="pre">e.g.</named-content></xref>. Particle rejuvenation methods such as jitter <xref ref-type="bibr" rid="bib1.bibx45" id="paren.291"/> or the resample-move approach using  MCMC <xref ref-type="bibr" rid="bib1.bibx51" id="paren.292"/> also help avoid degeneracy and are frequently used together with tempering for advanced particle methods <xref ref-type="bibr" rid="bib1.bibx21" id="paren.293"/>. Hybrid methods are also widely used <xref ref-type="bibr" rid="bib1.bibx126" id="paren.294"/>, for example, employing the ES-MDA to design the proposal distribution <xref ref-type="bibr" rid="bib1.bibx102" id="paren.295"/>. Crucially, tempering, rejuvenation, and hybridization can all be used together with adaptation to design more robust particle methods for cryospheric DA. Furthermore, as shown theoretically in <xref ref-type="bibr" rid="bib1.bibx36" id="text.296"/>, particle methods relying on AMIS, such as AdaPBS, can be readily applied to the auxiliary particle filter, which may further improve operational snow hydrological forecasting systems <xref ref-type="bibr" rid="bib1.bibx99" id="paren.297"/>. Using adaptive particle methods for spatio-temporal DA <xref ref-type="bibr" rid="bib1.bibx89 bib1.bibx9" id="paren.298"><named-content content-type="pre">e.g.</named-content></xref>, will require further progress on particle localization techniques <xref ref-type="bibr" rid="bib1.bibx45" id="paren.299"/> for cryospheric DA by building on the efforts of <xref ref-type="bibr" rid="bib1.bibx24" id="text.300"/>. As an extension of the PBS, AdaPBS has the potential to be directly applied to large-scale snow reanalysis efforts <xref ref-type="bibr" rid="bib1.bibx86 bib1.bibx6 bib1.bibx81" id="paren.301"/>, where <inline-formula><mml:math id="M489" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">ℓ</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> provides an upper bound on computational cost.</p>
</sec>
</sec>
<sec id="Ch1.S5" sec-type="conclusions">
  <label>5</label><title>Conclusions</title>
      <p id="d2e12874">We introduced the AdaPBS algorithm, an iterative and adaptive particle-based data assimilation scheme that combines a Particle Batch Smoother <xref ref-type="bibr" rid="bib1.bibx85" id="paren.302"/> with Adaptive Multiple Importance Sampling <xref ref-type="bibr" rid="bib1.bibx26" id="paren.303"/>. This formulation improves the resilience against ensemble collapse typically associated with particle-based methods while avoiding the Gaussian linear assumptions associated with ensemble Kalman-based methods. It also allows for the implementation of early stopping strategies, adapting the computational effort automatically to the problem complexity at the local grid-cell and water-year level in distributed multi-year simulations with potentially large cost reductions.</p>
      <p id="d2e12883">In a simple temperature index snow model assimilating snow depth observations, AdaPBS agreed closely with a costly RAM-based MCMC gold-standard benchmark, consistently outperforming or at least matching commonly used particle and ensemble Kalman-based batch smoothing methods. In more complex experiments assimilating hourly snow depth observations at SnowMIP sites into ensembles generated with the FSM2 model, AdaPBS managed to sample from the posterior simulations in a challenging high dimensional space of 19 uncertain parameters with a similar performance to ES-MDA. All experiments were developed with the open source MuSA toolbox <xref ref-type="bibr" rid="bib1.bibx7" id="paren.304"/>, in which both the proposed AdaPBS scheme and the RAM MCMC algorithm used for benchmarking are now available. This should facilitate the use and adaptation of AdaPBS to the plethora of data assimilation problems related to the terrestrial cryosphere <xref ref-type="bibr" rid="bib1.bibx53 bib1.bibx131 bib1.bibx111" id="paren.305"/> and adjacent fields <xref ref-type="bibr" rid="bib1.bibx102 bib1.bibx71" id="paren.306"/>. Contemporaneous and promising implementations of AdaPBS already exist in the CryoGrid community model <xref ref-type="bibr" rid="bib1.bibx131" id="paren.307"/> as detailed in <xref ref-type="bibr" rid="bib1.bibx133" id="text.308"/> and the coupled PyGEM-OGGM global glacier models <xref ref-type="bibr" rid="bib1.bibx111 bib1.bibx88" id="paren.309"/>  as detailed in <xref ref-type="bibr" rid="bib1.bibx134" id="text.310"/>. We encourage cryospheric researchers to help test, develop, refine, and remix data assimilation methods such as AdaPBS so that we as a community can continue to evolve in symbiosis with rapid advances in Bayesian methods <xref ref-type="bibr" rid="bib1.bibx21 bib1.bibx44 bib1.bibx95" id="paren.311"><named-content content-type="pre">e.g.</named-content></xref>.</p>
</sec>

      
      </body>
    <back><notes notes-type="codedataavailability"><title>Code and data availability</title>

      <p id="d2e12917">The latest release of MuSA (v2.3) containing the AdaPBS in conjunction with this paper is open source and can be found at <ext-link xlink:href="https://doi.org/10.5281/zenodo.17292981" ext-link-type="DOI">10.5281/zenodo.17292981</ext-link> <xref ref-type="bibr" rid="bib1.bibx3" id="paren.312"/>. The code and data needed to reproduce all the results and figures in this study are openly available through <ext-link xlink:href="https://doi.org/10.5281/zenodo.21244337" ext-link-type="DOI">10.5281/zenodo.21244337</ext-link> <xref ref-type="bibr" rid="bib1.bibx4" id="paren.313"/>. The original Izas input data needed to run the first set of experiments in this study can be found at <ext-link xlink:href="https://doi.org/10.5281/zenodo.7248635" ext-link-type="DOI">10.5281/zenodo.7248635</ext-link> <xref ref-type="bibr" rid="bib1.bibx2" id="paren.314"/>. The original SnowMIP input data for the second set of experiments can be found at <ext-link xlink:href="https://doi.org/10.1594/PANGAEA.897575" ext-link-type="DOI">10.1594/PANGAEA.897575</ext-link> <xref ref-type="bibr" rid="bib1.bibx90" id="paren.315"/>. Future versions of MuSA will continue to be submitted to <uri>https://github.com/ealonsogzl/MuSA</uri> (last access: 9 September 2026).</p>
  </notes><app-group>
        <supplementary-material position="anchor"><p id="d2e12948">The supplement related to this article is available online at <inline-supplementary-material xlink:href="https://doi.org/10.5194/gmd-19-8565-2026-supplement" xlink:title="zip">https://doi.org/10.5194/gmd-19-8565-2026-supplement</inline-supplementary-material>.</p></supplementary-material>
        </app-group><notes notes-type="authorcontribution"><title>Author contributions</title>

      <p id="d2e12957">Conceptualization: KA and EAG. Data curation: EAG. Formal analysis: EAG and KA. Funding acquisition: EAG, KA. Investigation: EAG and KA. Methodology: KA and EAG with contributions from NP, CW, SW, RY. Resources: EAG. Software: EAG and KA. Validation: EAG and KA. Visualization: EAG. Writing (original draft preparation): KA and EAG. Writing (review and editing): KA, EAG, NP, CW, SW, RY.</p>
  </notes><notes notes-type="competinginterests"><title>Competing interests</title>

      <p id="d2e12963">The contact author has declared that none of the authors has any competing interests.</p>
  </notes><notes notes-type="disclaimer"><title>Disclaimer</title>

      <p id="d2e12969">Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. The authors bear the ultimate responsibility for providing appropriate place names. Views expressed in the text are those of the authors and do not necessarily reflect the views of the publisher.</p>
  </notes><ack><title>Acknowledgements</title><p id="d2e12975">We are grateful to the European Centre for Medium-Range Weather Forecasts (ECMWF) for openly providing global atmospheric reanalysis data. In particular, all experiments carried out in this study were forced using ERA5 reanalysis data <xref ref-type="bibr" rid="bib1.bibx64" id="paren.316"/> obtained from the Copernicus Climate Change Service (C3S) Climate Data Store. These data were generated using modified Copernicus Climate Change Service information. Neither the European Commission nor ECMWF is responsible for any use that may be made of the Copernicus information or data it contains. All simulations were carried out using the MuSA toolbox <xref ref-type="bibr" rid="bib1.bibx7" id="paren.317"/> which is built on open source Python libraries and the Fortran-based snow model FSM2 <xref ref-type="bibr" rid="bib1.bibx41" id="paren.318"/>. We are also grateful to ESM-SnowMIP effort for making snow depth data readily available from several snow reference sites through <xref ref-type="bibr" rid="bib1.bibx91" id="text.319"/>. This research was made possible through the access granted by the Galician Supercomputing Center (CESGA) to its supercomputing infrastructure. The supercomputer FinisTerrae III and its permanent data storage system have been funded by the NextGeneration EU 2021 Recovery, Transformation and Resilience Plan, ICT2021-006904, and also from the Pluriregional Operational Programme of Spain 2014–2020 of the European Regional Development Fund (ERDF), ICTS-2019-02-CESGA-3, and from the State Programme for the Promotion of Scientific and Technical Research of Excellence of the State Plan for Scientific and Technical Research and Innovation 2013–2016 State subprogramme for scientific and technical infrastructures and equipment of ERDF, CESG15-DE-3114.</p></ack><notes notes-type="financialsupport"><title>Financial support</title>

      <p id="d2e12992">Kristoffer Aalstad acknowledges funding from an ESA-CCI Research Fellowship (PATCHES project) and the ERC-2022-ADG under grant agreement no. 101096057 GLACMASS. Esteban Alonso-González acknowledges funding from an ESA-CCI Research Fellowship (SnowHotspots project) and the “Ramon y Cajal” Fellowship RYC2023-044416-I. Norbert Pirk acknowledeges funding from the European Research Council (ACTIVATE project #101116083). Ruitang Yang acknowledges funding from the Research Council of Norway (GLACMOD project #324131). This work is a contribution to the strategic research initiative LATICE (#UiO/GEO103920), the Center for Biogeochemistry in the Anthropocene, as well as the Center for Computational and Data Science at the University of Oslo. The article processing charges for this open-access publication were covered in part by the CSIC Open Access Publication Support Initiative through its Unit of Information Resources for Research (URICI).</p>
  </notes><notes notes-type="reviewstatement"><title>Review statement</title>

      <p id="d2e13001">This paper was edited by Fabien Maussion and reviewed by Steven Margulis and Richard L. H. Essery.</p>
  </notes><ref-list>
    <title>References</title>

      <ref id="bib1.bibx1"><label>Aalstad et al.(2018)Aalstad, Westermann, Schuler, Boike, and Bertino</label><mixed-citation>Aalstad, K., Westermann, S., Schuler, T. V., Boike, J., and Bertino, L.: Ensemble-based assimilation of fractional snow-covered area satellite retrievals to estimate the snow distribution at Arctic sites, The Cryosphere, 12, 247–270, <ext-link xlink:href="https://doi.org/10.5194/tc-12-247-2018" ext-link-type="DOI">10.5194/tc-12-247-2018</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx2"><label>Alonso-González(2022)</label><mixed-citation>Alonso-González, E.: Inputs (forcing and observations) ready for use by “MuSA: The Multiscale Snow Data Assimilation System (v1.0)”, Zenodo [code], <ext-link xlink:href="https://doi.org/10.5281/zenodo.7248635" ext-link-type="DOI">10.5281/zenodo.7248635</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx3"><label>Alonso-González and Aalstad(2025)</label><mixed-citation>Alonso-González, E. and Aalstad, K.: MuSA: v2.3 AdaPBS submission, Zenodo [code], <ext-link xlink:href="https://doi.org/10.5281/zenodo.17292981" ext-link-type="DOI">10.5281/zenodo.17292981</ext-link>,  2025.</mixed-citation></ref>
      <ref id="bib1.bibx4"><label>Alonso González and Aalstad(2026)</label><mixed-citation>Alonso González, E. and Aalstad, K.: Code and data to reproduce the results and figures in: Evolving beyond collapse: An adaptive particle batch smoother for cryospheric data assimilation, Zenodo [code and data set], <ext-link xlink:href="https://doi.org/10.5281/zenodo.21244337" ext-link-type="DOI">10.5281/zenodo.21244337</ext-link>, 2026.</mixed-citation></ref>
      <ref id="bib1.bibx5"><label>Alonso-González et al.(2018)Alonso-González, López-Moreno, Gascoin, García-Valdecasas Ojeda, Sanmiguel-Vallelado, Navarro-Serrano, Revuelto, Ceballos, Esteban-Parra, and Essery</label><mixed-citation>Alonso-González, E., López-Moreno, J. I., Gascoin, S., García-Valdecasas Ojeda, M., Sanmiguel-Vallelado, A., Navarro-Serrano, F., Revuelto, J., Ceballos, A., Esteban-Parra, M. J., and Essery, R.: Daily gridded datasets of snow depth and snow water equivalent for the Iberian Peninsula from 1980 to 2014, Earth Syst. Sci. Data, 10, 303–315, <ext-link xlink:href="https://doi.org/10.5194/essd-10-303-2018" ext-link-type="DOI">10.5194/essd-10-303-2018</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx6"><label>Alonso-González et al.(2021)Alonso-González, Gutmann, Aalstad, Fayad, Bouchet, and Gascoin</label><mixed-citation>Alonso-González, E., Gutmann, E., Aalstad, K., Fayad, A., Bouchet, M., and Gascoin, S.: Snowpack dynamics in the Lebanese mountains from quasi-dynamically downscaled ERA5 reanalysis updated by assimilating remotely sensed fractional snow-covered area, Hydrol. Earth Syst. Sci., 25, 4455–4471, <ext-link xlink:href="https://doi.org/10.5194/hess-25-4455-2021" ext-link-type="DOI">10.5194/hess-25-4455-2021</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx7"><label>Alonso-González et al.(2022)Alonso-González, Aalstad, Baba, Revuelto, López-Moreno, Fiddes, Essery, and Gascoin</label><mixed-citation>Alonso-González, E., Aalstad, K., Baba, M. W., Revuelto, J., López-Moreno, J. I., Fiddes, J., Essery, R., and Gascoin, S.: The Multiple Snow Data Assimilation System (MuSA v1.0), Geosci. Model Dev., 15, 9127–9155, <ext-link xlink:href="https://doi.org/10.5194/gmd-15-9127-2022" ext-link-type="DOI">10.5194/gmd-15-9127-2022</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx8"><label>Alonso-González et al.(2023)Alonso-González, Aalstad, Pirk, Mazzolini, Treichler, Leclercq, Westermann, López-Moreno, and Gascoin</label><mixed-citation>Alonso-González, E., Aalstad, K., Pirk, N., Mazzolini, M., Treichler, D., Leclercq, P., Westermann, S., López-Moreno, J. I., and Gascoin, S.: Spatio-temporal information propagation using sparse observations in hyper-resolution ensemble-based snow data assimilation, Hydrol. Earth Syst. Sci., 27, 4637–4659, <ext-link xlink:href="https://doi.org/10.5194/hess-27-4637-2023" ext-link-type="DOI">10.5194/hess-27-4637-2023</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx9"><label>Alonso-González et al.(2026)Alonso-González, Harpold, Lundquist, Piske, Sourp, Aalstad, and Gascoin</label><mixed-citation>Alonso-González, E., Harpold, A., Lundquist, J. D., Piske, C., Sourp, L., Aalstad, K., and Gascoin, S.: Ensemble-based data assimilation improves hyperresolution snowpack simulations in forests, The Cryosphere, 20, 209–225, <ext-link xlink:href="https://doi.org/10.5194/tc-20-209-2026" ext-link-type="DOI">10.5194/tc-20-209-2026</ext-link>, 2026.</mixed-citation></ref>
      <ref id="bib1.bibx10"><label>Bannister(2017)</label><mixed-citation>Bannister, R. N.: A review of operational methods of variational and ensemble‐variational data assimilation, Q. J. Roy. Meteor. Soc., 143, 607–633, <ext-link xlink:href="https://doi.org/10.1002/qj.2982" ext-link-type="DOI">10.1002/qj.2982</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx11"><label>Barnett et al.(2005)Barnett, Adam, and Lettenmaier</label><mixed-citation>Barnett, T. P., Adam, J. C., and Lettenmaier, D. P.: Potential impacts of a warming climate on water availability in snow-dominated regions, Nature, 438, 303–309, <ext-link xlink:href="https://doi.org/10.1038/nature04141" ext-link-type="DOI">10.1038/nature04141</ext-link>, 2005.</mixed-citation></ref>
      <ref id="bib1.bibx12"><label>Bertino et al.(2003)Bertino, Evensen, and Wackernagel</label><mixed-citation>Bertino, L., Evensen, G., and Wackernagel, H.: Sequential Data Assimilation Techniques in Oceanography, Int. Stat. Rev., 71, 223–241, <ext-link xlink:href="https://doi.org/10.1111/j.1751-5823.2003.tb00194.x" ext-link-type="DOI">10.1111/j.1751-5823.2003.tb00194.x</ext-link>, 2003.</mixed-citation></ref>
      <ref id="bib1.bibx13"><label>Beven(2006)</label><mixed-citation>Beven, K.: A manifesto for the equifinality thesis, J. Hydrol., 320, 18–36, <ext-link xlink:href="https://doi.org/10.1016/j.jhydrol.2005.07.007" ext-link-type="DOI">10.1016/j.jhydrol.2005.07.007</ext-link>, 2006.</mixed-citation></ref>
      <ref id="bib1.bibx14"><label>Beven and Binley(1992)</label><mixed-citation>Beven, K. and Binley, A.: The future of distributed models: Model calibration and uncertainty prediction, Hydrol. Process., 6, 279–298, <ext-link xlink:href="https://doi.org/10.1002/hyp.3360060305" ext-link-type="DOI">10.1002/hyp.3360060305</ext-link>, 1992.</mixed-citation></ref>
      <ref id="bib1.bibx15"><label>Bugallo et al.(2017)Bugallo, Elvira, Martino, Luengo, Miguez, and Djuric</label><mixed-citation>Bugallo, M. F., Elvira, V., Martino, L., Luengo, D., Miguez, J., and Djuric, P. M.: Adaptive Importance Sampling: The past, the present, and the future, IEEE Signal Process. Mag., 34, 60–79, <ext-link xlink:href="https://doi.org/10.1109/MSP.2017.2699226" ext-link-type="DOI">10.1109/MSP.2017.2699226</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx16"><label>Campelo and Aranha(2023)</label><mixed-citation>Campelo, F. and Aranha, C.: Lessons from the Evolutionary Computation Bestiary, Artificial Life, 29, 421–432, <ext-link xlink:href="https://doi.org/10.1162/artl_a_00402" ext-link-type="DOI">10.1162/artl_a_00402</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx17"><label>Cao et al.(2025)Cao, Aalstad, Schmidt, Westermann, and Schuler</label><mixed-citation>Cao, W., Aalstad, K., Schmidt, L., Westermann, S., and Schuler, T.: Bayesian data assimilation on an Arctic glacier: learning from large ensemble twin experiments, J. Glaciol., 71, e121, <ext-link xlink:href="https://doi.org/10.1017/jog.2025.10101" ext-link-type="DOI">10.1017/jog.2025.10101</ext-link>, 2025.</mixed-citation></ref>
      <ref id="bib1.bibx18"><label>Carrassi et al.(2018)Carrassi, Bocquet, Bertino, and Evensen</label><mixed-citation>Carrassi, A., Bocquet, M., Bertino, L., and Evensen, G.: Data assimilation in the geosciences: An overview of methods, issues, and perspectives, WIREs Clim. Change, 9, e535, <ext-link xlink:href="https://doi.org/10.1002/wcc.535" ext-link-type="DOI">10.1002/wcc.535</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx19"><label>Charrois et al.(2016)Charrois, Cosme, Dumont, Lafaysse, Morin, Libois, and Picard</label><mixed-citation>Charrois, L., Cosme, E., Dumont, M., Lafaysse, M., Morin, S., Libois, Q., and Picard, G.: On the assimilation of optical reflectances and snow depth observations into a detailed snowpack model, The Cryosphere, 10, 1021–1038, <ext-link xlink:href="https://doi.org/10.5194/tc-10-1021-2016" ext-link-type="DOI">10.5194/tc-10-1021-2016</ext-link>, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx20"><label>Chopin(2002)</label><mixed-citation>Chopin, N.: A sequential particle filter method for static models, Biometrika, 89, 539–552, <ext-link xlink:href="https://doi.org/10.1093/biomet/89.3.539" ext-link-type="DOI">10.1093/biomet/89.3.539</ext-link>, 2002.</mixed-citation></ref>
      <ref id="bib1.bibx21"><label>Chopin and Papaspiliopoulos(2020)</label><mixed-citation>Chopin, N. and Papaspiliopoulos, O.: An Introduction to Sequential Monte Carlo, Springer, <ext-link xlink:href="https://doi.org/10.1007/978-3-030-47845-2" ext-link-type="DOI">10.1007/978-3-030-47845-2</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx22"><label>Clark et al.(2006)Clark, Slater, Barrett, Hay, McCabe, Rajagopalan, and Leavesley</label><mixed-citation>Clark, M. P., Slater, A. G., Barrett, A. P., Hay, L. E., McCabe, G. J., Rajagopalan, B., and Leavesley, G. H.: Assimilation of snow covered area information into hydrologic and land-surface models, Adv. Water Resour., 29, 1209–1221, <ext-link xlink:href="https://doi.org/10.1016/j.advwatres.2005.10.001" ext-link-type="DOI">10.1016/j.advwatres.2005.10.001</ext-link>, 2006.</mixed-citation></ref>
      <ref id="bib1.bibx23"><label>Cleary et al.(2021)Cleary, Garbuno-Inigo, Lan, Schneider, and Stuart</label><mixed-citation>Cleary, E., Garbuno-Inigo, A., Lan, S., Schneider, T., and Stuart, A.: Calibrate, emulate, sample, J. Comput. Phys., 424, 109716, <ext-link xlink:href="https://doi.org/10.1016/j.jcp.2020.109716" ext-link-type="DOI">10.1016/j.jcp.2020.109716</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx24"><label>Cluzet et al.(2021)Cluzet, Lafaysse, Cosme, Albergel, Meunier, and Dumont</label><mixed-citation>Cluzet, B., Lafaysse, M., Cosme, E., Albergel, C., Meunier, L.-F., and Dumont, M.: CrocO_v1.0: a particle filter to assimilate snowpack observations in a spatialised framework, Geosci. Model Dev., 14, 1595–1614, <ext-link xlink:href="https://doi.org/10.5194/gmd-14-1595-2021" ext-link-type="DOI">10.5194/gmd-14-1595-2021</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx25"><label>Cluzet et al.(2024)Cluzet, Magnusson, Quéno, Mazzotti, Mott, and Jonas</label><mixed-citation>Cluzet, B., Magnusson, J., Quéno, L., Mazzotti, G., Mott, R., and Jonas, T.: Exploring how Sentinel-1 wet-snow maps can inform fully distributed physically based snowpack models, The Cryosphere, 18, 5753–5767, <ext-link xlink:href="https://doi.org/10.5194/tc-18-5753-2024" ext-link-type="DOI">10.5194/tc-18-5753-2024</ext-link>, 2024.</mixed-citation></ref>
      <ref id="bib1.bibx26"><label>Cornuet et al.(2012)Cornuet, Marin, Antonietta, and Robert</label><mixed-citation>Cornuet, J.-M., Marin, J.-M., Antonietta, M., and Robert, C.: Adaptive Multiple Importance Sampling, Scand. J. Stat., 39, 798–812, <ext-link xlink:href="https://doi.org/10.1111/j.1467-9469.2011.00756.x" ext-link-type="DOI">10.1111/j.1467-9469.2011.00756.x</ext-link>, 2012.</mixed-citation></ref>
      <ref id="bib1.bibx27"><label>Cortés and Margulis(2017)</label><mixed-citation>Cortés, G. and Margulis, S.: Impacts of El Niño and La Niña on interannual snow accumulation in the Andes: Results from a high-resolution 31 year reanalysis: El Niño Effects on Andes Snow, Geophys. Res. Lett., 44, 6859–6867, <ext-link xlink:href="https://doi.org/10.1002/2017GL073826" ext-link-type="DOI">10.1002/2017GL073826</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx28"><label>De Lannoy et al.(2012)De Lannoy, Reichle, Arsenault, Houser, Kumar, Verhoest, and Pauwels</label><mixed-citation>De Lannoy, G. J. M., Reichle, R. H., Arsenault, K. R., Houser, P. R., Kumar, S., Verhoest, N. E. C., and Pauwels, V. R. N.: Multiscale assimilation of Advanced Microwave Scanning Radiometer-EOS snow water equivalent and Moderate Resolution Imaging Spectroradiometer snow cover fraction observations in northern Colorado, Water Resour. Res., 48, W01522, <ext-link xlink:href="https://doi.org/10.1029/2011WR010588" ext-link-type="DOI">10.1029/2011WR010588</ext-link>, 2012.</mixed-citation></ref>
      <ref id="bib1.bibx29"><label>De Lannoy et al.(2024)De Lannoy, Bechtold, Busschaert, Heyvaert, Modanesi, Dunmire, Lievens, Getirana, and Massari</label><mixed-citation>De Lannoy, G. J. M., Bechtold, M., Busschaert, L., Heyvaert, Z., Modanesi, S., Dunmire, D., Lievens, H., Getirana, A., and Massari, C.: Contributions of Irrigation Modeling, Soil Moisture and Snow Data Assimilation to High-Resolution Water Budget Estimates Over the Po Basin: Progress Towards Digital Replicas, J. Adv. Model. Earth Sy., 16, e2024MS004433, <ext-link xlink:href="https://doi.org/10.1029/2024MS004433" ext-link-type="DOI">10.1029/2024MS004433</ext-link>, 2024.</mixed-citation></ref>
      <ref id="bib1.bibx30"><label>Del Moral(2004)</label><mixed-citation>Del Moral, P.: Feynman-Kac Formulae, Springer, <ext-link xlink:href="https://doi.org/10.1007/978-1-4684-9393-1" ext-link-type="DOI">10.1007/978-1-4684-9393-1</ext-link>, 2004.</mixed-citation></ref>
      <ref id="bib1.bibx31"><label>Doucet et al.(2000)Doucet, Godsill, and Andrieu</label><mixed-citation>Doucet, A., Godsill, S., and Andrieu, C.: On sequential Monte Carlo sampling methods for Bayesian filtering, Stat. Comput., 10, 197–208, <ext-link xlink:href="https://doi.org/10.1023/A:1008935410038" ext-link-type="DOI">10.1023/A:1008935410038</ext-link>, 2000.</mixed-citation></ref>
      <ref id="bib1.bibx32"><label>Dozier et al.(2016)Dozier, Bair, and Davis</label><mixed-citation>Dozier, J., Bair, E. H., and Davis, R. E.: Estimating the spatial distribution of snow water equivalent in the world’s mountains, Wires Water, 3, 461–474, <ext-link xlink:href="https://doi.org/10.1002/wat2.1140" ext-link-type="DOI">10.1002/wat2.1140</ext-link>, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx33"><label>Duan et al.(2023)Duan, Li, and Xu</label><mixed-citation>Duan, J.-C., Li, S., and Xu, Y.: Sequential Monte Carlo optimization and statistical inference, WIREs Comput. Stat., 15, e1598, <ext-link xlink:href="https://doi.org/10.1002/wics.1598" ext-link-type="DOI">10.1002/wics.1598</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx34"><label>Durand et al.(2008)Durand, Molotch, and Margulis</label><mixed-citation>Durand, M., Molotch, N. P., and Margulis, S. A.: A Bayesian approach to snow water equivalent reconstruction, J. Geophys. Res.-Atmos., 113, D20117, <ext-link xlink:href="https://doi.org/10.1029/2008JD009894" ext-link-type="DOI">10.1029/2008JD009894</ext-link>, 2008.</mixed-citation></ref>
      <ref id="bib1.bibx35"><label>Elias Chereque et al.(2024)Elias Chereque, Kushner, Mudryk, Derksen, and Mortimer</label><mixed-citation>Elias Chereque, A., Kushner, P. J., Mudryk, L., Derksen, C., and Mortimer, C.: A simple snow temperature index model exposes discrepancies between reanalysis snow water equivalent products, The Cryosphere, 18, 4955–4969, <ext-link xlink:href="https://doi.org/10.5194/tc-18-4955-2024" ext-link-type="DOI">10.5194/tc-18-4955-2024</ext-link>, 2024.</mixed-citation></ref>
      <ref id="bib1.bibx36"><label>Elvira et al.(2019)Elvira, Martino, Bugallo, and Djuric</label><mixed-citation>Elvira, V., Martino, L., Bugallo, M. F., and Djuric, P. M.: Elucidating the Auxiliary Particle Filter via Multiple Importance Sampling [Lecture Notes], IEEE Signal Process. Mag., 36, 145–152, <ext-link xlink:href="https://doi.org/10.1109/MSP.2019.2938026" ext-link-type="DOI">10.1109/MSP.2019.2938026</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx37"><label>Elvira et al.(2022)Elvira, Martino, and Robert</label><mixed-citation>Elvira, V., Martino, L., and Robert, C.: Rethinking the Effective Sample Size, Int. Stat. Rev., 90, 525–550, <ext-link xlink:href="https://doi.org/10.1111/insr.12500" ext-link-type="DOI">10.1111/insr.12500</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx38"><label>Emerick and Reynolds(2011)</label><mixed-citation>Emerick, A. A. and Reynolds, A. C.: Combining the Ensemble Kalman Filter with Markov Chain Monte Carlo for Improved History Matching and Uncertainty Characterization,  SPE Reservoir Simulation Symposium, The Woodlands, Texas, USA, February 2011,  SPE-141336-MS, <ext-link xlink:href="https://doi.org/10.2118/141336-MS" ext-link-type="DOI">10.2118/141336-MS</ext-link>, 2011.</mixed-citation></ref>
      <ref id="bib1.bibx39"><label>Emerick and Reynolds(2013)</label><mixed-citation>Emerick, A. A. and Reynolds, A. C.: Ensemble smoother with multiple data assimilation, Comput. Geosci., 55, 3–15, <ext-link xlink:href="https://doi.org/10.1016/j.cageo.2012.03.011" ext-link-type="DOI">10.1016/j.cageo.2012.03.011</ext-link>, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx40"><label>Essery(2015)</label><mixed-citation>Essery, R.: A factorial snowpack model (FSM 1.0), Geosci. Model Dev., 8, 3867–3876, <ext-link xlink:href="https://doi.org/10.5194/gmd-8-3867-2015" ext-link-type="DOI">10.5194/gmd-8-3867-2015</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx41"><label>Essery et al.(2025)Essery, Mazzotti, Barr, Jonas, Quaife, and Rutter</label><mixed-citation>Essery, R., Mazzotti, G., Barr, S., Jonas, T., Quaife, T., and Rutter, N.: A Flexible Snow Model (FSM 2.1.1) including a forest canopy, Geosci. Model Dev., 18, 3583–3605, <ext-link xlink:href="https://doi.org/10.5194/gmd-18-3583-2025" ext-link-type="DOI">10.5194/gmd-18-3583-2025</ext-link>, 2025.</mixed-citation></ref>
      <ref id="bib1.bibx42"><label>Euskirchen et al.(2013)Euskirchen, Goodstein, and Huntington</label><mixed-citation>Euskirchen, E. S., Goodstein, E. S., and Huntington, H. P.: An estimated cost of lost climate regulation services caused by thawing of the Arctic cryosphere, Ecol. Appl., 23, 1869–1880, <ext-link xlink:href="https://doi.org/10.1890/11-0858.1" ext-link-type="DOI">10.1890/11-0858.1</ext-link>, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx43"><label>Evensen(1994)</label><mixed-citation>Evensen, G.: Sequential data assimilation with a nonlinear quasi-geostrophic model using Monte Carlo methods to forecast error statistics, J. Geophys. Res., 99, 10143, <ext-link xlink:href="https://doi.org/10.1029/94JC00572" ext-link-type="DOI">10.1029/94JC00572</ext-link>, 1994.</mixed-citation></ref>
      <ref id="bib1.bibx44"><label>Evensen et al.(2022)Evensen, Vossepoel, and van Leeuwen</label><mixed-citation>Evensen, G., Vossepoel, F. C., and van Leeuwen, P. J.: Data Assimilation Fundamentals, Springer, <ext-link xlink:href="https://doi.org/10.1007/978-3-030-96709-3" ext-link-type="DOI">10.1007/978-3-030-96709-3</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx45"><label>Farchi and Bocquet(2018)</label><mixed-citation>Farchi, A. and Bocquet, M.: Review article: Comparison of local particle filters and new implementations, Nonlin. Processes Geophys., 25, 765–807, <ext-link xlink:href="https://doi.org/10.5194/npg-25-765-2018" ext-link-type="DOI">10.5194/npg-25-765-2018</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx46"><label>Fiddes et al.(2019)Fiddes, Aalstad, and Westermann</label><mixed-citation>Fiddes, J., Aalstad, K., and Westermann, S.: Hyper-resolution ensemble-based snow reanalysis in mountain regions using clustering, Hydrol. Earth Syst. Sci., 23, 4717–4736, <ext-link xlink:href="https://doi.org/10.5194/hess-23-4717-2019" ext-link-type="DOI">10.5194/hess-23-4717-2019</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx47"><label>Garbuno-Inigo et al.(2020)Garbuno-Inigo, Hoffmann, Li, and Stuart</label><mixed-citation>Garbuno-Inigo, A., Hoffmann, F., Li, W., and Stuart, A.: Interacting Langevin Diffusions: Gradient Structure and Ensemble Kalman Sampler, SIAM J. Appl. Dynam. Syst., 19, 412–441, <ext-link xlink:href="https://doi.org/10.1137/19M1251655" ext-link-type="DOI">10.1137/19M1251655</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx48"><label>Gascoin(2024)</label><mixed-citation>Gascoin, S.: A call for an accurate presentation of glaciers as water resources, WIREs Water, 11, e1705, <ext-link xlink:href="https://doi.org/10.1002/wat2.1705" ext-link-type="DOI">10.1002/wat2.1705</ext-link>, 2024.</mixed-citation></ref>
      <ref id="bib1.bibx49"><label>Gascoin et al.(2024)Gascoin, Luojus, Nagler, Lievens, Masiokas, Jonas, Zheng, and De Rosnay</label><mixed-citation>Gascoin, S., Luojus, K., Nagler, T., Lievens, H., Masiokas, M., Jonas, T., Zheng, Z., and De Rosnay, P.: Remote sensing of mountain snow from space: status and recommendations, Front. Earth Sci., 12, 138323, <ext-link xlink:href="https://doi.org/10.3389/feart.2024.1381323" ext-link-type="DOI">10.3389/feart.2024.1381323</ext-link>, 2024.</mixed-citation></ref>
      <ref id="bib1.bibx50"><label>Gelman et al.(2013)Gelman, Carlin, Stern, Dunson, Vehtari, and Rubin</label><mixed-citation>Gelman, A., Carlin, J., Stern, H., Dunson, D., Vehtari, A., and Rubin, D.: Bayesian Data Analysis, CRC Press, 3rd Edn., 675 pp., <ext-link xlink:href="https://doi.org/10.1201/b16018" ext-link-type="DOI">10.1201/b16018</ext-link>, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx51"><label>Gilks and Berzuini(2001)</label><mixed-citation>Gilks, W. R. and Berzuini, C.: Following a moving target-Monte Carlo inference for dynamic Bayesian models, J. Roy. Stat. Soc. Ser. B, 63, 127–146, <ext-link xlink:href="https://doi.org/10.1111/1467-9868.00280" ext-link-type="DOI">10.1111/1467-9868.00280</ext-link>, 2001.</mixed-citation></ref>
      <ref id="bib1.bibx52"><label>Girotto et al.(2014)Girotto, Margulis, and Durand</label><mixed-citation>Girotto, M., Margulis, S., and Durand, M.: Probabilistic SWE reanalysis as a generalization of deterministic SWE reconstruction techniques, Hydrol. Process., 28, 3875–3895, <ext-link xlink:href="https://doi.org/10.1002/hyp.9887" ext-link-type="DOI">10.1002/hyp.9887</ext-link>, 2014.</mixed-citation></ref>
      <ref id="bib1.bibx53"><label>Girotto et al.(2020)Girotto, Musselman, and Essery</label><mixed-citation>Girotto, M., Musselman, K. N., and Essery, R. L. H.: Data Assimilation Improves Estimates of Climate-Sensitive Seasonal Snow, Curr. Clim. Change Rep., 6, 81–94, <ext-link xlink:href="https://doi.org/10.1007/s40641-020-00159-7" ext-link-type="DOI">10.1007/s40641-020-00159-7</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx54"><label>Girotto et al.(2024)Girotto, Formetta, Azimi, Bachand, Cowherd, De Lannoy, Lievens, Modanesi, Raleigh, Rigon, and Massari</label><mixed-citation>Girotto, M., Formetta, G., Azimi, S., Bachand, C., Cowherd, M., De Lannoy, G., Lievens, H., Modanesi, S., Raleigh, M. S., Rigon, R., and Massari, C.: Identifying snowfall elevation patterns by assimilating satellite-based snow depth retrievals, Sci. Total Environ., 906, 167312, <ext-link xlink:href="https://doi.org/10.1016/j.scitotenv.2023.167312" ext-link-type="DOI">10.1016/j.scitotenv.2023.167312</ext-link>, 2024.</mixed-citation></ref>
      <ref id="bib1.bibx55"><label>Gneiting and Raftery(2007)</label><mixed-citation>Gneiting, T. and Raftery, A. E.: Strictly Proper Scoring Rules, Prediction, and Estimation, J. Am. Stat. Assoc., 102, 359–378, <ext-link xlink:href="https://doi.org/10.1198/016214506000001437" ext-link-type="DOI">10.1198/016214506000001437</ext-link>, 2007.</mixed-citation></ref>
      <ref id="bib1.bibx56"><label>Gneiting et al.(2005)Gneiting, Raftery, Westveld III, and Goldman</label><mixed-citation> Gneiting, T., Raftery, A. E., Westveld III, A. H., and Goldman, T.: Calibrated Probabilistic Forecasting Using Ensemble Model Output Statistics and Minimum CRPS Estimation, Mon. Weather Rev., 133, 1098–1118, 2005.</mixed-citation></ref>
      <ref id="bib1.bibx57"><label>Gordon et al.(1993)Gordon, Salmond, and Smith</label><mixed-citation>Gordon, N. J., Salmond, D. J., and Smith, A. F. M.: Novel approach to nonlinear/non-Gaussian Bayesian state estimation, IEE Proc.-F, 140, 107, <ext-link xlink:href="https://doi.org/10.1049/ip-f-2.1993.0015" ext-link-type="DOI">10.1049/ip-f-2.1993.0015</ext-link>, 1993.</mixed-citation></ref>
      <ref id="bib1.bibx58"><label>Gottlieb and Mankin(2024)</label><mixed-citation>Gottlieb, A. R. and Mankin, J. S.: Evidence of human influence on Northern Hemisphere snow loss, Nature, 625, 293–300, <ext-link xlink:href="https://doi.org/10.1038/s41586-023-06794-y" ext-link-type="DOI">10.1038/s41586-023-06794-y</ext-link>, 2024.</mixed-citation></ref>
      <ref id="bib1.bibx59"><label>Groenke et al.(2023)Groenke, Langer, Nitzbon, Westermann, Gallego, and Boike</label><mixed-citation>Groenke, B., Langer, M., Nitzbon, J., Westermann, S., Gallego, G., and Boike, J.: Investigating the thermal state of permafrost with Bayesian inverse modeling of heat transfer, The Cryosphere, 17, 3505–3533, <ext-link xlink:href="https://doi.org/10.5194/tc-17-3505-2023" ext-link-type="DOI">10.5194/tc-17-3505-2023</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx60"><label>Günther et al.(2019)Günther, Marke, Essery, and Strasser</label><mixed-citation>Günther, D., Marke, T., Essery, R., and Strasser, U.: Uncertainties in Snowpack Simulations – Assessing the Impact of Model Structure, Parameter Choice, and Forcing Data Error on Point-Scale Energy Balance Snow Model Performance, Water Resour. Res., 55, 2779–2800, <ext-link xlink:href="https://doi.org/10.1029/2018WR023403" ext-link-type="DOI">10.1029/2018WR023403</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx61"><label>Hammersley and Morton(1954)</label><mixed-citation>Hammersley, J. and Morton, K.: Poor Man's Monte Carlo, J. Roy. Stat. Soc. Ser. B, 16, 23–38, <ext-link xlink:href="https://doi.org/10.1111/j.2517-6161.1954.tb00145.x" ext-link-type="DOI">10.1111/j.2517-6161.1954.tb00145.x</ext-link>, 1954.</mixed-citation></ref>
      <ref id="bib1.bibx62"><label>Hastings(1970)</label><mixed-citation>Hastings, W.: Monte Carlo Sampling Methods Using Markov Chains and Their Applications, Biometrika, 57, <ext-link xlink:href="https://doi.org/10.2307/2334940" ext-link-type="DOI">10.2307/2334940</ext-link>, 1970.</mixed-citation></ref>
      <ref id="bib1.bibx63"><label>Hennig et al.(2022)Hennig, Osborne, and Kersting</label><mixed-citation>Hennig, P., Osborne, M., and Kersting, H.: Probabilistic Numerics, Cambridge, <ext-link xlink:href="https://doi.org/10.1017/9781316681411" ext-link-type="DOI">10.1017/9781316681411</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx64"><label>Hersbach et al.(2020)Hersbach, Bell, Berrisford, Hirahara, Horányi, Muñoz-Sabater et al.</label><mixed-citation>Hersbach, H., Bell, B., Berrisford, P., Hirahara, S., Horányi, A., Muñoz-Sabater, J., Nicolas, J., Peubey, C., Radu, R., Schepers, D., Simmons, A., Soci, C., Abdalla, S., Abellan, X., Balsamo, G., Bechtold, P., Biavati, G., Bidlot, J., Bonavita, M., De Chiara, G., Dahlgren, P., Dee, D., Diamantakis, M., Dragani, R., Flemming, J., Forbes, R., Fuentes, M., Geer, A., Haimberger, L., Healy, S., Hogan, R. J., Hólm, E., Janisková, M., Keeley, S., Laloyaux, P., Lopez, P., Lupu, C., Radnoti, G., de Rosnay, P., Rozum, I., Vamborg, F., Villaume, S., and Thépaut, J.-N.: The ERA5 global reanalysis, Q.  J. Roy. Meteor. Soc., 146, 1999–2049, <ext-link xlink:href="https://doi.org/10.1002/qj.3803" ext-link-type="DOI">10.1002/qj.3803</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx65"><label>Hock(2003)</label><mixed-citation>Hock, R.: Temperature index melt modelling in mountain areas, J. Hydrol., 282, 104–115, <ext-link xlink:href="https://doi.org/10.1016/S0022-1694(03)00257-9" ext-link-type="DOI">10.1016/S0022-1694(03)00257-9</ext-link>, 2003.</mixed-citation></ref>
      <ref id="bib1.bibx66"><label>Hock et al.(2019)</label><mixed-citation>Hock, R., Rasul, G.,  Adler, C., Cáceres, B., Gruber, S., Hirabayashi, Y., Jackson, M., Kääb, A., Kang, S., Kutuzov, S., Milner, A., Molau, U., Morin, S., Orlove, B., and Steltzer, H.: High Mountain Areas, in: IPCC Special Report on the Ocean and Cryosphere in a Changing Climate, edited by: Pörtner, H.-O., Roberts, C. C., Masson-Delmotte, V., Zhai, P., Tignor, M., Poloczanska, E., Mintenbeck, K., Alegría, A., Nicolai, M., Okem, A., Petzold, J., Rama, B., and Weyer, N. M., Cambridge, 131–202, <ext-link xlink:href="https://doi.org/10.1017/9781009157964.004" ext-link-type="DOI">10.1017/9781009157964.004</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx67"><label>Holland(1992)</label><mixed-citation> Holland, J.: Adaptation in Natural and Artificial Systems, MIT, 2nd Edn., ISBN 9780262581110, 1992.</mixed-citation></ref>
      <ref id="bib1.bibx68"><label>Immerzeel et al.(2020)Immerzeel, Lutz, Andrade, Bahl, Biemans, Bolch, Hyde, Brumby, Davies, Elmore, Emmer, Feng, Fernánde, Haritashya, Kargel, Koppes, Kraaijenbrink, Kulkarni, Mayewski, Nepal, Pacheco, Painter, Pellicciotti, Rupper, Sinisalo, Srestha, Viviroli, Wada, Xiao, Yao, and Baillie</label><mixed-citation>Immerzeel, W., Lutz, A., Andrade, M., Bahl, A., Biemans, H., Bolch, T., Hyde, S., Brumby, S., Davies, B., Elmore, A., Emmer, A., Feng, M., Fernánde, A., Haritashya, U., Kargel, J., Koppes, M., Kraaijenbrink, P., Kulkarni, A., Mayewski, P., Nepal, S., Pacheco, P., Painter, T., Pellicciotti, F. Rajaram, H., Rupper, S., Sinisalo, A., Srestha, A., Viviroli, D., Wada, Y., Xiao, C., Yao, T., and Baillie, J.: Importance and vulnerability of the world’s water towers, Nature, 577, 364–369, <ext-link xlink:href="https://doi.org/10.1038/s41586-019-1822-y" ext-link-type="DOI">10.1038/s41586-019-1822-y</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx69"><label>Jaynes(2003)</label><mixed-citation>Jaynes, E.: Probability Theory, Cambridge, <ext-link xlink:href="https://doi.org/10.1017/CBO9780511790423" ext-link-type="DOI">10.1017/CBO9780511790423</ext-link>, 2003.</mixed-citation></ref>
      <ref id="bib1.bibx70"><label>Jazwinski(1970)</label><mixed-citation>Jazwinski, A.: Stochastic Processes and Filtering Theory, Academic Press, Vol. 64, 376 pp., <ext-link xlink:href="https://doi.org/10.1016/S0076-5392(09)X6022-4" ext-link-type="DOI">10.1016/S0076-5392(09)X6022-4</ext-link>, 1970.</mixed-citation></ref>
      <ref id="bib1.bibx71"><label>Keetz et al.(2025)Keetz, Aalstad, Fisher, Poppe Terán, Naz, Pirk, Yilmaz, and Skarpaas</label><mixed-citation>Keetz, L. T., Aalstad, K., Fisher, R. A., Poppe Terán, C., Naz, B., Pirk, N., Yilmaz, Y. A., and Skarpaas, O.: Inferring Parameters in a Complex Land Surface Model by Combining Data Assimilation and Machine Learning, J. Adv. Model. Earth Sy., 17, e2024MS004542, <ext-link xlink:href="https://doi.org/10.1029/2024MS004542" ext-link-type="DOI">10.1029/2024MS004542</ext-link>, 2025.</mixed-citation></ref>
      <ref id="bib1.bibx72"><label>Kitagawa(1996)</label><mixed-citation>Kitagawa, G.: Monte Carlo Filter and Smoother for Non-Gaussian Nonlinear State Space Models, J. Comput. Graph. Stat., 5, 1–25, <ext-link xlink:href="https://doi.org/10.1080/10618600.1996.10474692" ext-link-type="DOI">10.1080/10618600.1996.10474692</ext-link>, 1996.</mixed-citation></ref>
      <ref id="bib1.bibx73"><label>Koblents and Míguez(2015)</label><mixed-citation>Koblents, E. and Míguez, J.: A population Monte Carlo scheme with transformed weights and its application to stochastic kinetic models, Stat. Comput., 25, 407–425, <ext-link xlink:href="https://doi.org/10.1007/s11222-013-9440-2" ext-link-type="DOI">10.1007/s11222-013-9440-2</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx74"><label>Krinner et al.(2018)</label><mixed-citation>Krinner, G., Derksen, C., Essery, R., Flanner, M., Hagemann, S., Clark, M., Hall, A., Rott, H., Brutel-Vuilmet, C., Kim, H., Ménard, C. B., Mudryk, L., Thackeray, C., Wang, L., Arduini, G., Balsamo, G., Bartlett, P., Boike, J., Boone, A., Chéruy, F., Colin, J., Cuntz, M., Dai, Y., Decharme, B., Derry, J., Ducharne, A., Dutra, E., Fang, X., Fierz, C., Ghattas, J., Gusev, Y., Haverd, V., Kontu, A., Lafaysse, M., Law, R., Lawrence, D., Li, W., Marke, T., Marks, D., Ménégoz, M., Nasonova, O., Nitta, T., Niwano, M., Pomeroy, J., Raleigh, M. S., Schaedler, G., Semenov, V., Smirnova, T. G., Stacke, T., Strasser, U., Svenson, S., Turkov, D., Wang, T., Wever, N., Yuan, H., Zhou, W., and Zhu, D.: ESM-SnowMIP: assessing snow models and quantifying snow-related climate feedbacks, Geosci. Model Dev., 11, 5027–5049, <ext-link xlink:href="https://doi.org/10.5194/gmd-11-5027-2018" ext-link-type="DOI">10.5194/gmd-11-5027-2018</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx75"><label>Landmann et al.(2021)Landmann, Künsch, Huss, Ogier, Kalisch, and Farinotti</label><mixed-citation>Landmann, J. M., Künsch, H. R., Huss, M., Ogier, C., Kalisch, M., and Farinotti, D.: Assimilating near-real-time mass balance stake readings into a model ensemble using a particle filter, The Cryosphere, 15, 5017–5040, <ext-link xlink:href="https://doi.org/10.5194/tc-15-5017-2021" ext-link-type="DOI">10.5194/tc-15-5017-2021</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx76"><label>Largeron et al.(2020)Largeron, Dumont, Morin, Boone, Lafaysse, Metref, Cosme, Jonas, Wisntral, and Margulis</label><mixed-citation>Largeron, C., Dumont, M., Morin, S., Boone, A., Lafaysse, M., Metref, S., Cosme, E., Jonas, T., Wisntral, A., and Margulis, S.: Toward Snow Cover Estimation in Mountainous Areas Using Modern Data Assimilation Methods: A Review, Front. Earth Sci., 8, 325, <ext-link xlink:href="https://doi.org/10.3389/feart.2020.00325" ext-link-type="DOI">10.3389/feart.2020.00325</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx77"><label>Law and Stuart(2012)</label><mixed-citation>Law, K. and Stuart, A.: Evaluating Data Assimilation Algorithms, Mon. Weather Rev., 140, 3757–3782, <ext-link xlink:href="https://doi.org/10.1175/MWR-D-11-00257.1" ext-link-type="DOI">10.1175/MWR-D-11-00257.1</ext-link>, 2012.</mixed-citation></ref>
      <ref id="bib1.bibx78"><label>Leisenring and Moradkhani(2011)</label><mixed-citation>Leisenring, M. and Moradkhani, H.: Snow water equivalent prediction using Bayesian data assimilation methods, Stoch. Env. Res. Risk Assess., 25, 253–270, <ext-link xlink:href="https://doi.org/10.1007/s00477-010-0445-5" ext-link-type="DOI">10.1007/s00477-010-0445-5</ext-link>, 2011.</mixed-citation></ref>
      <ref id="bib1.bibx79"><label>Li et al.(2015)Li, Bolic, and Djuric</label><mixed-citation>Li, T., Bolic, M., and Djuric, P.: Resampling Methods for Particle Filtering: Classification, implementation, and strategies, IEEE Signal Process. Mag., 32, 70–86, <ext-link xlink:href="https://doi.org/10.1109/MSP.2014.2330626" ext-link-type="DOI">10.1109/MSP.2014.2330626</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx80"><label>Liston and Elder(2006b)</label><mixed-citation>Liston, G. E. and Elder, K.: A meteorological distribution system for high-resolution terrestrial modeling (MicroMet, J. Hydrometeorol., 7, 217–234, <ext-link xlink:href="https://doi.org/10.1175/JHM486.1" ext-link-type="DOI">10.1175/JHM486.1</ext-link>, 2006b.</mixed-citation></ref>
      <ref id="bib1.bibx81"><label>Liu et al.(2021)Liu, Fang, and Margulis</label><mixed-citation>Liu, Y., Fang, Y., and Margulis, S. A.: Spatiotemporal distribution of seasonal snow water equivalent in High Mountain Asia from an 18-year Landsat–MODIS era snow reanalysis dataset, The Cryosphere, 15, 5261–5280, <ext-link xlink:href="https://doi.org/10.5194/tc-15-5261-2021" ext-link-type="DOI">10.5194/tc-15-5261-2021</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx82"><label>Lussana et al.(2018)Lussana, Saloranta, Skaugen, Magnusson, Tveito, and Andersen</label><mixed-citation>Lussana, C., Saloranta, T., Skaugen, T., Magnusson, J., Tveito, O. E., and Andersen, J.: seNorge2 daily precipitation, an observational gridded dataset over Norway from 1957 to the present day, Earth Syst. Sci. Data, 10, 235–249, <ext-link xlink:href="https://doi.org/10.5194/essd-10-235-2018" ext-link-type="DOI">10.5194/essd-10-235-2018</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx83"><label>MacKay(2003)</label><mixed-citation> MacKay, D. J. C.: Information Theory, Inference, and Learning Algorithms, Cambridge,  ISBN 9780521642989,  2003.</mixed-citation></ref>
      <ref id="bib1.bibx84"><label>Magnusson et al.(2017)Magnusson, Winstral, Stordal, Essery, and Jonas</label><mixed-citation>Magnusson, J., Winstral, A., Stordal, A. S., Essery, R., and Jonas, T.: Improving physically based snow simulations by assimilating snow depths using the particle filter, Water Resour. Res., 53, 1125–1143, <ext-link xlink:href="https://doi.org/10.1002/2016WR019092" ext-link-type="DOI">10.1002/2016WR019092</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx85"><label>Margulis et al.(2015)Margulis, Girotto, Cortés, and Durand</label><mixed-citation>Margulis, S. A., Girotto, M., Cortés, G., and Durand, M.: A particle batch smoother approach to snow water equivalent estimation, J. Hydrometeorol., 16, 1752–1772, <ext-link xlink:href="https://doi.org/10.1175/JHM-D-14-0177.1" ext-link-type="DOI">10.1175/JHM-D-14-0177.1</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx86"><label>Margulis et al.(2016)Margulis, Cortés, Girotto, and Durand</label><mixed-citation>Margulis, S. A., Cortés, G., Girotto, M., and Durand, M.: A landsat-era Sierra Nevada snow reanalysis (1985-2015, J. Hydrometeorol., 17, 1203–1221, <ext-link xlink:href="https://doi.org/10.1175/JHM-D-15-0177.1" ext-link-type="DOI">10.1175/JHM-D-15-0177.1</ext-link>, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx87"><label>Marin et al.(2019)Marin, Pudlo, and Sedki</label><mixed-citation>Marin, J.-M., Pudlo, P., and Sedki, M.: Consistency of adaptive importance sampling and recycling schemes, Bernoulli, 25, 1977–1998, <ext-link xlink:href="https://doi.org/10.3150/18-BEJ1042" ext-link-type="DOI">10.3150/18-BEJ1042</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx88"><label>Maussion et al.(2019)Maussion, Butenko, Champollion, Dusch, Eis, Fourteau, Gregor, Jarosch, Landmann, Oesterle, Recinos, Rothenpieler, Vlug, Wild, and Marzeion</label><mixed-citation>Maussion, F., Butenko, A., Champollion, N., Dusch, M., Eis, J., Fourteau, K., Gregor, P., Jarosch, A. H., Landmann, J., Oesterle, F., Recinos, B., Rothenpieler, T., Vlug, A., Wild, C. T., and Marzeion, B.: The Open Global Glacier Model (OGGM) v1.1, Geosci. Model Dev., 12, 909–931, <ext-link xlink:href="https://doi.org/10.5194/gmd-12-909-2019" ext-link-type="DOI">10.5194/gmd-12-909-2019</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx89"><label>Mazzolini et al.(2025)Mazzolini, Aalstad, Alonso-González, Westermann, and Treichler</label><mixed-citation>Mazzolini, M., Aalstad, K., Alonso-González, E., Westermann, S., and Treichler, D.: Spatio-temporal snow data assimilation with the ICESat-2 laser altimeter, The Cryosphere, 19, 3831–3848, <ext-link xlink:href="https://doi.org/10.5194/tc-19-3831-2025" ext-link-type="DOI">10.5194/tc-19-3831-2025</ext-link>, 2025.</mixed-citation></ref>
      <ref id="bib1.bibx90"><label>Ménard and Essery(2019)</label><mixed-citation>Ménard, C. and Essery, R.: ESM-SnowMIP meteorological and evaluation datasets at ten reference sites (in situ and bias corrected reanalysis data), PANGAEA [data set], <ext-link xlink:href="https://doi.org/10.1594/PANGAEA.897575" ext-link-type="DOI">10.1594/PANGAEA.897575</ext-link>,   2019.</mixed-citation></ref>
      <ref id="bib1.bibx91"><label>Ménard et al.(2019)Ménard, Essery, Barr, Bartlett, Derry, Dumont, Fierz, Kim, Kontu, Lejeune, Marks, Niwano, Raleigh, Wang, and Wever</label><mixed-citation>Ménard, C. B., Essery, R., Barr, A., Bartlett, P., Derry, J., Dumont, M., Fierz, C., Kim, H., Kontu, A., Lejeune, Y., Marks, D., Niwano, M., Raleigh, M., Wang, L., and Wever, N.: Meteorological and evaluation datasets for snow modelling at 10 reference sites: description of in situ and bias-corrected reanalysis data, Earth Syst. Sci. Data, 11, 865–880, <ext-link xlink:href="https://doi.org/10.5194/essd-11-865-2019" ext-link-type="DOI">10.5194/essd-11-865-2019</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx92"><label>Meredith et al.(2019)</label><mixed-citation>Meredith, M.,  Sommerkorn, M., Cassotta, S., Derksen, C., Ekaykin, A., Hollowed, A., Kofinas, G., Mackintosh, A., Melbourne-Thomas, J., Muelbert, M. M. C.,  Ottersen, G.,  Pritchard, H., and Schuur, E. A. G.: Polar Regions, in: IPCC Special Report on the Ocean and Cryosphere in a Changing Climate, edited by: Pörtner, H.-O.,  Roberts, C. C.,  Masson-Delmotte, V., Zhai, P.,  Tignor, M., Poloczanska, E., Mintenbeck, K., Alegría, A., Nicolai, M., Okem, A., Petzold, J., Rama, B., and Weyer, N. M., Cambridge, 203–320, <ext-link xlink:href="https://doi.org/10.1017/9781009157964.005" ext-link-type="DOI">10.1017/9781009157964.005</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx93"><label>Metropolis et al.(1953)Metropolis, Rosenbluth, Rosenbluth, Teller, and Teller</label><mixed-citation>Metropolis, N., Rosenbluth, A., Rosenbluth, M., Teller, A., and Teller, E.: Equation of State Calculations by Fast Computing Machines, J. Chem. Phys., 26, 1087–1092, <ext-link xlink:href="https://doi.org/10.1063/1.1699114" ext-link-type="DOI">10.1063/1.1699114</ext-link>, 1953.</mixed-citation></ref>
      <ref id="bib1.bibx94"><label>Morzfeld et al.(2017)Morzfeld, Hodyss, and Snyder</label><mixed-citation>Morzfeld, M., Hodyss, D., and Snyder, C.: What the collapse of the ensemble Kalman filter tells us about particle filters, Tellus A, 69, 1–14, <ext-link xlink:href="https://doi.org/10.1080/16000870.2017.1283809" ext-link-type="DOI">10.1080/16000870.2017.1283809</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx95"><label>Murphy(2023)</label><mixed-citation>Murphy, K.: Probabilistic Machine Learning: Advanced Topics, MIT, <uri>https://probml.github.io/book2</uri> (last access: 18 December 2023), 2023.</mixed-citation></ref>
      <ref id="bib1.bibx96"><label>Navari et al.(2016)Navari, Margulis, Bateni, Tedesco, Alexander, and Fettweis</label><mixed-citation>Navari, M., Margulis, S. A., Bateni, S. M., Tedesco, M., Alexander, P., and Fettweis, X.: Feasibility of improving a priori regional climate model estimates of Greenland ice sheet surface mass loss through assimilation of measured ice surface temperatures, The Cryosphere, 10, 103–120, <ext-link xlink:href="https://doi.org/10.5194/tc-10-103-2016" ext-link-type="DOI">10.5194/tc-10-103-2016</ext-link>, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx97"><label>Neal(2001)</label><mixed-citation>Neal, R.: Annealed importance sampling, Stat. Comput., 11, 125–139, <ext-link xlink:href="https://doi.org/10.1023/A:1008923215028" ext-link-type="DOI">10.1023/A:1008923215028</ext-link>, 2001.</mixed-citation></ref>
      <ref id="bib1.bibx98"><label>Neal(2011)</label><mixed-citation>Neal, R.: MCMC Using Hamiltonian Dynamics, in: Handbook of Markov Chain Monte Carlo, edited by: Brooks, S.,   Gelman, A., Jones, G., and Meng, X.-L., chap. 5,   CRC Press, 113–162, <ext-link xlink:href="https://doi.org/10.1201/b10905" ext-link-type="DOI">10.1201/b10905</ext-link>, 2011.</mixed-citation></ref>
      <ref id="bib1.bibx99"><label>Oberrauch et al.(2024)Oberrauch, Cluzet, Magnusson, and Jonas</label><mixed-citation>Oberrauch, M., Cluzet, B., Magnusson, J., and Jonas, T.: Improving Fully Distributed Snowpack Simulations by Mapping Perturbations of Meteorological Forcings Inferred From Particle Filter Assimilation of Snow Monitoring Data, Water Resour. Res., 60, e2023WR036994, <ext-link xlink:href="https://doi.org/10.1029/2023WR036994" ext-link-type="DOI">10.1029/2023WR036994</ext-link>, 2024.</mixed-citation></ref>
      <ref id="bib1.bibx100"><label>Owen and Zhou(2000)</label><mixed-citation>Owen, A. and Zhou, Y.: Safe and Effective Importance Sampling, J. Am. Stat. Assoc., 95, 135–143, <ext-link xlink:href="https://doi.org/10.1080/01621459.2000.10473909" ext-link-type="DOI">10.1080/01621459.2000.10473909</ext-link>, 2000.</mixed-citation></ref>
      <ref id="bib1.bibx101"><label>Piazzi et al.(2018)Piazzi, Thirel, Campo, and Gabellani</label><mixed-citation>Piazzi, G., Thirel, G., Campo, L., and Gabellani, S.: A particle filter scheme for multivariate data assimilation into a point-scale snowpack model in an Alpine environment, The Cryosphere, 12, 2287–2306, <ext-link xlink:href="https://doi.org/10.5194/tc-12-2287-2018" ext-link-type="DOI">10.5194/tc-12-2287-2018</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx102"><label>Pirk et al.(2022)Pirk, Aalstad, Westermann, Vatne, van Hove, Tallaksen, Cassiani, and Katul</label><mixed-citation>Pirk, N., Aalstad, K., Westermann, S., Vatne, A., van Hove, A., Tallaksen, L. M., Cassiani, M., and Katul, G.: Inferring surface energy fluxes using drone data assimilation in large eddy simulations, Atmos. Meas. Tech., 15, 7293–7314, <ext-link xlink:href="https://doi.org/10.5194/amt-15-7293-2022" ext-link-type="DOI">10.5194/amt-15-7293-2022</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx103"><label>Pirk et al.(2024)Pirk, Aalstad, Mannerfelt, Clayer, de Wit, Christiansen, Althuizen, Lee, and Westermann</label><mixed-citation>Pirk, N., Aalstad, K., Mannerfelt, E. S., Clayer, F., de Wit, H., Christiansen, C. T., Althuizen, I., Lee, H., and Westermann, S.: Disaggregating the Carbon Exchange of Degrading Permafrost Peatlands Using Bayesian Deep Learning, Geophys. Res. Lett., 51, e2024GL109283, <ext-link xlink:href="https://doi.org/10.1029/2024GL109283" ext-link-type="DOI">10.1029/2024GL109283</ext-link>,  2024.</mixed-citation></ref>
      <ref id="bib1.bibx104"><label>Rainforth et al.(2020)Rainforth, Golinski, Wood, and Zaidi</label><mixed-citation> Rainforth, T., Golinski, A., Wood, F., and Zaidi, S.: Target–Aware Bayesian Inference: How to Beat Optimal Conventional Estimators, JMLR, 21, 1–54, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx105"><label>Reich and Cotter(2015)</label><mixed-citation>Reich, S. and Cotter, C.: Probabilistic Forecasting and Bayesian Data Assimilation, Cambridge, <ext-link xlink:href="https://doi.org/10.1017/CBO9781107706804" ext-link-type="DOI">10.1017/CBO9781107706804</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx106"><label>Revuelto et al.(2017)Revuelto, Azorin-Molina, Alonso-González, Sanmiguel-Vallelado, Navarro-Serrano, Rico, and Ignacio López-Moreno</label><mixed-citation>Revuelto, J., Azorin-Molina, C., Alonso-González, E., Sanmiguel-Vallelado, A., Navarro-Serrano, F., Rico, I., and López-Moreno, J. I.: Meteorological and snow distribution data in the Izas Experimental Catchment (Spanish Pyrenees) from 2011 to 2017, Earth Syst. Sci. Data, 9, 993–1005, <ext-link xlink:href="https://doi.org/10.5194/essd-9-993-2017" ext-link-type="DOI">10.5194/essd-9-993-2017</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx107"><label>Revuelto et al.(2021)Revuelto, Alonso-Gonzalez, Vidaller-Gayan, Lacroix, Izagirre, Rodríguez-López, and López-Moreno</label><mixed-citation>Revuelto, J., Alonso-Gonzalez, E., Vidaller-Gayan, I., Lacroix, E., Izagirre, E., Rodríguez-López, G., and López-Moreno, J.: Intercomparison of UAV platforms for mapping snow depth distribution in complex alpine terrain, Cold Reg. Sci. Technol., 190, 103344, <ext-link xlink:href="https://doi.org/10.1016/j.coldregions.2021.103344" ext-link-type="DOI">10.1016/j.coldregions.2021.103344</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx108"><label>Riihelä et al.(2021)Riihelä, Bright, and Anttila</label><mixed-citation>Riihelä, A., Bright, R. M., and Anttila, K.: Recent strengthening of snow and ice albedo feedback driven by Antarctic sea-ice loss, Nat. Geosci., 14, 832–836, <ext-link xlink:href="https://doi.org/10.1038/s41561-021-00841-x" ext-link-type="DOI">10.1038/s41561-021-00841-x</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx109"><label>Robert(2007)</label><mixed-citation>Robert, C.: The Bayesian Choice, Springer, 2nd Edn., <ext-link xlink:href="https://doi.org/10.1007/0-387-71599-1" ext-link-type="DOI">10.1007/0-387-71599-1</ext-link>, 2007.</mixed-citation></ref>
      <ref id="bib1.bibx110"><label>Robert and Casella(2004)</label><mixed-citation>Robert, C. and Casella, G.: Monte Carlo Statistical Methods, Springer, <ext-link xlink:href="https://doi.org/10.1007/978-1-4757-4145-2" ext-link-type="DOI">10.1007/978-1-4757-4145-2</ext-link>, 2004.</mixed-citation></ref>
      <ref id="bib1.bibx111"><label>Rounce et al.(2023)Rounce, Hock, Maussion, Hugonnet, Kochtitzky, Huss, Berthier, Brinkerhoff, Compagno, Copland, Farinotti, Menounos, and McNabb</label><mixed-citation>Rounce, D. R., Hock, R., Maussion, F., Hugonnet, R., Kochtitzky, W., Huss, M., Berthier, E., Brinkerhoff, D., Compagno, L., Copland, L., Farinotti, D., Menounos, B., and McNabb, R. W.: Global glacier change in the 21st century: Every increase in temperature matters, Science, 379, 78–83, <ext-link xlink:href="https://doi.org/10.1126/science.abo1324" ext-link-type="DOI">10.1126/science.abo1324</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx112"><label>Sanz-Alonso et al.(2023)Sanz-Alonso, Stuart, and Taeb</label><mixed-citation>Sanz-Alonso, D., Stuart, A. M., and Taeb, A.: Inverse problems and data assimilation, vol. 107, Cambridge, <ext-link xlink:href="https://doi.org/10.1017/9781009414319" ext-link-type="DOI">10.1017/9781009414319</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx113"><label>Särkkä and Svensson(2023)</label><mixed-citation>Särkkä, S. and Svensson, L.: Bayesian Filtering and Smoothing, Cambridge, 2nd Edn., <ext-link xlink:href="https://doi.org/10.1017/9781108917407" ext-link-type="DOI">10.1017/9781108917407</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx114"><label>Smith and Gelfand(1992)</label><mixed-citation>Smith, A. F. M. and Gelfand, A. E.: Bayesian Statistics without Tears: A Sampling–Resampling Perspective,   Am. Stat., 46, 84–88, <ext-link xlink:href="https://doi.org/10.1080/00031305.1992.10475856" ext-link-type="DOI">10.1080/00031305.1992.10475856</ext-link>, 1992.</mixed-citation></ref>
      <ref id="bib1.bibx115"><label>Smith et al.(2010)Smith, Sharma, Marshall, Mehrotra, and Sisson</label><mixed-citation>Smith, T., Sharma, A., Marshall, L., Mehrotra, R., and Sisson, S.: Development of a formal likelihood function for improved Bayesian inference of ephemeral catchments, Water Resour. Res., 46, <ext-link xlink:href="https://doi.org/10.1029/2010WR009514" ext-link-type="DOI">10.1029/2010WR009514</ext-link>, 2010.</mixed-citation></ref>
      <ref id="bib1.bibx116"><label>Smyth et al.(2019)Smyth, Raleigh, and Small</label><mixed-citation>Smyth, E. J., Raleigh, M. S., and Small, E. E.: Particle Filter Data Assimilation of Monthly Snow Depth Observations Improves Estimation of Snow Density and SWE, Water Resour. Res., 55, 1296–1311, <ext-link xlink:href="https://doi.org/10.1029/2018WR023400" ext-link-type="DOI">10.1029/2018WR023400</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx117"><label>Snyder et al.(2008)Snyder, Bengtsson, Bickel, and Anderson</label><mixed-citation>Snyder, C., Bengtsson, T., Bickel, P., and Anderson, J.: Obstacles to High-Dimensional Particle Filtering, Mon. Weather Rev., 136, 4629–4640, <ext-link xlink:href="https://doi.org/10.1175/2008MWR2529.1" ext-link-type="DOI">10.1175/2008MWR2529.1</ext-link>, 2008.</mixed-citation></ref>
      <ref id="bib1.bibx118"><label>Sörensen(2015)</label><mixed-citation>Sörensen, K.: Metaheuristics – the metaphor exposed, Int. Trans. Oper. Res., 22, 3–18, <ext-link xlink:href="https://doi.org/10.1111/itor.12001" ext-link-type="DOI">10.1111/itor.12001</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx119"><label>Stordal and Elsheikh(2015)</label><mixed-citation>Stordal, A. S. and Elsheikh, A. H.: Iterative ensemble smoothers in the annealed importance sampling framework, Adv. Water Resour., 86, 231–239, <ext-link xlink:href="https://doi.org/10.1016/j.advwatres.2015.09.030" ext-link-type="DOI">10.1016/j.advwatres.2015.09.030</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx120"><label>Sun et al.(2025)Sun, Fang, Margulis, Mortimer, Mudryk, and Derksen</label><mixed-citation>Sun, H., Fang, Y., Margulis, S. A., Mortimer, C., Mudryk, L., and Derksen, C.: Evaluation of the Snow Climate Change Initiative (Snow CCI) snow-covered area product within a mountain snow water equivalent reanalysis, The Cryosphere, 19, 2017–2036, <ext-link xlink:href="https://doi.org/10.5194/tc-19-2017-2025" ext-link-type="DOI">10.5194/tc-19-2017-2025</ext-link>, 2025.</mixed-citation></ref>
      <ref id="bib1.bibx121"><label>Tang et al.(2023)Tang, Frye, Gelfand, and Silander</label><mixed-citation> Tang, B., Frye, H., Gelfand, A., and Silander, J. A.: Zero-Inflated Beta Distribution Regression Modeling, J. Agr., Biol. Environ. Stat., 28, 117–137, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx122"><label>van Hove et al.(2025)van Hove, Aalstad, Lind, Arndt, Odongo, Ceriani, Fava, Hulth, and Pirk</label><mixed-citation>van Hove, A., Aalstad, K., Lind, V., Arndt, C., Odongo, V., Ceriani, R., Fava, F., Hulth, J., and Pirk, N.: Inferring methane emissions from African livestock by fusing drone, tower, and satellite data, Biogeosciences, 22, 4163–4186, <ext-link xlink:href="https://doi.org/10.5194/bg-22-4163-2025" ext-link-type="DOI">10.5194/bg-22-4163-2025</ext-link>, 2025.</mixed-citation></ref>
      <ref id="bib1.bibx123"><label>van Hove et al.(2026)van Hove, Aalstad, and Pirk</label><mixed-citation>van Hove, A., Aalstad, K., and Pirk, N.: Actively inferring methane sources with drones, Environ. Data Sci., 5, e2, <ext-link xlink:href="https://doi.org/10.1017/eds.2026.10029" ext-link-type="DOI">10.1017/eds.2026.10029</ext-link>, 2026.</mixed-citation></ref>
      <ref id="bib1.bibx124"><label>van Leeuwen(2009)</label><mixed-citation>van Leeuwen, P. J.: Particle Filtering in Geophysical Systems, Mon. Weather Rev., 137, 4089–4114, <ext-link xlink:href="https://doi.org/10.1175/2009MWR2835.1" ext-link-type="DOI">10.1175/2009MWR2835.1</ext-link>, 2009.</mixed-citation></ref>
      <ref id="bib1.bibx125"><label>van Leeuwen and Evensen(1996)</label><mixed-citation>van Leeuwen, P. J. and Evensen, G.: Data Assimilation and Inverse Methods in Terms of a Probabilistic Formulation, Mon. Weather Rev., 124, 2898–2913, <ext-link xlink:href="https://doi.org/10.1175/1520-0493(1996)124&lt;2898:DAAIMI&gt;2.0.CO;2" ext-link-type="DOI">10.1175/1520-0493(1996)124&lt;2898:DAAIMI&gt;2.0.CO;2</ext-link>, 1996.</mixed-citation></ref>
      <ref id="bib1.bibx126"><label>van Leeuwen et al.(2019)van Leeuwen, Künsch, Nerger, Potthast, and Reich</label><mixed-citation>van Leeuwen, P. J., Künsch, H. R., Nerger, L., Potthast, R., and Reich, S.: Particle filters for high-dimensional geoscience applications: A review, Q. J. Roy. Meteor. Soc., 145, 2335–2365, <ext-link xlink:href="https://doi.org/10.1002/qj.3551" ext-link-type="DOI">10.1002/qj.3551</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx127"><label>Vihola(2012)</label><mixed-citation>Vihola, M.: Robust adaptive Metropolis algorithm with coerced acceptance rate, Stat. Comput., 22, 997–1008, <ext-link xlink:href="https://doi.org/10.1007/s11222-011-9269-5" ext-link-type="DOI">10.1007/s11222-011-9269-5</ext-link>, 2012.</mixed-citation></ref>
      <ref id="bib1.bibx128"><label>Virtanen et al.(2020)</label><mixed-citation>Virtanen, P., Gommers, R., Oliphant, T. E., Haberland, M., Reddy, T., Cournapeau, D., Burovski, E., Peterson, P., Weckesser, W., Bright, J., van der Walt, S. J., Brett, M., Wilson, J., Millman, K. J., Mayorov, N., Nelson, A. R. J., Jones, E., Kern, R., Larson, E., Carey, C. J., Polat, I., Feng, Y., Moore, E. W., VanderPlas, J., Laxalde, D., Perktold, J., Cimrman, R., Henriksen, I., Quintero, E. A., Harris, C. R., Archibald, A. M., Ribeiro, A. H., Pedregosa, F., van Mulbregt, P., Vijaykumar, A., Bardelli, A. P., Rothberg, A., Hilboll, A., Kloeckner, A., Scopatz, A., Lee, A., Rokem, A., Woods, C. N., Fulton, C., Masson, C., Häggström, C., Fitzgerald, C., Nicholson, D. A., Hagen, D. R., Pasechnik, D. V., Olivetti, E., Martin, E., Wieser, E., Silva, F., Lenders, F., Wilhelm, F., Young, G., Price, G. A., Ingold, G.-L., Allen, G. E., Lee, G. R., Audren, H., Probst, I., Dietrich, J. P., Silterra, J., Webber, J. T., Slavič, J., Nothman, J., Buchner, J., Kulick, J., Schönberger, J. L., de Miranda Cardoso, J. V., Reimer, J., Harrington, J., Rodríguez, J. L. C., Nunez-Iglesias, J., Kuczynski, J., Tritz, K., Thoma, M., Newville, M., Kümmerer, M., Bolingbroke, M., Tartre, M., Pak, M., Smith, N. J., Nowaczyk, N., Shebanov, N., Pavlyk, O., Brodtkorb, P. A., Lee, P., McGibbon, R. T., Feldbauer, R., Lewis, S., Tygier, S., Sievert, S., Vigna, S., Peterson, S., More, S., Pudlik, T., Oshima, T., Pingel, T. J., Robitaille, T. P., Spura, T., Jones, T. R., Cera, T., Leslie, T., Zito, T., Krauss, T., Upadhyay, U., Halchenko, Y. O., and Vázquez-Baeza, Y.: SciPy 1.0: fundamental algorithms for scientific computing in Python, Nat. Meth., 17, 261–272, <ext-link xlink:href="https://doi.org/10.1038/s41592-019-0686-2" ext-link-type="DOI">10.1038/s41592-019-0686-2</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx129"><label>Vrugt et al.(2003)Vrugt, Gupta, Bouten, and Sorooshian</label><mixed-citation>Vrugt, J. A., Gupta, H. V., Bouten, W., and Sorooshian, S.: A Shuffled Complex Evolution Metropolis algorithm for optimization and uncertainty assessment of hydrologic model parameters, Water Resour. Res., 39, <ext-link xlink:href="https://doi.org/10.1029/2002WR001642" ext-link-type="DOI">10.1029/2002WR001642</ext-link>, 2003.</mixed-citation></ref>
      <ref id="bib1.bibx130"><label>Vrugt et al.(2008)Vrugt, ter Braak, Clark, Hyman, and Robinson</label><mixed-citation>Vrugt, J. A., ter Braak, C. J. F., Clark, M. P., Hyman, J. M., and Robinson, B. A.: Treatment of input uncertainty in hydrologic modeling: Doing hydrology backward with Markov chain Monte Carlo simulation, Water Resour. Res., 44, W00B09, <ext-link xlink:href="https://doi.org/10.1029/2007WR006720" ext-link-type="DOI">10.1029/2007WR006720</ext-link>, 2008.</mixed-citation></ref>
      <ref id="bib1.bibx131"><label>Westermann et al.(2023)Westermann, Ingeman-Nielsen, Scheer, Aalstad, Aga, Chaudhary, Etzelmüller, Filhol, Kääb, Renette, Schmidt, Schuler, Zweigel, Martin, Morard, Ben-Asher, Angelopoulos, Boike, Groenke, Miesner, Nitzbon, Overduin, Stuenzi, and Langer</label><mixed-citation>Westermann, S., Ingeman-Nielsen, T., Scheer, J., Aalstad, K., Aga, J., Chaudhary, N., Etzelmüller, B., Filhol, S., Kääb, A., Renette, C., Schmidt, L. S., Schuler, T. V., Zweigel, R. B., Martin, L., Morard, S., Ben-Asher, M., Angelopoulos, M., Boike, J., Groenke, B., Miesner, F., Nitzbon, J., Overduin, P., Stuenzi, S. M., and Langer, M.: The CryoGrid community model (version 1.0) – a multi-physics toolbox for climate-driven simulations in the terrestrial cryosphere, Geosci. Model Dev., 16, 2607–2647, <ext-link xlink:href="https://doi.org/10.5194/gmd-16-2607-2023" ext-link-type="DOI">10.5194/gmd-16-2607-2023</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx132"><label>Wikle and Berliner(2007)</label><mixed-citation>Wikle, C. K. and Berliner, L. M.: A Bayesian tutorial for data assimilation, Phys. D, 230, 1–16, <ext-link xlink:href="https://doi.org/10.1016/j.physd.2006.09.017" ext-link-type="DOI">10.1016/j.physd.2006.09.017</ext-link>, 2007. </mixed-citation></ref>
      <ref id="bib1.bibx133"><label>Willmes et al.(2025)Willmes, Aalstad, and Westermann</label><mixed-citation>Willmes, C., Aalstad, K., and Westermann, S.: Assimilating high-resolution satellite snow cover data in a permafrost model, EGUsphere [preprint], <ext-link xlink:href="https://doi.org/10.5194/egusphere-2025-3142" ext-link-type="DOI">10.5194/egusphere-2025-3142</ext-link>, 2025.</mixed-citation></ref>
      <ref id="bib1.bibx134"><label>Yang et al.(2026)Yang, Ultee, Aalstad, Debolsky, Hock, Schmitt, and Rounce</label><mixed-citation>Yang, R., Ultee, L., Aalstad, K., Debolskiy, M. V., Hock, R., Schmitt, P., Rounce, D., and Li, T.: Joint Bayesian Calibration of Frontal Ablation and Surface Mass Balance in Global Glacier Models, EGUsphere [preprint], <ext-link xlink:href="https://doi.org/10.5194/egusphere-2026-1081" ext-link-type="DOI">10.5194/egusphere-2026-1081</ext-link>, 2026.</mixed-citation></ref>
      <ref id="bib1.bibx135"><label>Zschenderlein et al.(2023)Zschenderlein, Luojus, Takala, Venäläinen, and Pulliainen</label><mixed-citation>Zschenderlein, L., Luojus, K., Takala, M., Venäläinen, P., and Pulliainen, J.: Evaluation of passive microwave dry snow detection algorithms and application to SWE retrieval during seasonal snow accumulation, Remote Sens. Environ., 288, 113476, <ext-link xlink:href="https://doi.org/10.1016/j.rse.2023.113476" ext-link-type="DOI">10.1016/j.rse.2023.113476</ext-link>, 2023.</mixed-citation></ref>

  </ref-list></back>
    <!--<article-title-html>Evolving beyond collapse: an adaptive particle batch smoother for cryospheric data assimilation</article-title-html>
<abstract-html/>
<ref-html id="bib1.bib1"><label>Aalstad et al.(2018)Aalstad, Westermann, Schuler, Boike, and
Bertino</label><mixed-citation>
      
Aalstad, K., Westermann, S., Schuler, T. V., Boike, J., and Bertino, L.: Ensemble-based assimilation of fractional snow-covered area satellite retrievals to estimate the snow distribution at Arctic sites, The Cryosphere, 12, 247–270, <a href="https://doi.org/10.5194/tc-12-247-2018" target="_blank">https://doi.org/10.5194/tc-12-247-2018</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib2"><label>Alonso-González(2022)</label><mixed-citation>
      
Alonso-González, E.: Inputs (forcing and observations) ready for use by “MuSA:
The Multiscale Snow Data Assimilation System (v1.0)”, Zenodo [code],
<a href="https://doi.org/10.5281/zenodo.7248635" target="_blank">https://doi.org/10.5281/zenodo.7248635</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib3"><label>Alonso-González and Aalstad(2025)</label><mixed-citation>
      
Alonso-González, E. and Aalstad, K.: MuSA: v2.3 AdaPBS submission, Zenodo [code],
<a href="https://doi.org/10.5281/zenodo.17292981" target="_blank">https://doi.org/10.5281/zenodo.17292981</a>,  2025.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib4"><label>Alonso González and Aalstad(2026)</label><mixed-citation>
      
Alonso González, E. and Aalstad, K.: Code and data to reproduce the results
and figures in: Evolving beyond collapse: An adaptive particle batch
smoother for cryospheric data assimilation, Zenodo [code and data set], <a href="https://doi.org/10.5281/zenodo.21244337" target="_blank">https://doi.org/10.5281/zenodo.21244337</a>,
2026.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib5"><label>Alonso-González et al.(2018)Alonso-González,
López-Moreno, Gascoin, García-Valdecasas Ojeda, Sanmiguel-Vallelado,
Navarro-Serrano, Revuelto, Ceballos, Esteban-Parra, and
Essery</label><mixed-citation>
      
Alonso-González, E., López-Moreno, J. I., Gascoin, S., García-Valdecasas Ojeda, M., Sanmiguel-Vallelado, A., Navarro-Serrano, F., Revuelto, J., Ceballos, A., Esteban-Parra, M. J., and Essery, R.: Daily gridded datasets of snow depth and snow water equivalent for the Iberian Peninsula from 1980 to 2014, Earth Syst. Sci. Data, 10, 303–315, <a href="https://doi.org/10.5194/essd-10-303-2018" target="_blank">https://doi.org/10.5194/essd-10-303-2018</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib6"><label>Alonso-González et al.(2021)Alonso-González, Gutmann,
Aalstad, Fayad, Bouchet, and Gascoin</label><mixed-citation>
      
Alonso-González, E., Gutmann, E., Aalstad, K., Fayad, A., Bouchet, M., and Gascoin, S.: Snowpack dynamics in the Lebanese mountains from quasi-dynamically downscaled ERA5 reanalysis updated by assimilating remotely sensed fractional snow-covered area, Hydrol. Earth Syst. Sci., 25, 4455–4471, <a href="https://doi.org/10.5194/hess-25-4455-2021" target="_blank">https://doi.org/10.5194/hess-25-4455-2021</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib7"><label>Alonso-González et al.(2022)Alonso-González, Aalstad, Baba,
Revuelto, López-Moreno, Fiddes, Essery, and Gascoin</label><mixed-citation>
      
Alonso-González, E., Aalstad, K., Baba, M. W., Revuelto, J., López-Moreno, J. I., Fiddes, J., Essery, R., and Gascoin, S.: The Multiple Snow Data Assimilation System (MuSA v1.0), Geosci. Model Dev., 15, 9127–9155, <a href="https://doi.org/10.5194/gmd-15-9127-2022" target="_blank">https://doi.org/10.5194/gmd-15-9127-2022</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib8"><label>Alonso-González et al.(2023)Alonso-González, Aalstad, Pirk,
Mazzolini, Treichler, Leclercq, Westermann, López-Moreno, and
Gascoin</label><mixed-citation>
      
Alonso-González, E., Aalstad, K., Pirk, N., Mazzolini, M., Treichler, D., Leclercq, P., Westermann, S., López-Moreno, J. I., and Gascoin, S.: Spatio-temporal information propagation using sparse observations in hyper-resolution ensemble-based snow data assimilation, Hydrol. Earth Syst. Sci., 27, 4637–4659, <a href="https://doi.org/10.5194/hess-27-4637-2023" target="_blank">https://doi.org/10.5194/hess-27-4637-2023</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib9"><label>Alonso-González et al.(2026)Alonso-González, Harpold, Lundquist,
Piske, Sourp, Aalstad, and Gascoin</label><mixed-citation>
      
Alonso-González, E., Harpold, A., Lundquist, J. D., Piske, C., Sourp, L., Aalstad, K., and Gascoin, S.: Ensemble-based data assimilation improves hyperresolution snowpack simulations in forests, The Cryosphere, 20, 209–225, <a href="https://doi.org/10.5194/tc-20-209-2026" target="_blank">https://doi.org/10.5194/tc-20-209-2026</a>, 2026.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib10"><label>Bannister(2017)</label><mixed-citation>
      
Bannister, R. N.: A review of operational methods of variational and
ensemble‐variational data assimilation, Q. J. Roy.
Meteor. Soc., 143, 607–633, <a href="https://doi.org/10.1002/qj.2982" target="_blank">https://doi.org/10.1002/qj.2982</a>, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib11"><label>Barnett et al.(2005)Barnett, Adam, and Lettenmaier</label><mixed-citation>
      
Barnett, T. P., Adam, J. C., and Lettenmaier, D. P.: Potential impacts of a
warming climate on water availability in snow-dominated regions, Nature,
438, 303–309, <a href="https://doi.org/10.1038/nature04141" target="_blank">https://doi.org/10.1038/nature04141</a>, 2005.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib12"><label>Bertino et al.(2003)Bertino, Evensen, and Wackernagel</label><mixed-citation>
      
Bertino, L., Evensen, G., and Wackernagel, H.: Sequential Data Assimilation
Techniques in Oceanography, Int. Stat. Rev., 71, 223–241,
<a href="https://doi.org/10.1111/j.1751-5823.2003.tb00194.x" target="_blank">https://doi.org/10.1111/j.1751-5823.2003.tb00194.x</a>, 2003.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib13"><label>Beven(2006)</label><mixed-citation>
      
Beven, K.: A manifesto for the equifinality thesis, J. Hydrol.,
320, 18–36, <a href="https://doi.org/10.1016/j.jhydrol.2005.07.007" target="_blank">https://doi.org/10.1016/j.jhydrol.2005.07.007</a>, 2006.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib14"><label>Beven and Binley(1992)</label><mixed-citation>
      
Beven, K. and Binley, A.: The future of distributed models: Model calibration
and uncertainty prediction, Hydrol. Process., 6, 279–298,
<a href="https://doi.org/10.1002/hyp.3360060305" target="_blank">https://doi.org/10.1002/hyp.3360060305</a>, 1992.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib15"><label>Bugallo et al.(2017)Bugallo, Elvira, Martino, Luengo, Miguez, and
Djuric</label><mixed-citation>
      
Bugallo, M. F., Elvira, V., Martino, L., Luengo, D., Miguez, J., and Djuric,
P. M.: Adaptive Importance Sampling: The past, the present, and the future,
IEEE Signal Process. Mag., 34, 60–79, <a href="https://doi.org/10.1109/MSP.2017.2699226" target="_blank">https://doi.org/10.1109/MSP.2017.2699226</a>,
2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib16"><label>Campelo and Aranha(2023)</label><mixed-citation>
      
Campelo, F. and Aranha, C.: Lessons from the Evolutionary Computation
Bestiary, Artificial Life, 29, 421–432, <a href="https://doi.org/10.1162/artl_a_00402" target="_blank">https://doi.org/10.1162/artl_a_00402</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib17"><label>Cao et al.(2025)Cao, Aalstad, Schmidt, Westermann, and
Schuler</label><mixed-citation>
      
Cao, W., Aalstad, K., Schmidt, L., Westermann, S., and Schuler, T.: Bayesian
data assimilation on an Arctic glacier: learning from large ensemble twin
experiments, J. Glaciol., 71, e121, <a href="https://doi.org/10.1017/jog.2025.10101" target="_blank">https://doi.org/10.1017/jog.2025.10101</a>,
2025.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib18"><label>Carrassi et al.(2018)Carrassi, Bocquet, Bertino, and
Evensen</label><mixed-citation>
      
Carrassi, A., Bocquet, M., Bertino, L., and Evensen, G.: Data assimilation in
the geosciences: An overview of methods, issues, and perspectives, WIREs
Clim. Change, 9, e535, <a href="https://doi.org/10.1002/wcc.535" target="_blank">https://doi.org/10.1002/wcc.535</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib19"><label>Charrois et al.(2016)Charrois, Cosme, Dumont, Lafaysse, Morin,
Libois, and Picard</label><mixed-citation>
      
Charrois, L., Cosme, E., Dumont, M., Lafaysse, M., Morin, S., Libois, Q., and Picard, G.: On the assimilation of optical reflectances and snow depth observations into a detailed snowpack model, The Cryosphere, 10, 1021–1038, <a href="https://doi.org/10.5194/tc-10-1021-2016" target="_blank">https://doi.org/10.5194/tc-10-1021-2016</a>, 2016.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib20"><label>Chopin(2002)</label><mixed-citation>
      
Chopin, N.: A sequential particle filter method for static models,
Biometrika, 89, 539–552, <a href="https://doi.org/10.1093/biomet/89.3.539" target="_blank">https://doi.org/10.1093/biomet/89.3.539</a>, 2002.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib21"><label>Chopin and Papaspiliopoulos(2020)</label><mixed-citation>
      
Chopin, N. and Papaspiliopoulos, O.: An Introduction to Sequential Monte
Carlo, Springer, <a href="https://doi.org/10.1007/978-3-030-47845-2" target="_blank">https://doi.org/10.1007/978-3-030-47845-2</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib22"><label>Clark et al.(2006)Clark, Slater, Barrett, Hay, McCabe, Rajagopalan,
and Leavesley</label><mixed-citation>
      
Clark, M. P., Slater, A. G., Barrett, A. P., Hay, L. E., McCabe, G. J.,
Rajagopalan, B., and Leavesley, G. H.: Assimilation of snow covered area
information into hydrologic and land-surface models, Adv. Water
Resour., 29, 1209–1221, <a href="https://doi.org/10.1016/j.advwatres.2005.10.001" target="_blank">https://doi.org/10.1016/j.advwatres.2005.10.001</a>, 2006.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib23"><label>Cleary et al.(2021)Cleary, Garbuno-Inigo, Lan, Schneider, and
Stuart</label><mixed-citation>
      
Cleary, E., Garbuno-Inigo, A., Lan, S., Schneider, T., and Stuart, A.:
Calibrate, emulate, sample, J. Comput. Phys., 424, 109716,
<a href="https://doi.org/10.1016/j.jcp.2020.109716" target="_blank">https://doi.org/10.1016/j.jcp.2020.109716</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib24"><label>Cluzet et al.(2021)Cluzet, Lafaysse, Cosme, Albergel, Meunier, and
Dumont</label><mixed-citation>
      
Cluzet, B., Lafaysse, M., Cosme, E., Albergel, C., Meunier, L.-F., and Dumont, M.: CrocO_v1.0: a particle filter to assimilate snowpack observations in a spatialised framework, Geosci. Model Dev., 14, 1595–1614, <a href="https://doi.org/10.5194/gmd-14-1595-2021" target="_blank">https://doi.org/10.5194/gmd-14-1595-2021</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib25"><label>Cluzet et al.(2024)Cluzet, Magnusson, Quéno, Mazzotti, Mott, and
Jonas</label><mixed-citation>
      
Cluzet, B., Magnusson, J., Quéno, L., Mazzotti, G., Mott, R., and Jonas, T.: Exploring how Sentinel-1 wet-snow maps can inform fully distributed physically based snowpack models, The Cryosphere, 18, 5753–5767, <a href="https://doi.org/10.5194/tc-18-5753-2024" target="_blank">https://doi.org/10.5194/tc-18-5753-2024</a>, 2024.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib26"><label>Cornuet et al.(2012)Cornuet, Marin, Antonietta, and
Robert</label><mixed-citation>
      
Cornuet, J.-M., Marin, J.-M., Antonietta, M., and Robert, C.: Adaptive
Multiple Importance Sampling, Scand. J. Stat., 39,
798–812, <a href="https://doi.org/10.1111/j.1467-9469.2011.00756.x" target="_blank">https://doi.org/10.1111/j.1467-9469.2011.00756.x</a>, 2012.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib27"><label>Cortés and Margulis(2017)</label><mixed-citation>
      
Cortés, G. and Margulis, S.: Impacts of El Niño and La Niña on
interannual snow accumulation in the Andes: Results from a high-resolution 31
year reanalysis: El Niño Effects on Andes Snow, Geophys. Res.
Lett., 44, 6859–6867, <a href="https://doi.org/10.1002/2017GL073826" target="_blank">https://doi.org/10.1002/2017GL073826</a>, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib28"><label>De Lannoy et al.(2012)De Lannoy, Reichle, Arsenault, Houser, Kumar,
Verhoest, and Pauwels</label><mixed-citation>
      
De Lannoy, G. J. M., Reichle, R. H., Arsenault, K. R., Houser, P. R., Kumar,
S., Verhoest, N. E. C., and Pauwels, V. R. N.: Multiscale assimilation of
Advanced Microwave Scanning Radiometer-EOS snow water equivalent and Moderate
Resolution Imaging Spectroradiometer snow cover fraction observations in
northern Colorado, Water Resour. Res., 48, W01522,
<a href="https://doi.org/10.1029/2011WR010588" target="_blank">https://doi.org/10.1029/2011WR010588</a>, 2012.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib29"><label>De Lannoy et al.(2024)De Lannoy, Bechtold, Busschaert, Heyvaert,
Modanesi, Dunmire, Lievens, Getirana, and Massari</label><mixed-citation>
      
De Lannoy, G. J. M., Bechtold, M., Busschaert, L., Heyvaert, Z., Modanesi, S.,
Dunmire, D., Lievens, H., Getirana, A., and Massari, C.: Contributions of
Irrigation Modeling, Soil Moisture and Snow Data Assimilation to
High-Resolution Water Budget Estimates Over the Po Basin: Progress Towards
Digital Replicas, J. Adv. Model. Earth Sy., 16,
e2024MS004433, <a href="https://doi.org/10.1029/2024MS004433" target="_blank">https://doi.org/10.1029/2024MS004433</a>, 2024.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib30"><label>Del Moral(2004)</label><mixed-citation>
      
Del Moral, P.: Feynman-Kac Formulae, Springer,
<a href="https://doi.org/10.1007/978-1-4684-9393-1" target="_blank">https://doi.org/10.1007/978-1-4684-9393-1</a>, 2004.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib31"><label>Doucet et al.(2000)Doucet, Godsill, and Andrieu</label><mixed-citation>
      
Doucet, A., Godsill, S., and Andrieu, C.: On sequential Monte Carlo sampling
methods for Bayesian filtering, Stat. Comput., 10, 197–208,
<a href="https://doi.org/10.1023/A:1008935410038" target="_blank">https://doi.org/10.1023/A:1008935410038</a>, 2000.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib32"><label>Dozier et al.(2016)Dozier, Bair, and Davis</label><mixed-citation>
      
Dozier, J., Bair, E. H., and Davis, R. E.: Estimating the spatial distribution
of snow water equivalent in the world’s mountains, Wires Water, 3, 461–474, <a href="https://doi.org/10.1002/wat2.1140" target="_blank">https://doi.org/10.1002/wat2.1140</a>, 2016.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib33"><label>Duan et al.(2023)Duan, Li, and Xu</label><mixed-citation>
      
Duan, J.-C., Li, S., and Xu, Y.: Sequential Monte Carlo optimization and
statistical inference, WIREs Comput. Stat., 15, e1598,
<a href="https://doi.org/10.1002/wics.1598" target="_blank">https://doi.org/10.1002/wics.1598</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib34"><label>Durand et al.(2008)Durand, Molotch, and Margulis</label><mixed-citation>
      
Durand, M., Molotch, N. P., and Margulis, S. A.: A Bayesian approach to snow
water equivalent reconstruction, J. Geophys. Res.-Atmos., 113, D20117, <a href="https://doi.org/10.1029/2008JD009894" target="_blank">https://doi.org/10.1029/2008JD009894</a>, 2008.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib35"><label>Elias Chereque et al.(2024)Elias Chereque, Kushner, Mudryk, Derksen,
and Mortimer</label><mixed-citation>
      
Elias Chereque, A., Kushner, P. J., Mudryk, L., Derksen, C., and Mortimer, C.: A simple snow temperature index model exposes discrepancies between reanalysis snow water equivalent products, The Cryosphere, 18, 4955–4969, <a href="https://doi.org/10.5194/tc-18-4955-2024" target="_blank">https://doi.org/10.5194/tc-18-4955-2024</a>, 2024.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib36"><label>Elvira et al.(2019)Elvira, Martino, Bugallo, and Djuric</label><mixed-citation>
      
Elvira, V., Martino, L., Bugallo, M. F., and Djuric, P. M.: Elucidating the
Auxiliary Particle Filter via Multiple Importance Sampling [Lecture Notes],
IEEE Signal Process. Mag., 36, 145–152,
<a href="https://doi.org/10.1109/MSP.2019.2938026" target="_blank">https://doi.org/10.1109/MSP.2019.2938026</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib37"><label>Elvira et al.(2022)Elvira, Martino, and Robert</label><mixed-citation>
      
Elvira, V., Martino, L., and Robert, C.: Rethinking the Effective Sample
Size, Int. Stat. Rev., 90, 525–550,
<a href="https://doi.org/10.1111/insr.12500" target="_blank">https://doi.org/10.1111/insr.12500</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib38"><label>Emerick and Reynolds(2011)</label><mixed-citation>
      
Emerick, A. A. and Reynolds, A. C.: Combining the Ensemble Kalman Filter with
Markov Chain Monte Carlo for Improved History Matching and Uncertainty
Characterization,  SPE Reservoir Simulation Symposium, The Woodlands, Texas, USA, February 2011,  SPE-141336-MS, <a href="https://doi.org/10.2118/141336-MS" target="_blank">https://doi.org/10.2118/141336-MS</a>, 2011.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib39"><label>Emerick and Reynolds(2013)</label><mixed-citation>
      
Emerick, A. A. and Reynolds, A. C.: Ensemble smoother with multiple data
assimilation, Comput. Geosci., 55, 3–15,
<a href="https://doi.org/10.1016/j.cageo.2012.03.011" target="_blank">https://doi.org/10.1016/j.cageo.2012.03.011</a>, 2013.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib40"><label>Essery(2015)</label><mixed-citation>
      
Essery, R.: A factorial snowpack model (FSM 1.0), Geosci. Model Dev., 8, 3867–3876, <a href="https://doi.org/10.5194/gmd-8-3867-2015" target="_blank">https://doi.org/10.5194/gmd-8-3867-2015</a>, 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib41"><label>Essery et al.(2025)Essery, Mazzotti, Barr, Jonas, Quaife, and
Rutter</label><mixed-citation>
      
Essery, R., Mazzotti, G., Barr, S., Jonas, T., Quaife, T., and Rutter, N.: A Flexible Snow Model (FSM 2.1.1) including a forest canopy, Geosci. Model Dev., 18, 3583–3605, <a href="https://doi.org/10.5194/gmd-18-3583-2025" target="_blank">https://doi.org/10.5194/gmd-18-3583-2025</a>, 2025.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib42"><label>Euskirchen et al.(2013)Euskirchen, Goodstein, and
Huntington</label><mixed-citation>
      
Euskirchen, E. S., Goodstein, E. S., and Huntington, H. P.: An estimated cost
of lost climate regulation services caused by thawing of the Arctic
cryosphere, Ecol. Appl., 23, 1869–1880, <a href="https://doi.org/10.1890/11-0858.1" target="_blank">https://doi.org/10.1890/11-0858.1</a>,
2013.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib43"><label>Evensen(1994)</label><mixed-citation>
      
Evensen, G.: Sequential data assimilation with a nonlinear quasi-geostrophic
model using Monte Carlo methods to forecast error statistics, J.
Geophys. Res., 99, 10143, <a href="https://doi.org/10.1029/94JC00572" target="_blank">https://doi.org/10.1029/94JC00572</a>, 1994.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib44"><label>Evensen et al.(2022)Evensen, Vossepoel, and van
Leeuwen</label><mixed-citation>
      
Evensen, G., Vossepoel, F. C., and van Leeuwen, P. J.: Data Assimilation
Fundamentals, Springer, <a href="https://doi.org/10.1007/978-3-030-96709-3" target="_blank">https://doi.org/10.1007/978-3-030-96709-3</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib45"><label>Farchi and Bocquet(2018)</label><mixed-citation>
      
Farchi, A. and Bocquet, M.: Review article: Comparison of local particle filters and new implementations, Nonlin. Processes Geophys., 25, 765–807, <a href="https://doi.org/10.5194/npg-25-765-2018" target="_blank">https://doi.org/10.5194/npg-25-765-2018</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib46"><label>Fiddes et al.(2019)Fiddes, Aalstad, and Westermann</label><mixed-citation>
      
Fiddes, J., Aalstad, K., and Westermann, S.: Hyper-resolution ensemble-based snow reanalysis in mountain regions using clustering, Hydrol. Earth Syst. Sci., 23, 4717–4736, <a href="https://doi.org/10.5194/hess-23-4717-2019" target="_blank">https://doi.org/10.5194/hess-23-4717-2019</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib47"><label>Garbuno-Inigo et al.(2020)Garbuno-Inigo, Hoffmann, Li, and
Stuart</label><mixed-citation>
      
Garbuno-Inigo, A., Hoffmann, F., Li, W., and Stuart, A.: Interacting Langevin
Diffusions: Gradient Structure and Ensemble Kalman Sampler, SIAM J.
Appl. Dynam. Syst., 19, 412–441, <a href="https://doi.org/10.1137/19M1251655" target="_blank">https://doi.org/10.1137/19M1251655</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib48"><label>Gascoin(2024)</label><mixed-citation>
      
Gascoin, S.: A call for an accurate presentation of glaciers as water
resources, WIREs Water, 11, e1705, <a href="https://doi.org/10.1002/wat2.1705" target="_blank">https://doi.org/10.1002/wat2.1705</a>, 2024.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib49"><label>Gascoin et al.(2024)Gascoin, Luojus, Nagler, Lievens, Masiokas,
Jonas, Zheng, and De Rosnay</label><mixed-citation>
      
Gascoin, S., Luojus, K., Nagler, T., Lievens, H., Masiokas, M., Jonas, T.,
Zheng, Z., and De Rosnay, P.: Remote sensing of mountain snow from space:
status and recommendations, Front. Earth Sci., 12, 138323, <a href="https://doi.org/10.3389/feart.2024.1381323" target="_blank">https://doi.org/10.3389/feart.2024.1381323</a>, 2024.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib50"><label>Gelman et al.(2013)Gelman, Carlin, Stern, Dunson, Vehtari, and
Rubin</label><mixed-citation>
      
Gelman, A., Carlin, J., Stern, H., Dunson, D., Vehtari, A., and Rubin, D.:
Bayesian Data Analysis, CRC Press, 3rd Edn., 675 pp., <a href="https://doi.org/10.1201/b16018" target="_blank">https://doi.org/10.1201/b16018</a>, 2013.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib51"><label>Gilks and Berzuini(2001)</label><mixed-citation>
      
Gilks, W. R. and Berzuini, C.: Following a moving target-Monte Carlo inference
for dynamic Bayesian models, J. Roy. Stat. Soc.
Ser. B, 63, 127–146,
<a href="https://doi.org/10.1111/1467-9868.00280" target="_blank">https://doi.org/10.1111/1467-9868.00280</a>, 2001.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib52"><label>Girotto et al.(2014)Girotto, Margulis, and Durand</label><mixed-citation>
      
Girotto, M., Margulis, S., and Durand, M.: Probabilistic SWE reanalysis as a
generalization of deterministic SWE reconstruction techniques, Hydrol.
Process., 28, 3875–3895, <a href="https://doi.org/10.1002/hyp.9887" target="_blank">https://doi.org/10.1002/hyp.9887</a>, 2014.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib53"><label>Girotto et al.(2020)Girotto, Musselman, and Essery</label><mixed-citation>
      
Girotto, M., Musselman, K. N., and Essery, R. L. H.: Data Assimilation
Improves Estimates of Climate-Sensitive Seasonal Snow, Curr. Clim.
Change Rep., 6, 81–94, <a href="https://doi.org/10.1007/s40641-020-00159-7" target="_blank">https://doi.org/10.1007/s40641-020-00159-7</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib54"><label>Girotto et al.(2024)Girotto, Formetta, Azimi, Bachand, Cowherd, De
Lannoy, Lievens, Modanesi, Raleigh, Rigon, and Massari</label><mixed-citation>
      
Girotto, M., Formetta, G., Azimi, S., Bachand, C., Cowherd, M., De Lannoy,
G., Lievens, H., Modanesi, S., Raleigh, M. S., Rigon, R., and Massari, C.:
Identifying snowfall elevation patterns by assimilating satellite-based snow
depth retrievals, Sci. Total Environ., 906, 167312,
<a href="https://doi.org/10.1016/j.scitotenv.2023.167312" target="_blank">https://doi.org/10.1016/j.scitotenv.2023.167312</a>, 2024.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib55"><label>Gneiting and Raftery(2007)</label><mixed-citation>
      
Gneiting, T. and Raftery, A. E.: Strictly Proper Scoring Rules, Prediction,
and Estimation, J. Am. Stat. Assoc., 102,
359–378, <a href="https://doi.org/10.1198/016214506000001437" target="_blank">https://doi.org/10.1198/016214506000001437</a>, 2007.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib56"><label>Gneiting et al.(2005)Gneiting, Raftery, Westveld III, and
Goldman</label><mixed-citation>
      
Gneiting, T., Raftery, A. E., Westveld III, A. H., and Goldman, T.: Calibrated
Probabilistic Forecasting Using Ensemble Model Output Statistics and Minimum
CRPS Estimation, Mon. Weather Rev., 133, 1098–1118, 2005.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib57"><label>Gordon et al.(1993)Gordon, Salmond, and Smith</label><mixed-citation>
      
Gordon, N. J., Salmond, D. J., and Smith, A. F. M.: Novel approach to
nonlinear/non-Gaussian Bayesian state estimation, IEE Proc.-F, 140, 107, <a href="https://doi.org/10.1049/ip-f-2.1993.0015" target="_blank">https://doi.org/10.1049/ip-f-2.1993.0015</a>, 1993.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib58"><label>Gottlieb and Mankin(2024)</label><mixed-citation>
      
Gottlieb, A. R. and Mankin, J. S.: Evidence of human influence on Northern
Hemisphere snow loss, Nature, 625, 293–300,
<a href="https://doi.org/10.1038/s41586-023-06794-y" target="_blank">https://doi.org/10.1038/s41586-023-06794-y</a>, 2024.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib59"><label>Groenke et al.(2023)Groenke, Langer, Nitzbon, Westermann, Gallego,
and Boike</label><mixed-citation>
      
Groenke, B., Langer, M., Nitzbon, J., Westermann, S., Gallego, G., and Boike, J.: Investigating the thermal state of permafrost with Bayesian inverse modeling of heat transfer, The Cryosphere, 17, 3505–3533, <a href="https://doi.org/10.5194/tc-17-3505-2023" target="_blank">https://doi.org/10.5194/tc-17-3505-2023</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib60"><label>Günther et al.(2019)Günther, Marke, Essery, and
Strasser</label><mixed-citation>
      
Günther, D., Marke, T., Essery, R., and Strasser, U.: Uncertainties in
Snowpack Simulations – Assessing the Impact of Model Structure, Parameter
Choice, and Forcing Data Error on Point-Scale Energy Balance Snow Model
Performance, Water Resour. Res., 55, 2779–2800,
<a href="https://doi.org/10.1029/2018WR023403" target="_blank">https://doi.org/10.1029/2018WR023403</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib61"><label>Hammersley and Morton(1954)</label><mixed-citation>
      
Hammersley, J. and Morton, K.: Poor Man's Monte Carlo, J. Roy.
Stat. Soc. Ser. B, 16, 23–38,
<a href="https://doi.org/10.1111/j.2517-6161.1954.tb00145.x" target="_blank">https://doi.org/10.1111/j.2517-6161.1954.tb00145.x</a>, 1954.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib62"><label>Hastings(1970)</label><mixed-citation>
      
Hastings, W.: Monte Carlo Sampling Methods Using Markov Chains and Their
Applications, Biometrika, 57, <a href="https://doi.org/10.2307/2334940" target="_blank">https://doi.org/10.2307/2334940</a>, 1970.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib63"><label>Hennig et al.(2022)Hennig, Osborne, and Kersting</label><mixed-citation>
      
Hennig, P., Osborne, M., and Kersting, H.: Probabilistic Numerics, Cambridge,
<a href="https://doi.org/10.1017/9781316681411" target="_blank">https://doi.org/10.1017/9781316681411</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib64"><label>Hersbach et al.(2020)Hersbach, Bell, Berrisford, Hirahara,
Horányi, Muñoz-Sabater et al.</label><mixed-citation>
      
Hersbach, H., Bell, B., Berrisford, P., Hirahara, S., Horányi, A.,
Muñoz-Sabater, J., Nicolas, J., Peubey, C., Radu, R., Schepers, D.,
Simmons, A., Soci, C., Abdalla, S., Abellan, X., Balsamo, G., Bechtold, P., Biavati, G., Bidlot, J., Bonavita, M., De Chiara, G., Dahlgren, P., Dee, D., Diamantakis, M., Dragani, R., Flemming, J., Forbes, R., Fuentes, M., Geer, A., Haimberger, L., Healy, S., Hogan, R. J., Hólm, E., Janisková, M., Keeley, S., Laloyaux, P., Lopez, P., Lupu, C., Radnoti, G., de Rosnay, P., Rozum, I., Vamborg, F., Villaume, S., and Thépaut, J.-N.: The ERA5 global reanalysis, Q.  J.
Roy. Meteor. Soc., 146, 1999–2049, <a href="https://doi.org/10.1002/qj.3803" target="_blank">https://doi.org/10.1002/qj.3803</a>,
2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib65"><label>Hock(2003)</label><mixed-citation>
      
Hock, R.: Temperature index melt modelling in mountain areas, J.
Hydrol., 282, 104–115, <a href="https://doi.org/10.1016/S0022-1694(03)00257-9" target="_blank">https://doi.org/10.1016/S0022-1694(03)00257-9</a>, 2003.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib66"><label>Hock et al.(2019)</label><mixed-citation>
      
Hock, R., Rasul, G.,  Adler, C., Cáceres, B., Gruber, S., Hirabayashi, Y., Jackson, M., Kääb, A., Kang, S., Kutuzov, S., Milner, A., Molau, U., Morin, S., Orlove, B., and Steltzer, H.: High Mountain Areas, in: IPCC Special Report on the Ocean
and Cryosphere in a Changing Climate, edited by: Pörtner, H.-O., Roberts, C. C., Masson-Delmotte, V., Zhai, P., Tignor, M., Poloczanska, E., Mintenbeck, K., Alegría, A., Nicolai, M., Okem, A., Petzold, J., Rama, B., and Weyer, N. M.,
Cambridge, 131–202, <a href="https://doi.org/10.1017/9781009157964.004" target="_blank">https://doi.org/10.1017/9781009157964.004</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib67"><label>Holland(1992)</label><mixed-citation>
      
Holland, J.: Adaptation in Natural and Artificial Systems, MIT, 2nd Edn., ISBN
9780262581110, 1992.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib68"><label>Immerzeel et al.(2020)Immerzeel, Lutz, Andrade, Bahl, Biemans, Bolch,
Hyde, Brumby, Davies, Elmore, Emmer, Feng, Fernánde, Haritashya, Kargel,
Koppes, Kraaijenbrink, Kulkarni, Mayewski, Nepal, Pacheco, Painter,
Pellicciotti, Rupper, Sinisalo, Srestha, Viviroli, Wada, Xiao, Yao, and
Baillie</label><mixed-citation>
      
Immerzeel, W., Lutz, A., Andrade, M., Bahl, A., Biemans, H., Bolch, T., Hyde,
S., Brumby, S., Davies, B., Elmore, A., Emmer, A., Feng, M., Fernánde,
A., Haritashya, U., Kargel, J., Koppes, M., Kraaijenbrink, P., Kulkarni, A.,
Mayewski, P., Nepal, S., Pacheco, P., Painter, T., Pellicciotti, F. Rajaram,
H., Rupper, S., Sinisalo, A., Srestha, A., Viviroli, D., Wada, Y., Xiao, C.,
Yao, T., and Baillie, J.: Importance and vulnerability of the world’s
water towers, Nature, 577, 364–369, <a href="https://doi.org/10.1038/s41586-019-1822-y" target="_blank">https://doi.org/10.1038/s41586-019-1822-y</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib69"><label>Jaynes(2003)</label><mixed-citation>
      
Jaynes, E.: Probability Theory, Cambridge, <a href="https://doi.org/10.1017/CBO9780511790423" target="_blank">https://doi.org/10.1017/CBO9780511790423</a>,
2003.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib70"><label>Jazwinski(1970)</label><mixed-citation>
      
Jazwinski, A.: Stochastic Processes and Filtering Theory, Academic Press, Vol. 64, 376 pp., <a href="https://doi.org/10.1016/S0076-5392(09)X6022-4" target="_blank">https://doi.org/10.1016/S0076-5392(09)X6022-4</a>,
1970.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib71"><label>Keetz et al.(2025)Keetz, Aalstad, Fisher, Poppe Terán, Naz, Pirk,
Yilmaz, and Skarpaas</label><mixed-citation>
      
Keetz, L. T., Aalstad, K., Fisher, R. A., Poppe Terán, C., Naz, B., Pirk, N.,
Yilmaz, Y. A., and Skarpaas, O.: Inferring Parameters in a Complex Land
Surface Model by Combining Data Assimilation and Machine Learning, J.
Adv. Model. Earth Sy., 17, e2024MS004542,
<a href="https://doi.org/10.1029/2024MS004542" target="_blank">https://doi.org/10.1029/2024MS004542</a>, 2025.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib72"><label>Kitagawa(1996)</label><mixed-citation>
      
Kitagawa, G.: Monte Carlo Filter and Smoother for Non-Gaussian Nonlinear State
Space Models, J. Comput. Graph. Stat., 5, 1–25,
<a href="https://doi.org/10.1080/10618600.1996.10474692" target="_blank">https://doi.org/10.1080/10618600.1996.10474692</a>, 1996.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib73"><label>Koblents and Míguez(2015)</label><mixed-citation>
      
Koblents, E. and Míguez, J.: A population Monte Carlo scheme with transformed
weights and its application to stochastic kinetic models, Stat.
Comput., 25, 407–425, <a href="https://doi.org/10.1007/s11222-013-9440-2" target="_blank">https://doi.org/10.1007/s11222-013-9440-2</a>, 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib74"><label>Krinner et al.(2018)</label><mixed-citation>
      
Krinner, G., Derksen, C., Essery, R., Flanner, M., Hagemann, S., Clark, M., Hall, A., Rott, H., Brutel-Vuilmet, C., Kim, H., Ménard, C. B., Mudryk, L., Thackeray, C., Wang, L., Arduini, G., Balsamo, G., Bartlett, P., Boike, J., Boone, A., Chéruy, F., Colin, J., Cuntz, M., Dai, Y., Decharme, B., Derry, J., Ducharne, A., Dutra, E., Fang, X., Fierz, C., Ghattas, J., Gusev, Y., Haverd, V., Kontu, A., Lafaysse, M., Law, R., Lawrence, D., Li, W., Marke, T., Marks, D., Ménégoz, M., Nasonova, O., Nitta, T., Niwano, M., Pomeroy, J., Raleigh, M. S., Schaedler, G., Semenov, V., Smirnova, T. G., Stacke, T., Strasser, U., Svenson, S., Turkov, D., Wang, T., Wever, N., Yuan, H., Zhou, W., and Zhu, D.: ESM-SnowMIP: assessing snow models and quantifying snow-related climate feedbacks, Geosci. Model Dev., 11, 5027–5049, <a href="https://doi.org/10.5194/gmd-11-5027-2018" target="_blank">https://doi.org/10.5194/gmd-11-5027-2018</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib75"><label>Landmann et al.(2021)Landmann, Künsch, Huss, Ogier, Kalisch, and
Farinotti</label><mixed-citation>
      
Landmann, J. M., Künsch, H. R., Huss, M., Ogier, C., Kalisch, M., and Farinotti, D.: Assimilating near-real-time mass balance stake readings into a model ensemble using a particle filter, The Cryosphere, 15, 5017–5040, <a href="https://doi.org/10.5194/tc-15-5017-2021" target="_blank">https://doi.org/10.5194/tc-15-5017-2021</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib76"><label>Largeron et al.(2020)Largeron, Dumont, Morin, Boone, Lafaysse,
Metref, Cosme, Jonas, Wisntral, and Margulis</label><mixed-citation>
      
Largeron, C., Dumont, M., Morin, S., Boone, A., Lafaysse, M., Metref, S.,
Cosme, E., Jonas, T., Wisntral, A., and Margulis, S.: Toward Snow Cover
Estimation in Mountainous Areas Using Modern Data Assimilation Methods: A
Review, Front. Earth Sci., 8, 325, <a href="https://doi.org/10.3389/feart.2020.00325" target="_blank">https://doi.org/10.3389/feart.2020.00325</a>,
2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib77"><label>Law and Stuart(2012)</label><mixed-citation>
      
Law, K. and Stuart, A.: Evaluating Data Assimilation Algorithms, Mon.
Weather Rev., 140, 3757–3782, <a href="https://doi.org/10.1175/MWR-D-11-00257.1" target="_blank">https://doi.org/10.1175/MWR-D-11-00257.1</a>, 2012.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib78"><label>Leisenring and Moradkhani(2011)</label><mixed-citation>
      
Leisenring, M. and Moradkhani, H.: Snow water equivalent prediction using
Bayesian data assimilation methods, Stoch. Env. Res.
Risk Assess., 25, 253–270, <a href="https://doi.org/10.1007/s00477-010-0445-5" target="_blank">https://doi.org/10.1007/s00477-010-0445-5</a>, 2011.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib79"><label>Li et al.(2015)Li, Bolic, and Djuric</label><mixed-citation>
      
Li, T., Bolic, M., and Djuric, P.: Resampling Methods for Particle Filtering:
Classification, implementation, and strategies, IEEE Signal Process.
Mag., 32, 70–86, <a href="https://doi.org/10.1109/MSP.2014.2330626" target="_blank">https://doi.org/10.1109/MSP.2014.2330626</a>, 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib80"><label>Liston and Elder(2006b)</label><mixed-citation>
      
Liston, G. E. and Elder, K.: A meteorological distribution system for
high-resolution terrestrial modeling (MicroMet, J. Hydrometeorol.,
7, 217–234, <a href="https://doi.org/10.1175/JHM486.1" target="_blank">https://doi.org/10.1175/JHM486.1</a>, 2006b.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib81"><label>Liu et al.(2021)Liu, Fang, and Margulis</label><mixed-citation>
      
Liu, Y., Fang, Y., and Margulis, S. A.: Spatiotemporal distribution of seasonal snow water equivalent in High Mountain Asia from an 18-year Landsat–MODIS era snow reanalysis dataset, The Cryosphere, 15, 5261–5280, <a href="https://doi.org/10.5194/tc-15-5261-2021" target="_blank">https://doi.org/10.5194/tc-15-5261-2021</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib82"><label>Lussana et al.(2018)Lussana, Saloranta, Skaugen, Magnusson, Tveito,
and Andersen</label><mixed-citation>
      
Lussana, C., Saloranta, T., Skaugen, T., Magnusson, J., Tveito, O. E., and Andersen, J.: seNorge2 daily precipitation, an observational gridded dataset over Norway from 1957 to the present day, Earth Syst. Sci. Data, 10, 235–249, <a href="https://doi.org/10.5194/essd-10-235-2018" target="_blank">https://doi.org/10.5194/essd-10-235-2018</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib83"><label>MacKay(2003)</label><mixed-citation>
      
MacKay, D. J. C.: Information Theory, Inference, and Learning Algorithms,
Cambridge,  ISBN 9780521642989,  2003.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib84"><label>Magnusson et al.(2017)Magnusson, Winstral, Stordal, Essery, and
Jonas</label><mixed-citation>
      
Magnusson, J., Winstral, A., Stordal, A. S., Essery, R., and Jonas, T.:
Improving physically based snow simulations by assimilating snow depths
using the particle filter, Water Resour. Res., 53, 1125–1143,
<a href="https://doi.org/10.1002/2016WR019092" target="_blank">https://doi.org/10.1002/2016WR019092</a>, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib85"><label>Margulis et al.(2015)Margulis, Girotto, Cortés, and
Durand</label><mixed-citation>
      
Margulis, S. A., Girotto, M., Cortés, G., and Durand, M.: A particle batch
smoother approach to snow water equivalent estimation, J.
Hydrometeorol., 16, 1752–1772, <a href="https://doi.org/10.1175/JHM-D-14-0177.1" target="_blank">https://doi.org/10.1175/JHM-D-14-0177.1</a>, 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib86"><label>Margulis et al.(2016)Margulis, Cortés, Girotto, and
Durand</label><mixed-citation>
      
Margulis, S. A., Cortés, G., Girotto, M., and Durand, M.: A landsat-era
Sierra Nevada snow reanalysis (1985-2015, J. Hydrometeorol., 17,
1203–1221, <a href="https://doi.org/10.1175/JHM-D-15-0177.1" target="_blank">https://doi.org/10.1175/JHM-D-15-0177.1</a>, 2016.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib87"><label>Marin et al.(2019)Marin, Pudlo, and Sedki</label><mixed-citation>
      
Marin, J.-M., Pudlo, P., and Sedki, M.: Consistency of adaptive importance
sampling and recycling schemes, Bernoulli, 25, 1977–1998,
<a href="https://doi.org/10.3150/18-BEJ1042" target="_blank">https://doi.org/10.3150/18-BEJ1042</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib88"><label>Maussion et al.(2019)Maussion, Butenko, Champollion, Dusch, Eis,
Fourteau, Gregor, Jarosch, Landmann, Oesterle, Recinos, Rothenpieler, Vlug,
Wild, and Marzeion</label><mixed-citation>
      
Maussion, F., Butenko, A., Champollion, N., Dusch, M., Eis, J., Fourteau, K., Gregor, P., Jarosch, A. H., Landmann, J., Oesterle, F., Recinos, B., Rothenpieler, T., Vlug, A., Wild, C. T., and Marzeion, B.: The Open Global Glacier Model (OGGM) v1.1, Geosci. Model Dev., 12, 909–931, <a href="https://doi.org/10.5194/gmd-12-909-2019" target="_blank">https://doi.org/10.5194/gmd-12-909-2019</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib89"><label>Mazzolini et al.(2025)Mazzolini, Aalstad, Alonso-González,
Westermann, and Treichler</label><mixed-citation>
      
Mazzolini, M., Aalstad, K., Alonso-González, E., Westermann, S., and Treichler, D.: Spatio-temporal snow data assimilation with the ICESat-2 laser altimeter, The Cryosphere, 19, 3831–3848, <a href="https://doi.org/10.5194/tc-19-3831-2025" target="_blank">https://doi.org/10.5194/tc-19-3831-2025</a>, 2025.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib90"><label>Ménard and Essery(2019)</label><mixed-citation>
      
Ménard, C. and Essery, R.: ESM-SnowMIP meteorological and evaluation
datasets at ten reference sites (in situ and bias corrected reanalysis
data), PANGAEA [data set], <a href="https://doi.org/10.1594/PANGAEA.897575" target="_blank">https://doi.org/10.1594/PANGAEA.897575</a>,   2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib91"><label>Ménard et al.(2019)Ménard, Essery, Barr, Bartlett, Derry, Dumont,
Fierz, Kim, Kontu, Lejeune, Marks, Niwano, Raleigh, Wang, and
Wever</label><mixed-citation>
      
Ménard, C. B., Essery, R., Barr, A., Bartlett, P., Derry, J., Dumont, M., Fierz, C., Kim, H., Kontu, A., Lejeune, Y., Marks, D., Niwano, M., Raleigh, M., Wang, L., and Wever, N.: Meteorological and evaluation datasets for snow modelling at 10 reference sites: description of in situ and bias-corrected reanalysis data, Earth Syst. Sci. Data, 11, 865–880, <a href="https://doi.org/10.5194/essd-11-865-2019" target="_blank">https://doi.org/10.5194/essd-11-865-2019</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib92"><label>Meredith et al.(2019)</label><mixed-citation>
      
Meredith, M.,  Sommerkorn, M., Cassotta, S., Derksen, C., Ekaykin, A., Hollowed, A., Kofinas, G., Mackintosh, A., Melbourne-Thomas, J., Muelbert, M. M. C.,  Ottersen, G.,  Pritchard, H., and Schuur, E. A. G.: Polar Regions, in: IPCC Special Report on the Ocean and
Cryosphere in a Changing Climate, edited by: Pörtner, H.-O.,  Roberts, C. C.,  Masson-Delmotte, V., Zhai, P.,  Tignor, M., Poloczanska, E., Mintenbeck, K., Alegría, A., Nicolai, M., Okem, A., Petzold, J., Rama, B., and Weyer, N. M.,
Cambridge, 203–320, <a href="https://doi.org/10.1017/9781009157964.005" target="_blank">https://doi.org/10.1017/9781009157964.005</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib93"><label>Metropolis et al.(1953)Metropolis, Rosenbluth, Rosenbluth, Teller,
and Teller</label><mixed-citation>
      
Metropolis, N., Rosenbluth, A., Rosenbluth, M., Teller, A., and Teller, E.:
Equation of State Calculations by Fast Computing Machines, J. Chem. Phys.,
26, 1087–1092, <a href="https://doi.org/10.1063/1.1699114" target="_blank">https://doi.org/10.1063/1.1699114</a>, 1953.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib94"><label>Morzfeld et al.(2017)Morzfeld, Hodyss, and Snyder</label><mixed-citation>
      
Morzfeld, M., Hodyss, D., and Snyder, C.: What the collapse of the ensemble
Kalman filter tells us about particle filters, Tellus A, 69, 1–14, <a href="https://doi.org/10.1080/16000870.2017.1283809" target="_blank">https://doi.org/10.1080/16000870.2017.1283809</a>, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib95"><label>Murphy(2023)</label><mixed-citation>
      
Murphy, K.: Probabilistic Machine Learning: Advanced Topics, MIT,
<a href="https://probml.github.io/book2" target="_blank"/> (last access: 18 December 2023),
2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib96"><label>Navari et al.(2016)Navari, Margulis, Bateni, Tedesco, Alexander, and
Fettweis</label><mixed-citation>
      
Navari, M., Margulis, S. A., Bateni, S. M., Tedesco, M., Alexander, P., and Fettweis, X.: Feasibility of improving a priori regional climate model estimates of Greenland ice sheet surface mass loss through assimilation of measured ice surface temperatures, The Cryosphere, 10, 103–120, <a href="https://doi.org/10.5194/tc-10-103-2016" target="_blank">https://doi.org/10.5194/tc-10-103-2016</a>, 2016.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib97"><label>Neal(2001)</label><mixed-citation>
      
Neal, R.: Annealed importance sampling, Stat. Comput., 11,
125–139, <a href="https://doi.org/10.1023/A:1008923215028" target="_blank">https://doi.org/10.1023/A:1008923215028</a>, 2001.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib98"><label>Neal(2011)</label><mixed-citation>
      
Neal, R.: MCMC Using Hamiltonian Dynamics, in: Handbook of Markov Chain
Monte Carlo, edited by: Brooks, S.,   Gelman, A., Jones, G., and Meng, X.-L., chap. 5,   CRC Press, 113–162,
<a href="https://doi.org/10.1201/b10905" target="_blank">https://doi.org/10.1201/b10905</a>, 2011.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib99"><label>Oberrauch et al.(2024)Oberrauch, Cluzet, Magnusson, and
Jonas</label><mixed-citation>
      
Oberrauch, M., Cluzet, B., Magnusson, J., and Jonas, T.: Improving Fully
Distributed Snowpack Simulations by Mapping Perturbations of Meteorological
Forcings Inferred From Particle Filter Assimilation of Snow Monitoring Data,
Water Resour. Res., 60, e2023WR036994, <a href="https://doi.org/10.1029/2023WR036994" target="_blank">https://doi.org/10.1029/2023WR036994</a>,
2024.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib100"><label>Owen and Zhou(2000)</label><mixed-citation>
      
Owen, A. and Zhou, Y.: Safe and Effective Importance Sampling, J.
Am. Stat. Assoc., 95, 135–143,
<a href="https://doi.org/10.1080/01621459.2000.10473909" target="_blank">https://doi.org/10.1080/01621459.2000.10473909</a>, 2000.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib101"><label>Piazzi et al.(2018)Piazzi, Thirel, Campo, and Gabellani</label><mixed-citation>
      
Piazzi, G., Thirel, G., Campo, L., and Gabellani, S.: A particle filter scheme for multivariate data assimilation into a point-scale snowpack model in an Alpine environment, The Cryosphere, 12, 2287–2306, <a href="https://doi.org/10.5194/tc-12-2287-2018" target="_blank">https://doi.org/10.5194/tc-12-2287-2018</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib102"><label>Pirk et al.(2022)Pirk, Aalstad, Westermann, Vatne, van Hove,
Tallaksen, Cassiani, and Katul</label><mixed-citation>
      
Pirk, N., Aalstad, K., Westermann, S., Vatne, A., van Hove, A., Tallaksen, L. M., Cassiani, M., and Katul, G.: Inferring surface energy fluxes using drone data assimilation in large eddy simulations, Atmos. Meas. Tech., 15, 7293–7314, <a href="https://doi.org/10.5194/amt-15-7293-2022" target="_blank">https://doi.org/10.5194/amt-15-7293-2022</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib103"><label>Pirk et al.(2024)Pirk, Aalstad, Mannerfelt, Clayer, de Wit,
Christiansen, Althuizen, Lee, and Westermann</label><mixed-citation>
      
Pirk, N., Aalstad, K., Mannerfelt, E. S., Clayer, F., de Wit, H., Christiansen,
C. T., Althuizen, I., Lee, H., and Westermann, S.: Disaggregating the Carbon
Exchange of Degrading Permafrost Peatlands Using Bayesian Deep Learning,
Geophys. Res. Lett., 51, e2024GL109283,
<a href="https://doi.org/10.1029/2024GL109283" target="_blank">https://doi.org/10.1029/2024GL109283</a>,  2024.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib104"><label>Rainforth et al.(2020)Rainforth, Golinski, Wood, and
Zaidi</label><mixed-citation>
      
Rainforth, T., Golinski, A., Wood, F., and Zaidi, S.: Target–Aware Bayesian
Inference: How to Beat Optimal Conventional Estimators, JMLR, 21, 1–54,
2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib105"><label>Reich and Cotter(2015)</label><mixed-citation>
      
Reich, S. and Cotter, C.: Probabilistic Forecasting and Bayesian Data
Assimilation, Cambridge, <a href="https://doi.org/10.1017/CBO9781107706804" target="_blank">https://doi.org/10.1017/CBO9781107706804</a>, 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib106"><label>Revuelto et al.(2017)Revuelto, Azorin-Molina, Alonso-González,
Sanmiguel-Vallelado, Navarro-Serrano, Rico, and Ignacio
López-Moreno</label><mixed-citation>
      
Revuelto, J., Azorin-Molina, C., Alonso-González, E., Sanmiguel-Vallelado, A., Navarro-Serrano, F., Rico, I., and López-Moreno, J. I.: Meteorological and snow distribution data in the Izas Experimental Catchment (Spanish Pyrenees) from 2011 to 2017, Earth Syst. Sci. Data, 9, 993–1005, <a href="https://doi.org/10.5194/essd-9-993-2017" target="_blank">https://doi.org/10.5194/essd-9-993-2017</a>, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib107"><label>Revuelto et al.(2021)Revuelto, Alonso-Gonzalez, Vidaller-Gayan,
Lacroix, Izagirre, Rodríguez-López, and
López-Moreno</label><mixed-citation>
      
Revuelto, J., Alonso-Gonzalez, E., Vidaller-Gayan, I., Lacroix, E., Izagirre,
E., Rodríguez-López, G., and López-Moreno, J.: Intercomparison
of UAV platforms for mapping snow depth distribution in complex alpine
terrain, Cold Reg. Sci. Technol., 190, 103344,
<a href="https://doi.org/10.1016/j.coldregions.2021.103344" target="_blank">https://doi.org/10.1016/j.coldregions.2021.103344</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib108"><label>Riihelä et al.(2021)Riihelä, Bright, and
Anttila</label><mixed-citation>
      
Riihelä, A., Bright, R. M., and Anttila, K.: Recent strengthening of snow
and ice albedo feedback driven by Antarctic sea-ice loss, Nat. Geosci.,
14, 832–836, <a href="https://doi.org/10.1038/s41561-021-00841-x" target="_blank">https://doi.org/10.1038/s41561-021-00841-x</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib109"><label>Robert(2007)</label><mixed-citation>
      
Robert, C.: The Bayesian Choice, Springer, 2nd Edn.,
<a href="https://doi.org/10.1007/0-387-71599-1" target="_blank">https://doi.org/10.1007/0-387-71599-1</a>, 2007.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib110"><label>Robert and Casella(2004)</label><mixed-citation>
      
Robert, C. and Casella, G.: Monte Carlo Statistical Methods, Springer,
<a href="https://doi.org/10.1007/978-1-4757-4145-2" target="_blank">https://doi.org/10.1007/978-1-4757-4145-2</a>, 2004.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib111"><label>Rounce et al.(2023)Rounce, Hock, Maussion, Hugonnet, Kochtitzky,
Huss, Berthier, Brinkerhoff, Compagno, Copland, Farinotti, Menounos, and
McNabb</label><mixed-citation>
      
Rounce, D. R., Hock, R., Maussion, F., Hugonnet, R., Kochtitzky, W., Huss, M.,
Berthier, E., Brinkerhoff, D., Compagno, L., Copland, L., Farinotti, D.,
Menounos, B., and McNabb, R. W.: Global glacier change in the 21st century:
Every increase in temperature matters, Science, 379, 78–83,
<a href="https://doi.org/10.1126/science.abo1324" target="_blank">https://doi.org/10.1126/science.abo1324</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib112"><label>Sanz-Alonso et al.(2023)Sanz-Alonso, Stuart, and
Taeb</label><mixed-citation>
      
Sanz-Alonso, D., Stuart, A. M., and Taeb, A.: Inverse problems and data
assimilation, vol. 107, Cambridge, <a href="https://doi.org/10.1017/9781009414319" target="_blank">https://doi.org/10.1017/9781009414319</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib113"><label>Särkkä and Svensson(2023)</label><mixed-citation>
      
Särkkä, S. and Svensson, L.: Bayesian Filtering and Smoothing,
Cambridge, 2nd Edn., <a href="https://doi.org/10.1017/9781108917407" target="_blank">https://doi.org/10.1017/9781108917407</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib114"><label>Smith and Gelfand(1992)</label><mixed-citation>
      
Smith, A. F. M. and Gelfand, A. E.: Bayesian Statistics without Tears: A
Sampling–Resampling Perspective,   Am. Stat., 46, 84–88,
<a href="https://doi.org/10.1080/00031305.1992.10475856" target="_blank">https://doi.org/10.1080/00031305.1992.10475856</a>, 1992.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib115"><label>Smith et al.(2010)Smith, Sharma, Marshall, Mehrotra, and
Sisson</label><mixed-citation>
      
Smith, T., Sharma, A., Marshall, L., Mehrotra, R., and Sisson, S.: Development
of a formal likelihood function for improved Bayesian inference of ephemeral
catchments, Water Resour. Res., 46, <a href="https://doi.org/10.1029/2010WR009514" target="_blank">https://doi.org/10.1029/2010WR009514</a>, 2010.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib116"><label>Smyth et al.(2019)Smyth, Raleigh, and Small</label><mixed-citation>
      
Smyth, E. J., Raleigh, M. S., and Small, E. E.: Particle Filter Data
Assimilation of Monthly Snow Depth Observations Improves Estimation of Snow
Density and SWE, Water Resour. Res., 55, 1296–1311,
<a href="https://doi.org/10.1029/2018WR023400" target="_blank">https://doi.org/10.1029/2018WR023400</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib117"><label>Snyder et al.(2008)Snyder, Bengtsson, Bickel, and
Anderson</label><mixed-citation>
      
Snyder, C., Bengtsson, T., Bickel, P., and Anderson, J.: Obstacles to
High-Dimensional Particle Filtering, Mon. Weather Rev., 136, 4629–4640, <a href="https://doi.org/10.1175/2008MWR2529.1" target="_blank">https://doi.org/10.1175/2008MWR2529.1</a>, 2008.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib118"><label>Sörensen(2015)</label><mixed-citation>
      
Sörensen, K.: Metaheuristics – the metaphor exposed, Int.
Trans. Oper. Res., 22, 3–18, <a href="https://doi.org/10.1111/itor.12001" target="_blank">https://doi.org/10.1111/itor.12001</a>,
2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib119"><label>Stordal and Elsheikh(2015)</label><mixed-citation>
      
Stordal, A. S. and Elsheikh, A. H.: Iterative ensemble smoothers in the
annealed importance sampling framework, Adv. Water Resour., 86,
231–239, <a href="https://doi.org/10.1016/j.advwatres.2015.09.030" target="_blank">https://doi.org/10.1016/j.advwatres.2015.09.030</a>, 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib120"><label>Sun et al.(2025)Sun, Fang, Margulis, Mortimer, Mudryk, and
Derksen</label><mixed-citation>
      
Sun, H., Fang, Y., Margulis, S. A., Mortimer, C., Mudryk, L., and Derksen, C.: Evaluation of the Snow Climate Change Initiative (Snow CCI) snow-covered area product within a mountain snow water equivalent reanalysis, The Cryosphere, 19, 2017–2036, <a href="https://doi.org/10.5194/tc-19-2017-2025" target="_blank">https://doi.org/10.5194/tc-19-2017-2025</a>, 2025.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib121"><label>Tang et al.(2023)Tang, Frye, Gelfand, and Silander</label><mixed-citation>
      
Tang, B., Frye, H., Gelfand, A., and Silander, J. A.: Zero-Inflated Beta
Distribution Regression Modeling, J. Agr., Biol.
Environ. Stat., 28, 117–137, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib122"><label>van Hove et al.(2025)van Hove, Aalstad, Lind, Arndt, Odongo, Ceriani,
Fava, Hulth, and Pirk</label><mixed-citation>
      
van Hove, A., Aalstad, K., Lind, V., Arndt, C., Odongo, V., Ceriani, R., Fava, F., Hulth, J., and Pirk, N.: Inferring methane emissions from African livestock by fusing drone, tower, and satellite data, Biogeosciences, 22, 4163–4186, <a href="https://doi.org/10.5194/bg-22-4163-2025" target="_blank">https://doi.org/10.5194/bg-22-4163-2025</a>, 2025.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib123"><label>van Hove et al.(2026)van Hove, Aalstad, and Pirk</label><mixed-citation>
      
van Hove, A., Aalstad, K., and Pirk, N.: Actively inferring methane sources
with drones, Environ. Data Sci., 5, e2, <a href="https://doi.org/10.1017/eds.2026.10029" target="_blank">https://doi.org/10.1017/eds.2026.10029</a>,
2026.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib124"><label>van Leeuwen(2009)</label><mixed-citation>
      
van Leeuwen, P. J.: Particle Filtering in Geophysical Systems, Mon.
Weather Rev., 137, 4089–4114, <a href="https://doi.org/10.1175/2009MWR2835.1" target="_blank">https://doi.org/10.1175/2009MWR2835.1</a>, 2009.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib125"><label>van Leeuwen and Evensen(1996)</label><mixed-citation>
      
van Leeuwen, P. J. and Evensen, G.: Data Assimilation and Inverse Methods in
Terms of a Probabilistic Formulation, Mon. Weather Rev., 124,
2898–2913, <a href="https://doi.org/10.1175/1520-0493(1996)124&lt;2898:DAAIMI&gt;2.0.CO;2" target="_blank">https://doi.org/10.1175/1520-0493(1996)124&lt;2898:DAAIMI&gt;2.0.CO;2</a>, 1996.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib126"><label>van Leeuwen et al.(2019)van Leeuwen, Künsch, Nerger, Potthast,
and Reich</label><mixed-citation>
      
van Leeuwen, P. J., Künsch, H. R., Nerger, L., Potthast, R., and Reich, S.:
Particle filters for high-dimensional geoscience applications: A review,
Q. J. Roy. Meteor. Soc., 145, 2335–2365,
<a href="https://doi.org/10.1002/qj.3551" target="_blank">https://doi.org/10.1002/qj.3551</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib127"><label>Vihola(2012)</label><mixed-citation>
      
Vihola, M.: Robust adaptive Metropolis algorithm with coerced acceptance
rate, Stat. Comput., 22, 997–1008,
<a href="https://doi.org/10.1007/s11222-011-9269-5" target="_blank">https://doi.org/10.1007/s11222-011-9269-5</a>, 2012.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib128"><label>Virtanen et al.(2020)</label><mixed-citation>
      
Virtanen, P., Gommers, R., Oliphant, T. E., Haberland, M., Reddy, T.,
Cournapeau, D., Burovski, E., Peterson, P., Weckesser, W., Bright, J.,
van der Walt, S. J., Brett, M., Wilson, J., Millman, K. J., Mayorov, N.,
Nelson, A. R. J., Jones, E., Kern, R., Larson, E., Carey, C. J., Polat, I.,
Feng, Y., Moore, E. W., VanderPlas, J., Laxalde, D., Perktold, J., Cimrman,
R., Henriksen, I., Quintero, E. A., Harris, C. R., Archibald, A. M., Ribeiro,
A. H., Pedregosa, F., van Mulbregt, P., Vijaykumar, A., Bardelli, A. P.,
Rothberg, A., Hilboll, A., Kloeckner, A., Scopatz, A., Lee, A., Rokem, A.,
Woods, C. N., Fulton, C., Masson, C., Häggström, C., Fitzgerald, C.,
Nicholson, D. A., Hagen, D. R., Pasechnik, D. V., Olivetti, E., Martin, E.,
Wieser, E., Silva, F., Lenders, F., Wilhelm, F., Young, G., Price, G. A.,
Ingold, G.-L., Allen, G. E., Lee, G. R., Audren, H., Probst, I., Dietrich,
J. P., Silterra, J., Webber, J. T., Slavič, J., Nothman, J., Buchner, J.,
Kulick, J., Schönberger, J. L., de Miranda Cardoso, J. V., Reimer, J.,
Harrington, J., Rodríguez, J. L. C., Nunez-Iglesias, J., Kuczynski, J.,
Tritz, K., Thoma, M., Newville, M., Kümmerer, M., Bolingbroke, M., Tartre,
M., Pak, M., Smith, N. J., Nowaczyk, N., Shebanov, N., Pavlyk, O., Brodtkorb,
P. A., Lee, P., McGibbon, R. T., Feldbauer, R., Lewis, S., Tygier, S.,
Sievert, S., Vigna, S., Peterson, S., More, S., Pudlik, T., Oshima, T.,
Pingel, T. J., Robitaille, T. P., Spura, T., Jones, T. R., Cera, T., Leslie,
T., Zito, T., Krauss, T., Upadhyay, U., Halchenko, Y. O., and Vázquez-Baeza,
Y.: SciPy 1.0: fundamental algorithms for scientific computing in Python,
Nat. Meth., 17, 261–272, <a href="https://doi.org/10.1038/s41592-019-0686-2" target="_blank">https://doi.org/10.1038/s41592-019-0686-2</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib129"><label>Vrugt et al.(2003)Vrugt, Gupta, Bouten, and Sorooshian</label><mixed-citation>
      
Vrugt, J. A., Gupta, H. V., Bouten, W., and Sorooshian, S.: A Shuffled Complex
Evolution Metropolis algorithm for optimization and uncertainty assessment of
hydrologic model parameters, Water Resour. Res., 39,
<a href="https://doi.org/10.1029/2002WR001642" target="_blank">https://doi.org/10.1029/2002WR001642</a>, 2003.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib130"><label>Vrugt et al.(2008)Vrugt, ter Braak, Clark, Hyman, and
Robinson</label><mixed-citation>
      
Vrugt, J. A., ter Braak, C. J. F., Clark, M. P., Hyman, J. M., and Robinson,
B. A.: Treatment of input uncertainty in hydrologic modeling: Doing hydrology
backward with Markov chain Monte Carlo simulation, Water Resour. Res.,
44, W00B09, <a href="https://doi.org/10.1029/2007WR006720" target="_blank">https://doi.org/10.1029/2007WR006720</a>, 2008.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib131"><label>Westermann et al.(2023)Westermann, Ingeman-Nielsen, Scheer, Aalstad,
Aga, Chaudhary, Etzelmüller, Filhol, Kääb, Renette, Schmidt, Schuler,
Zweigel, Martin, Morard, Ben-Asher, Angelopoulos, Boike, Groenke, Miesner,
Nitzbon, Overduin, Stuenzi, and Langer</label><mixed-citation>
      
Westermann, S., Ingeman-Nielsen, T., Scheer, J., Aalstad, K., Aga, J., Chaudhary, N., Etzelmüller, B., Filhol, S., Kääb, A., Renette, C., Schmidt, L. S., Schuler, T. V., Zweigel, R. B., Martin, L., Morard, S., Ben-Asher, M., Angelopoulos, M., Boike, J., Groenke, B., Miesner, F., Nitzbon, J., Overduin, P., Stuenzi, S. M., and Langer, M.: The CryoGrid community model (version 1.0) – a multi-physics toolbox for climate-driven simulations in the terrestrial cryosphere, Geosci. Model Dev., 16, 2607–2647, <a href="https://doi.org/10.5194/gmd-16-2607-2023" target="_blank">https://doi.org/10.5194/gmd-16-2607-2023</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib132"><label>Wikle and Berliner(2007)</label><mixed-citation>
      
Wikle, C. K. and Berliner, L. M.: A Bayesian tutorial for data assimilation,
Phys. D, 230, 1–16,
<a href="https://doi.org/10.1016/j.physd.2006.09.017" target="_blank">https://doi.org/10.1016/j.physd.2006.09.017</a>, 2007.


    </mixed-citation></ref-html>
<ref-html id="bib1.bib133"><label>Willmes et al.(2025)Willmes, Aalstad, and Westermann</label><mixed-citation>
      
Willmes, C., Aalstad, K., and Westermann, S.: Assimilating high-resolution satellite snow cover data in a permafrost model, EGUsphere [preprint], <a href="https://doi.org/10.5194/egusphere-2025-3142" target="_blank">https://doi.org/10.5194/egusphere-2025-3142</a>, 2025.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib134"><label>Yang et al.(2026)Yang, Ultee, Aalstad, Debolsky, Hock, Schmitt, and
Rounce</label><mixed-citation>
      
Yang, R., Ultee, L., Aalstad, K., Debolskiy, M. V., Hock, R., Schmitt, P., Rounce, D., and Li, T.: Joint Bayesian Calibration of Frontal Ablation and Surface Mass Balance in Global Glacier Models, EGUsphere [preprint], <a href="https://doi.org/10.5194/egusphere-2026-1081" target="_blank">https://doi.org/10.5194/egusphere-2026-1081</a>, 2026.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib135"><label>Zschenderlein et al.(2023)Zschenderlein, Luojus, Takala,
Venäläinen, and Pulliainen</label><mixed-citation>
      
Zschenderlein, L., Luojus, K., Takala, M., Venäläinen, P., and Pulliainen,
J.: Evaluation of passive microwave dry snow detection algorithms and
application to SWE retrieval during seasonal snow accumulation, Remote
Sens. Environ., 288, 113476,
<a href="https://doi.org/10.1016/j.rse.2023.113476" target="_blank">https://doi.org/10.1016/j.rse.2023.113476</a>, 2023.

    </mixed-citation></ref-html>--></article>
