<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing with OASIS Tables v3.0 20080202//EN" "https://jats.nlm.nih.gov/nlm-dtd/publishing/3.0/journalpub-oasis3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:oasis="http://docs.oasis-open.org/ns/oasis-exchange/table" xml:lang="en" dtd-version="3.0" article-type="research-article">
  <front>
    <journal-meta><journal-id journal-id-type="publisher">GMD</journal-id><journal-title-group>
    <journal-title>Geoscientific Model Development</journal-title>
    <abbrev-journal-title abbrev-type="publisher">GMD</abbrev-journal-title><abbrev-journal-title abbrev-type="nlm-ta">Geosci. Model Dev.</abbrev-journal-title>
  </journal-title-group><issn pub-type="epub">1991-9603</issn><publisher>
    <publisher-name>Copernicus Publications</publisher-name>
    <publisher-loc>Göttingen, Germany</publisher-loc>
  </publisher></journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.5194/gmd-19-7893-2026</article-id><title-group><article-title>Paleoclimate data assimilation with adaptive observation error inflation and adaptive localization</article-title><alt-title>Paleoclimate data assimilation with AOEI and adaptive localization</alt-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author" corresp="no" rid="aff1 aff2">
          <name><surname>Luo</surname><given-names>Ge</given-names></name>
          
        </contrib>
        <contrib contrib-type="author" corresp="yes" rid="aff1 aff2">
          <name><surname>Zeng</surname><given-names>Yuefei</given-names></name>
          <email>yuefei.zeng@nuist.edu.cn</email>
        <ext-link>https://orcid.org/0000-0003-2927-7049</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff3">
          <name><surname>Zhu</surname><given-names>Feng</given-names></name>
          
        <ext-link>https://orcid.org/0000-0002-9969-2953</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1 aff2">
          <name><surname>Zhao</surname><given-names>Jiuwei</given-names></name>
          
        </contrib>
        <aff id="aff1"><label>1</label><institution>State Key Laboratory of Climate System Prediction and Risk Management/Key Laboratory of Meteorological Disaster, Ministry of Education/Collaborative Innovation Center on Forecast and Evaluation of Meteorological Disasters, Nanjing University of Information Science and Technology, Nanjing 210044, China</institution>
        </aff>
        <aff id="aff2"><label>2</label><institution>School of Atmospheric Sciences, Nanjing University of Information Science and Technology, Nanjing 210044, China</institution>
        </aff>
        <aff id="aff3"><label>3</label><institution>Climate and Global Dynamics Laboratory, NSF National Center for Atmospheric Research, Boulder, CO, USA</institution>
        </aff>
      </contrib-group>
      <author-notes><corresp id="corr1">Yuefei Zeng (yuefei.zeng@nuist.edu.cn)</corresp></author-notes><pub-date><day>25</day><month>August</month><year>2026</year></pub-date>
      
      <volume>19</volume>
      <issue>16</issue>
      <fpage>7893</fpage><lpage>7909</lpage>
      <history>
        <date date-type="received"><day>14</day><month>March</month><year>2026</year></date>
           <date date-type="rev-request"><day>1</day><month>June</month><year>2026</year></date>
           <date date-type="rev-recd"><day>29</day><month>July</month><year>2026</year></date>
           <date date-type="accepted"><day>16</day><month>August</month><year>2026</year></date>
      </history>
      <permissions>
        <copyright-statement>Copyright: © 2026 Ge Luo et al.</copyright-statement>
        <copyright-year>2026</copyright-year>
      <license license-type="open-access"><license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p></license></permissions><self-uri xlink:href="https://gmd.copernicus.org/articles/19/7893/2026/gmd-19-7893-2026.html">This article is available from https://gmd.copernicus.org/articles/19/7893/2026/gmd-19-7893-2026.html</self-uri><self-uri xlink:href="https://gmd.copernicus.org/articles/19/7893/2026/gmd-19-7893-2026.pdf">The full text article is available as a PDF file from https://gmd.copernicus.org/articles/19/7893/2026/gmd-19-7893-2026.pdf</self-uri>
      <abstract><title>Abstract</title>

      <p id="d2e124">Paleoclimate data assimilation methods significantly enhance the accuracy, spatiotemporal continuity, and global relevance of climate reconstructions by integrating Earth system models with proxy records. In this study, we further improve the algorithm by implementing two adaptive strategies – adaptive observation error inflation and adaptive localization – and systematically evaluate their performance in reconstructing temperature data over equatorial regions. For the adaptive observation error inflation experiments, two distinct methods were employed: the adaptive observation error inflation (AOEI) method yields significant improvements in specific regions but also introduces local biases, whereas the Huber Robust Estimation (HAOEI) method provides more robust and spatially consistent enhancements overall. In the adaptive localization experiments, the localization radius and weight matrix at each grid point are dynamically adjusted based on observational density and correlation information. This strategy effectively utilizes sparse observational data, suppresses spurious teleconnections, accurately reproduces the spatial structure of dominant climate variability modes, and thereby enhances the overall stability of the analyzed field.</p>
  </abstract>
    
<funding-group>
<award-group id="gs1">
<funding-source>National Key Research and Development Program of China</funding-source>
<award-id>2023YFF0804703</award-id>
</award-group>
</funding-group>
</article-meta>
  </front>
<body>
      

<sec id="Ch1.S1" sec-type="intro">
  <label>1</label><title>Introduction</title>
      <p id="d2e136">Reconstructing paleoclimate states is a crucial link in understanding the mechanisms and impacts of Earth system evolution. Currently, this field primarily relies on two methods: paleoclimate proxy records and Earth system model simulations. However, both approaches have significant limitations when reconstructing past climate states. On one hand, proxy records (such as tree rings, ice cores, speleothems, corals and etc.) play an indispensable role in reconstructing climate over the past millennia due to their long temporal span, but they often suffer from issues such as uneven spatiotemporal distribution, discontinuity, and sparsity. Furthermore, climate signals are easily contaminated by noise during preservation, making it extremely challenging to establish quantitative relationships between proxy indicators and the true climate state <xref ref-type="bibr" rid="bib1.bibx23 bib1.bibx24 bib1.bibx35 bib1.bibx20 bib1.bibx11 bib1.bibx38" id="paren.1"/>, and reconstructions of the same climate state from different proxy sources often show inconsistencies. On the other hand, climate models based on physical processes can provide physically consistent, globally complete climate fields with full spatiotemporal coverage, effectively capturing large-scale features of the climate system, however, due to inaccuracies in parameterizing physical processes and uncertainties in response to external forcings, model simulations may exhibit systematic biases, struggling to fully reproduce observed climate variability and leading to limited predictive skills <xref ref-type="bibr" rid="bib1.bibx40 bib1.bibx31 bib1.bibx16" id="paren.2"/>. Therefore, to overcome the limitations of individual methods, paleoclimate data assimilation (PDA) techniques have emerged as a powerful tool to optimally merge information from proxy records – sparse, noisy, and indirect indicators of past climate – with the dynamical constraints of climate models <xref ref-type="bibr" rid="bib1.bibx37 bib1.bibx12 bib1.bibx10" id="paren.3"/>. By doing so, the PDA method produces spatially complete and physically consistent climate field estimates, akin to reanalysis products for the instrumental era, while also quantifying reconstruction uncertainties. Among other methods, the Ensemble Kalman Filter <xref ref-type="bibr" rid="bib1.bibx8" id="paren.4"><named-content content-type="pre">EnKF,</named-content></xref>, has gained prominence in PDA due to its easy implementation and capability to handle high-dimensional nonlinear systems <xref ref-type="bibr" rid="bib1.bibx12 bib1.bibx34 bib1.bibx28 bib1.bibx33" id="paren.5"/>. However, EnKF also suffers from sampling errors due to finite ensemble size, assumes near-Gaussian error distributions, and requires covariance localization to suppress spurious long-range correlations – issues that are particularly acute in paleoclimate applications with sparse proxies.</p>
      <p id="d2e156">Despite the great potential of PDA, it faces unique challenges distinct from modern meteorological data assimilation, primarily due to the fact that proxy records are fairly sparse in time and space, and often represent time-averaged climatic signals (e.g., annual means) rather than instantaneous observations. Furthermore, proxy data are subject to complex errors arising from measurement inaccuracies, chronological uncertainties, and the imperfect relationship between proxy signals and the target climate variables. In the PDA, observation error variance is often specified empirically from the residual variance of a Proxy System Model <xref ref-type="bibr" rid="bib1.bibx7 bib1.bibx6 bib1.bibx42" id="paren.6"><named-content content-type="pre">PSM,</named-content></xref> calibrated against instrumental data. However, this method typically underestimates true proxy uncertainty, as it fails to account for calibration error (such as non-stationarity between past and instrumental periodes), structural error (arising from missing physics, biology, or chemistry in the PSM), and representation error (due to the mismatch between point-scale proxies and gridded state variables). To compensate, observation error inflation is frequently used <xref ref-type="bibr" rid="bib1.bibx33" id="paren.7"/>. Given the strong spatial and temporal heterogeneity of proxy errors and information content, in PDA adaptive inflation methods are generally favored over a fixed inflation approach. While modern data assimilation has adopted techniques like Adaptive Observation Error Inflation <xref ref-type="bibr" rid="bib1.bibx27" id="paren.8"><named-content content-type="pre">AOEI,</named-content></xref>, its applicability to the specific challenges of the PDA remains insufficiently discussed and requires further exploration.</p>
      <p id="d2e172">A critical factor for the success of the EnKF is its use of covariance localization. Early implementations typically employed a fixed localization radius <xref ref-type="bibr" rid="bib1.bibx13 bib1.bibx14" id="paren.9"/>. Within this framework, a common approach is to taper the sampled covariances – between observations and model states, or between different model states – using a smooth, distance-dependent function. A standard choice for this function is the Gaussian-like Gaspari-Cohn (GC) function <xref ref-type="bibr" rid="bib1.bibx9" id="paren.10"/>. However, the Gaussian-like tapering function is not necessarily optimal as shown in several studies <xref ref-type="bibr" rid="bib1.bibx1 bib1.bibx17 bib1.bibx18" id="paren.11"/>. To address limitations of the fixed localization, a range of adaptive localization methods have been developed. Those methods adjust the localization often in response to the state-dependent correlation patterns or ensemble-estimated errors. Examples include techniques developed by <xref ref-type="bibr" rid="bib1.bibx1 bib1.bibx2" id="text.12"/>, <xref ref-type="bibr" rid="bib1.bibx3 bib1.bibx4" id="text.13"/>, and <xref ref-type="bibr" rid="bib1.bibx25 bib1.bibx26" id="text.14"/>. However, these methods have yet to be applied to PDA context. In PDA, a large covariance localization radius is typically employed due to the sparse distribution of proxy records <xref ref-type="bibr" rid="bib1.bibx34" id="paren.15"/>. While necessary, an excessively large localization radius can introduce spurious long-distance correlations, ultimately degrading analysis quality. This limitation suggests that localization with a fixed radius is suboptimal for PDA applications. Instead, an adaptive localization, which can optimize the influence radius of observations based on the underlying spatial correlation structures, applying broader scales in regions with strong large-scale covariability and tighter constraints in data-rich or locally forced regions, is preferred. Currently, research on adaptive localization techniques in the PDA is still very limited.</p>
      <p id="d2e197">In this paper, adaptive observation error inflation and adaptive covariance localization will be developed and explored within an ensemble-based PDA framework, aiming to create a more flexible and accurate assimilation system. As a methodological paper, a pseudoproxy experiment might be more informative than a real-word one. But we directly conduct real-world proxy experiments using a climate model to systematically evaluate whether these adaptive strategies can enhance reconstruction skill – particularly in sparse data regimes – improve uncertainty quantification, and provide a more generalizable approach for assimilating the heterogeneous and uncertain proxy records that characterize paleoclimatology.</p>
</sec>
<sec id="Ch1.S2">
  <label>2</label><title>Methods</title>
<sec id="Ch1.S2.SS1">
  <label>2.1</label><title>Offline Data Assimilation</title>
      <p id="d2e215">The offline EnKF is a widely used approach for paleoclimate data  assimilation, popularized in the paleoclimate context by <xref ref-type="bibr" rid="bib1.bibx32" id="text.16"/>, <xref ref-type="bibr" rid="bib1.bibx12" id="text.17"/>, and <xref ref-type="bibr" rid="bib1.bibx34" id="text.18"/>. “Offline” means that the background ensemble is static – typically drawn from climate model simulations covering the entire reconstruction period – rather than evolving through a continuous forecast cycle. The state update equation is given by

            <disp-formula id="Ch1.E1" content-type="numbered"><label>1</label><mml:math id="M1" display="block"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>b</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:mi mathvariant="bold">K</mml:mi><mml:mfenced close="]" open="["><mml:mrow><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>-</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>e</mml:mi></mml:msub></mml:mrow></mml:mfenced><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

          where <inline-formula><mml:math id="M2" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>b</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the ensemble prior. The vector <inline-formula><mml:math id="M3" display="inline"><mml:mi mathvariant="bold-italic">y</mml:mi></mml:math></inline-formula> represents the assimilated proxy data, and <inline-formula><mml:math id="M4" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>e</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M5" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M6" display="inline"><mml:mrow><mml:mi>H</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>b</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the vector of proxy estimates derived from the prior through the forward operator <inline-formula><mml:math id="M7" display="inline"><mml:mi>H</mml:mi></mml:math></inline-formula>. <inline-formula><mml:math id="M8" display="inline"><mml:mi mathvariant="bold">K</mml:mi></mml:math></inline-formula> is the Kalman gain matrix:

            <disp-formula id="Ch1.E2" content-type="numbered"><label>2</label><mml:math id="M9" display="block"><mml:mrow><mml:mi mathvariant="bold">K</mml:mi><mml:mo>=</mml:mo><mml:msup><mml:mi mathvariant="bold">BH</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:msup><mml:mfenced open="[" close="]"><mml:mrow><mml:msup><mml:mi mathvariant="bold">HBH</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:mo>+</mml:mo><mml:mi mathvariant="bold">R</mml:mi></mml:mrow></mml:mfenced><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

          where <inline-formula><mml:math id="M10" display="inline"><mml:mi mathvariant="bold">B</mml:mi></mml:math></inline-formula> is the prior covariance matrix, <inline-formula><mml:math id="M11" display="inline"><mml:mi mathvariant="bold">R</mml:mi></mml:math></inline-formula> is the error covariance matrix of the proxy data, and <inline-formula><mml:math id="M12" display="inline"><mml:mi mathvariant="bold">H</mml:mi></mml:math></inline-formula> is the linearization of the forward operator <inline-formula><mml:math id="M13" display="inline"><mml:mi>H</mml:mi></mml:math></inline-formula> about the prior mean. The above Eq. (<xref ref-type="disp-formula" rid="Ch1.E1"/>) is solved using the ensemble square-root filter <xref ref-type="bibr" rid="bib1.bibx39" id="paren.19"><named-content content-type="pre">EnSRF,</named-content></xref>. <inline-formula><mml:math id="M14" display="inline"><mml:mi mathvariant="bold">R</mml:mi></mml:math></inline-formula> is taken as a diagonal matrix (uncorrelated observation errors), with the diagonal elements representing the error variance for each assimilated proxy record <xref ref-type="bibr" rid="bib1.bibx34" id="paren.20"/>. This allows for serial processing of observations, in which observations are assimilated one at a time, greatly simplifying the implementation of covariance localization. For a single <inline-formula><mml:math id="M15" display="inline"><mml:mi>j</mml:mi></mml:math></inline-formula>th proxy data <inline-formula><mml:math id="M16" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, the ensemble mean is updated by

            <disp-formula id="Ch1.E3" content-type="numbered"><label>3</label><mml:math id="M17" display="block"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mi>a</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mi>b</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mi mathvariant="normal">loc</mml:mi></mml:msub><mml:mo>∘</mml:mo><mml:mtext>cov</mml:mtext><mml:mfenced close=")" open="("><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>b</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mrow><mml:mi mathvariant="normal">e</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced></mml:mrow><mml:mrow><mml:mtext>var</mml:mtext><mml:mfenced close=")" open="("><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mrow><mml:mi mathvariant="normal">e</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced><mml:mo>+</mml:mo><mml:msub><mml:mi>R</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:mfenced close=")" open="("><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mrow><mml:mi mathvariant="normal">e</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

          where <inline-formula><mml:math id="M18" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mrow><mml:mi mathvariant="normal">e</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mover accent="true"><mml:mrow><mml:mi>H</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>b</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the prior estimate of the <inline-formula><mml:math id="M19" display="inline"><mml:mi>j</mml:mi></mml:math></inline-formula>th proxy from the ensemble mean, <inline-formula><mml:math id="M20" display="inline"><mml:mrow><mml:msub><mml:mi>R</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the observation error variance for the <inline-formula><mml:math id="M21" display="inline"><mml:mi>j</mml:mi></mml:math></inline-formula>th proxy record, and <inline-formula><mml:math id="M22" display="inline"><mml:mrow><mml:mtext>cov</mml:mtext><mml:mo>(</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M23" display="inline"><mml:mrow><mml:mtext>var</mml:mtext><mml:mo>(</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> denote the covariance and variance functions. Covariance localization is represented by the Schur product (denoted by <inline-formula><mml:math id="M24" display="inline"><mml:mo>∘</mml:mo></mml:math></inline-formula>, i.e., element-wise multiplication) in the equation, which acts as a distance-weighted filter <inline-formula><mml:math id="M25" display="inline"><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mi mathvariant="normal">loc</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> on the prior covariance matrix to suppress spurious long-distance correlations. The <inline-formula><mml:math id="M26" display="inline"><mml:mi>j</mml:mi></mml:math></inline-formula>th ensemble perturbation is updated by

            <disp-formula id="Ch1.E4" content-type="numbered"><label>4</label><mml:math id="M27" display="block"><mml:mrow><mml:msubsup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>a</mml:mi><mml:mo>′</mml:mo></mml:msubsup><mml:mo>=</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>b</mml:mi><mml:mo>′</mml:mo></mml:msubsup><mml:mo>-</mml:mo><mml:msup><mml:mfenced close="]" open="["><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>+</mml:mo><mml:msqrt><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msub><mml:mi>R</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mtext>var</mml:mtext><mml:mfenced open="(" close=")"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mrow><mml:mi mathvariant="normal">e</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced><mml:mo>+</mml:mo><mml:msub><mml:mi>R</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mstyle></mml:msqrt><mml:mspace width="0.125em" linebreak="nobreak"/></mml:mrow></mml:mfenced><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mi mathvariant="normal">loc</mml:mi></mml:msub><mml:mo>∘</mml:mo><mml:mtext>cov</mml:mtext><mml:mfenced close=")" open="("><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>b</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mrow><mml:mi mathvariant="normal">e</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced></mml:mrow><mml:mrow><mml:mtext>var</mml:mtext><mml:mfenced open="(" close=")"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mrow><mml:mi mathvariant="normal">e</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced><mml:mo>+</mml:mo><mml:msub><mml:mi>R</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:mfenced close=")" open="("><mml:mrow><mml:msubsup><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mrow><mml:mi mathvariant="normal">e</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mo>′</mml:mo></mml:msubsup></mml:mrow></mml:mfenced><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula></p>
</sec>
<sec id="Ch1.S2.SS2">
  <label>2.2</label><title>Adaptive observation error inflation</title>
      <p id="d2e793">The AOEI was first introduced and systematically applied in the context of satellite radiance data assimilation for numerical weather prediction <xref ref-type="bibr" rid="bib1.bibx27" id="paren.21"/>. For the <inline-formula><mml:math id="M28" display="inline"><mml:mi>j</mml:mi></mml:math></inline-formula>th proxy, it operates as in Eq. (<xref ref-type="disp-formula" rid="Ch1.E5"/>) by inflating the observation error variance when the squared innovation – the difference between the observation and the model simulation – exceeds the ensemble spread of the model simulation:

            <disp-formula id="Ch1.E5" content-type="numbered"><label>5</label><mml:math id="M29" display="block"><mml:mrow><mml:msub><mml:mi>R</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mo movablelimits="false">max⁡</mml:mo><mml:mfenced open="{" close="}"><mml:mrow><mml:msubsup><mml:mi mathvariant="italic">σ</mml:mi><mml:mrow><mml:mi mathvariant="normal">o</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>,</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:msup><mml:mfenced close="]" open="["><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mrow><mml:mi mathvariant="normal">e</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>-</mml:mo><mml:msubsup><mml:mi mathvariant="italic">σ</mml:mi><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>e</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:mfenced><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

          where <inline-formula><mml:math id="M30" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="italic">σ</mml:mi><mml:mrow><mml:mi mathvariant="normal">o</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:math></inline-formula> is original observation error variance, <inline-formula><mml:math id="M31" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the <inline-formula><mml:math id="M32" display="inline"><mml:mi>j</mml:mi></mml:math></inline-formula>th proxy value, <inline-formula><mml:math id="M33" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mrow><mml:mi mathvariant="normal">e</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mover accent="true"><mml:mrow><mml:mi>H</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>b</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the prior estimate of the <inline-formula><mml:math id="M34" display="inline"><mml:mi>j</mml:mi></mml:math></inline-formula>th proxy from the ensemble mean, <inline-formula><mml:math id="M35" display="inline"><mml:mi>H</mml:mi></mml:math></inline-formula> is the observation forward operator, <inline-formula><mml:math id="M36" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mrow><mml:mi mathvariant="normal">e</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> is the innovation, and <inline-formula><mml:math id="M37" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="italic">σ</mml:mi><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mrow><mml:mi mathvariant="normal">e</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:math></inline-formula> is the ensemble spread in observation space.</p>
      <p id="d2e1024">The AOEI is designed to limit erroneous analysis increments where there are large representativeness errors in either the forecast model and/or from the observation itself. However, its aggressive adjustment strategy (inflation grows with the square of innovation) requires careful application in data-sparse regions or areas with uncertain observations to avoid amplifying noise.</p>
</sec>
<sec id="Ch1.S2.SS3">
  <label>2.3</label><title>Huber Robust Estimation</title>
      <p id="d2e1035">In this study, a new adaptive observation error inflation scheme based on the Huber robust estimation method <xref ref-type="bibr" rid="bib1.bibx15" id="paren.22"/> is introduced (hereafter called “HAOEI”), which aims to mitigate the undue influence of large outliers. This method computes a normalized innovation – the difference between the observation and the model forecast, normalized by the square root of the sum of the model forecast error variance and the baseline observational error variance – and compares it against a threshold <inline-formula><mml:math id="M38" display="inline"><mml:mi mathvariant="italic">δ</mml:mi></mml:math></inline-formula>. The baseline observational error variance is then inflated for values exceeding this threshold. For the <inline-formula><mml:math id="M39" display="inline"><mml:mi>j</mml:mi></mml:math></inline-formula>th proxy, the simplified formulation is as follows:

                <disp-formula specific-use="gather" content-type="numbered"><mml:math id="M40" display="block"><mml:mtable displaystyle="true"><mml:mlabeledtr id="Ch1.E6"><mml:mtd><mml:mtext>6</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:msub><mml:mi>r</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mrow><mml:mi mathvariant="normal">e</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:msqrt><mml:mrow><mml:msubsup><mml:mi mathvariant="italic">σ</mml:mi><mml:mrow><mml:mi mathvariant="normal">o</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>+</mml:mo><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mrow><mml:msubsup><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mrow><mml:mi mathvariant="normal">e</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:msub></mml:mrow></mml:msqrt></mml:mfrac></mml:mstyle><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr><mml:mlabeledtr id="Ch1.E7"><mml:mtd><mml:mtext>7</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:msub><mml:mi>R</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mfenced close="" open="{"><mml:mtable class="array" columnalign="left left"><mml:mtr><mml:mtd><mml:mrow><mml:msubsup><mml:mi mathvariant="italic">σ</mml:mi><mml:mrow><mml:mi mathvariant="normal">o</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>,</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mtext>for</mml:mtext><mml:mspace linebreak="nobreak" width="0.25em"/><mml:mspace width="0.25em" linebreak="nobreak"/><mml:mo>|</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>|</mml:mo><mml:mo>≤</mml:mo><mml:mi mathvariant="italic">δ</mml:mi><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:msubsup><mml:mi mathvariant="italic">σ</mml:mi><mml:mrow><mml:mi mathvariant="normal">o</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>⋅</mml:mo><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>|</mml:mo></mml:mrow><mml:mi mathvariant="italic">δ</mml:mi></mml:mfrac></mml:mstyle><mml:mo>,</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mtext>for</mml:mtext><mml:mspace width="0.25em" linebreak="nobreak"/><mml:mspace linebreak="nobreak" width="0.25em"/><mml:mo>|</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>|</mml:mo><mml:mo>&gt;</mml:mo><mml:mi mathvariant="italic">δ</mml:mi><mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mfenced></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula>

          where <inline-formula><mml:math id="M41" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="italic">σ</mml:mi><mml:mrow><mml:mi mathvariant="normal">o</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:math></inline-formula> is the uninflated observation error variance, <inline-formula><mml:math id="M42" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the <inline-formula><mml:math id="M43" display="inline"><mml:mi>j</mml:mi></mml:math></inline-formula>th proxy value, <inline-formula><mml:math id="M44" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mrow><mml:mi mathvariant="normal">e</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mover accent="true"><mml:mrow><mml:mi>H</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>b</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the model estimated proxy using the ensemble mean of climatological samples, <inline-formula><mml:math id="M45" display="inline"><mml:mi>H</mml:mi></mml:math></inline-formula> is the observation forward operator and <inline-formula><mml:math id="M46" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="italic">σ</mml:mi><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>e</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:math></inline-formula> is the ensemble spread in observation space. The threshold parameter <inline-formula><mml:math id="M47" display="inline"><mml:mi mathvariant="italic">δ</mml:mi></mml:math></inline-formula> defines the boundary between normal observations and potential outliers, typically chosen based on statistical significance levels (e.g., <inline-formula><mml:math id="M48" display="inline"><mml:mi mathvariant="italic">δ</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M49" display="inline"><mml:mo>≈</mml:mo></mml:math></inline-formula> 1.645 corresponds to the 90 % confidence interval of a standard normal distribution). Observations with normalized innovations within <inline-formula><mml:math id="M50" display="inline"><mml:mrow><mml:mo>±</mml:mo><mml:mi mathvariant="italic">δ</mml:mi></mml:mrow></mml:math></inline-formula> are considered statistically consistent with the forecast and retain their original error variance, while those beyond this range are treated as potential outliers and undergo error variance inflation.</p>
      <p id="d2e1376">Although <inline-formula><mml:math id="M51" display="inline"><mml:mi mathvariant="italic">δ</mml:mi></mml:math></inline-formula> introduces an additional degree of freedom, <inline-formula><mml:math id="M52" display="inline"><mml:mi mathvariant="italic">δ</mml:mi></mml:math></inline-formula> can be chosen based on statistical significance levels. For example, <inline-formula><mml:math id="M53" display="inline"><mml:mi mathvariant="italic">δ</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M54" display="inline"><mml:mo>≈</mml:mo></mml:math></inline-formula> 1.645 corresponds to the 90 % confidence interval of a standard normal distribution. This provides a theoretically grounded starting point, rather than purely empirical tuning. In practice, <inline-formula><mml:math id="M55" display="inline"><mml:mi mathvariant="italic">δ</mml:mi></mml:math></inline-formula> can be selected using cross-validation or calibrated against reference periods (e.g., the instrumental era) when available. Compared to AOEI that inflates the observation error variance quadratically with the innovation, which can over-penalize observations in regions with poor model priors and thus lead to local degradation, HAOEI addresses this by first normalizing the innovation by the combined forecast and observation uncertainty. Only when this normalized residual exceeds a predefined threshold <inline-formula><mml:math id="M56" display="inline"><mml:mi mathvariant="italic">δ</mml:mi></mml:math></inline-formula> (e.g., corresponding to a 90 % confidence interval) does it inflate the error variance, and the inflation is linear rather than quadratic. This follows the logic of Huber's estimationr: observations within the expected range are trusted, while those outside are down-weighted gradually, not abruptly or excessively. Therefore, the HAOEI applies a more moderate, linear inflation only after a threshold is crossed. This design makes it a robust stabilizer: it consistently improves performance across wider areas by avoiding extreme adjustments, even if it does not achieve the same peak correction in the most mismatched locations.</p>
</sec>
<sec id="Ch1.S2.SS4">
  <label>2.4</label><title>Adaptive localization</title>
      <p id="d2e1430">Due to the uneven spatial distribution of proxy data, localization is required to constrain their spurious influence on remote regions. Within a fixed localization radius PDA framework, a weight matrix was introduced by <xref ref-type="bibr" rid="bib1.bibx34" id="text.23"/> to adjust the covariance. This weight matrix is defined by the Gaspari-Cohn function <xref ref-type="bibr" rid="bib1.bibx9" id="paren.24"/>:

            <disp-formula id="Ch1.E8" content-type="numbered"><label>8</label><mml:math id="M57" display="block"><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mtext>GC</mml:mtext><mml:mo>(</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>L</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>

          where <inline-formula><mml:math id="M58" display="inline"><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> is the weight assigned to observation <inline-formula><mml:math id="M59" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> when updating grid point <inline-formula><mml:math id="M60" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M61" display="inline"><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> is the distance between grid point and observation, <inline-formula><mml:math id="M62" display="inline"><mml:mi>L</mml:mi></mml:math></inline-formula> is a specified cut-off (or localization) radius, and GC denotes the Gaspari-Cohn fifth-order polynomial:

            <disp-formula id="Ch1.E9" content-type="numbered"><label>9</label><mml:math id="M63" display="block"><mml:mrow><mml:mtable class="split" rowspacing="0.2ex" displaystyle="true" columnalign="right left"><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mtext>GC</mml:mtext><mml:mo>(</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo><mml:mo>=</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mfenced close="" open="{"><mml:mtable rowspacing="12pt 12pt" class="array" columnalign="left left"><mml:mtr><mml:mtd><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:mstyle displaystyle="false"><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mn mathvariant="normal">5</mml:mn><mml:mn mathvariant="normal">3</mml:mn></mml:mfrac></mml:mstyle></mml:mstyle><mml:msubsup><mml:mi>z</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>+</mml:mo><mml:mstyle displaystyle="false"><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mn mathvariant="normal">5</mml:mn><mml:mn mathvariant="normal">8</mml:mn></mml:mfrac></mml:mstyle></mml:mstyle><mml:msubsup><mml:mi>z</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mn mathvariant="normal">3</mml:mn></mml:msubsup><mml:mo>+</mml:mo><mml:mstyle displaystyle="false"><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:mfrac></mml:mstyle></mml:mstyle><mml:msubsup><mml:mi>z</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mn mathvariant="normal">4</mml:mn></mml:msubsup><mml:mo>-</mml:mo><mml:mstyle displaystyle="false"><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">4</mml:mn></mml:mfrac></mml:mstyle></mml:mstyle><mml:msubsup><mml:mi>z</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mn mathvariant="normal">5</mml:mn></mml:msubsup><mml:mo>,</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mn mathvariant="normal">0</mml:mn><mml:mo>≤</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>≤</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">2</mml:mn><mml:mn mathvariant="normal">3</mml:mn></mml:mfrac></mml:mstyle><mml:msubsup><mml:mi>z</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:mn mathvariant="normal">4</mml:mn><mml:mo>-</mml:mo><mml:mn mathvariant="normal">5</mml:mn><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">5</mml:mn><mml:mn mathvariant="normal">3</mml:mn></mml:mfrac></mml:mstyle><mml:msubsup><mml:mi>z</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>+</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">5</mml:mn><mml:mn mathvariant="normal">8</mml:mn></mml:mfrac></mml:mstyle><mml:msubsup><mml:mi>z</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mn mathvariant="normal">3</mml:mn></mml:msubsup><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:mfrac></mml:mstyle><mml:msubsup><mml:mi>z</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mn mathvariant="normal">4</mml:mn></mml:msubsup><mml:mo>+</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">12</mml:mn></mml:mfrac></mml:mstyle><mml:msubsup><mml:mi>z</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mn mathvariant="normal">5</mml:mn></mml:msubsup><mml:mo>,</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>&lt;</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>≤</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mfenced></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:math></disp-formula>

          where <inline-formula><mml:math id="M64" display="inline"><mml:mrow><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M65" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M66" display="inline"><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>L</mml:mi><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:mfrac></mml:mstyle></mml:math></inline-formula>.</p>
      <p id="d2e1867">This method uses a fixed localization radius with the GC function, which tapers covariances based solely on physical distance. However, this method has two major limitations when applied to PDA: (1) Fixed radius ignores varying data density: In data-dense regions, a large fixed radius unnecessarily smooths fine-scale variability and introduces spurious correlations. In data-sparse regions, a small radius would discard too much information, while a large radius risks including irrelevant remote proxies. An adaptive radius that responds to local observational density can better balance these competing needs. (2) Distance alone misses teleconnections: Important climate phenomena such as ENSO involve strong correlations over large distances. A purely distance-based cutoff (like GC) may either exclude these physically meaningful teleconnections (if the radius is too small) or include many spurious ones (if the radius is too large). Thus,  a two-component adaptive strategy is proposed: (1) a density-dependent localization radius to handle heterogeneous data coverage. (2) Correlation-based weighting to identify and preserve genuine teleconnections beyond the distance cutoff.</p>
      <p id="d2e1870">Observational density is estimated at each grid point using kernel density estimation (KDE) with a Gaussian kernel <xref ref-type="bibr" rid="bib1.bibx30" id="paren.25"/>. For a given grid point <inline-formula><mml:math id="M67" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, the density is computed as:

            <disp-formula id="Ch1.E10" content-type="numbered"><label>10</label><mml:math id="M68" display="block"><mml:mrow><mml:msup><mml:mi mathvariant="italic">ρ</mml:mi><mml:mo>*</mml:mo></mml:msup><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:mi>n</mml:mi><mml:msup><mml:mi>h</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:munderover><mml:mi>K</mml:mi><mml:mfenced open="(" close=")"><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mo>‖</mml:mo><mml:msub><mml:mi mathvariant="bold">o</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mi mathvariant="bold">o</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>‖</mml:mo></mml:mrow><mml:mi>h</mml:mi></mml:mfrac></mml:mstyle></mml:mfenced></mml:mrow></mml:math></disp-formula>

          where <inline-formula><mml:math id="M69" display="inline"><mml:mi>n</mml:mi></mml:math></inline-formula> is the total number of proxy observations, <inline-formula><mml:math id="M70" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">o</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> denotes the location of the grid point <inline-formula><mml:math id="M71" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M72" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">o</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> denotes the location of the <inline-formula><mml:math id="M73" display="inline"><mml:mi>j</mml:mi></mml:math></inline-formula>th observation, <inline-formula><mml:math id="M74" display="inline"><mml:mrow><mml:mo>‖</mml:mo><mml:mo>⋅</mml:mo><mml:mo>‖</mml:mo></mml:mrow></mml:math></inline-formula> denotes the euclidean distance between two points, <inline-formula><mml:math id="M75" display="inline"><mml:mi>h</mml:mi></mml:math></inline-formula> is the bandwidth parameter, and <inline-formula><mml:math id="M76" display="inline"><mml:mrow><mml:mi>K</mml:mi><mml:mo>(</mml:mo><mml:mo>⋅</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the Gaussian kernel function:

            <disp-formula id="Ch1.E11" content-type="numbered"><label>11</label><mml:math id="M77" display="block"><mml:mrow><mml:mi>K</mml:mi><mml:mo>(</mml:mo><mml:mi>u</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mi mathvariant="italic">π</mml:mi></mml:mrow></mml:mfrac></mml:mstyle><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mstyle scriptlevel="+1"><mml:mfrac><mml:mrow><mml:msup><mml:mi>u</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:mfrac></mml:mstyle></mml:mrow></mml:msup><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>

          The bandwidth <inline-formula><mml:math id="M78" display="inline"><mml:mi>h</mml:mi></mml:math></inline-formula> is determined using Scott's rule:

            <disp-formula id="Ch1.E12" content-type="numbered"><label>12</label><mml:math id="M79" display="block"><mml:mrow><mml:mi>h</mml:mi><mml:mo>=</mml:mo><mml:msup><mml:mi>n</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">6</mml:mn></mml:mrow></mml:msup><mml:mo>⋅</mml:mo><mml:mover accent="true"><mml:mi mathvariant="italic">σ</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

          where <inline-formula><mml:math id="M80" display="inline"><mml:mover accent="true"><mml:mi mathvariant="italic">σ</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover></mml:math></inline-formula> is the standard deviation of the observation locations (computed from their longitudes and latitudes). This choice provides a data-adaptive smoothing that balances bias and variance in density estimation, yielding a continuous density field. This density field is then normalized by its global maximum value to obtain <inline-formula><mml:math id="M81" display="inline"><mml:mrow><mml:mi mathvariant="italic">ρ</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>∈</mml:mo><mml:mo>[</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>, representing the relative data density at grid point <inline-formula><mml:math id="M82" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. Based on this, the localization radius <inline-formula><mml:math id="M83" display="inline"><mml:mrow><mml:mi>L</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> at grid point <inline-formula><mml:math id="M84" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is defined as a function of this normalized density.

            <disp-formula id="Ch1.E13" content-type="numbered"><label>13</label><mml:math id="M85" display="block"><mml:mrow><mml:mtable rowspacing="0.2ex" class="split" displaystyle="true" columnalign="right left"><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mi>L</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mo>=</mml:mo><mml:mfenced close="" open="{"><mml:mtable class="array" rowspacing="6pt" columnalign="left left"><mml:mtr><mml:mtd><mml:mrow><mml:mi>L</mml:mi><mml:mo>⋅</mml:mo><mml:mo mathsize="1.1em">(</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:mi mathvariant="italic">ρ</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msup><mml:mo>)</mml:mo><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:msup><mml:mo mathsize="1.1em">)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mtext>if</mml:mtext><mml:mspace width="0.25em" linebreak="nobreak"/><mml:mspace width="0.25em" linebreak="nobreak"/><mml:mi>L</mml:mi><mml:mo>⋅</mml:mo><mml:mo mathsize="1.1em">(</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:mi mathvariant="italic">ρ</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msup><mml:mo>)</mml:mo><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:msup><mml:mo mathsize="1.1em">)</mml:mo><mml:mo>≥</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mo>min⁡</mml:mo></mml:msub></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mo>min⁡</mml:mo></mml:msub><mml:mo>,</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mtext>otherwise</mml:mtext></mml:mtd></mml:mtr></mml:mtable></mml:mfenced></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:math></disp-formula>

          where <inline-formula><mml:math id="M86" display="inline"><mml:mi>L</mml:mi></mml:math></inline-formula> is the original fixed localization radius (attained in  observation-sparse regions where <inline-formula><mml:math id="M87" display="inline"><mml:mrow><mml:mi mathvariant="italic">ρ</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M88" display="inline"><mml:mo>→</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M89" display="inline"><mml:mn mathvariant="normal">0</mml:mn></mml:math></inline-formula>), the exponent <inline-formula><mml:math id="M90" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula> controls the curvature of the density-to-radius mapping, and <inline-formula><mml:math id="M91" display="inline"><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mo>min⁡</mml:mo></mml:msub></mml:mrow></mml:math></inline-formula> is a lower bound enforced to maintain a minimum influence range even in observation-dense regions.</p>
      <p id="d2e2387">However, adjusting the localization radius solely based on density may further reduce the number of effective observations available for some grid points. While this helps suppress spurious teleconnections, it can also limit the use of already sparse paleoclimate proxy data, potentially undermining assimilation performance.</p>
      <p id="d2e2391">To better balance the use of observational information and the suppression of spurious correlations, we further incorporate correlation information into the weighting strategy, i.e., for observations outside the localization radius, their influence weight depends on the correlation (<inline-formula><mml:math id="M92" display="inline"><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M93" display="inline"><mml:mo>∼</mml:mo></mml:math></inline-formula> corr). The correlation coefficient <inline-formula><mml:math id="M94" display="inline"><mml:mrow><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> between grid point <inline-formula><mml:math id="M95" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and observation <inline-formula><mml:math id="M96" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is computed as the Pearson correlation coefficient between their respective time series over the reconstruction period:

            <disp-formula id="Ch1.E14" content-type="numbered"><label>14</label><mml:math id="M97" display="block"><mml:mrow><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msubsup><mml:mo>∑</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>T</mml:mi></mml:msubsup><mml:mfenced open="(" close=")"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:mfenced><mml:mfenced close=")" open="("><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:mfenced></mml:mrow><mml:msqrt><mml:mrow><mml:msubsup><mml:mo>∑</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>T</mml:mi></mml:msubsup><mml:msup><mml:mfenced open="(" close=")"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:msubsup><mml:mo>∑</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>T</mml:mi></mml:msubsup><mml:msup><mml:mfenced open="(" close=")"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:msqrt></mml:mfrac></mml:mstyle><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

          where <inline-formula><mml:math id="M98" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> is the value at grid point <inline-formula><mml:math id="M99" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> at time <inline-formula><mml:math id="M100" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M101" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> is the value at observation <inline-formula><mml:math id="M102" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> at time <inline-formula><mml:math id="M103" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M104" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M105" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> are the temporal means of the respective series, and <inline-formula><mml:math id="M106" display="inline"><mml:mi>T</mml:mi></mml:math></inline-formula> is the length of the time series. This coefficient quantifies the statistical relationship between distant locations and helps identify teleconnections that are physically meaningful rather than spurious.</p>
      <p id="d2e2709">The final weight <inline-formula><mml:math id="M107" display="inline"><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> assigned to observation <inline-formula><mml:math id="M108" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> when updating grid point <inline-formula><mml:math id="M109" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is then defined as a hybrid of distance-based and correlation-based weighting:

            <disp-formula id="Ch1.E15" content-type="numbered"><label>15</label><mml:math id="M110" display="block"><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfenced close="" open="{"><mml:mtable class="array" rowspacing="8pt" columnalign="left left"><mml:mtr><mml:mtd><mml:mrow><mml:mstyle displaystyle="false"><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:mfrac></mml:mstyle></mml:mstyle><mml:mo>⋅</mml:mo><mml:mo>|</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mtext>if</mml:mtext><mml:mspace width="0.25em" linebreak="nobreak"/><mml:mspace linebreak="nobreak" width="0.25em"/><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&gt;</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mspace width="0.25em" linebreak="nobreak"/><mml:mspace linebreak="nobreak" width="0.25em"/><mml:mtext>and</mml:mtext><mml:mspace linebreak="nobreak" width="0.25em"/><mml:mspace width="0.25em" linebreak="nobreak"/><mml:mo>|</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">0.2</mml:mn></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mtext>GC</mml:mtext><mml:mo>(</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mtext>else</mml:mtext></mml:mtd></mml:mtr></mml:mtable></mml:mfenced></mml:mrow></mml:math></disp-formula>

          where <inline-formula><mml:math id="M111" display="inline"><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> is the distance between grid point <inline-formula><mml:math id="M112" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and observation <inline-formula><mml:math id="M113" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M114" display="inline"><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M115" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M116" display="inline"><mml:mrow><mml:mi>L</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the adaptive localization radius at grid point <inline-formula><mml:math id="M117" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> computed from Eq. (<xref ref-type="disp-formula" rid="Ch1.E13"/>), <inline-formula><mml:math id="M118" display="inline"><mml:mrow><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> is the correlation coefficient, 0.2 is the set expression threshold.</p>
      <p id="d2e2965">Figure <xref ref-type="fig" rid="F1"/> illustrates the flowchart of this method. Compared to the standard fixed-radius GC function, the new method is superior because it is spatially adaptive and information-aware. The GC applies the same tapering everywhere regardless of how many proxies are available, which is suboptimal for heterogeneous networks. The KDE-based radius shrinks in data-rich areas to preserve local details and expands in data-poor areas to gather more information. The correlation-based component further adds value by recovering remote signals that GC would completely zero out. Compared to other existing adaptive localization methods <xref ref-type="bibr" rid="bib1.bibx2 bib1.bibx3 bib1.bibx4 bib1.bibx25 bib1.bibx26" id="paren.26"><named-content content-type="pre">e.g.,</named-content></xref>, this approach is specifically tailored to the challenges of the PDA. Many of those methods were developed for dense observing networks (e.g., satellite radiances or radiosondes). The new method preserves the stability of distance-based localization within the adaptive radius while selectively adding distant, highly correlated observations – a feature not present in most existing adaptive schemes. It is specifically designed for sparse, unevenly distributed proxy networks and directly addresses the two key weaknesses of fixed-radius GC (ignoring data density and teleconnections).</p>

      <fig id="F1"><label>Figure 1</label><caption><p id="d2e2977">Flowchart of the adaptive localization method.</p></caption>
          <graphic xlink:href="https://gmd.copernicus.org/articles/19/7893/2026/gmd-19-7893-2026-f01.png"/>

        </fig>

      <p id="d2e2986">Finally, it is worth noting that <inline-formula><mml:math id="M119" display="inline"><mml:mi>h</mml:mi></mml:math></inline-formula> is a parameter in the mathematical sense, which is determined uniquely from the data and does not introduce additional user-specific uncertainty. <inline-formula><mml:math id="M120" display="inline"><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mo>min⁡</mml:mo></mml:msub></mml:mrow></mml:math></inline-formula> is physically meaningful and easy to set. The minimum localization radius <inline-formula><mml:math id="M121" display="inline"><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mo>min⁡</mml:mo></mml:msub></mml:mrow></mml:math></inline-formula> is introduced to prevent the radius from becoming zero in very data-dense regions. Even in the densest proxy network, a zero radius would discard all observations, which is clearly undesirable. <inline-formula><mml:math id="M122" display="inline"><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mo>min⁡</mml:mo></mml:msub></mml:mrow></mml:math></inline-formula> does not rely on an exquisitely precise value but it should be set to a reasonably small value that can be guided by the typical decorrelation length scale of the climate variables.</p>
</sec>
</sec>
<sec id="Ch1.S3">
  <label>3</label><title>Experimental design and results</title>
<sec id="Ch1.S3.SS1">
  <label>3.1</label><title>Experimental design</title>
      <p id="d2e3045">To evaluate the effectiveness of adaptive inflation and adaptive localization methods, a set of sensitivity experiments (Table <xref ref-type="table" rid="T1"/>) is conducted using the cfr framework <xref ref-type="bibr" rid="bib1.bibx44" id="paren.27"><named-content content-type="pre">v2024.1.26;</named-content></xref>, which is a Python-based package that provides a suite of ensemble-based data assimilation and reconstruction tools, including EnKF implementation, proxy system models, and verification diagnostics. It is specifically designed for paleoclimate applications and is openly available. The goal is to reconstruct tropical (30° S–30° N) surface temperature anomalies from 1880 to 2000, with respect to the 1951–1980 climatological baseline. The prior ensemble is drawn from the “iCESM1” last millennium and historical simulations <xref ref-type="bibr" rid="bib1.bibx5" id="paren.28"/>, which provide physically consistent climate backgrounds spanning 850–2000 CE. Coral-based proxy records from the PAGES 2k Phase 2 database <xref ref-type="bibr" rid="bib1.bibx29" id="paren.29"/> – calibrated against NASA GISTEMP v4 <xref ref-type="bibr" rid="bib1.bibx19" id="paren.30"/> – are assimilated. These records are selected due to their strong temperature sensitivity and predominant distribution within the tropics (Fig. <xref ref-type="fig" rid="F2"/>). Focusing on the tropics is motivated by two main reasons: first, the high temperature sensitivity of coral proxies enhances reconstruction reliability, whereas incorporating other proxy types (e.g., tree rings, ice cores) at a global scale could increase uncertainty due to their generally lower temperature sensitivity. Second, the tropics contain key climate systems such as ENSO, frequently studied in paleoclimatology, allowing for a clear assessment of methodological performance while reducing computational costs. It is important to note that the methodology presented here can, in principle, be extended to global applications. In the following, the impact of adaptive observation error is first assessed by comparing two inflation schemes – AOEI and HAOEI. Next, sensitivity experiments investigate the role of localization. Initial tests use fixed localization radii, which are then compared to two experiments employing the adaptive localization method described in Sect. <xref ref-type="sec" rid="Ch1.S2.SS4"/> – one estimating proxy-state correlations based on reanalysis data GISTEMP v4, and the other one based on the model priors from “ICESM1” last millennium simulations. All data assimilation experiments cover 1880–2000, use a 1-year assimilation window, and an ensemble size of 100.</p>

<table-wrap id="T1" specific-use="star"><label>Table 1</label><caption><p id="d2e3072">Experimental setups, which are divided into two groups, sensitivity experiments on observation error inflation and on localization.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="5">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="left"/>
     <oasis:colspec colnum="3" colname="col3" align="left"/>
     <oasis:colspec colnum="4" colname="col4" align="left"/>
     <oasis:colspec colnum="5" colname="col5" align="left"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Group</oasis:entry>
         <oasis:entry colname="col2">EXP</oasis:entry>
         <oasis:entry colname="col3">Localization</oasis:entry>
         <oasis:entry colname="col4">Observation error inflation</oasis:entry>
         <oasis:entry colname="col5">Estimation of correlations in Eq. (<xref ref-type="disp-formula" rid="Ch1.E14"/>)</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry namest="col1" nameend="col5">Adaptive observation error inflation </oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">E_20000</oasis:entry>
         <oasis:entry colname="col3">20 000 km</oasis:entry>
         <oasis:entry colname="col4">No</oasis:entry>
         <oasis:entry colname="col5">–</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">E_AOEI</oasis:entry>
         <oasis:entry colname="col3">20 000 km</oasis:entry>
         <oasis:entry colname="col4">AOEI</oasis:entry>
         <oasis:entry colname="col5">–</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">E_HAOEI0.67</oasis:entry>
         <oasis:entry colname="col3">20 000 km</oasis:entry>
         <oasis:entry colname="col4">HAOEI, <inline-formula><mml:math id="M123" display="inline"><mml:mi mathvariant="italic">δ</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M124" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 0.67</oasis:entry>
         <oasis:entry colname="col5">–</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">E_HAOEI0.38</oasis:entry>
         <oasis:entry colname="col3">20 000 km</oasis:entry>
         <oasis:entry colname="col4">HAOEI, <inline-formula><mml:math id="M125" display="inline"><mml:mi mathvariant="italic">δ</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M126" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 0.38</oasis:entry>
         <oasis:entry colname="col5">–</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">E_HAOEI0.13</oasis:entry>
         <oasis:entry colname="col3">20 000 km</oasis:entry>
         <oasis:entry colname="col4">HAOEI, <inline-formula><mml:math id="M127" display="inline"><mml:mi mathvariant="italic">δ</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M128" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 0.13</oasis:entry>
         <oasis:entry colname="col5">–</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry namest="col1" nameend="col5">Adaptive localization </oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">E_20000</oasis:entry>
         <oasis:entry colname="col3">20 000 km</oasis:entry>
         <oasis:entry colname="col4">No</oasis:entry>
         <oasis:entry colname="col5">–</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">E_10000</oasis:entry>
         <oasis:entry colname="col3">10 000 km</oasis:entry>
         <oasis:entry colname="col4">No</oasis:entry>
         <oasis:entry colname="col5">–</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">E_5000</oasis:entry>
         <oasis:entry colname="col3">5000 km</oasis:entry>
         <oasis:entry colname="col4">No</oasis:entry>
         <oasis:entry colname="col5">–</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">E_1000</oasis:entry>
         <oasis:entry colname="col3">1000 km</oasis:entry>
         <oasis:entry colname="col4">No</oasis:entry>
         <oasis:entry colname="col5">–</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">E_AL1</oasis:entry>
         <oasis:entry colname="col3">(Adaptive)</oasis:entry>
         <oasis:entry colname="col4">No</oasis:entry>
         <oasis:entry colname="col5">Reanalyses</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">E_AL2</oasis:entry>
         <oasis:entry colname="col3">(Adaptive)</oasis:entry>
         <oasis:entry colname="col4">No</oasis:entry>
         <oasis:entry colname="col5">Model priors</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">E_HAOEI0.38_AL1</oasis:entry>
         <oasis:entry colname="col3">(Adaptive)</oasis:entry>
         <oasis:entry colname="col4">HAOEI, <inline-formula><mml:math id="M129" display="inline"><mml:mi mathvariant="italic">δ</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M130" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 0.38</oasis:entry>
         <oasis:entry colname="col5">Reanalyses</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <fig id="F2" specific-use="star"><label>Figure 2</label><caption><p id="d2e3393">Spatial and temporal distributions of proxy records from the PAGES 2k dataset, visualized with cfr <xref ref-type="bibr" rid="bib1.bibx44" id="paren.31"/>.</p></caption>
          <graphic xlink:href="https://gmd.copernicus.org/articles/19/7893/2026/gmd-19-7893-2026-f02.png"/>

        </fig>

      <p id="d2e3406">Three metrics are computed to evaluate the reconstruction analyses: the Root Mean Square Error (RMSE), the Coefficient of Efficiency (CE), and the Empirical Orthogonal Function (EOF). The RMSE for a given model grid point <inline-formula><mml:math id="M131" display="inline"><mml:mi>q</mml:mi></mml:math></inline-formula> is defined as:

            <disp-formula id="Ch1.E16" content-type="numbered"><label>16</label><mml:math id="M132" display="block"><mml:mrow><mml:mtext>RMSE</mml:mtext><mml:mo>(</mml:mo><mml:mi>q</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:msqrt><mml:mrow><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mi>n</mml:mi></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:munderover><mml:msup><mml:mfenced close=")" open="("><mml:mrow><mml:msub><mml:mi mathvariant="bold">v</mml:mi><mml:mrow><mml:mi>q</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold">x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mrow><mml:mi>q</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:msqrt><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

          where <inline-formula><mml:math id="M133" display="inline"><mml:mi>n</mml:mi></mml:math></inline-formula> is the number of time samples, <inline-formula><mml:math id="M134" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">v</mml:mi><mml:mrow><mml:mi>q</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> is the  verification data at grid point <inline-formula><mml:math id="M135" display="inline"><mml:mi>q</mml:mi></mml:math></inline-formula> and time <inline-formula><mml:math id="M136" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>, and <inline-formula><mml:math id="M137" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold">x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mrow><mml:mi>q</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> is the corresponding prior or posterior ensemble mean.</p>
      <p id="d2e3541">The CE evaluates the skill of the reconstruction relative to the climatological mean of the verification data. For a given model grid point <inline-formula><mml:math id="M138" display="inline"><mml:mi>q</mml:mi></mml:math></inline-formula>, it is calculated as:

            <disp-formula id="Ch1.E17" content-type="numbered"><label>17</label><mml:math id="M139" display="block"><mml:mrow><mml:mtext>CE</mml:mtext><mml:mo>(</mml:mo><mml:mi>q</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msubsup><mml:mo>∑</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:msubsup><mml:msup><mml:mfenced open="(" close=")"><mml:mrow><mml:msub><mml:mi mathvariant="bold">v</mml:mi><mml:mrow><mml:mi>q</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold">x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mrow><mml:mi>q</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:msubsup><mml:mo>∑</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:msubsup><mml:msup><mml:mfenced close=")" open="("><mml:mrow><mml:msub><mml:mi mathvariant="bold">v</mml:mi><mml:mrow><mml:mi>q</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold">v</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mi>q</mml:mi></mml:msub></mml:mrow></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

          where <inline-formula><mml:math id="M140" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold">v</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mi>q</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the temporal mean of the verification data at grid point <inline-formula><mml:math id="M141" display="inline"><mml:mi>q</mml:mi></mml:math></inline-formula>. A CE value close to 1.0 indicates a highly accurate reconstruction, a value of 0 means the reconstruction is only as accurate as the climatological mean, and negative values suggest lower skill.</p>
      <p id="d2e3677">Beyond pointwise accuracy metrics, we also examine the ability of the reconstructions to capture the dominant large-scale modes of climate variability. For this purpose, we apply Empirical Orthogonal Function (EOF) analysis <xref ref-type="bibr" rid="bib1.bibx21" id="paren.32"/> to the reconstructed spatial fields The EOF method decomposes a spatiotemporal field <inline-formula><mml:math id="M142" display="inline"><mml:mrow><mml:mi mathvariant="bold">X</mml:mi><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> into a set of orthogonal spatial patterns (EOFs) and their associated principal component (PC) time series. The leading EOF mode explains the largest possible fraction of the total variance, with each subsequent mode explaining the maximum remaining variance under the constraint of orthogonality to all previous modes. The corresponding PC represent the temporal evolution of these spatial patterns. In the following analysis, we compare the leading EOF modes – their spatial patterns, explained variance – obtained from different reconstruction methods against those derived from the verification dataset, to assess which method better recovers the key climate teleconnections and oscillatory signals.</p>
      <p id="d2e3701">As validation data, the spatially completed version of the near-surface air temperature and sea-surface temperature analyses product HadCRUT4.6 <xref ref-type="bibr" rid="bib1.bibx36" id="paren.33"/> is employed.</p>
</sec>
<sec id="Ch1.S3.SS2">
  <label>3.2</label><title>Results</title>
<sec id="Ch1.S3.SS2.SSS1">
  <label>3.2.1</label><title>Adaptive observation error inflation</title>
      <p id="d2e3723">Figure <xref ref-type="fig" rid="F3"/> shows the RMSE for the surface temperature analyses in experiment E_20000, as well as the RMSE and percentage RMSE differences between E_20000 and E_AOEI, E_HAOEI(0.67), E_HAOEI(0.38), and E_HAOEI(0.13). The largest RMSE values in E_20000 are located in the equatorial ENSO region. E_AOEI produces a noticeable reduction in RMSE relative to E_20000 (mean difference <inline-formula><mml:math id="M143" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M144" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.0195 °C, approximately 4.5 %). This improvement stems primarily from its aggressive error-inflation strategy, which strongly down-weights observations that deviate sharply from the model prior-most effectively in regions such as the ENSO zone where model biases are large. However, the same mechanism can also suppress accurate observations where the prior is poor, leading to local error increases, as seen over the North Pacific and Atlantic.The HAOEI method applies a more moderate, threshold-dependent inflation. E_HAOEI(0.67) yields widespread and modest improvements, with a mean RMSE difference of <inline-formula><mml:math id="M145" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.0210 °C (a reduction of 4.8 %). E_HAOEI(0.38) delivers a stronger overall error reduction (mean difference <inline-formula><mml:math id="M146" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M147" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.0253 °C, a reduction of approximately 5.8 %) and outperforming both E_AOEI and E_HAOEI(0.67). However, E_HAOEI(0.13) still achieves some improvement (RMSE difference <inline-formula><mml:math id="M148" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M149" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.0208 °C, a reduction of 4.7 %), its performance is inferior to that of E_HAOEI(0.38). Similar performance patterns are reflected in the CE results (Fig. <xref ref-type="fig" rid="F4"/>). While E_AOEI yields considerably higher CE values than E_20000 (0.1201 compared to <inline-formula><mml:math id="M150" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.2106), tuning the <inline-formula><mml:math id="M151" display="inline"><mml:mi mathvariant="italic">δ</mml:mi></mml:math></inline-formula> parameter can adjust the skill, the mean CE values in  E_HAOEI(0.67), E_HAOEI(0.38) and E_HAOEI(0.13) are 0.1275, 0.1514 and 0.1297, respectively, demonstrating a comprehensively better performance of E_HAOEI(0.38) overall. Hence, the choice of the threshold <inline-formula><mml:math id="M152" display="inline"><mml:mi mathvariant="italic">δ</mml:mi></mml:math></inline-formula> in HAOEI involves a fundamental trade-off: a value that is too small (e.g., 0.13) makes the method overly sensitive, flagging even moderate model-observation mismatches as outliers and down-weighting them, which leads to excessive loss of useful information and degraded analysis quality; conversely, a value that is too large (e.g., 0.67) makes the method overly conservative, affecting only the most extreme observations and failing to adequately control many genuinely large errors, thereby yielding only limited improvements. The optimal value (e.g., 0.38) strikes the best balance – maximizing the correction of harmful errors while minimizing the unnecessary loss of informative observations.</p>

      <fig id="F3" specific-use="star"><label>Figure 3</label><caption><p id="d2e3804"><bold>(a)</bold> The RMSE for the surface temperature analyses in E_20000, <bold>(b)</bold> the RMSE differences and the percentage RMSE differences between E_AOEI and E_20000, <bold>(c)</bold> between E_HAOEI0.67 and E_20000, <bold>(d)</bold> between E_HAOEI0.38 and E_20000, and <bold>(e)</bold> between E_HAOEI0.13 and E_20000.</p></caption>
            <graphic xlink:href="https://gmd.copernicus.org/articles/19/7893/2026/gmd-19-7893-2026-f03.png"/>

          </fig>

      <fig id="F4" specific-use="star"><label>Figure 4</label><caption><p id="d2e3829">Same as Fig. <xref ref-type="fig" rid="F3"/> but for the CE.</p></caption>
            <graphic xlink:href="https://gmd.copernicus.org/articles/19/7893/2026/gmd-19-7893-2026-f04.png"/>

          </fig>

      <p id="d2e3841">Figure <xref ref-type="fig" rid="F5"/> compares the spatial pattern and PC time series of the leading EOF mode derived from the observations with those from the reconstructed analyses of E_AOEI, E_HAOEI(<inline-formula><mml:math id="M153" display="inline"><mml:mi mathvariant="italic">δ</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M154" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 0.67), E_HAOEI(<inline-formula><mml:math id="M155" display="inline"><mml:mi mathvariant="italic">δ</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M156" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 0.38) and E_HAOEI(<inline-formula><mml:math id="M157" display="inline"><mml:mi mathvariant="italic">δ</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M158" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 0.13). All experiments display a similar spatial structure to the observations, yet they show markedly stronger variability in the ENSO region and produce a substantially higher explained variance. This discrepancy suggests that within the current assimilation framework, the reconstructed temperature field may be overly focused on a single dominant spatial mode, thereby failing to capture the more dispersed, multi-scale variability present in the actual observations. This effect may be partly due to the use of an excessively large localization radius, a point that will be further discussed in the following section.</p>

      <fig id="F5" specific-use="star"><label>Figure 5</label><caption><p id="d2e3891">The leading EOF mode of <bold>(a)</bold> observations <bold>(b)</bold> E_AOEI analyses <bold>(c)</bold> E_HAOEI (<inline-formula><mml:math id="M159" display="inline"><mml:mrow><mml:mi mathvariant="italic">δ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.67</mml:mn></mml:mrow></mml:math></inline-formula>) analyses <bold>(d)</bold> E_HAOEI (<inline-formula><mml:math id="M160" display="inline"><mml:mi mathvariant="italic">δ</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M161" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 0.38) analyses <bold>(e)</bold> E_HAOEI (<inline-formula><mml:math id="M162" display="inline"><mml:mi mathvariant="italic">δ</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M163" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 0.13) analyses <bold>(f)</bold> corresponding PC time series for all experiments.</p></caption>
            <graphic xlink:href="https://gmd.copernicus.org/articles/19/7893/2026/gmd-19-7893-2026-f05.png"/>

          </fig>

</sec>
<sec id="Ch1.S3.SS2.SSS2">
  <label>3.2.2</label><title>Adaptive localization</title>
      <p id="d2e3967">Figure <xref ref-type="fig" rid="F6"/> presents the time series of RMSE for surface temperature analyses across experiments using different fixed localization radii (1000–20 000 km). As the localization radius grows, the RMSE decreases from 0.4377 in E_20000 to 0.4234 in E_5000, and increases again to 0.4307 in E_1000. This behavior is consistent with other studies <xref ref-type="bibr" rid="bib1.bibx41" id="paren.34"><named-content content-type="pre">e.g.,</named-content></xref>. More importantly, regarding the leading EOF mode as shown in Fig. <xref ref-type="fig" rid="F7"/>, smaller radii significantly distort the spatial pattern of variability, although they  explained variance values closer to those of the observations (66.7 % in E_20000 and 27.3 % in E_1000 compared to 27.4 % in observations). This seemingly paradoxical behavior-where RMSE first decreases and then increases as the localization radius shrinks, while the EOF spatial pattern becomes progressively distorted-reflects the fundamental trade-off between sampling error and the retention of physically meaningful large-scale teleconnections in ensemble-based data assimilation. A large radius (e.g., 20 000 km) allows each grid point to be influenced by distant observations, preserving coherent large-scale structures (such as the ENSO pattern) and yielding a spatially reasonable EOF mode, but it also introduces spurious long-distance  correlations from sampling noise, which inflates both the RMSE and the explained variance of the leading EOF. Reducing the radius to an intermediate value (e.g., 5000 km) effectively suppresses many of these spurious correlations, leading to a lower RMSE and a more physical analysis. However, when the radius becomes too small (e.g., 1000 km), the assimilation overly relies on local observations, cutting off genuine large-scale connections and fragmenting the spatial structure; although this yields a notably lower RMSE and an explained variance that matches the observations, the leading EOF mode becomes spatially distorted and physically unrealistic-demonstrating that numerical metrics like RMSE and variance explained are insufficient indicators of physical fidelity.</p>

      <fig id="F6" specific-use="star"><label>Figure 6</label><caption><p id="d2e3981">The time series of the RMSE for the surface temperature analyses in experiments with varying fixed localization radii.</p></caption>
            <graphic xlink:href="https://gmd.copernicus.org/articles/19/7893/2026/gmd-19-7893-2026-f06.png"/>

          </fig>

      <fig id="F7" specific-use="star"><label>Figure 7</label><caption><p id="d2e3992">The leading EOF mode of the tropics: <bold>(a)</bold> reconstructed EOF field with 20 000 km localization radius; <bold>(b)</bold> reconstructed EOF field with 10 000 km localization radius; <bold>(c)</bold> reconstructed EOF field with 5000 km localization radius; <bold>(d)</bold> reconstructed EOF field with 1000 km localization radius; <bold>(e)</bold> corresponding PC time series for all experiments. All results are based on the period 1880–2000. EOF units are °C.</p></caption>
            <graphic xlink:href="https://gmd.copernicus.org/articles/19/7893/2026/gmd-19-7893-2026-f07.png"/>

          </fig>

      <p id="d2e4017">To implement the adaptive localization method described in Sect. <xref ref-type="sec" rid="Ch1.S2.SS4"/>, the spatial density of proxy records is first estimated using kernel density estimation (KDE), as visualized in Fig. <xref ref-type="fig" rid="F8"/>. While this density varies from year to year, the time-averaged distribution reveals distinct regional patterns: Australia and Central America are data-rich, the equatorial ENSO region exhibits medium proxy density, and areas such as South America remain relatively data-sparse. Furthermore, proxy-state correlations given in Eq. (<xref ref-type="disp-formula" rid="Ch1.E14"/>) are given by using the reanalyses and model priors, respectively. Figure <xref ref-type="fig" rid="F9"/> illustrates the correlations for one representative proxy record (ID: Ocn_090). Both estimated correlations reflect similar spatial structures but with detailed differences, e.g., the former one shows generally higher values and distinct discrepancies over Africa and the adjacent Indian Ocean region.</p>

      <fig id="F8" specific-use="star"><label>Figure 8</label><caption><p id="d2e4030">Gaussian kernel density estimation maps for different years: <bold>(a)</bold> 1880; <bold>(b)</bold> 1940; <bold>(c)</bold> 1990; <bold>(d)</bold> time averaged.</p></caption>
            <graphic xlink:href="https://gmd.copernicus.org/articles/19/7893/2026/gmd-19-7893-2026-f08.png"/>

          </fig>

      <fig id="F9" specific-use="star"><label>Figure 9</label><caption><p id="d2e4053">Spatial correlation maps between coral proxy data (ID: Ocn_090) located at 4.2° S, 144° E and <bold>(a)</bold> instrumental observations, <bold>(b)</bold> prior data.</p></caption>
            <graphic xlink:href="https://gmd.copernicus.org/articles/19/7893/2026/gmd-19-7893-2026-f09.png"/>

          </fig>

      <fig id="F10" specific-use="star"><label>Figure 10</label><caption><p id="d2e4070"><bold>(a)</bold> The RMSE differences and the percentage RMSE differences between E_AL1 and E_20000 <bold>(b)</bold> between E_HAOEI0.38_AL1 and E_HAOEI0.38.</p></caption>
            <graphic xlink:href="https://gmd.copernicus.org/articles/19/7893/2026/gmd-19-7893-2026-f10.png"/>

          </fig>

      <p id="d2e4085">As shown in Figs. <xref ref-type="fig" rid="F10"/> and <xref ref-type="fig" rid="F11"/>, E_AL1 outperforms E_20000, yielding smaller RMSE (E_20000: 0.4377 °C; E_AL1: 0.4211 °C; mean difference <inline-formula><mml:math id="M164" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M165" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.0166 °C, approximately 3.8 %) and generally higher CE (E_20000: <inline-formula><mml:math id="M166" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.2106; E_AL1: <inline-formula><mml:math id="M167" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.1179; mean difference <inline-formula><mml:math id="M168" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 0.0927). The improvements scale with data density: moderate in data-rich regions (e.g., Australia, Central America), more substantial in medium-density areas like the equatorial ENSO region, and largely neutral in sparse-data regions such as South America. Similarly, E_HAOEI0.38_AL1 achieves lower RMSE (0.4166 °C) and higher CE (0.1638) than its adaptive-localization counterpart E_HAOEI0.38 (RMSE: 0.4233 °C; CE: 0.1514), indicating that the combination of HAOEI and adaptive localization yields further improvements. Furthermore, Fig. <xref ref-type="fig" rid="F12"/> reveals that the leading EOF modes of E_AL1 and E_HAOEI0.38_AL1 exhibit a spatial structure similar to observations but with amplified variability in the ENSO region. Although their explained variance (E_AL1: 42.7 %; E_HAOEI0.38_AL1: <inline-formula><mml:math id="M169" display="inline"><mml:mo>∼</mml:mo></mml:math></inline-formula> 40 %) remains substantially higher than observed (27.4 %), it is considerably lower than that of E_20000 (66.7 %, Fig. <xref ref-type="fig" rid="F7"/>), indicating a more faithful representation of the spatial variance distribution. Overall, the application of adaptive localization is beneficial for enhancing reconstruction quality, both in terms of statistical accuracy and the fidelity of dominant climate modes.</p>

      <fig id="F11" specific-use="star"><label>Figure 11</label><caption><p id="d2e4141">Same as Fig. <xref ref-type="fig" rid="F10"/> but for the CE.</p></caption>
            <graphic xlink:href="https://gmd.copernicus.org/articles/19/7893/2026/gmd-19-7893-2026-f11.png"/>

          </fig>

      <fig id="F12" specific-use="star"><label>Figure 12</label><caption><p id="d2e4154">The leading EOF mode of <bold>(a)</bold> observations <bold>(b)</bold> adaptive localization analyses <bold>(c)</bold> adaptive localization and HAOEI (<inline-formula><mml:math id="M170" display="inline"><mml:mi mathvariant="italic">δ</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M171" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 0.38) analyses <bold>(d)</bold> corresponding PC time series for all experiments.</p></caption>
            <graphic xlink:href="https://gmd.copernicus.org/articles/19/7893/2026/gmd-19-7893-2026-f12.png"/>

          </fig>

      <fig id="F13" specific-use="star"><label>Figure 13</label><caption><p id="d2e4192"><bold>(a)</bold> The RMSE differences and the percentage RMSE differences between E_AL2 and E_20000 <bold>(b)</bold> between E_AL2 and E_AL1 <bold>(c)</bold> the CE differences between E_AL2 and E_20000 <bold>(d)</bold> between E_AL2 and E_AL1.</p></caption>
            <graphic xlink:href="https://gmd.copernicus.org/articles/19/7893/2026/gmd-19-7893-2026-f13.png"/>

          </fig>

      <p id="d2e4213">Finally, Fig. <xref ref-type="fig" rid="F13"/> illustrates that adaptive localization using model-prior correlations E_AL2 improves upon the fixed-radius baseline E_20000, as evidenced by its lower RMSEs and higher CEs, even if it does not match the performance of the reanalysis-driven version E_AL1. This demonstrates the intrinsic value of the model prior: its physically constrained correlations offer a viable and effective guide for adaptive localization in the absence of reanalysis data. Consequently, this approach ensures the method's applicability to pre-industrial and deeper-time paleoclimate periods, where observational constraints are otherwise unavailable.</p>
</sec>
</sec>
</sec>
<sec id="Ch1.S4" sec-type="conclusions">
  <label>4</label><title>Conclusion and outlook</title>
      <p id="d2e4228">This study has developed and systematically evaluated two adaptive strategies within an ensemble-based paleoclimate data assimilation framework: adaptive observation error inflation and adaptive covariance localization. Through sensitivity experiments using coral-based proxy records over the tropical region (1880–2000), we have quantified their impacts on reconstruction quality in terms of pointwise accuracy (RMSE and CE) and fidelity of large-scale climate modes (EOF analysis).</p>
      <p id="d2e4231">For adaptive observation error inflation, the aggressive AOEI method yields moderate improvements but introduces local degradations where accurate observations are mistakenly down-weighted. The HAOEI method, with its threshold-based linear inflation, provides more balanced and robust performance. Systematic experiments with three <inline-formula><mml:math id="M172" display="inline"><mml:mi mathvariant="italic">δ</mml:mi></mml:math></inline-formula> values (0.67, 0.38, and 0.13) reveal a clear non-monotonic behavior, with <inline-formula><mml:math id="M173" display="inline"><mml:mi mathvariant="italic">δ</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M174" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 0.38 achieving the largest RMSE reduction (approximately 5.8 %) and the highest CE among all inflation experiments. The existence of an optimal <inline-formula><mml:math id="M175" display="inline"><mml:mi mathvariant="italic">δ</mml:mi></mml:math></inline-formula> indicates a fundamental trade-off: a threshold that is too small over-penalizes moderately inconsistent observations, while a threshold that is too large fails to adequately control genuinely large errors. HAOEI with optimal tuning delivers spatially consistent improvements across most of the tropical domain with substantially fewer local degradations than AOEI.</p>
      <p id="d2e4262">For covariance localization, the fixed-radius experiments demonstrate that the localization radius critically controls the balance between sampling error and spatial coherence. A large radius preserves large-scale structures but introduces spurious correlations with over-amplified EOF variance, while a small radius distorts the spatial pattern of the leading mode despite yielding a coincidentally close variance percentage. Our proposed adaptive localization method addresses these limitations by dynamically adjusting the radius based on local data density and incorporating correlation-based weights. Compared to the fixed large-radius baseline, adaptive localization reduces RMSE by approximately 3.8 % and substantially improves CE. When combined with HAOEI, the full adaptive system achieves the best overall performance, with the lowest RMSE and highest CE among all experiments. Moreover, the leading EOF modes of the adaptive experiments exhibit spatial structures that closely resemble observations, with explained variances substantially lower than the fixed large-radius experiment and much closer to the observed value. Importantly, the adaptive method remains effective even when correlation information is derived solely from the model prior rather than reanalysis data, confirming its applicability to pre-industrial periods.</p>
      <p id="d2e4265">In summary, the optimal combination – HAOEI with <inline-formula><mml:math id="M176" display="inline"><mml:mi mathvariant="italic">δ</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M177" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 0.38 and adaptive localization – yields meaningful improvements in both statistical accuracy and the fidelity of dominant climate modes, reducing RMSE by approximately 5 %–6 % and substantially correcting the over-amplified EOF variance. These improvements, while modest in magnitude, are statistically robust, spatially consistent, and physically meaningful. They demonstrate that the adaptive strategies developed here offer a practical and effective pathway for improving paleoclimate reconstructions, particularly in data-sparse regions and for capturing large-scale climate variability modes such as ENSO.</p>
      <p id="d2e4283">While this study demonstrates the considerable promise of adaptive strategies in paleoclimate data assimilation, several important avenues merit further investigation to advance the methodology and expand its applications. For the adaptive localization, one could explore time-varying correlation estimates to account for non-stationary teleconnections. Moreover, this study focused on coral-based temperature reconstructions within the tropics; extending the framework to a global scale and to multiple proxy types and climate variables would provide a rigorous test of its generalizability and enable more holistic climate field reconstructions. The methods should be also tested during periods of abrupt climate change (e.g., the Last Glacial Termination, Dansgaard-Oeschger events) or past warm periods (e.g., the Last Interglacial, the Pliocene), where data constraints are particularly challenging but scientific stakes are high.</p>
</sec>

      
      </body>
    <back><notes notes-type="codedataavailability"><title>Code and data availability</title>

      <p id="d2e4291">The code and data that support the findings of this study are openly available. All input datasets actually used in this study (including the PAGES 2k Phase 2 global multiproxy database, the “iCESM1” last millennium simulation, the NASA Goddard's Global Surface Temperature Analysis, as well as the spatially completed version of the near-surface air temperature and sea-surface temperature analysis product HadCRUT4.6) are hosted within the cfr and can be accessed at <ext-link xlink:href="https://doi.org/10.5281/zenodo.10575537" ext-link-type="DOI">10.5281/zenodo.10575537</ext-link> <xref ref-type="bibr" rid="bib1.bibx43" id="paren.35"/>. The exact version of the code, the source code of the assimilation system, and the validation results of the reconstructed estimates used in this study are also archived in a trusted permanent repository at Zenodo under the following DOI: <ext-link xlink:href="https://doi.org/10.5281/zenodo.19015635" ext-link-type="DOI">10.5281/zenodo.19015635</ext-link> <xref ref-type="bibr" rid="bib1.bibx22" id="paren.36"/>.</p>
  </notes><notes notes-type="authorcontribution"><title>Author contributions</title>

      <p id="d2e4309">G. Luo conducted the experiments and wrote the first draft of the manuscript. Y. Zeng and F. Zhu proposed the ideas of adaptive observation error inflation and adaptive localization, and implemented them. J. Zhao contributed to the analysis of the results. All authors contributed to revising the text and defining the structure of the paper.</p>
  </notes><notes notes-type="competinginterests"><title>Competing interests</title>

      <p id="d2e4315">At least one of the (co-)authors is a member of the editorial board of <italic>Geoscientific Model Development</italic>. The peer-review process was guided by an independent editor, and the authors also have no other competing interests to declare.</p>
  </notes><notes notes-type="disclaimer"><title>Disclaimer</title>

      <p id="d2e4324">Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. The authors bear the ultimate responsibility for providing appropriate place names. Views expressed in the text are those of the authors and do not necessarily reflect the views of the publisher.</p>
  </notes><ack><title>Acknowledgements</title><p id="d2e4330">We acknowledge financial support from the National Key Research and Development Program of China.</p></ack><notes notes-type="financialsupport"><title>Financial support</title>

      <p id="d2e4335">This research has been supported by the National Key Research and Development Program of China (grant no. 2023YFF0804703).</p>
  </notes><notes notes-type="reviewstatement"><title>Review statement</title>

      <p id="d2e4341">This paper was edited by Benjamin Gaubert and reviewed by three anonymous referees.</p>
  </notes><ref-list>
    <title>References</title>

      <ref id="bib1.bibx1"><label>Anderson(2007)</label><mixed-citation> Anderson, J. L.: Exploring the need for localization in ensemble data  assimilation using a hierarchical ensemble filter, Physica D, 230, 99–111, 2007.</mixed-citation></ref>
      <ref id="bib1.bibx2"><label>Anderson(2012)</label><mixed-citation> Anderson, J. L.: Localization and sampling error correction in ensemble Kalman filter data assimilation, Mon. Weather Rev., 140, 2359–2371, 2012.</mixed-citation></ref>
      <ref id="bib1.bibx3"><label>Bishop and Hodyss(2009a)</label><mixed-citation> Bishop, C. and Hodyss, D.: Ensemble covariances adaptively localized with  ECO-RAP. Part 1: Tests on simple error models, Tellus A, 61, 84–96, 2009a.</mixed-citation></ref>
      <ref id="bib1.bibx4"><label>Bishop and Hodyss(2009b)</label><mixed-citation> Bishop, C. and Hodyss, D.: Ensemble covariances adaptively localized with  ECO-RAP. Part 2: A strategy for the atmosphere, Tellus A, 61, 97–111, 2009b.</mixed-citation></ref>
      <ref id="bib1.bibx5"><label>Brady et al.(2019)</label><mixed-citation> Brady, E., Stevenson, S., Bailey, D., Liu, Z., Noone, D., Nusbaumer, J.,  Otto-Bliesner, B. L., Tabor, C., Tomas, R., Wong, T., Zhang, J., and Zhu, J.: The connected  isotopic water cycle in the Community Earth System Model version 1, J. Adv. Model. Earth Sy., 11, 2547–2566, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx6"><label>Dee et al.(2015)</label><mixed-citation> Dee, S., Emile-Geay, J., Evans, M., Allam, A., Steig, E., and Thompson, D.:  PRYSM: An open-source framework for PRoxY System Modeling, with applications  to oxygen-isotope systems, J. Adv. Model. Earth Sy., 7, 1220–1247, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx7"><label>Evans et al.(2013)</label><mixed-citation> Evans, M. N., Tolwinski-Ward, S. E., Thompson, D. M., and Anchukaitis, K. J.:  Applications of proxy system modeling in high resolution paleoclimatology,  Quaternary Sci. Rev., 76, 16–28, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx8"><label>Evensen(1994)</label><mixed-citation> Evensen, G.: Sequential data assimilation with a nonlinear quasi-geostrophic  model using Monte Carlo methods to forecast error statistics, J. Geophys. Res.-Oceans, 99, 10143–10162, 1994.</mixed-citation></ref>
      <ref id="bib1.bibx9"><label>Gaspari and Cohn(1999)</label><mixed-citation> Gaspari, G. and Cohn, S. E.: Construction of correlation functions in two and  three dimensions, Q. J. Roy. Meteor. Soc., 125, 723–757, 1999.</mixed-citation></ref>
      <ref id="bib1.bibx10"><label>Goosse(2017)</label><mixed-citation> Goosse, H.: Reconstructed and simulated temperature asymmetry between  continents in both hemispheres over the last centuries, Clim. Dynam., 48,  1483–1501, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx11"><label>Guillot et al.(2015)</label><mixed-citation> Guillot, D., Rajaratnam, B., and Emile-Geay, J.: Statistical paleoclimate  reconstructions via Markov random fields, Ann. Appl. Stat., 9, 324–352, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx12"><label>Hakim et al.(2016)</label><mixed-citation> Hakim, G. J., Emile-Geay, J., Steig, E. J., Noone, D., Anderson, D. M., Tardif, R., Steiger, N., and Perkins, W. A.: The last millennium climate reanalysis project: Framework and first results, J. Geophys. Res.-Atmos., 121, 6745–6764, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx13"><label>Houtekamer and Mitchell(1998)</label><mixed-citation> Houtekamer, P. L. and Mitchell, H. L.: Data assimilation using an ensemble  Kalman filter technique, Mon. Weather Rev., 126, 796–811, 1998.</mixed-citation></ref>
      <ref id="bib1.bibx14"><label>Houtekamer and Mitchell(2001)</label><mixed-citation> Houtekamer, P. L. and Mitchell, H. L.: A sequential ensemble Kalman filter for atmospheric data assimilation, Mon. Weather Rev., 129, 123–137, 2001.</mixed-citation></ref>
      <ref id="bib1.bibx15"><label>Huber(1992)</label><mixed-citation>Huber, P. J.: Robust estimation of a location parameter, in: Breakthroughs in  statistics: Methodology and distribution, Springer, 492–518, <ext-link xlink:href="https://doi.org/10.1007/978-1-4612-4380-9_35" ext-link-type="DOI">10.1007/978-1-4612-4380-9_35</ext-link>, 1992.</mixed-citation></ref>
      <ref id="bib1.bibx16"><label>Kageyama et al.(2021)</label><mixed-citation>Kageyama, M., Harrison, S. P., Kapsch, M.-L., Lofverstrom, M., Lora, J. M., Mikolajewicz, U., Sherriff-Tadano, S., Vadsaria, T., Abe-Ouchi, A., Bouttes, N., Chandan, D., Gregoire, L. J., Ivanovic, R. F., Izumi, K., LeGrande, A. N., Lhardy, F., Lohmann, G., Morozova, P. A., Ohgaito, R., Paul, A., Peltier, W. R., Poulsen, C. J., Quiquet, A., Roche, D. M., Shi, X., Tierney, J. E., Valdes, P. J., Volodin, E., and Zhu, J.: The PMIP4 Last Glacial Maximum experiments: preliminary results and comparison with the PMIP3 simulations, Clim. Past, 17, 1065–1089, <ext-link xlink:href="https://doi.org/10.5194/cp-17-1065-2021" ext-link-type="DOI">10.5194/cp-17-1065-2021</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx17"><label>Lei and Anderson(2014a)</label><mixed-citation> Lei, L. and Anderson, J. L.: Comparisons of empirical localization techniques  for serial ensemble Kalman filters in a simple atmospheric general  circulation model, Mon. Weather Rev., 142, 739–754, 2014a.</mixed-citation></ref>
      <ref id="bib1.bibx18"><label>Lei and Anderson(2014b)</label><mixed-citation> Lei, L. and Anderson, J. L.: Empirical localization of observations for serial ensemble Kalman filter data assimilation in an atmospheric general  circulation model, Mon. Weather Rev., 142, 1835–1851, 2014b.</mixed-citation></ref>
      <ref id="bib1.bibx19"><label>Lenssen et al.(2019)</label><mixed-citation> Lenssen, N. J., Schmidt, G. A., Hansen, J. E., Menne, M. J., Persin, A., Ruedy, R., and Zyss, D.: Improvements in the GISTEMP uncertainty model, J.  Geophys. Res.-Atmos., 124, 6307–6326, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx20"><label>Liu et al.(2014)</label><mixed-citation> Liu, Z., Zhu, J., Rosenthal, Y., Zhang, X., Otto-Bliesner, B. L., Timmermann,  A., Smith, R. S., Lohmann, G., Zheng, W., and Elison Timm, O.: The Holocene  temperature conundrum, P. Natl. Acad. Sci. USA, 111, E3501–E3505, 2014.</mixed-citation></ref>
      <ref id="bib1.bibx21"><label>Lorenz(1956)</label><mixed-citation> Lorenz, E. N.: Empirical orthogonal functions and statistical weather  prediction, Vol. 1, Department of  Meteorology, Massachusetts Institute of Technology, Cambridge, 1956.</mixed-citation></ref>
      <ref id="bib1.bibx22"><label>Luo et al.(2026)</label><mixed-citation>Luo, G., Zeng, Y., Zhu, F., and Zhao, J.: Paleoclimate data assimilation with  adaptive observation error inflation and adaptive localization, Version v2, Zenodo [data set/code], <ext-link xlink:href="https://doi.org/10.5281/zenodo.19015635" ext-link-type="DOI">10.5281/zenodo.19015635</ext-link>, 2026.</mixed-citation></ref>
      <ref id="bib1.bibx23"><label>Mann et al.(1998)</label><mixed-citation> Mann, M. E., Bradley, R. S., and Hughes, M. K.: Global-scale temperature  patterns and climate forcing over the past six centuries, Nature, 392,  779–787, 1998.</mixed-citation></ref>
      <ref id="bib1.bibx24"><label>Mann et al.(2008)</label><mixed-citation> Mann, M. E., Zhang, Z., Hughes, M. K., Bradley, R. S., Miller, S. K.,  Rutherford, S., and Ni, F.: Proxy-based reconstructions of hemispheric and  global surface temperature variations over the past two millennia, P. Natl. Acad. Sci. USA, 105, 13252–13257, 2008.</mixed-citation></ref>
      <ref id="bib1.bibx25"><label>Ménétrier et al.(2015a)</label><mixed-citation> Ménétrier, B., Montmerle, T., Michel, Y., and Berre, L.: Linear  filtering of sample covariances for ensemble-based data assimilation. Part I:  Optimality criteria and application to variance filtering and covariance  localization, Mon. Weather Rev., 143, 1622–1643, 2015a.</mixed-citation></ref>
      <ref id="bib1.bibx26"><label>Ménétrier et al.(2015b)</label><mixed-citation> Ménétrier, B., Montmerle, T., Michel, Y., and Berre, L.: Linear  filtering of sample covariances for ensemble-based data assimilation. Part  II: Application to a convective-scale NWP model, Mon. Weather Rev., 143,  1644–1664, 2015b.</mixed-citation></ref>
      <ref id="bib1.bibx27"><label>Minamide and Zhang(2017)</label><mixed-citation> Minamide, M. and Zhang, F.: Adaptive observation error inflation for  assimilating all-sky satellite radiance, Mon. Weather Rev., 145, 1063–1081, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx28"><label>Osman et al.(2021)</label><mixed-citation> Osman, M. B., Tierney, J. E., Zhu, J., Tardif, R., Hakim, G. J., King, J., and Poulsen, C. J.: Globally resolved surface temperatures since the Last Glacial Maximum, Nature, 599, 239–244, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx29"><label>PAGES 2k Consortium(2017)</label><mixed-citation>PAGES 2k Consortium: A global multiproxy database for temperature  reconstructions of the Common Era, Scientific Data, 4, 170088,  <ext-link xlink:href="https://doi.org/10.1038/sdata.2017.88" ext-link-type="DOI">10.1038/sdata.2017.88</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx30"><label>Parzen(1962)</label><mixed-citation> Parzen, E.: On estimation of a probability density function and mode, Ann. Math. Stat., 33, 1065–1076, 1962.</mixed-citation></ref>
      <ref id="bib1.bibx31"><label>Phipps et al.(2013)</label><mixed-citation> Phipps, S. J., McGregor, H. V., Gergis, J., Gallant, A. J., Neukom, R.,  Stevenson, S., Ackerley, D., Brown, J. R., Fischer, M. J., and Van Ommen,  T. D.: Paleoclimate data–model comparison and the role of climate forcings  over the past 1500 years, J. Climate, 26, 6915–6936, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx32"><label>Steiger et al.(2014)</label><mixed-citation> Steiger, N. J., Hakim, G. J., Steig, E. J., Battisti, D. S., and Roe, G. H.:  Assimilation of time-averaged pseudoproxies for climate reconstruction, J. Climate, 27, 426–441, 2014.</mixed-citation></ref>
      <ref id="bib1.bibx33"><label>Sun et al.(2025)</label><mixed-citation>Sun, H., Lei, L., Liu, Z., Ning, L., and Tan, Z.-M.: An online paleoclimate  data assimilation with a deep learning-based network, J. Adv. Model. Earth Sy., 17, e2024MS004675, <ext-link xlink:href="https://doi.org/10.1029/2024MS004675" ext-link-type="DOI">10.1029/2024MS004675</ext-link>, 2025.</mixed-citation></ref>
      <ref id="bib1.bibx34"><label>Tardif et al.(2019)</label><mixed-citation>Tardif, R., Hakim, G. J., Perkins, W. A., Horlick, K. A., Erb, M. P., Emile-Geay, J., Anderson, D. M., Steig, E. J., and Noone, D.: Last Millennium Reanalysis with an expanded proxy database and seasonal proxy modeling, Clim. Past, 15, 1251–1273, <ext-link xlink:href="https://doi.org/10.5194/cp-15-1251-2019" ext-link-type="DOI">10.5194/cp-15-1251-2019</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx35"><label>Tingley et al.(2012)</label><mixed-citation> Tingley, M. P., Craigmile, P. F., Haran, M., Li, B., Mannshardt, E., and  Rajaratnam, B.: Piecing together the past: statistical insights into  paleoclimatic reconstructions, Quaternary Sci. Rev., 35, 1–22, 2012.</mixed-citation></ref>
      <ref id="bib1.bibx36"><label>Vaccaro et al.(2021)</label><mixed-citation> Vaccaro, A., Emile-Geay, J., Guillot, D., Verna, R., Morice, C., Kennedy, J.,  and Rajaratnam, B.: Climate field completion via Markov random fields:  Application to the HadCRUT4. 6 temperature dataset, J. Climate, 34,  4169–4188, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx37"><label>Von Storch et al.(2000)</label><mixed-citation> von Storch, H., Cubasch, U., Gonzalez-Rouco, J. F., Jones, J. M., Voss, R., Widmann, M., and Zorita, E.: Combining paleoclimatic evidence and GCMs by means of data assimilation through upscaling and nudging (DATUN), in: Proceedings of the 11th symposium on global change studies, 28–31, 2000.</mixed-citation></ref>
      <ref id="bib1.bibx38"><label>Wang et al.(2015)</label><mixed-citation> Wang, J., Emile-Geay, J., Guillot, D., McKay, N. P., and Rajaratnam, B.:  Fragility of reconstructed temperature patterns over the Common Era:  Implications for model evaluation, Geophys. Res. Lett., 42, 7162–7170, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx39"><label>Whitaker and Hamill(2002)</label><mixed-citation>Whitaker, J. S. and Hamill, T. M.: Ensemble data assimilation without perturbed observations, Mon. Weather Rev., 130, 1913–1924, 2002.  </mixed-citation></ref>
      <ref id="bib1.bibx40"><label>Widmann et al.(2010)</label><mixed-citation>Widmann, M., Goosse, H., van der Schrier, G., Schnur, R., and Barkmeijer, J.: Using data assimilation to study extratropical Northern Hemisphere climate over the last millennium, Clim. Past, 6, 627–644, <ext-link xlink:href="https://doi.org/10.5194/cp-6-627-2010" ext-link-type="DOI">10.5194/cp-6-627-2010</ext-link>, 2010.</mixed-citation></ref>
      <ref id="bib1.bibx41"><label>Zeng and Janjić(2016)</label><mixed-citation> Zeng, Y. and Janjić, T.: Study of conservation laws with the local  ensemble transform kalman filter, Q. J. Roy. Meteor. Soc., 699, 2359–2372, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx42"><label>Zhu et al.(2023)</label><mixed-citation>Zhu, F., Emile-Geay, J., Anchukaitis, K. J., McKay, N. P., Stevenson, S., and  Meng, Z.: A pseudoproxy emulation of the PAGES 2k database using a hierarchy  of proxy system models, Scientific Data, 10, 624, <ext-link xlink:href="https://doi.org/10.1038/s41597-023-02489-1" ext-link-type="DOI">10.1038/s41597-023-02489-1</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx43"><label>Zhu et al.(2024a)</label><mixed-citation>Zhu, F., Emile-Geay, J., Hakim, G. J., Guillot, D., Khider, D., Tardif, R., and Perkins, W. A.: cfr: a Python package for Climate Field Reconstruction, Version v2024.1.26, Zenodo [code], <ext-link xlink:href="https://doi.org/10.5281/zenodo.10575537" ext-link-type="DOI">10.5281/zenodo.10575537</ext-link>, 2024a.</mixed-citation></ref>
      <ref id="bib1.bibx44"><label>Zhu et al.(2024b)</label><mixed-citation>Zhu, F., Emile-Geay, J., Hakim, G. J., Guillot, D., Khider, D., Tardif, R., and Perkins, W. A.: cfr (v2024.1.26): a Python package for climate field reconstruction, Geosci. Model Dev., 17, 3409–3431, <ext-link xlink:href="https://doi.org/10.5194/gmd-17-3409-2024" ext-link-type="DOI">10.5194/gmd-17-3409-2024</ext-link>, 2024b.</mixed-citation></ref>

  </ref-list></back>
    <!--<article-title-html>Paleoclimate data assimilation with adaptive observation error inflation and adaptive localization</article-title-html>
<abstract-html/>
<ref-html id="bib1.bib1"><label>Anderson(2007)</label><mixed-citation>
      
Anderson, J. L.: Exploring the need for localization in ensemble data  assimilation using a hierarchical ensemble filter, Physica D, 230, 99–111, 2007.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib2"><label>Anderson(2012)</label><mixed-citation>
      
Anderson, J. L.: Localization and sampling error correction in ensemble Kalman filter data assimilation, Mon. Weather Rev., 140, 2359–2371, 2012.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib3"><label>Bishop and Hodyss(2009a)</label><mixed-citation>
      
Bishop, C. and Hodyss, D.: Ensemble covariances adaptively localized with  ECO-RAP. Part 1: Tests on simple error models, Tellus A, 61, 84–96, 2009a.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib4"><label>Bishop and Hodyss(2009b)</label><mixed-citation>
      
Bishop, C. and Hodyss, D.: Ensemble covariances adaptively localized with  ECO-RAP. Part 2: A strategy for the atmosphere, Tellus A, 61, 97–111, 2009b.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib5"><label>Brady et al.(2019)</label><mixed-citation>
      
Brady, E., Stevenson, S., Bailey, D., Liu, Z., Noone, D., Nusbaumer, J.,  Otto-Bliesner, B. L., Tabor, C., Tomas, R., Wong, T., Zhang, J., and Zhu, J.: The connected  isotopic water cycle in the Community Earth System Model version 1, J. Adv. Model. Earth Sy., 11, 2547–2566, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib6"><label>Dee et al.(2015)</label><mixed-citation>
      
Dee, S., Emile-Geay, J., Evans, M., Allam, A., Steig, E., and Thompson, D.:  PRYSM: An open-source framework for PRoxY System Modeling, with applications  to oxygen-isotope systems, J. Adv. Model. Earth Sy., 7, 1220–1247, 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib7"><label>Evans et al.(2013)</label><mixed-citation>
      
Evans, M. N., Tolwinski-Ward, S. E., Thompson, D. M., and Anchukaitis, K. J.:  Applications of proxy system modeling in high resolution paleoclimatology,  Quaternary Sci. Rev., 76, 16–28, 2013.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib8"><label>Evensen(1994)</label><mixed-citation>
      
Evensen, G.: Sequential data assimilation with a nonlinear quasi-geostrophic  model using Monte Carlo methods to forecast error statistics, J. Geophys. Res.-Oceans, 99, 10143–10162, 1994.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib9"><label>Gaspari and Cohn(1999)</label><mixed-citation>
      
Gaspari, G. and Cohn, S. E.: Construction of correlation functions in two and  three dimensions, Q. J. Roy. Meteor. Soc., 125, 723–757, 1999.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib10"><label>Goosse(2017)</label><mixed-citation>
      
Goosse, H.: Reconstructed and simulated temperature asymmetry between  continents in both hemispheres over the last centuries, Clim. Dynam., 48,  1483–1501, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib11"><label>Guillot et al.(2015)</label><mixed-citation>
      
Guillot, D., Rajaratnam, B., and Emile-Geay, J.: Statistical paleoclimate  reconstructions via Markov random fields, Ann. Appl. Stat., 9, 324–352, 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib12"><label>Hakim et al.(2016)</label><mixed-citation>
      
Hakim, G. J., Emile-Geay, J., Steig, E. J., Noone, D., Anderson, D. M., Tardif, R., Steiger, N., and Perkins, W. A.: The last millennium climate reanalysis project: Framework and first results, J. Geophys. Res.-Atmos., 121, 6745–6764, 2016.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib13"><label>Houtekamer and Mitchell(1998)</label><mixed-citation>
      
Houtekamer, P. L. and Mitchell, H. L.: Data assimilation using an ensemble  Kalman filter technique, Mon. Weather Rev., 126, 796–811, 1998.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib14"><label>Houtekamer and Mitchell(2001)</label><mixed-citation>
      
Houtekamer, P. L. and Mitchell, H. L.: A sequential ensemble Kalman filter for atmospheric data assimilation, Mon. Weather Rev., 129, 123–137, 2001.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib15"><label>Huber(1992)</label><mixed-citation>
      
Huber, P. J.: Robust estimation of a location parameter, in: Breakthroughs in  statistics: Methodology and distribution, Springer, 492–518, <a href="https://doi.org/10.1007/978-1-4612-4380-9_35" target="_blank">https://doi.org/10.1007/978-1-4612-4380-9_35</a>, 1992.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib16"><label>Kageyama et al.(2021)</label><mixed-citation>
      
Kageyama, M., Harrison, S. P., Kapsch, M.-L., Lofverstrom, M., Lora, J. M., Mikolajewicz, U., Sherriff-Tadano, S., Vadsaria, T., Abe-Ouchi, A., Bouttes, N., Chandan, D., Gregoire, L. J., Ivanovic, R. F., Izumi, K., LeGrande, A. N., Lhardy, F., Lohmann, G., Morozova, P. A., Ohgaito, R., Paul, A., Peltier, W. R., Poulsen, C. J., Quiquet, A., Roche, D. M., Shi, X., Tierney, J. E., Valdes, P. J., Volodin, E., and Zhu, J.: The PMIP4 Last Glacial Maximum experiments: preliminary results and comparison with the PMIP3 simulations, Clim. Past, 17, 1065–1089, <a href="https://doi.org/10.5194/cp-17-1065-2021" target="_blank">https://doi.org/10.5194/cp-17-1065-2021</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib17"><label>Lei and Anderson(2014a)</label><mixed-citation>
      
Lei, L. and Anderson, J. L.: Comparisons of empirical localization techniques  for serial ensemble Kalman filters in a simple atmospheric general  circulation model, Mon. Weather Rev., 142, 739–754, 2014a.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib18"><label>Lei and Anderson(2014b)</label><mixed-citation>
      
Lei, L. and Anderson, J. L.: Empirical localization of observations for serial ensemble Kalman filter data assimilation in an atmospheric general  circulation model, Mon. Weather Rev., 142, 1835–1851, 2014b.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib19"><label>Lenssen et al.(2019)</label><mixed-citation>
      
Lenssen, N. J., Schmidt, G. A., Hansen, J. E., Menne, M. J., Persin, A., Ruedy, R., and Zyss, D.: Improvements in the GISTEMP uncertainty model, J.  Geophys. Res.-Atmos., 124, 6307–6326, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib20"><label>Liu et al.(2014)</label><mixed-citation>
      
Liu, Z., Zhu, J., Rosenthal, Y., Zhang, X., Otto-Bliesner, B. L., Timmermann,  A., Smith, R. S., Lohmann, G., Zheng, W., and Elison Timm, O.: The Holocene  temperature conundrum, P. Natl. Acad. Sci. USA, 111, E3501–E3505, 2014.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib21"><label>Lorenz(1956)</label><mixed-citation>
      
Lorenz, E. N.: Empirical orthogonal functions and statistical weather  prediction, Vol. 1, Department of  Meteorology, Massachusetts Institute of Technology, Cambridge, 1956.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib22"><label>Luo et al.(2026)</label><mixed-citation>
      
Luo, G., Zeng, Y., Zhu, F., and Zhao, J.: Paleoclimate data assimilation with  adaptive observation error inflation and adaptive localization, Version v2, Zenodo [data set/code], <a href="https://doi.org/10.5281/zenodo.19015635" target="_blank">https://doi.org/10.5281/zenodo.19015635</a>, 2026.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib23"><label>Mann et al.(1998)</label><mixed-citation>
      
Mann, M. E., Bradley, R. S., and Hughes, M. K.: Global-scale temperature  patterns and climate forcing over the past six centuries, Nature, 392,  779–787, 1998.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib24"><label>Mann et al.(2008)</label><mixed-citation>
      
Mann, M. E., Zhang, Z., Hughes, M. K., Bradley, R. S., Miller, S. K.,  Rutherford, S., and Ni, F.: Proxy-based reconstructions of hemispheric and  global surface temperature variations over the past two millennia, P. Natl. Acad. Sci. USA, 105, 13252–13257, 2008.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib25"><label>Ménétrier et al.(2015a)</label><mixed-citation>
      
Ménétrier, B., Montmerle, T., Michel, Y., and Berre, L.: Linear  filtering of sample covariances for ensemble-based data assimilation. Part I:  Optimality criteria and application to variance filtering and covariance  localization, Mon. Weather Rev., 143, 1622–1643, 2015a.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib26"><label>Ménétrier et al.(2015b)</label><mixed-citation>
      
Ménétrier, B., Montmerle, T., Michel, Y., and Berre, L.: Linear  filtering of sample covariances for ensemble-based data assimilation. Part  II: Application to a convective-scale NWP model, Mon. Weather Rev., 143,  1644–1664, 2015b.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib27"><label>Minamide and Zhang(2017)</label><mixed-citation>
      
Minamide, M. and Zhang, F.: Adaptive observation error inflation for  assimilating all-sky satellite radiance, Mon. Weather Rev., 145, 1063–1081, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib28"><label>Osman et al.(2021)</label><mixed-citation>
      
Osman, M. B., Tierney, J. E., Zhu, J., Tardif, R., Hakim, G. J., King, J., and Poulsen, C. J.: Globally resolved surface temperatures since the Last Glacial Maximum, Nature, 599, 239–244, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib29"><label>PAGES 2k Consortium(2017)</label><mixed-citation>
      
PAGES 2k Consortium: A global multiproxy database for temperature  reconstructions of the Common Era, Scientific Data, 4, 170088,  <a href="https://doi.org/10.1038/sdata.2017.88" target="_blank">https://doi.org/10.1038/sdata.2017.88</a>, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib30"><label>Parzen(1962)</label><mixed-citation>
      
Parzen, E.: On estimation of a probability density function and mode, Ann. Math. Stat., 33, 1065–1076, 1962.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib31"><label>Phipps et al.(2013)</label><mixed-citation>
      
Phipps, S. J., McGregor, H. V., Gergis, J., Gallant, A. J., Neukom, R.,  Stevenson, S., Ackerley, D., Brown, J. R., Fischer, M. J., and Van Ommen,  T. D.: Paleoclimate data–model comparison and the role of climate forcings  over the past 1500 years, J. Climate, 26, 6915–6936, 2013.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib32"><label>Steiger et al.(2014)</label><mixed-citation>
      
Steiger, N. J., Hakim, G. J., Steig, E. J., Battisti, D. S., and Roe, G. H.:  Assimilation of time-averaged pseudoproxies for climate reconstruction, J. Climate, 27, 426–441, 2014.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib33"><label>Sun et al.(2025)</label><mixed-citation>
      
Sun, H., Lei, L., Liu, Z., Ning, L., and Tan, Z.-M.: An online paleoclimate  data assimilation with a deep learning-based network, J. Adv. Model. Earth Sy., 17, e2024MS004675, <a href="https://doi.org/10.1029/2024MS004675" target="_blank">https://doi.org/10.1029/2024MS004675</a>, 2025.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib34"><label>Tardif et al.(2019)</label><mixed-citation>
      
Tardif, R., Hakim, G. J., Perkins, W. A., Horlick, K. A., Erb, M. P., Emile-Geay, J., Anderson, D. M., Steig, E. J., and Noone, D.: Last Millennium Reanalysis with an expanded proxy database and seasonal proxy modeling, Clim. Past, 15, 1251–1273, <a href="https://doi.org/10.5194/cp-15-1251-2019" target="_blank">https://doi.org/10.5194/cp-15-1251-2019</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib35"><label>Tingley et al.(2012)</label><mixed-citation>
      
Tingley, M. P., Craigmile, P. F., Haran, M., Li, B., Mannshardt, E., and  Rajaratnam, B.: Piecing together the past: statistical insights into  paleoclimatic reconstructions, Quaternary Sci. Rev., 35, 1–22, 2012.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib36"><label>Vaccaro et al.(2021)</label><mixed-citation>
      
Vaccaro, A., Emile-Geay, J., Guillot, D., Verna, R., Morice, C., Kennedy, J.,  and Rajaratnam, B.: Climate field completion via Markov random fields:  Application to the HadCRUT4. 6 temperature dataset, J. Climate, 34,  4169–4188, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib37"><label>Von Storch et al.(2000)</label><mixed-citation>
      
von Storch, H., Cubasch, U., Gonzalez-Rouco, J. F., Jones, J. M., Voss, R., Widmann, M., and Zorita, E.: Combining paleoclimatic evidence and GCMs by means of data assimilation through upscaling and nudging (DATUN), in: Proceedings of the 11th symposium on global change studies, 28–31, 2000.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib38"><label>Wang et al.(2015)</label><mixed-citation>
      
Wang, J., Emile-Geay, J., Guillot, D., McKay, N. P., and Rajaratnam, B.:  Fragility of reconstructed temperature patterns over the Common Era:  Implications for model evaluation, Geophys. Res. Lett., 42, 7162–7170, 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib39"><label>Whitaker and Hamill(2002)</label><mixed-citation>
      
Whitaker, J. S. and Hamill, T. M.: Ensemble data assimilation without perturbed observations, Mon. Weather Rev., 130, 1913–1924, 2002.


    </mixed-citation></ref-html>
<ref-html id="bib1.bib40"><label>Widmann et al.(2010)</label><mixed-citation>
      
Widmann, M., Goosse, H., van der Schrier, G., Schnur, R., and Barkmeijer, J.: Using data assimilation to study extratropical Northern Hemisphere climate over the last millennium, Clim. Past, 6, 627–644, <a href="https://doi.org/10.5194/cp-6-627-2010" target="_blank">https://doi.org/10.5194/cp-6-627-2010</a>, 2010.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib41"><label>Zeng and Janjić(2016)</label><mixed-citation>
      
Zeng, Y. and Janjić, T.: Study of conservation laws with the local  ensemble transform kalman filter, Q. J. Roy. Meteor. Soc., 699, 2359–2372, 2016.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib42"><label>Zhu et al.(2023)</label><mixed-citation>
      
Zhu, F., Emile-Geay, J., Anchukaitis, K. J., McKay, N. P., Stevenson, S., and  Meng, Z.: A pseudoproxy emulation of the PAGES 2k database using a hierarchy  of proxy system models, Scientific Data, 10, 624, <a href="https://doi.org/10.1038/s41597-023-02489-1" target="_blank">https://doi.org/10.1038/s41597-023-02489-1</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib43"><label>Zhu et al.(2024a)</label><mixed-citation>
      
Zhu, F., Emile-Geay, J., Hakim, G. J., Guillot, D., Khider, D., Tardif, R., and Perkins, W. A.: cfr: a Python package for Climate Field Reconstruction, Version v2024.1.26, Zenodo [code], <a href="https://doi.org/10.5281/zenodo.10575537" target="_blank">https://doi.org/10.5281/zenodo.10575537</a>, 2024a.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib44"><label>Zhu et al.(2024b)</label><mixed-citation>
      
Zhu, F., Emile-Geay, J., Hakim, G. J., Guillot, D., Khider, D., Tardif, R., and Perkins, W. A.: cfr (v2024.1.26): a Python package for climate field reconstruction, Geosci. Model Dev., 17, 3409–3431, <a href="https://doi.org/10.5194/gmd-17-3409-2024" target="_blank">https://doi.org/10.5194/gmd-17-3409-2024</a>, 2024b.

    </mixed-citation></ref-html>--></article>
