<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing with OASIS Tables v3.0 20080202//EN" "https://jats.nlm.nih.gov/nlm-dtd/publishing/3.0/journalpub-oasis3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:oasis="http://docs.oasis-open.org/ns/oasis-exchange/table" xml:lang="en" dtd-version="3.0" article-type="research-article">
  <front>
    <journal-meta><journal-id journal-id-type="publisher">GMD</journal-id><journal-title-group>
    <journal-title>Geoscientific Model Development</journal-title>
    <abbrev-journal-title abbrev-type="publisher">GMD</abbrev-journal-title><abbrev-journal-title abbrev-type="nlm-ta">Geosci. Model Dev.</abbrev-journal-title>
  </journal-title-group><issn pub-type="epub">1991-9603</issn><publisher>
    <publisher-name>Copernicus Publications</publisher-name>
    <publisher-loc>Göttingen, Germany</publisher-loc>
  </publisher></journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.5194/gmd-18-5549-2025</article-id><title-group><article-title>CRITER 1.0: a coarse reconstruction with iterative refinement network for sparse spatio-temporal satellite data</article-title><alt-title>Sparse satellite data reconstruction</alt-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author" corresp="yes" rid="aff1">
          <name><surname>Zupančič Muc</surname><given-names>Matjaž</given-names></name>
          <email>matjazzupancicmuc@gmail.com</email>
        </contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1">
          <name><surname>Zavrtanik</surname><given-names>Vitjan</given-names></name>
          
        </contrib>
        <contrib contrib-type="author" corresp="no" rid="aff2">
          <name><surname>Barth</surname><given-names>Alexander</given-names></name>
          
        <ext-link>https://orcid.org/0000-0003-2952-5997</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff2">
          <name><surname>Alvera-Azcarate</surname><given-names>Aida</given-names></name>
          
        <ext-link>https://orcid.org/0000-0002-0484-4791</ext-link></contrib>
        <contrib contrib-type="author" equal-contrib="yes" corresp="no" rid="aff3 aff4">
          <name><surname>Ličer</surname><given-names>Matjaž</given-names></name>
          
        <ext-link>https://orcid.org/0000-0003-2304-2505</ext-link></contrib>
        <contrib contrib-type="author" equal-contrib="yes" corresp="no" rid="aff1">
          <name><surname>Kristan</surname><given-names>Matej</given-names></name>
          
        <ext-link>https://orcid.org/0000-0002-4252-4342</ext-link></contrib>
        <aff id="aff1"><label>1</label><institution>Faculty of Computer and Information Science, Visual Cognitive Systems Lab, University of Ljubljana, Ljubljana, Slovenia</institution>
        </aff>
        <aff id="aff2"><label>2</label><institution>Department of Astrophysics, Geophysics and Oceanography, Geohydrodynamics and Environment Research,  University of Liège, Liège, Belgium</institution>
        </aff>
        <aff id="aff3"><label>3</label><institution>Slovenian Environment Agency, Office for Meteorology, Hydrology and Oceanography, Ljubljana, Slovenia</institution>
        </aff>
        <aff id="aff4"><label>4</label><institution>National Institute of Biology, Marine Biology Station, Piran, Slovenia</institution>
        </aff><author-comment content-type="econtrib"><p>These authors contributed equally to this work.</p></author-comment>
      </contrib-group>
      <author-notes><corresp id="corr1">Matjaž Zupančič Muc (matjazzupancicmuc@gmail.com)</corresp></author-notes><pub-date><day>4</day><month>September</month><year>2025</year></pub-date>
      
      <volume>18</volume>
      <issue>17</issue>
      <fpage>5549</fpage><lpage>5573</lpage>
      <history>
        <date date-type="received"><day>8</day><month>November</month><year>2024</year></date>
           <date date-type="rev-request"><day>6</day><month>February</month><year>2025</year></date>
           <date date-type="rev-recd"><day>28</day><month>June</month><year>2025</year></date>
           <date date-type="accepted"><day>1</day><month>July</month><year>2025</year></date>
      </history>
      <permissions>
        <copyright-statement>Copyright: © 2025 Matjaž Zupančič Muc et al.</copyright-statement>
        <copyright-year>2025</copyright-year>
      <license license-type="open-access"><license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p></license></permissions><self-uri xlink:href="https://gmd.copernicus.org/articles/18/5549/2025/gmd-18-5549-2025.html">This article is available from https://gmd.copernicus.org/articles/18/5549/2025/gmd-18-5549-2025.html</self-uri><self-uri xlink:href="https://gmd.copernicus.org/articles/18/5549/2025/gmd-18-5549-2025.pdf">The full text article is available as a PDF file from https://gmd.copernicus.org/articles/18/5549/2025/gmd-18-5549-2025.pdf</self-uri>
      <abstract><title>Abstract</title>

      <p id="d2e151">Satellite observations of sea surface temperature (SST) are essential for accurate weather forecasting and climate modeling. However, these data often suffer from incomplete coverage due to cloud obstruction and limited satellite swath width, which requires development of dense reconstruction algorithms. The current state of the art struggles to accurately recover high-frequency variability, particularly in SST gradients in ocean fronts, eddies, and filaments, which are crucial for downstream processing and predictive tasks. To address this challenge, we propose a novel two-stage method CRITER (Coarse Reconstruction with ITerative Refinement Network), which consists of two stages. First, it reconstructs low-frequency SST components utilizing a Vision Transformer-based model, leveraging global spatio-temporal correlations in the available observations. Second, a UNet type of network iteratively refines the estimate by recovering high-frequency details. Extensive analysis on datasets from the Mediterranean, Adriatic, and Atlantic seas demonstrates CRITER's superior performance over the current state of the art. Specifically, CRITER achieves up to 44 % lower reconstruction errors of the missing values and over 80 % lower reconstruction errors of the observed values compared to the state of the art.</p>
  </abstract>
    
<funding-group>
<award-group id="gs1">
<funding-source>The Slovenian Research and Innovation Agency</funding-source>
<award-id>P1-0237</award-id>
<award-id>J2-2506</award-id>
<award-id>P2-0214</award-id>
</award-group>
</funding-group>
</article-meta>
  </front>
<body>
      

      
<sec id="Ch1.S1" sec-type="intro">
  <label>1</label><title>Introduction</title>
      <p id="d2e165">Infrared satellite sea surface temperature (SST) data are critical for ocean modeling, climate monitoring, fisheries management, and marine ecology <xref ref-type="bibr" rid="bib1.bibx29" id="paren.1"/>. On the one hand, the SST is a key boundary condition for atmospheric models extending from classical numerical weather prediction <xref ref-type="bibr" rid="bib1.bibx35 bib1.bibx9" id="paren.2"/> to extreme storms <xref ref-type="bibr" rid="bib1.bibx33" id="paren.3"/> and climate variability <xref ref-type="bibr" rid="bib1.bibx17" id="paren.4"/>. In the ocean realm, continuous description of SST is vital for analyses of mesoscale <xref ref-type="bibr" rid="bib1.bibx6" id="paren.5"/> and submesoscale baroclinic processes like fronts and eddies but also for implementations of atmosphere–ocean couplings through turbulent heat fluxes <xref ref-type="bibr" rid="bib1.bibx36 bib1.bibx23" id="paren.6"/>. Furthermore, vertical temperature profiles are a critical driver of heat, carbon, and nutrient exchange between the surface and the deep ocean and thus for a wide plethora of biogeochemical processes <xref ref-type="bibr" rid="bib1.bibx28" id="paren.7"/> in the ocean surface boundary layer which depend on the temperatures above the pycnocline. Last but not least, SST is a key parameter for the detection, mapping, and analysis of marine heatwaves <xref ref-type="bibr" rid="bib1.bibx22" id="paren.8"/>, and reconstructed satellite fields are imperative for determining the regional extent and intensity of such extreme events <xref ref-type="bibr" rid="bib1.bibx30 bib1.bibx10" id="paren.9"/>, which can have enormous impacts on aquaculture, fisheries, and other aspects of economy <xref ref-type="bibr" rid="bib1.bibx20 bib1.bibx18" id="paren.10"/>.</p>
      <p id="d2e199">Such downstream applications therefore often require complete, dense SST fields but cloud cover and sparse satellite coverage invariably lead to gappy and sparse data in both space and time. Reconstruction of gaps in the observations is therefore essential for a continuous description of ocean temperature fields and for many daily operational processes. These can be categorized into two groups: (i) extensions of the optimal interpolation (OI) scheme <xref ref-type="bibr" rid="bib1.bibx37 bib1.bibx38" id="paren.11"/>, and (ii) data-driven approaches. The latter includes methods based on empirical orthogonal functions (EOFs), such as DINEOF <xref ref-type="bibr" rid="bib1.bibx1" id="paren.12"/>, and, more recently, end-to-end deep learning techniques. Notable deep learning methods include DINCAE1 <xref ref-type="bibr" rid="bib1.bibx2" id="paren.13"/>, dADRSR <xref ref-type="bibr" rid="bib1.bibx7 bib1.bibx15" id="paren.14"/>, TS-RBFNN <xref ref-type="bibr" rid="bib1.bibx39" id="paren.15"/>, DINCAE2 <xref ref-type="bibr" rid="bib1.bibx3" id="paren.16"/>, 4DVarNet <xref ref-type="bibr" rid="bib1.bibx14" id="paren.17"/>, 4DVarNet-SSH <xref ref-type="bibr" rid="bib1.bibx5" id="paren.18"/>, the SSH reconstruction method by <xref ref-type="bibr" rid="bib1.bibx26" id="text.19"/>, NeurOST <xref ref-type="bibr" rid="bib1.bibx27" id="paren.20"/>, and MAESSTRO <xref ref-type="bibr" rid="bib1.bibx19" id="paren.21"/>.</p>
      <p id="d2e236">Traditional methods like DINEOF <xref ref-type="bibr" rid="bib1.bibx1" id="paren.22"/> have been widely adopted, iteratively filling in missing data using truncated EOF decomposition. While effective for large-scale patterns, DINEOF struggles with fine-scale features, mostly because of their transient nature. Deep learning approaches have since emerged, surpassing traditional methods' performance. DINCAE1 <xref ref-type="bibr" rid="bib1.bibx2" id="paren.23"/> introduced a UNet-based <xref ref-type="bibr" rid="bib1.bibx34" id="paren.24"/> model with probabilistic output, while 4DVarNet <xref ref-type="bibr" rid="bib1.bibx14" id="paren.25"/> proposed an energy-based formulation for interpolation, achieving comparable SST reconstruction performance to a convolutional autoencoder architecturally similar to DINCAE1. Recently, <xref ref-type="bibr" rid="bib1.bibx39" id="text.26"/> proposed a physically informed neural network that reconstructs daily SSTs in both cloudy and cloud-free areas, outperforming DINEOF. Beyond gap filling, super-resolution techniques have been developed to enhance SST resolution: <xref ref-type="bibr" rid="bib1.bibx24" id="text.27"/> designed a network that fuses optical and thermal satellite imagery, and, more recently, <xref ref-type="bibr" rid="bib1.bibx15" id="text.28"/> applied a convolutional super-resolution network (originally proposed by <xref ref-type="bibr" rid="bib1.bibx7" id="altparen.29"/>) to super-resolve small low-resolution SST tiles obtained through optimal interpolation, improving fine-scale feature reconstruction.</p>
      <p id="d2e264">DINCAE2 <xref ref-type="bibr" rid="bib1.bibx3" id="paren.30"/>, the current state of the art and successor to DINCAE1, extended the original implementation with an additional refinement UNet. It operates on temporally consecutive partial SST observations, gradually improving central SST field reconstruction. However, its finite receptive field limits long-range spatio-temporal dependency exploitation, resulting in oversmoothed reconstructions lacking high-frequency details. Recently, MAESSTRO <xref ref-type="bibr" rid="bib1.bibx19" id="paren.31"/> addressed some limitations by adapting the Masked Autoencoder (MAE) <xref ref-type="bibr" rid="bib1.bibx21" id="paren.32"/> framework for SST reconstruction. It employs a Vision Transformer (ViT)  <xref ref-type="bibr" rid="bib1.bibx11" id="paren.33"/> architecture to capture global spatial dependencies. However, its single-time-step approach neglects temporal correlations, potentially compromising reconstruction quality for large, contiguous cloud occlusions. Furthermore, MAESSTRO's random patch masking strategy during training and evaluation may inadequately represent real cloud patterns, potentially yielding optimistic error estimates.</p>
      <p id="d2e280">To address these limitations, we propose a two-stage Coarse Reconstruction with ITerative Refinement network (CRITER). A transformer-based module first leverages long-range spatio-temporal dependencies to estimate a low-frequency reconstruction. Subsequently, an iterative refinement module enhances high-frequency content. Unlike previous methods, which attempt full signal reconstruction in each block, CRITER decomposes the problem into a sequence of networks, each reducing the residual error of its predecessor, thus optimizing network capacity for local error reduction.</p>
      <p id="d2e283">The paper is structured as follows. Section <xref ref-type="sec" rid="Ch1.S2"/> contains descriptions of employed datasets together with preprocessing steps executed prior to the training. Section <xref ref-type="sec" rid="Ch1.S3"/> describes the CRITER architecture, focusing on coarse reconstruction step in Sect. <xref ref-type="sec" rid="Ch1.S3.SS1"/>, its iterative refinement in Sect. <xref ref-type="sec" rid="Ch1.S3.SS2"/>, and residual estimation network in the “Residual estimation network (REN)” section. The training strategy is described in Sect. <xref ref-type="sec" rid="Ch1.S3.SS3"/>, and results are listed in Sect. <xref ref-type="sec" rid="Ch1.S4"/>, including an in-depth ablation study investigating the role of individual architectural components (Sect. <xref ref-type="sec" rid="Ch1.S4.SS5"/>).</p>
</sec>
<sec id="Ch1.S2">
  <label>2</label><title>Input data: sea surface temperature</title>
<sec id="Ch1.S2.SS1">
  <label>2.1</label><title>Evaluation datasets</title>
      <p id="d2e316">For our study we utilize Level 3 (L3) sea surface temperature (SST) satellite observation products. L3 level of product refers to the  satellite product where spatially sparse and irregular point observations of the ocean surface are gridded into a fixed grid across space and/or time. Such products may combine multiple satellite overpasses or even multiple sensors for the same observed quantity.</p>
      <p id="d2e319">Specifically we consider the following three datasets corresponding to three different geographic regions: <list list-type="order"><list-item>
      <p id="d2e324"><italic>Central Mediterranean</italic>: The SST_MED_SST_L3S_NRT_OBSERVATIONS_010_ 012_a <xref ref-type="bibr" rid="bib1.bibx12" id="paren.34"/> dataset contains daily near-real-time (NRT) SST measurements over the Mediterranean sea from 1 January 2008 to 31 December 2021. The dataset is provided on a remapped grid with a spatial resolution of 0.0625° <inline-formula><mml:math id="M1" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 0.0625°.</p></list-item><list-item>
      <p id="d2e340"><italic>Adriatic</italic>: The SST_MED_PHY_L3S_MY_010 _042 <xref ref-type="bibr" rid="bib1.bibx32 bib1.bibx8" id="paren.35"/> dataset contains daily multi-year reprocessed (MY) SST measurements over the Adriatic Sea from 25 August 1981 to 31 December 2022.</p>
      <p id="d2e348">The dataset is provided on a remapped grid with a spatial resolution of 0.05° <inline-formula><mml:math id="M2" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 0.05°.</p></list-item><list-item>
      <p id="d2e359"><italic>Atlantic</italic>: The SST_ATL_PHY_L3S_MY_010 _038 <xref ref-type="bibr" rid="bib1.bibx13" id="paren.36"/> dataset contains daily multi-year reprocessed (MY) SST measurements from 1 January 1982–1 January 2022. The dataset is provided on a remapped grid with a spatial resolution of 0.05° <inline-formula><mml:math id="M3" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 0.05°.</p></list-item></list></p>
      <p id="d2e374">These regions were chosen due to their oceanographic variety. The Adriatic is an elongated semi-enclosed basin with correspondingly poor satellite coverage, and the Central Mediterranean exhibits a wide variety of oceanographic regimes (from regions of freshwater influence in the northern Adriatic to a much deeper Ionian where  Levantine and Adriatic water masses communicate), while the Atlantic region is essentially an open ocean region, very different from the Adriatic. These regions should demonstrate generalization abilities of CRITER under a variety of oceanographic conditions. The geographic areas of the three datasets are shown in Fig. <xref ref-type="fig" rid="F1"/>. It is worth noting that two different satellite products are used in this study, a near-real-time (NRT) and a multi-year (MY) reprocessed dataset. This was done to show that like DINCAE2, CRITER also generalizes well across various datasets of SST. Furthermore, multi-year reprocessed datasets come at a higher resolution and span significantly longer periods of time, which gives access to a larger train and, more importantly, test set.</p>

      <fig id="F1"><label>Figure 1</label><caption><p id="d2e382">The map shows the spatial extent of the Central Mediterranean, Adriatic, and Atlantic datasets, highlighting the distinct geographic areas covered by each dataset.</p></caption>
          <graphic xlink:href="https://gmd.copernicus.org/articles/18/5549/2025/gmd-18-5549-2025-f01.png"/>

        </fig>

</sec>
<sec id="Ch1.S2.SS2">
  <label>2.2</label><title>Input data preprocessing</title>
<sec id="Ch1.S2.SS2.SSS1">
  <label>2.2.1</label><title>Filtering out days with excessive cloud coverage</title>
      <p id="d2e406">The satellite products corresponding to Level 3 (L3) SST are provided on a fixed grid but are spatially sparse over a subset of spatial locations (mainly due to clouds and land pixels). For training and evaluation of the method in this work, additional missing values need to be simulated to test network performance on values which are hidden to the network but are otherwise known. If the original SST observation field already contains a large number of missing measurements, it becomes difficult to effectively simulate additional missing data. Consequently, to ensure that the dataset is suitable for training and evaluating models, observations that are too sparse need to be filtered out. In the preprocessing stage, we first construct sequences of 3 temporally consecutive days of observed SST fields, as proposed by <xref ref-type="bibr" rid="bib1.bibx2" id="text.37"/>. The observation sequences are then filtered. Specifically, any 3 d observation sequence is discarded according to the following rule: if the cloud coverage, defined as the fraction of pixels that are missing in the central observation field, relative to the total number of pixels belonging to the sea, is greater than or equal to a certain threshold, the corresponding observation sequence is discarded. The appropriate threshold is selected by considering the total number of samples in each dataset. Specifically, we use a threshold of 100 % for the Mediterranean dataset, resulting in a total of <inline-formula><mml:math id="M4" display="inline"><mml:mn mathvariant="normal">5114</mml:mn></mml:math></inline-formula> samples. For the Adriatic dataset, we apply a threshold of 60 %, which yields 7800 samples. Finally, we use a threshold of 75 % for the Atlantic dataset, resulting in <inline-formula><mml:math id="M5" display="inline"><mml:mn mathvariant="normal">3454</mml:mn></mml:math></inline-formula> samples.</p>
</sec>
<sec id="Ch1.S2.SS2.SSS2">
  <label>2.2.2</label><title>Train, validation, and test datasets</title>
      <p id="d2e434">The filtered satellite SST observations are chronologically split into three subsets: the train set, which comprises the first 90 % of the samples; the validation set, which comprises the next 5 % of the samples; and the test set, which consists of the last 5 % of the samples. The models are trained on the train set, the hyper-parameters are tuned on the validation set, and the performance is assessed on the test set. This approach ensures evaluation on future, unseen data with no temporal overlap between training and test phases.</p>
</sec>
</sec>
</sec>
<sec id="Ch1.S3">
  <label>3</label><title>CRITER – Coarse Reconstruction with ITerative Refinement network</title>
      <p id="d2e447">Given a sequence of spatially sparse sea surface temperature observations <inline-formula><mml:math id="M6" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">X</mml:mi><mml:mi>m</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mo>[</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:msub><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:msub><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>, where <inline-formula><mml:math id="M7" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>∈</mml:mo><mml:msup><mml:mi mathvariant="double-struck">R</mml:mi><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>×</mml:mo><mml:mi>W</mml:mi><mml:mo>×</mml:mo><mml:mi>H</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> is the potentially sparse observation field of width <inline-formula><mml:math id="M8" display="inline"><mml:mi>W</mml:mi></mml:math></inline-formula> and height <inline-formula><mml:math id="M9" display="inline"><mml:mi>H</mml:mi></mml:math></inline-formula> at time step <inline-formula><mml:math id="M10" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M11" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:msub><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:msub><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> defines the observed time interval, the task is to estimate the dense reconstruction <inline-formula><mml:math id="M12" display="inline"><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo stretchy="false" mathvariant="normal">̃</mml:mo></mml:mover></mml:math></inline-formula> at time step <inline-formula><mml:math id="M13" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> and the uncertainty specified by the variance <inline-formula><mml:math id="M14" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>. Following <xref ref-type="bibr" rid="bib1.bibx3" id="text.38"/>, we set the temporal horizon to <inline-formula><mml:math id="M15" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M16" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">d</mml:mi></mml:mrow></mml:math></inline-formula>, thus in reconstruction of <inline-formula><mml:math id="M17" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, the days before and after day <inline-formula><mml:math id="M18" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> are considered.</p>
      <p id="d2e663">The proposed Coarse Reconstruction with ITerative Refinement network (CRITER) is a two-stage method composed of a Coarse Reconstruction Module (CRM), described in Sect. <xref ref-type="sec" rid="Ch1.S3.SS1"/>, and an Iterative Refinement Module (IRM), described in Sect. <xref ref-type="sec" rid="Ch1.S3.SS2"/>. An overview of the architecture is provided in Fig. <xref ref-type="fig" rid="F2"/>.</p>

      <fig id="F2" specific-use="star"><label>Figure 2</label><caption><p id="d2e674">Given observations for 3 consecutive days <inline-formula><mml:math id="M19" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> and a binary mask <inline-formula><mml:math id="M20" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> indicating missing pixels, CRITER densely reconstructs <inline-formula><mml:math id="M21" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> in two phases. First, the CRM estimates a coarse reconstruction <inline-formula><mml:math id="M22" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, which the IRM then iteratively refines to produce the final reconstruction <inline-formula><mml:math id="M23" display="inline"><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo stretchy="false" mathvariant="normal">̃</mml:mo></mml:mover></mml:math></inline-formula> and uncertainty <inline-formula><mml:math id="M24" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>. CRM tokenizes the input into tokens requiring reconstruction <inline-formula><mml:math id="M25" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">T</mml:mi><mml:mi mathvariant="normal">r</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and contextual tokens <inline-formula><mml:math id="M26" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">T</mml:mi><mml:mi mathvariant="normal">c</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. These contextual tokens are encoded by a ViT-based encoder into <inline-formula><mml:math id="M27" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="bold">T</mml:mi><mml:mi mathvariant="normal">c</mml:mi><mml:mi mathvariant="script">E</mml:mi></mml:msubsup></mml:mrow></mml:math></inline-formula>, combined with <inline-formula><mml:math id="M28" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">T</mml:mi><mml:mi mathvariant="normal">r</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, and decoded by a ViT-based decoder into decoded tokens <inline-formula><mml:math id="M29" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="bold">T</mml:mi><mml:mi mathvariant="normal">t</mml:mi><mml:mi mathvariant="script">D</mml:mi></mml:msubsup></mml:mrow></mml:math></inline-formula>, which are finally mapped to <inline-formula><mml:math id="M30" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. In the IRM, dashed lines indicate the iterative refinement process. At each iteration <inline-formula><mml:math id="M31" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula>, the current reconstruction estimate <inline-formula><mml:math id="M32" display="inline"><mml:mrow><mml:msup><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo stretchy="false" mathvariant="normal">̃</mml:mo></mml:mover><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> and uncertainty estimate <inline-formula><mml:math id="M33" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> are refined by adding the predicted residuals: reconstruction residual <inline-formula><mml:math id="M34" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="bold-italic">δ</mml:mi><mml:mi>x</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula> and uncertainty residual <inline-formula><mml:math id="M35" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="bold-italic">δ</mml:mi><mml:mrow><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula>. The index in REN<sup>(<italic>i</italic>)</sup> indicates the change in network parameters in each iteration.</p></caption>
        <graphic xlink:href="https://gmd.copernicus.org/articles/18/5549/2025/gmd-18-5549-2025-f02.png"/>

      </fig>

<sec id="Ch1.S3.SS1">
  <label>3.1</label><title>Coarse Reconstruction Module (CRM)</title>
      <p id="d2e961">The Coarse Reconstruction Module (CRM, Fig. <xref ref-type="fig" rid="F2"/>) follows the ViT encoder–decoder architecture <xref ref-type="bibr" rid="bib1.bibx11" id="paren.39"/>, similar to spatio-temporal MAE <xref ref-type="bibr" rid="bib1.bibx16" id="paren.40"/>.  The input observation fields <inline-formula><mml:math id="M37" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">X</mml:mi><mml:mi>m</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mo>[</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>]</mml:mo><mml:mo>∈</mml:mo><mml:msup><mml:mi mathvariant="double-struck">R</mml:mi><mml:mrow><mml:mn mathvariant="normal">3</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>×</mml:mo><mml:mi>W</mml:mi><mml:mo>×</mml:mo><mml:mi>H</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> are first fed to a tokenization process. To encode information about the yearly temperature cycle, each observation field <inline-formula><mml:math id="M38" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is concatenated channel-wise with a day-of-the-year auxiliary tensor <inline-formula><mml:math id="M39" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">a</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mo>[</mml:mo><mml:mi>sin⁡</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mi mathvariant="italic">π</mml:mi></mml:mrow><mml:mn mathvariant="normal">365</mml:mn></mml:mfrac></mml:mstyle><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:mi>cos⁡</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mi mathvariant="italic">π</mml:mi></mml:mrow><mml:mn mathvariant="normal">365</mml:mn></mml:mfrac></mml:mstyle><mml:mo>)</mml:mo><mml:mo>]</mml:mo><mml:mo>∈</mml:mo><mml:msup><mml:mi mathvariant="double-struck">R</mml:mi><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mo>×</mml:mo><mml:mi>W</mml:mi><mml:mo>×</mml:mo><mml:mi>H</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>, where the two channels contain constants, and <inline-formula><mml:math id="M40" display="inline"><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the numerical day of year index (between 1 and 365). The resulting fields are split into non-overlapping <inline-formula><mml:math id="M41" display="inline"><mml:mrow><mml:mn mathvariant="normal">3</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">8</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">8</mml:mn></mml:mrow></mml:math></inline-formula> patches, which are then flattened and linearly projected into tokens of shape <inline-formula><mml:math id="M42" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>×</mml:mo><mml:msub><mml:mi>D</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, where <inline-formula><mml:math id="M43" display="inline"><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the dimension of tokens used in ViT blocks, thus creating the list of tokens <inline-formula><mml:math id="M44" display="inline"><mml:mrow><mml:mi mathvariant="bold">T</mml:mi><mml:mo>=</mml:mo><mml:mo mathvariant="italic">{</mml:mo><mml:msub><mml:mi mathvariant="bold">T</mml:mi><mml:mi mathvariant="normal">r</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold">T</mml:mi><mml:mi mathvariant="normal">c</mml:mi></mml:msub><mml:mo mathvariant="italic">}</mml:mo></mml:mrow></mml:math></inline-formula>. Tokens <inline-formula><mml:math id="M45" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">T</mml:mi><mml:mi mathvariant="normal">r</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> correspond to patches in <inline-formula><mml:math id="M46" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> with at least one unobserved pixel and thus have to be reconstructed. Tokens <inline-formula><mml:math id="M47" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">T</mml:mi><mml:mi mathvariant="normal">c</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> are the remaining tokens, and they are used as a context for reconstruction. To encode the extent of missing values in a token, all tokens in <inline-formula><mml:math id="M48" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> are summed with their corresponding mask tokens. These are obtained by splitting the binary mask indicating missing pixels <inline-formula><mml:math id="M49" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:msub><mml:mo>∈</mml:mo><mml:mo mathvariant="italic">{</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:msup><mml:mo mathvariant="italic">}</mml:mo><mml:mrow><mml:mi>W</mml:mi><mml:mo>×</mml:mo><mml:mi>H</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> into <inline-formula><mml:math id="M50" display="inline"><mml:mrow><mml:mn mathvariant="normal">8</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">8</mml:mn></mml:mrow></mml:math></inline-formula> non-overlapping patches, which are then flattened and projected into mask tokens of shape <inline-formula><mml:math id="M51" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>×</mml:mo><mml:msub><mml:mi>D</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. To maintain the necessary spatio-temporal location of each token, all tokens in <inline-formula><mml:math id="M52" display="inline"><mml:mi mathvariant="bold">T</mml:mi></mml:math></inline-formula> are summed with a spatio-temporal positional embedding as in <xref ref-type="bibr" rid="bib1.bibx16" id="text.41"/>.</p>
      <p id="d2e1319">After obtaining tokens <inline-formula><mml:math id="M53" display="inline"><mml:mi mathvariant="bold">T</mml:mi></mml:math></inline-formula>, the context tokens <inline-formula><mml:math id="M54" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">T</mml:mi><mml:mi mathvariant="normal">c</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> are encoded by a ViT <xref ref-type="bibr" rid="bib1.bibx11" id="paren.42"/> encoder <inline-formula><mml:math id="M55" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="script">E</mml:mi><mml:mi mathvariant="normal">ViT</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> into <inline-formula><mml:math id="M56" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="bold">T</mml:mi><mml:mi mathvariant="normal">c</mml:mi><mml:mi mathvariant="script">E</mml:mi></mml:msubsup></mml:mrow></mml:math></inline-formula> (Fig. <xref ref-type="fig" rid="F2"/>). Then, the list of tokens <inline-formula><mml:math id="M57" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">T</mml:mi><mml:mi mathvariant="normal">r</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> requiring reconstruction is concatenated with the list of the encoded tokens <inline-formula><mml:math id="M58" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="bold">T</mml:mi><mml:mi mathvariant="normal">c</mml:mi><mml:mi mathvariant="script">E</mml:mi></mml:msubsup></mml:mrow></mml:math></inline-formula>. The set of all tokens is again summed with the spatio-temporal positional embedding and passed through a ViT decoder <inline-formula><mml:math id="M59" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mi mathvariant="normal">ViT</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, producing the decoded tokens <inline-formula><mml:math id="M60" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold">T</mml:mi><mml:mi mathvariant="script">D</mml:mi></mml:msup></mml:mrow></mml:math></inline-formula>. The decoded tokens not corresponding to the central observation <inline-formula><mml:math id="M61" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> are removed from <inline-formula><mml:math id="M62" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold">T</mml:mi><mml:mi mathvariant="script">D</mml:mi></mml:msup></mml:mrow></mml:math></inline-formula>, resulting in <inline-formula><mml:math id="M63" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="bold">T</mml:mi><mml:mi mathvariant="normal">t</mml:mi><mml:mi mathvariant="script">D</mml:mi></mml:msubsup></mml:mrow></mml:math></inline-formula>. Tokens in <inline-formula><mml:math id="M64" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="bold">T</mml:mi><mml:mi mathvariant="normal">t</mml:mi><mml:mi mathvariant="script">D</mml:mi></mml:msubsup></mml:mrow></mml:math></inline-formula> are then linearly projected into <inline-formula><mml:math id="M65" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">8</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> vectors and reshaped into <inline-formula><mml:math id="M66" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">8</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">8</mml:mn></mml:mrow></mml:math></inline-formula> patches. Finally, the patches are reassembled into a grid to form the coarse reconstruction <inline-formula><mml:math id="M67" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. All pixel values corresponding to land areas are set to zero using the land mask <inline-formula><mml:math id="M68" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">l</mml:mi></mml:msub><mml:mo>∈</mml:mo><mml:mo mathvariant="italic">{</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:msup><mml:mo mathvariant="italic">}</mml:mo><mml:mrow><mml:mi>W</mml:mi><mml:mo>×</mml:mo><mml:mi>H</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> that accompanies the data.</p>
</sec>
<sec id="Ch1.S3.SS2">
  <label>3.2</label><title>Iterative refinement module (IRM)</title>
      <p id="d2e1550">To improve the reconstruction accuracy, the coarse reconstruction <inline-formula><mml:math id="M69" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is refined by an iterative refinement module (IRM, Fig. <xref ref-type="fig" rid="F2"/>) through a sequence of residual improvements, producing the final reconstruction <inline-formula><mml:math id="M70" display="inline"><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal" stretchy="false">̃</mml:mo></mml:mover></mml:math></inline-formula> and the corresponding uncertainty characterized by the variance <inline-formula><mml:math id="M71" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>. Per pixel <inline-formula><mml:math id="M72" display="inline"><mml:mi>j</mml:mi></mml:math></inline-formula>, we model the reconstructed SST as a Gaussian distribution parameterized by predicted mean <inline-formula><mml:math id="M73" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal" stretchy="false">̃</mml:mo></mml:mover><mml:mrow><mml:mo>(</mml:mo><mml:mi>j</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> and standard deviation <inline-formula><mml:math id="M74" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>j</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>, following <xref ref-type="bibr" rid="bib1.bibx2" id="text.43"/>. Note that <inline-formula><mml:math id="M75" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> emerges from training the model to minimize Eq. (<xref ref-type="disp-formula" rid="Ch1.E4"/>), which penalizes over- and underestimation of the error variance <inline-formula><mml:math id="M76" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>.</p>
      <p id="d2e1660">Let <inline-formula><mml:math id="M77" display="inline"><mml:mrow><mml:msup><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo stretchy="false" mathvariant="normal">̃</mml:mo></mml:mover><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M78" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> be the reconstruction of the observation field <inline-formula><mml:math id="M79" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and its estimated uncertainty at <inline-formula><mml:math id="M80" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula>th refinement iteration. An iteration of IRM proceeds as follows. The input observation fields <inline-formula><mml:math id="M81" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">X</mml:mi><mml:mi>m</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mo>[</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>]</mml:mo><mml:mo>∈</mml:mo><mml:msup><mml:mi mathvariant="double-struck">R</mml:mi><mml:mrow><mml:mn mathvariant="normal">3</mml:mn><mml:mo>×</mml:mo><mml:mi>W</mml:mi><mml:mo>×</mml:mo><mml:mi>H</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> and the refined estimates <inline-formula><mml:math id="M82" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:msup><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal" stretchy="false">̃</mml:mo></mml:mover><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>]</mml:mo><mml:mo>∈</mml:mo><mml:msup><mml:mi mathvariant="double-struck">R</mml:mi><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mo>×</mml:mo><mml:mi>W</mml:mi><mml:mo>×</mml:mo><mml:mi>H</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> from the previous iteration are concatenated channel-wise and passed to a residual estimation network REN<sup>(<italic>i</italic>)</sup> (detailed in Sect. “Residual estimation network (REN)”) alongside the tokens <inline-formula><mml:math id="M84" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="bold">T</mml:mi><mml:mi mathvariant="normal">t</mml:mi><mml:mi mathvariant="script">D</mml:mi></mml:msubsup></mml:mrow></mml:math></inline-formula> produced by CRM, to produce a two-channel output <inline-formula><mml:math id="M85" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold">Y</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mo>[</mml:mo><mml:msubsup><mml:mi mathvariant="bold">Y</mml:mi><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi mathvariant="bold">Y</mml:mi><mml:mn mathvariant="normal">2</mml:mn><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>]</mml:mo><mml:mo>∈</mml:mo><mml:msup><mml:mi mathvariant="double-struck">R</mml:mi><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mo>×</mml:mo><mml:mi>W</mml:mi><mml:mo>×</mml:mo><mml:mi>H</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>. Following the formulation of <xref ref-type="bibr" rid="bib1.bibx2" id="text.44"/>, <inline-formula><mml:math id="M86" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold">Y</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mo>[</mml:mo><mml:msubsup><mml:mi mathvariant="bold">Y</mml:mi><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi mathvariant="bold">Y</mml:mi><mml:mn mathvariant="normal">2</mml:mn><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> are decoded into reconstruction <inline-formula><mml:math id="M87" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="bold-italic">δ</mml:mi><mml:mi>x</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula> and uncertainty <inline-formula><mml:math id="M88" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="bold-italic">δ</mml:mi><mml:mrow><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula> residuals:

                <disp-formula specific-use="gather" content-type="numbered"><mml:math id="M89" display="block"><mml:mtable displaystyle="true"><mml:mlabeledtr id="Ch1.E1"><mml:mtd><mml:mtext>1</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:msubsup><mml:mi mathvariant="bold-italic">δ</mml:mi><mml:mrow><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:mtext>max</mml:mtext><mml:mo>(</mml:mo><mml:mtext>exp</mml:mtext><mml:mo>(</mml:mo><mml:mtext>min</mml:mtext><mml:mo>(</mml:mo><mml:msubsup><mml:mi mathvariant="bold">Y</mml:mi><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>)</mml:mo><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr><mml:mlabeledtr id="Ch1.E2"><mml:mtd><mml:mtext>2</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:msubsup><mml:mi mathvariant="bold-italic">δ</mml:mi><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msubsup><mml:mi mathvariant="bold">Y</mml:mi><mml:mn mathvariant="normal">2</mml:mn><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>⊙</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">δ</mml:mi><mml:mrow><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula>

          where <inline-formula><mml:math id="M90" display="inline"><mml:mo>⊙</mml:mo></mml:math></inline-formula> denotes element-wise tensor multiplication (the Hadamard product), while <inline-formula><mml:math id="M91" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M92" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M93" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>&gt;</mml:mo><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula> are hyperparameters ensuring training stability. The reconstruction and uncertainty estimates at iteration <inline-formula><mml:math id="M94" display="inline"><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula> are initialized with the coarse reconstruction <inline-formula><mml:math id="M95" display="inline"><mml:mrow><mml:msup><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal" stretchy="false">̃</mml:mo></mml:mover><mml:mrow><mml:mo>(</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and a zero <inline-formula><mml:math id="M96" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mo>(</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mn mathvariant="bold">0</mml:mn></mml:mrow></mml:math></inline-formula>.  The reconstruction and uncertainty estimated at the <inline-formula><mml:math id="M97" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>th refinement iteration are thus <inline-formula><mml:math id="M98" display="inline"><mml:mrow><mml:msup><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal" stretchy="false">̃</mml:mo></mml:mover><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:msup><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal" stretchy="false">̃</mml:mo></mml:mover><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">δ</mml:mi><mml:mi>x</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M99" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">δ</mml:mi><mml:mrow><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula>, respectively. IRM runs for <inline-formula><mml:math id="M100" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">IRM</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> iterations, with each REN<sup>(<italic>i</italic>)</sup> having its own set of trained parameters, allowing each to specialize to its respective residual estimation, finally producing the refined reconstruction <inline-formula><mml:math id="M102" display="inline"><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo stretchy="false" mathvariant="normal">̃</mml:mo></mml:mover></mml:math></inline-formula> and uncertainty <inline-formula><mml:math id="M103" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>.</p>
<sec id="Ch1.S3.SS2.SSSx1" specific-use="unnumbered">
  <title>Residual estimation network (REN)</title>
      <p id="d2e2442">The residual estimation network REN<sup>(<italic>i</italic>)</sup> is a UNet-type architecture <xref ref-type="bibr" rid="bib1.bibx34" id="paren.45"/>. The encoder <inline-formula><mml:math id="M105" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="script">E</mml:mi><mml:mi mathvariant="normal">REN</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> takes the reconstruction and uncertainty estimates <inline-formula><mml:math id="M106" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:msup><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal" stretchy="false">̃</mml:mo></mml:mover><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> as well as the observation fields <inline-formula><mml:math id="M107" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">X</mml:mi><mml:mi>m</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mo>[</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> as input and produces the latent features <inline-formula><mml:math id="M108" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> with an <inline-formula><mml:math id="M109" display="inline"><mml:mn mathvariant="normal">8</mml:mn></mml:math></inline-formula>-fold reduction in spatial resolution compared to the input. The latent features are then enriched with spatio-temporally aggregated features <inline-formula><mml:math id="M110" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="bold">T</mml:mi><mml:mi mathvariant="normal">t</mml:mi><mml:mi mathvariant="script">D</mml:mi></mml:msubsup></mml:mrow></mml:math></inline-formula> from CRM. Specifically, the tokens <inline-formula><mml:math id="M111" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="bold">T</mml:mi><mml:mi mathvariant="normal">t</mml:mi><mml:mi mathvariant="script">D</mml:mi></mml:msubsup></mml:mrow></mml:math></inline-formula> (see Fig. <xref ref-type="fig" rid="F2"/>) are spatially reshaped and bilinearly upsampled to match the dimensions of <inline-formula><mml:math id="M112" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>. The two tensors are concatenated and fused by the feature fusion module (FFM) <xref ref-type="bibr" rid="bib1.bibx40" id="paren.46"/>, yielding the enriched bottleneck features <inline-formula><mml:math id="M113" display="inline"><mml:mrow><mml:msup><mml:mover accent="true"><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mo mathvariant="normal" stretchy="false">̃</mml:mo></mml:mover><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>.</p>
      <p id="d2e2647">The resulting features are then input in the decoder <inline-formula><mml:math id="M114" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mi mathvariant="normal">REN</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and decoded to the same dimensions as the input <inline-formula><mml:math id="M115" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:msup><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo stretchy="false" mathvariant="normal">̃</mml:mo></mml:mover><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> via convolutional and upsampling blocks, while incorporating intermediate encoder features at multiple scales through UNet skip connections. The resulting decoded features are transformed with two <inline-formula><mml:math id="M116" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> convolutional layers to produce a two-channel output <inline-formula><mml:math id="M117" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold">Y</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>∈</mml:mo><mml:msup><mml:mi mathvariant="double-struck">R</mml:mi><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mo>×</mml:mo><mml:mi>W</mml:mi><mml:mo>×</mml:mo><mml:mi>H</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>.</p>
</sec>
</sec>
<sec id="Ch1.S3.SS3">
  <label>3.3</label><title>Training strategy</title>
      <p id="d2e2751">CRITER is trained in two stages to train both CRM and IRM. First, supervised learning with automatically generated targets is used to train CRM. In this setup, part of the input signal is deleted, and the network is trained to reconstruct the entire input signal. The input training samples are created by sampling triples of consecutive observations <inline-formula><mml:math id="M118" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> and deleting parts of the central observation <inline-formula><mml:math id="M119" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>,  resulting in <inline-formula><mml:math id="M120" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>⊙</mml:mo><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">m</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>, where <inline-formula><mml:math id="M121" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">m</mml:mi></mml:msub><mml:mo>∈</mml:mo><mml:mo mathvariant="italic">{</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:msup><mml:mo mathvariant="italic">}</mml:mo><mml:mrow><mml:mi>W</mml:mi><mml:mo>×</mml:mo><mml:mi>H</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> is a generated binary mask with 0 corresponding to missing values. Following <xref ref-type="bibr" rid="bib1.bibx3" id="text.47"/>, the masks <inline-formula><mml:math id="M122" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">m</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> are generated by copying clouds from a random day not included in the triplet to maintain mask simulation realism. CRM is trained to minimize the following reconstruction error:

            <disp-formula id="Ch1.E3" content-type="numbered"><label>3</label><mml:math id="M123" display="block"><mml:mrow><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:mi mathvariant="normal">CRM</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:msub><mml:mo>⊙</mml:mo><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">l</mml:mi></mml:msub><mml:mo>|</mml:mo></mml:mrow></mml:mfrac></mml:mstyle><mml:msubsup><mml:mo>∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:msubsup><mml:mfenced open="[" close="]"><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mrow><mml:mi>t</mml:mi><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msub><mml:msup><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mrow><mml:mi mathvariant="normal">t</mml:mi><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msub><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mrow><mml:mi mathvariant="normal">l</mml:mi><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msub></mml:mrow></mml:mfenced><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

          where <inline-formula><mml:math id="M124" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the coarse reconstruction generated by CRM, and mask <inline-formula><mml:math id="M125" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> has zeros at locations where ground truth measurements within the observation field <inline-formula><mml:math id="M126" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> are missing, while <inline-formula><mml:math id="M127" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">l</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> has zeros at spatial locations belonging to land, and <inline-formula><mml:math id="M128" display="inline"><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:msub><mml:mo>⊙</mml:mo><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">l</mml:mi></mml:msub><mml:mo>|</mml:mo></mml:mrow></mml:math></inline-formula> denotes the number of ground truth measurements. The summation goes over the <inline-formula><mml:math id="M129" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula> pixels in each of <inline-formula><mml:math id="M130" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M131" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M132" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, and <inline-formula><mml:math id="M133" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">l</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. The operator <inline-formula><mml:math id="M134" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mo>⋅</mml:mo><mml:msub><mml:mo>)</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> indexes the <inline-formula><mml:math id="M135" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula>th element of a matrix. The consecutive observations used as the model input and the masks <inline-formula><mml:math id="M136" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M137" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">l</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, and <inline-formula><mml:math id="M138" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">m</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> used in the training process are visualized in Fig. <xref ref-type="fig" rid="F3"/>.</p>
      <p id="d2e3195">In the second stage, the parameters of CRM are fixed and only the parameters of IRM are trained. The training samples are generated as in CRM training, but since IRM produces the mean and variance of the reconstruction, the following negative log-likelihood loss is minimized as in DINCAE <xref ref-type="bibr" rid="bib1.bibx2 bib1.bibx3" id="paren.48"/>:

            <disp-formula id="Ch1.E4" content-type="numbered"><label>4</label><mml:math id="M139" display="block"><mml:mtable rowspacing="0.2ex" class="split" displaystyle="true" columnalign="right left"><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:mi mathvariant="normal">IRM</mml:mi></mml:msub></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:msub><mml:mo>⊙</mml:mo><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">l</mml:mi></mml:msub><mml:mo>|</mml:mo></mml:mrow></mml:mfrac></mml:mstyle><mml:msubsup><mml:mo>∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:msubsup></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mfenced open="[" close="]"><mml:mrow><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo stretchy="false" mathvariant="normal">̃</mml:mo></mml:mover><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msub><mml:msup><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:msubsup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>+</mml:mo><mml:mi>log⁡</mml:mi><mml:mo>(</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:mfenced><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mrow><mml:mi mathvariant="normal">t</mml:mi><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msub><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mrow><mml:mi mathvariant="normal">l</mml:mi><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>

          where <inline-formula><mml:math id="M140" display="inline"><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo stretchy="false" mathvariant="normal">̃</mml:mo></mml:mover></mml:math></inline-formula> and <inline-formula><mml:math id="M141" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> are the reconstruction and variance estimated after the last iteration in IRM, and the summation goes over the <inline-formula><mml:math id="M142" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula> pixels in each of <inline-formula><mml:math id="M143" display="inline"><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo stretchy="false" mathvariant="normal">̃</mml:mo></mml:mover></mml:math></inline-formula>, <inline-formula><mml:math id="M144" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>, and <inline-formula><mml:math id="M145" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. This loss thus trains the model to assign higher variance to areas with greater than expected reconstruction error. We validate the variance prediction quality in Sect. <xref ref-type="sec" rid="Ch1.S4.SS4"/>,  by demonstrating its correlation with empirical errors.</p>

      <fig id="F3" specific-use="star"><label>Figure 3</label><caption><p id="d2e3419">Top row: a sequence of three consecutive observation fields <inline-formula><mml:math id="M146" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>, and the central observation <inline-formula><mml:math id="M147" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>⊙</mml:mo><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">m</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, with additional missing values deleted by the sampled mask <inline-formula><mml:math id="M148" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">m</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. Bottom row: the land mask <inline-formula><mml:math id="M149" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">l</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> with zeros at land locations; the missing data mask <inline-formula><mml:math id="M150" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> with zeros at locations with missing measurements in <inline-formula><mml:math id="M151" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>; and <inline-formula><mml:math id="M152" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">m</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, which is a randomly sampled <inline-formula><mml:math id="M153" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> from an observation field not included in the input.</p></caption>
          <graphic xlink:href="https://gmd.copernicus.org/articles/18/5549/2025/gmd-18-5549-2025-f03.jpg"/>

        </fig>

</sec>
<sec id="Ch1.S3.SS4">
  <label>3.4</label><title>Implementation details</title>
      <p id="d2e3556">CRM (Sect. <xref ref-type="sec" rid="Ch1.S3.SS1"/>) consists of <inline-formula><mml:math id="M154" display="inline"><mml:mn mathvariant="normal">12</mml:mn></mml:math></inline-formula> encoder and decoder transformer blocks, with <inline-formula><mml:math id="M155" display="inline"><mml:mn mathvariant="normal">3</mml:mn></mml:math></inline-formula> multi-head attention (MHA) heads, a token dimension of <inline-formula><mml:math id="M156" display="inline"><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">192</mml:mn></mml:mrow></mml:math></inline-formula>, and a patch size of <inline-formula><mml:math id="M157" display="inline"><mml:mrow><mml:mn mathvariant="normal">3</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">8</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">8</mml:mn></mml:mrow></mml:math></inline-formula>, where <inline-formula><mml:math id="M158" display="inline"><mml:mn mathvariant="normal">3</mml:mn></mml:math></inline-formula> denotes the number of channels, while <inline-formula><mml:math id="M159" display="inline"><mml:mrow><mml:mn mathvariant="normal">8</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">8</mml:mn></mml:mrow></mml:math></inline-formula> represents the width and height, respectively. IRM (Sect. <xref ref-type="sec" rid="Ch1.S3.SS2"/>) consists of a CNN-based encoder with <inline-formula><mml:math id="M160" display="inline"><mml:mn mathvariant="normal">3</mml:mn></mml:math></inline-formula> <italic>double conv</italic> blocks, each followed by a <inline-formula><mml:math id="M161" display="inline"><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula> max pooling operation. The <italic>double conv</italic> block is composed of two <inline-formula><mml:math id="M162" display="inline"><mml:mrow><mml:mn mathvariant="normal">3</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:math></inline-formula> convolutional layers, each followed by a batch normalization layer and a ReLU activation function. The number of convolutional kernels in each block is <inline-formula><mml:math id="M163" display="inline"><mml:mrow><mml:mn mathvariant="normal">32</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">64</mml:mn></mml:mrow></mml:math></inline-formula>, and <inline-formula><mml:math id="M164" display="inline"><mml:mn mathvariant="normal">128</mml:mn></mml:math></inline-formula>, respectively. This is followed by another <italic>double conv</italic> block, with <inline-formula><mml:math id="M165" display="inline"><mml:mn mathvariant="normal">256</mml:mn></mml:math></inline-formula> kernels, at the bottleneck of the network, a Feature Fusion Module (FFM), and a decoder with <inline-formula><mml:math id="M166" display="inline"><mml:mn mathvariant="normal">3</mml:mn></mml:math></inline-formula> transpose convolution layers, each followed by a concatenation based skip connection and a <italic>double conv</italic> block. The number of kernels in each block is  <inline-formula><mml:math id="M167" display="inline"><mml:mrow><mml:mn mathvariant="normal">128</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">64</mml:mn></mml:mrow></mml:math></inline-formula>, and <inline-formula><mml:math id="M168" display="inline"><mml:mn mathvariant="normal">32</mml:mn></mml:math></inline-formula>, respectively.  IRM utilizes <inline-formula><mml:math id="M169" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mtext>IRM</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:math></inline-formula> refinement iterations – this value is selected based on the results of the ablation study in Sect. <xref ref-type="sec" rid="Ch1.S4.SS5.SSS4"/>. Hyperparameters  <inline-formula><mml:math id="M170" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M171" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> are set as <inline-formula><mml:math id="M172" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi mathvariant="italic">θ</mml:mi><mml:mo stretchy="false" mathvariant="normal">̃</mml:mo></mml:mover><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mtext>ln</mml:mtext><mml:mo>(</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mtext>IRM</mml:mtext></mml:msub><mml:mo>)</mml:mo><mml:mo>+</mml:mo><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M173" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi mathvariant="italic">θ</mml:mi><mml:mo stretchy="false" mathvariant="normal">̃</mml:mo></mml:mover><mml:mn mathvariant="normal">2</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mtext>IRM</mml:mtext></mml:msub><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> to ensure that the variance <inline-formula><mml:math id="M174" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> is bounded between <inline-formula><mml:math id="M175" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:mtext>exp</mml:mtext><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M176" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> for an arbitrary number of refinement iterations <inline-formula><mml:math id="M177" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mtext>IRM</mml:mtext></mml:msub><mml:mo>≥</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula>.</p>
</sec>
</sec>
<sec id="Ch1.S4">
  <label>4</label><title>Results</title>
<sec id="Ch1.S4.SS1">
  <label>4.1</label><title>Implementation details</title>
      <p id="d2e3905">CRITER is implemented using the PyTorch library <xref ref-type="bibr" rid="bib1.bibx31" id="paren.49"/> and trained on an NVIDIA Tesla V100 GPU. The CRM block is trained with a batch size of <inline-formula><mml:math id="M178" display="inline"><mml:mn mathvariant="normal">8</mml:mn></mml:math></inline-formula> using the AdamW optimizer with a learning rate <inline-formula><mml:math id="M179" display="inline"><mml:mrow><mml:mi mathvariant="italic">α</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">3</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">4</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M180" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">β</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.9</mml:mn></mml:mrow></mml:math></inline-formula>, and <inline-formula><mml:math id="M181" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">β</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.95</mml:mn></mml:mrow></mml:math></inline-formula> for <inline-formula><mml:math id="M182" display="inline"><mml:mn mathvariant="normal">60</mml:mn></mml:math></inline-formula> epochs (warm-up period) and then with a cosine decay scheduler <xref ref-type="bibr" rid="bib1.bibx25" id="paren.50"/> with step size <inline-formula><mml:math id="M183" display="inline"><mml:mn mathvariant="normal">30</mml:mn></mml:math></inline-formula> for another <inline-formula><mml:math id="M184" display="inline"><mml:mn mathvariant="normal">140</mml:mn></mml:math></inline-formula> epochs. In the next phase the IRM block is trained using the pre-trained CRM with fixed parameters. We train IRM using the Adam optimizer, with <inline-formula><mml:math id="M185" display="inline"><mml:mrow><mml:mi mathvariant="italic">α</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">3</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">4</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M186" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">β</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.9</mml:mn></mml:mrow></mml:math></inline-formula>, and <inline-formula><mml:math id="M187" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">β</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.999</mml:mn></mml:mrow></mml:math></inline-formula> for <inline-formula><mml:math id="M188" display="inline"><mml:mn mathvariant="normal">300</mml:mn></mml:math></inline-formula> epochs, using a step learning rate scheduler with step size <inline-formula><mml:math id="M189" display="inline"><mml:mn mathvariant="normal">50</mml:mn></mml:math></inline-formula> and multiplicative factor <inline-formula><mml:math id="M190" display="inline"><mml:mrow><mml:mi mathvariant="italic">γ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.5</mml:mn></mml:mrow></mml:math></inline-formula>.</p>
</sec>
<sec id="Ch1.S4.SS2">
  <label>4.2</label><title>Performance measures</title>
      <p id="d2e4082">The performance of CRITER is assessed on an independent test set. Reconstruction quality is computed in terms of root-mean-squared error (RMSE) between the ground truth  <inline-formula><mml:math id="M191" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and the reconstruction <inline-formula><mml:math id="M192" display="inline"><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal" stretchy="false">̃</mml:mo></mml:mover></mml:math></inline-formula>. In particular, the overall reconstruction error RMSE<sub>all</sub> is defined as

            <disp-formula id="Ch1.E5" content-type="numbered"><label>5</label><mml:math id="M194" display="block"><mml:mrow><mml:msub><mml:mtext>RMSE</mml:mtext><mml:mtext>all</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:msqrt><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msubsup><mml:mo>∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:msubsup><mml:mfenced close="]" open="["><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal" stretchy="false">̃</mml:mo></mml:mover><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msub><mml:msup><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mrow><mml:mi mathvariant="normal">t</mml:mi><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msub><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mrow><mml:mi mathvariant="normal">l</mml:mi><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msub></mml:mrow></mml:mfenced></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:msub><mml:mo>⊙</mml:mo><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">l</mml:mi></mml:msub><mml:mo>|</mml:mo></mml:mrow></mml:mfrac></mml:mstyle></mml:msqrt><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula></p>
      <p id="d2e4222">For additional insights we compute the RMSE separately for (i) deleted regions, corresponding to observations artificially removed by simulated clouds in the L3 SST product and thus withheld during the training, and (ii) visible regions, corresponding to remaining observations post-deletion in <inline-formula><mml:math id="M195" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>.</p>
      <p id="d2e4236">The reconstruction error of deleted regions is defined as

            <disp-formula id="Ch1.E6" content-type="numbered"><label>6</label><mml:math id="M196" display="block"><mml:mrow><mml:msub><mml:mtext mathvariant="normal">RMSE</mml:mtext><mml:mtext>mis</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:msqrt><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msubsup><mml:mo>∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:msubsup><mml:mfenced close="]" open="["><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal" stretchy="false">̃</mml:mo></mml:mover><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msub><mml:msup><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mrow><mml:mi mathvariant="normal">t</mml:mi><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msub><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mrow><mml:mi mathvariant="normal">l</mml:mi><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mn mathvariant="bold">1</mml:mn><mml:mo>-</mml:mo><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mrow><mml:mi mathvariant="normal">m</mml:mi><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mfenced></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:msub><mml:mo>⊙</mml:mo><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">l</mml:mi></mml:msub><mml:mo>⊙</mml:mo><mml:mo>(</mml:mo><mml:mn mathvariant="bold">1</mml:mn><mml:mo>-</mml:mo><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">m</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>|</mml:mo></mml:mrow></mml:mfrac></mml:mstyle></mml:msqrt><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

          <inline-formula><mml:math id="M197" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">m</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the mask of deleted regions, and <inline-formula><mml:math id="M198" display="inline"><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:msub><mml:mo>⊙</mml:mo><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">l</mml:mi></mml:msub><mml:mo>⊙</mml:mo><mml:mo>(</mml:mo><mml:mn mathvariant="bold">1</mml:mn><mml:mo>-</mml:mo><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">m</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>|</mml:mo></mml:mrow></mml:math></inline-formula> denotes the number of deleted ground truth measurements. The reconstruction error of visible regions is defined as

            <disp-formula id="Ch1.E7" content-type="numbered"><label>7</label><mml:math id="M199" display="block"><mml:mrow><mml:msub><mml:mtext>RMSE</mml:mtext><mml:mtext>vis</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:msqrt><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msubsup><mml:mo>∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:msubsup><mml:mfenced close="]" open="["><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal" stretchy="false">̃</mml:mo></mml:mover><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msub><mml:msup><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mrow><mml:mi mathvariant="normal">t</mml:mi><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msub><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mrow><mml:mi mathvariant="normal">l</mml:mi><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msub><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mrow><mml:mi mathvariant="normal">m</mml:mi><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msub></mml:mrow></mml:mfenced></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:msub><mml:mo>⊙</mml:mo><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">l</mml:mi></mml:msub><mml:mo>⊙</mml:mo><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">m</mml:mi></mml:msub><mml:mo>|</mml:mo></mml:mrow></mml:mfrac></mml:mstyle></mml:msqrt><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

          where <inline-formula><mml:math id="M200" display="inline"><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:msub><mml:mo>⊙</mml:mo><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">l</mml:mi></mml:msub><mml:mo>⊙</mml:mo><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">m</mml:mi></mml:msub><mml:mo>|</mml:mo></mml:mrow></mml:math></inline-formula> is the number of visible ground truth measurements. To enhance the metric stability, we sample <inline-formula><mml:math id="M201" display="inline"><mml:mn mathvariant="normal">10</mml:mn></mml:math></inline-formula> distinct cloud masks for each test SST field, simulating realistic observational variability. We thus evaluate the performance on <inline-formula><mml:math id="M202" display="inline"><mml:mn mathvariant="normal">2560</mml:mn></mml:math></inline-formula>, <inline-formula><mml:math id="M203" display="inline"><mml:mn mathvariant="normal">3900</mml:mn></mml:math></inline-formula>, and <inline-formula><mml:math id="M204" display="inline"><mml:mn mathvariant="normal">1720</mml:mn></mml:math></inline-formula> masked SST fields for the respective regions, ensuring robust statistical validation.</p>
</sec>
<sec id="Ch1.S4.SS3">
  <label>4.3</label><title>Comparison with state of the art</title>
      <p id="d2e4626">We compare CRITER with DINCAE2 <xref ref-type="bibr" rid="bib1.bibx3" id="paren.51"/>, a well-known and highly competitive SST reconstruction method, serving as a widely recognized benchmark in recent studies <xref ref-type="bibr" rid="bib1.bibx4" id="paren.52"/>, and with the recently presented MAESSTRO <xref ref-type="bibr" rid="bib1.bibx19" id="paren.53"/> on the three datasets from Sect. <xref ref-type="sec" rid="Ch1.S2.SS1"/>. We reimplemented both DINCAE2 (originally in Julia) following <xref ref-type="bibr" rid="bib1.bibx3" id="text.54"/> and MAESSTRO (public implementation unavailable) following <xref ref-type="bibr" rid="bib1.bibx19" id="text.55"/> in Pytorch. To ensure a fair evaluation, both methods were trained using the same dataset splits, with tuned hyperparameters, and employed the same loss function computed over identical regions to CRITER. For MAESSTRO, architectural modifications were necessary to ensure comparability. Please refer to Appendix <xref ref-type="sec" rid="App1.Ch1.S4"/> for the implementation details of baseline models.</p>
      <p id="d2e4649">Results in Table <xref ref-type="table" rid="T1"/> demonstrate CRITER's consistent superior performance across all datasets. Compared to the current state of the art, DINCAE2, CRITER achieves error reductions in deleted and visible regions of 20 % and 89 % for the Mediterranean, 44 % and 80 % for the Adriatic, and 1 % and 88 % for the Atlantic dataset, respectively. MAESSTRO's significantly lower performance is attributed to its single time step reconstruction approach. This hypothesis is confirmed by our ablation study, detailed in Sect. <xref ref-type="sec" rid="Ch1.S4.SS5"/>, which examines the importance of modeling spatio-temporal data dependencies.</p>
      <p id="d2e4656">The relative improvements of CRITER compared to the related methods vary across the datasets. This can be attributed to the differing amounts of information available for reconstruction, which is inversely proportional with the extent of missing values. Our analysis of missing values (Appendix <xref ref-type="sec" rid="App1.Ch1.S1"/>) reveals that the datasets can be ranked by the average amount of information available in each observation triplet, from highest to lowest: Mediterranean, Adriatic, and Atlantic. Notably, the Adriatic dataset shows the greatest decrease in reconstruction error, suggesting that CRITER achieves optimal improvement when the available information is moderate. In contrast, the Atlantic dataset, with the lowest amount of available information, likely requires additional data to be effectively reconstructed. To address this, we propose increasing the temporal horizon <inline-formula><mml:math id="M205" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and incorporating supplementary or proxy variables, such as chlorophyll <inline-formula><mml:math id="M206" display="inline"><mml:mi>a</mml:mi></mml:math></inline-formula> and surface winds. We leave the exploration of this approach to future work.</p>

<table-wrap id="T1" specific-use="star"><label>Table 1</label><caption><p id="d2e4683">Comparison of CRITER, DINCAE2, and MAESSTRO. We report the overall reconstruction error (RMSE<sub>all</sub>), as well as the error over deleted (RMSE<sub>mis</sub>) and observed regions (RMSE<sub>vis</sub>), where the two numbers in parentheses correspond to the 10 % and 90 % percentiles of the error. Bold font indicates the best result for each metric and dataset.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="5">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="left"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:colspec colnum="5" colname="col5" align="right"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Dataset</oasis:entry>
         <oasis:entry colname="col2">Model</oasis:entry>
         <oasis:entry colname="col3">RMSE<sub>all</sub> (°C)</oasis:entry>
         <oasis:entry colname="col4">RMSE<sub>mis</sub> (°C)</oasis:entry>
         <oasis:entry colname="col5">RMSE<sub>vis</sub>  (°C)</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">Mediterranean</oasis:entry>
         <oasis:entry colname="col2">MAESSTRO</oasis:entry>
         <oasis:entry colname="col3">0.487 (0.320, 0.657)</oasis:entry>
         <oasis:entry colname="col4">0.607 (0.394, 0.856)</oasis:entry>
         <oasis:entry colname="col5">0.434 (0.299, 0.564)</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">DINCAE2</oasis:entry>
         <oasis:entry colname="col3">0.209 (0.140, 0.300)</oasis:entry>
         <oasis:entry colname="col4">0.319 (0.226, 0.418)</oasis:entry>
         <oasis:entry colname="col5">0.148 (0.112, 0.184)</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">CRITER (ours)</oasis:entry>
         <oasis:entry colname="col3"><bold>0.127 (0.037, 0.235)</bold></oasis:entry>
         <oasis:entry colname="col4"><bold>0.255 (0.168, 0.352)</bold></oasis:entry>
         <oasis:entry colname="col5"><bold>0.017 (0.013, 0.021)</bold></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Adriatic</oasis:entry>
         <oasis:entry colname="col2">MAESSTRO</oasis:entry>
         <oasis:entry colname="col3">0.456 (0.296, 0.635)</oasis:entry>
         <oasis:entry colname="col4">0.583 (0.362, 0.844)</oasis:entry>
         <oasis:entry colname="col5">0.392 (0.261, 0.539)</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">DINCAE2</oasis:entry>
         <oasis:entry colname="col3">0.270 (0.111, 0.522)</oasis:entry>
         <oasis:entry colname="col4">0.433 (0.203, 0.769)</oasis:entry>
         <oasis:entry colname="col5">0.106 (0.087, 0.129)</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">CRITER (ours)</oasis:entry>
         <oasis:entry colname="col3"><bold>0.130 (0.045, 0.222)</bold></oasis:entry>
         <oasis:entry colname="col4"><bold>0.243 (0.140, 0.358)</bold></oasis:entry>
         <oasis:entry colname="col5"><bold>0.021 (0.014, 0.030)</bold></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Atlantic</oasis:entry>
         <oasis:entry colname="col2">MAESSTRO</oasis:entry>
         <oasis:entry colname="col3">0.802 (0.508, 1.239)</oasis:entry>
         <oasis:entry colname="col4">0.832 (0.514, 1.301)</oasis:entry>
         <oasis:entry colname="col5">0.764 (0.479, 1.137)</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">DINCAE2</oasis:entry>
         <oasis:entry colname="col3">0.444 (0.332, 0.581)</oasis:entry>
         <oasis:entry colname="col4">0.525 (0.396, 0.692)</oasis:entry>
         <oasis:entry colname="col5">0.302 (0.236, 0.364)</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">CRITER (ours)</oasis:entry>
         <oasis:entry colname="col3"><bold>0.391 (0.249, 0.542)</bold></oasis:entry>
         <oasis:entry colname="col4"><bold>0.518 (0.386, 0.692)</bold></oasis:entry>
         <oasis:entry colname="col5"><bold>0.036 (0.019, 0.046)</bold></oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>


<sec id="Ch1.S4.SS3.SSS1">
  <label>4.3.1</label><title>Qualitative comparison</title>
      <p id="d2e4955">For further insights we visualize the CRITER and DINCAE2 reconstructions in Figs. <xref ref-type="fig" rid="F4"/> and <xref ref-type="fig" rid="F5"/>. We showcase examples from the Mediterranean and the Adriatic test set, respectively, highlighting the masked SST (<inline-formula><mml:math id="M213" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>⊙</mml:mo><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">m</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>), target SST (<inline-formula><mml:math id="M214" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>), full reconstruction (<inline-formula><mml:math id="M215" display="inline"><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo stretchy="false" mathvariant="normal">̃</mml:mo></mml:mover></mml:math></inline-formula>), standard deviation (<inline-formula><mml:math id="M216" display="inline"><mml:mi mathvariant="bold-italic">σ</mml:mi></mml:math></inline-formula>), and RMSE computed over the entire target (RMSE<sub>all</sub>). Notice that CRITER preserves fine details in cloud-free regions, ensuring minimal distortion of the original input data. In contrast, obscured (deleted) regions require the model to infer missing SST values using spatio-temporal context from adjacent days/pixels. These reconstructed regions exhibit reduced sharpness as a result of the inherent uncertainty caused by sparse observations. However, CRITER demonstrates an excellent ability to reconstruct high-frequency components of the target SST under deleted regions compared to DINCAE2. Additionally, CRITER proves robust to clouds of arbitrary shape, whether small and scattered (Fig. <xref ref-type="fig" rid="F4"/>, first and last comparison) or large and contiguous (Fig. <xref ref-type="fig" rid="F4"/>, second and third comparisons). Similar observations can be drawn from the comparisons on the Adriatic dataset presented in Fig. <xref ref-type="fig" rid="F5"/>. On the Atlantic test set, both models face challenges in reconstructing high-frequency components under deleted regions, as illustrated in Fig. <xref ref-type="fig" rid="F6"/>. However, we observe that CRITER is able to preserve the SST measurements over visible regions, whereas DINCAE2 introduces significant smoothing. Additional comparison figures are shown in Appendix <xref ref-type="sec" rid="App1.Ch1.S2"/> (Figs. <xref ref-type="fig" rid="FB1"/>, <xref ref-type="fig" rid="FB2"/>, and <xref ref-type="fig" rid="FB3"/>).</p>

      <fig id="F4" specific-use="star"><label>Figure 4</label><caption><p id="d2e5037">Comparison of sea surface temperature (SST) reconstructions generated by CRITER and DINCAE2 on the Mediterranean dataset. The columns display (1) the original SST field with simulated missing values, (2) the original SST field, (3, 4) full reconstruction of the SST field and the associated standard deviation, and (5) the absolute error map, highlighting the differences between the original and reconstructed fields. All panel values are in °C. Note that color scales for <inline-formula><mml:math id="M218" display="inline"><mml:mi mathvariant="bold-italic">σ</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M219" display="inline"><mml:mrow><mml:msub><mml:mtext>RMSE</mml:mtext><mml:mtext>all</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> are truncated at the <inline-formula><mml:math id="M220" display="inline"><mml:mn mathvariant="normal">90</mml:mn></mml:math></inline-formula>th percentile of the data to improve visibility.</p></caption>
            <graphic xlink:href="https://gmd.copernicus.org/articles/18/5549/2025/gmd-18-5549-2025-f04.jpg"/>

          </fig>

      <fig id="F5" specific-use="star"><label>Figure 5</label><caption><p id="d2e5073">Same as Fig. <xref ref-type="fig" rid="F4"/> but for the Adriatic domain.</p></caption>
            <graphic xlink:href="https://gmd.copernicus.org/articles/18/5549/2025/gmd-18-5549-2025-f05.jpg"/>

          </fig>

      <fig id="F6" specific-use="star"><label>Figure 6</label><caption><p id="d2e5087">Same as Fig. <xref ref-type="fig" rid="F4"/> but for the Atlantic domain.</p></caption>
            <graphic xlink:href="https://gmd.copernicus.org/articles/18/5549/2025/gmd-18-5549-2025-f06.jpg"/>

          </fig>

</sec>
<sec id="Ch1.S4.SS3.SSS2">
  <label>4.3.2</label><title>Spatial spectral analysis</title>
      <p id="d2e5106">We conduct spatial spectral analysis by comparing the power spectral density (PSD) of ground truth observations against reconstructions from CRITER and DINCAE2, focusing on the Ionian Sea region due to its significant SST variability.</p>
      <p id="d2e5109">First, we identify observation fields with maximum number of known measurements within the ROI (Region Of Interest) and compute their PSDs over the ROI. Following <xref ref-type="bibr" rid="bib1.bibx15" id="text.56"/>, we compute PSD using a fast Fourier transform (FFT) with a Blackman–Harris window.  We then sample <inline-formula><mml:math id="M221" display="inline"><mml:mn mathvariant="normal">30</mml:mn></mml:math></inline-formula> cloud masks with distinct coverage over the ROI, with the fraction of missing values ranging from 50 % to 98 %. For each mask, we simulate missing data in the observation fields, reconstruct them using both methods, and compute PSD over the reconstructed ROI.</p>
      <p id="d2e5122">Figure <xref ref-type="fig" rid="F7"/> shows an observation sequence with few available measurements. Both methods maintain PSD values near the target at low wavenumbers, indicating comparable low-frequency reconstruction. For wavenumbers <inline-formula><mml:math id="M222" display="inline"><mml:mrow><mml:mi>k</mml:mi><mml:mo>≥</mml:mo><mml:mn mathvariant="normal">4</mml:mn></mml:mrow></mml:math></inline-formula> cycles per degree, however, CRITER's PSD remains closer to the target than DINCAE2's, demonstrating its superior ability to resolve high-frequency components. Figure <xref ref-type="fig" rid="F8"/> depicts a case with more measurements, where both methods generally align closer to the target. Nevertheless, CRITER still outperforms DINCAE2 at high wavenumbers (<inline-formula><mml:math id="M223" display="inline"><mml:mrow><mml:mi>k</mml:mi><mml:mo>≥</mml:mo><mml:mn mathvariant="normal">5</mml:mn><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mtext>cycles</mml:mtext><mml:mi>deg⁡</mml:mi></mml:mfrac></mml:mstyle></mml:mrow></mml:math></inline-formula>). Additional results are provided in Appendix <xref ref-type="sec" rid="App1.Ch1.S3"/>.</p>

      <fig id="F7" specific-use="star"><label>Figure 7</label><caption><p id="d2e5164">Visualization of reconstruction performance. Row 1 shows the full fields (left to right: masked SST, target SST, CRITER reconstruction, and DINCAE2 reconstruction) with the Region of Interest (ROI) marked by a dashed black rectangle. Row 2 displays the corresponding ROI fields: target SST, CRITER reconstruction, and DINCAE2 reconstruction. Row 3 presents gradient magnitudes within the ROI for target, CRITER, and DINCAE2 outputs. Row 4 compares power spectral densities: target ROI (black), CRITER mean <inline-formula><mml:math id="M224" display="inline"><mml:mo>±</mml:mo></mml:math></inline-formula> SD (orange band), and DINCAE2 mean <inline-formula><mml:math id="M225" display="inline"><mml:mo>±</mml:mo></mml:math></inline-formula> SD (blue band), with solid orange and dotted blue lines showing CRITER's and DINCAE2's PSDs for the selected example.</p></caption>
            <graphic xlink:href="https://gmd.copernicus.org/articles/18/5549/2025/gmd-18-5549-2025-f07.png"/>

          </fig>

      <fig id="F8" specific-use="star"><label>Figure 8</label><caption><p id="d2e5189">Same as Fig. <xref ref-type="fig" rid="F7"/> but for another sample.</p></caption>
            <graphic xlink:href="https://gmd.copernicus.org/articles/18/5549/2025/gmd-18-5549-2025-f08.png"/>

          </fig>

</sec>
<sec id="Ch1.S4.SS3.SSS3">
  <label>4.3.3</label><title>Comparison under different cloud coverage levels</title>
      <p id="d2e5208">The qualitative results presented in Sect. <xref ref-type="sec" rid="Ch1.S4.SS3.SSS1"/> suggest that CRITER is robust to clouds of various size. To test this, we compare the reconstruction error of CRITER and DINCAE2 on images with different coverage levels. The cloud coverage is given by the fraction of pixels that are missing or deleted relative to the total number of pixels belonging to the sea. Specifically, we categorize clouds into three distinct groups based on their coverage: low coverage (0 %, 60 %], moderate coverage (60 %, 75 %], and high coverage (75 %, 100 %). We then compute the reconstruction error within each group to assess the performance of both models under varying cloud conditions.</p>
      <p id="d2e5213">On the Mediterranean test set, the cloud coverage ranged from a minimum of 8.7 % to a maximum of 99 %. CRITER outperformed DINCAE2 across all cloud coverage groups, achieving significant reductions in reconstruction error over deleted regions. Specifically, the error was reduced by 21 % in the low-coverage group, 18 % in the moderate-coverage group, and 16 % in the high-coverage group. Similarly, on the Adriatic test set, the cloud coverage ranged from a minimum of 3.4 % to a maximum of 93 %. Here, CRITER substantially reduced the reconstruction error over deleted regions by 38 % in the low-coverage group, 49 % in the moderate-coverage group, and 54 % in the high-coverage group. Finally, on the Atlantic test set, the cloud coverage ranged from a minimum of 37 % to a maximum of 97 %. CRITER achieved a 4 % decrease in the low coverage group and around a 1.3 % decrease in moderate- and high-coverage groups.</p>

      <fig id="F9" specific-use="star"><label>Figure 9</label><caption><p id="d2e5218">Reconstruction error comparison between CRITER and DINCAE2 across different cloud coverage groups (low, moderate, and high) on the Mediterranean, Adriatic, and Atlantic test sets. The three rows correspond to the RMSE computed over (1) all ground truth measurements, (2) missing measurements, and (3) observed measurements. The error bars indicate the 10 % percentile, mean, and 90 % percentile of the error, respectively.</p></caption>
            <graphic xlink:href="https://gmd.copernicus.org/articles/18/5549/2025/gmd-18-5549-2025-f09.png"/>

          </fig>

</sec>
</sec>
<sec id="Ch1.S4.SS4">
  <label>4.4</label><title>Uncertainty estimation and bias analysis</title>
      <p id="d2e5236">CRITER and DINCAE2 estimate both the reconstruction of missing values and the associated uncertainty (i.e., the standard deviation) for each pixel. To assess the reliability of the estimated standard deviation, we employ the scaled error metric

            <disp-formula id="Ch1.E8" content-type="numbered"><label>8</label><mml:math id="M226" display="block"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">ϵ</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal" stretchy="false">̃</mml:mo></mml:mover><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

          as proposed by <xref ref-type="bibr" rid="bib1.bibx2" id="text.57"/>. This metric quantifies the difference between the ground truth observation <inline-formula><mml:math id="M227" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> and the reconstruction <inline-formula><mml:math id="M228" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo stretchy="false" mathvariant="normal">̃</mml:mo></mml:mover><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>, normalized by the estimated standard deviation <inline-formula><mml:math id="M229" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">σ</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>, where <inline-formula><mml:math id="M230" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula> is the pixel index. We calculate the mean, <inline-formula><mml:math id="M231" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi mathvariant="italic">ϵ</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, and standard deviation, <inline-formula><mml:math id="M232" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi mathvariant="italic">ϵ</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, of the scaled error over the entire test set. Furthermore, we compute the bias, defined as the (non-normalized) mean difference between the ground truth observations and reconstructions. An ideal reconstruction method would thus have the bias equal to zero (i.e., predicted values are not globally under or over estimated). and standard deviation of the scaled error <inline-formula><mml:math id="M233" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi mathvariant="italic">ϵ</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> equal to 1 (i.e., per-pixel disparities match the predicted uncertainties).  Standard deviation of the scaled error <inline-formula><mml:math id="M234" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi mathvariant="italic">ϵ</mml:mi></mml:msub><mml:mo>&lt;</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> indicates that the predicted standard deviation <inline-formula><mml:math id="M235" display="inline"><mml:mi mathvariant="bold-italic">σ</mml:mi></mml:math></inline-formula> is overestimated, while <inline-formula><mml:math id="M236" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi mathvariant="italic">ϵ</mml:mi></mml:msub><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> indicates that <inline-formula><mml:math id="M237" display="inline"><mml:mi mathvariant="bold-italic">σ</mml:mi></mml:math></inline-formula> is underestimated.</p>
      <p id="d2e5438">Figure <xref ref-type="fig" rid="F10"/> displays the histogram of the scaled error metric <inline-formula><mml:math id="M238" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">ϵ</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> for each test set, along with the corresponding Gaussian distribution, characterized by the estimated mean <inline-formula><mml:math id="M239" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi mathvariant="italic">ϵ</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and standard deviation <inline-formula><mml:math id="M240" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi mathvariant="italic">ϵ</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. The mean (<inline-formula><mml:math id="M241" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi mathvariant="italic">ϵ</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>), standard deviation (<inline-formula><mml:math id="M242" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi mathvariant="italic">ϵ</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>), and the bias for each dataset are provided in Table <xref ref-type="table" rid="T2"/>. Notably, CRITER moderately underestimates the standard deviation, with standard deviation of the scaled error <inline-formula><mml:math id="M243" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi mathvariant="italic">ϵ</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> values of 1.116, 1.082, and 1.156 on the Mediterranean, Adriatic, and Atlantic datasets, respectively, ranging from 8 % to 16 %. In contrast, on average, DINCAE2 significantly overestimates the standard deviation, with <inline-formula><mml:math id="M244" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi mathvariant="italic">ϵ</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> values of <inline-formula><mml:math id="M245" display="inline"><mml:mrow><mml:mn mathvariant="normal">0.334</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">0.996</mml:mn><mml:mo>,</mml:mo></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M246" display="inline"><mml:mn mathvariant="normal">0.801</mml:mn></mml:math></inline-formula> across the three datasets. The over-estimation thus ranges from as little as <inline-formula><mml:math id="M247" display="inline"><mml:mrow><mml:mn mathvariant="normal">0.4</mml:mn><mml:mi mathvariant="italic">%</mml:mi></mml:mrow></mml:math></inline-formula> to substantial over-estimates of 66 %. CRITER consistently exhibits a very low bias (of the order of <inline-formula><mml:math id="M248" display="inline"><mml:mn mathvariant="normal">0.01</mml:mn></mml:math></inline-formula> °C or lower) over all datasets. Furthermore, CRITER exhibits a significantly smaller bias on the Mediterranean and Adriatic datasets than DINCAE2, whereas DINCAE2 achieves a smaller bias on the Atlantic dataset. Note that, on the Adriatic dataset, DINCAE2 exhibits 18<inline-formula><mml:math id="M249" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> larger bias than CRITER.</p>

<table-wrap id="T2"><label>Table 2</label><caption><p id="d2e5577">Comparison of CRITER and DINCAE2 on each test set, showing the mean of the scaled error (<inline-formula><mml:math id="M250" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi mathvariant="italic">ϵ</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>), standard deviation of the scaled error (<inline-formula><mml:math id="M251" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi mathvariant="italic">ϵ</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) – both unitless and bias in °C. Bold font indicates the best result for each metric and dataset.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="5">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:colspec colnum="5" colname="col5" align="right"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Dataset</oasis:entry>
         <oasis:entry colname="col2">Model</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M252" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi mathvariant="italic">ϵ</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> (/)</oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M253" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi mathvariant="italic">ϵ</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> (/)</oasis:entry>
         <oasis:entry colname="col5">bias (°C)</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">Mediterranean</oasis:entry>
         <oasis:entry colname="col2">DINCAE2</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M254" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.060</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">0.334</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M255" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.060</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">CRITER (ours)</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M256" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.022</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4"><bold>1.116</bold></oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M257" display="inline"><mml:mrow><mml:mo mathvariant="bold">-</mml:mo><mml:mn mathvariant="bold">0.007</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Adriatic</oasis:entry>
         <oasis:entry colname="col2">DINCAE2</oasis:entry>
         <oasis:entry colname="col3">0.198</oasis:entry>
         <oasis:entry colname="col4"><bold>0.996</bold></oasis:entry>
         <oasis:entry colname="col5">0.128</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">CRITER (ours)</oasis:entry>
         <oasis:entry colname="col3">0.041</oasis:entry>
         <oasis:entry colname="col4">1.082</oasis:entry>
         <oasis:entry colname="col5"><bold>0.007</bold></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Atlantic</oasis:entry>
         <oasis:entry colname="col2">DINCAE2</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M258" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.017</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">0.801</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M259" display="inline"><mml:mrow><mml:mo mathvariant="bold">-</mml:mo><mml:mn mathvariant="bold">0.006</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">CRITER (ours)</oasis:entry>
         <oasis:entry colname="col3">0.118</oasis:entry>
         <oasis:entry colname="col4"><bold>1.156</bold></oasis:entry>
         <oasis:entry colname="col5">0.047</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <fig id="F10" specific-use="star"><label>Figure 10</label><caption><p id="d2e5823">Histograms of the scaled error <inline-formula><mml:math id="M260" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">ϵ</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> for the Mediterranean, Adriatic, and Atlantic datasets, overlaid with the corresponding Gaussian distributions, which are characterized by the estimated mean (<inline-formula><mml:math id="M261" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi mathvariant="italic">ϵ</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) and standard deviation (<inline-formula><mml:math id="M262" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi mathvariant="italic">ϵ</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>). Additionally, an ideal model is shown in black.</p></caption>
          <graphic xlink:href="https://gmd.copernicus.org/articles/18/5549/2025/gmd-18-5549-2025-f10.png"/>

        </fig>

</sec>
<sec id="Ch1.S4.SS5">
  <label>4.5</label><title>Ablation study</title>
      <p id="d2e5879">We analyze the proposed CRITER architecture by ablating or replacing individual parts. All model variants are trained for a total of <inline-formula><mml:math id="M263" display="inline"><mml:mn mathvariant="normal">500</mml:mn></mml:math></inline-formula> epochs (CRM and IRM are trained for <inline-formula><mml:math id="M264" display="inline"><mml:mn mathvariant="normal">200</mml:mn></mml:math></inline-formula> and <inline-formula><mml:math id="M265" display="inline"><mml:mn mathvariant="normal">300</mml:mn></mml:math></inline-formula> epochs, respectively) using the hyper-parameters described in the “Implementation details” section (Sect. <xref ref-type="sec" rid="Ch1.S4.SS1"/>). The variants are evaluated on the Mediterranean dataset.</p>
<sec id="Ch1.S4.SS5.SSS1">
  <label>4.5.1</label><title>Importance of the coarse reconstruction stage</title>
      <p id="d2e5912">We evaluate CRITER<sub>CRM‾</sub>, a variant without CRM. In this configuration, IRM is initialized with an uninformative prior <inline-formula><mml:math id="M267" display="inline"><mml:mrow><mml:msup><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo stretchy="false" mathvariant="normal">̃</mml:mo></mml:mover><mml:mrow><mml:mo>(</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mn mathvariant="bold">0</mml:mn></mml:mrow></mml:math></inline-formula> and operates without FFM. Consequently, the first Residual Estimation Network (REN) assumes responsibility for the low-frequency reconstruction, a task previously performed by the transformer-based CRM. Table <xref ref-type="table" rid="T3"/> demonstrates that incorporating CRM reduces the error over deleted and visible regions by 24 % and 87 %, respectively, validating the use of a transformer-based model for estimating the low frequency components.</p>

<table-wrap id="T3"><label>Table 3</label><caption><p id="d2e5955">Performance of CRITER and CRITER<sub>CRM‾</sub>, which does not utilize CRM. Bold font indicates the best result for each metric.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="4">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Variant</oasis:entry>
         <oasis:entry colname="col2">RMSE<sub>all</sub> (°C)</oasis:entry>
         <oasis:entry colname="col3">RMSE<sub>mis</sub> (°C)</oasis:entry>
         <oasis:entry colname="col4">RMSE<sub>vis</sub>  (°C)</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">CRITER<sub>CRM‾</sub></oasis:entry>
         <oasis:entry colname="col2">0.205</oasis:entry>
         <oasis:entry colname="col3">0.336</oasis:entry>
         <oasis:entry colname="col4">0.129</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">CRITER</oasis:entry>
         <oasis:entry colname="col2"><bold>0.127</bold></oasis:entry>
         <oasis:entry colname="col3"><bold>0.255</bold></oasis:entry>
         <oasis:entry colname="col4"><bold>0.017</bold></oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

</sec>
<sec id="Ch1.S4.SS5.SSS2">
  <label>4.5.2</label><title>Architectural design of CRM</title>
</sec>
<sec id="Ch1.S4.SS5.SSSx1" specific-use="unnumbered">
  <title>Vision Transformer-based backbone</title>
      <p id="d2e6090">CRM (Sect. <xref ref-type="sec" rid="Ch1.S3.SS1"/>) utilizes a Vision Transformer-based architecture to compute a coarse reconstruction. The main argument for the transformer-based design is to allow direct information flow from all observed measurements into all corresponding tokens and the final coarse reconstruction. To evaluate the transformer design choice, we replace it by a convolutional counterpart, which maintains the same spatial reduction as the original CRM.</p>
      <p id="d2e6095">The convolutional variant, denoted with CRITER<sub>CNN</sub>, utilizes a CNN-based CRM which accepts the same input as the original CRM and reduces the spatial resolution <inline-formula><mml:math id="M274" display="inline"><mml:mrow><mml:mn mathvariant="normal">8</mml:mn><mml:mo>×</mml:mo></mml:mrow></mml:math></inline-formula> by three <italic>double conv</italic> blocks, each followed by a <inline-formula><mml:math id="M275" display="inline"><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula> max pooling operation. The number of convolutional kernels in each block is <inline-formula><mml:math id="M276" display="inline"><mml:mrow><mml:mn mathvariant="normal">64</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">128</mml:mn></mml:mrow></mml:math></inline-formula>, and <inline-formula><mml:math id="M277" display="inline"><mml:mn mathvariant="normal">256</mml:mn></mml:math></inline-formula>, respectively. This is followed by a bottleneck layer, consisting of a <italic>double conv</italic> block with <inline-formula><mml:math id="M278" display="inline"><mml:mn mathvariant="normal">512</mml:mn></mml:math></inline-formula> convolutional kernels. The output latent features are upsampled to the original spatial resolution by applying <inline-formula><mml:math id="M279" display="inline"><mml:mn mathvariant="normal">3</mml:mn></mml:math></inline-formula> transpose convolution layers, each followed by a <italic>double conv</italic> block. The number of convolutional kernels in each block is <inline-formula><mml:math id="M280" display="inline"><mml:mrow><mml:mn mathvariant="normal">256</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">128</mml:mn></mml:mrow></mml:math></inline-formula>, and <inline-formula><mml:math id="M281" display="inline"><mml:mn mathvariant="normal">64</mml:mn></mml:math></inline-formula>, respectively. Finally, a single <inline-formula><mml:math id="M282" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> convolutional layer computes the coarse reconstruction, which is passed to the IRM along with the latent features.</p>
      <p id="d2e6204">Results in Table <xref ref-type="table" rid="T4"/> show that using CRITER<sub>CNN</sub> leads to a substantial increase in reconstruction error over deleted and visible regions. This verifies the importance of the transformer-based design of CRM and suggests that global information flow plays an important role in obtaining good latent features and the coarse reconstruction.</p>

<table-wrap id="T4"><label>Table 4</label><caption><p id="d2e6221">Performance of CRITER variants with different backbones. CRITER<sub>CNN</sub> utilizes a CNN-based CRM, while CRITER utilizes the proposed ViT-based CRM. Bold font indicates the best result for each metric.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="4">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Variant</oasis:entry>
         <oasis:entry colname="col2">RMSE<sub>all</sub> (°C)</oasis:entry>
         <oasis:entry colname="col3">RMSE<sub>mis</sub> (°C)</oasis:entry>
         <oasis:entry colname="col4">RMSE<sub>vis</sub>  (°C)</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">CRITER<sub>CNN</sub></oasis:entry>
         <oasis:entry colname="col2">0.203</oasis:entry>
         <oasis:entry colname="col3">0.345</oasis:entry>
         <oasis:entry colname="col4">0.115</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">CRITER</oasis:entry>
         <oasis:entry colname="col2"><bold>0.127</bold></oasis:entry>
         <oasis:entry colname="col3"><bold>0.255</bold></oasis:entry>
         <oasis:entry colname="col4"><bold>0.017</bold></oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

</sec>
<sec id="Ch1.S4.SS5.SSSx2" specific-use="unnumbered">
  <title>Modeling spatio-temporal data dependencies</title>
      <p id="d2e6343">We next inspect the importance of using spatio-temporal information in the coarse reconstruction. For this reason we remove the IRM from CRITER, leading to only using our proposed spatio-temporal masked-auto-encoder-based CRM architecture for reconstruction. We compare the reconstruction capabilities of CRM with the recent MAESSTRO <xref ref-type="bibr" rid="bib1.bibx19" id="paren.58"/>, which also employs a Vision Transformer (ViT)  <xref ref-type="bibr" rid="bib1.bibx11" id="paren.59"/> and is based on a masked autoencoder <xref ref-type="bibr" rid="bib1.bibx21" id="paren.60"/>. In fact, the major difference is that CRM utilizes three temporally consecutive SST fields to reconstruct the central field, while MAESSTRO uses only the central field.</p>
      <p id="d2e6355">Table <xref ref-type="table" rid="T5"/> demonstrates that CRITER<sub>IRM‾</sub> reduces reconstruction error by 44 % and 56 % over deleted and visible regions, respectively, compared to MAESSTRO. These results confirm the CRM modeling capability of spatio-temporal data dependencies, which considerably improves reconstruction performance.</p>

<table-wrap id="T5"><label>Table 5</label><caption><p id="d2e6375">Performance of MAESSTRO and CRM. Bold font indicates the best result for each metric.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="4">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Model</oasis:entry>
         <oasis:entry colname="col2">RMSE<sub>all</sub> (°C)</oasis:entry>
         <oasis:entry colname="col3">RMSE<sub>mis</sub> (°C)</oasis:entry>
         <oasis:entry colname="col4">RMSE<sub>vis</sub>  (°C)</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">MAESSTRO</oasis:entry>
         <oasis:entry colname="col2">0.487</oasis:entry>
         <oasis:entry colname="col3">0.607</oasis:entry>
         <oasis:entry colname="col4">0.434</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">CRITER<sub>IRM‾</sub></oasis:entry>
         <oasis:entry colname="col2"><bold>0.242</bold></oasis:entry>
         <oasis:entry colname="col3"><bold>0.337</bold></oasis:entry>
         <oasis:entry colname="col4"><bold>0.190</bold></oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

</sec>
<sec id="Ch1.S4.SS5.SSS3">
  <label>4.5.3</label><title>Importance of the refinement stage</title>
      <p id="d2e6492">To investigate the importance of refinement, we compare CRITER with two variants. The first variant, CRITER<sub>IRM‾</sub>, does not utilize refinement and takes the output of CRM as the final reconstruction. The second variant CRITER<sub>res‾</sub> modifies IRM to estimate the full reconstruction at each iteration (in contrast to the proposed IRM that estimates a sequence of residuals).</p>
      <p id="d2e6519">Table <xref ref-type="table" rid="T6"/> shows that utilizing refinement consistently leads to improved reconstruction. In particular the proposed IRM reduces CRM reconstruction error by 24 % and 91 % over deleted and visible regions, respectively. Furthermore, the results confirm that our proposed approach of consecutive residual estimation leads to lower errors than when the full signal is reconstructed at each refinement step. We hypothesize two reasons for this result. First, the residual estimation approach better exploits the individual REN networks in IRM, allowing each network to dedicate the full capacity for correction of the errors from the previous REN, thus gradually focusing on the high-frequency content reconstruction. Secondly, since the final reconstruction is obtained by summing the residuals, this enables a better gradient flow directly to each REN, thus enabling better training.</p>

<table-wrap id="T6"><label>Table 6</label><caption><p id="d2e6527">Performance of CRITER variants using different refinement approaches. CRITER<sub>IRM‾</sub> does not utilize refinement, and CRITER<sub>res‾</sub> modifies IRM to estimate the full reconstruction at each iteration, while CRITER utilizes the proposed IRM. Bold font indicates the best result for each metric.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="4">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Variant</oasis:entry>
         <oasis:entry colname="col2">RMSE<sub>all</sub> (°C)</oasis:entry>
         <oasis:entry colname="col3">RMSE<sub>mis</sub> (°C)</oasis:entry>
         <oasis:entry colname="col4">RMSE<sub>vis</sub>  (°C)</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">CRITER<sub>IRM‾</sub></oasis:entry>
         <oasis:entry colname="col2">0.242</oasis:entry>
         <oasis:entry colname="col3">0.337</oasis:entry>
         <oasis:entry colname="col4">0.190</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">CRITER<sub>res‾</sub></oasis:entry>
         <oasis:entry colname="col2">0.156</oasis:entry>
         <oasis:entry colname="col3">0.286</oasis:entry>
         <oasis:entry colname="col4">0.062</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">CRITER</oasis:entry>
         <oasis:entry colname="col2"><bold>0.127</bold></oasis:entry>
         <oasis:entry colname="col3"><bold>0.255</bold></oasis:entry>
         <oasis:entry colname="col4"><bold>0.017</bold></oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

</sec>
<sec id="Ch1.S4.SS5.SSS4">
  <label>4.5.4</label><title>Influence of refinement iteration steps</title>
      <p id="d2e6694">We next investigate the impact of varying the number of refinement steps in IRM (Sect. <xref ref-type="sec" rid="Ch1.S3.SS2"/>) on the reconstruction quality. Figure <xref ref-type="fig" rid="F11"/> shows results of CRITER retrained with different number of steps in IRM. The lowest reconstruction error is reached at three refinement steps. In particular the RMSE<sub>all</sub> is reduced by 8 % compared to using a single refinement step. Using more refinement steps does not improve performance but leads to increased error. This is likely due to the parameter increase, since each refinement step introduces a new REN network, which makes training less efficient on the limited dataset size. We defer explorations of more resilient IRM architectures to future work.</p>

      <fig id="F11"><label>Figure 11</label><caption><p id="d2e6712">Performance of CRITER variants with increasing number of refinement iterations.</p></caption>
            <graphic xlink:href="https://gmd.copernicus.org/articles/18/5549/2025/gmd-18-5549-2025-f11.png"/>

          </fig>


</sec>
<sec id="Ch1.S4.SS5.SSS5">
  <label>4.5.5</label><title>Importance of the CRM latent features</title>
      <p id="d2e6731">In IRM (Sect. <xref ref-type="sec" rid="Ch1.S3.SS2"/>), the latent features computed by CRM are fused with the bottleneck features to improve injection of global coarse information in the refinement steps. To evaluate the importance of this, we retrained CRITER without the coarse latent features fusion in IRMs – this variant is denoted as CRITER<sub>fus‾</sub>.</p>
      <p id="d2e6748">Results in Table <xref ref-type="table" rid="T7"/> show that the reconstruction error over deleted and visible regions of CRITER with feature fusion reduces by 1.9 % and 19 %, respectively, compared to the feature fusion free counterpart.</p>

<table-wrap id="T7"><label>Table 7</label><caption><p id="d2e6756">Comparison of CRITER, which utilizes latent features computed by CRM, with CRITER<sub>fus‾</sub>, which does not. Bold font indicates the best result for each metric.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="4">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Variant</oasis:entry>
         <oasis:entry colname="col2">RMSE<sub>all</sub> (°C)</oasis:entry>
         <oasis:entry colname="col3">RMSE<sub>mis</sub> (°C)</oasis:entry>
         <oasis:entry colname="col4">RMSE<sub>vis</sub>  (°C)</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">CRITER<sub>fus‾</sub></oasis:entry>
         <oasis:entry colname="col2">0.130</oasis:entry>
         <oasis:entry colname="col3">0.260</oasis:entry>
         <oasis:entry colname="col4">0.021</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">CRITER</oasis:entry>
         <oasis:entry colname="col2"><bold>0.127</bold></oasis:entry>
         <oasis:entry colname="col3"><bold>0.255</bold></oasis:entry>
         <oasis:entry colname="col4"><bold>0.017</bold></oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

</sec>
<sec id="Ch1.S4.SS5.SSS6">
  <label>4.5.6</label><title>Importance of time auxiliary features</title>
      <p id="d2e6885">CRM (Sect. <xref ref-type="sec" rid="Ch1.S3.SS1"/>) takes as input a sequence of consecutive observation fields, that are concatenated with auxiliary features, particularly the cosine and sine of the day of the year that encode the yearly cycle of SST. The auxiliary features offer additional information which CRM can incorporate when computing the latent features and generating the coarse reconstruction. To evaluate the importance of this, we train a CRM variant which does not leverage auxiliary features, denoted by CRITER<sub>aux‾</sub>. Results in Table <xref ref-type="table" rid="T8"/> show that augmenting the input with auxiliary features leads to a 1.5 % and 26 % decrease in reconstruction error over deleted and visible regions, respectively.</p>

<table-wrap id="T8"><label>Table 8</label><caption><p id="d2e6907">Comparison of CRITER, which utilizes auxiliary features (cosine and sine of the day of the year), with CRITER<sub>aux‾</sub>, which does not. Bold font indicates the best result for each metric.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="4">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Variant</oasis:entry>
         <oasis:entry colname="col2">RMSE<sub>all</sub> (°C)</oasis:entry>
         <oasis:entry colname="col3">RMSE<sub>mis</sub> (°C)</oasis:entry>
         <oasis:entry colname="col4">RMSE<sub>vis</sub>  (°C)</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">CRITER<sub>aux‾</sub></oasis:entry>
         <oasis:entry colname="col2">0.130</oasis:entry>
         <oasis:entry colname="col3">0.259</oasis:entry>
         <oasis:entry colname="col4">0.023</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">CRITER</oasis:entry>
         <oasis:entry colname="col2"><bold>0.127</bold></oasis:entry>
         <oasis:entry colname="col3"><bold>0.255</bold></oasis:entry>
         <oasis:entry colname="col4"><bold>0.017</bold></oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

</sec>
</sec>
</sec>
<sec id="Ch1.S5" sec-type="conclusions">
  <label>5</label><title>Conclusions</title>
      <p id="d2e7038">This study introduced CRITER, a novel two-stage model for reconstructing sea surface temperature (SST) from sparse satellite observations. High performance of the CRITER method stems from a Coarse Reconstruction Module (CRM) utilizing a Vision Transformer (ViT) architecture for initial reconstruction, followed by an Iterative Refinement Module (IRM) to refine the reconstruction with a focus on high-frequency information. The global receptive field of the ViT enables modeling of long-range dependencies in the data, while iterative refinement allows each network to focus its full capacity on modeling high-frequency corrections. This combination leads to significant enhancements in overall performance. The introduction of CRM's ViT global attention mechanism proved crucial for effective long-range dependency modeling, addressing limitations of convolutional architectures.</p>
      <p id="d2e7041">Our results show that CRITER surpasses the state-of-the-art DINCAE2 model by a significant margin across three diverse SST datasets: Mediterranean, Adriatic, and Atlantic. Notably, CRITER achieves substantial reductions in reconstruction error, with improvements of up to 89 % in observed regions and up to 44 % in missing regions.</p>
      <p id="d2e7044">The iterative refinement process of IRM, focusing on residual estimation, further enhanced reconstruction accuracy by efficiently utilizing model capacity for high-frequency variability in the SST observations. Ablation studies confirmed the importance of CRM's transformer-based design, the effectiveness of iterative residual estimation in IRM, and the utility of incorporating auxiliary features such as the day-of-year encoding.</p>
      <p id="d2e7047">Overall, CRITER sets a new benchmark for SST reconstruction, providing a robust framework that leverages the strengths of both transformer and convolutional architectures to deliver superior performance. Future work will explore extending CRITER's applicability by incorporating additional environmental proxy  variables (like chlorophyll <inline-formula><mml:math id="M316" display="inline"><mml:mi>a</mml:mi></mml:math></inline-formula>, which often serves as a complementary variable to SST in ocean state estimates) and increasing the temporal horizon for even more accurate sparse data reconstructions.</p>
</sec>

      
      </body>
    <back><app-group>

<app id="App1.Ch1.S1">
  <label>Appendix A</label><title>Analysis of missing values in evaluation datasets</title>
      <p id="d2e7069">We analyze the extent of missing values in each dataset described in Sect. <xref ref-type="sec" rid="Ch1.S2.SS1"/>. To quantify the number of missing data, we define the cloud coverage <inline-formula><mml:math id="M317" display="inline"><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> of an observation <inline-formula><mml:math id="M318" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> as

          <disp-formula id="App1.Ch1.S1.E9" content-type="numbered"><label>A1</label><mml:math id="M319" display="block"><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mo>|</mml:mo><mml:mn mathvariant="bold">1</mml:mn><mml:mo>-</mml:mo><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:msub><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">l</mml:mi></mml:msub><mml:mo>|</mml:mo></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

        where <inline-formula><mml:math id="M320" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:msub><mml:mo>∈</mml:mo><mml:mo mathvariant="italic">{</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:msup><mml:mo mathvariant="italic">}</mml:mo><mml:mrow><mml:mi>W</mml:mi><mml:mo>×</mml:mo><mml:mi>H</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> is the missing data mask corresponding to observation <inline-formula><mml:math id="M321" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, and <inline-formula><mml:math id="M322" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">l</mml:mi></mml:msub><mml:mo>∈</mml:mo><mml:mo mathvariant="italic">{</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:msup><mml:mo mathvariant="italic">}</mml:mo><mml:mrow><mml:mi>W</mml:mi><mml:mo>×</mml:mo><mml:mi>H</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> is the land mask. Cloud coverage is computed as the fraction of pixels that are missing relative to the number of pixels belonging to sea areas. We then calculate the mean cloud coverage over <inline-formula><mml:math id="M323" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:math></inline-formula> consecutive observation fields as <inline-formula><mml:math id="M324" display="inline"><mml:mrow><mml:mi>A</mml:mi><mml:mo>=</mml:mo><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">3</mml:mn></mml:mfrac></mml:mstyle><mml:mo>(</mml:mo><mml:msub><mml:mi>A</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>A</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>A</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. Note that the proportion of available information in the entire observation triplet is thus given by <inline-formula><mml:math id="M325" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:mi>A</mml:mi></mml:mrow></mml:math></inline-formula>. Figure <xref ref-type="fig" rid="FA1"/> presents a histogram of the mean cloud coverage <inline-formula><mml:math id="M326" display="inline"><mml:mi>A</mml:mi></mml:math></inline-formula> for all three filtered datasets. The results show that the datasets can be ranked by the average amount of available information in each observation triplet, from highest to lowest: Mediterranean, Adriatic, and Atlantic.</p>

      <fig id="FA1"><label>Figure A1</label><caption><p id="d2e7299">Histogram (<inline-formula><mml:math id="M327" display="inline"><mml:mn mathvariant="normal">100</mml:mn></mml:math></inline-formula> bins) of cloud coverage for all three (filtered) datasets.</p></caption>
        <graphic xlink:href="https://gmd.copernicus.org/articles/18/5549/2025/gmd-18-5549-2025-f12.png"/>

      </fig>


</app>

<app id="App1.Ch1.S2">
  <label>Appendix B</label><title>Additional qualitative analysis figures</title>
      <p id="d2e7325">This section presents additional reconstructions generated by CRITER and DINCAE2 (<xref ref-type="bibr" rid="bib1.bibx3" id="altparen.61"/>). For a detailed discussion of the qualitative comparison, refer to Sect. <xref ref-type="sec" rid="Ch1.S4.SS3.SSS1"/>.</p>

      <fig id="FB1"><label>Figure B1</label><caption><p id="d2e7335">Same as Fig. <xref ref-type="fig" rid="F4"/> on different samples.</p></caption>
        
        <graphic xlink:href="https://gmd.copernicus.org/articles/18/5549/2025/gmd-18-5549-2025-f13.jpg"/>

      </fig>

<fig id="FB2"><label>Figure B2</label><caption><p id="d2e7352">Same as Fig. <xref ref-type="fig" rid="F4"/> but for the Adriatic domain.</p></caption>
        
        <graphic xlink:href="https://gmd.copernicus.org/articles/18/5549/2025/gmd-18-5549-2025-f14.jpg"/>

      </fig>

<fig id="FB3"><label>Figure B3</label><caption><p id="d2e7368">Same as Fig. <xref ref-type="fig" rid="F4"/> but for the Atlantic domain.</p></caption>
        
        <graphic xlink:href="https://gmd.copernicus.org/articles/18/5549/2025/gmd-18-5549-2025-f15.jpg"/>

      </fig>


</app>

<app id="App1.Ch1.S3">
  <label>Appendix C</label><title>Extended spatial spectral analysis</title>
      <p id="d2e7391">This section presents supplementary power spectral density (PSD) comparisons. Figure <xref ref-type="fig" rid="FC1"/> shows a challenging case with sparse measurements, where CRITER's PSD remains closer to the target (on average) for wavenumbers <inline-formula><mml:math id="M328" display="inline"><mml:mrow><mml:mi>k</mml:mi><mml:mo>≥</mml:mo><mml:mn mathvariant="normal">4</mml:mn><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mtext>cycles</mml:mtext><mml:mi>deg⁡</mml:mi></mml:mfrac></mml:mstyle></mml:mrow></mml:math></inline-formula>. Figure <xref ref-type="fig" rid="FC2"/> depicts a high-measurement scenario featuring a failure case for CRITER: minor noise amplification beyond <inline-formula><mml:math id="M329" display="inline"><mml:mrow><mml:mi>k</mml:mi><mml:mo>≥</mml:mo><mml:mn mathvariant="normal">5</mml:mn><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mtext>cycles</mml:mtext><mml:mi>deg⁡</mml:mi></mml:mfrac></mml:mstyle></mml:mrow></mml:math></inline-formula>. A similar issue occurs with DINCAE2 but in a different wavenumber band: Fig. <xref ref-type="fig" rid="FC3"/> shows significant noise amplification within <inline-formula><mml:math id="M330" display="inline"><mml:mrow><mml:mi>k</mml:mi><mml:mo>∈</mml:mo><mml:mo>[</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">4</mml:mn><mml:mo>]</mml:mo><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mtext>cycles</mml:mtext><mml:mi>deg⁡</mml:mi></mml:mfrac></mml:mstyle></mml:mrow></mml:math></inline-formula>. For a detailed discussion of the comparison, refer to Sect. <xref ref-type="sec" rid="Ch1.S4.SS3.SSS2"/>.</p>

      <fig id="FC1"><label>Figure C1</label><caption><p id="d2e7467">Same as Fig. <xref ref-type="fig" rid="F7"/> but for a different sample.</p></caption>
        
        <graphic xlink:href="https://gmd.copernicus.org/articles/18/5549/2025/gmd-18-5549-2025-f16.jpg"/>

      </fig>

<fig id="FC2"><label>Figure C2</label><caption><p id="d2e7484">Same as Fig. <xref ref-type="fig" rid="F7"/> but for a different sample.</p></caption>
        
        <graphic xlink:href="https://gmd.copernicus.org/articles/18/5549/2025/gmd-18-5549-2025-f17.jpg"/>

      </fig>

<fig id="FC3"><label>Figure C3</label><caption><p id="d2e7500">Same as Fig. <xref ref-type="fig" rid="F7"/> but for a different sample.</p></caption>
        
        <graphic xlink:href="https://gmd.copernicus.org/articles/18/5549/2025/gmd-18-5549-2025-f18.jpg"/>

      </fig>

</app>

<app id="App1.Ch1.S4">
  <label>Appendix D</label><title>Implementation details of baseline models</title>
      <p id="d2e7521">MAESSTRO <xref ref-type="bibr" rid="bib1.bibx19" id="paren.62"/> is trained using mean squared error loss, as described in Sect. <xref ref-type="sec" rid="Ch1.S3.SS3"/>, for consistency in our comparison. The model processes only the current time step SST field without auxiliary features. We modified MAESSTRO's original random patch masking to inject sampled real cloud masks, enhancing real-world applicability. An SST patch is masked if its corresponding cloud mask patch contains any zero values. MAESSTRO employs a ViT-Tiny backbone with <inline-formula><mml:math id="M331" display="inline"><mml:mn mathvariant="normal">12</mml:mn></mml:math></inline-formula> encoder and decoder layers, <inline-formula><mml:math id="M332" display="inline"><mml:mn mathvariant="normal">3</mml:mn></mml:math></inline-formula> multi-head attention (MHA) heads, a token dimension of <inline-formula><mml:math id="M333" display="inline"><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">192</mml:mn></mml:mrow></mml:math></inline-formula>, layer-norm epsilon of <inline-formula><mml:math id="M334" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">12</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>, and patch size of <inline-formula><mml:math id="M335" display="inline"><mml:mrow><mml:mn mathvariant="normal">8</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">8</mml:mn></mml:mrow></mml:math></inline-formula>. MAESSTRO is trained with a batch size of <inline-formula><mml:math id="M336" display="inline"><mml:mn mathvariant="normal">8</mml:mn></mml:math></inline-formula> using the AdamW optimizer with a learning rate <inline-formula><mml:math id="M337" display="inline"><mml:mrow><mml:mi mathvariant="italic">α</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">3</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">4</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M338" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">β</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.9</mml:mn></mml:mrow></mml:math></inline-formula>, and <inline-formula><mml:math id="M339" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">β</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.95</mml:mn></mml:mrow></mml:math></inline-formula> for <inline-formula><mml:math id="M340" display="inline"><mml:mn mathvariant="normal">100</mml:mn></mml:math></inline-formula> epochs (warm-up period) and then with a cosine decay scheduler <xref ref-type="bibr" rid="bib1.bibx25" id="paren.63"/> with a step size of <inline-formula><mml:math id="M341" display="inline"><mml:mn mathvariant="normal">50</mml:mn></mml:math></inline-formula> for another <inline-formula><mml:math id="M342" display="inline"><mml:mn mathvariant="normal">300</mml:mn></mml:math></inline-formula> epochs.</p>
      <p id="d2e7674">DINCAE2 <xref ref-type="bibr" rid="bib1.bibx3" id="paren.64"/> is trained using the negative log-likelihood loss, as described in Sect. <xref ref-type="sec" rid="Ch1.S3.SS3"/>, to maintain consistency in our comparison. The model utilizes a sequence of three temporally consecutive SST fields, along with day-of-the-year auxiliary features, to reconstruct the central SST field. Hyperparameters of the re-implemented DINCAE2 differ slightly between the datasets. On the Mediterranean and Atlantic, DINCAE2 is trained using the Adam optimizer, with an initial learning rate of <inline-formula><mml:math id="M343" display="inline"><mml:mrow><mml:mi mathvariant="italic">α</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">4</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M344" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">β</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.90</mml:mn></mml:mrow></mml:math></inline-formula>, and <inline-formula><mml:math id="M345" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">β</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.999</mml:mn></mml:mrow></mml:math></inline-formula> and a batch size of <inline-formula><mml:math id="M346" display="inline"><mml:mn mathvariant="normal">8</mml:mn></mml:math></inline-formula> for a total of <inline-formula><mml:math id="M347" display="inline"><mml:mn mathvariant="normal">1000</mml:mn></mml:math></inline-formula> epochs, using a step learning rate scheduler with a step size of <inline-formula><mml:math id="M348" display="inline"><mml:mn mathvariant="normal">100</mml:mn></mml:math></inline-formula> epochs and a multiplicative factor of <inline-formula><mml:math id="M349" display="inline"><mml:mrow><mml:mi mathvariant="italic">γ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.5</mml:mn></mml:mrow></mml:math></inline-formula>. On the Adriatic we use an initial learning rate of <inline-formula><mml:math id="M350" display="inline"><mml:mrow><mml:mi mathvariant="italic">α</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">7</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> and a step size of <inline-formula><mml:math id="M351" display="inline"><mml:mn mathvariant="normal">150</mml:mn></mml:math></inline-formula>; all other hyperparameters remain unchanged.</p>
</app>
  </app-group><notes notes-type="codedataavailability"><title>Code and data availability</title>

      <p id="d2e7803">Implementation of CRITER and the code to train and evaluate the model are available in the GitHub repository at <uri>https://github.com/Matjaz12/CRITER</uri>. We also include CRITER weights pretrained on the <italic>Mediterranean</italic>, <italic>Adriatic</italic>, and <italic>Atlantic</italic> datasets. The persistent version of our GitHub repository containing code under MIT license is available at <ext-link xlink:href="https://doi.org/10.5281/zenodo.13923156" ext-link-type="DOI">10.5281/zenodo.13923156</ext-link> <xref ref-type="bibr" rid="bib1.bibx41" id="paren.65"/>. We publish all three datasets at <ext-link xlink:href="https://doi.org/10.5281/zenodo.13923189" ext-link-type="DOI">10.5281/zenodo.13923189</ext-link> <xref ref-type="bibr" rid="bib1.bibx42" id="paren.66"/>.</p>
  </notes><notes notes-type="authorcontribution"><title>Author contributions</title>

      <p id="d2e7834">MZM was the main designer of CRITER; he implemented and evaluated the method. MK and VZ consulted on machine learning methodology, while ML, AB, and AAA consulted on the geophysical background and the datasets. MZM wrote the first draft of the paper, while MK, VZ, ML, AB, and AAA contributed to the final version of the paper.</p>
  </notes><notes notes-type="competinginterests"><title>Competing interests</title>

      <p id="d2e7840">The contact author has declared that none of the authors has any competing interests.</p>
  </notes><notes notes-type="disclaimer"><title>Disclaimer</title>

      <p id="d2e7846">Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. While Copernicus Publications makes every effort to include appropriate place names, the final responsibility lies with the authors.</p>
  </notes><ack><title>Acknowledgements</title><p id="d2e7852">The authors would like to thank the Academic and Research Network of Slovenia (ARNES) and the Slovenian National Supercomputing Network (SLING) consortium (ARNES, EuroHPC Vega – IZUM) for making the research possible by using their supercomputer clusters. This study has been conducted using E.U. Copernicus Marine Service Information; see <ext-link xlink:href="https://doi.org/10.48670/moi-00171" ext-link-type="DOI">10.48670/moi-00171</ext-link> <xref ref-type="bibr" rid="bib1.bibx12" id="paren.67"/> and <ext-link xlink:href="https://doi.org/10.48670/moi-00310" ext-link-type="DOI">10.48670/moi-00310</ext-link> <xref ref-type="bibr" rid="bib1.bibx13" id="paren.68"/>.</p></ack><notes notes-type="financialsupport"><title>Financial support</title>

      <p id="d2e7869">Matjaž Ličer acknowledges the financial support from the Slovenian Research and Innovation Agency ARIS (contract no. P1-0237). This research was supported in part by ARIS programme J2-2506 and project P2-0214.</p>
  </notes><notes notes-type="reviewstatement"><title>Review statement</title>

      <p id="d2e7876">This paper was edited by Nicola Bodini and reviewed by two anonymous referees.</p>
  </notes><ref-list>
    <title>References</title>

      <ref id="bib1.bibx1"><label>Alvera-Azcárate et al.(2005)</label><mixed-citation>Alvera-Azcárate, A., Barth, A., Rixen, M., and Beckers, J.: Reconstruction of incomplete oceanographic data sets using empirical orthogonal functions: application to the Adriatic Sea surface temperature, Ocean Model., 9, 325–346, <ext-link xlink:href="https://doi.org/10.1016/j.ocemod.2004.08.001" ext-link-type="DOI">10.1016/j.ocemod.2004.08.001</ext-link>,  2005.</mixed-citation></ref>
      <ref id="bib1.bibx2"><label>Barth et al.(2020)</label><mixed-citation>Barth, A., Alvera-Azcárate, A., Licer, M., and Beckers, J.-M.: DINCAE 1.0: a convolutional neural network with error estimates to reconstruct sea surface temperature satellite observations, Geosci. Model Dev., 13, 1609–1622, <ext-link xlink:href="https://doi.org/10.5194/gmd-13-1609-2020" ext-link-type="DOI">10.5194/gmd-13-1609-2020</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx3"><label>Barth et al.(2022)</label><mixed-citation>Barth, A., Alvera-Azcárate, A., Troupin, C., and Beckers, J.-M.: DINCAE 2.0: multivariate convolutional neural network with error estimates to reconstruct sea surface temperature satellite and altimetry observations, Geosci. Model Dev., 15, 2183–2196, <ext-link xlink:href="https://doi.org/10.5194/gmd-15-2183-2022" ext-link-type="DOI">10.5194/gmd-15-2183-2022</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx4"><label>Barth et al.(2024)</label><mixed-citation>Barth, A., Brajard, J., Alvera-Azcárate, A., Mohamed, B., Troupin, C., and Beckers, J.-M.: Ensemble reconstruction of missing satellite data using a denoising diffusion model: application to chlorophyll a concentration in the Black Sea, Ocean Sci., 20, 1567–1584, <ext-link xlink:href="https://doi.org/10.5194/os-20-1567-2024" ext-link-type="DOI">10.5194/os-20-1567-2024</ext-link>, 2024.</mixed-citation></ref>
      <ref id="bib1.bibx5"><label>Beauchamp et al.(2023)</label><mixed-citation>Beauchamp, M., Febvre, Q., Georgenthum, H., and Fablet, R.: 4DVarNet-SSH: end-to-end learning of variational interpolation schemes for nadir and wide-swath satellite altimetry, Geosci. Model Dev., 16, 2119–2147, <ext-link xlink:href="https://doi.org/10.5194/gmd-16-2119-2023" ext-link-type="DOI">10.5194/gmd-16-2119-2023</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx6"><label>Bishop et al.(2017)</label><mixed-citation>Bishop, S. P., Small, R. J., Bryan, F. O., and Tomas, R. A.: Scale Dependence of Midlatitude Air–Sea Interaction, J. Climate, 30, 8207–8221, <ext-link xlink:href="https://doi.org/10.1175/JCLI-D-17-0159.1" ext-link-type="DOI">10.1175/JCLI-D-17-0159.1</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx7"><label>Buongiorno Nardelli et al.(2022)</label><mixed-citation>Buongiorno Nardelli, B., Cavaliere, D., Charles, E., and Ciani, D.: Super-resolving ocean dynamics from space with computer vision algorithms, Remote Sens., 14, 1159, <ext-link xlink:href="https://doi.org/10.3390/rs14051159" ext-link-type="DOI">10.3390/rs14051159</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx8"><label>Casey et al.(2010)</label><mixed-citation>Casey, K., Brandon, T., Cornillon, P., and Evans, R.: The Past, Present and Future of the AVHRR Pathfinder SST Program, in: Oceanography from Space: Revisited, edited by: Barale, V., Gower, J., and Alberotanza, L., Springer, <ext-link xlink:href="https://doi.org/10.1007/978-90-481-8681-5_16" ext-link-type="DOI">10.1007/978-90-481-8681-5_16</ext-link>, 2010.</mixed-citation></ref>
      <ref id="bib1.bibx9"><label>Chelton(2005)</label><mixed-citation>Chelton, D. B.: The Impact of SST Specification on ECMWF Surface Wind Stress Fields in the Eastern Tropical Pacific, J. Climate, 18, 530–550, <ext-link xlink:href="https://doi.org/10.1175/JCLI-3275.1" ext-link-type="DOI">10.1175/JCLI-3275.1</ext-link>, 2005.</mixed-citation></ref>
      <ref id="bib1.bibx10"><label>Darmaraki et al.(2019)</label><mixed-citation> Darmaraki, S., Somot, S., Sevault, F., and Nabat, P.: Past variability of Mediterranean Sea marine heatwaves, Geophys. Res. Lett., 46, 9813–9823, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx11"><label>Dosovitskiy et al.(2021)</label><mixed-citation>Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N.: An Image is Worth <inline-formula><mml:math id="M352" display="inline"><mml:mrow><mml:mn mathvariant="normal">16</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">16</mml:mn></mml:mrow></mml:math></inline-formula> Words: Transformers for Image Recognition at Scale, in: International Conference on Learning Representations,  <uri>https://openreview.net/forum?id=YicbFdNTTy</uri> (last access: 26 January 2025), 2021.</mixed-citation></ref>
      <ref id="bib1.bibx12"><label>E.U. Copernicus Marine Service Information(2023a)</label><mixed-citation>E.U. Copernicus Marine Service Information (CMEMS):  Mediterranean Sea – High Resolution and Ultra High Resolution L3S Sea Surface Temperature, CEMS [data set], <ext-link xlink:href="https://doi.org/10.48670/moi-00171" ext-link-type="DOI">10.48670/moi-00171</ext-link>, 2023a.</mixed-citation></ref>
      <ref id="bib1.bibx13"><label>E.U. Copernicus Marine Service Information(2023b)</label><mixed-citation>E.U. Copernicus Marine Service Information (CMEMS): European North West Shelf/Iberia Biscay Irish Seas – High Resolution ODYSSEA Sea Surface Temperature Multi-sensor L3 Observations, CEMS [data set], <ext-link xlink:href="https://doi.org/10.48670/moi-00310" ext-link-type="DOI">10.48670/moi-00310</ext-link>, 2023b.</mixed-citation></ref>
      <ref id="bib1.bibx14"><label>Fablet et al.(2021)</label><mixed-citation>Fablet, R., Beauchamp, M., Drumetz, L., and Rousseau, F.: Joint interpolation and representation learning for irregularly sampled satellite-derived geophysical fields, Frontiers in Applied Mathematics and Statistics, 7, 655224, <ext-link xlink:href="https://doi.org/10.3389/fams.2021.655224" ext-link-type="DOI">10.3389/fams.2021.655224</ext-link>,  2021.</mixed-citation></ref>
      <ref id="bib1.bibx15"><label>Fanelli et al.(2024)</label><mixed-citation>Fanelli, C., Ciani, D., Pisano, A., and Buongiorno Nardelli, B.: Deep learning for the super resolution of Mediterranean sea surface temperature fields, Ocean Sci., 20, 1035–1050, <ext-link xlink:href="https://doi.org/10.5194/os-20-1035-2024" ext-link-type="DOI">10.5194/os-20-1035-2024</ext-link>, 2024.</mixed-citation></ref>
      <ref id="bib1.bibx16"><label>Feichtenhofer et al.(2022)</label><mixed-citation>Feichtenhofer, C., Fan, H., Li, Y., and He, K.: Masked autoencoders as spatiotemporal learners, Adv. Neur. In., 35, 35946–35958, <uri>https://proceedings.neurips.cc/paper_files/paper/2022/file/e97d1081481a4017df96b51be31001d3-Paper-Conference.pdf</uri> (last access: 26 January 2025), 2022.</mixed-citation></ref>
      <ref id="bib1.bibx17"><label>Garcia-Soto et al.(2021)</label><mixed-citation>Garcia-Soto, C., Cheng, L., Caesar, L., Schmidtko, S., Jewett, E. B., Cheripka, A., Rigor, I., Caballero, A., Chiba, S., Báez, J. C., Zielinski, T., and Abraham, J. P.: An Overview of Ocean Climate Change Indicators: Sea Surface Temperature, Ocean Heat Content, Ocean pH, Dissolved Oxygen Concentration, Arctic Sea Ice Extent, Thickness and Volume, Sea Level and Strength of the AMOC (Atlantic Meridional Overturning Circulation), Front. Mar. Sci., 8, 642372, <ext-link xlink:href="https://doi.org/10.3389/fmars.2021.642372" ext-link-type="DOI">10.3389/fmars.2021.642372</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx18"><label>Garrabou et al.(2022)</label><mixed-citation> Garrabou, J., Gómez-Gras, D., Medrano, A., Cerrano, C., Ponti, M., Schlegel, R., Bensoussan, N., Turicchia, E., Sini, M., Gerovasileiou, V., Teixido, N., Mirasole, A., Tamburello, L., Cebrian, E., Rilov, G., Ledoux, J.-B., Ben Souissi, J., Khamassi, F., Ghanem, R., Benabdi, M., Grimes, S., Ocaña, O., Bazairi, H., Hereu, B., Linares, C., Kersting, D. K., Rovira, G., Ortega, J., Casals, D., Pagès-Escolà, M., Margarit, N., Capdevila, P., Verdura, J., Ramos, A., Izquierdo, A., Barbera, C., Rubio-Portillo, E., Anton, I., López-Sendino, P., Díaz, D., Vázquez-Luis, M., Duarte, C., Marbà, N., Aspillaga, E., Espinosa, F., Grech, D., Guala, I., Azzurro, E., Farina, S., Gambi, M. C., Chimienti, G., Montefalcone, M., Azzola, A., Pulido Mantas, T., Fraschetti, S., Ceccherelli, G., Kipson, S., Bakran-Petricioli, T., Petricioli, D., Jimenez, C., Katsanevakis, S., Tuney Kizilkaya, I., Kizilkaya, Z., Sartoretto, S., Rouanet, E., Ruitton, S., Comeau, S., Gattuso, J.-P., and Harmelin, J.-G.: Marine heatwaves drive recurrent mass mortalities in the Mediterranean Sea, Glob. Change Biol., 28, 5708–5725, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx19"><label>Goh et al.(2024)</label><mixed-citation>Goh, E., Yepremyan, A., Wang, J., and Wilson, B.: MAESSTRO: Masked Autoencoders for Sea Surface Temperature Reconstruction under Occlusion, Ocean Sci., 20, 1309–1323, <ext-link xlink:href="https://doi.org/10.5194/os-20-1309-2024" ext-link-type="DOI">10.5194/os-20-1309-2024</ext-link>, 2024.</mixed-citation></ref>
      <ref id="bib1.bibx20"><label>Gómez-Gras et al.(2021)</label><mixed-citation>Gómez-Gras, D., Linares, C., López-Sanz, A., Amate, R., Ledoux, J. B., Bensoussan, N., Drap, P., Bianchimani, O., Marschal, C., Torrents, O., Zuberer, F., Cebrian, E., Teixidó, N., Zabala, M., Kipson, S., Kersting, D. K., Montero-Serra, I., Pagès-Escolà, M., Medrano, A., Frleta-Valić, M., Dimarchopoulou, D., López-Sendino, P., and Garrabou, J.: Population collapse of habitat-forming species in the Mediterranean: a long-term study of gorgonian populations affected by recurrent marine heatwaves, P. Roy. Soc. B, 288, 20212384, <ext-link xlink:href="https://doi.org/10.1098/rspb.2021.2384" ext-link-type="DOI">10.1098/rspb.2021.2384</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx21"><label>He et al.(2022)</label><mixed-citation>He, K., Chen, X., Xie, S., Li, Y., Dollár, P., and Girshick, R.: Masked autoencoders are scalable vision learners, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition,  16000–16009, <uri>https://openaccess.thecvf.com/content/CVPR2022/papers/He_Masked_Autoencoders_Are_Scalable_Vision_Learners_CVPR_2022_paper.pdf</uri> (last access: 26 January 2025), 2022.</mixed-citation></ref>
      <ref id="bib1.bibx22"><label>Hobday et al.(2016)</label><mixed-citation>Hobday, A. J., Alexander, L. V., Perkins, S. E., Smale, D. A., Straub, S. C., Oliver, E. C., Benthuysen, J. A., Burrows, M. T., Donat, M. G., Feng, M., Holbrook, N. J., Moore, P. J., Scannell, H. A., Sen Gupta, A., and Wernberg, T.: A hierarchical approach to defining marine heatwaves, Prog. Oceanogr., 141, 227–238, <ext-link xlink:href="https://doi.org/10.1016/j.pocean.2015.12.014" ext-link-type="DOI">10.1016/j.pocean.2015.12.014</ext-link>, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx23"><label>Ličer et al.(2016)</label><mixed-citation>Ličer, M., Smerkol, P., Fettich, A., Ravdas, M., Papapostolou, A., Mantziafou, A., Strajnar, B., Cedilnik, J., Jeromel, M., Jerman, J., Petan, S., Malačič, V., and Sofianos, S.: Modeling the ocean and atmosphere during an extreme bora event in northern Adriatic using one-way and two-way atmosphere–ocean coupling, Ocean Sci., 12, 71–86, <ext-link xlink:href="https://doi.org/10.5194/os-12-71-2016" ext-link-type="DOI">10.5194/os-12-71-2016</ext-link>, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx24"><label>Lloyd et al.(2021)</label><mixed-citation> Lloyd, D. T., Abela, A., Farrugia, R. A., Galea, A., and Valentino, G.: Optically enhanced super-resolution of sea surface temperature using deep learning, IEEE T. Geosci. Remote Sens., 60, 1–14, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx25"><label>Loshchilov and Hutter(2016)</label><mixed-citation>Loshchilov, I. and Hutter, F.: Sgdr: Stochastic gradient descent with warm restarts, arXiv [preprint], <ext-link xlink:href="https://doi.org/10.48550/arXiv.1608.03983" ext-link-type="DOI">10.48550/arXiv.1608.03983</ext-link>, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx26"><label>Martin et al.(2023)</label><mixed-citation>Martin, S. A., Manucharyan, G. E., and Klein, P.: Synthesizing sea surface temperature and satellite altimetry observations using deep learning improves the accuracy and resolution of gridded sea surface height anomalies, J. Adv. Model. Earth Sy., 15, e2022MS003589, <ext-link xlink:href="https://doi.org/10.1029/2022MS003589" ext-link-type="DOI">10.1029/2022MS003589</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx27"><label>Martin et al.(2024)</label><mixed-citation>Martin, S. A., Manucharyan, G. E., and Klein, P.: Deep learning improves global satellite observations of ocean eddy dynamics, Geophys. Res. Lett., 51, e2024GL110059, <ext-link xlink:href="https://doi.org/10.1029/2024GL110059" ext-link-type="DOI">10.1029/2024GL110059</ext-link>, 2024.</mixed-citation></ref>
      <ref id="bib1.bibx28"><label>Mogen et al.(2022)</label><mixed-citation>Mogen, S. C., Lovenduski, N. S., Dallmann, A. R., Gregor, L., Sutton, A. J., Bograd, S. J., Quiros, N. C., Di Lorenzo, E., Hazen, E. L., Jacox, M. G., Buil, M. P., and Yeager, S.: Ocean Biogeochemical Signatures of the North Pacific Blob, Geophys. Res. Lett., 49, e2021GL096938, <ext-link xlink:href="https://doi.org/10.1029/2021GL096938" ext-link-type="DOI">10.1029/2021GL096938</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx29"><label>O'Carroll et al.(2019)</label><mixed-citation>O'Carroll, A. G., Armstrong, E. M., Beggs, H. M., Bouali, M., Casey, K. S., Corlett, G. K., Dash, P., Donlon, C. J., Gentemann, C. L., Høyer, J. L., Ignatov, A., Kabobah, K., Kachi, M., Kurihara, Y., Karagali, I., Maturi, E., Merchant, C. J., Marullo, S., Minnett, P. J., Pennybacker, M., Ramakrishnan, B., Ramsankaran, R., Santoleri, R., Sunder, S., Saux Picart, S., Vázquez-Cuervo, J., and Wimmer, W.: Observational Needs of Sea Surface Temperature, Front. Mar. Sci., 6, 420, <ext-link xlink:href="https://doi.org/10.3389/fmars.2019.00420" ext-link-type="DOI">10.3389/fmars.2019.00420</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx30"><label>Pastor and Khodayar(2023)</label><mixed-citation>Pastor, F. and Khodayar, S.: Marine heat waves: Characterizing a major climate impact in the Mediterranean, Sci. Total Environ., 861, 160621, <ext-link xlink:href="https://doi.org/10.1016/j.scitotenv.2022.160621" ext-link-type="DOI">10.1016/j.scitotenv.2022.160621</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx31"><label>Paszke et al.(2017)</label><mixed-citation>Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., and Lerer, A.: Automatic differentiation in PyTorch, <uri>https://pytorch.org/</uri> (last access: 26 January 2025), 2017.</mixed-citation></ref>
      <ref id="bib1.bibx32"><label>Pisano et al.(2016)</label><mixed-citation>Pisano, A., Buongiorno Nardelli, B., Tronconi, C., and Santoleri, R.: The new Mediterranean optimally interpolated pathfinder AVHRR SST Dataset (1982–2012), Remote Sens. Environ., 176, 107–116, <ext-link xlink:href="https://doi.org/10.1016/j.rse.2016.01.019" ext-link-type="DOI">10.1016/j.rse.2016.01.019</ext-link>, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx33"><label>Ricchi et al.(2023)</label><mixed-citation>Ricchi, A., Sangelantoni, L., Redaelli, G., Mazzarella, V., Montopoli, M., Miglietta, M. M., Tiesi, A., Mazzà, S., Rotunno, R., and Ferretti, R.: Impact of the SST and topography on the development of a large-hail storm event, on the Adriatic Sea, Atmos. Res., 296, 107078, <ext-link xlink:href="https://doi.org/10.1016/j.atmosres.2023.107078" ext-link-type="DOI">10.1016/j.atmosres.2023.107078</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx34"><label>Ronneberger et al.(2015)</label><mixed-citation>Ronneberger, O., Fischer, P., and Brox, T.: U-net: Convolutional networks for biomedical image segmentation, in: Medical image computing and computer-assisted intervention–MICCAI 2015: 18th international conference, Munich, Germany, 5–9 October 2015, proceedings, part III 18, 234–241, Springer, <ext-link xlink:href="https://doi.org/10.1007/978-3-319-24574-4_28" ext-link-type="DOI">10.1007/978-3-319-24574-4_28</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx35"><label>Senatore et al.(2020)</label><mixed-citation>Senatore, A., Furnari, L., and Mendicino, G.: Impact of high-resolution sea surface temperature representation on the forecast of small Mediterranean catchments' hydrological responses to heavy precipitation, Hydrol. Earth Syst. Sci., 24, 269–291, <ext-link xlink:href="https://doi.org/10.5194/hess-24-269-2020" ext-link-type="DOI">10.5194/hess-24-269-2020</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx36"><label>Strajnar et al.(2019)</label><mixed-citation>Strajnar, B., Cedilnik, J., Fettich, A., Ličer, M., Pristov, N., Smerkol, P., and Jerman, J.: Impact of two-way coupling and sea-surface temperature on precipitation forecasts in regional atmosphere and ocean models, Q. J. Roy. Meteor. Soc., 145, 228–242, <ext-link xlink:href="https://doi.org/10.1002/qj.3425" ext-link-type="DOI">10.1002/qj.3425</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx37"><label>Taburet et al.(2019)</label><mixed-citation>Taburet, G., Sanchez-Roman, A., Ballarotta, M., Pujol, M.-I., Legeais, J.-F., Fournier, F., Faugere, Y., and Dibarboure, G.: DUACS DT2018: 25 years of reprocessed sea level altimetry products, Ocean Sci., 15, 1207–1224, <ext-link xlink:href="https://doi.org/10.5194/os-15-1207-2019" ext-link-type="DOI">10.5194/os-15-1207-2019</ext-link>, 2019. </mixed-citation></ref>
      <ref id="bib1.bibx38"><label>Ubelmann et al.(2021)</label><mixed-citation>Ubelmann, C., Dibarboure, G., Gaultier, L., Ponte, A., Ardhuin, F., Ballarotta, M., and Faugère, Y.: Reconstructing ocean surface current combining altimetry and future spaceborne Doppler data, J. Geophys. Res.-Oceans, 126, e2020JC016560, <ext-link xlink:href="https://doi.org/10.1029/2020JC016560" ext-link-type="DOI">10.1029/2020JC016560</ext-link>,  2021.</mixed-citation></ref>
      <ref id="bib1.bibx39"><label>Young et al.(2024)</label><mixed-citation>Young, C.-C., Cheng, Y.-C., Lee, M.-A., and Wu, J.-H.: Accurate reconstruction of satellite-derived SST under cloud and cloud-free areas using a physically-informed machine learning approach, Remote Sens. Environ., 313, 114339, <ext-link xlink:href="https://doi.org/10.1016/j.rse.2024.114339" ext-link-type="DOI">10.1016/j.rse.2024.114339</ext-link>, 2024.</mixed-citation></ref>
      <ref id="bib1.bibx40"><label>Yu et al.(2018)</label><mixed-citation>Yu, C., Wang, J., Peng, C., Gao, C., Yu, G., and Sang, N.: Bisenet: Bilateral segmentation network for real-time semantic segmentation, in: Proceedings of the European conference on computer vision (ECCV),   325–341, <uri>https://openaccess.thecvf.com/content_ECCV_2018/papers/Changqian_Yu_BiSeNet_Bilateral_Segmentation_ECCV_2018_paper.pdf</uri> (last access: 26 January 2025), 2018.</mixed-citation></ref>
      <ref id="bib1.bibx41"><label>Zupančič Muc(2025)</label><mixed-citation>Zupančič Muc, M.: CRITER – Coarse Reconstruction with ITerative Refinement network, Zenodo [code], <ext-link xlink:href="https://doi.org/10.5281/zenodo.15066015" ext-link-type="DOI">10.5281/zenodo.15066015</ext-link>, 2025.</mixed-citation></ref>
      <ref id="bib1.bibx42"><label>Zupančič Muc et al.(2024)</label><mixed-citation>Zupančič Muc, M., Zavrtanik, V., Barth, A., Alvera-Azcarate, A., Licer, M., and Kristan, M.: CRITER 1.0: Sea Surface Temperature Evaluation Datasets, Zenodo [data set], <ext-link xlink:href="https://doi.org/10.5281/zenodo.13923189" ext-link-type="DOI">10.5281/zenodo.13923189</ext-link>, 2024.</mixed-citation></ref>

  </ref-list></back>
    <!--<article-title-html>CRITER 1.0: a coarse reconstruction with iterative refinement network for sparse spatio-temporal satellite data</article-title-html>
<abstract-html/>
<ref-html id="bib1.bib1"><label>Alvera-Azcárate et al.(2005)</label><mixed-citation>
      
Alvera-Azcárate, A., Barth, A., Rixen, M., and Beckers, J.: Reconstruction of
incomplete oceanographic data sets using empirical orthogonal functions:
application to the Adriatic Sea surface temperature, Ocean Model., 9,
325–346, <a href="https://doi.org/10.1016/j.ocemod.2004.08.001" target="_blank">https://doi.org/10.1016/j.ocemod.2004.08.001</a>,  2005.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib2"><label>Barth et al.(2020)</label><mixed-citation>
      
Barth, A., Alvera-Azcárate, A., Licer, M., and Beckers, J.-M.: DINCAE 1.0: a convolutional neural network with error estimates to reconstruct sea surface temperature satellite observations, Geosci. Model Dev., 13, 1609–1622, <a href="https://doi.org/10.5194/gmd-13-1609-2020" target="_blank">https://doi.org/10.5194/gmd-13-1609-2020</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib3"><label>Barth et al.(2022)</label><mixed-citation>
      
Barth, A., Alvera-Azcárate, A., Troupin, C., and Beckers, J.-M.: DINCAE 2.0: multivariate convolutional neural network with error estimates to reconstruct sea surface temperature satellite and altimetry observations, Geosci. Model Dev., 15, 2183–2196, <a href="https://doi.org/10.5194/gmd-15-2183-2022" target="_blank">https://doi.org/10.5194/gmd-15-2183-2022</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib4"><label>Barth et al.(2024)</label><mixed-citation>
      
Barth, A., Brajard, J., Alvera-Azcárate, A., Mohamed, B., Troupin, C., and Beckers, J.-M.: Ensemble reconstruction of missing satellite data using a denoising diffusion model: application to chlorophyll a concentration in the Black Sea, Ocean Sci., 20, 1567–1584, <a href="https://doi.org/10.5194/os-20-1567-2024" target="_blank">https://doi.org/10.5194/os-20-1567-2024</a>, 2024.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib5"><label>Beauchamp et al.(2023)</label><mixed-citation>
      
Beauchamp, M., Febvre, Q., Georgenthum, H., and Fablet, R.: 4DVarNet-SSH: end-to-end learning of variational interpolation schemes for nadir and wide-swath satellite altimetry, Geosci. Model Dev., 16, 2119–2147, <a href="https://doi.org/10.5194/gmd-16-2119-2023" target="_blank">https://doi.org/10.5194/gmd-16-2119-2023</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib6"><label>Bishop et al.(2017)</label><mixed-citation>
      
Bishop, S. P., Small, R. J., Bryan, F. O., and Tomas, R. A.: Scale Dependence
of Midlatitude Air–Sea Interaction, J. Climate, 30, 8207–8221,
<a href="https://doi.org/10.1175/JCLI-D-17-0159.1" target="_blank">https://doi.org/10.1175/JCLI-D-17-0159.1</a>, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib7"><label>Buongiorno Nardelli et al.(2022)</label><mixed-citation>
      
Buongiorno Nardelli, B., Cavaliere, D., Charles, E., and Ciani, D.:
Super-resolving ocean dynamics from space with computer vision algorithms,
Remote Sens., 14, 1159, <a href="https://doi.org/10.3390/rs14051159" target="_blank">https://doi.org/10.3390/rs14051159</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib8"><label>Casey et al.(2010)</label><mixed-citation>
      
Casey, K., Brandon, T., Cornillon, P., and Evans, R.: The Past, Present and
Future of the AVHRR Pathfinder SST Program, in: Oceanography from Space:
Revisited, edited by: Barale, V., Gower, J., and Alberotanza, L., Springer,
<a href="https://doi.org/10.1007/978-90-481-8681-5_16" target="_blank">https://doi.org/10.1007/978-90-481-8681-5_16</a>, 2010.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib9"><label>Chelton(2005)</label><mixed-citation>
      
Chelton, D. B.: The Impact of SST Specification on ECMWF Surface Wind Stress
Fields in the Eastern Tropical Pacific, J. Climate, 18, 530–550,
<a href="https://doi.org/10.1175/JCLI-3275.1" target="_blank">https://doi.org/10.1175/JCLI-3275.1</a>, 2005.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib10"><label>Darmaraki et al.(2019)</label><mixed-citation>
      
Darmaraki, S., Somot, S., Sevault, F., and Nabat, P.: Past variability of
Mediterranean Sea marine heatwaves, Geophys. Res. Lett., 46,
9813–9823, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib11"><label>Dosovitskiy et al.(2021)</label><mixed-citation>
      
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X.,
Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S.,
Uszkoreit, J., and Houlsby, N.: An Image is Worth 16×16 Words: Transformers
for Image Recognition at Scale, in: International Conference on Learning
Representations,  <a href="https://openreview.net/forum?id=YicbFdNTTy" target="_blank"/> (last access: 26 January 2025), 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib12"><label>E.U. Copernicus Marine Service Information(2023a)</label><mixed-citation>
      
E.U. Copernicus Marine Service Information (CMEMS):  Mediterranean Sea – High Resolution and Ultra High Resolution L3S Sea Surface Temperature, CEMS [data set], <a href="https://doi.org/10.48670/moi-00171" target="_blank">https://doi.org/10.48670/moi-00171</a>, 2023a.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib13"><label>E.U. Copernicus Marine Service Information(2023b)</label><mixed-citation>
      
E.U. Copernicus Marine Service Information (CMEMS):
European North West Shelf/Iberia Biscay Irish Seas – High Resolution ODYSSEA Sea Surface Temperature Multi-sensor L3 Observations, CEMS [data set], <a href="https://doi.org/10.48670/moi-00310" target="_blank">https://doi.org/10.48670/moi-00310</a>, 2023b.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib14"><label>Fablet et al.(2021)</label><mixed-citation>
      
Fablet, R., Beauchamp, M., Drumetz, L., and Rousseau, F.: Joint interpolation
and representation learning for irregularly sampled satellite-derived
geophysical fields, Frontiers in Applied Mathematics and Statistics, 7,
655224, <a href="https://doi.org/10.3389/fams.2021.655224" target="_blank">https://doi.org/10.3389/fams.2021.655224</a>,  2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib15"><label>Fanelli et al.(2024)</label><mixed-citation>
      
Fanelli, C., Ciani, D., Pisano, A., and Buongiorno Nardelli, B.: Deep learning for the super resolution of Mediterranean sea surface temperature fields, Ocean Sci., 20, 1035–1050, <a href="https://doi.org/10.5194/os-20-1035-2024" target="_blank">https://doi.org/10.5194/os-20-1035-2024</a>, 2024.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib16"><label>Feichtenhofer et al.(2022)</label><mixed-citation>
      
Feichtenhofer, C., Fan, H., Li, Y., and He, K.: Masked autoencoders as
spatiotemporal learners, Adv. Neur. In.,
35, 35946–35958, <a href="https://proceedings.neurips.cc/paper_files/paper/2022/file/e97d1081481a4017df96b51be31001d3-Paper-Conference.pdf" target="_blank"/>
(last access: 26 January 2025), 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib17"><label>Garcia-Soto et al.(2021)</label><mixed-citation>
      
Garcia-Soto, C., Cheng, L., Caesar, L., Schmidtko, S., Jewett, E. B., Cheripka,
A., Rigor, I., Caballero, A., Chiba, S., Báez, J. C., Zielinski, T., and
Abraham, J. P.: An Overview of Ocean Climate Change Indicators: Sea Surface
Temperature, Ocean Heat Content, Ocean pH, Dissolved Oxygen Concentration,
Arctic Sea Ice Extent, Thickness and Volume, Sea Level and Strength of the
AMOC (Atlantic Meridional Overturning Circulation), Front. Mar.
Sci., 8, 642372, <a href="https://doi.org/10.3389/fmars.2021.642372" target="_blank">https://doi.org/10.3389/fmars.2021.642372</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib18"><label>Garrabou et al.(2022)</label><mixed-citation>
      
Garrabou, J., Gómez-Gras, D., Medrano, A., Cerrano, C., Ponti, M.,
Schlegel, R., Bensoussan, N., Turicchia, E., Sini, M., Gerovasileiou, V., Teixido, N., Mirasole, A., Tamburello, L., Cebrian, E., Rilov, G., Ledoux, J.-B., Ben Souissi, J., Khamassi, F., Ghanem, R., Benabdi, M., Grimes, S., Ocaña, O., Bazairi, H., Hereu, B., Linares, C., Kersting, D. K., Rovira, G., Ortega, J., Casals, D., Pagès-Escolà, M., Margarit, N., Capdevila, P., Verdura, J., Ramos, A., Izquierdo, A., Barbera, C., Rubio-Portillo, E., Anton, I., López-Sendino, P., Díaz, D., Vázquez-Luis, M., Duarte, C., Marbà, N., Aspillaga, E., Espinosa, F., Grech, D., Guala, I., Azzurro, E., Farina, S., Gambi, M. C., Chimienti, G., Montefalcone, M., Azzola, A., Pulido Mantas, T., Fraschetti, S., Ceccherelli, G., Kipson, S., Bakran-Petricioli, T., Petricioli, D., Jimenez, C., Katsanevakis, S., Tuney Kizilkaya, I., Kizilkaya, Z., Sartoretto, S., Rouanet, E., Ruitton, S., Comeau, S., Gattuso, J.-P., and Harmelin, J.-G.: Marine heatwaves drive recurrent mass mortalities in the
Mediterranean Sea, Glob. Change Biol., 28, 5708–5725, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib19"><label>Goh et al.(2024)</label><mixed-citation>
      
Goh, E., Yepremyan, A., Wang, J., and Wilson, B.: MAESSTRO: Masked Autoencoders for Sea Surface Temperature Reconstruction under Occlusion, Ocean Sci., 20, 1309–1323, <a href="https://doi.org/10.5194/os-20-1309-2024" target="_blank">https://doi.org/10.5194/os-20-1309-2024</a>, 2024.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib20"><label>Gómez-Gras et al.(2021)</label><mixed-citation>
      
Gómez-Gras, D., Linares, C., López-Sanz, A., Amate, R., Ledoux, J. B.,
Bensoussan, N., Drap, P., Bianchimani, O., Marschal, C., Torrents, O.,
Zuberer, F., Cebrian, E., Teixidó, N., Zabala, M., Kipson, S., Kersting,
D. K., Montero-Serra, I., Pagès-Escolà, M., Medrano, A., Frleta-Valić, M.,
Dimarchopoulou, D., López-Sendino, P., and Garrabou, J.: Population collapse
of habitat-forming species in the Mediterranean: a long-term study of
gorgonian populations affected by recurrent marine heatwaves, P. Roy. Soc. B, 288, 20212384,
<a href="https://doi.org/10.1098/rspb.2021.2384" target="_blank">https://doi.org/10.1098/rspb.2021.2384</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib21"><label>He et al.(2022)</label><mixed-citation>
      
He, K., Chen, X., Xie, S., Li, Y., Dollár, P., and Girshick, R.: Masked
autoencoders are scalable vision learners, in: Proceedings of the IEEE/CVF
conference on computer vision and pattern recognition,  16000–16009,
<a href="https://openaccess.thecvf.com/content/CVPR2022/papers/He_Masked_Autoencoders_Are_Scalable_Vision_Learners_CVPR_2022_paper.pdf" target="_blank"/> (last access: 26 January 2025), 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib22"><label>Hobday et al.(2016)</label><mixed-citation>
      
Hobday, A. J., Alexander, L. V., Perkins, S. E., Smale, D. A., Straub, S. C.,
Oliver, E. C., Benthuysen, J. A., Burrows, M. T., Donat, M. G., Feng, M.,
Holbrook, N. J., Moore, P. J., Scannell, H. A., Sen Gupta, A., and
Wernberg, T.: A hierarchical approach to defining marine heatwaves, Prog. Oceanogr., 141, 227–238,
<a href="https://doi.org/10.1016/j.pocean.2015.12.014" target="_blank">https://doi.org/10.1016/j.pocean.2015.12.014</a>, 2016.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib23"><label>Ličer et al.(2016)</label><mixed-citation>
      
Ličer, M., Smerkol, P., Fettich, A., Ravdas, M., Papapostolou, A., Mantziafou, A., Strajnar, B., Cedilnik, J., Jeromel, M., Jerman, J., Petan, S., Malačič, V., and Sofianos, S.: Modeling the ocean and atmosphere during an extreme bora event in northern Adriatic using one-way and two-way atmosphere–ocean coupling, Ocean Sci., 12, 71–86, <a href="https://doi.org/10.5194/os-12-71-2016" target="_blank">https://doi.org/10.5194/os-12-71-2016</a>, 2016.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib24"><label>Lloyd et al.(2021)</label><mixed-citation>
      
Lloyd, D. T., Abela, A., Farrugia, R. A., Galea, A., and Valentino, G.:
Optically enhanced super-resolution of sea surface temperature using deep
learning, IEEE T. Geosci. Remote Sens., 60, 1–14,
2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib25"><label>Loshchilov and Hutter(2016)</label><mixed-citation>
      
Loshchilov, I. and Hutter, F.: Sgdr: Stochastic gradient descent with warm
restarts, arXiv [preprint], <a href="https://doi.org/10.48550/arXiv.1608.03983" target="_blank">https://doi.org/10.48550/arXiv.1608.03983</a>, 2016.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib26"><label>Martin et al.(2023)</label><mixed-citation>
      
Martin, S. A., Manucharyan, G. E., and Klein, P.: Synthesizing sea surface
temperature and satellite altimetry observations using deep learning improves
the accuracy and resolution of gridded sea surface height anomalies, J. Adv. Model. Earth Sy., 15, e2022MS003589, <a href="https://doi.org/10.1029/2022MS003589" target="_blank">https://doi.org/10.1029/2022MS003589</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib27"><label>Martin et al.(2024)</label><mixed-citation>
      
Martin, S. A., Manucharyan, G. E., and Klein, P.: Deep learning improves global
satellite observations of ocean eddy dynamics, Geophys. Res. Lett.,
51, e2024GL110059, <a href="https://doi.org/10.1029/2024GL110059" target="_blank">https://doi.org/10.1029/2024GL110059</a>, 2024.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib28"><label>Mogen et al.(2022)</label><mixed-citation>
      
Mogen, S. C., Lovenduski, N. S., Dallmann, A. R., Gregor, L., Sutton, A. J.,
Bograd, S. J., Quiros, N. C., Di Lorenzo, E., Hazen, E. L., Jacox, M. G.,
Buil, M. P., and Yeager, S.: Ocean Biogeochemical Signatures of the North
Pacific Blob, Geophys. Res. Lett., 49, e2021GL096938,
<a href="https://doi.org/10.1029/2021GL096938" target="_blank">https://doi.org/10.1029/2021GL096938</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib29"><label>O'Carroll et al.(2019)</label><mixed-citation>
      
O'Carroll, A. G., Armstrong, E. M., Beggs, H. M., Bouali, M., Casey, K. S.,
Corlett, G. K., Dash, P., Donlon, C. J., Gentemann, C. L., Høyer, J. L.,
Ignatov, A., Kabobah, K., Kachi, M., Kurihara, Y., Karagali, I., Maturi, E.,
Merchant, C. J., Marullo, S., Minnett, P. J., Pennybacker, M., Ramakrishnan,
B., Ramsankaran, R., Santoleri, R., Sunder, S., Saux Picart, S.,
Vázquez-Cuervo, J., and Wimmer, W.: Observational Needs of Sea Surface
Temperature, Front. Mar. Sci., 6, 420,
<a href="https://doi.org/10.3389/fmars.2019.00420" target="_blank">https://doi.org/10.3389/fmars.2019.00420</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib30"><label>Pastor and Khodayar(2023)</label><mixed-citation>
      
Pastor, F. and Khodayar, S.: Marine heat waves: Characterizing a major climate
impact in the Mediterranean, Sci. Total Environ., 861, 160621,
<a href="https://doi.org/10.1016/j.scitotenv.2022.160621" target="_blank">https://doi.org/10.1016/j.scitotenv.2022.160621</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib31"><label>Paszke et al.(2017)</label><mixed-citation>
      
Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z.,
Desmaison, A., Antiga, L., and Lerer, A.: Automatic differentiation in
PyTorch, <a href="https://pytorch.org/" target="_blank"/> (last access: 26 January 2025), 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib32"><label>Pisano et al.(2016)</label><mixed-citation>
      
Pisano, A., Buongiorno Nardelli, B., Tronconi, C., and Santoleri, R.: The new
Mediterranean optimally interpolated pathfinder AVHRR SST Dataset
(1982–2012), Remote Sens. Environ., 176, 107–116,
<a href="https://doi.org/10.1016/j.rse.2016.01.019" target="_blank">https://doi.org/10.1016/j.rse.2016.01.019</a>, 2016.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib33"><label>Ricchi et al.(2023)</label><mixed-citation>
      
Ricchi, A., Sangelantoni, L., Redaelli, G., Mazzarella, V., Montopoli, M.,
Miglietta, M. M., Tiesi, A., Mazzà, S., Rotunno, R., and Ferretti, R.:
Impact of the SST and topography on the development of a large-hail storm
event, on the Adriatic Sea, Atmos. Res., 296, 107078,
<a href="https://doi.org/10.1016/j.atmosres.2023.107078" target="_blank">https://doi.org/10.1016/j.atmosres.2023.107078</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib34"><label>Ronneberger et al.(2015)</label><mixed-citation>
      
Ronneberger, O., Fischer, P., and Brox, T.: U-net: Convolutional networks for
biomedical image segmentation, in: Medical image computing and
computer-assisted intervention–MICCAI 2015: 18th international conference,
Munich, Germany, 5–9 October 2015, proceedings, part III 18, 234–241,
Springer, <a href="https://doi.org/10.1007/978-3-319-24574-4_28" target="_blank">https://doi.org/10.1007/978-3-319-24574-4_28</a>, 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib35"><label>Senatore et al.(2020)</label><mixed-citation>
      
Senatore, A., Furnari, L., and Mendicino, G.: Impact of high-resolution sea surface temperature representation on the forecast of small Mediterranean catchments' hydrological responses to heavy precipitation, Hydrol. Earth Syst. Sci., 24, 269–291, <a href="https://doi.org/10.5194/hess-24-269-2020" target="_blank">https://doi.org/10.5194/hess-24-269-2020</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib36"><label>Strajnar et al.(2019)</label><mixed-citation>
      
Strajnar, B., Cedilnik, J., Fettich, A., Ličer, M., Pristov, N., Smerkol, P.,
and Jerman, J.: Impact of two-way coupling and sea-surface temperature on
precipitation forecasts in regional atmosphere and ocean models, Q.
J. Roy. Meteor. Soc., 145, 228–242,
<a href="https://doi.org/10.1002/qj.3425" target="_blank">https://doi.org/10.1002/qj.3425</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib37"><label>Taburet et al.(2019)</label><mixed-citation>
      
Taburet, G., Sanchez-Roman, A., Ballarotta, M., Pujol, M.-I., Legeais, J.-F., Fournier, F., Faugere, Y., and Dibarboure, G.: DUACS DT2018: 25 years of reprocessed sea level altimetry products, Ocean Sci., 15, 1207–1224, <a href="https://doi.org/10.5194/os-15-1207-2019" target="_blank">https://doi.org/10.5194/os-15-1207-2019</a>, 2019.


    </mixed-citation></ref-html>
<ref-html id="bib1.bib38"><label>Ubelmann et al.(2021)</label><mixed-citation>
      
Ubelmann, C., Dibarboure, G., Gaultier, L., Ponte, A., Ardhuin, F., Ballarotta,
M., and Faugère, Y.: Reconstructing ocean surface current combining
altimetry and future spaceborne Doppler data, J. Geophys.
Res.-Oceans, 126, e2020JC016560, <a href="https://doi.org/10.1029/2020JC016560" target="_blank">https://doi.org/10.1029/2020JC016560</a>,  2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib39"><label>Young et al.(2024)</label><mixed-citation>
      
Young, C.-C., Cheng, Y.-C., Lee, M.-A., and Wu, J.-H.: Accurate reconstruction
of satellite-derived SST under cloud and cloud-free areas using a
physically-informed machine learning approach, Remote Sens. Environ.,
313, 114339, <a href="https://doi.org/10.1016/j.rse.2024.114339" target="_blank">https://doi.org/10.1016/j.rse.2024.114339</a>, 2024.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib40"><label>Yu et al.(2018)</label><mixed-citation>
      
Yu, C., Wang, J., Peng, C., Gao, C., Yu, G., and Sang, N.: Bisenet: Bilateral
segmentation network for real-time semantic segmentation, in: Proceedings of
the European conference on computer vision (ECCV),   325–341,
<a href="https://openaccess.thecvf.com/content_ECCV_2018/papers/Changqian_Yu_BiSeNet_Bilateral_Segmentation_ECCV_2018_paper.pdf" target="_blank"/>
(last access: 26 January 2025), 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib41"><label>Zupančič Muc(2025)</label><mixed-citation>
      
Zupančič Muc, M.: CRITER – Coarse Reconstruction with ITerative Refinement network, Zenodo [code], <a href="https://doi.org/10.5281/zenodo.15066015" target="_blank">https://doi.org/10.5281/zenodo.15066015</a>, 2025.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib42"><label>Zupančič Muc et al.(2024)</label><mixed-citation>
      
Zupančič Muc, M., Zavrtanik, V., Barth, A., Alvera-Azcarate, A., Licer, M., and Kristan, M.: CRITER 1.0: Sea Surface Temperature Evaluation Datasets, Zenodo [data set], <a href="https://doi.org/10.5281/zenodo.13923189" target="_blank">https://doi.org/10.5281/zenodo.13923189</a>, 2024.

    </mixed-citation></ref-html>--></article>
