<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing with OASIS Tables v3.0 20080202//EN" "journalpub-oasis3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:oasis="http://docs.oasis-open.org/ns/oasis-exchange/table" xml:lang="en" dtd-version="3.0" article-type="research-article"><?xmltex \makeatother\@nolinetrue\makeatletter?><?xmltex \bartext{Model description paper}?>
  <front>
    <journal-meta><journal-id journal-id-type="publisher">GMD</journal-id><journal-title-group>
    <journal-title>Geoscientific Model Development</journal-title>
    <abbrev-journal-title abbrev-type="publisher">GMD</abbrev-journal-title><abbrev-journal-title abbrev-type="nlm-ta">Geosci. Model Dev.</abbrev-journal-title>
  </journal-title-group><issn pub-type="epub">1991-9603</issn><publisher>
    <publisher-name>Copernicus Publications</publisher-name>
    <publisher-loc>Göttingen, Germany</publisher-loc>
  </publisher></journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.5194/gmd-16-1553-2023</article-id><title-group><article-title>Evaluating a global soil moisture dataset from a multitask model (GSM3 v1.0)
with potential applications for crop threats</article-title><alt-title>Evaluating a global soil moisture dataset from a multitask model</alt-title>
      </title-group><?xmltex \runningtitle{Evaluating a global soil moisture dataset from a multitask model}?><?xmltex \runningauthor{J. Liu et al.}?>
      <contrib-group>
        <contrib contrib-type="author" corresp="no" rid="aff1">
          <name><surname>Liu</surname><given-names>Jiangtao</given-names></name>
          
        <ext-link>https://orcid.org/0000-0002-9219-8354</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff2 aff3 aff4">
          <name><surname>Hughes</surname><given-names>David</given-names></name>
          
        </contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1">
          <name><surname>Rahmani</surname><given-names>Farshid</given-names></name>
          
        <ext-link>https://orcid.org/0000-0001-9241-7206</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1">
          <name><surname>Lawson</surname><given-names>Kathryn</given-names></name>
          
        </contrib>
        <contrib contrib-type="author" corresp="yes" rid="aff1">
          <name><surname>Shen</surname><given-names>Chaopeng</given-names></name>
          <email>cshen@engr.psu.edu</email>
        <ext-link>https://orcid.org/0000-0002-0685-1901</ext-link></contrib>
        <aff id="aff1"><label>1</label><institution>Department of Civil and Environmental Engineering, The Pennsylvania
State University, University Park, PA, USA</institution>
        </aff>
        <aff id="aff2"><label>2</label><institution>Department of Entomology, The Pennsylvania State University,
University Park, PA, USA</institution>
        </aff>
        <aff id="aff3"><label>3</label><institution>Department of Biology, The Pennsylvania State University, University
Park, PA, USA</institution>
        </aff>
        <aff id="aff4"><label>4</label><institution>The Current and Emerging Threat to Crop Innovation Lab, The
Pennsylvania State University, University Park, PA, USA</institution>
        </aff>
      </contrib-group>
      <author-notes><corresp id="corr1">Chaopeng Shen (cshen@engr.psu.edu)</corresp></author-notes><pub-date><day>17</day><month>March</month><year>2023</year></pub-date>
      
      <volume>16</volume>
      <issue>5</issue>
      <fpage>1553</fpage><lpage>1567</lpage>
      <history>
        <date date-type="received"><day>26</day><month>August</month><year>2022</year></date>
           <date date-type="rev-request"><day>30</day><month>September</month><year>2022</year></date>
           <date date-type="rev-recd"><day>27</day><month>January</month><year>2023</year></date>
           <date date-type="accepted"><day>9</day><month>February</month><year>2023</year></date>
      </history>
      <permissions>
        <copyright-statement>Copyright: © 2023 Jiangtao Liu et al.</copyright-statement>
        <copyright-year>2023</copyright-year>
      <license license-type="open-access"><license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p></license></permissions><self-uri xlink:href="https://gmd.copernicus.org/articles/16/1553/2023/gmd-16-1553-2023.html">This article is available from https://gmd.copernicus.org/articles/16/1553/2023/gmd-16-1553-2023.html</self-uri><self-uri xlink:href="https://gmd.copernicus.org/articles/16/1553/2023/gmd-16-1553-2023.pdf">The full text article is available as a PDF file from https://gmd.copernicus.org/articles/16/1553/2023/gmd-16-1553-2023.pdf</self-uri>
      <abstract><title>Abstract</title>

      <p id="d1e139">Climate change threatens our ability to grow food for an ever-increasing population. There is a
need for high-quality soil moisture predictions in under-monitored regions
like Africa. However, it is unclear if soil moisture processes are globally
similar enough to allow our models trained on available in situ data to
maintain accuracy in unmonitored regions. We present a multitask long
short-term memory (LSTM) model that learns simultaneously from global
satellite-based data and in situ soil moisture data. This model is evaluated in
both random spatial holdout mode and continental holdout mode (trained on
some continents, tested on a different one). The model compared favorably to
current land surface models, satellite products, and a candidate machine
learning model, reaching a global median correlation of 0.792 for the random
spatial holdout test. It behaved surprisingly well in Africa and Australia,
showing high correlation even when we excluded their sites from the training
set, but it performed relatively poorly in Alaska where rapid changes are
occurring. In all but one continent (Asia), the multitask model in the
worst-case scenario test performed better than the soil moisture active
passive (SMAP) 9 km product. Factorial analysis has shown that the LSTM model's
accuracy varies with terrain aspect, resulting in lower performance for dry
and south-facing slopes or wet and north-facing slopes. This knowledge
helps us apply the model while understanding its limitations. This model is
being integrated into an operational agricultural assistance application
which currently provides information to 13 million African farmers.</p>
  </abstract>
    
<funding-group>
<award-group id="gs1">
<funding-source>Google</funding-source>
<award-id>1904-57775</award-id>
</award-group>
<award-group id="gs2">
<funding-source>Bill and Melinda Gates Foundation</funding-source>
<award-id>INV-018429</award-id>
</award-group>
</funding-group>
</article-meta>
  </front>
<body>
      

<sec id="Ch1.S1" sec-type="intro">
  <label>1</label><title>Background</title>
      <p id="d1e151">Soil moisture is a critical variable that influences a number of natural
disasters. As a result, widely available, high-quality soil moisture
products can be vital for regions that need aid. Too much soil moisture can
prime the landscape for floods (Norbiato et al., 2008), and too
little of it for too long can damage or kill crops and native vegetation
(Narasimhan
and Srinivasan, 2005; Sheffield and Wood, 2008). Moreover, many insect pests
lay eggs in soils with certain soil moisture conditions – for example,
locusts prefer to lay their eggs in sandy, wet soils
(Hunter-Jones, 1964). In the year 2020, disastrous
locust swarms terrorized large swaths of eastern Africa and Southeast Asia
(Baraniuk, 2020; UN WFP, 2020). The knowledge of soil
moisture levels can be critical in planning pest control activities, as the
immature stages are the best targets for effective control
(Ellenburg et al., 2021; Nuwer,
2021). Besides insect pests, pathogenic fungi and bacteria can be heavily
influenced by soil moisture, resulting in crop losses. In all of these
cases, soil moisture products can be highly valuable in reducing both
current and emerging threats to crops. Finally, the current global crisis in
fertilizer availability following the ongoing war in Europe
(Bentley et al., 2022) necessitates strategies that
increase the efficient use of fertilizer, for which a precise understanding
of soil moisture is critical because water availability in the soil affects
both plant uptake of fertilizer and fertilizer loss.</p>
      <p id="d1e154">Soil moisture is monitored globally by a number of satellite missions and is simulated globally by multiple land surface hydrologic models, but
these products have their<?pagebreak page1554?> respective limitations. Satellite missions like
Soil Moisture Active Passive (SMAP) (Entekhabi, 2010) and Soil
Moisture and Ocean Salinity (SMOS) (Kerr et al., 2010) have
limited spatial resolution and accuracy. When evaluated in comparison to
in situ data, especially on sparsely instrumented sites that are outside of
the missions' core calibration and validation sites, their error can be high
(Al-Yaari
et al., 2017) (also demonstrated later in this work). Land surface models
can also produce decent simulations with seamless spatiotemporal coverage
(Albergel
et al., 2018; Beaudoing and Rodell, 2019; Yang et al., 2011), but they may not
be fully exploiting available information, as evidenced by the better
performance produced by machine learning models where data are available
(Liu et al., 2022a; O and
Orth, 2021). Both satellite and model products may also have a large bias
compared to in situ data.</p>
      <p id="d1e157">Recently, we developed a multiscale time series deep learning (DL) model
that learns simultaneously from satellite and in situ data and can
substantially outperform satellite-based products, a model trained on
in situ data alone, and traditional land surface model simulations
(Liu et al., 2022a). In a spatial cross-validation
test (trained on some sites and tested on others), the multiscale DL model
obtained a median correlation (<inline-formula><mml:math id="M1" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula>) of 0.901 when evaluated by the sparse soil
moisture network over the conterminous United States (CONUS), comparing
favorably to the SMAP 9 km product's <inline-formula><mml:math id="M2" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula> value of 0.762 and the Noah model's <inline-formula><mml:math id="M3" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula> value of 0.761,
and it had minor bias. This work suggested that many previous simulations
have not fully leveraged the available information. In addition, it
demonstrated that multiple sources of datasets may each constrain certain
aspects of a network and train models that outperform each one of its
supervising datasets; i.e., learning from two teachers can be better than one. This multiscale approach can overcome the
limitations with each single dataset.</p>
      <p id="d1e181">However, it is uncertain if the robust model performance from deep networks
in the data-dense CONUS can generalize well to other regions in the world where
hydrological variables are of interest due to potential natural disasters. Typically, the
performance of all kinds of models declines somewhat when applied to
neighboring untrained sites (as in a random holdout test) and then declines
more substantially when applied in a large region without training data
(Feng
et al., 2021; Gauch et al., 2021; Hrachowitz et al., 2013).
Sequence-to-sequence deep networks like long short-term memory (LSTM)
(Hochreiter and Schmidhuber, 1997) can give
us high predictive performance in a range of hydrologic tasks
(Fang
et al., 2017, 2019; Feng et al., 2020; Kratzert et al., 2019; Meyal et al.,
2020; Rahmani et al., 2021b; Shen, 2018; Zhi et al., 2021) because they do
not have rigid model structures and can absorb information more exhaustively
from big data. Their functional behaviors are completely shaped by data, and
thus they can be exempt from many errors in previous models' assumptions. On
the flip side, in data-sparse regions, there is a chance that such a
strength could potentially become a weakness. In Africa especially, there
are very few in situ sites to constrain a model. Recent work has trained
LSTM-based global soil moisture models completely on in situ sites, for
example, the SoMo.ml model (O and Orth, 2021),
but this only learns from in situ locations. It was not clear if optimality had been reached by such
models or if a multitask model learning from both satellite and in situ
data could provide further advantages.</p>
      <p id="d1e185">Regarding the potential for data-driven models in data-scarce regions, there
are two competing hypotheses. The optimistic hypothesis is that surface
soil moisture dynamics are relatively simple to grasp (compared to the
streamflow prediction problem), quite uniform around the world, and well
described by available surface characterization datasets (soil texture);
as a result, the hundreds of publicly available sites can thoroughly train a
DL model that generalizes well in space. The more pessimistic hypothesis is
that the quality of available inputs, e.g., soil texture, is low, meaning that the
number of sites in the world is far from being sufficient to train a
global-scale DL soil moisture model. Confirming one hypothesis or the other
not only influences how we choose a model but may also alter our
understanding about the complexity of the soil moisture prediction problem.</p>
      <p id="d1e188">Given that we would like to have a high-quality product in data-sparse
regions like Africa, we asked three research questions regarding not only the
performance of a global-scale LSTM-based soil moisture model but also the
nature of the soil moisture dynamics:
<list list-type="order"><list-item>
      <p id="d1e193">How well can a LSTM-based soil moisture model perform on the global scale for untrained sites in comparison to existing satellite-based and model-based products?</p></list-item><list-item>
      <p id="d1e197">How well can such a model generalize to highly data-sparse regions; e.g., in an entire continent without data, are soil moisture processes homogeneous enough to permit cross-continental model applications?</p></list-item><list-item>
      <p id="d1e201">What factors control the success or failure of such a model; i.e., can we predict, a priori, if this model can be successful?</p></list-item></list>
We developed and trained a multitask LSTM-based model that learns
simultaneously from both satellite and in situ data. We tested the model in
random holdout and cross-continental experiments to learn its strengths and
weaknesses. We then used a stratified analysis to diagnose where the model
would likely be successful or challenged. In the end, we produced a
globally operational surface soil moisture product that can be leveraged by
non-profit organizations at 9 km resolution.
<?xmltex \hack{\newpage}?></p>
</sec>
<?pagebreak page1555?><sec id="Ch1.S2">
  <label>2</label><title>Data and methods</title>
<sec id="Ch1.S2.SS1">
  <label>2.1</label><title>The multitask LSTM model</title>
      <p id="d1e221">The multitask model based on the long short-term memory (LSTM) algorithm can
be described succinctly as follows:

                <disp-formula specific-use="gather" content-type="numbered"><mml:math id="M4" display="block"><mml:mtable displaystyle="true"><mml:mlabeledtr id="Ch1.E1"><mml:mtd><mml:mtext>1</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:mi mathvariant="normal">LSTM</mml:mi><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>A</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr><mml:mlabeledtr id="Ch1.E2"><mml:mtd><mml:mtext>2</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mi>L</mml:mi><mml:mo>=</mml:mo><mml:mi mathvariant="normal">RMSE</mml:mi><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mi>y</mml:mi><mml:mi mathvariant="normal">s</mml:mi></mml:msup><mml:mo>)</mml:mo><mml:mo>+</mml:mo><mml:mi mathvariant="normal">RMSE</mml:mi><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mi>y</mml:mi><mml:mi mathvariant="normal">in</mml:mi></mml:msup><mml:mo>)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula>

            where <inline-formula><mml:math id="M5" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula> represents simulated soil moisture, <inline-formula><mml:math id="M6" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> represents dynamic atmospheric
forcings, and <inline-formula><mml:math id="M7" display="inline"><mml:mi>A</mml:mi></mml:math></inline-formula> represents static landscape attributes. <inline-formula><mml:math id="M8" display="inline"><mml:mi>L</mml:mi></mml:math></inline-formula> is the loss
function the model tries to minimize, which is based on root-mean-square
error (RMSE). <inline-formula><mml:math id="M9" display="inline"><mml:mrow><mml:msup><mml:mi>y</mml:mi><mml:mi mathvariant="normal">s</mml:mi></mml:msup></mml:mrow></mml:math></inline-formula> represents satellite-based soil moisture products
(SMAP L3, 9 km resolution), and <inline-formula><mml:math id="M10" display="inline"><mml:mrow><mml:msup><mml:mi>y</mml:mi><mml:mi mathvariant="normal">in</mml:mi></mml:msup></mml:mrow></mml:math></inline-formula> represents in situ data (from the
International Soil Moisture Network, ISMN). This model does not use recent
observations and is thus suitable for long-term simulations or trend
predictions but could be enhanced for short-term forecasting via data
assimilation or data integration
(Fang and
Shen, 2020; Feng et al., 2020). This multitask loss function means that the
simulations will attempt to respect both in situ data and satellite data.
Since LSTM has been described extensively in previous work
(Fang
et al., 2019; Feng et al., 2020; Liu et al., 2022a), we omit its mathematical
descriptions here for brevity. Here, because we are now applying it on a
global scale, we chose this multitask scheme over our previous multiscale
scheme (Liu et al., 2022a), which aggregates many
fine-resolution grid cells to match a coarse-resolution grid cell to reduce
computational demand. To avoid over-tuning the hyperparameters, we inherited
most of the parameters from our multi-scale model. Our final parameters were
as follows: a mini-batch size of 128, a hidden-state size of 256, a dropout
rate of 0.5, an epoch length of 100, and a sequence length (rho) of 365 d.</p>
</sec>
<sec id="Ch1.S2.SS2">
  <label>2.2</label><title>The input and training datasets</title>
      <p id="d1e358">We used the SMAP Enhanced L3 Radiometer Global and Polar Grid Daily 9 km
EASE-Grid Soil Moisture version 5 (SPL3SMP_E) product
(O'Neill et al., 2021) as our
satellite target and the International Soil Moisture Network (ISMN) product
as our in situ target
(Dorigo
et al., 2011, 2013). The input data include 18 different meteorological
forcings and 17 different static attributes. We obtained daily leaf area
index (LAI), soil temperature, and surface pressure, among other variables (Table S1 in
the Supplement), from the ECMWF Reanalysis v5 (ERA5)
(Muñoz Sabater, 2019). We tried multiple sources of precipitation
data, including Multi-Source Weighted-Ensemble Precipitation (MSWEP)
(Beck
et al., 2019), Global Precipitation Measurement (GPM)
(Huffman et al., 2019), and ERA5 precipitation
data. Our preliminary results suggested that, in terms of the correlations
of the resulting models, we had the following order: MSWEP<inline-formula><mml:math id="M11" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>GPM<inline-formula><mml:math id="M12" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>ERA5<inline-formula><mml:math id="M13" display="inline"><mml:mo>≈</mml:mo></mml:math></inline-formula>MSWEP<inline-formula><mml:math id="M14" display="inline"><mml:mo>&gt;</mml:mo></mml:math></inline-formula>GPM<inline-formula><mml:math id="M15" display="inline"><mml:mo>&gt;</mml:mo></mml:math></inline-formula>ERA5. Thus, to allow the model to fully
absorb the precipitation information, we include all of MSWEP, GPM, and ERA5
in the input data. Albedo data included black-sky albedo and white-sky albedo
from the Moderate Resolution Imaging Spectroradiometer (MODIS) MCD43A3
version 6 (Schaaf and Wang, 2021). The Land
Surface Temperature (LST) dataset includes LST day and LST night data from
MODIS Land Surface Temperature/Emissivity Daily (MYD11A1) version 6.1
(Wan et al., 2021).</p>
      <p id="d1e396">Static terrain attributes included slope, aspect, plane curvature (pcurv),
elevation, and roughness from the global 1, 5, 10, and 100 km topography database
(Amatulli et al., 2018), and we changed their
resolution from 10 to 9 km using the bilinear interpolation method.
Aspect was determined using the aspect cosine, which is <inline-formula><mml:math id="M16" display="inline"><mml:mrow><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula> for
north-facing slopes and <inline-formula><mml:math id="M17" display="inline"><mml:mrow><mml:mo>&lt;</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula> for south-facing slopes in the Northern
Hemisphere. We further multiplied the aspect cosine in the Southern
Hemisphere by <inline-formula><mml:math id="M18" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> to reflect the Sun's position. Soil physiographic
attributes included sand, clay, and silt fractions and bulk density from
the Harmonized World Soil Database v1.2 (HWSD)
(FAO et al., 2012; Fischer et al., 2008). Other
attributes, including land cover (ESA, 2017) and normalized
difference vegetation index (NDVI; Didan, 2015), were derived from
several satellite products. We averaged all NDVI data from 1 April 2015 to
31 March 2022 to obtain multiple years of static NDVI data and resampled
the data to 9 km using bilinear interpolation. To perform factorial
importance analysis, we also calculated long-term averages of daily LST,
albedo, LST, and SMAP data and used them along with other static attributes
as inputs in the LSTM model and random forest model (using <inline-formula><mml:math id="M19" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula> of the tests as the target, and removing duplication caused by multi-year average values). All attributes and their sources are
listed in Table S1 in the Supplement.</p>
      <p id="d1e436">To train the model and evaluate its performance, we used soil moisture
measurements (m<inline-formula><mml:math id="M20" display="inline"><mml:msup><mml:mi/><mml:mn mathvariant="normal">3</mml:mn></mml:msup></mml:math></inline-formula> m<inline-formula><mml:math id="M21" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>) from the International Soil Moisture Network
(ISMN) (Table S2). The ISMN is an international
collaboration where soil moisture measurements are collected from dozens of
soil moisture networks across the world. We selected site data from ISMN
with <inline-formula><mml:math id="M22" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">5</mml:mn></mml:mrow></mml:math></inline-formula> cm depth and aggregated the hourly data into daily
data. We used a total of 1317 sites, located across Africa (18), Asia (115),
Europe (129), the CONUS (969), Alaska (44), and Australia (19). The remaining sites are distributed across some islands. Based on the
site clustering in Africa, we divided the data for Africa into northern Africa
and southern Africa according to latitudes 1.8 to 19.3 and <inline-formula><mml:math id="M23" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">38.9</mml:mn></mml:mrow></mml:math></inline-formula> to <inline-formula><mml:math id="M24" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">22.0</mml:mn></mml:mrow></mml:math></inline-formula>,
respectively.</p>
</sec>
<sec id="Ch1.S2.SS3">
  <label>2.3</label><title>The models and products for comparisons</title>
      <p id="d1e498">We compared the results with a wealth of data products and algorithms to put
the proposed method into context. These include the SMAP-L3 enhanced 9 km
product (O'Neill et al., 2021),
the SMOS-L3 product
(Al Bitar et al.,
2017; Support CATDS, 2022), the LPRM_AMSR2_DS_ A_SOILM3 product
(de Jeu and Owe, 2013; Owe et al.,
2008), the NOAH025<?pagebreak page1556?> (10 cm depth) model from the Global Land Data
Assimilation System (GLDAS)
(Beaudoing and Rodell,
2019; Rodell et al., 2004), and another machine learning model, SoMo.ml
(O and Orth, 2021). SMAP-L3 and SMOS-L3 are the
low-frequency-pass microwave products that provide a composite of daily
estimates of global land surface soil moisture retrieved by the L band at
9 and 25 km resolution, respectively. LPRM_AMSR2_ DS_ A_SOILM3 (denoted as
AMSR2) is a high-frequency-pass microwave product, and we used the X-band
data to estimate global soil moisture (de
Jeu and  Owe, 2013; Owe et al., 2008). GLDAS_NOAH025
integrates ground-based observation data and satellite data to drive land
surface models to estimate hydrologic variables including soil moisture. It
is to be noted that SMAP and GLDAS products were not optimized to match the
sparse in situ networks, meaning that this comparison is not entirely fair, but they
were shown to provide context.</p>
      <p id="d1e501">Another machine-learning-based model, SoMo.ml, obtained by an LSTM model
trained solely on global in situ networks (O and
Orth, 2021), has been evaluated on global in situ networks using the spatial
cross-validation method. Notably, the SoMo.ml product provides soil moisture
estimation from 0–10 cm depth rather than 0–5 cm depth. Its final product was
obtained by retraining the model using all available sites and times rather
than by using spatial cross-validation (spatial cross-validation is regarded
as a more rigorous test, so this comparison puts our model at a
disadvantage). The SoMo.ml model also differs from the multitask model as it
uses different input data, only in situ data in calculating the loss
function, and a sequence-to-one structure. Despite these differences, we
still think a best effort at comparison could be useful to the community.
The model performance under different experiments is compared with the ISMN
in situ data, while the final product input and output data are both global
9 km grid data. All of the comparison datasets and results are listed in
Tables S3 and S4, respectively. We also
resampled the model's input data and the other products to retrain a new
model. They were compared at the same resolution of 0.25<inline-formula><mml:math id="M25" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula>. The
model's performance dropped slightly, but the results supported the same
conclusions as the 9 km resolution (Table S5).</p>
</sec>
<sec id="Ch1.S2.SS4">
  <label>2.4</label><title>The experiments</title>
      <p id="d1e521">To understand the model's performance for short-distance spatial
interpolation, we ran random 5-fold cross-validation for random spatial
tests. To understand performance for long-distance spatial extrapolation, we
slightly modified this procedure and ran cross-continental tests. We also
ran a 7-fold cross-validation experiment. However, there was no significant
difference in their results. To save computational resources, we showed the
results from the 5-fold experiments. We randomly separated the in situ and
satellite data into five groups. In each round, we used four of the five groups to
train the multitask model and used the remaining one for testing. We
repeated this for a total of five rounds so that each point was tested. In the
cross-continental test, we divided the global data into seven large regions. In
each round, we kept one region's data as the test set and used the rest as
the training dataset. We repeated this process seven times so that each region
was treated as the test region once. Both the spatial training and test
periods were from 1 April 2015 to 31 December 2020. We also ran temporal
tests, for which the training period was from 1 April 2016 to 31 December 2020, and the test period was from 1 April 2015 to 31 March 2016.</p>
</sec>
<sec id="Ch1.S2.SS5">
  <label>2.5</label><title>Analysis of controls of model performance</title>
      <p id="d1e533">We used a stratified analysis to explain which variables may have had
control over the model's performance. We first trained a random forest model
from the scikit-learn library (Pedregosa et al., 2011) in order to
identify the first few important factors influencing correlation (<inline-formula><mml:math id="M26" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula>) of the
LSTM model in these experiments. A random forest (RF) model is a
classification and regression algorithm consisting of many decision trees that
use bagging and randomness of features to create a series of decision trees.
It is suitable for nonlinear data and reduces the risk of overfitting.
Briefly, RF uses a collection of decision trees to predict the <inline-formula><mml:math id="M27" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula> values. At
the nodes of each tree, the data are split into two bins to minimize the
variance of the bins after the split. Therefore, we could calculate the
average contribution of each factor to the reduction in variance and then
obtain the ranking of importance. Note that the importance ranking is not about factor A being important for predicting soil moisture, but rather whether there are certain ranges of a factor, or joint ranges of multiple factors, where the model behaves more poorly than other ranges. From the importance results, we chose two
importance factors and plotted <inline-formula><mml:math id="M28" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula> as a function of these factors to explore
and interpret how they controlled model performance. The goals were to gain a
physical interpretation of why the model sometimes produced lower-quality
outputs and offer some possible guidance about when to be more cautious in
relying on model results.</p>
</sec>
<sec id="Ch1.S2.SS6">
  <label>2.6</label><title>Evaluation metrics</title>
      <p id="d1e565">The metrics used to evaluate the multitask model's performance include
Pearson's correlation coefficient (Corr), bias, root-mean-square error
(RMSE), and unbiased RMSE (ubRMSE), in which RMSE is calculated after bias
is removed. These metrics are the median value of all satellite grids and
in situ measurements. When we calculated these metrics, we removed the observed and
predicted data when there was a “nan” value (not a number; an error) in the
observation.</p><?xmltex \hack{\newpage}?>
</sec>
</sec>
<?pagebreak page1557?><sec id="Ch1.S3">
  <label>3</label><title>Results and discussion</title>
<sec id="Ch1.S3.SS1">
  <label>3.1</label><title>Error types and temporal tests</title>
      <p id="d1e585">Before we dive into the results, we first need to discuss several error
types so it is easier to interpret the results. We can roughly separate soil
moisture modeling errors into multiple components: (i) climatic forcing
errors, (ii) training data limitations and nonstationarity (e.g., the model
being unable to learn the correct response to drastic changes that have
never been seen before), (iii) errors due to uncaptured spatial heterogeneity
in soil properties, and (iv) model training errors (i.e., overfitting or
underfitting to the training data resulting in mismatches for the testing
data). Among these, components (i) and (ii) are likely to manifest as errors in the temporal
tests. Component (ii) especially appears as large temporal test errors when compared to
the spatial test errors, which would indicate strong nonstationarity. Components (ii)
and (iii) can both be reduced when there are more numerous or more accurate training
data. Component (iii) appears as a large error in the spatial test, indicating either that
the available soil property data are not accurate or diverse enough to
reflect the impacts of soil texture or that there are local hydrologic
processes, e.g., riverine inundation or irrigation, that are unknown to the
LSTM (not contained in the inputs). Component (iii) will also modestly decrease as data
density increases (as training sites inherently become closer together) but
typically cannot be removed entirely. Component (iv) appears as a large difference
between training and testing metrics. It is worthwhile to note that
LSTM models typically (although not always) perform better for each site when
given data from more numerous or more diverse sites due to a
“data synergy” effect (Fang et al., 2022).</p>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T1" specific-use="star"><?xmltex \currentcnt{1}?><label>Table 1</label><caption><p id="d1e591">Model performance in three scenarios. <bold>(a)</bold> The model's temporal
testing in different regions. <bold>(b)</bold> The model's spatial cross-validation
testing in different regions. <bold>(c)</bold> The model's continental cross-validation
testing in different regions.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="9">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:colspec colnum="5" colname="col5" align="right"/>
     <oasis:colspec colnum="6" colname="col6" align="right"/>
     <oasis:colspec colnum="7" colname="col7" align="right"/>
     <oasis:colspec colnum="8" colname="col8" align="right"/>
     <oasis:colspec colnum="9" colname="col9" align="right"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry namest="col1" nameend="col9"><bold>(a)</bold> Temporal testing </oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Median metrics</oasis:entry>
         <oasis:entry colname="col2">CONUS</oasis:entry>
         <oasis:entry colname="col3">Europe</oasis:entry>
         <oasis:entry colname="col4">Africa_North</oasis:entry>
         <oasis:entry colname="col5">Alaska</oasis:entry>
         <oasis:entry colname="col6">Asia</oasis:entry>
         <oasis:entry colname="col7">Africa_South</oasis:entry>
         <oasis:entry colname="col8">Australia</oasis:entry>
         <oasis:entry colname="col9">Global</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Bias</oasis:entry>
         <oasis:entry colname="col2">0.003</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M29" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.005</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M30" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.014</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col5">0.001</oasis:entry>
         <oasis:entry colname="col6"><inline-formula><mml:math id="M31" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.007</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col7"><inline-formula><mml:math id="M32" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.044</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col8">0.001</oasis:entry>
         <oasis:entry colname="col9">0.001</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">RMSE</oasis:entry>
         <oasis:entry colname="col2">0.051</oasis:entry>
         <oasis:entry colname="col3">0.058</oasis:entry>
         <oasis:entry colname="col4">0.031</oasis:entry>
         <oasis:entry colname="col5">0.075</oasis:entry>
         <oasis:entry colname="col6">0.049</oasis:entry>
         <oasis:entry colname="col7">0.056</oasis:entry>
         <oasis:entry colname="col8">0.055</oasis:entry>
         <oasis:entry colname="col9">0.051</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">ubRMSE</oasis:entry>
         <oasis:entry colname="col2">0.043</oasis:entry>
         <oasis:entry colname="col3">0.037</oasis:entry>
         <oasis:entry colname="col4">0.026</oasis:entry>
         <oasis:entry colname="col5">0.056</oasis:entry>
         <oasis:entry colname="col6">0.035</oasis:entry>
         <oasis:entry colname="col7">0.048</oasis:entry>
         <oasis:entry colname="col8">0.044</oasis:entry>
         <oasis:entry colname="col9">0.043</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Corr</oasis:entry>
         <oasis:entry colname="col2">0.847</oasis:entry>
         <oasis:entry colname="col3">0.808</oasis:entry>
         <oasis:entry colname="col4">0.881</oasis:entry>
         <oasis:entry colname="col5">0.654</oasis:entry>
         <oasis:entry colname="col6">0.873</oasis:entry>
         <oasis:entry colname="col7">0.656</oasis:entry>
         <oasis:entry colname="col8">0.877</oasis:entry>
         <oasis:entry colname="col9">0.837</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry namest="col1" nameend="col9"><bold>(b)</bold> Spatial cross-validation testing </oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Bias</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M33" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.004</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">0.004</oasis:entry>
         <oasis:entry colname="col4">0.029</oasis:entry>
         <oasis:entry colname="col5">0.007</oasis:entry>
         <oasis:entry colname="col6">0.011</oasis:entry>
         <oasis:entry colname="col7">0.021</oasis:entry>
         <oasis:entry colname="col8"><inline-formula><mml:math id="M34" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.010</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col9"><inline-formula><mml:math id="M35" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.0003</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">RMSE</oasis:entry>
         <oasis:entry colname="col2">0.075</oasis:entry>
         <oasis:entry colname="col3">0.080</oasis:entry>
         <oasis:entry colname="col4">0.067</oasis:entry>
         <oasis:entry colname="col5">0.079</oasis:entry>
         <oasis:entry colname="col6">0.067</oasis:entry>
         <oasis:entry colname="col7">0.074</oasis:entry>
         <oasis:entry colname="col8">0.096</oasis:entry>
         <oasis:entry colname="col9">0.075</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">ubRMSE</oasis:entry>
         <oasis:entry colname="col2">0.057</oasis:entry>
         <oasis:entry colname="col3">0.057</oasis:entry>
         <oasis:entry colname="col4">0.048</oasis:entry>
         <oasis:entry colname="col5">0.053</oasis:entry>
         <oasis:entry colname="col6">0.052</oasis:entry>
         <oasis:entry colname="col7">0.071</oasis:entry>
         <oasis:entry colname="col8">0.055</oasis:entry>
         <oasis:entry colname="col9">0.056</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Corr</oasis:entry>
         <oasis:entry colname="col2">0.790</oasis:entry>
         <oasis:entry colname="col3">0.791</oasis:entry>
         <oasis:entry colname="col4">0.861</oasis:entry>
         <oasis:entry colname="col5">0.789</oasis:entry>
         <oasis:entry colname="col6">0.762</oasis:entry>
         <oasis:entry colname="col7">0.647</oasis:entry>
         <oasis:entry colname="col8">0.778</oasis:entry>
         <oasis:entry colname="col9">0.792</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry namest="col1" nameend="col9"><bold>(c)</bold> Continental cross-validation </oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Bias</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M36" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.009</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M37" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.016</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">0.041</oasis:entry>
         <oasis:entry colname="col5">0.039</oasis:entry>
         <oasis:entry colname="col6">0.052</oasis:entry>
         <oasis:entry colname="col7"><inline-formula><mml:math id="M38" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.067</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col8"><inline-formula><mml:math id="M39" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.016</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col9"><inline-formula><mml:math id="M40" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.002</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">RMSE</oasis:entry>
         <oasis:entry colname="col2">0.099</oasis:entry>
         <oasis:entry colname="col3">0.104</oasis:entry>
         <oasis:entry colname="col4">0.047</oasis:entry>
         <oasis:entry colname="col5">0.119</oasis:entry>
         <oasis:entry colname="col6">0.092</oasis:entry>
         <oasis:entry colname="col7">0.078</oasis:entry>
         <oasis:entry colname="col8">0.065</oasis:entry>
         <oasis:entry colname="col9">0.098</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">ubRMSE</oasis:entry>
         <oasis:entry colname="col2">0.071</oasis:entry>
         <oasis:entry colname="col3">0.062</oasis:entry>
         <oasis:entry colname="col4">0.032</oasis:entry>
         <oasis:entry colname="col5">0.075</oasis:entry>
         <oasis:entry colname="col6">0.055</oasis:entry>
         <oasis:entry colname="col7">0.052</oasis:entry>
         <oasis:entry colname="col8">0.061</oasis:entry>
         <oasis:entry colname="col9">0.068</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Corr</oasis:entry>
         <oasis:entry colname="col2">0.605</oasis:entry>
         <oasis:entry colname="col3">0.646</oasis:entry>
         <oasis:entry colname="col4">0.87</oasis:entry>
         <oasis:entry colname="col5">0.581</oasis:entry>
         <oasis:entry colname="col6">0.711</oasis:entry>
         <oasis:entry colname="col7">0.718</oasis:entry>
         <oasis:entry colname="col8">0.806</oasis:entry>
         <oasis:entry colname="col9">0.624</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <?xmltex \floatpos{t}?><fig id="Ch1.F1" specific-use="star"><?xmltex \currentcnt{1}?><?xmltex \def\figurename{Figure}?><label>Figure 1</label><caption><p id="d1e1161">Comparison of model performances for different continents in
data-rich regions. Models from left to right are ranked from lowest to
highest global correlation. We plotted results for the training period, as
well as temporal, spatial, and cross-continental tests.
“Multitask_exclude” means the cross-continent test: the
models were tested on a continent, but sites from that continent were
excluded from training. The SoMo.ml product shown here was trained on all
sites in all time periods, and thus it is most comparable to our
“Multitask_train” product.</p></caption>
          <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/16/1553/2023/gmd-16-1553-2023-f01.png"/>

        </fig>

      <p id="d1e1171">The temporal tests (trained on some sites in one time period and tested on
the same sites in another time period), which are used to establish a
reference performance level, showed a strong ability for LSTM to capture
soil moisture dynamics around the world, posting a global median correlation
(<inline-formula><mml:math id="M41" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula>) of 0.837 (Table 1a and the sky-blue box, Multitask_temporal, in
Figs. 1 and 2). Because LSTM has learned from the history of the sites,
these test-region-aggregated temporal test metrics are normally higher than
spatial tests (except for Alaska, which is to be discussed below) and
reflect the inherent and geographically varying difficulties of soil
moisture modeling in different regions. The temporal test <inline-formula><mml:math id="M42" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula> values for
different regions are organized in the following order: Africa_North <inline-formula><mml:math id="M43" display="inline"><mml:mo>&gt;</mml:mo></mml:math></inline-formula> Australia <inline-formula><mml:math id="M44" display="inline"><mml:mo>&gt;</mml:mo></mml:math></inline-formula> Asia <inline-formula><mml:math id="M45" display="inline"><mml:mo>&gt;</mml:mo></mml:math></inline-formula> CONUS <inline-formula><mml:math id="M46" display="inline"><mml:mo>&gt;</mml:mo></mml:math></inline-formula> Europe <inline-formula><mml:math id="M47" display="inline"><mml:mo>&gt;</mml:mo></mml:math></inline-formula> Africa_South <inline-formula><mml:math id="M48" display="inline"><mml:mo>≅</mml:mo></mml:math></inline-formula> Alaska. One immediately
apparent observation is that this order is not related to the number of
sites in each region or the density of sites. For example, the
highest-ranking (in terms of <inline-formula><mml:math id="M49" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula>) regions are Africa_North,
Australia, and Asia, which all are among the regions with the lowest counts
of sites. Alaska has a relatively high site density but had the lowest
median <inline-formula><mml:math id="M50" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula>, which could be attributed to the unique difficulties associated
with frozen soil and thawing permafrost. This observation suggests that more
training sites in these regions may not result in significantly better
temporal test results at existing sites. Africa_South was
more difficult than Africa_North, presumably because more
sites are located in arid environments (LSTM has previously shown lower
performance in such regions in the CONUS, as discussed in Feng et al.,
2020). While these results show that
there are some regions in the world that are more difficult to capture than
others for the prediction of soil moisture, the overall results are
encouraging. The model's performance over these regions indicates that the
quality of the forcing (MSWEP<inline-formula><mml:math id="M51" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>GPM<inline-formula><mml:math id="M52" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>ERA5 precipitation) and soil
characterization data is globally consistent.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F2" specific-use="star"><?xmltex \currentcnt{2}?><?xmltex \def\figurename{Figure}?><label>Figure 2</label><caption><p id="d1e1262">The same as Fig. 1 but for data-sparse regions, i.e., Africa, Asia, and
Australia.</p></caption>
          <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/16/1553/2023/gmd-16-1553-2023-f02.png"/>

        </fig>

      <p id="d1e1271">Apart from Alaska, there were no particularly strong spatial patterns in
either <inline-formula><mml:math id="M53" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula> or RMSE in the random spatial (cross-validation) tests (Fig. 3).
Over the CONUS, there was a mild concentration of poorly performing sites in
the northwest. In Europe, we found poorly performing sites in the central
region, e.g., Hungary and Romania. Other than that, poorly performing sites
were interspersed among the well-performing sites, suggesting that most of
the causes of poor performance are local rather than climatic effects, which
we will explore in Sect. 3.3. The cross-continental tests led to a
widespread decrease in <inline-formula><mml:math id="M54" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula>, in comparison to the random spatial tests (Fig. 4). While some African sites, like those immediately south of the Sahara (Fig. 4c), had noticeably deteriorated performance, some other
sites in fact improved, such as the three most southern sites in
Africa_South (Fig. 4f).</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F3" specific-use="star"><?xmltex \currentcnt{3}?><?xmltex \def\figurename{Figure}?><label>Figure 3</label><caption><p id="d1e1290">Metric distributions for the multitask model random spatial
cross-validation tests. <bold>(I)</bold> Correlation and <bold>(II)</bold> RMSE of spatial
cross-validation tests for <bold>(a)</bold> the CONUS, <bold>(b)</bold> Europe, <bold>(c)</bold> Africa_North, <bold>(d)</bold> Alaska, <bold>(e)</bold> Asia, <bold>(f)</bold> Africa_South, and <bold>(g)</bold> Australia. The training and testing periods were both from
1 April 2015 to 31 December 2020. Maps are made with Natural Earth
imagery, no permission needed (Natural Earth, 2022).</p></caption>
          <?xmltex \igopts{width=483.69685pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/16/1553/2023/gmd-16-1553-2023-f03.png"/>

        </fig>

</sec>
<sec id="Ch1.S3.SS2">
  <label>3.2</label><title>Randomly sampled spatial cross-validation</title>
      <p id="d1e1335">The random spatial (randomly sampled cross-validation) tests, which examined
the effect of spatial interpolation, showed record-breaking results despite
their slight performance decline compared to the temporal tests. The global
median <inline-formula><mml:math id="M55" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula> was 0.792, ubRMSE was 0.056, and RMSE was 0.075, all of which were
slightly better than the CONUS median values (Table 1b and the
wheat-colored box, Multitask_spatial, in Figs. 1 and 2). In contrast,
the SMAP 9 km product and GLDAS had global median <inline-formula><mml:math id="M56" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula> values of 0.621 and
0.608, respectively. It should be noted that the numbers are not entirely
comparable: SMAP 9 km and GLDAS were not calibrated fully on the sparse
in situ sites. As expected, at the global scale, the training metrics were
slightly better and had a smaller spread than those for the temporal and
spatial tests. The LSTM-based SoMo.ml model obtained a median <inline-formula><mml:math id="M57" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula> of
<inline-formula><mml:math id="M58" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">0.6</mml:mn></mml:mrow></mml:math></inline-formula> for spatial cross-validation (Fig. 7 in O and Orth,
2021), while the downloadable SoMo.ml product (0.805, as shown in Figs. 1
and 2) was obtained based on training on all the sites and time periods and
thus should in fact be compared to the our training period results (the
rightmost box in each panel, Multitask_train; <inline-formula><mml:math id="M59" display="inline"><mml:mrow><mml:mi>R</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.853</mml:mn></mml:mrow></mml:math></inline-formula>). It should be noted that SoMo.ml has
soil moisture for multiple depths, but we only explored the 5 cm product here.
The closest model to the multitask LSTM is the one from Beck et al. (2021) (we<?pagebreak page1558?> do
not have the data to plot their results), which was calibrated on 177 of the
soil moisture sites and tested on the others. Their MSWEP<inline-formula><mml:math id="M60" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>HBV (Hydrologiska Byråns Vattenbalansavdelning) model
obtained a median <inline-formula><mml:math id="M61" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula> value of 0.78. Their performance is competitive and
quite impressive for a process-based model, but unfortunately the HBV model
only outputs a water storage value (in mm) that can be correlated to the
fluctuation of observed soil moisture and not the soil moisture itself, and
thus other metrics like bias cannot be calculated (additional linear
transformations are required to obtain soil moisture, which introduces
uncertainty). It would be interesting to explore how HBV or similar models
would react to the cross-continental test below, where it may show some
advantages.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F4" specific-use="star"><?xmltex \currentcnt{4}?><?xmltex \def\figurename{Figure}?><label>Figure 4</label><caption><p id="d1e1398">Metric distributions for the multitask model continental
cross-validation tests. <bold>(I)</bold> Correlation and <bold>(II)</bold> RMSE of continental
cross-validation tests for <bold>(a)</bold> the CONUS, <bold>(b)</bold> Europe, <bold>(c)</bold> Africa_North, <bold>(d)</bold> Alaska, <bold>(e)</bold> Asia, <bold>(f)</bold> Africa_South, and <bold>(g)</bold> Australia. The training and testing periods were both from
1 April 2015 to 31 December 2020. Maps are made with Natural Earth
imagery, no permission needed (Natural Earth, 2022).</p></caption>
          <?xmltex \igopts{width=483.69685pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/16/1553/2023/gmd-16-1553-2023-f04.png"/>

        </fig>

      <p id="d1e1435">In general, the difference between the training and temporal tests is small,
and thus we regard the model training error to be small. Switching from the
temporal test to the random spatial test, most regions suffered a small
decline in performance, suggesting that the impact of spatial heterogeneity is
larger than the impact of temporal nonstationarity for soil moisture
predictions. Regions seeing noticeable declines include the CONUS (from
0.847 to 0.790), Asia (0.873 to 0.762), and Australia (0.877 to 0.778),
which could reflect the limited quality of soil texture data and
processes that cannot be described by the input attributes. Alaska stood out
as the exception (temporal test <inline-formula><mml:math id="M62" display="inline"><mml:mrow><mml:mi>R</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.654</mml:mn></mml:mrow></mml:math></inline-formula>; spatial test <inline-formula><mml:math id="M63" display="inline"><mml:mrow><mml:mi>R</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.789</mml:mn></mml:mrow></mml:math></inline-formula>), which is
in fact consistent with our theory of errors discussed earlier and
highlights the rapid changes facing Arctic regions. Alaska is challenging
because it is the frontier of rapid changes in permafrost thawing and months
of frozen ground conditions. As a result, temporal nonstationarity in that
region trumps spatial heterogeneity. This observation suggests that the soil
moisture dynamics in Arctic regions in the coming years will differ
materially from those in the past decades.</p>
      <p id="d1e1463">Precipitation data quality exerts an important influence on the performance
of the model but does not materially change the model comparisons. MSWEP is
a high-quality global precipitation dataset (daily, 0.1<inline-formula><mml:math id="M64" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula>
resolution) arising from blending multiple forcing datasets and correcting
their biases
(Beck
et al., 2019). To support a fair comparison, we also ran our multitask model
with the more widely used ERA5 precipitation data, which gave slightly
lower-performing results. The multiscale model still outperformed reference
products (Table S6). We further note that our
previous CONUS results used the NLDAS-2 (North American Land Data Assimilation System Phase 2) forcing data, which are more customized
toward North America, and obtained an <inline-formula><mml:math id="M65" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula> of 0.901
(Liu et al., 2022a). We thus conclude that the
forcing dataset used has a moderate impact on results and needs to be the
same for models to be fully comparable.</p>
</sec>
<sec id="Ch1.S3.SS3">
  <label>3.3</label><title>Cross-continental tests</title>
      <?pagebreak page1559?><p id="d1e1490">As expected, model performances dropped significantly in the
cross-continental test (testing on a continent where no training data was
provided), but even under this adverse situation, the multitask LSTM model
surpassed or equaled the performance of SMAP in all regions except Asia
(Tables 1c, S4, Figs. 1 and 2). When the CONUS was included as a
training region, the <inline-formula><mml:math id="M66" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula> value in all regions except Alaska stayed above 0.64.
When both the CONUS and Europe were included (again, except for Alaska),
there seemed to be a baseline performance level (<inline-formula><mml:math id="M67" display="inline"><mml:mrow><mml:mi>R</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.70</mml:mn></mml:mrow></mml:math></inline-formula>) that the model
would not fall below, despite there being no training data from the test
continent. For Africa and Australia, the advantages of the multitask model
(multitask_exclude in Fig. 2) over SMAP or GLDAS are
prominent. This suggests that we could consider the
multitask LSTM model to be a viable product even in the no-data scenario.</p>
      <p id="d1e1512">Interestingly, the fewer sites a region had, the less impact there was by
switching from the random spatial to the cross-continental test (Table 1c
and the bright red boxes, Multitask_exclude, in Figs. 1 and 2). For Africa_North, Africa_South, and Australia, there was no decline from
the random spatial test (Table 1b) to the cross-continental test (Table 1c).
Asia saw a larger impact, with <inline-formula><mml:math id="M68" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula> dropping from 0.762 in the spatial test to
0.711 in the cross-continental test. We noticed precipitous drops for Alaska
(median <inline-formula><mml:math id="M69" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula> from 0.789 to 0.581 – again suggesting soil moisture dynamics
there are materially different from other parts of the world), Europe (0.791
to 0.646), and the CONUS (0.790 to 0.605). We thus conclude that when a
region had very few sites but high heterogeneity, these sites only played a
minor role in training the model, and thus removing them did not materially
change the model. When a region had a large number of sites, like the CONUS
or Europe, removing them substantially reduced the training data diversity.
The quality of a DL model is a strong function of its training data – thus
it would be severely weakened if a large part of its training data were
removed.</p>
      <?pagebreak page1561?><p id="d1e1529">It is also interesting that African, Australian, and Asian sites had good
performance in this experiment. It seems to suggest their soil types and
rainfall moisture responses have already been covered by similar sites in
the CONUS and Europe, and thus the model was already sufficient. With
diverse climates and landscapes ranging from desert to temperate forest and
from croplands to wetlands, the CONUS networks play a dominant role in
providing training on how soil moisture responds to different forcings, as
modulated by the soil and landscape characteristics. We cannot know for
certain that the model will work well in other parts of Africa and Asia
until we have more in situ sites there. However, the current results at
least can make us hopeful that the model will likely produce good results
in some parts of the untrained world and will likely add value beyond
satellite products.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F5" specific-use="star"><?xmltex \currentcnt{5}?><?xmltex \def\figurename{Figure}?><label>Figure 5</label><caption><p id="d1e1535">The feature importance determined from a random forest (RF) model
constructed to predict temporal test <inline-formula><mml:math id="M70" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula> using all the unduplicated category
data presented in this paper as inputs. The correlation of this model is
0.6. Aspect, average soil moisture, and downward radiation are the top three
factors. A separate gradient-boosted decision model was also trained and obtained
a correlation of 0.77, and the top three important factors were similar:
slope aspect, precipitation, and downward solar radiation.</p></caption>
          <?xmltex \igopts{width=398.338583pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/16/1553/2023/gmd-16-1553-2023-f05.png"/>

        </fig>

</sec>
<sec id="Ch1.S3.SS4">
  <label>3.4</label><title>Factorial influences on model performance</title>
      <p id="d1e1559">Due to LSTM's strong ability to fit to data, it can serve as a probe for
process complexity
(Liu et al., 2022a; Feng et al., 2022, 2020; Tsai et al., 2021): those sites that
LSTM cannot adequately capture may contain complicated processes that are
not well described by the inputs. The factorial importance analysis
indicates that slope aspect, average soil moisture, and surface solar
radiation downwards are the top three factors that influence the multitask
LSTM model's <inline-formula><mml:math id="M71" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula> in the temporal test (Fig. 5). The RF model has a test
correlation of 0.6 (with 80 % training data and 20 % test data), but its
only purpose here is to provide a reading on the top three factors. We have
also tried using gradient-boosted decision trees (Friedman,
2001), which produced a test correlation of 0.77, and the top three important
factors were slope aspect, precipitation, and surface solar radiation
downwards. Therefore, this model choice does not qualitatively affect our
conclusions. As a reminder, the feature importance test is based on training
random forest (RF) models with the inputs listed in Fig. 5, and <inline-formula><mml:math id="M72" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula> from
the temporal test serves as the target. A high-ranking factor in Fig. 5
implies that it has influence not only on soil moisture but also on the
predictability of soil moisture. It could be that in a certain range of this
factor (the range may be conditional on other factors due to factorial
interactions) there are not that many sites of this kind (it is a minority
class that is not well represented in the training dataset) or that some
latent processes become important. Nevertheless, due to the inherent
limitations of machine learning, factorial importance is only a hypothesis
rather than confirmed truth (Tsai et al.,
2021). As a result, human interpretation of the results will be required.
Because the sensitivity to radiation is somewhat difficult to interpret,
here we focus on average soil moisture and aspect.</p>

      <?xmltex \floatpos{p}?><fig id="Ch1.F6" specific-use="star"><?xmltex \currentcnt{6}?><?xmltex \def\figurename{Figure}?><label>Figure 6</label><caption><p id="d1e1578">Stratified analysis of the distribution of <inline-formula><mml:math id="M73" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula> values from <bold>(I)</bold> temporal and <bold>(II)</bold> spatial tests. <bold>(a–g)</bold> The maps show the global
distribution of test sites as a function of average SMAP soil moisture value
and aspect. The colors on the map represent aspect cosine. The average
SMAP <inline-formula><mml:math id="M74" display="inline"><mml:mrow><mml:mo>&lt;</mml:mo><mml:mn mathvariant="normal">0.08</mml:mn></mml:mrow></mml:math></inline-formula> sites are a minority class and are represented by
squares. <bold>(h)</bold> The SMAP boxplot shows the distribution of <inline-formula><mml:math id="M75" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula> under different
average soil moisture values (SMAP). <bold>(i)</bold> The aspect boxplot shows the
distribution of <inline-formula><mml:math id="M76" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula> in different aspect cosine bins, where the left boxplot in each bin
indicates SMAP <inline-formula><mml:math id="M77" display="inline"><mml:mrow><mml:mo>≦</mml:mo><mml:mn mathvariant="normal">0.08</mml:mn></mml:mrow></mml:math></inline-formula>, and the right boxplot in each bin indicates
SMAP <inline-formula><mml:math id="M78" display="inline"><mml:mrow><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">0.08</mml:mn></mml:mrow></mml:math></inline-formula>. The upper half of the figure, part I, shows temporal test <inline-formula><mml:math id="M79" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula> values (which
characterize temporal nonstationarity), while the lower half of the figure, part II, shows spatial
test <inline-formula><mml:math id="M80" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula> values, which characterize the effect of spatial heterogeneity. Maps
are made with Natural Earth imagery, no permission needed
(Natural Earth, 2022).</p></caption>
          <?xmltex \igopts{width=426.791339pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/16/1553/2023/gmd-16-1553-2023-f06.png"/>

        </fig>

      <p id="d1e1669">The model correlation in the temporal test generally rises as soil moisture
goes up, until reaching the wettest regime (0.48–0.6), where its variability
increases (Fig. 6Ih). The sites in the middle range tend to have
continuity in soil moisture and regular rainfall patterns, which are most
ideal for LSTM. There is a clear rising trend in <inline-formula><mml:math id="M81" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula> for the temporal test
from dry to wet sites. The driest sites may be difficult to predict due to
scarce but sudden rainfall events that quickly dry out, which reduces the
usefulness of LSTM's memory capability. When we plotted the spatial test <inline-formula><mml:math id="M82" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula>
(Fig. 6IIh), the pattern is similar but less pronounced, which suggests
the driest sites are also more impacted by temporal non-stationarity than
spatial heterogeneity because they have seen limited storm events. Toward
the wettest regime, saturation often occurs, and soil moisture may be
influenced by groundwater or flooding processes, which are difficult to
account for.</p>
      <p id="d1e1687">Interestingly, aspect has a nonlinear effect that varies in different soil
moisture regimes (Fig. 6Ii) due to its impact on shading and solar
insolation. It is well known that aspect can have a predominant control on
soil moisture and plants for dry sites, as witnessed by different vegetation
densities and species and microbial communities on south-facing and
north-facing slopes
(Armesto
and Martnez, 1978; Bennie et<?pagebreak page1562?> al., 2006; Xue et al., 2018). For the very dry
sites (average SMAP <inline-formula><mml:math id="M83" display="inline"><mml:mrow><mml:mo>&lt;</mml:mo><mml:mn mathvariant="normal">0.08</mml:mn></mml:mrow></mml:math></inline-formula>), only those with mid-range aspects tended
to have a decent correlation. The temporal test <inline-formula><mml:math id="M84" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula> (Fig. 6Ii) had a
larger response to aspect than the spatial test <inline-formula><mml:math id="M85" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula> (Fig. 6IIi), which
suggests this difficulty is not a result of too few training sites in space
but a result of highly complex and nonstationary temporal trends in this
combined range of average soil moisture and aspect. The north-facing dry
slopes have a lower <inline-formula><mml:math id="M86" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula>, perhaps because of complex vegetation–soil moisture
interactions in this regime, which may shift from year to year. The most
south-facing dry slopes also have low <inline-formula><mml:math id="M87" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula>, perhaps because they approach the
lower limit of soil moisture and can see large changes due to individual
storm events. In contrast, for the wetter soil regimes, the role of
aspect is reduced; we see noticeably reduced <inline-formula><mml:math id="M88" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula> only for the most
south-facing slope (Fig. 6IIi). This reduced impact may be because soil
moisture is no longer such a strong selector of vegetation species on these
slopes, and thus the distinction of aspect becomes less important.</p>
      <p id="d1e1736">In the vast parts of Africa or Asia where soil moisture predictions are
required but not well supported by in situ measurements, the analysis above
can help us to anticipate challenges. At the hillslope scale, our
predictions may have a larger error for those north-facing slopes in the dry
regime and also straight south-facing slopes for the Northern Hemisphere (to
be reversed for the Southern Hemisphere). The results highlight the
importance of aspect controls on soil moisture and suggest that future
models will need to represent its effect well before they can be accurate.</p>
</sec>
<sec id="Ch1.S3.SS5">
  <label>3.5</label><title>Further discussion</title>
      <p id="d1e1748">Our correlation is modestly higher than the previous state-of-the-art model,
the well-calibrated conceptual hydrologic model, HBV. Even though that model
does not simulate the physical quantity of soil moisture, it could be
modified to have a module that does. However, to obtain suitable parameters
on the global scale and improve the physical processes, we think adding
differentiable programming to the model will give it the adaptive capability
to learn from big data
(Feng
et al., 2022; Shen et al., 2023; Aboelyazeed et al., 2022; Bindas et al.,
2022). It is possible that such a model may generalize better than LSTM over
long distances due to the imposed physical constraints.</p>
      <?pagebreak page1564?><p id="d1e1751">Typically, for many hydrologic applications
(Fang
et al., 2022; Feng et al., 2021; Liu et al., 2022a; Rahmani et al., 2021a), a
spatial test is a tougher test than a temporal test for fully data-driven
models, showing the strong impacts of spatial heterogeneity. This could
either mean the inputs of the model do not completely describe the problem
or that there are not enough sites in space with different combinations of input
attributes for the model to fully resolve their impacts. Typically, spatial
error can be gradually reduced if there are more training sites in space.
However, in both Alaska (Fig. 1) and north-facing dry slopes around the
world (Fig. 6IIi), temporal errors have exceeded spatial errors.
Consistently, when we ran K-fold experiments with a higher K, it also did
not result in noticeably different performances for the models (data not
shown). These observations not only highlight the unique challenges of these
places (rapid climate-driven changes and strong nonstationarity) but also
suggest that the number of training sites is not a predominant issue for
limiting the accuracy of soil moisture predictions.</p>
      <p id="d1e1754">While our product did
not surpass SMAP by a large margin under the most stringent test (cross-continental), it is suitable as a long-term simulation
tool, as it does not require near-real-time observations. Thus it can be
used to assess future climate change impacts. It is also easy to further
expand LSTM networks to enable “data integration” or “data
assimilation”, which absorbs information from recent observations to
improve future forecasts
(Fang and
Shen, 2020; Feng et al., 2020). Satellite observations could also be
employed as the recent observations, as they could help to update LSTM's
hidden states. Such assimilation typically results in a significant boost in
performance and the elimination of bias. Compared to data assimilation with
traditional models, we could skip the bias correction procedure, as LSTM
models tend to have little bias and will adaptively learn to remove the bias
by themselves. Data assimilation only has short-term impacts, however, and
the value of the information content of the data will eventually wane as the
simulation proceeds.</p>
      <p id="d1e1757">The LSTM-based SMAP modeling product is already deployed at scale via the
operational agricultural advice application of PlantVillage, a nonprofit
organization based at Penn State. We intend to put the multitask model into
production alongside alternative estimates. This service is provided
free of charge to farmers and extension services in Africa through the USAID
Current and Emerging Threats to Crops Innovation Lab (CETC IL). PlantVillage
currently scales out precipitation data to 13 million farmers per week in Kenya
and Burkina Faso and believes the ability to complement this with more
accurate information on soil moisture will be of large assistance to farmers
coping with droughts and erratic weather as a result of climate change. It
is also valuable to help farmers optimize fertilizer application rates,
which has become even more critical due to the massive increase in
fertilizer prices over the last 12 months.</p>
</sec>
</sec>
<sec id="Ch1.S4" sec-type="conclusions">
  <label>4</label><title>Conclusions</title>
      <p id="d1e1769">When evaluated against sparse in situ soil moisture networks, the multitask
LSTM model outperformed currently available satellite-based products, land
surface models, and an alternative DL model across most continents. Judging
by the 5-fold spatial test model results, the model had not only
dramatically lower bias but also the highest correlation with in situ soil
moisture networks. Learning from multiple data sources, the model can be
deployed at large scales at a small computational cost and can be expanded
to incorporate data assimilation capabilities. These features make it a
suitable operational tool to democratize access to information for
agriculture in developing regions. While we wish for more measurements in
Africa for model training and validation, the results are at least
encouraging. The model can utilize satellite-estimated soil moisture as one
of the learning targets while also learning from in situ data, and thus it is
well poised to provide higher-resolution outputs than the satellite-based
products.</p>
      <p id="d1e1772">The LSTM model served as a probe for process complexity and showed that mean
soil moisture and aspect have important controls on soil moisture
predictability, while also showing that Arctic regions are inherently more difficult to predict due to
rapid soil changes. For the dry slopes (average SMAP soil moisture <inline-formula><mml:math id="M89" display="inline"><mml:mrow><mml:mo>&lt;</mml:mo><mml:mn mathvariant="normal">0.08</mml:mn></mml:mrow></mml:math></inline-formula>) that face north, there could be complicated vegetation–soil moisture
interactions that are difficult to predict. For the wetter slopes, the role
of aspect becomes less prominent. Error analysis suggests that in these
difficult regions, temporal errors can outweigh spatial errors, and thus having
longer data records and monitoring most recent changes can be more important
than adding more sites.</p>
      <p id="d1e1785">The multitask LSTM model can generalize well in highly data-sparse
regions. Even in the worst-case scenario (no training data on a whole
continent), the model was able to surpass SMAP's accuracy on most
continents. It did seem to have some trouble generalizing to Alaska, where
the soil dynamics are much different from other regions and are also
experiencing rapid changes. However, it provided decent performance when
tested in data-sparse continents where it has not been trained, like Africa
and Australia, showing that these predictions can be beneficial for such
regions where there are not a lot of published soil moisture datasets. This
modeling success is partially due to the strong ability of the model to
generalize but also because the soils in the known sites in Africa are
similar to those in the training set. It is fortunate that the more
intensively instrumented CONUS and Europe already contain a wide variety of
soils and climates for training, without which the model would suffer
greatly.</p>
</sec>

      
      </body>
    <back><notes notes-type="codedataavailability"><title>Code and data availability</title>

      <p id="d1e1793">The multitask LSTM code and GSM3 soil moisture dataset can be downloaded at
<ext-link xlink:href="https://doi.org/10.5281/zenodo.7026036" ext-link-type="DOI">10.5281/zenodo.7026036</ext-link> (Liu et al., 2022b). Links to data sources
have been provided in the Methods section.</p>
  </notes><app-group>
        <supplementary-material position="anchor"><p id="d1e1799">The supplement related to this article is available online at: <inline-supplementary-material xlink:href="https://doi.org/10.5194/gmd-16-1553-2023-supplement" xlink:title="pdf">https://doi.org/10.5194/gmd-16-1553-2023-supplement</inline-supplementary-material>.</p></supplementary-material>
        </app-group><notes notes-type="authorcontribution"><title>Author contributions</title>

      <p id="d1e1808">CS conceived the study. JL ran the experiments and wrote an early draft.
JL, DH, FR, KL, and CS edited the manuscript.</p>
  </notes><notes notes-type="competinginterests"><title>Competing interests</title>

      <p id="d1e1814">Chaopeng Shen and Kathryn Lawson have financial interests in HydroSapient, Inc., a company that could potentially benefit from the results of this research. This interest has been reviewed by the university in accordance with its individual conflict of interest policy for the purpose of maintaining the objectivity and integrity of research at The Pennsylvania State University.</p>
  </notes><notes notes-type="disclaimer"><title>Disclaimer</title>

      <p id="d1e1820">Publisher’s note: Copernicus Publications remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.</p>
  </notes><ack><title>Acknowledgements</title><p id="d1e1826">We thank the two reviewers and the editor for their constructive comments. We also thank the dataset providers for their effort in making data publicly accessible. Data access is a critical step toward enabling research on the global scale.</p></ack><notes notes-type="financialsupport"><title>Financial support</title>

      <p id="d1e1831">This work was supported by Google.org's AI
Impacts Challenge Grant 1904-57775 and Gates Foundation award INV-018429. Chaopeng Shen was partially supported by National Science Foundation Award EAR-2221880 and OAC no. 1940190. Computation was partially supported by National Science Foundation Major Research Instrumentation Award PHY-2018280.</p>
  </notes><notes notes-type="reviewstatement"><title>Review statement</title>

      <p id="d1e1837">This paper was edited by Christoph Müller and reviewed by Richard Mills and one anonymous referee.</p>
  </notes><ref-list>
    <title>References</title>

      <ref id="bib1.bib1"><label>1</label><?label 1?><mixed-citation>Aboelyazeed, D., Xu, C., Hoffman, F. M., Jones, A. W., Rackauckas, C., Lawson, K. E., and Shen, C.: A differentiable ecosystem modeling framework for large-scale inverse problems: demonstration with photosynthesis simulations, Biogeosciences Discuss. [preprint], <ext-link xlink:href="https://doi.org/10.5194/bg-2022-211" ext-link-type="DOI">10.5194/bg-2022-211</ext-link>, in review, 2022.</mixed-citation></ref>
      <ref id="bib1.bib2"><label>2</label><?label 1?><mixed-citation>Al Bitar, A., Mialon, A., Kerr, Y. H., Cabot, F., Richaume, P., Jacquette, E., Quesney, A., Mahmoodi, A., Tarot, S., Parrens, M., Al-Yaari, A., Pellarin, T., Rodriguez-Fernandez, N., and Wigneron, J.-P.: The global SMOS Level 3 daily soil moisture and brightness temperature maps, Earth Syst. Sci. Data, 9, 293–315, <ext-link xlink:href="https://doi.org/10.5194/essd-9-293-2017" ext-link-type="DOI">10.5194/essd-9-293-2017</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bib3"><label>3</label><?label 1?><mixed-citation>Albergel, C., Dutra, E., Munier, S., Calvet, J.-C., Munoz-Sabater, J., de Rosnay, P., and Balsamo, G.: ERA-5 and ERA-Interim driven ISBA land surface model simulations: which one performs better?, Hydrol. Earth Syst. Sci., 22, 3515–3532, <ext-link xlink:href="https://doi.org/10.5194/hess-22-3515-2018" ext-link-type="DOI">10.5194/hess-22-3515-2018</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bib4"><label>4</label><?label 1?><mixed-citation>Al-Yaari, A., Wigneron, J.-P., Kerr, Y., Rodriguez-Fernandez, N., O'Neill,
P. E., Jackson, T. J., De Lannoy, G. J. M., Al Bitar, A., Mialon, A.,
Richaume, P., Walker, J. P., Mahmoodi, A., and Yueh, S.: Evaluating soil
moisture retrievals from ESA's SMOS and NASA's SMAP brightness temperature
datasets, Remote Sens. Environ., 193, 257–273,
<ext-link xlink:href="https://doi.org/10.1016/j.rse.2017.03.010" ext-link-type="DOI">10.1016/j.rse.2017.03.010</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bib5"><label>5</label><?label 1?><mixed-citation>Amatulli, G., Domisch, S., Tuanmu, M.-N., Parmentier, B., Ranipeta, A.,
Malczyk, J., and Jetz, W.: A suite of global, cross-scale topographic
variables for environmental and biodiversity modeling, Sci. Data, 5, 180040,
<ext-link xlink:href="https://doi.org/10.1038/sdata.2018.40" ext-link-type="DOI">10.1038/sdata.2018.40</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bib6"><label>6</label><?label 1?><mixed-citation>Armesto, J. J. and Martnez, J. A.: Relations between vegetation structure
and slope aspect in the Mediterranean region of Chile, J. Ecol., 66,
881–889, <ext-link xlink:href="https://doi.org/10.2307/2259301" ext-link-type="DOI">10.2307/2259301</ext-link>, 1978.</mixed-citation></ref>
      <ref id="bib1.bib7"><label>7</label><?label 1?><mixed-citation>Baraniuk, C.: Locust Swarms Are Getting So Big That We Need Radar to Track
Them, Medium, <ext-link xlink:href="https://onezero.medium.com/locust-swarms-are-getting-so-big-that-we-need-radar-to-track-them-dc79c06496a0">https://onezero.medium.com/locust-swarms-are-getting-so-big-that-we-need-radar-to-track-them-dc79c06496a0</ext-link> (last access: 1 August 2022),  2020.</mixed-citation></ref>
      <ref id="bib1.bib8"><label>8</label><?label 1?><mixed-citation>Beaudoing, H. and Rodell, M.: GLDAS Noah Land Surface Model
L4 3 hourly 0.25 x 0.25 degree V2.0 (GLDAS_NOAH025_3H 2.0),   Greenbelt, Maryland, USA, Goddard Earth Sciences Data and Information Services Center (GES DISC) [data set], <ext-link xlink:href="https://doi.org/10.5067/342OHQM9AK6Q" ext-link-type="DOI">10.5067/342OHQM9AK6Q</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bib9"><label>9</label><?label 1?><mixed-citation>Beck, H. E., Wood, E. F., Pan, M., Fisher, C. K., Miralles, D. G., Dijk, A.
I. J. M. van, McVicar, T. R., and Adler, R. F.: MSWEP V2 Global 3-Hourly
0.1<inline-formula><mml:math id="M90" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> Precipitation: Methodology and Quantitative Assessment, B.
Am. Meteorol. Soc., 100, 473–500, <ext-link xlink:href="https://doi.org/10.1175/BAMS-D-17-0138.1" ext-link-type="DOI">10.1175/BAMS-D-17-0138.1</ext-link>,
2019.</mixed-citation></ref>
      <ref id="bib1.bib10"><label>10</label><?label 1?><mixed-citation>Beck, H. E., Pan, M., Miralles, D. G., Reichle, R. H., Dorigo, W. A., Hahn, S., Sheffield, J., Karthikeyan, L., Balsamo, G., Parinussa, R. M., van Dijk, A. I. J. M., Du, J., Kimball, J. S., Vergopolan, N., and Wood, E. F.: Evaluation of 18 satellite- and model-based soil moisture products using in situ measurements from 826 sensors, Hydrol. Earth Syst. Sci., 25, 17–40, <ext-link xlink:href="https://doi.org/10.5194/hess-25-17-2021" ext-link-type="DOI">10.5194/hess-25-17-2021</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bib11"><label>11</label><?label 1?><mixed-citation>Bennie, J., Hill, M. O., Baxter, R., and Huntley, B.: Influence of slope and
aspect on long-term vegetation change in British chalk grasslands, J. Ecol.,
94, 355–368, <ext-link xlink:href="https://doi.org/10.1111/j.1365-2745.2006.01104.x" ext-link-type="DOI">10.1111/j.1365-2745.2006.01104.x</ext-link>, 2006.</mixed-citation></ref>
      <ref id="bib1.bib12"><label>12</label><?label 1?><mixed-citation>Bentley, A. R., Donovan, J., Sonder, K., Baudron, F., Lewis, J. M., Voss,
R., Rutsaert, P., Poole, N., Kamoun, S., Saunders, D. G. O., Hodson, D.,
Hughes, D. P., Negra, C., Ibba, M. I., Snapp, S., Sida, T. S., Jaleta, M.,
Tesfaye, K., Becker-Reshef, I., and Govaerts, B.: Near- to long-term
measures to stabilize global wheat supplies and food security, Nat. Food, 3,
483–486, <ext-link xlink:href="https://doi.org/10.1038/s43016-022-00559-y" ext-link-type="DOI">10.1038/s43016-022-00559-y</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bib13"><label>13</label><?label 1?><mixed-citation>Bindas, T., Tsai, W.-P., Liu, J., Rahmani, F., Feng, D., Bian, Y., Lawson,
K., and Shen, C.: Improving large-basin streamflow simulation using a
modular, differentiable, learnable graph model for routing,  ESS Open Archive [preprint],
<ext-link xlink:href="https://doi.org/10.1002/essoar.10512512.1" ext-link-type="DOI">10.1002/essoar.10512512.1</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bib14"><label>14</label><?label 1?><mixed-citation>de Jeu, R. and Owe, M.: AMSR2/GCOM-W1 surface soil moisture (LPRM)
L3 1 day 10 km x 10 km ascending V001, Goddard Earth Sciences Data and Information Services Center (GES DISC) (Bill Teng), Greenbelt, MD, USA, Goddard Earth Sciences Data and Information Services Center (GES DISC) [data set], <ext-link xlink:href="https://doi.org/10.5067/B0GHODHJLDA8" ext-link-type="DOI">10.5067/B0GHODHJLDA8</ext-link>,
2013.</mixed-citation></ref>
      <ref id="bib1.bib15"><label>15</label><?label 1?><mixed-citation>Didan, K.: MOD13C2: MODIS/Terra Vegetation Indices Monthly L3 Global 0.05Deg
CMG version 6,   NASA EOSDIS Land Processes DAAC [data set], <ext-link xlink:href="https://doi.org/10.5067/MODIS/MOD13C2.006" ext-link-type="DOI">10.5067/MODIS/MOD13C2.006</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bib16"><label>16</label><?label 1?><mixed-citation>Dorigo, W. A., Wagner, W., Hohensinn, R., Hahn, S., Paulik, C., Xaver, A., Gruber, A., Drusch, M., Mecklenburg, S., van Oevelen, P., Robock, A., and Jackson, T.: The International Soil Moisture Network: a data hosting facility for global in situ soil moisture measurements, Hydrol. Earth Syst. Sci., 15, 1675–1698, <ext-link xlink:href="https://doi.org/10.5194/hess-15-1675-2011" ext-link-type="DOI">10.5194/hess-15-1675-2011</ext-link>, 2011.</mixed-citation></ref>
      <ref id="bib1.bib17"><label>17</label><?label 1?><mixed-citation>Dorigo, W. A., Xaver, A., Vreugdenhil, M., Gruber, A., Hegyiová, A.,
Sanchis-Dufau, A. D., Zamojski, D., Cordes, C., Wagner, W., and Drusch, M.:
Global automated quality control of in situ soil moisture data from the
international soil moisture network, Vadose Zone J., 12, vzj2012.0097,
<ext-link xlink:href="https://doi.org/10.2136/vzj2012.0097" ext-link-type="DOI">10.2136/vzj2012.0097</ext-link>, 2013.</mixed-citation></ref>
      <?pagebreak page1566?><ref id="bib1.bib18"><label>18</label><?label 1?><mixed-citation>Ellenburg, W. L., Mishra, V., Roberts, J. B., Limaye, A. S., Case, J. L.,
Blankenship, C. B., and Cressman, K.: Detecting desert locust breeding
grounds: A satellite-assisted modeling approach, Remote Sens., 13, 1276,
<ext-link xlink:href="https://doi.org/10.3390/rs13071276" ext-link-type="DOI">10.3390/rs13071276</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bib19"><label>19</label><?label 1?><mixed-citation>Entekhabi, D.: The Soil Moisture Active Passive (SMAP) mission, P. IEEE,
98, 704–716, <ext-link xlink:href="https://doi.org/10/bz3xhb" ext-link-type="DOI">10/bz3xhb</ext-link>, 2010.</mixed-citation></ref>
      <ref id="bib1.bib20"><label>20</label><?label 1?><mixed-citation>ESA: Land Cover CCI Product User Guide Version 2, <uri>http://maps.elie.ucl.ac.be/CCI/viewer/download/ESACCI-LC-Ph2-PUGv2_2.0.pdf</uri> (last access: 1 August 2022), 2017.</mixed-citation></ref>
      <ref id="bib1.bib21"><label>21</label><?label 1?><mixed-citation>Fang, K. and Shen, C.: Near-real-time forecast of satellite-based soil
moisture using long short-term memory with an adaptive data integration
kernel, J. Hydrometeorol., 21, 399–413,
<ext-link xlink:href="https://doi.org/10.1175/jhm-d-19-0169.1" ext-link-type="DOI">10.1175/jhm-d-19-0169.1</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bib22"><label>22</label><?label 1?><mixed-citation>Fang, K., Shen, C., Kifer, D., and Yang, X.: Prolongation of SMAP to
spatiotemporally seamless coverage of continental U.S. using a deep learning
neural network, Geophys. Res. Lett., 44, 11030–11039,
<ext-link xlink:href="https://doi.org/10.1002/2017gl075619" ext-link-type="DOI">10.1002/2017gl075619</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bib23"><label>23</label><?label 1?><mixed-citation>Fang, K., Pan, M., and Shen, C.: The value of SMAP for long-term soil
moisture estimation with the help of deep learning, IEEE T. Geosci.
Remote, 57, 2221–2233, <ext-link xlink:href="https://doi.org/10/gghp3v" ext-link-type="DOI">10/gghp3v</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bib24"><label>24</label><?label 1?><mixed-citation>Fang, K., Kifer, D., Lawson, K., Feng, D., and Shen, C.: The data synergy
effects of time-series deep learning models in hydrology, Water Resour.
Res., 58, e2021WR029583, <ext-link xlink:href="https://doi.org/10.1029/2021WR029583" ext-link-type="DOI">10.1029/2021WR029583</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bib25"><label>25</label><?label 1?><mixed-citation>FAO, IIASA, ISRIC, ISSCAS, and JRC: Harmonized World Soil Database (version
1.2), FAO IIASA, ISRIC, ISSCAS, and JRC [data set], <uri>http://www.fao.org/soils-portal/data-hub/soil-maps-and-databases/harmonized-world-soil-database-v12/en/</uri> (last access: 1 August 2022),  2012.</mixed-citation></ref>
      <ref id="bib1.bib26"><label>26</label><?label 1?><mixed-citation>Feng, D., Fang, K., and Shen, C.: Enhancing streamflow forecast and
extracting insights using long-short term memory networks with data
integration at continental scales, Water Resour. Res., 56, e2019WR026793,
<ext-link xlink:href="https://doi.org/10.1029/2019WR026793" ext-link-type="DOI">10.1029/2019WR026793</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bib27"><label>27</label><?label 1?><mixed-citation>Feng, D., Lawson, K., and Shen, C.: Mitigating prediction error of deep
learning streamflow models in large data-sparse regions with ensemble
modeling and soft data, Geophys. Res. Lett., 48, e2021GL092999,
<ext-link xlink:href="https://doi.org/10.1029/2021GL092999" ext-link-type="DOI">10.1029/2021GL092999</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bib28"><label>28</label><?label 1?><mixed-citation>Feng, D., Liu, J., Lawson, K., and Shen, C.: Differentiable, learnable,
regionalized process-based models with multiphysical outputs can approach
state-of-the-art hydrologic prediction accuracy, Water Resour. Res., 58,
e2022WR032404, <ext-link xlink:href="https://doi.org/10.1029/2022WR032404" ext-link-type="DOI">10.1029/2022WR032404</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bib29"><label>29</label><?label 1?><mixed-citation>
Fischer, G., Nachtergaele, F., Prieler, S., van Velthuizen, H. T., Verelst,
L., and Wiberg, D.: Global Agro-Ecological Zones Assessment for Agriculture
(GAEZ 2008), IIASA Laxenburg, Austria and FAO, Rome, Italy, 2008.</mixed-citation></ref>
      <ref id="bib1.bib30"><label>30</label><?label 1?><mixed-citation>
Friedman, J. H.: Greedy Function Approximation: A Gradient Boosting Machine,
Ann. Stat., 29, 1189–1232, 2001.</mixed-citation></ref>
      <ref id="bib1.bib31"><label>31</label><?label 1?><mixed-citation>Gauch, M., Mai, J., and Lin, J.: The Proper Care and Feeding of CAMELS: How
Limited Training Data Affects Streamflow Prediction,  Environ. Modell. Softw., 135, 104926, <ext-link xlink:href="https://doi.org/10.1016/j.envsoft.2020.104926" ext-link-type="DOI">10.1016/j.envsoft.2020.104926</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bib32"><label>32</label><?label 1?><mixed-citation>Hochreiter, S. and Schmidhuber, J.: Long Short-Term Memory, Neural. Comput.,
9, 1735–1780, <ext-link xlink:href="https://doi.org/10.1162/neco.1997.9.8.1735" ext-link-type="DOI">10.1162/neco.1997.9.8.1735</ext-link>, 1997.</mixed-citation></ref>
      <ref id="bib1.bib33"><label>33</label><?label 1?><mixed-citation>Hrachowitz, M., Savenije, H. H. G., Blöschl, G., McDonnell, J. J.,
Sivapalan, M., Pomeroy, J. W., Arheimer, B., Blume, T., Clark, M. P., Ehret,
U., Fenicia, F., Freer, J. E., Gelfan, A., Gupta, H. V., Hughes, D. A., Hut,
R. W., Montanari, A., Pande, S., Tetzlaff, D., Troch, P. A., Uhlenbrook, S.,
Wagener, T., Winsemius, H. C., Woods, R. A., Zehe, E., and Cudennec, C.: A
decade of Predictions in Ungauged Basins (PUB)–a review, Hydrol. Sci. J.,
58, 1198–1255, <ext-link xlink:href="https://doi.org/10/gfsq5q" ext-link-type="DOI">10/gfsq5q</ext-link>, 2013.</mixed-citation></ref>
      <ref id="bib1.bib34"><label>34</label><?label 1?><mixed-citation>Huffman, G. J., Stocker, E. F., Bolvin, D. T., Nelkin, E. J., and Tan, J.:
GPM IMERG Final Precipitation L3 1 day 0.1 degree x 0.1 degree V06
(GPM_3IMERGDF 06), edited by: Andrey Savtchenko, Greenbelt, MD, Goddard Earth Sciences Data and Information Services Center (GES DISC) [data set],
<ext-link xlink:href="https://doi.org/10.5067/GPM/IMERGDF/DAY/06" ext-link-type="DOI">10.5067/GPM/IMERGDF/DAY/06</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bib35"><label>35</label><?label 1?><mixed-citation>Hunter-Jones, P.: Egg development in the Desert Locust (Schistocerca
gregaria Forsk.) in relation to the availability of water, Proc. R. Entomol.
Soc. A, 39, 25–33,
<ext-link xlink:href="https://doi.org/10.1111/j.1365-3032.1964.tb00781.x" ext-link-type="DOI">10.1111/j.1365-3032.1964.tb00781.x</ext-link>, 1964.</mixed-citation></ref>
      <ref id="bib1.bib36"><label>36</label><?label 1?><mixed-citation>Kerr, Y. H., Waldteufel, P., Wigneron, J.-P., Delwart, S., Cabot, F.,
Boutin, J., Escorihuela, M.-J., Font, J., Reul, N., Gruhier, C., Juglea, S.
E., Drinkwater, M. R., Hahne, A., Martín-Neira, M., and Mecklenburg,
S.: The SMOS mission: New tool for monitoring key elements of the global
water cycle, P. IEEE, 98, 666–687, <ext-link xlink:href="https://doi.org/10/b9szx6" ext-link-type="DOI">10/b9szx6</ext-link>, 2010.</mixed-citation></ref>
      <ref id="bib1.bib37"><label>37</label><?label 1?><mixed-citation>Kratzert, F., Klotz, D., Shalev, G., Klambauer, G., Hochreiter, S., and Nearing, G.: Towards learning universal, regional, and local hydrological behaviors via machine learning applied to large-sample datasets, Hydrol. Earth Syst. Sci., 23, 5089–5110, <ext-link xlink:href="https://doi.org/10.5194/hess-23-5089-2019" ext-link-type="DOI">10.5194/hess-23-5089-2019</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bib38"><label>38</label><?label 1?><mixed-citation>Liu, J., Rahmani, F., Lawson, K., and Shen, C.: A multiscale deep learning
model for soil moisture integrating satellite and in situ data, Geophys.
Res. Lett., 49, e2021GL096847, <ext-link xlink:href="https://doi.org/10.1029/2021GL096847" ext-link-type="DOI">10.1029/2021GL096847</ext-link>, 2022a.</mixed-citation></ref>
      <ref id="bib1.bib39"><label>39</label><?label 1?><mixed-citation>Liu, J., Hughes, D., Rahmani, F., Lawson, K., and Shen, C.: Global Soil Moisture Dataset From a Multitask Model (GSM3), Zenodo [data set], <ext-link xlink:href="https://doi.org/10.5281/zenodo.7344484" ext-link-type="DOI">10.5281/zenodo.7344484</ext-link>, 2022b.</mixed-citation></ref>
      <ref id="bib1.bib40"><label>40</label><?label 1?><mixed-citation>Meyal, A. Y., Versteeg, R., Alper, E., Johnson, D., Rodzianko, A., Franklin,
M., and Wainwright, H.: Automated cloud based long short-term memory neural
network based SWE prediction, Front. Water, 2, 1–12,
<ext-link xlink:href="https://doi.org/10.3389/frwa.2020.574917" ext-link-type="DOI">10.3389/frwa.2020.574917</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bib41"><label>41</label><?label 1?><mixed-citation>Muñoz Sabater, J.: ERA5-Land hourly data from 1950 to present, Copernicus Climate Change Service (C3S) Climate Data Store (CDS) [data set], <ext-link xlink:href="https://doi.org/10.24381/cds.e2161bac" ext-link-type="DOI">10.24381/cds.e2161bac</ext-link>,   2019.</mixed-citation></ref>
      <ref id="bib1.bib42"><label>42</label><?label 1?><mixed-citation>Narasimhan, B. and Srinivasan, R.: Development and evaluation of Soil
Moisture Deficit Index (SMDI) and Evapotranspiration Deficit Index (ETDI)
for agricultural drought monitoring, Agr. Forest Meteorol., 133, 69–88,
<ext-link xlink:href="https://doi.org/10/fq9wdv" ext-link-type="DOI">10/fq9wdv</ext-link>, 2005.</mixed-citation></ref>
      <ref id="bib1.bib43"><label>43</label><?label 1?><mixed-citation>Natural Earth: Free vector and raster map data,  Natural Earth [data set],  <uri>https://www.naturalearthdata.com/</uri>, last access: 1 August 2022.</mixed-citation></ref>
      <ref id="bib1.bib44"><label>44</label><?label 1?><mixed-citation>Norbiato, D., Borga, M., Degli Esposti, S., Gaume, E., and Anquetin, S.:
Flash flood warning based on rainfall thresholds and soil moisture
conditions: An assessment for gauged and ungauged basins, J. Hydrol., 362,
274–290, <ext-link xlink:href="https://doi.org/10/dtw4hm" ext-link-type="DOI">10/dtw4hm</ext-link>, 2008.</mixed-citation></ref>
      <ref id="bib1.bib45"><label>45</label><?label 1?><mixed-citation>Nuwer, R.: As Locusts Swarmed East Africa, This Tech Helped Squash Them, N.
Y. Times, 8th April,  <uri>https://www.nytimes.com/2021/04/08/science/locust-swarms-africa.html</uri> (last access: 1 August 2022), 2021.</mixed-citation></ref>
      <?pagebreak page1567?><ref id="bib1.bib46"><label>46</label><?label 1?><mixed-citation>O, S. and Orth, R.: Global soil moisture data derived through machine
learning trained with in-situ measurements, Sci. Data, 8, 170,
<ext-link xlink:href="https://doi.org/10.1038/s41597-021-00964-1" ext-link-type="DOI">10.1038/s41597-021-00964-1</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bib47"><label>47</label><?label 1?><mixed-citation>O'Neill, P. E., Chan, S., Njoku, E. G., Jackson, T., Bindlish, R., Chaubell,
J., and Colliander, A.: SMAP Enhanced L3 Radiometer Global and Polar Grid
Daily 9 km EASE-Grid Soil Moisture, Version 5 (SPL3SMP_E),  Boulder, Colorado USA, NASA National Snow and Ice Data Center Distributed Active Archive Center [data set],
<ext-link xlink:href="https://doi.org/10.5067/4DQ54OUIJ9DL" ext-link-type="DOI">10.5067/4DQ54OUIJ9DL</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bib48"><label>48</label><?label 1?><mixed-citation>Owe, M., de Jeu, R., and Holmes, T.: Multisensor historical climatology of
satellite-derived global land surface moisture, J. Geophys. Res., 113,
1–17, <ext-link xlink:href="https://doi.org/10.1029/2007JF000769" ext-link-type="DOI">10.1029/2007JF000769</ext-link>, 2008.</mixed-citation></ref>
      <ref id="bib1.bib49"><label>49</label><?label 1?><mixed-citation>
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel,
O., Blondel, M., Prettenhofer, P., Weiss, R., and Dubourg, V.: Scikit-learn:
Machine learning in Python, J. Mach. Learn. Res., 12, 2825–2830, 2011.</mixed-citation></ref>
      <ref id="bib1.bib50"><label>50</label><?label 1?><mixed-citation>Rahmani, F., Shen, C., Oliver, S., Lawson, K., and Appling, A.: Deep
learning approaches for improving prediction of daily stream temperature in
data-scarce, unmonitored, and dammed basins, Hydrol. Process., 35, e14400,
<ext-link xlink:href="https://doi.org/10.1002/hyp.14400" ext-link-type="DOI">10.1002/hyp.14400</ext-link>, 2021a.</mixed-citation></ref>
      <ref id="bib1.bib51"><label>51</label><?label 1?><mixed-citation>Rahmani, F., Lawson, K., Ouyang, W., Appling, A., Oliver, S., and Shen, C.:
Exploring the exceptional performance of a deep learning stream temperature
model and the value of streamflow data, Environ. Res. Lett., 16, 024025,
<ext-link xlink:href="https://doi.org/10.1088/1748-9326/abd501" ext-link-type="DOI">10.1088/1748-9326/abd501</ext-link>, 2021b.</mixed-citation></ref>
      <ref id="bib1.bib52"><label>52</label><?label 1?><mixed-citation>Rodell, M., Houser, P. R., Jambor, U., Gottschalck, J., Mitchell, K., Meng,
C.-J., Arsenault, K., Cosgrove, B., Radakovich, J., Bosilovich, M., Entin,
J. K., Walker, J. P., Lohmann, D., and Toll, D.: The Global Land Data
Assimilation System, B. Am. Meteorol. Soc., 85, 381–394,
<ext-link xlink:href="https://doi.org/10.1175/BAMS-85-3-381" ext-link-type="DOI">10.1175/BAMS-85-3-381</ext-link>, 2004.</mixed-citation></ref>
      <ref id="bib1.bib53"><label>53</label><?label 1?><mixed-citation>Schaaf, C. and Wang, Z.: MODIS/Terra<inline-formula><mml:math id="M91" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>Aqua BRDF/Albedo Daily L3
Global – 500m V061,   NASA EOSDIS Land Processes DAAC [data set], <ext-link xlink:href="https://doi.org/10.5067/MODIS/MCD43A3.061" ext-link-type="DOI">10.5067/MODIS/MCD43A3.061</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bib54"><label>54</label><?label 1?><mixed-citation>Sheffield, J. and Wood, E. F.: Global trends and variability in soil
moisture and drought characteristics, 1950–2000, from observation-driven
simulations of the terrestrial hydrologic cycle, J. Climate, 21, 432–458,
<ext-link xlink:href="https://doi.org/10.1175/2007JCLI1822.1" ext-link-type="DOI">10.1175/2007JCLI1822.1</ext-link>, 2008.</mixed-citation></ref>
      <ref id="bib1.bib55"><label>55</label><?label 1?><mixed-citation>Shen, C.: A transdisciplinary review of deep learning research and its
relevance for water resources scientists, Water Resour. Res., 54,
8558–8593, <ext-link xlink:href="https://doi.org/10.1029/2018wr022643" ext-link-type="DOI">10.1029/2018wr022643</ext-link>, 2018.
</mixed-citation></ref><?xmltex \hack{\newpage}?>
      <ref id="bib1.bib56"><label>56</label><?label 1?><mixed-citation>Shen, C., Appling, A. P., Gentine, P., Bandai, T., Gupta, H., Tartakovsky,
A., Baity-Jesi, M., Fenicia, F., Kifer, D., Li, L., Liu, X., Ren, W., Zheng,
Y., Harman, C. J., Clark, M., Farthing, M., Feng, D., Kumar, P.,
Aboelyazeed, D., Rahmani, F., Beck, H. E., Bindas, T., Dwivedi, D., Fang,
K., Höge, M., Rackauckas, C., Roy, T., Xu, C., and Lawson, K.:
Differentiable modeling to unify machine learning and physical models and
advance Geosciences, arXiv [preprint], <ext-link xlink:href="https://doi.org/10.48550/arXiv.2301.04027" ext-link-type="DOI">10.48550/arXiv.2301.04027</ext-link>, 10 January
2023.</mixed-citation></ref>
      <ref id="bib1.bib57"><label>57</label><?label 1?><mixed-citation>Support CATDS: CATDS-PDC L3SM Simple UDP – 1 day soil moisture Simple User
Data Product from SMOS satellite,  CATDS (CNES, IFREMER, CESBIO) [data set]
<ext-link xlink:href="https://doi.org/10.12770/8db7102b-1b22-4db3-949d-e51269417aae" ext-link-type="DOI">10.12770/8db7102b-1b22-4db3-949d-e51269417aae</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bib58"><label>58</label><?label 1?><mixed-citation>Tsai, W.-P., Feng, D., Pan, M., Beck, H., Lawson, K., Yang, Y., Liu, J., and
Shen, C.: From calibration to parameter learning: Harnessing the scaling
effects of big data in geoscientific modeling, Nat. Commun., 12, 5988,
<ext-link xlink:href="https://doi.org/10.1038/s41467-021-26107-z" ext-link-type="DOI">10.1038/s41467-021-26107-z</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bib59"><label>59</label><?label 1?><mixed-citation>UN WFP: Stop locusts in East Africa now or pay much more to help people
later, U. N. UN World Food Programme WFP, 14th February,  <ext-link xlink:href="https://www.wfp.org/news/stop-locusts-east-africa-now-or-pay-much-more-help-people-later-wfp">https://www.wfp.org/news/stop-locusts-east-africa-now-or-pay-much-more-help-people-later-wfp</ext-link>  (last access: 1 August 2022),  2020.</mixed-citation></ref>
      <ref id="bib1.bib60"><label>60</label><?label 1?><mixed-citation>Wan, Z., Hook, S., and Hulley, G.: MODIS/Aqua Land Surface
Temperature/Emissivity Daily L3 Global 1km SIN Grid V061,  NASA EOSDIS Land Processes DAAC [data set],
<ext-link xlink:href="https://doi.org/10.5067/MODIS/MYD11A1.061" ext-link-type="DOI">10.5067/MODIS/MYD11A1.061</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bib61"><label>61</label><?label 1?><mixed-citation>Xue, R., Yang, Q., Miao, F., Wang, X., Shen, Y., Xue, R., Yang, Q., Miao,
F., Wang, X., and Shen, Y.: Slope aspect influences plant biomass, soil
properties and microbial composition in alpine meadow on the Qinghai-Tibetan
plateau, J. Soil Sci. Plant Nut., 18, 1–12,
<ext-link xlink:href="https://doi.org/10.4067/S0718-95162018005000101" ext-link-type="DOI">10.4067/S0718-95162018005000101</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bib62"><label>62</label><?label 1?><mixed-citation>Yang, Z.-L., Niu, G.-Y., Mitchell, K. E., Chen, F., Ek, M. B., Barlage, M.,
Longuevergne, L., Manning, K., Niyogi, D., Tewari, M., and Xia, Y.: The
community Noah land surface model with multiparameterization options
(Noah-MP): 2. Evaluation over global river basins, J. Geophys. Res.-Atmos., 116, D12110, <ext-link xlink:href="https://doi.org/10.1029/2010JD015140" ext-link-type="DOI">10.1029/2010JD015140</ext-link>, 2011.</mixed-citation></ref>
      <ref id="bib1.bib63"><label>63</label><?label 1?><mixed-citation>Zhi, W., Feng, D., Tsai, W.-P., Sterle, G., Harpold, A., Shen, C., and Li,
L.: From hydrometeorology to river water quality: Can a deep learning model
predict dissolved oxygen at the continental scale?, Environ. Sci. Technol.,
55, 2357–2368, <ext-link xlink:href="https://doi.org/10.1021/acs.est.0c06783" ext-link-type="DOI">10.1021/acs.est.0c06783</ext-link>, 2021.</mixed-citation></ref>

  </ref-list></back>
    <!--<article-title-html>Evaluating a global soil moisture dataset from a multitask model (GSM3 v1.0) with potential applications for crop threats</article-title-html>
<abstract-html/>
<ref-html id="bib1.bib1"><label>1</label><mixed-citation>
      
Aboelyazeed, D., Xu, C., Hoffman, F. M., Jones, A. W., Rackauckas, C., Lawson, K. E., and Shen, C.: A differentiable ecosystem modeling framework for large-scale inverse problems: demonstration with photosynthesis simulations, Biogeosciences Discuss. [preprint], <a href="https://doi.org/10.5194/bg-2022-211" target="_blank">https://doi.org/10.5194/bg-2022-211</a>, in review, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib2"><label>2</label><mixed-citation>
      
Al Bitar, A., Mialon, A., Kerr, Y. H., Cabot, F., Richaume, P., Jacquette, E., Quesney, A., Mahmoodi, A., Tarot, S., Parrens, M., Al-Yaari, A., Pellarin, T., Rodriguez-Fernandez, N., and Wigneron, J.-P.: The global SMOS Level 3 daily soil moisture and brightness temperature maps, Earth Syst. Sci. Data, 9, 293–315, <a href="https://doi.org/10.5194/essd-9-293-2017" target="_blank">https://doi.org/10.5194/essd-9-293-2017</a>, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib3"><label>3</label><mixed-citation>
      
Albergel, C., Dutra, E., Munier, S., Calvet, J.-C., Munoz-Sabater, J., de Rosnay, P., and Balsamo, G.: ERA-5 and ERA-Interim driven ISBA land surface model simulations: which one performs better?, Hydrol. Earth Syst. Sci., 22, 3515–3532, <a href="https://doi.org/10.5194/hess-22-3515-2018" target="_blank">https://doi.org/10.5194/hess-22-3515-2018</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib4"><label>4</label><mixed-citation>
      
Al-Yaari, A., Wigneron, J.-P., Kerr, Y., Rodriguez-Fernandez, N., O'Neill,
P. E., Jackson, T. J., De Lannoy, G. J. M., Al Bitar, A., Mialon, A.,
Richaume, P., Walker, J. P., Mahmoodi, A., and Yueh, S.: Evaluating soil
moisture retrievals from ESA's SMOS and NASA's SMAP brightness temperature
datasets, Remote Sens. Environ., 193, 257–273,
<a href="https://doi.org/10.1016/j.rse.2017.03.010" target="_blank">https://doi.org/10.1016/j.rse.2017.03.010</a>, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib5"><label>5</label><mixed-citation>
      
Amatulli, G., Domisch, S., Tuanmu, M.-N., Parmentier, B., Ranipeta, A.,
Malczyk, J., and Jetz, W.: A suite of global, cross-scale topographic
variables for environmental and biodiversity modeling, Sci. Data, 5, 180040,
<a href="https://doi.org/10.1038/sdata.2018.40" target="_blank">https://doi.org/10.1038/sdata.2018.40</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib6"><label>6</label><mixed-citation>
      
Armesto, J. J. and Martnez, J. A.: Relations between vegetation structure
and slope aspect in the Mediterranean region of Chile, J. Ecol., 66,
881–889, <a href="https://doi.org/10.2307/2259301" target="_blank">https://doi.org/10.2307/2259301</a>, 1978.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib7"><label>7</label><mixed-citation>
      
Baraniuk, C.: Locust Swarms Are Getting So Big That We Need Radar to Track
Them, Medium, <a href="https://onezero.medium.com/locust-swarms-are-getting-so-big-that-we-need-radar-to-track-them-dc79c06496a0" target="_blank">https://onezero.medium.com/locust-swarms-are-getting-so-big-that-we-need-radar-to-track-them-dc79c06496a0</a> (last access: 1 August 2022),  2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib8"><label>8</label><mixed-citation>
      
Beaudoing, H. and Rodell, M.: GLDAS Noah Land Surface Model
L4 3 hourly 0.25 x 0.25 degree V2.0 (GLDAS_NOAH025_3H 2.0),   Greenbelt, Maryland, USA, Goddard Earth Sciences Data and Information Services Center (GES DISC) [data set], <a href="https://doi.org/10.5067/342OHQM9AK6Q" target="_blank">https://doi.org/10.5067/342OHQM9AK6Q</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib9"><label>9</label><mixed-citation>
      
Beck, H. E., Wood, E. F., Pan, M., Fisher, C. K., Miralles, D. G., Dijk, A.
I. J. M. van, McVicar, T. R., and Adler, R. F.: MSWEP V2 Global 3-Hourly
0.1° Precipitation: Methodology and Quantitative Assessment, B.
Am. Meteorol. Soc., 100, 473–500, <a href="https://doi.org/10.1175/BAMS-D-17-0138.1" target="_blank">https://doi.org/10.1175/BAMS-D-17-0138.1</a>,
2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib10"><label>10</label><mixed-citation>
      
Beck, H. E., Pan, M., Miralles, D. G., Reichle, R. H., Dorigo, W. A., Hahn, S., Sheffield, J., Karthikeyan, L., Balsamo, G., Parinussa, R. M., van Dijk, A. I. J. M., Du, J., Kimball, J. S., Vergopolan, N., and Wood, E. F.: Evaluation of 18 satellite- and model-based soil moisture products using in situ measurements from 826 sensors, Hydrol. Earth Syst. Sci., 25, 17–40, <a href="https://doi.org/10.5194/hess-25-17-2021" target="_blank">https://doi.org/10.5194/hess-25-17-2021</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib11"><label>11</label><mixed-citation>
      
Bennie, J., Hill, M. O., Baxter, R., and Huntley, B.: Influence of slope and
aspect on long-term vegetation change in British chalk grasslands, J. Ecol.,
94, 355–368, <a href="https://doi.org/10.1111/j.1365-2745.2006.01104.x" target="_blank">https://doi.org/10.1111/j.1365-2745.2006.01104.x</a>, 2006.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib12"><label>12</label><mixed-citation>
      
Bentley, A. R., Donovan, J., Sonder, K., Baudron, F., Lewis, J. M., Voss,
R., Rutsaert, P., Poole, N., Kamoun, S., Saunders, D. G. O., Hodson, D.,
Hughes, D. P., Negra, C., Ibba, M. I., Snapp, S., Sida, T. S., Jaleta, M.,
Tesfaye, K., Becker-Reshef, I., and Govaerts, B.: Near- to long-term
measures to stabilize global wheat supplies and food security, Nat. Food, 3,
483–486, <a href="https://doi.org/10.1038/s43016-022-00559-y" target="_blank">https://doi.org/10.1038/s43016-022-00559-y</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib13"><label>13</label><mixed-citation>
      
Bindas, T., Tsai, W.-P., Liu, J., Rahmani, F., Feng, D., Bian, Y., Lawson,
K., and Shen, C.: Improving large-basin streamflow simulation using a
modular, differentiable, learnable graph model for routing,  ESS Open Archive [preprint],
<a href="https://doi.org/10.1002/essoar.10512512.1" target="_blank">https://doi.org/10.1002/essoar.10512512.1</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib14"><label>14</label><mixed-citation>
      
de Jeu, R. and Owe, M.: AMSR2/GCOM-W1 surface soil moisture (LPRM)
L3 1 day 10&thinsp;km x 10&thinsp;km ascending V001, Goddard Earth Sciences Data and Information Services Center (GES DISC) (Bill Teng), Greenbelt, MD, USA, Goddard Earth Sciences Data and Information Services Center (GES DISC) [data set], <a href="https://doi.org/10.5067/B0GHODHJLDA8" target="_blank">https://doi.org/10.5067/B0GHODHJLDA8</a>,
2013.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib15"><label>15</label><mixed-citation>
      
Didan, K.: MOD13C2: MODIS/Terra Vegetation Indices Monthly L3 Global 0.05Deg
CMG version 6,   NASA EOSDIS Land Processes DAAC [data set], <a href="https://doi.org/10.5067/MODIS/MOD13C2.006" target="_blank">https://doi.org/10.5067/MODIS/MOD13C2.006</a>, 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib16"><label>16</label><mixed-citation>
      
Dorigo, W. A., Wagner, W., Hohensinn, R., Hahn, S., Paulik, C., Xaver, A., Gruber, A., Drusch, M., Mecklenburg, S., van Oevelen, P., Robock, A., and Jackson, T.: The International Soil Moisture Network: a data hosting facility for global in situ soil moisture measurements, Hydrol. Earth Syst. Sci., 15, 1675–1698, <a href="https://doi.org/10.5194/hess-15-1675-2011" target="_blank">https://doi.org/10.5194/hess-15-1675-2011</a>, 2011.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib17"><label>17</label><mixed-citation>
      
Dorigo, W. A., Xaver, A., Vreugdenhil, M., Gruber, A., Hegyiová, A.,
Sanchis-Dufau, A. D., Zamojski, D., Cordes, C., Wagner, W., and Drusch, M.:
Global automated quality control of in situ soil moisture data from the
international soil moisture network, Vadose Zone J., 12, vzj2012.0097,
<a href="https://doi.org/10.2136/vzj2012.0097" target="_blank">https://doi.org/10.2136/vzj2012.0097</a>, 2013.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib18"><label>18</label><mixed-citation>
      
Ellenburg, W. L., Mishra, V., Roberts, J. B., Limaye, A. S., Case, J. L.,
Blankenship, C. B., and Cressman, K.: Detecting desert locust breeding
grounds: A satellite-assisted modeling approach, Remote Sens., 13, 1276,
<a href="https://doi.org/10.3390/rs13071276" target="_blank">https://doi.org/10.3390/rs13071276</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib19"><label>19</label><mixed-citation>
      
Entekhabi, D.: The Soil Moisture Active Passive (SMAP) mission, P. IEEE,
98, 704–716, <a href="https://doi.org/10/bz3xhb" target="_blank">https://doi.org/10/bz3xhb</a>, 2010.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib20"><label>20</label><mixed-citation>
      
ESA: Land Cover CCI Product User Guide Version 2, <a href="http://maps.elie.ucl.ac.be/CCI/viewer/download/ESACCI-LC-Ph2-PUGv2_2.0.pdf" target="_blank"/> (last access: 1 August 2022), 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib21"><label>21</label><mixed-citation>
      
Fang, K. and Shen, C.: Near-real-time forecast of satellite-based soil
moisture using long short-term memory with an adaptive data integration
kernel, J. Hydrometeorol., 21, 399–413,
<a href="https://doi.org/10.1175/jhm-d-19-0169.1" target="_blank">https://doi.org/10.1175/jhm-d-19-0169.1</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib22"><label>22</label><mixed-citation>
      
Fang, K., Shen, C., Kifer, D., and Yang, X.: Prolongation of SMAP to
spatiotemporally seamless coverage of continental U.S. using a deep learning
neural network, Geophys. Res. Lett., 44, 11030–11039,
<a href="https://doi.org/10.1002/2017gl075619" target="_blank">https://doi.org/10.1002/2017gl075619</a>, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib23"><label>23</label><mixed-citation>
      
Fang, K., Pan, M., and Shen, C.: The value of SMAP for long-term soil
moisture estimation with the help of deep learning, IEEE T. Geosci.
Remote, 57, 2221–2233, <a href="https://doi.org/10/gghp3v" target="_blank">https://doi.org/10/gghp3v</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib24"><label>24</label><mixed-citation>
      
Fang, K., Kifer, D., Lawson, K., Feng, D., and Shen, C.: The data synergy
effects of time-series deep learning models in hydrology, Water Resour.
Res., 58, e2021WR029583, <a href="https://doi.org/10.1029/2021WR029583" target="_blank">https://doi.org/10.1029/2021WR029583</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib25"><label>25</label><mixed-citation>
      
FAO, IIASA, ISRIC, ISSCAS, and JRC: Harmonized World Soil Database (version
1.2), FAO IIASA, ISRIC, ISSCAS, and JRC [data set], <a href="http://www.fao.org/soils-portal/data-hub/soil-maps-and-databases/harmonized-world-soil-database-v12/en/" target="_blank"/> (last access: 1 August 2022),  2012.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib26"><label>26</label><mixed-citation>
      
Feng, D., Fang, K., and Shen, C.: Enhancing streamflow forecast and
extracting insights using long-short term memory networks with data
integration at continental scales, Water Resour. Res., 56, e2019WR026793,
<a href="https://doi.org/10.1029/2019WR026793" target="_blank">https://doi.org/10.1029/2019WR026793</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib27"><label>27</label><mixed-citation>
      
Feng, D., Lawson, K., and Shen, C.: Mitigating prediction error of deep
learning streamflow models in large data-sparse regions with ensemble
modeling and soft data, Geophys. Res. Lett., 48, e2021GL092999,
<a href="https://doi.org/10.1029/2021GL092999" target="_blank">https://doi.org/10.1029/2021GL092999</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib28"><label>28</label><mixed-citation>
      
Feng, D., Liu, J., Lawson, K., and Shen, C.: Differentiable, learnable,
regionalized process-based models with multiphysical outputs can approach
state-of-the-art hydrologic prediction accuracy, Water Resour. Res., 58,
e2022WR032404, <a href="https://doi.org/10.1029/2022WR032404" target="_blank">https://doi.org/10.1029/2022WR032404</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib29"><label>29</label><mixed-citation>
      
Fischer, G., Nachtergaele, F., Prieler, S., van Velthuizen, H. T., Verelst,
L., and Wiberg, D.: Global Agro-Ecological Zones Assessment for Agriculture
(GAEZ 2008), IIASA Laxenburg, Austria and FAO, Rome, Italy, 2008.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib30"><label>30</label><mixed-citation>
      
Friedman, J. H.: Greedy Function Approximation: A Gradient Boosting Machine,
Ann. Stat., 29, 1189–1232, 2001.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib31"><label>31</label><mixed-citation>
      
Gauch, M., Mai, J., and Lin, J.: The Proper Care and Feeding of CAMELS: How
Limited Training Data Affects Streamflow Prediction,  Environ. Modell. Softw., 135, 104926, <a href="https://doi.org/10.1016/j.envsoft.2020.104926" target="_blank">https://doi.org/10.1016/j.envsoft.2020.104926</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib32"><label>32</label><mixed-citation>
      
Hochreiter, S. and Schmidhuber, J.: Long Short-Term Memory, Neural. Comput.,
9, 1735–1780, <a href="https://doi.org/10.1162/neco.1997.9.8.1735" target="_blank">https://doi.org/10.1162/neco.1997.9.8.1735</a>, 1997.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib33"><label>33</label><mixed-citation>
      
Hrachowitz, M., Savenije, H. H. G., Blöschl, G., McDonnell, J. J.,
Sivapalan, M., Pomeroy, J. W., Arheimer, B., Blume, T., Clark, M. P., Ehret,
U., Fenicia, F., Freer, J. E., Gelfan, A., Gupta, H. V., Hughes, D. A., Hut,
R. W., Montanari, A., Pande, S., Tetzlaff, D., Troch, P. A., Uhlenbrook, S.,
Wagener, T., Winsemius, H. C., Woods, R. A., Zehe, E., and Cudennec, C.: A
decade of Predictions in Ungauged Basins (PUB)–a review, Hydrol. Sci. J.,
58, 1198–1255, <a href="https://doi.org/10/gfsq5q" target="_blank">https://doi.org/10/gfsq5q</a>, 2013.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib34"><label>34</label><mixed-citation>
      
Huffman, G. J., Stocker, E. F., Bolvin, D. T., Nelkin, E. J., and Tan, J.:
GPM IMERG Final Precipitation L3 1 day 0.1 degree x 0.1 degree V06
(GPM_3IMERGDF 06), edited by: Andrey Savtchenko, Greenbelt, MD, Goddard Earth Sciences Data and Information Services Center (GES DISC) [data set],
<a href="https://doi.org/10.5067/GPM/IMERGDF/DAY/06" target="_blank">https://doi.org/10.5067/GPM/IMERGDF/DAY/06</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib35"><label>35</label><mixed-citation>
      
Hunter-Jones, P.: Egg development in the Desert Locust (Schistocerca
gregaria Forsk.) in relation to the availability of water, Proc. R. Entomol.
Soc. A, 39, 25–33,
<a href="https://doi.org/10.1111/j.1365-3032.1964.tb00781.x" target="_blank">https://doi.org/10.1111/j.1365-3032.1964.tb00781.x</a>, 1964.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib36"><label>36</label><mixed-citation>
      
Kerr, Y. H., Waldteufel, P., Wigneron, J.-P., Delwart, S., Cabot, F.,
Boutin, J., Escorihuela, M.-J., Font, J., Reul, N., Gruhier, C., Juglea, S.
E., Drinkwater, M. R., Hahne, A., Martín-Neira, M., and Mecklenburg,
S.: The SMOS mission: New tool for monitoring key elements of the global
water cycle, P. IEEE, 98, 666–687, <a href="https://doi.org/10/b9szx6" target="_blank">https://doi.org/10/b9szx6</a>, 2010.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib37"><label>37</label><mixed-citation>
      
Kratzert, F., Klotz, D., Shalev, G., Klambauer, G., Hochreiter, S., and Nearing, G.: Towards learning universal, regional, and local hydrological behaviors via machine learning applied to large-sample datasets, Hydrol. Earth Syst. Sci., 23, 5089–5110, <a href="https://doi.org/10.5194/hess-23-5089-2019" target="_blank">https://doi.org/10.5194/hess-23-5089-2019</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib38"><label>38</label><mixed-citation>
      
Liu, J., Rahmani, F., Lawson, K., and Shen, C.: A multiscale deep learning
model for soil moisture integrating satellite and in situ data, Geophys.
Res. Lett., 49, e2021GL096847, <a href="https://doi.org/10.1029/2021GL096847" target="_blank">https://doi.org/10.1029/2021GL096847</a>, 2022a.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib39"><label>39</label><mixed-citation>
      
Liu, J., Hughes, D., Rahmani, F., Lawson, K., and Shen, C.: Global Soil Moisture Dataset From a Multitask Model (GSM3), Zenodo [data set], <a href="https://doi.org/10.5281/zenodo.7344484" target="_blank">https://doi.org/10.5281/zenodo.7344484</a>, 2022b.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib40"><label>40</label><mixed-citation>
      
Meyal, A. Y., Versteeg, R., Alper, E., Johnson, D., Rodzianko, A., Franklin,
M., and Wainwright, H.: Automated cloud based long short-term memory neural
network based SWE prediction, Front. Water, 2, 1–12,
<a href="https://doi.org/10.3389/frwa.2020.574917" target="_blank">https://doi.org/10.3389/frwa.2020.574917</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib41"><label>41</label><mixed-citation>
      
Muñoz Sabater, J.: ERA5-Land hourly data from 1950 to present, Copernicus Climate Change Service (C3S) Climate Data Store (CDS) [data set], <a href="https://doi.org/10.24381/cds.e2161bac" target="_blank">https://doi.org/10.24381/cds.e2161bac</a>,   2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib42"><label>42</label><mixed-citation>
      
Narasimhan, B. and Srinivasan, R.: Development and evaluation of Soil
Moisture Deficit Index (SMDI) and Evapotranspiration Deficit Index (ETDI)
for agricultural drought monitoring, Agr. Forest Meteorol., 133, 69–88,
<a href="https://doi.org/10/fq9wdv" target="_blank">https://doi.org/10/fq9wdv</a>, 2005.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib43"><label>43</label><mixed-citation>
      
Natural Earth: Free vector and raster map data,  Natural Earth [data set],  <a href="https://www.naturalearthdata.com/" target="_blank"/>, last access: 1 August 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib44"><label>44</label><mixed-citation>
      
Norbiato, D., Borga, M., Degli Esposti, S., Gaume, E., and Anquetin, S.:
Flash flood warning based on rainfall thresholds and soil moisture
conditions: An assessment for gauged and ungauged basins, J. Hydrol., 362,
274–290, <a href="https://doi.org/10/dtw4hm" target="_blank">https://doi.org/10/dtw4hm</a>, 2008.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib45"><label>45</label><mixed-citation>
      
Nuwer, R.: As Locusts Swarmed East Africa, This Tech Helped Squash Them, N.
Y. Times, 8th April,  <a href="https://www.nytimes.com/2021/04/08/science/locust-swarms-africa.html" target="_blank"/> (last access: 1 August 2022), 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib46"><label>46</label><mixed-citation>
      
O, S. and Orth, R.: Global soil moisture data derived through machine
learning trained with in-situ measurements, Sci. Data, 8, 170,
<a href="https://doi.org/10.1038/s41597-021-00964-1" target="_blank">https://doi.org/10.1038/s41597-021-00964-1</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib47"><label>47</label><mixed-citation>
      
O'Neill, P. E., Chan, S., Njoku, E. G., Jackson, T., Bindlish, R., Chaubell,
J., and Colliander, A.: SMAP Enhanced L3 Radiometer Global and Polar Grid
Daily 9&thinsp;km EASE-Grid Soil Moisture, Version 5 (SPL3SMP_E),  Boulder, Colorado USA, NASA National Snow and Ice Data Center Distributed Active Archive Center [data set],
<a href="https://doi.org/10.5067/4DQ54OUIJ9DL" target="_blank">https://doi.org/10.5067/4DQ54OUIJ9DL</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib48"><label>48</label><mixed-citation>
      
Owe, M., de Jeu, R., and Holmes, T.: Multisensor historical climatology of
satellite-derived global land surface moisture, J. Geophys. Res., 113,
1–17, <a href="https://doi.org/10.1029/2007JF000769" target="_blank">https://doi.org/10.1029/2007JF000769</a>, 2008.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib49"><label>49</label><mixed-citation>
      
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel,
O., Blondel, M., Prettenhofer, P., Weiss, R., and Dubourg, V.: Scikit-learn:
Machine learning in Python, J. Mach. Learn. Res., 12, 2825–2830, 2011.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib50"><label>50</label><mixed-citation>
      
Rahmani, F., Shen, C., Oliver, S., Lawson, K., and Appling, A.: Deep
learning approaches for improving prediction of daily stream temperature in
data-scarce, unmonitored, and dammed basins, Hydrol. Process., 35, e14400,
<a href="https://doi.org/10.1002/hyp.14400" target="_blank">https://doi.org/10.1002/hyp.14400</a>, 2021a.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib51"><label>51</label><mixed-citation>
      
Rahmani, F., Lawson, K., Ouyang, W., Appling, A., Oliver, S., and Shen, C.:
Exploring the exceptional performance of a deep learning stream temperature
model and the value of streamflow data, Environ. Res. Lett., 16, 024025,
<a href="https://doi.org/10.1088/1748-9326/abd501" target="_blank">https://doi.org/10.1088/1748-9326/abd501</a>, 2021b.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib52"><label>52</label><mixed-citation>
      
Rodell, M., Houser, P. R., Jambor, U., Gottschalck, J., Mitchell, K., Meng,
C.-J., Arsenault, K., Cosgrove, B., Radakovich, J., Bosilovich, M., Entin,
J. K., Walker, J. P., Lohmann, D., and Toll, D.: The Global Land Data
Assimilation System, B. Am. Meteorol. Soc., 85, 381–394,
<a href="https://doi.org/10.1175/BAMS-85-3-381" target="_blank">https://doi.org/10.1175/BAMS-85-3-381</a>, 2004.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib53"><label>53</label><mixed-citation>
      
Schaaf, C. and Wang, Z.: MODIS/Terra+Aqua BRDF/Albedo Daily L3
Global – 500m V061,   NASA EOSDIS Land Processes DAAC [data set], <a href="https://doi.org/10.5067/MODIS/MCD43A3.061" target="_blank">https://doi.org/10.5067/MODIS/MCD43A3.061</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib54"><label>54</label><mixed-citation>
      
Sheffield, J. and Wood, E. F.: Global trends and variability in soil
moisture and drought characteristics, 1950–2000, from observation-driven
simulations of the terrestrial hydrologic cycle, J. Climate, 21, 432–458,
<a href="https://doi.org/10.1175/2007JCLI1822.1" target="_blank">https://doi.org/10.1175/2007JCLI1822.1</a>, 2008.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib55"><label>55</label><mixed-citation>
      
Shen, C.: A transdisciplinary review of deep learning research and its
relevance for water resources scientists, Water Resour. Res., 54,
8558–8593, <a href="https://doi.org/10.1029/2018wr022643" target="_blank">https://doi.org/10.1029/2018wr022643</a>, 2018.


    </mixed-citation></ref-html>
<ref-html id="bib1.bib56"><label>56</label><mixed-citation>
      
Shen, C., Appling, A. P., Gentine, P., Bandai, T., Gupta, H., Tartakovsky,
A., Baity-Jesi, M., Fenicia, F., Kifer, D., Li, L., Liu, X., Ren, W., Zheng,
Y., Harman, C. J., Clark, M., Farthing, M., Feng, D., Kumar, P.,
Aboelyazeed, D., Rahmani, F., Beck, H. E., Bindas, T., Dwivedi, D., Fang,
K., Höge, M., Rackauckas, C., Roy, T., Xu, C., and Lawson, K.:
Differentiable modeling to unify machine learning and physical models and
advance Geosciences, arXiv [preprint], <a href="https://doi.org/10.48550/arXiv.2301.04027" target="_blank">https://doi.org/10.48550/arXiv.2301.04027</a>, 10 January
2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib57"><label>57</label><mixed-citation>
      
Support CATDS: CATDS-PDC L3SM Simple UDP – 1 day soil moisture Simple User
Data Product from SMOS satellite,  CATDS (CNES, IFREMER, CESBIO) [data set]
<a href="https://doi.org/10.12770/8db7102b-1b22-4db3-949d-e51269417aae" target="_blank">https://doi.org/10.12770/8db7102b-1b22-4db3-949d-e51269417aae</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib58"><label>58</label><mixed-citation>
      
Tsai, W.-P., Feng, D., Pan, M., Beck, H., Lawson, K., Yang, Y., Liu, J., and
Shen, C.: From calibration to parameter learning: Harnessing the scaling
effects of big data in geoscientific modeling, Nat. Commun., 12, 5988,
<a href="https://doi.org/10.1038/s41467-021-26107-z" target="_blank">https://doi.org/10.1038/s41467-021-26107-z</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib59"><label>59</label><mixed-citation>
      
UN WFP: Stop locusts in East Africa now or pay much more to help people
later, U. N. UN World Food Programme WFP, 14th February,  <a href="https://www.wfp.org/news/stop-locusts-east-africa-now-or-pay-much-more-help-people-later-wfp" target="_blank">https://www.wfp.org/news/stop-locusts-east-africa-now-or-pay-much-more-help-people-later-wfp</a>  (last access: 1 August 2022),  2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib60"><label>60</label><mixed-citation>
      
Wan, Z., Hook, S., and Hulley, G.: MODIS/Aqua Land Surface
Temperature/Emissivity Daily L3 Global 1km SIN Grid V061,  NASA EOSDIS Land Processes DAAC [data set],
<a href="https://doi.org/10.5067/MODIS/MYD11A1.061" target="_blank">https://doi.org/10.5067/MODIS/MYD11A1.061</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib61"><label>61</label><mixed-citation>
      
Xue, R., Yang, Q., Miao, F., Wang, X., Shen, Y., Xue, R., Yang, Q., Miao,
F., Wang, X., and Shen, Y.: Slope aspect influences plant biomass, soil
properties and microbial composition in alpine meadow on the Qinghai-Tibetan
plateau, J. Soil Sci. Plant Nut., 18, 1–12,
<a href="https://doi.org/10.4067/S0718-95162018005000101" target="_blank">https://doi.org/10.4067/S0718-95162018005000101</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib62"><label>62</label><mixed-citation>
      
Yang, Z.-L., Niu, G.-Y., Mitchell, K. E., Chen, F., Ek, M. B., Barlage, M.,
Longuevergne, L., Manning, K., Niyogi, D., Tewari, M., and Xia, Y.: The
community Noah land surface model with multiparameterization options
(Noah-MP): 2. Evaluation over global river basins, J. Geophys. Res.-Atmos., 116, D12110, <a href="https://doi.org/10.1029/2010JD015140" target="_blank">https://doi.org/10.1029/2010JD015140</a>, 2011.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib63"><label>63</label><mixed-citation>
      
Zhi, W., Feng, D., Tsai, W.-P., Sterle, G., Harpold, A., Shen, C., and Li,
L.: From hydrometeorology to river water quality: Can a deep learning model
predict dissolved oxygen at the continental scale?, Environ. Sci. Technol.,
55, 2357–2368, <a href="https://doi.org/10.1021/acs.est.0c06783" target="_blank">https://doi.org/10.1021/acs.est.0c06783</a>, 2021.

    </mixed-citation></ref-html>--></article>
