<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing with OASIS Tables v3.0 20080202//EN" "https://jats.nlm.nih.gov/nlm-dtd/publishing/3.0/journalpub-oasis3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:oasis="http://docs.oasis-open.org/ns/oasis-exchange/table" xml:lang="en" dtd-version="3.0" article-type="research-article">
  <front>
    <journal-meta><journal-id journal-id-type="publisher">GMD</journal-id><journal-title-group>
    <journal-title>Geoscientific Model Development</journal-title>
    <abbrev-journal-title abbrev-type="publisher">GMD</abbrev-journal-title><abbrev-journal-title abbrev-type="nlm-ta">Geosci. Model Dev.</abbrev-journal-title>
  </journal-title-group><issn pub-type="epub">1991-9603</issn><publisher>
    <publisher-name>Copernicus Publications</publisher-name>
    <publisher-loc>Göttingen, Germany</publisher-loc>
  </publisher></journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.5194/gmd-18-2521-2025</article-id><title-group><article-title>NN-TOC v1: global prediction of total organic carbon in marine sediments using deep neural networks</article-title><alt-title>NN-TOC v1</alt-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author" corresp="yes" rid="aff1 aff2">
          <name><surname>Parameswaran</surname><given-names>Naveenkumar</given-names></name>
          <email>nparameswaran@geomar.de</email>
        <ext-link>https://orcid.org/0000-0003-3817-2863</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1">
          <name><surname>González</surname><given-names>Everardo</given-names></name>
          
        </contrib>
        <contrib contrib-type="author" corresp="no" rid="aff3">
          <name><surname>Burwicz-Galerne</surname><given-names>Ewa</given-names></name>
          
        <ext-link>https://orcid.org/0000-0003-4551-5609</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff2">
          <name><surname>Braack</surname><given-names>Malte</given-names></name>
          
        </contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1">
          <name><surname>Wallmann</surname><given-names>Klaus</given-names></name>
          
        </contrib>
        <aff id="aff1"><label>1</label><institution>GEOMAR Helmholtz Centre for Ocean Research Kiel, Kiel, Germany</institution>
        </aff>
        <aff id="aff2"><label>2</label><institution>Department of Mathematics, Kiel University, Kiel, Germany</institution>
        </aff>
        <aff id="aff3"><label>3</label><institution>MARUM – Center for Marine Environmental Sciences, University of Bremen, Bremen, Germany</institution>
        </aff>
      </contrib-group>
      <author-notes><corresp id="corr1">Naveenkumar Parameswaran (nparameswaran@geomar.de)</corresp></author-notes><pub-date><day>8</day><month>May</month><year>2025</year></pub-date>
      
      <volume>18</volume>
      <issue>9</issue>
      <fpage>2521</fpage><lpage>2544</lpage>
      <history>
        <date date-type="received"><day>16</day><month>May</month><year>2024</year></date>
           <date date-type="rev-request"><day>4</day><month>June</month><year>2024</year></date>
           <date date-type="rev-recd"><day>13</day><month>November</month><year>2024</year></date>
           <date date-type="accepted"><day>11</day><month>February</month><year>2025</year></date>
      </history>
      <permissions>
        <copyright-statement>Copyright: © 2025 Naveenkumar Parameswaran et al.</copyright-statement>
        <copyright-year>2025</copyright-year>
      <license license-type="open-access"><license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p></license></permissions><self-uri xlink:href="https://gmd.copernicus.org/articles/18/2521/2025/gmd-18-2521-2025.html">This article is available from https://gmd.copernicus.org/articles/18/2521/2025/gmd-18-2521-2025.html</self-uri><self-uri xlink:href="https://gmd.copernicus.org/articles/18/2521/2025/gmd-18-2521-2025.pdf">The full text article is available as a PDF file from https://gmd.copernicus.org/articles/18/2521/2025/gmd-18-2521-2025.pdf</self-uri>
      <abstract><title>Abstract</title>

      <p id="d2e131">Spatial predictions of total organic carbon (TOC) concentrations and stocks are crucial for understanding marine sediments’ role as a significant carbon sink in the global carbon cycle. In this study, we present a geospatial prediction of global TOC concentrations and stocks on a 5 <inline-formula><mml:math id="M1" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 5 arcmin grid, using a novel neural network approach. We also provide and apply a new compilation of over 21 000 global TOC measurements and a new set of predictors, including features such as seafloor lithologies, benthic oxygen fluxes, and chlorophyll-<inline-formula><mml:math id="M2" display="inline"><mml:mi>a</mml:mi></mml:math></inline-formula> satellite data. Moreover, we compare different machine learning models based on their performance metrics and predictions and assess their strengths and limitations. For the dataset used, we find that the performance metrics of the models are comparable and that the neural network approach outperforms, on unseen data, methods such as <inline-formula><mml:math id="M3" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>-nearest neighbours and random forests, which tend to overfit the training data. We provide estimates of mean TOC concentrations and stocks, both on continental shelves and in deep-sea settings across various marine regions and oceans. Our model suggests that the upper 10 cm of oceanic sediments harbour approximately 156 Pg of TOC stocks and have a mean TOC concentration of 0.61 %. Furthermore, we introduce a standardized methodology for quantifying predictive uncertainty using Monte Carlo dropout. The method was applied to our neural network model and underlying features to generate a map of information gain that measures the expected increase in model knowledge, achieved through additional sampling at specific locations, which is pivotal for sampling strategy planning.</p>
  </abstract>
    </article-meta>
  </front>
<body>
      

<sec id="Ch1.S1" sec-type="intro">
  <label>1</label><title>Introduction</title>
      <p id="d2e164">Burial of particulate organic carbon in marine sediments removes carbon dioxide (<inline-formula><mml:math id="M4" display="inline"><mml:mrow class="chem"><mml:msub><mml:mi mathvariant="normal">CO</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>) from the atmosphere and generates molecular oxygen (<inline-formula><mml:math id="M5" display="inline"><mml:mrow class="chem"><mml:msub><mml:mi mathvariant="normal">O</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>) that accumulates in the atmosphere <xref ref-type="bibr" rid="bib1.bibx6 bib1.bibx25" id="paren.1"/>. It is a key process in the global carbon cycle that largely controls the atmospheric partial pressures of <inline-formula><mml:math id="M6" display="inline"><mml:mrow class="chem"><mml:msub><mml:mi mathvariant="normal">O</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M7" display="inline"><mml:mrow class="chem"><mml:msub><mml:mi mathvariant="normal">CO</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> on geological timescales <xref ref-type="bibr" rid="bib1.bibx6 bib1.bibx7" id="paren.2"/>. The mechanisms controlling concentrations, standing stocks, and degradation and accumulation rates of organic carbon at the seabed are, however, complex and remain a topic of active research <xref ref-type="bibr" rid="bib1.bibx1 bib1.bibx12 bib1.bibx25 bib1.bibx31 bib1.bibx10" id="paren.3"/>.  Furthermore, present estimates of the spatial distribution of sedimentary carbon concentrations and stocks across the global ocean, including shelf regions, are limited due to sparse data and the high spatial variability observed in shelf deposits <xref ref-type="bibr" rid="bib1.bibx2 bib1.bibx15 bib1.bibx33 bib1.bibx35 bib1.bibx55" id="paren.4"/>. An improved map of global organic carbon concentrations and stocks in marine surface sediments, including the continental shelf, could, hence, help to better understand processes governing the turnover and accumulation of organic carbon at the seabed.</p>
      <p id="d2e224">Sedimentary organic carbon concentrations are typically reported as total organic carbon (TOC in weight percent), which includes particulate organic carbon bound to sediment grains and a minor contribution by organic carbon dissolved in sediment porewater <xref ref-type="bibr" rid="bib1.bibx25" id="paren.5"/>. TOC varies between different geological environments <xref ref-type="bibr" rid="bib1.bibx16" id="paren.6"/>. Fine-grained shelf and delta sediments deposited close to river mouths typically contain 0.5 %–1.0 % TOC at 0–10 cm sediment depth <xref ref-type="bibr" rid="bib1.bibx6" id="paren.7"/>. A major fraction of TOC deposited in these environments (up to 67 %) is not formed by marine plankton but is produced by land plants <xref ref-type="bibr" rid="bib1.bibx11" id="paren.8"/>. Shelf regions where neritic carbonates are formed by corals and other organisms at the seabed contain about 1 % TOC <xref ref-type="bibr" rid="bib1.bibx6" id="paren.9"/>. However, large parts of the continent shelf (about 50 %–70 %) do not receive sediment inputs and are covered by relict sands <xref ref-type="bibr" rid="bib1.bibx17 bib1.bibx22" id="paren.10"/> that contain only minor amounts of TOC (about 0.1 %). Typical deep-sea sediments, which are not associated with high-productivity regions, contain about 0.2 %–0.4 % TOC <xref ref-type="bibr" rid="bib1.bibx3 bib1.bibx6 bib1.bibx33 bib1.bibx55" id="paren.11"/>. In oceanic upwelling regions with high productivity, large amounts of TOC are rapidly deposited at the seabed such that sedimentary TOC concentrations are usually higher than 1 % and may reach up to 10 % <xref ref-type="bibr" rid="bib1.bibx6 bib1.bibx33 bib1.bibx55" id="paren.12"/>. Elevated TOC values are also reported for surface sediments deposited in the Arctic Ocean (1.0 %) and the deep basins of the Black Sea (2.0 %) <xref ref-type="bibr" rid="bib1.bibx6 bib1.bibx33 bib1.bibx55" id="paren.13"/>. Considering these observations, the global mean TOC concentration in both shelf and deep-sea sediments seems to be close to 0.5 % to 1.0 %.</p>
      <p id="d2e255">The inventory or standing stock of TOC in surface sediments (in mass of carbon per seafloor area) is calculated by multiplying TOC concentrations by the dry bulk density of sediments and the thickness of the considered surface layer. Different methods have been applied to derive the standing stock of TOC at regional and global scales. An early estimate based on limited data and expert knowledge concluded that the global TOC stock is 146 Pg TOC for a 30 cm surface layer <xref ref-type="bibr" rid="bib1.bibx16" id="paren.14"/>. The first estimate of the global TOC inventory derived by a machine learning approach (<inline-formula><mml:math id="M8" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>-nearest neighbours – kNNs) using an extended database (5623 data points) yielded a global inventory of 87 <inline-formula><mml:math id="M9" display="inline"><mml:mo>±</mml:mo></mml:math></inline-formula> 43 Pg TOC in the top 5 cm layer <xref ref-type="bibr" rid="bib1.bibx33" id="paren.15"/>. In subsequent publications with an extended database (11 574 sediment cores) and a more advanced machine learning approach (random forest model), the global inventory was estimated as 2322 Pg TOC for the top 1 m of the sediment column <xref ref-type="bibr" rid="bib1.bibx2" id="paren.16"/>. This inventory exceeds the global TOC inventory in terrestrial soils and suggests that TOC in marine surface sediments is the largest TOC pool at Earth's surface <xref ref-type="bibr" rid="bib1.bibx2" id="paren.17"/>. Another estimate of the global TOC inventory was derived by reactive transport modelling of sedimentary processes employing a range of global datasets <xref ref-type="bibr" rid="bib1.bibx30" id="paren.18"/>. This model yields a global inventory of 170 Pg TOC for the top 10 cm affected by biological mixing processes.</p>
      <p id="d2e288">Since about 70 % of Earth's surface is covered by oceans and sampling sediments at the seafloor is costly, data coverage will always be sparse. Therefore, advanced methods are required to derive spatial information on sediment properties from a limited number of point measurements. Machine learning approaches, which have rapidly advanced in recent years, are the most promising approach to tackling this challenge. So far, <inline-formula><mml:math id="M10" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>-nearest neighbours and random forest models have been applied to derive global maps of sediment porosity <xref ref-type="bibr" rid="bib1.bibx39" id="paren.19"/>, TOC concentration <xref ref-type="bibr" rid="bib1.bibx33" id="paren.20"/>, TOC inventory <xref ref-type="bibr" rid="bib1.bibx2" id="paren.21"/>, sedimentation rate <xref ref-type="bibr" rid="bib1.bibx52 bib1.bibx51" id="paren.22"/>, and regional estimates of TOC accumulation rates <xref ref-type="bibr" rid="bib1.bibx15" id="paren.23"/>. However, machine learning techniques have their own challenges and limitations. Overfitting issues are often encountered, and a standardized approach for estimating predictive uncertainty has not yet been established <xref ref-type="bibr" rid="bib1.bibx33" id="paren.24"/>.</p>
      <p id="d2e318">Given these challenges, this paper aims to derive more robust maps of TOC concentrations and inventories for the global ocean. These maps, including the continental shelf, are based on a new, larger TOC measurement database and an extended collection of predictors to improve the accuracy of predictions for highly heterogeneous and undersampled geological settings. We compiled an enlarged database of TOC concentrations in surface sediments with 21 125 entries and applied a deep neural network (DNN) as a more advanced machine learning approach that considers the non-linear relationships between TOC and other geological features. The global ocean was divided into two different domains (shelf and deep sea), and the network was trained separately for each of these domains. Moreover, we introduced a standardized methodology called Monte Carlo dropout to quantify predictive uncertainties in the DNN model and derive information gain to guide future sampling efforts.</p>
</sec>
<sec id="Ch1.S2">
  <label>2</label><title>Materials</title>
<sec id="Ch1.S2.SS1">
  <label>2.1</label><title>Features</title>
      <p id="d2e336">An extensive repository of features from both the sea surface and the seafloor at a 5 <inline-formula><mml:math id="M11" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 5 arcmin grid resolution has been compiled using previously reported feature lists <xref ref-type="bibr" rid="bib1.bibx33 bib1.bibx52 bib1.bibx23" id="paren.25"/> that include a range of oceanographic, geological, geographic, biological, and biogeochemical parameters. It is worth noting that oceanographic features are updated very often from newer models and measurements, and some of the features used here might be outdated. Features deemed irrelevant to TOC distributions (e.g. crustal and mantle properties, distance to the plate boundary, continental ridges, and trenches) were excluded. Additional features that may influence TOC distributions were added to improve TOC predictions. These include total oxygen uptake (respiration rates) at the seabed  <xref ref-type="bibr" rid="bib1.bibx27" id="paren.26"/>, sediment lithology <xref ref-type="bibr" rid="bib1.bibx20" id="paren.27"/>, tidal velocities <xref ref-type="bibr" rid="bib1.bibx23" id="paren.28"/>, and chlorophyll-<inline-formula><mml:math id="M12" display="inline"><mml:mi>a</mml:mi></mml:math></inline-formula> concentrations at the sea surface <xref ref-type="bibr" rid="bib1.bibx41" id="paren.29"/>.</p>
      <p id="d2e369">Ninety-nine raw feature grids are compiled for a comprehensive representation of the marine environment, providing the necessary input for the neural network analysis in this study to predict TOC concentrations in marine sediments. Most of these features are easily measurable from the sea surface by e.g. satellite observations, making them a reliable dataset compared to the less accessible properties of the seafloor. Some of the seafloor feature grids used in this work were previously generated from raw data using machine learning methods (e.g. a porosity grid provided by <xref ref-type="bibr" rid="bib1.bibx39" id="altparen.30"/>). Others were reprocessed in this work to achieve global coverage at a resolution of 5 <inline-formula><mml:math id="M13" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 5 arcmin (e.g. sediment lithology, <xref ref-type="bibr" rid="bib1.bibx20" id="altparen.31"/>).</p>
      <p id="d2e385">Neighbourhood information was incorporated for a subset of the features. Specifically, 40 of the 99 initial features were averaged spatially using a 50 km radius <xref ref-type="bibr" rid="bib1.bibx33" id="paren.32"/>. Spatial averaging was applied when TOC concentrations are assumed to be affected not only by the local feature value, but also feature values in the surrounding area. This approach was used for selected physical (e.g. current velocity), chemical (e.g. dissolved compounds), and biological (e.g. bio-fauna abundance) parameters.</p>
      <p id="d2e391">Overall, a total of 139 features, including 99 original features and 40 additional spatially averaged features, is used in the model. The complete feature list is presented in Appendix <xref ref-type="sec" rid="App1.Ch1.S1"/>.</p>
</sec>
<sec id="Ch1.S2.SS2">
  <label>2.2</label><title>TOC data</title>
      <p id="d2e404">The dataset for TOC concentrations (in weight percent) utilized in this study has been compiled from multiple sources. It includes global datasets <xref ref-type="bibr" rid="bib1.bibx55 bib1.bibx53 bib1.bibx61 bib1.bibx45" id="paren.33"/> and regional datasets for the northern Gulf of Mexico <xref ref-type="bibr" rid="bib1.bibx4" id="paren.34"/> and the North Sea (Wenyen Zhang, personal communication, 2023, HEREON). Each label represents a known measurement (TOC concentration) and is paired with the nearest grid point on the 139-feature grids via L2 distance computation, resulting in the association of a feature vector with each label. For those stations where TOC is reported as function of sediment depth, we calculated the mean TOC concentration for the top 10 cm and used this mean as the model label. For many stations, values are only reported for the top 1–2 cm (around 19 000 measurements). We included these stations in our model since they contain valuable information, but we acknowledge that they may be somewhat higher than those integrated over the top 10 cm since TOC concentrations tend to decrease with sediment depth due to ongoing TOC degradation. However, most sediments deposited on the continental shelf and in high-productivity regions of the open ocean are affected by intense biogenic and physical mixing processes <xref ref-type="bibr" rid="bib1.bibx8" id="paren.35"/> such that the downcore TOC decrease is usually small within the mixed surface layer (0–10 cm sediment depth). The labelled data are pre-processed to enhance the reliability and robustness of the dataset for subsequent model development and validation. We first searched for duplicates in our combined database that may arise when the same data are reported in multiple databases. They were removed from the combined database when longitudes, latitudes, and TOC concentrations were identical. Moreover, coastal regions often exhibit clustered measurements, potentially resulting in shared feature vectors, as all the measurements lie in the same feature grid cell. To mitigate this, a variance assessment is conducted. Labels that share the same feature vectors, exhibiting high variance (the standard deviation of these labels is higher than 20 % of the maximum of these labels) are excluded, while those with low variance are averaged, and the shared feature vector is assigned. Also, some data points situated in close proximity to land were not captured adequately by the 5 <inline-formula><mml:math id="M14" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 5 arcmin grid. To address this, reasonable values are assigned by interpolating from the nearest points, ensuring the overall quality of the dataset. Our database includes a total of 110 149 data points that have been consolidated as discussed above such that the final TOC database employed in the model is composed of 21 125 entries (Fig. <xref ref-type="fig" rid="Ch1.F1"/>). Both the datasets for the labels and features can be downloaded at <ext-link xlink:href="https://doi.org/10.5281/zenodo.11186224" ext-link-type="DOI">10.5281/zenodo.11186224</ext-link> <xref ref-type="bibr" rid="bib1.bibx47" id="paren.36"/>.</p>

      <fig id="Ch1.F1" specific-use="star"><label>Figure 1</label><caption><p id="d2e434">Quantitative TOC measurements (i.e. labels) acquired from various sources <xref ref-type="bibr" rid="bib1.bibx55 bib1.bibx53 bib1.bibx61 bib1.bibx4 bib1.bibx45" id="paren.37"/>. Notably, data point clusters are observed in close proximity to coastal regions. The colour maps used for the figures in this paper are from <xref ref-type="bibr" rid="bib1.bibx13" id="text.38"/> and <xref ref-type="bibr" rid="bib1.bibx60" id="text.39"/>.</p></caption>
          <graphic xlink:href="https://gmd.copernicus.org/articles/18/2521/2025/gmd-18-2521-2025-f01.png"/>

        </fig>

</sec>
</sec>
<sec id="Ch1.S3">
  <label>3</label><title>Methods</title>
      <p id="d2e461">The primary objective of this study is to build a supervised prediction model that uses feature grid maps as inputs to predict TOC concentrations as outputs. Additionally, we aim to quantify prediction uncertainties using Monte Carlo dropout and information theory techniques. The supervised model is trained using the set of labels (TOC data) and their corresponding feature vectors. Due to the non-linearity in the relationships between data and features, we choose deep-learning models, which are good at understanding such patterns. Deep neural networks (DNNs) transform data non-linearly with non-linear activation functions such as ReLU (rectified linear unit), a piecewise linear function that outputs 0 for negative inputs and the input itself for positive inputs, introducing non-linearity into the DNN. Therefore, even after one layer, multi-collinearity in the data is eliminated. In our case of a deep neural network, the final output is controlled by  numerous combinations of ReLU functions involving higher-order interactions of original features <xref ref-type="bibr" rid="bib1.bibx14" id="paren.40"/>.</p>
<sec id="Ch1.S3.SS1">
  <label>3.1</label><title>Deep-learning model</title>
      <p id="d2e474">Deep neural networks have achieved state-of-the-art results in a variety of tasks in ocean observation, prediction, and forecasting of ocean phenomena <xref ref-type="bibr" rid="bib1.bibx58" id="paren.41"/>. DNN architectures, which are intrinsically non-parametric and non-linear, are less susceptible to the curse of dimensionality. They capture complex relationships between data and features at different levels of abstraction through their hierarchical nature, which makes them well-suited to resolving highly complex geoscientific problems <xref ref-type="bibr" rid="bib1.bibx32" id="paren.42"/>.</p>
      <p id="d2e483">Here, we use a multi-layer perceptron (MLP), feed-forward DNN to predict global TOC in sediments and introduce a new approach to mapping uncertainty in predictions that serves as a quantifiable measure of information gain from sampling. To further improve our predictions,  the global ocean was separated into a continental shelf and deep-sea region using the 200 m water depth horizon as a boundary. Two separate models were trained for these regions (shelf: 0–200 m, deep sea: <inline-formula><mml:math id="M15" display="inline"><mml:mo>&gt;</mml:mo></mml:math></inline-formula> 200 m) to consider the different processes that drive sedimentation and control TOC values in the deep-sea and shelf environments. The same set of features is used for both regions, but the interplay of these features differs between the contrasting environments. The weights and biases in the DNN are initialized using the technique proposed by <xref ref-type="bibr" rid="bib1.bibx24" id="text.43"/>. Batch normalization (which normalizes the inputs of each layer for faster and more stable training) and dropout (which assigns a probability of being deactivated to each node during training and thus prevents overfitting) are applied to each layer for regularization. ReLU is used as the activation function.</p>
      <p id="d2e496">The Monte Carlo dropout method is implemented here to estimate uncertainty in the DNN model, leveraging dropout layers as approximate Bayesian inferences <xref ref-type="bibr" rid="bib1.bibx19" id="paren.44"/>. It gives us an ensemble of predictions from different subsets of neurons in the same DNN model. Kullback–Leibler (KL) divergence is used to map information gain from the quantified predictive uncertainty. In the field of information theory, KL divergence represents the information gain and is defined as the difference of the cross-entropy between the observation and prediction of an event and the entropy in the observation of the event <xref ref-type="bibr" rid="bib1.bibx29" id="paren.45"/>. In our context, the predicted distribution arises from the Monte Carlo dropout prediction ensemble, while the reconstructed observed distribution is modelled with a normal distribution, with the predicted value as a mean and a standard deviation of 0.05 TOC % arising from both technical handling and the precision of the weighting tool <xref ref-type="bibr" rid="bib1.bibx44" id="paren.46"/>.</p>
      <p id="d2e508">Uncertainty and information gain are inherently associated insofar as there cannot be high information gain without high uncertainty. However, information gain also depends on the observation probability distribution and is constrained by it. In other words, information gain measures the expected increase in model knowledge achieved through field sampling at a specific location. This concept provides a strategic guide for determining optimal sampling strategies: taking samples in regions with the highest information gain values is the most efficient way of refining our model’s representation of the real world. The mathematical formulation of entropy, cross-entropy, and information gain is detailed in Appendix <xref ref-type="sec" rid="App1.Ch1.S2"/>.</p>
</sec>
</sec>
<sec id="Ch1.S4">
  <label>4</label><title>Results and discussions</title>
      <p id="d2e522">Understanding the global distribution of TOC concentrations and stocks is crucial for advancing our knowledge of the carbon cycle and sedimentary environments worldwide. Before delving into the prediction maps from the DNN, we first compare the performance of three methods: DNNs, kNNs, and random forests. Separate models are run for the deep-sea and continental shelf regions, and the outcomes are summarized in Table <xref ref-type="table" rid="Ch1.T1"/>. For kNN, five neighbours were utilized for the continental shelves and four for the deep sea, based on a sensitivity analysis with respect to model performance. Random forests employed 100 estimators for both marine regions. The DNN consists of 10 layers with 128 nodes each. The choice of hyperparameters in the models is discussed in Appendix <xref ref-type="sec" rid="App1.Ch1.S3"/>. This comparison sets the groundwork for a detailed exploration of DNN results. All of the methods were run with the same training–testing splits of the dataset, and the random split is seeded to make the methods reproducible.</p>

<table-wrap id="Ch1.T1" specific-use="star"><label>Table 1</label><caption><p id="d2e532">Comparison of machine learning methods based on performance metrics: Pearson correlation coefficient (Pearson CC), coefficient of determination (<inline-formula><mml:math id="M16" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>), and mean squared error (MSE) for predicted values vs. observed labels for the training and testing data. The <inline-formula><mml:math id="M17" display="inline"><mml:mrow><mml:mi mathvariant="normal">train</mml:mi><mml:mo>:</mml:mo><mml:mi mathvariant="normal">test</mml:mi></mml:mrow></mml:math></inline-formula> data ratio is <inline-formula><mml:math id="M18" display="inline"><mml:mrow><mml:mn mathvariant="normal">85</mml:mn><mml:mo>:</mml:mo><mml:mn mathvariant="normal">15</mml:mn></mml:mrow></mml:math></inline-formula>. The better-performing method, based on the metric, is highlighted in bold.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="7">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="center"/>
     <oasis:colspec colnum="3" colname="col3" align="center"/>
     <oasis:colspec colnum="4" colname="col4" align="center" colsep="1"/>
     <oasis:colspec colnum="5" colname="col5" align="center"/>
     <oasis:colspec colnum="6" colname="col6" align="center"/>
     <oasis:colspec colnum="7" colname="col7" align="center"/>
     <oasis:thead>
       <oasis:row>
         <oasis:entry colname="col1">Method</oasis:entry>
         <oasis:entry rowsep="1" namest="col2" nameend="col4" colsep="1">Training data </oasis:entry>
         <oasis:entry rowsep="1" namest="col5" nameend="col7">Testing data (15 % of all the data) </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">Pearson CC</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M19" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">MSE</oasis:entry>
         <oasis:entry colname="col5">Pearson CC</oasis:entry>
         <oasis:entry colname="col6"><inline-formula><mml:math id="M20" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col7">MSE</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">kNN</oasis:entry>
         <oasis:entry colname="col2">0.921</oasis:entry>
         <oasis:entry colname="col3">0.45</oasis:entry>
         <oasis:entry colname="col4">0.517</oasis:entry>
         <oasis:entry colname="col5">0.852</oasis:entry>
         <oasis:entry colname="col6">0.719</oasis:entry>
         <oasis:entry colname="col7">0.541</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Random forest</oasis:entry>
         <oasis:entry colname="col2"><bold>0.986</bold></oasis:entry>
         <oasis:entry colname="col3"><bold> 0.966</bold></oasis:entry>
         <oasis:entry colname="col4"><bold>0.239</bold></oasis:entry>
         <oasis:entry colname="col5">0.867</oasis:entry>
         <oasis:entry colname="col6">0.745</oasis:entry>
         <oasis:entry colname="col7"><bold>0.499</bold></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">DNN</oasis:entry>
         <oasis:entry colname="col2">0.928</oasis:entry>
         <oasis:entry colname="col3">0.844</oasis:entry>
         <oasis:entry colname="col4">0.492</oasis:entry>
         <oasis:entry colname="col5"><bold>0.888</bold></oasis:entry>
         <oasis:entry colname="col6"><bold> 0.737</bold></oasis:entry>
         <oasis:entry colname="col7">0.537</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d2e731">The results of this model comparison show that random forest and kNN algorithms exhibit higher correlation coefficients and superior overall performance in the training dataset than the DNN. However, the DNN outperforms the other two algorithms in the testing data performance (Table <xref ref-type="table" rid="Ch1.T1"/>) for the dataset used. This discrepancy suggests a potential overfitting issue, where the kNN and random forest models may have become specialized in learning the training data. The emphasis on generalization capabilities is crucial in our context due to data scarcity in many regions, making predictions in unexplored  areas a priority. The correlation plot between the measured and predicted data shows similar errors for the training and testing datasets, which confirms that the DNN model largely avoids overfitting (Fig. <xref ref-type="fig" rid="Ch1.F2"/>). The observed underestimation of TOC concentrations at higher values is likely due to the distribution of the ground truth dataset, which is predominantly composed of low TOC concentrations (<inline-formula><mml:math id="M21" display="inline"><mml:mo lspace="0mm">&lt;</mml:mo></mml:math></inline-formula> 1 %). Training an NN model on such an imbalanced dataset often results in a model that is biased towards predicting lower values, effectively “erring on the side of caution”. Several approaches could be employed to address this issue, such as weighting the gradient descent steps based on concentration values, applying a logarithmic transformation to the TOC scale, or balancing the dataset by withholding low-value labels. However, each of these methods is likely to introduce trade-offs, potentially reducing accuracy in other areas. Ultimately, the most effective way of improving the model's performance in predicting higher TOC concentrations is to obtain additional TOC samples within this higher range.</p>

      <fig id="Ch1.F2" specific-use="star"><label>Figure 2</label><caption><p id="d2e748">Heatmap of the correlation plot between measured (labels) and predicted data (targets) using DNN for <bold>(a)</bold> training data and <bold>(b)</bold> testing data in order to assess the model performance. The minimal difference observed between the training and testing errors serves as an indicator of the model's ability to avoid overfitting.</p></caption>
        <graphic xlink:href="https://gmd.copernicus.org/articles/18/2521/2025/gmd-18-2521-2025-f02.png"/>

      </fig>

      <p id="d2e763">The prediction map of the DNN is presented in Fig. <xref ref-type="fig" rid="Ch1.F3"/>, while maps generated by kNN and random forests are provided in Appendix <xref ref-type="sec" rid="App1.Ch1.S3"/> (Figs. <xref ref-type="fig" rid="App1.Ch1.S3.F7"/> and <xref ref-type="fig" rid="App1.Ch1.S3.F8"/>). Both the kNN and random forests showed artifacts, particularly in the equatorial Pacific and Atlantic oceans, similar to the map published by <xref ref-type="bibr" rid="bib1.bibx33" id="text.47"/>. As stated by <xref ref-type="bibr" rid="bib1.bibx33" id="text.48"/>, there is no standard means of quantifying uncertainty in kNN. In random forests, the variance or standard deviation of all the sub-output values for measuring the regression uncertainty is considered an uncertainty quantification method, but it is difficult to provide the uncertainty for an individual base learner <xref ref-type="bibr" rid="bib1.bibx36" id="paren.49"/>. Estimating the confidence of the predictions should be an important factor in deciding which model to use. On the other hand, uncertainty quantification in the DNN is an active field of research and has standardized methods. Nonetheless, kNNs and random forests are useful learning algorithms when computational resources are constrained and require an out-of-the-box solution.</p>

      <fig id="Ch1.F3" specific-use="star"><label>Figure 3</label><caption><p id="d2e786">Global prediction map of the TOC concentration using a DNN. A higher TOC concentration is observed in the Arctic region and in upwelling areas located along the western continental margins of America and Africa, the equatorial Pacific, and the Arabian Sea.</p></caption>
        <graphic xlink:href="https://gmd.copernicus.org/articles/18/2521/2025/gmd-18-2521-2025-f03.png"/>

      </fig>

      <p id="d2e795">We also tested a DNN model where the global ocean was not separated into shelf and deep-sea regions but treated as one entity. The resulting TOC map shows spurious features in the Pacific Ocean (Appendix <xref ref-type="sec" rid="App1.Ch1.S5"/>), similar to those found in previous predictions. This additional model shows that the separation of the ocean into shelf and deep-sea regions improves the model results.</p>
      <p id="d2e800">Our DNN-based map of TOC concentrations (Fig. <xref ref-type="fig" rid="Ch1.F3"/>) shows similarities to maps previously published by <xref ref-type="bibr" rid="bib1.bibx55" id="text.50"/> and <xref ref-type="bibr" rid="bib1.bibx33" id="text.51"/>, who used geostatistical methods and a kNN model. All of the maps show elevated concentrations in the Arctic region and in upwelling areas located along the western continental margins of North America, South America, Africa, the equatorial Pacific, and the Arabian Sea. This pattern can be explained by elevated rates of marine primary and export production in upwelling regions delivering large fluxes of TOC to the seabed. The low TOC values in the open oceans are related to lower productivity and the large water depths limiting the TOC flux to the deep-sea floor. The predictions in Fig. <xref ref-type="fig" rid="Ch1.F3"/> are also consistent with the early work on TOC distributions by <xref ref-type="bibr" rid="bib1.bibx6" id="text.52"/> and <xref ref-type="bibr" rid="bib1.bibx16" id="text.53"/>, showing low TOC values in the open oceans and elevated values for upwelling regions and the Arctic region. The high TOC concentrations predicted for the Black Sea and Baltic Sea (Fig. <xref ref-type="fig" rid="Ch1.F3"/>) are probably related to the lack of oxygen in the bottom waters of these marginal seas that promotes TOC preservation <xref ref-type="bibr" rid="bib1.bibx25" id="paren.54"/>. The map published by <xref ref-type="bibr" rid="bib1.bibx33" id="text.55"/> shows several large areas in the open Pacific that have unusually high TOC concentrations.  These patches are probably not realistic since they do not appear in other maps and are not consistent with our understanding of the TOC cycle. They may be artifacts generated by the kNN method and the sparse data coverage in these regions. Our new map avoids these artifacts and presents a pattern that corresponds better to our understanding of TOC accumulation on the seafloor for both the deep sea and the continental shelf, which were never modelled individually on previous maps.</p>
      <p id="d2e829">We also produced a map of TOC stocks for the global ocean (Fig. <xref ref-type="fig" rid="Ch1.F4"/>). The TOC stocks were calculated using the global porosity grid provided by <xref ref-type="bibr" rid="bib1.bibx39" id="text.56"/> and a density of dry solids (<inline-formula><mml:math id="M22" display="inline"><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mi mathvariant="normal">s</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) of 2.6 g cm<sup>−3</sup>. We performed the calculation for the top 10 cm of the sediment column since our TOC data have been measured within this mixed surface layer. Moreover, the top 10 cm are the most vulnerable and dynamic part of the sedimentary TOC pool since they are subject to frequent biological and physical mixing  processes <xref ref-type="bibr" rid="bib1.bibx57" id="paren.57"/> and are affected by human interventions such as bottom trawling <xref ref-type="bibr" rid="bib1.bibx54" id="paren.58"/>.
          <disp-formula id="Ch1.E1" content-type="numbered"><label>1</label><mml:math id="M24" display="block"><mml:mtable class="split" rowspacing="0.2ex" displaystyle="true" columnalign="right left"><mml:mtr><mml:mtd><mml:mrow><mml:mtext>TOC stock</mml:mtext></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mo>=</mml:mo><mml:mo>(</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:mspace linebreak="nobreak" width="0.25em"/><mml:mtext>porosity</mml:mtext><mml:mo>)</mml:mo><mml:mo>×</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi mathvariant="normal">s</mml:mi></mml:msub></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mo>×</mml:mo><mml:mspace width="0.25em" linebreak="nobreak"/><mml:mtext>TOC concentration</mml:mtext><mml:mspace linebreak="nobreak" width="0.25em"/><mml:mo>×</mml:mo><mml:mspace linebreak="nobreak" width="0.25em"/><mml:mtext>10 cm</mml:mtext></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>

      <fig id="Ch1.F4" specific-use="star"><label>Figure 4</label><caption><p id="d2e916">TOC stock map using the global porosity grid provided by <xref ref-type="bibr" rid="bib1.bibx39" id="text.59"/>. The colour map is shown on a logarithmic scale.</p></caption>
        <graphic xlink:href="https://gmd.copernicus.org/articles/18/2521/2025/gmd-18-2521-2025-f04.png"/>

      </fig>

      <p id="d2e928">The TOC stock is computed for global oceans and major seas <xref ref-type="bibr" rid="bib1.bibx18" id="paren.60"/>,  considering both continental shelves and deep-sea regions within each ocean and sea (Table <xref ref-type="table" rid="Ch1.T2"/>). Notably, the mean TOC concentration on continental shelves exhibits significant variability across regions. A visualization of the TOC stock in the oceans is provided in Appendix <xref ref-type="sec" rid="App1.Ch1.S4"/>.</p>

<table-wrap id="Ch1.T2" specific-use="star"><label>Table 2</label><caption><p id="d2e941">TOC stock in the continental shelf  and deep-sea regions.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="7">
     <oasis:colspec colnum="1" colname="col1" align="justify" colwidth="3cm"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right" colsep="1"/>
     <oasis:colspec colnum="5" colname="col5" align="right"/>
     <oasis:colspec colnum="6" colname="col6" align="right"/>
     <oasis:colspec colnum="7" colname="col7" align="right"/>
     <oasis:thead>
       <oasis:row>
         <oasis:entry colname="col1" align="left"/>
         <oasis:entry rowsep="1" namest="col2" nameend="col4" align="center" colsep="1">Continental shelves </oasis:entry>
         <oasis:entry rowsep="1" namest="col5" nameend="col7" align="center">Deep sea </oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1" align="left">Region</oasis:entry>
         <oasis:entry colname="col2">Sum of the TOC</oasis:entry>
         <oasis:entry colname="col3">Area</oasis:entry>
         <oasis:entry colname="col4">Mean TOC</oasis:entry>
         <oasis:entry colname="col5">Sum of the</oasis:entry>
         <oasis:entry colname="col6">Area</oasis:entry>
         <oasis:entry colname="col7">Mean TOC</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left"/>
         <oasis:entry colname="col2">stock (Pg)</oasis:entry>
         <oasis:entry colname="col3">(<inline-formula><mml:math id="M28" display="inline"><mml:mrow><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mn mathvariant="normal">6</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> km<sup>2</sup>)</oasis:entry>
         <oasis:entry colname="col4">concentration (%)</oasis:entry>
         <oasis:entry colname="col5">TOC stock (Pg)</oasis:entry>
         <oasis:entry colname="col6">(<inline-formula><mml:math id="M30" display="inline"><mml:mrow><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mn mathvariant="normal">6</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> km<sup>2</sup>)</oasis:entry>
         <oasis:entry colname="col7">concentration (%)</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Arctic Ocean</oasis:entry>
         <oasis:entry colname="col2">5.57</oasis:entry>
         <oasis:entry colname="col3">5.72</oasis:entry>
         <oasis:entry colname="col4">0.94</oasis:entry>
         <oasis:entry colname="col5">7.18</oasis:entry>
         <oasis:entry colname="col6">9.46</oasis:entry>
         <oasis:entry colname="col7">0.88</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Indian Ocean</oasis:entry>
         <oasis:entry colname="col2">2.63</oasis:entry>
         <oasis:entry colname="col3">4.06</oasis:entry>
         <oasis:entry colname="col4">0.61</oasis:entry>
         <oasis:entry colname="col5">25.86</oasis:entry>
         <oasis:entry colname="col6">67.10</oasis:entry>
         <oasis:entry colname="col7">0.55</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Mediterranean Sea</oasis:entry>
         <oasis:entry colname="col2">0.61</oasis:entry>
         <oasis:entry colname="col3">0.65</oasis:entry>
         <oasis:entry colname="col4">0.98</oasis:entry>
         <oasis:entry colname="col5">1.95</oasis:entry>
         <oasis:entry colname="col6">2.27</oasis:entry>
         <oasis:entry colname="col7">1.04</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">North Atlantic Ocean</oasis:entry>
         <oasis:entry colname="col2">2.82</oasis:entry>
         <oasis:entry colname="col3">4.26</oasis:entry>
         <oasis:entry colname="col4">0.63</oasis:entry>
         <oasis:entry colname="col5">14.96</oasis:entry>
         <oasis:entry colname="col6">37.46</oasis:entry>
         <oasis:entry colname="col7">0.58</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">North Pacific Ocean</oasis:entry>
         <oasis:entry colname="col2">2.74</oasis:entry>
         <oasis:entry colname="col3">3.83</oasis:entry>
         <oasis:entry colname="col4">0.66</oasis:entry>
         <oasis:entry colname="col5">31.16</oasis:entry>
         <oasis:entry colname="col6">73.42</oasis:entry>
         <oasis:entry colname="col7">0.67</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">South Atlantic Ocean</oasis:entry>
         <oasis:entry colname="col2">1.25</oasis:entry>
         <oasis:entry colname="col3">1.86</oasis:entry>
         <oasis:entry colname="col4">0.82</oasis:entry>
         <oasis:entry colname="col5">12.76</oasis:entry>
         <oasis:entry colname="col6">38.67</oasis:entry>
         <oasis:entry colname="col7">0.51</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">South China and Easter Archipelago seas</oasis:entry>
         <oasis:entry colname="col2">1.56</oasis:entry>
         <oasis:entry colname="col3">3.00</oasis:entry>
         <oasis:entry colname="col4">0.48</oasis:entry>
         <oasis:entry colname="col5">2.58</oasis:entry>
         <oasis:entry colname="col6">3.74</oasis:entry>
         <oasis:entry colname="col7">0.83</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">South Pacific Ocean</oasis:entry>
         <oasis:entry colname="col2">1.26</oasis:entry>
         <oasis:entry colname="col3">1.46</oasis:entry>
         <oasis:entry colname="col4">1.02</oasis:entry>
         <oasis:entry colname="col5">32.96</oasis:entry>
         <oasis:entry colname="col6">83.81</oasis:entry>
         <oasis:entry colname="col7">0.58</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Southern Ocean</oasis:entry>
         <oasis:entry colname="col2">0.21</oasis:entry>
         <oasis:entry colname="col3">0.57</oasis:entry>
         <oasis:entry colname="col4">0.59</oasis:entry>
         <oasis:entry colname="col5">6.39</oasis:entry>
         <oasis:entry colname="col6">20.16</oasis:entry>
         <oasis:entry colname="col7">0.43</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Baltic Sea<sup>*</sup></oasis:entry>
         <oasis:entry colname="col2">0.77</oasis:entry>
         <oasis:entry colname="col3">0.39</oasis:entry>
         <oasis:entry colname="col4">3.03</oasis:entry>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6"/>
         <oasis:entry colname="col7"/>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Caspian Sea<sup>*</sup></oasis:entry>
         <oasis:entry colname="col2">0.72</oasis:entry>
         <oasis:entry colname="col3">0.38</oasis:entry>
         <oasis:entry colname="col4">2.27</oasis:entry>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6"/>
         <oasis:entry colname="col7"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1" align="left">Total</oasis:entry>
         <oasis:entry colname="col2">20.15</oasis:entry>
         <oasis:entry colname="col3">26.20</oasis:entry>
         <oasis:entry colname="col4">0.79</oasis:entry>
         <oasis:entry colname="col5">135.80</oasis:entry>
         <oasis:entry colname="col6">336.08</oasis:entry>
         <oasis:entry colname="col7">0.59</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table><table-wrap-foot><p id="d2e944"><sup>*</sup> The total sums and the mean concentrations in the continental shelves include the Baltic Sea and the Caspian Sea. Without these regions, the total TOC stock in the continental shelves is 18.66 Pg, the area of the continental shelves is 25.42 <inline-formula><mml:math id="M26" display="inline"><mml:mrow><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mn mathvariant="normal">6</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> km<sup>2</sup>, and the mean TOC concentration is 0.66 %.</p></table-wrap-foot></table-wrap>

      <p id="d2e1419">According to our model, most the TOC stock can be found in the vast deep-sea basins of the Pacific, Indian, and Atlantic oceans, which is due to the large area of these basins (Table <xref ref-type="table" rid="Ch1.T2"/>). The shelf region harbours 12.1 % of the global stock (Table <xref ref-type="table" rid="Ch1.T2"/>, excluding the Baltic Sea and Caspian Sea), similar to the fraction previously derived by <xref ref-type="bibr" rid="bib1.bibx2" id="text.61"/>, who suggested that 11.5 % of the global TOC stock is located on the continental shelves. The global TOC stock derived from our model amounts to 155.8 Pg carbon for the 10 cm layer considered in our calculations (Table <xref ref-type="table" rid="Ch1.T2"/>). This value is close to the global stock in the top 10 cm derived by reactive transport modelling (170 Pg carbon, <xref ref-type="bibr" rid="bib1.bibx30" id="altparen.62"/>). The other stock estimates were calculated by applying a range of sediment thicknesses. When normalized to 10 cm, the stock reported by <xref ref-type="bibr" rid="bib1.bibx33" id="text.63"/> amounts to 174 Pg carbon, while the stock derived by <xref ref-type="bibr" rid="bib1.bibx2" id="text.64"/> amounts to 232 Pg carbon. The first stock estimate, which was based on expert knowledge and a limited database, corresponds to only 49 Pg carbon when normalized to 10 cm <xref ref-type="bibr" rid="bib1.bibx16" id="paren.65"/>, which is lower than our estimate. Our new global stock assessment, hence, falls into the range of previous estimates.</p>
      <p id="d2e1445">According to our DNN model, the mean TOC concentration in continental shelf sediments, excluding the Baltic Sea and the Caspian Sea (0.70 %), is close to the concentration in deep-sea sediments (0.59 %, Table <xref ref-type="table" rid="Ch1.T2"/>). This is a surprising result since the high marine productivity and low water depths on the shelf induce high TOC fluxes to the seabed that should result in elevated TOC concentrations in surface sediments. Moreover, large amounts of terrestrial particulate organic carbon (POC) produced by land plants are deposited in shelf sediments <xref ref-type="bibr" rid="bib1.bibx11" id="paren.66"/>, which should further increase TOC concentrations in these deposits. However, TOC concentrations in shelf surface sediments are diminished by a number of factors: (i) frequent biological and physical reworking that accelerates TOC degradation processes <xref ref-type="bibr" rid="bib1.bibx57" id="paren.67"/>, (ii) dilution of TOC by inorganic material (clay, silt, and sand) in delta deposits and other shelf regions with high sedimentation rates <xref ref-type="bibr" rid="bib1.bibx6" id="paren.68"/>, (iii) strong bottom currents that inhibit sediment deposition such that large shelf areas are covered by relict coarse-grained sediments that were deposited in the geological past and that do not contain significant amounts of TOC <xref ref-type="bibr" rid="bib1.bibx17" id="paren.69"/>, and (iv) frequent bottom trawling that exposes sedimentary TOC to oxygen and accelerates TOC degradation <xref ref-type="bibr" rid="bib1.bibx2" id="paren.70"/>. According to our DNN model, these factors could potentially decrease TOC concentrations in shelf sediments to such a degree that they attain mean values that are close to those observed in deep-sea sediments (Table <xref ref-type="table" rid="Ch1.T2"/>). It should, however, be noted that most TOC burial occurs on the shelf, where sedimentation rates are elevated due to the deposition of riverine particles <xref ref-type="bibr" rid="bib1.bibx10" id="paren.71"/>.</p>
      <p id="d2e1471">A method based on cooperative game theory (SHAP, SHapley Additive exPlanations) is used to analyse our results further and identify features that have a large effect on the predicted distribution of TOC concentrations <xref ref-type="bibr" rid="bib1.bibx38" id="paren.72"/>. The higher the SHAP value for a feature, the more important the feature is for the predictions of that particular model. According to our model analysis, the total oxygen uptake feature <xref ref-type="bibr" rid="bib1.bibx27" id="paren.73"/> has the largest effect (SHAP value) on predicted TOC concentrations in shelf sediments, while the global porosity grid <xref ref-type="bibr" rid="bib1.bibx39" id="paren.74"/> was the most important feature for deep-sea sediments. It should, however, be noted that the feature importance ranking is only valid for our specific model set-up and might not be representative of the real world. Model interpretability and feature importance ranking are discussed further in Appendix <xref ref-type="sec" rid="App1.Ch1.S6"/>.</p>
      <p id="d2e1485">To guide future sampling, a new information gain map is provided (Fig. <xref ref-type="fig" rid="Ch1.F5"/>). It identifies the regions that should be explored in order to improve the current model predictions. Some of the main takeaways from the information gain map are that (i) regions with high information gain are found in parts of the equatorial Pacific Ocean, in Zealandia, and around Papua New Guinea. These regions are less explored geographically, and hence the model is not trained with the features in this region. (ii) The continental slopes on the western coast of North America, in the east of Iceland, and in parts of the eastern coast of Africa have a higher information gain, though they have more measurements. This could be due to the steep slopes and rough topography in these regions that may induce a high spatial heterogeneity in TOC values that is not yet resolved by the model. (iii) Though the Southern Ocean is not well-explored, regions with higher information gain are only found with relatively steep terrain, such as areas located close to islands and ocean ridges. These examples show that an abundance of measurements does not necessarily correspond to lower information gain and vice versa. Information gain depends not only on the geographical proximity of measurements, but also on their proximity in the parameter space and the congruence of the measurements made there. Including measurements from a region of higher information gain should lead to higher model knowledge, and hence this is more valuable compared to regions of low information gain. An experiment showing this is presented in Appendix <xref ref-type="sec" rid="App1.Ch1.S2"/>.</p>

      <fig id="Ch1.F5" specific-use="star"><label>Figure 5</label><caption><p id="d2e1494">The information gain map serves as a guide for determining optimal sampling locations, i.e. those with high information gain values. The colour scheme highlights the high information gain regions with brighter colours. Information gain does not have any units and is non-negative – <inline-formula><mml:math id="M34" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">inf</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>.</p></caption>
        <graphic xlink:href="https://gmd.copernicus.org/articles/18/2521/2025/gmd-18-2521-2025-f05.jpg"/>

      </fig>

</sec>
<sec id="Ch1.S5" sec-type="conclusions">
  <label>5</label><title>Conclusions</title>
      <p id="d2e1528">The comparison between different modelling approaches, including DNNs, kNNs, and random forests, highlights the effectiveness of each method in predicting TOC concentrations. While kNN and random forest models exhibit higher correlation coefficients and overall performance in the training dataset, the DNN outperforms them in testing data performance. This suggests a potential overfitting issue with the kNN and random forest models, where they may have become specialized in learning the training data. Nonetheless, these algorithms remain useful, especially when computational resources are limited.</p>
      <p id="d2e1531">Our DNN-based map of TOC concentrations shows elevated values in specific regions such as the Arctic and upwelling areas along continental margins. These patterns are consistent with known processes of marine primary and export production. Notably, our model that treats the shelf and deep-sea regions as separate entities captures their individual dynamics with higher accuracy and yields a better global map of TOC concentrations than a model version that simulates the entire ocean as one continuous system. It specifically avoids artifacts like unrealistically high TOC concentrations in  open-ocean regions with poor data coverage that have also been encountered in previous kNN and forest models.  The computed TOC stock for global oceans and major seas provides valuable insights into the distribution and magnitude of TOC storage. Despite significant variability in the mean TOC concentration across continental shelves, our model confirms that the majority of the TOC stock is found in deep-sea basins. Surprisingly, mean TOC concentrations in continental shelves are close to those in deep-sea sediments, suggesting complex processes at play that diminish TOC concentrations in shelf sediments.</p>
      <p id="d2e1536">The analysis of information gain highlights regions with sparse or contradicting measurements and higher uncertainty, providing guidance for future sampling efforts. It reveals that the abundance of measurements does not necessarily correspond to lower uncertainty, emphasizing the importance of considering both geographical proximity and parameter space proximity in sampling strategies.</p>
      <p id="d2e1539">In conclusion, our study contributes to a better understanding of global TOC distributions and stocks, shedding light on the complex interplay between biological, physical, and geological processes in marine sedimentary environments. The insights gained from our modelling approach can inform future research and management efforts aimed at preserving and managing marine carbon sinks.</p>
</sec>

      
      </body>
    <back><app-group>

<app id="App1.Ch1.S1">
  <label>Appendix A</label><title>Feature list</title>
      <p id="d2e1553">File names adhere to the naming conventions discussed below. The naming structure is partitioned by underscores and periods in the following order: the interface to which the gridded values refer, the quantity of values contained within the grid, the units and reference values or units (e.g. metres below sea level), the data source, the statistic calculated (if applicable), the grid pitch, and the file extension.</p>

        <table-wrap id="Taba" position="anchor"><oasis:table frame="topbot"><oasis:tgroup cols="2">
     <oasis:colspec colnum="1" colname="col1" align="justify" colwidth="4cm"/>
     <oasis:colspec colnum="2" colname="col2" align="justify" colwidth="4cm"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">SS</oasis:entry>
         <oasis:entry colname="col2" align="left">Sea surface–atmosphere interface (may also be the average of the entire water column)</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">SF</oasis:entry>
         <oasis:entry colname="col2" align="left">Seafloor–water interface (may also be denoted by GL – ground level)</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">(r50 km) – raw feature and feature averaged at a 50 km radius</oasis:entry>
         <oasis:entry colname="col2"/>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry namest="col1" nameend="col2" align="left">The units referenced are as follows. </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">KGM3</oasis:entry>
         <oasis:entry colname="col2" align="left">Kilograms per cubic metre</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">MS</oasis:entry>
         <oasis:entry colname="col2" align="left">Metres per second</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">KM</oasis:entry>
         <oasis:entry colname="col2" align="left">Kilometres</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">M_ASL</oasis:entry>
         <oasis:entry colname="col2" align="left">Metres above sea level (i.e. metres referenced to the sea level)</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">MWM2</oasis:entry>
         <oasis:entry colname="col2" align="left">Milliwatt per square metre</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">TGCYR</oasis:entry>
         <oasis:entry colname="col2" align="left">Teragrams of carbon per year</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">TGYR</oasis:entry>
         <oasis:entry colname="col2" align="left">Teragrams per year</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">MA</oasis:entry>
         <oasis:entry colname="col2" align="left">Mega-annum</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">M</oasis:entry>
         <oasis:entry colname="col2" align="left">Metres</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">MGCM2</oasis:entry>
         <oasis:entry colname="col2" align="left">Milligrams of carbon per square metre</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">DEG</oasis:entry>
         <oasis:entry colname="col2" align="left">Degrees</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1" align="left">S</oasis:entry>
         <oasis:entry colname="col2" align="left">Seconds</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d2e1715">Most of the features presented below were collected by  <xref ref-type="bibr" rid="bib1.bibx34" id="text.75"/> and <xref ref-type="bibr" rid="bib1.bibx49" id="text.76"/>. The new datasets, including the additions from this work, are available at <ext-link xlink:href="https://doi.org/10.5281/zenodo.11186224" ext-link-type="DOI">10.5281/zenodo.11186224</ext-link> <xref ref-type="bibr" rid="bib1.bibx47" id="paren.77"/>.</p>

<table-wrap id="App1.Ch1.S1.T3"><label>Table A1</label><caption><p id="d2e1736">Feature list with the descriptions and references used as input for all of the models in the paper.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="3">
     <oasis:colspec colnum="1" colname="col1" align="justify" colwidth="7cm"/>
     <oasis:colspec colnum="2" colname="col2" align="justify" colwidth="6.5cm"/>
     <oasis:colspec colnum="3" colname="col3" align="justify" colwidth="3cm"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Feature</oasis:entry>
         <oasis:entry colname="col2" align="left">Explanation</oasis:entry>
         <oasis:entry colname="col3" align="left">Data source</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">GL _COAST _FROM _LAND _IS _1.0 _ETOPO2v2.5m.nc (raw, r50 km)</oasis:entry>
         <oasis:entry colname="col2" align="left">Coastline, with a binary indicator for the presence of a coastline. This dataset is derived from ETOPO2v2, with 2 min gridded global relief data for land topography.</oasis:entry>
         <oasis:entry colname="col3" align="left">
                      <xref ref-type="bibr" rid="bib1.bibx43" id="text.78"/>
                    </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">GL _COAST _FROM _SEA _IS _1.0 _ETOPO2v2.r50km.men.5m.nc (raw, r50 km)</oasis:entry>
         <oasis:entry colname="col2" align="left">Coastline with a binary indicator for the presence of a coastline using ETOPO2v2 relief data for ocean bathymetry</oasis:entry>
         <oasis:entry colname="col3" align="left">
                      <xref ref-type="bibr" rid="bib1.bibx43" id="text.79"/>
                    </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">GL _DIST _TO _COAST _KM _ETOPO.r50km.men.5m.grd (raw, r50 km)</oasis:entry>
         <oasis:entry colname="col2" align="left">Distance of the ocean grid points to the nearest coast</oasis:entry>
         <oasis:entry colname="col3" align="left">
                      <xref ref-type="bibr" rid="bib1.bibx43" id="text.80"/>
                    </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">GL _ELEVATION _M _ASL _ETOPO2v2.r50km.men.5m.grd (raw, r50 km)</oasis:entry>
         <oasis:entry colname="col2" align="left">Elevation data from ETOPO2v2, representing heights above sea level</oasis:entry>
         <oasis:entry colname="col3" align="left">
                      <xref ref-type="bibr" rid="bib1.bibx43" id="text.81"/>
                    </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">GL _RIVERMOUTH _CO2 _TGCYR-1 _ORNL.r50km.men.5m.grd (raw, r50 km)</oasis:entry>
         <oasis:entry colname="col2" align="left">Carbon dioxide flux at river mouths, measured in teragrams of carbon per year (Tg C yr<sup>−1</sup>)</oasis:entry>
         <oasis:entry colname="col3" align="left">
                      <xref ref-type="bibr" rid="bib1.bibx37" id="text.82"/>
                    </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">GL _RIVERMOUTH _DOC _TGCYR-1 _ORNL.r50km.men.5m.grd (raw, r50 km)</oasis:entry>
         <oasis:entry colname="col2" align="left">Dissolved organic carbon flux at river mouths (Tg C yr<sup>−1</sup>)</oasis:entry>
         <oasis:entry colname="col3" align="left">
                      <xref ref-type="bibr" rid="bib1.bibx37" id="text.83"/>
                    </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">GL _RIVERMOUTH _HCO3 _TGCYR-1 _ORNL.r50km.men.5m.grd (raw, r50 km)</oasis:entry>
         <oasis:entry colname="col2" align="left">Bicarbonate <inline-formula><mml:math id="M37" display="inline"><mml:mrow class="chem"><mml:msub><mml:mi mathvariant="normal">HCO</mml:mi><mml:mn mathvariant="normal">3</mml:mn></mml:msub><mml:mo>-</mml:mo></mml:mrow></mml:math></inline-formula> flux at river mouths (Tg C yr<sup>−1</sup>)</oasis:entry>
         <oasis:entry colname="col3" align="left">
                      <xref ref-type="bibr" rid="bib1.bibx37" id="text.84"/>
                    </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">GL _RIVERMOUTH _POC _TGCYR-1 _ORNL.r50km.men.5m.grd (raw, r50 km)</oasis:entry>
         <oasis:entry colname="col2" align="left">Particulate organic carbon flux at river mouths (Tg C yr<sup>−1</sup>)</oasis:entry>
         <oasis:entry colname="col3" align="left">
                      <xref ref-type="bibr" rid="bib1.bibx37" id="text.85"/>
                    </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">GL _RIVERMOUTH _TSS _TGYR-1 _ORNL.r50km.men.5m.grd (raw, r50 km)</oasis:entry>
         <oasis:entry colname="col2" align="left">Total suspended solid flux at river mouths (50 km resolution) (Tg C yr<sup>−1</sup>). All of the riverine fluxes are binary features, with the magnitudes of the fluxes defined at the coastline and with a value of 0 for the ocean's interior.</oasis:entry>
         <oasis:entry colname="col3" align="left">
                      <xref ref-type="bibr" rid="bib1.bibx37" id="text.86"/>
                    </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">GL _TOT _SED _THICK _M _CRUST1 _NOAA.r50km.men.5m.grd (raw, r50 km)</oasis:entry>
         <oasis:entry colname="col2" align="left">Total sediment thickness in Earth's crust (m)</oasis:entry>
         <oasis:entry colname="col3" align="left">
                      <xref ref-type="bibr" rid="bib1.bibx63" id="text.87"/>
                    </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">2N2 _ocean _eot20 _modified.nc;K1 _ocean _eot20 _modified.nc;K2 _load _eot20 _modified.nc;K2 _ocean _eot20 _modified.nc;M2 _load _eot20 _modified.nc;M2 _ocean _eot20 _modified.nc;M4 _load _eot20 _modified.nc;M4 _ocean _eot20 _modified.nc;MF _load _eot20 _modified.nc;MF _ocean _eot20 _modified.nc;MM _load _eot20 _modified.nc;MM _ocean _eot20 _modified.nc;N2 _load _eot20 _modified.nc;N2 _ocean _eot20 _modified.nc;O1 _load _eot20 _modified.nc;O1 _ocean _eot20 _modified.nc;P1 _load _eot20 _modified.nc;P1 _ocean _eot20 _modified.nc;Q1 _load _eot20 _modified.nc;S1 _load _eot20 _modified.nc;S1 _ocean _eot20 _modified.nc;S2 _load _eot20 _modified.nc;S2 _ocean _eot20 _modified.nc;SA _load _eot20 _modified.nc;SA _ocean _eot20 _modified.nc; SSA_load_eot20_modified.nc; SSA_ocean_eot20_modified.nc</oasis:entry>
         <oasis:entry colname="col2" align="left"><xref ref-type="bibr" rid="bib1.bibx23" id="text.88"/> provided global atlases of both ocean and load tides, containing information about the amplitudes and phases of 17 tidal constituents (ocean and load) for the global ocean. These constituents include 2N2, J1, K1, K2, M2, M4, MF, MM, N2, O1, P1, Q1, S1, S2, SA, SSA, and T2, which extend across the entire global ocean and range from 66° S to 66° N. For higher latitudes, the FES2014b model is used to fill in the gaps. Eleven satellite altimetry missions contribute to this model.</oasis:entry>
         <oasis:entry colname="col3" align="left">
                      <xref ref-type="bibr" rid="bib1.bibx23" id="text.89"/>
                    </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">ChlorSummerMean.nc</oasis:entry>
         <oasis:entry colname="col2" align="left">Average chlorophyll-<inline-formula><mml:math id="M41" display="inline"><mml:mi>a</mml:mi></mml:math></inline-formula> concentration during summer (June to November), collected from July 2002 till July 2022</oasis:entry>
         <oasis:entry colname="col3" align="left">
                      <xref ref-type="bibr" rid="bib1.bibx41" id="text.90"/>
                    </oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1" align="left">ChlorWinterMean.nc</oasis:entry>
         <oasis:entry colname="col2" align="left">Average chlorophyll-<inline-formula><mml:math id="M42" display="inline"><mml:mi>a</mml:mi></mml:math></inline-formula> concentration during winter (December to May), collected from July 2002 till July 2022</oasis:entry>
         <oasis:entry colname="col3" align="left">
                      <xref ref-type="bibr" rid="bib1.bibx41" id="text.91"/>
                    </oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

<table-wrap id="App1.Ch1.S1.T4"><label>Table A1</label><caption><p id="d2e2059">Continued.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="3">
     <oasis:colspec colnum="1" colname="col1" align="justify" colwidth="7cm"/>
     <oasis:colspec colnum="2" colname="col2" align="justify" colwidth="6.5cm"/>
     <oasis:colspec colnum="3" colname="col3" align="justify" colwidth="3cm"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Feature</oasis:entry>
         <oasis:entry colname="col2" align="left">Explanation</oasis:entry>
         <oasis:entry colname="col3" align="left">Data source</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">DERIVATIVE _GL _ELEVATION _M _ASL _ETOPO2v2.5.nc</oasis:entry>
         <oasis:entry colname="col2" align="left">Slope from ETOPO2v2.5 data</oasis:entry>
         <oasis:entry colname="col3" align="left"/>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">GL _HEATFLUX _MWM2 _Becker.5m.nc</oasis:entry>
         <oasis:entry colname="col2" align="left">Oceanic heat flux data (exchange of heat energy between the ocean surface and the atmosphere) in megawatts per square metre (MW m<sup>−2</sup>)</oasis:entry>
         <oasis:entry colname="col3" align="left">
                    <xref ref-type="bibr" rid="bib1.bibx5" id="text.92"/>
                  </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">GL _LAND _IS _1.0 _ETOPO2v2.5m.nc</oasis:entry>
         <oasis:entry colname="col2" align="left">Land mask data</oasis:entry>
         <oasis:entry colname="col3" align="left">
                    <xref ref-type="bibr" rid="bib1.bibx43" id="text.93"/>
                  </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">POROSITY _global _prediction.grd</oasis:entry>
         <oasis:entry colname="col2" align="left">Global prediction map for porosity of surface sediments using a random forest method</oasis:entry>
         <oasis:entry colname="col3" align="left">
                    <xref ref-type="bibr" rid="bib1.bibx39" id="text.94"/>
                  </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">SF _ACTIVE _SEAMOUNTS _KIM.r10km.wct.5m.grd</oasis:entry>
         <oasis:entry colname="col2" align="left">Volcanically active seamount location data at 10 km resolution</oasis:entry>
         <oasis:entry colname="col3" align="left">
                    <xref ref-type="bibr" rid="bib1.bibx28" id="text.95"/>
                  </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">SF _AVG _SEA _DENSITY _KGM3 _DECADAL _MEAN _woa13x.5m.grd (raw, r50 km)</oasis:entry>
         <oasis:entry colname="col2" align="left">Sea density (kg m<sup>−3</sup>), averaged</oasis:entry>
         <oasis:entry colname="col3" align="left">
                    <xref ref-type="bibr" rid="bib1.bibx9" id="text.96"/>
                  </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">SF _COASTLINE _IS _1.0.5m.nc</oasis:entry>
         <oasis:entry colname="col2" align="left">Coastline data from Global Land One-kilometer Base Elevation (GLOBE)</oasis:entry>
         <oasis:entry colname="col3" align="left">
                    <xref ref-type="bibr" rid="bib1.bibx43" id="text.97"/>
                  </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">SF _CURRENT _EAST _MS _2012 _12 _HYCOMx.5m.grd;SF _CURRENT _NORTH _MS _2012 _12 _HYCOMx.5m.grd;SF _CURRENT _MAG _MS _2012 _12 _HYCOMx.5m.grd (raw, r50 km)</oasis:entry>
         <oasis:entry colname="col2" align="left">Ocean bottom current data for the east–west component, north–south component, and total magnitude using the HYCOM model (m s<sup>−1</sup>). The data are provided at 1/12° resolution. The dataset has a time range from 1 August 1995 to 31 December  2012, temporally averaged.</oasis:entry>
         <oasis:entry colname="col3" align="left">
                    <xref ref-type="bibr" rid="bib1.bibx59" id="text.98"/>
                  </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">SF _GRAINSIZE _D16 _MM _NGDC.5m.nc;SF _GRAINSIZE _D50 _MM _NGDC.5m.nc;SF _GRAINSIZE _D84 _MM _NGDC.5m.nc</oasis:entry>
         <oasis:entry colname="col2" align="left">Grain-size data with the 16th percentile (D16), median (D50), and 84th percentile (D84)</oasis:entry>
         <oasis:entry colname="col3" align="left">
                    <xref ref-type="bibr" rid="bib1.bibx42" id="text.99"/>
                  </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">SF _SEA _BULKMODULUS _MPA _DECADAL _MEAN _woa13x.5m.nc</oasis:entry>
         <oasis:entry colname="col2" align="left">Sea bulk modulus (MPa) averaged over 6 decades from the years 1955 to 2012. The sea bulk modulus is an important thermodynamic property and is a measure of resistance to the compressibility of a fluid. It is calculated from the International Equation of State of Seawater of <xref ref-type="bibr" rid="bib1.bibx26" id="text.100"/>.</oasis:entry>
         <oasis:entry colname="col3" align="left">
                    <xref ref-type="bibr" rid="bib1.bibx9" id="text.101"/>
                  </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">SF _SEA _CONDUCTIVITY _SM _DECADAL _MEAN _woa13v2x.5m.grd (raw, r50 km)</oasis:entry>
         <oasis:entry colname="col2" align="left">Average conductivity of seawater (dissolved ions) at the sea surface over 6 decades from the years 1955 to 2012. The units are Siemens per metre (S m<sup>−1</sup>).</oasis:entry>
         <oasis:entry colname="col3" align="left">
                    <xref ref-type="bibr" rid="bib1.bibx9" id="text.102"/>
                  </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">SF _SEA _OXYGEN _MLL _DECADAL _MEAN _woa13v2x.5m.grd (raw, r50 km)</oasis:entry>
         <oasis:entry colname="col2" align="left">Average dissolved oxygen concentration in seawater in millilitres per litre over a decadal mean</oasis:entry>
         <oasis:entry colname="col3" align="left">
                    <xref ref-type="bibr" rid="bib1.bibx9" id="text.103"/>
                  </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">SF _SEA _OXYGEN _PCTSAT _DECADAL _MEAN _woa13v2x.5m.grd (raw, r50 km)</oasis:entry>
         <oasis:entry colname="col2" align="left">Oxygen concentration in seawater percentage saturation averaged over 6 decades from the years 1955 to 2012</oasis:entry>
         <oasis:entry colname="col3" align="left">
                    <xref ref-type="bibr" rid="bib1.bibx9" id="text.104"/>
                  </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">SF _SEA _PRESSURE _MPA _DECADAL _MEAN _woa13x.5m.nc</oasis:entry>
         <oasis:entry colname="col2" align="left">Seawater pressure (MPa) averaged over 6 decades from the years 1955 to 2012</oasis:entry>
         <oasis:entry colname="col3" align="left">
                    <xref ref-type="bibr" rid="bib1.bibx9" id="text.105"/>
                  </oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1" align="left">SF _SEA _SALINITY _PSU _DECADAL _MEAN _woa13v2x.5m.nc</oasis:entry>
         <oasis:entry colname="col2" align="left">Seawater salinity in practical salinity units averaged over 6 decades from the years 1955 to 2012</oasis:entry>
         <oasis:entry colname="col3" align="left">
                    <xref ref-type="bibr" rid="bib1.bibx9" id="text.106"/>
                  </oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

<table-wrap id="App1.Ch1.S1.T5"><label>Table A1</label><caption><p id="d2e2370">Continued.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="3">
     <oasis:colspec colnum="1" colname="col1" align="justify" colwidth="6cm"/>
     <oasis:colspec colnum="2" colname="col2" align="justify" colwidth="8cm"/>
     <oasis:colspec colnum="3" colname="col3" align="justify" colwidth="1.5cm"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Feature</oasis:entry>
         <oasis:entry colname="col2" align="left">Explanation</oasis:entry>
         <oasis:entry colname="col3" align="left">Data source</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">SF _SEA _SEA _OXYGEN _UTILIZATION _MOLM3 _DECADAL _MEAN _woa13v2x.5m.grd (raw, r50 km)</oasis:entry>
         <oasis:entry colname="col2" align="left">Oxygen concentration in seawater utilization (mol m<sup>−3</sup>)  averaged over 6 decades from the years 1955 to 2012</oasis:entry>
         <oasis:entry colname="col3" align="left">
                    <xref ref-type="bibr" rid="bib1.bibx9" id="text.107"/>
                  </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">SF _SEA _TEMPERATURE _C _DECADAL _MEAN _woa13v2x.5m.grd (raw, r50 km)</oasis:entry>
         <oasis:entry colname="col2" align="left">Seawater temperature in degrees Celsius averaged over 6 decades from the years 1955 to 2012</oasis:entry>
         <oasis:entry colname="col3" align="left">
                    <xref ref-type="bibr" rid="bib1.bibx9" id="text.108"/>
                  </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">SL _GEOID _M _ABOVE _WGS84 _NGA _egm2008.5m.grd</oasis:entry>
         <oasis:entry colname="col2" align="left">Height of the geoid above the WGS84 reference ellipsoid (m) and referenced to the National Geospatial-Intelligence Agency (NGA)</oasis:entry>
         <oasis:entry colname="col3" align="left">
                    <xref ref-type="bibr" rid="bib1.bibx48" id="text.109"/>
                  </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">SS _BIOMASS _BACTERIA _LOG10 _MGCM2 _WEI2010x.5m.grd (raw, r50 km); SS _BIOMASS _FISH _LOG10 _MGCM2 _WEI2010x.5m.grd (raw, r50 km); SS _BIOMASS _INVERTEBRATE _LOG10 _MGCM2 _WEI2010x.5m.grd (raw, r50 km); SS _BIOMASS _MACROFAUNA _LOG10 _MGCM2 _WEI2010x.5m.grd (raw, r50 km); SS _BIOMASS _MEGAFAUNA _LOG10 _MGCM2 _WEI2010x.5m.grd (raw, r50 km); SS _BIOMASS _MEIOFAUNA _LOG10 _MGCM2 _WEI2010x.5m.grd (raw, r50 km); SS _BIOMASS _TOTAL _LOG10 _MGCM2 _WEI2010x.5m.grd (raw, r50 km)</oasis:entry>
         <oasis:entry colname="col2" align="left">Distribution of mean biomass predictions for (a) bacteria, (b) fish, (c) invertebrates, (d) macrofauna, (e) megafauna, and (f) meiofauna. The mean biomass was computed using a random forest algorithm. The total biomass was combined from predictions of bacteria, meiofauna, macrofauna, and megafauna biomasses. Predictions were smoothed by inverse distance weighting interpolation to 0.1° resolution and displayed on a logarithm scale (base 10), which was then converted into 5 arcmin grids by <xref ref-type="bibr" rid="bib1.bibx33" id="text.110"/>.</oasis:entry>
         <oasis:entry colname="col3" align="left">
                    <xref ref-type="bibr" rid="bib1.bibx62" id="text.111"/>
                  </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">SS _CHLOROPHYLL _LOG _MG _M3 _MODIS _Aqua _MISSION _MEANx.5m.grd (raw, r50 km); SS _PIC _LOG _MOL _M3-1 _MODIS _Aqua _MISSION _MEANx.5m.grd (raw, r50 km); SS _POC _LOG _MOL _M3-1 _MODIS _Aqua _MISSION _MEANx.5m.grd (raw, r50 km)</oasis:entry>
         <oasis:entry colname="col2" align="left">The Moderate Resolution Imaging Spectroradiometer (MODIS) is a 36-band spectroradiometer measuring visible and infrared radiation and obtaining data that are used to derive the near-surface concentration of chlorophyll <inline-formula><mml:math id="M48" display="inline"><mml:mi>a</mml:mi></mml:math></inline-formula> (chlor_a) (mg m<sup>−3</sup>). This is calculated using an empirical relationship derived from in situ measurements of chlor_a, concentrations of particulate organic carbon (POC) and particulate inorganic carbon (PIC) (i.e. calcium carbonate or calcite), and blue- to green-band ratios of in situ remote sensing reflectances (Rrs).</oasis:entry>
         <oasis:entry colname="col3" align="left">
                    <xref ref-type="bibr" rid="bib1.bibx41" id="text.112"/>
                  </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">SS _CORIOLIS.5m.nc</oasis:entry>
         <oasis:entry colname="col2" align="left">Coriolis data, generated using empirical means</oasis:entry>
         <oasis:entry colname="col3" align="left">
                    <xref ref-type="bibr" rid="bib1.bibx34" id="text.113"/>
                  </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">SS _DENSITY _KGM-3 _SACD _Aquarius _MISSION _MEANx.5m.grd</oasis:entry>
         <oasis:entry colname="col2" align="left">The Aquarius/SAC-D satellite mission, launched on 10 June 2011, was a joint venture between NASA and the Argentinian Space Agency (CONAE). The mission featured the sea surface salinity sensor Aquarius and was the first mission with the primary goal of measuring sea surface salinity (SSS) from space. The monthly maps of sea surface density are derived from Aquarius sea surface salinity and ancillary sea surface temperature. The time period used to average is between 25 August 2011 and 7 June 2015.</oasis:entry>
         <oasis:entry colname="col3" align="left">
                    <xref ref-type="bibr" rid="bib1.bibx40" id="text.114"/>
                  </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">SS _GEOID _ANOMALY _NGA _egm2008.5m.nc (raw, r50 km)</oasis:entry>
         <oasis:entry colname="col2" align="left">The regional Free-air and Bouguer gravity anomaly grids (averaged over 2.5 arcmin by 2.5 arcmin) are computed at BGI (Bureau Gravimetrique International) from the EGM2008 spherical harmonic coefficients.</oasis:entry>
         <oasis:entry colname="col3" align="left">
                    <xref ref-type="bibr" rid="bib1.bibx48" id="text.115"/>
                  </oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1" align="left">SS _MIXED _LAYER _DEPTH _MAX _M _Goyetx.5m.grd (raw, r50 km); SS _MIXED _LAYER _DEPTH _MIN _M _Goyetx.5m.grd (raw, r50 km)</oasis:entry>
         <oasis:entry colname="col2" align="left">This shows the geographical distribution of the maximum and minimum depths (m) of the mixed layer. The observations are in the time span between March 1995 and February 1996.</oasis:entry>
         <oasis:entry colname="col3" align="left">
                    <xref ref-type="bibr" rid="bib1.bibx21" id="text.116"/>
                  </oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

<table-wrap id="App1.Ch1.S1.T6"><label>Table A1</label><caption><p id="d2e2578">Continued.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="3">
     <oasis:colspec colnum="1" colname="col1" align="justify" colwidth="7cm"/>
     <oasis:colspec colnum="2" colname="col2" align="justify" colwidth="7cm"/>
     <oasis:colspec colnum="3" colname="col3" align="justify" colwidth="2cm"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Feature</oasis:entry>
         <oasis:entry colname="col2" align="left">Explanation</oasis:entry>
         <oasis:entry colname="col3" align="left">Data source</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">SS _PHOTO _AVAIL _RAD _EINSTEIN _M-2 _DAY _SNPP _VIIRS _MISSION _MEANx.5m.grd (raw, r50 km); SS _PHYTO _ABSORPTION _443NM _M-1 _SNPP _VIIRS _MISSION _MEANx.5m.grd</oasis:entry>
         <oasis:entry colname="col2" align="left">Daily average photosynthetically available radiation (PAR) at the ocean surface (Einstein m<sup>−2</sup> d<sup>−1</sup>). The Visible Infrared Imaging Radiometer Suite (VIIRS) on the Suomi National Polar-orbiting Partnership (SNPP) was developed for global ocean colour products. PAR is defined as the quantum energy flux from the Sun in the 400–700 nm range. For ocean colour applications, PAR is a common input used in modelling marine primary productivity. An average of the sensors and the 443 nm wavelength maps is used as a feature.</oasis:entry>
         <oasis:entry colname="col3" align="left">
                    <xref ref-type="bibr" rid="bib1.bibx41" id="text.117"/>
                  </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">SS _WAVE _DIRECTION _DEG _2012 _12 _WAVEWATCH3x.5m.grd (raw, r50 km); SS _WAVE _HEIGHT _M _2012 _12 _WAVEWATCH3x.5m.grd (raw, r50 km); SS _WAVE _PERIOD _S _2012 _12 _WAVEWATCH3x.5m.grd (raw, r50 km)</oasis:entry>
         <oasis:entry colname="col2" align="left">Mean wave direction (°), wave height (m), and wave period (<inline-formula><mml:math id="M52" display="inline"><mml:mi>s</mml:mi></mml:math></inline-formula>). The data are provided at 1/12° resolution. The dataset is averaged over the time range from 1 August  1995 to 31 December  2012. The features are based on the third-generation wave model WAVEWATCH III®.</oasis:entry>
         <oasis:entry colname="col3" align="left">
                    <xref ref-type="bibr" rid="bib1.bibx59" id="text.118"/>
                  </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">SS _WINDSPEED _MS-1 _SACD _Aquarius _MISSION _MEANx.5m.grd (raw, r50 km)</oasis:entry>
         <oasis:entry colname="col2" align="left">Mean wind speed (m s<sup>−1</sup>) from the Aquarius/SAC-D satellite mission. The time period used for averaging is between 25 August 2011 and 7 June 2015.</oasis:entry>
         <oasis:entry colname="col3" align="left">
                    <xref ref-type="bibr" rid="bib1.bibx40" id="text.119"/>
                  </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">TOU _Jorgenson2022.nc</oasis:entry>
         <oasis:entry colname="col2" align="left">Global map of the total oxygen uptake (TOU) of the seabed</oasis:entry>
         <oasis:entry colname="col3" align="left">
                    <xref ref-type="bibr" rid="bib1.bibx27" id="text.120"/>
                  </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">litho_maps_type1_.nc</oasis:entry>
         <oasis:entry colname="col2" align="left">Lithology map: mudflats binary map (median grain size <inline-formula><mml:math id="M54" display="inline"><mml:mo>&lt;</mml:mo></mml:math></inline-formula> 0.05 mm)</oasis:entry>
         <oasis:entry colname="col3" align="left">
                    <xref ref-type="bibr" rid="bib1.bibx20" id="text.121"/>
                  </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">litho_maps_type2_.nc</oasis:entry>
         <oasis:entry colname="col2" align="left">Lithology map: fine-sand binary map (median grain size 0.05–0.5 mm)</oasis:entry>
         <oasis:entry colname="col3" align="left">
                    <xref ref-type="bibr" rid="bib1.bibx20" id="text.122"/>
                  </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">litho_maps_type3_.nc</oasis:entry>
         <oasis:entry colname="col2" align="left">Lithology map: sand binary map (median grain size 0.5–2 mm)</oasis:entry>
         <oasis:entry colname="col3" align="left">
                    <xref ref-type="bibr" rid="bib1.bibx20" id="text.123"/>
                  </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">litho_maps_type4_.nc</oasis:entry>
         <oasis:entry colname="col2" align="left">Lithology map: clay binary map (median grain size <inline-formula><mml:math id="M55" display="inline"><mml:mo>&lt;</mml:mo></mml:math></inline-formula> 0.01 mm)</oasis:entry>
         <oasis:entry colname="col3" align="left">
                    <xref ref-type="bibr" rid="bib1.bibx20" id="text.124"/>
                  </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">litho_maps_type5_.nc</oasis:entry>
         <oasis:entry colname="col2" align="left">Lithology map: gravel and stone binary map (median grain size <inline-formula><mml:math id="M56" display="inline"><mml:mo>&gt;</mml:mo></mml:math></inline-formula> 2 mm)</oasis:entry>
         <oasis:entry colname="col3" align="left">
                    <xref ref-type="bibr" rid="bib1.bibx20" id="text.125"/>
                  </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">litho_maps_type6_.nc</oasis:entry>
         <oasis:entry colname="col2" align="left">Lithology map: bedrock binary map</oasis:entry>
         <oasis:entry colname="col3" align="left">
                    <xref ref-type="bibr" rid="bib1.bibx20" id="text.126"/>
                  </oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1" align="left">lithology _grain _size _global _8.nc</oasis:entry>
         <oasis:entry colname="col2" align="left">Global seabed sediment map with 24 different classes or types of sediments based on a logarithmic progression of the median grain size</oasis:entry>
         <oasis:entry colname="col3" align="left">
                    <xref ref-type="bibr" rid="bib1.bibx20" id="text.127"/>
                  </oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

</app>

<app id="App1.Ch1.S2">
  <label>Appendix B</label><title>Information gain</title>
      <p id="d2e2848">In this paper, KL divergence, also known as information gain or relative entropy, has been used to quantify model uncertainty. As <xref ref-type="bibr" rid="bib1.bibx50" id="text.128"/> pointed out, in the absence of observational information, the amount of information can be taken to be numerically equal to the amount of uncertainty concerning the model prediction. The mathematical derivation of KL divergence against the theoretical background of information theory <xref ref-type="bibr" rid="bib1.bibx56" id="paren.129"/> is presented below. The information entropy of a random variable <inline-formula><mml:math id="M57" display="inline"><mml:mi>X</mml:mi></mml:math></inline-formula> with a probability distribution <inline-formula><mml:math id="M58" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula> is represented as
          <disp-formula id="App1.Ch1.S2.E2" content-type="numbered"><label>B1</label><mml:math id="M59" display="block"><mml:mrow><mml:mi>H</mml:mi><mml:mo>(</mml:mo><mml:mi>P</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:munder><mml:mo movablelimits="false">∑</mml:mo><mml:mi>i</mml:mi></mml:munder><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mi>log⁡</mml:mi><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula></p>

      <fig id="App1.Ch1.S2.F6" specific-use="star"><label>Figure B1</label><caption><p id="d2e2927">Top left: difference in the prediction of the TOC concentration between <inline-formula><mml:math id="M60" display="inline"><mml:mrow><mml:mi mathvariant="italic">ϕ</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold">W</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold-italic">b</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M61" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">ϕ</mml:mi><mml:mrow><mml:mi mathvariant="normal">wh</mml:mi><mml:mo>+</mml:mo><mml:mi mathvariant="normal">low</mml:mi><mml:mi mathvariant="normal">_</mml:mi><mml:mi mathvariant="normal">ig</mml:mi></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold">W</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold-italic">b</mml:mi><mml:msub><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>. Top right: difference in the prediction of the TOC concentration between <inline-formula><mml:math id="M62" display="inline"><mml:mrow><mml:mi mathvariant="italic">ϕ</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold">W</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold-italic">b</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M63" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">ϕ</mml:mi><mml:mrow><mml:mi mathvariant="normal">wh</mml:mi><mml:mo>+</mml:mo><mml:mi mathvariant="normal">high</mml:mi><mml:mi mathvariant="normal">_</mml:mi><mml:mi mathvariant="normal">ig</mml:mi></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold">W</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold-italic">b</mml:mi><mml:msub><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>. Brighter colours (shades of yellow) show higher differences, and darker colours (shades of blue) show lower differences. Bottom left (red box): zoomed-in version of the equatorial Pacific region in panels <bold>(a)</bold> and <bold>(c)</bold>. Bottom right (green box): zoomed-in versions of the Caspian and Black seas in panels <bold>(b)</bold> and <bold>(d)</bold>.</p></caption>
        <graphic xlink:href="https://gmd.copernicus.org/articles/18/2521/2025/gmd-18-2521-2025-f06.jpg"/>

      </fig>

      <p id="d2e3067">The <xref ref-type="bibr" rid="bib1.bibx56" id="text.130"/> definition of entropy determines the minimum channel capacity required to reliably transmit the information as encoded binary digits. Usually, the true distribution <inline-formula><mml:math id="M64" display="inline"><mml:mrow><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi>X</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> denotes observed data, measurements, or an exact probability distribution. Here, <inline-formula><mml:math id="M65" display="inline"><mml:mrow><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi>X</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is constructed using a normal distribution with a mean value equal to Monte Carlo dropout prediction and a standard deviation of 0.05 TOC %, which arises from both technical handling and the precision of the weighting tool <xref ref-type="bibr" rid="bib1.bibx44" id="paren.131"/>. The predicted distribution <inline-formula><mml:math id="M66" display="inline"><mml:mrow><mml:mi>Q</mml:mi><mml:mo>(</mml:mo><mml:mi>X</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is derived from the Monte Carlo dropout prediction ensemble. The measure <inline-formula><mml:math id="M67" display="inline"><mml:mrow><mml:mi>Q</mml:mi><mml:mo>(</mml:mo><mml:mi>X</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> typically represents a theoretical framework, a model, a description, or an approximation of <inline-formula><mml:math id="M68" display="inline"><mml:mrow><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi>X</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. The cross-entropy between <inline-formula><mml:math id="M69" display="inline"><mml:mrow><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi>X</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M70" display="inline"><mml:mrow><mml:mi>Q</mml:mi><mml:mo>(</mml:mo><mml:mi>X</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> measures the average number of binary digits to represent an event from <inline-formula><mml:math id="M71" display="inline"><mml:mrow><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi>X</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> by <inline-formula><mml:math id="M72" display="inline"><mml:mrow><mml:mi>Q</mml:mi><mml:mo>(</mml:mo><mml:mi>X</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. It is represented as
          <disp-formula id="App1.Ch1.S2.E3" content-type="numbered"><label>B2</label><mml:math id="M73" display="block"><mml:mrow><mml:mi>H</mml:mi><mml:mo>(</mml:mo><mml:mi>P</mml:mi><mml:mo>,</mml:mo><mml:mi>Q</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:munder><mml:mo movablelimits="false">∑</mml:mo><mml:mi>i</mml:mi></mml:munder><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mi>log⁡</mml:mi><mml:mi>Q</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula></p>

<table-wrap id="App1.Ch1.S2.T7" specific-use="star"><label>Table B1</label><caption><p id="d2e3261">Performance metrics of models trained on different subsets of data based on information gain for different splits of data (seeds): Pearson CC, <inline-formula><mml:math id="M74" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>, and MSE for predicted values vs. observed labels for the training and testing data. The <inline-formula><mml:math id="M75" display="inline"><mml:mrow><mml:mi mathvariant="normal">train</mml:mi><mml:mo>:</mml:mo><mml:mi mathvariant="normal">test</mml:mi></mml:mrow></mml:math></inline-formula> data ratio is <inline-formula><mml:math id="M76" display="inline"><mml:mrow><mml:mn mathvariant="normal">85</mml:mn><mml:mo>:</mml:mo><mml:mn mathvariant="normal">15</mml:mn></mml:mrow></mml:math></inline-formula>. The better-performing dataset, based on the metric, is highlighted in bold. It shows that the performance on the test dataset, or the generalisation is mostly higher when using the dataset with higher information gain.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="7">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="center"/>
     <oasis:colspec colnum="3" colname="col3" align="center"/>
     <oasis:colspec colnum="4" colname="col4" align="center" colsep="1"/>
     <oasis:colspec colnum="5" colname="col5" align="center"/>
     <oasis:colspec colnum="6" colname="col6" align="center"/>
     <oasis:colspec colnum="7" colname="col7" align="center"/>
     <oasis:thead>
       <oasis:row>
         <oasis:entry colname="col1">DNN model with different training datasets</oasis:entry>
         <oasis:entry rowsep="1" namest="col2" nameend="col4" colsep="1">Training data </oasis:entry>
         <oasis:entry rowsep="1" namest="col5" nameend="col7">Testing data (15 % of all of the data) </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">Pearson CC</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M77" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">MSE</oasis:entry>
         <oasis:entry colname="col5">Pearson CC</oasis:entry>
         <oasis:entry colname="col6"><inline-formula><mml:math id="M78" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col7">MSE</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1"><inline-formula><mml:math id="M79" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">ϕ</mml:mi><mml:mrow><mml:mi mathvariant="normal">wh</mml:mi><mml:mo>+</mml:mo><mml:mi mathvariant="normal">low</mml:mi><mml:mi mathvariant="normal">_</mml:mi><mml:mi mathvariant="normal">ig</mml:mi></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold">W</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold-italic">b</mml:mi><mml:msub><mml:mo>)</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col2">0.918</oasis:entry>
         <oasis:entry colname="col3"><bold>0.840</bold></oasis:entry>
         <oasis:entry colname="col4"><bold>0.535</bold></oasis:entry>
         <oasis:entry colname="col5">0.814</oasis:entry>
         <oasis:entry colname="col6">0.641</oasis:entry>
         <oasis:entry colname="col7">0.692</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"><inline-formula><mml:math id="M80" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">ϕ</mml:mi><mml:mrow><mml:mi mathvariant="normal">wh</mml:mi><mml:mo>+</mml:mo><mml:mi mathvariant="normal">high</mml:mi><mml:mi mathvariant="normal">_</mml:mi><mml:mi mathvariant="normal">ig</mml:mi></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold">W</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold-italic">b</mml:mi><mml:msub><mml:mo>)</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col2">0.918</oasis:entry>
         <oasis:entry colname="col3">0.833</oasis:entry>
         <oasis:entry colname="col4">0.586</oasis:entry>
         <oasis:entry colname="col5"><bold>0.825</bold></oasis:entry>
         <oasis:entry colname="col6"><bold>0.679</bold></oasis:entry>
         <oasis:entry colname="col7"><bold>0.640</bold></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"><inline-formula><mml:math id="M81" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">ϕ</mml:mi><mml:mrow><mml:mi mathvariant="normal">wh</mml:mi><mml:mo>+</mml:mo><mml:mi mathvariant="normal">low</mml:mi><mml:mi mathvariant="normal">_</mml:mi><mml:mi mathvariant="normal">ig</mml:mi></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold">W</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold-italic">b</mml:mi><mml:msub><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col2"><bold>0.927</bold></oasis:entry>
         <oasis:entry colname="col3"><bold>0.856</bold></oasis:entry>
         <oasis:entry colname="col4"><bold>0.405</bold></oasis:entry>
         <oasis:entry colname="col5"><bold>0.887</bold></oasis:entry>
         <oasis:entry colname="col6"><bold>0.784</bold></oasis:entry>
         <oasis:entry colname="col7"><bold>0.525</bold></oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"><inline-formula><mml:math id="M82" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">ϕ</mml:mi><mml:mrow><mml:mi mathvariant="normal">wh</mml:mi><mml:mo>+</mml:mo><mml:mi mathvariant="normal">high</mml:mi><mml:mi mathvariant="normal">_</mml:mi><mml:mi mathvariant="normal">ig</mml:mi></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold">W</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold-italic">b</mml:mi><mml:msub><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col2">0.924</oasis:entry>
         <oasis:entry colname="col3">0.843</oasis:entry>
         <oasis:entry colname="col4">0.470</oasis:entry>
         <oasis:entry colname="col5">0.877</oasis:entry>
         <oasis:entry colname="col6">0.761</oasis:entry>
         <oasis:entry colname="col7">0.593</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"><inline-formula><mml:math id="M83" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">ϕ</mml:mi><mml:mrow><mml:mi mathvariant="normal">wh</mml:mi><mml:mo>+</mml:mo><mml:mi mathvariant="normal">low</mml:mi><mml:mi mathvariant="normal">_</mml:mi><mml:mi mathvariant="normal">ig</mml:mi></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold">W</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold-italic">b</mml:mi><mml:msub><mml:mo>)</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col2"><bold>0.935</bold></oasis:entry>
         <oasis:entry colname="col3"><bold>0.864</bold></oasis:entry>
         <oasis:entry colname="col4"><bold>0.410</bold></oasis:entry>
         <oasis:entry colname="col5">0.819</oasis:entry>
         <oasis:entry colname="col6">0.668</oasis:entry>
         <oasis:entry colname="col7">0.575</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"><inline-formula><mml:math id="M84" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">ϕ</mml:mi><mml:mrow><mml:mi mathvariant="normal">wh</mml:mi><mml:mo>+</mml:mo><mml:mi mathvariant="normal">high</mml:mi><mml:mi mathvariant="normal">_</mml:mi><mml:mi mathvariant="normal">ig</mml:mi></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold">W</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold-italic">b</mml:mi><mml:msub><mml:mo>)</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col2">0.932</oasis:entry>
         <oasis:entry colname="col3">0.853</oasis:entry>
         <oasis:entry colname="col4">0.428</oasis:entry>
         <oasis:entry colname="col5"><bold>0.855</bold></oasis:entry>
         <oasis:entry colname="col6"><bold>0.728</bold></oasis:entry>
         <oasis:entry colname="col7"><bold>0.475</bold></oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d2e3757">The information gain that measures the difference between the cross-entropy (Eq. <xref ref-type="disp-formula" rid="App1.Ch1.S2.E3"/>) and the entropy (Eq. <xref ref-type="disp-formula" rid="App1.Ch1.S2.E2"/>) is represented as <inline-formula><mml:math id="M85" display="inline"><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mtext>KL</mml:mtext></mml:msub><mml:mo>(</mml:mo><mml:mi>P</mml:mi><mml:mo>‖</mml:mo><mml:mi>Q</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>.
          <disp-formula id="App1.Ch1.S2.E4" content-type="numbered"><label>B3</label><mml:math id="M86" display="block"><mml:mtable rowspacing="0.2ex" class="split" displaystyle="true" columnalign="right left"><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mtext>KL</mml:mtext></mml:msub><mml:mo>(</mml:mo><mml:mi>P</mml:mi><mml:mo>‖</mml:mo><mml:mi>Q</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mo>=</mml:mo><mml:mi>H</mml:mi><mml:mo>(</mml:mo><mml:mi>P</mml:mi><mml:mo>,</mml:mo><mml:mi>Q</mml:mi><mml:mo>)</mml:mo><mml:mo>-</mml:mo><mml:mi>H</mml:mi><mml:mo>(</mml:mo><mml:mi>P</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mo>=</mml:mo><mml:munder><mml:mo movablelimits="false">∑</mml:mo><mml:mi>i</mml:mi></mml:munder><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mi>log⁡</mml:mi><mml:mfenced open="(" close=")"><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mi>Q</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mfrac></mml:mstyle></mml:mfenced></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
      <p id="d2e3886"><inline-formula><mml:math id="M87" display="inline"><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mtext>KL</mml:mtext></mml:msub><mml:mo>(</mml:mo><mml:mi>P</mml:mi><mml:mo>‖</mml:mo><mml:mi>Q</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is always non-negative and remains well-defined for continuous distributions. To obtain the continuous distribution for the predicted distribution <inline-formula><mml:math id="M88" display="inline"><mml:mrow><mml:mi>Q</mml:mi><mml:mo>(</mml:mo><mml:mi>X</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, the prediction ensemble is binned into histograms to obtain an approximate probability density function (PDF). This PDF is then modelled using curve-fitting techniques typically fitted to a Gaussian distribution (Algorithm <xref ref-type="other" rid="App1.Ch1.S7.Prog2"/>). <inline-formula><mml:math id="M89" display="inline"><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mtext>KL</mml:mtext></mml:msub><mml:mo>(</mml:mo><mml:mi>P</mml:mi><mml:mo>‖</mml:mo><mml:mi>Q</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is calculated globally for each prediction and plotted on the information gain map.</p>
      <p id="d2e3946">In supervised machine learning, a model's predictive performance is usually determined by withholding a test dataset during the training phase and comparing the final model outputs to these known values. Such a procedure is not possible when evaluating the performance of information gain: firstly, the concept of a ground truth for the information gain values does not exist. Secondly, we aim to measure the effect that data point selection guided by information gain has on the model output, not the information gain itself. Thus, in order to explore the effect that information gain has on data sampling and model refinement, we devised the following experiment: a DNN model with the same parameters as the original one was trained while withholding one-third of the original training dataset: <inline-formula><mml:math id="M90" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">ϕ</mml:mi><mml:mi mathvariant="normal">wh</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold">W</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold-italic">b</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. Afterwards, this model was used to calculate the information gain for each point in the withheld data. These additional data points were sorted according to their information gain values and divided into two subsets of equal size. Each of these subsets was used along with the initial two-thirds to train two new DNN models: one with added high information gain data points (<inline-formula><mml:math id="M91" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">ϕ</mml:mi><mml:mrow><mml:mi mathvariant="normal">wh</mml:mi><mml:mo>+</mml:mo><mml:mi mathvariant="normal">high</mml:mi><mml:mi mathvariant="normal">_</mml:mi><mml:mi mathvariant="normal">ig</mml:mi></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold">W</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold-italic">b</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>) and one with added low information gain data points (<inline-formula><mml:math id="M92" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">ϕ</mml:mi><mml:mrow><mml:mi mathvariant="normal">wh</mml:mi><mml:mo>+</mml:mo><mml:mi mathvariant="normal">low</mml:mi><mml:mi mathvariant="normal">_</mml:mi><mml:mi mathvariant="normal">ig</mml:mi></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold">W</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold-italic">b</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>). To validate the entirety of the training data, the process was repeated two more times, withholding a different third of the dataset each time.</p>
      <p id="d2e4042">In two of the three executions, the (test) performance of <inline-formula><mml:math id="M93" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">ϕ</mml:mi><mml:mrow><mml:mi mathvariant="normal">wh</mml:mi><mml:mo>+</mml:mo><mml:mi mathvariant="normal">high</mml:mi><mml:mi mathvariant="normal">_</mml:mi><mml:mi mathvariant="normal">ig</mml:mi></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold">W</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold-italic">b</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> was superior to that of <inline-formula><mml:math id="M94" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">ϕ</mml:mi><mml:mrow><mml:mi mathvariant="normal">wh</mml:mi><mml:mo>+</mml:mo><mml:mi mathvariant="normal">low</mml:mi><mml:mi mathvariant="normal">_</mml:mi><mml:mi mathvariant="normal">ig</mml:mi></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold">W</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold-italic">b</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> (Table <xref ref-type="table" rid="App1.Ch1.S2.T7"/>). While the difference in performance from the different data subsets might be small in magnitude, the selection of high information gain points also has a positive effect on the structure of the global inference patterns: in Fig. <xref ref-type="fig" rid="App1.Ch1.S2.F6"/> we took the prediction maps for both models of the worst-performing data subset, <inline-formula><mml:math id="M95" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">ϕ</mml:mi><mml:mrow><mml:mi mathvariant="normal">wh</mml:mi><mml:mo>+</mml:mo><mml:mi mathvariant="normal">high</mml:mi><mml:mi mathvariant="normal">_</mml:mi><mml:mi mathvariant="normal">ig</mml:mi></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold">W</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold-italic">b</mml:mi><mml:msub><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M96" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">ϕ</mml:mi><mml:mrow><mml:mi mathvariant="normal">wh</mml:mi><mml:mo>+</mml:mo><mml:mi mathvariant="normal">low</mml:mi><mml:mi mathvariant="normal">_</mml:mi><mml:mi mathvariant="normal">ig</mml:mi></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold">W</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold-italic">b</mml:mi><mml:msub><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>, and calculated the absolute difference between them and the inference map of the original model <inline-formula><mml:math id="M97" display="inline"><mml:mrow><mml:mi mathvariant="italic">ϕ</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold">W</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold-italic">b</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> in Fig. <xref ref-type="fig" rid="Ch1.F3"/>. Regardless of the performance metrics, the high information gain model resembles the output of the original model more closely than the low information gain model.</p>
</app>

<app id="App1.Ch1.S3">
  <label>Appendix C</label><title>Comparison of the methods</title>
      <p id="d2e4224">One of the drawbacks of using DNN is the number of hyperparameters that needs to be tuned. The number of layers and the nodes in each layer were decided using a trial-and-error method starting with the simplest configuration of three layers of eight neurons. The model complexity was increased till the validation and training performance were comparable, thus avoiding overfitting while still getting relatively good performance in the test dataset. The initial learning rate was chosen based on the model convergence. The DNN model had 10 layers of 128 nodes each with a learning rate of 0.01. The batch size, decided based on the amount of data, was set to 500 and was also chosen based on model convergence. On the other hand, the parameters that were tuned in the random forest algorithm and kNNs were the number of trees in the forest (controlled by the number of estimators in sklearn) and the number of neighbours, respectively. They are tuned using the performance metrics for 1–50 neighbours for kNN. The number of estimators is 10, 20, 30, <inline-formula><mml:math id="M98" display="inline"><mml:mi mathvariant="normal">…</mml:mi></mml:math></inline-formula>, 100 for random forests.</p>
      <p id="d2e4234">Though it is difficult to tune the DNN model, Table <xref ref-type="table" rid="Ch1.T1"/> highlights superior performance in the training dataset for kNNs and random forests, while their test performance or generalization capability lags behind that of DNNs. Figures <xref ref-type="fig" rid="App1.Ch1.S3.F7"/> and <xref ref-type="fig" rid="App1.Ch1.S3.F8"/> show artifacts of the global predictions from kNNs and random forests, particularly in the equatorial Pacific and Atlantic oceans. </p>

      <fig id="App1.Ch1.S3.F7"><label>Figure C1</label><caption><p id="d2e4246">Global prediction map of TOC concentrations using a <inline-formula><mml:math id="M99" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>-nearest neighbour algorithm, with the five nearest neighbours in the continental shelves and the four nearest neighbours in the deep sea. Spurious patches are observed in the equatorial Pacific Ocean and in the Atlantic Ocean.</p></caption>
        
        <graphic xlink:href="https://gmd.copernicus.org/articles/18/2521/2025/gmd-18-2521-2025-f07.png"/>

      </fig>

      <fig id="App1.Ch1.S3.F8"><label>Figure C2</label><caption><p id="d2e4267">Global prediction map of TOC concentrations using a random forest algorithm with 100 estimators. Spurious patches are observed in the Atlantic Ocean and the Bay of Bengal.</p></caption>
        
        <graphic xlink:href="https://gmd.copernicus.org/articles/18/2521/2025/gmd-18-2521-2025-f08.png"/>

      </fig>

</app>

<app id="App1.Ch1.S4">
  <label>Appendix D</label><title>TOC stock in different marine regions</title>
      <p id="d2e4286">Table <xref ref-type="table" rid="Ch1.T2"/> breaks down how much TOC stock is found in different parts of the ocean. Each region is listed, showing how much TOC is there. Here we show a visualization of the different regions in Fig. <xref ref-type="fig" rid="App1.Ch1.S4.F9"/>. </p>
      <p id="d2e4295">In Fig. <xref ref-type="fig" rid="App1.Ch1.S4.F10"/>, we use a waffle chart to make it easier to see how the TOC is split among these regions. It is like dividing a pie into slices, but here we use squares. With a total of about 156 Pg of TOC worldwide, the South Pacific Ocean gets the biggest share, while the Baltic Sea gets the smallest. </p>

      <fig id="App1.Ch1.S4.F9"><label>Figure D1</label><caption><p id="d2e4303">TOC stocks in different oceans.</p></caption>
        
        <graphic xlink:href="https://gmd.copernicus.org/articles/18/2521/2025/gmd-18-2521-2025-f09.png"/>

      </fig>

      <fig id="App1.Ch1.S4.F10"><label>Figure D2</label><caption><p id="d2e4317">TOC stocks in different oceans: waffle chart.</p></caption>
        
        <graphic xlink:href="https://gmd.copernicus.org/articles/18/2521/2025/gmd-18-2521-2025-f10.png"/>

      </fig>

</app>

<app id="App1.Ch1.S5">
  <label>Appendix E</label><title>DNN model run without separation of deep-sea and shelf environments</title>
      <p id="d2e4336">Here, we test a DNN model where the global ocean was not separated into shelf and deep-ocean regions but treated as one entity. The resulting TOC map shows spurious features in the Pacific Ocean, similar to those that occur in the map published by <xref ref-type="bibr" rid="bib1.bibx33" id="text.132"/>. These results underscore the importance of separating shelf and deep-ocean regions in order to achieve more accurate and realistic model outcomes. </p>

      <fig id="App1.Ch1.S5.F11"><label>Figure E1</label><caption><p id="d2e4345">TOC concentration map when the DNN model was not separated into shelf and deep-ocean regions. We see unrealistic TOC concentrations, especially in the Pacific Ocean.</p></caption>
        
        <graphic xlink:href="https://gmd.copernicus.org/articles/18/2521/2025/gmd-18-2521-2025-f11.png"/>

      </fig>

</app>

<app id="App1.Ch1.S6">
  <label>Appendix F</label><title>Model interpretability using SHAP values</title>
      <p id="d2e4365">Explaining and understanding why a model makes a certain prediction is as crucial as accuracy and uncertainty in the predictions. This becomes particularly challenging in high-dimensional spaces, where interpreting complex models can be more intricate compared to simpler yet less accurate models. <xref ref-type="bibr" rid="bib1.bibx38" id="text.133"/> proposed SHAP as a unified framework for interpreting predictions. SHAP assigns importance values to each feature for a particular prediction, providing a comprehensive understanding of the model's decision-making process. In our supervised learning model <inline-formula><mml:math id="M100" display="inline"><mml:mi>f</mml:mi></mml:math></inline-formula> trained on features <inline-formula><mml:math id="M101" display="inline"><mml:mrow><mml:mi>X</mml:mi><mml:mo>∈</mml:mo><mml:mi mathvariant="script">X</mml:mi><mml:mo>⊆</mml:mo><mml:msup><mml:mi mathvariant="double-struck">R</mml:mi><mml:mi>d</mml:mi></mml:msup></mml:mrow></mml:math></inline-formula> to predict outcomes <inline-formula><mml:math id="M102" display="inline"><mml:mrow><mml:mi>Y</mml:mi><mml:mo>∈</mml:mo><mml:mi mathvariant="script">Y</mml:mi><mml:mo>⊆</mml:mo><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:math></inline-formula>, SHAP, a feature attribution method, considers the model predictions to be decomposed as a sum <inline-formula><mml:math id="M103" display="inline"><mml:mrow><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="italic">ϕ</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:msubsup><mml:mo>∑</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>d</mml:mi></mml:msubsup><mml:mi mathvariant="italic">ϕ</mml:mi><mml:mo>(</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, where <inline-formula><mml:math id="M104" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">ϕ</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> is the baseline expectation (i.e. <inline-formula><mml:math id="M105" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">ϕ</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mi mathvariant="double-struck">E</mml:mi><mml:mo>[</mml:mo><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>) and <inline-formula><mml:math id="M106" display="inline"><mml:mrow><mml:mi mathvariant="italic">ϕ</mml:mi><mml:mo>(</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> denotes the Shapley value of feature <inline-formula><mml:math id="M107" display="inline"><mml:mi>j</mml:mi></mml:math></inline-formula> at point <inline-formula><mml:math id="M108" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula>.</p>
      <p id="d2e4532">In our analysis, we aim to simplify the interpretation process by presenting the average importance of features across all of the predictions, from the deep sea to the continental shelves. All of the effects describe the behaviour of the model and are not necessarily causal in the real world.</p>
      <p id="d2e4537">The summary plot in Figs. <xref ref-type="fig" rid="App1.Ch1.S6.F12"/> and <xref ref-type="fig" rid="App1.Ch1.S6.F13"/> combines the feature importance with feature effects. The summary plot displays Shapley values representing the impact of features on predictions. Each point represents a Shapley value for a feature and an instance. The <inline-formula><mml:math id="M109" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula> axis position indicates the feature, while the <inline-formula><mml:math id="M110" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> axis position corresponds to the Shapley value. Feature values are represented by colours ranging from low (blue) to high (red). To visualize feature importance, points are spread along the <inline-formula><mml:math id="M111" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula> axis to reveal the distribution of Shapley values per feature. The features are ordered based on their importance, determined by the mean absolute Shapley values across all of the predictions. The Shapley value is expressed in the same units as the TOC concentration. This indicates the extent to which a specific feature value influences the TOC concentration and whether it drives it towards higher or lower values.</p>

      <fig id="App1.Ch1.S6.F12"><label>Figure F1</label><caption><p id="d2e4569">Summary plot of Shapley values of the deep-sea DNN model. The global porosity grid <xref ref-type="bibr" rid="bib1.bibx39" id="paren.134"/> has the highest feature importance. Regions with high porosity lead to higher TOC concentrations and vice versa. In the mechanistic model of <xref ref-type="bibr" rid="bib1.bibx10" id="text.135"/>, porosity is positively correlated with the organic carbon flux through a specific depth. The biological features that include total biomass, meiofauna, fish biomass in the sea surface <xref ref-type="bibr" rid="bib1.bibx62" id="paren.136"/>, oxygen concentration in bottom waters <xref ref-type="bibr" rid="bib1.bibx9" id="paren.137"/>, and daily average PAR <xref ref-type="bibr" rid="bib1.bibx41" id="paren.138"/> show that higher biomass or marine productivity leads to higher TOC concentrations, as expected. On the other hand, higher oxygen saturation leads to oxic conditions, resulting in the oxidation of the organic carbon and hence a lower TOC concentration. The other features which dominate are the physical oceanographic features, where higher feature values result in lower TOC concentrations, such as tidal features (Q1 loading) <xref ref-type="bibr" rid="bib1.bibx23" id="paren.139"/>, sea bulk modulus <xref ref-type="bibr" rid="bib1.bibx9" id="paren.140"/>, and seafloor pressure <xref ref-type="bibr" rid="bib1.bibx9" id="paren.141"/>.</p></caption>
        
        <graphic xlink:href="https://gmd.copernicus.org/articles/18/2521/2025/gmd-18-2521-2025-f12.png"/>

      </fig>

<fig id="App1.Ch1.S6.F13"><label>Figure F2</label><caption><p id="d2e4608">Summary plot of Shapley values of the continental shelf DNN model. The total oxygen uptake <xref ref-type="bibr" rid="bib1.bibx27" id="paren.142"/> of the seabed has the highest feature importance, with regions of higher oxygen uptake resulting in lower TOC concentrations and denoting oxic conditions. Regions with higher porosity <xref ref-type="bibr" rid="bib1.bibx39" id="paren.143"/> result in higher TOC concentrations, while regions with lower porosity result in lower TOC concentrations but with less impact. The lithology map is a binary map. Regions with fine sand, with grain sizes between 0.05 and 0.5 mm (1 mm being the higher feature value), have low TOC concentrations. Higher sediment thicknesses in Earth's crust lead to lower TOC concentrations because of dilution <xref ref-type="bibr" rid="bib1.bibx6" id="paren.144"/>. The bottom current components, north–south and east–west, result in reduced TOC concentrations due to higher resuspension of sediments, the inhibition of sedimentation, and burial of organic carbon. Higher average seawater conductivity results in lower TOC concentrations. Higher particulate organic carbon (POC) in the water column has a positive impact on the TOC concentrations, as expected. It can be seen that the feature importance is not as clearly defined as in the deep ocean (as the high (red) and low (blue) feature values are mixed) because of the complex dynamics on continental shelves.</p></caption>
        
        <graphic xlink:href="https://gmd.copernicus.org/articles/18/2521/2025/gmd-18-2521-2025-f13.png"/>

      </fig>


</app>

<app id="App1.Ch1.S7">
  <label>Appendix G</label><title>Algorithms</title><boxed-text content-type="algorithm" position="float" id="App1.Ch1.S7.Prog1"><label>Algorithm G1</label><caption><p id="d2e4639">DNN training with batch normalization and dropout, including Monte Carlo dropout for inference.</p></caption><disp-quote content-type="algorithmic" specific-use="numbering{1}"><list>

    <list-item><label><bold>Require:</bold></label>

      <p id="d2e4649" specific-use="REQUIRE">Labelled dataset <inline-formula><mml:math id="M112" display="inline"><mml:mrow><mml:mi>D</mml:mi><mml:mo>=</mml:mo><mml:mo mathvariant="italic">{</mml:mo><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>N</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>N</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo mathvariant="italic">}</mml:mo></mml:mrow></mml:math></inline-formula></p>

      <p id="d2e4723" specific-use="REQUIRE"><inline-formula><mml:math id="M113" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>: feature vector for the <inline-formula><mml:math id="M114" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula>th label</p>

      <p id="d2e4744" specific-use="REQUIRE"><inline-formula><mml:math id="M115" display="inline"><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>: corresponding TOC %</p>
          </list-item>

    <list-item>

      <p id="d2e4761" specific-use="STATE"><bold>Input:</bold> feature vector <inline-formula><mml:math id="M116" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula></p>
          </list-item>

    <list-item>

      <p id="d2e4779" specific-use="STATE"><bold>Output:</bold> predicted TOC % predicted% for <inline-formula><mml:math id="M117" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula></p>
          </list-item>

    <list-item>

      <p id="d2e4797" specific-use="STATE"><bold>Method:</bold> construct a neural network with 10 layers and 128 nodes per layer: <inline-formula><mml:math id="M118" display="inline"><mml:mrow><mml:mi mathvariant="italic">ϕ</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold">W</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold-italic">b</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula></p>
          </list-item>

    <list-item>

      <p id="d2e4826" specific-use="STATE">Apply batch normalization and dropout to each layer.</p>
          </list-item>

    <list-item>

      <p id="d2e4833" specific-use="STATE">Initialize optimizer (e.g. Adam) with appropriate learning rate and parameters.</p>
          </list-item>

    <list-item>

      <p id="d2e4839" specific-use="STATE">Initialize loss function (e.g. mean squared error) for regression.</p>
          </list-item>

    <list-item>

      <p id="d2e4845" specific-use="STATE">Train the neural network with <inline-formula><mml:math id="M119" display="inline"><mml:mi>D</mml:mi></mml:math></inline-formula> for 1000 epochs:</p>
          </list-item>

    <list-item>

      <p id="d2e4858" specific-use="FOR"><bold>for</bold> <inline-formula><mml:math id="M120" display="inline"><mml:mrow><mml:mtext>epoch</mml:mtext><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> <bold>to</bold> num_epochs <bold>do</bold> <list>
    <list-item>
      <p id="d2e4884" specific-use="STATE">Randomly shuffle the training dataset..</p></list-item>
    <list-item>
      <p id="d2e4889" specific-use="FOR"><bold>for</bold> <inline-formula><mml:math id="M121" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> in <inline-formula><mml:math id="M122" display="inline"><mml:mi>D</mml:mi></mml:math></inline-formula> <bold>do</bold> <list>
    <list-item>
      <p id="d2e4929" specific-use="STATE">Forward pass: compute predictions <inline-formula><mml:math id="M123" display="inline"><mml:mrow><mml:mover accent="true"><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:mi mathvariant="italic">ϕ</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="bold">W</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold-italic">b</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>.</p></list-item>
    <list-item>
      <p id="d2e4970" specific-use="STATE">Compute target loss: <inline-formula><mml:math id="M124" display="inline"><mml:mrow><mml:msub><mml:mtext>loss</mml:mtext><mml:mtext>target</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:mtext>MSE</mml:mtext><mml:mo>(</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mover accent="true"><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>.</p></list-item>
    <list-item>
      <p id="d2e5010" specific-use="STATE">Back-propagation: update weights and biases using optimizer, with <inline-formula><mml:math id="M125" display="inline"><mml:mrow><mml:msub><mml:mtext>loss</mml:mtext><mml:mtext>target</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> as the cost function.</p></list-item></list></p></list-item>
    <list-item>
      <p id="d2e5026" specific-use="ENDFOR"><bold>end</bold> <bold>for</bold></p></list-item></list></p>
          </list-item>

    <list-item>

      <p id="d2e5036" specific-use="ENDFOR"><bold>end</bold> <bold>for</bold></p>
          </list-item>

    <list-item>

      <p id="d2e5046" specific-use="STATE">Set dropout to active during inference</p>
          </list-item>

    <list-item>

      <p id="d2e5053" specific-use="STATE">Perform Monte Carlo dropout for <inline-formula><mml:math id="M126" display="inline"><mml:mi>M</mml:mi></mml:math></inline-formula> forward runs:</p>
          </list-item>

    <list-item>

      <p id="d2e5066" specific-use="STATE"><inline-formula><mml:math id="M127" display="inline"><mml:mrow><mml:msubsup><mml:mover accent="true"><mml:mi>y</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mi>j</mml:mi><mml:mtext>ensemble</mml:mtext></mml:msubsup><mml:mo>=</mml:mo><mml:mi mathvariant="italic">ϕ</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="bold">W</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold-italic">b</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mtext>dropout_mask</mml:mtext><mml:mi>T</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula></p>
          </list-item>

    <list-item>

      <p id="d2e5116" specific-use="STATE">Predicted TOC %, <inline-formula><mml:math id="M128" display="inline"><mml:mover accent="true"><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover></mml:math></inline-formula> for <inline-formula><mml:math id="M129" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mn mathvariant="normal">1</mml:mn><mml:mi>M</mml:mi></mml:mfrac></mml:mstyle><mml:msubsup><mml:mo>∑</mml:mo><mml:mrow><mml:mi>m</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>M</mml:mi></mml:msubsup><mml:msubsup><mml:mover accent="true"><mml:mi>y</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mi>j</mml:mi><mml:mtext>ensemble</mml:mtext></mml:msubsup></mml:mrow></mml:math></inline-formula></p>
          </list-item>
        </list></disp-quote></boxed-text><boxed-text content-type="algorithm" position="float" id="App1.Ch1.S7.Prog2"><label>Algorithm G2</label><caption><p id="d2e5176">Calculating information gain for the predictions.</p></caption><disp-quote content-type="algorithmic" specific-use="numbering{1}"><list>

    <list-item><label><bold>Require:</bold></label>

      <p id="d2e5186" specific-use="REQUIRE">Monte Carlo dropout prediction ensemble, <inline-formula><mml:math id="M130" display="inline"><mml:mrow><mml:msup><mml:mover accent="true"><mml:mi>y</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mtext>ensemble</mml:mtext></mml:msup></mml:mrow></mml:math></inline-formula>, for each grid cell</p>
          </list-item>

    <list-item>

      <p id="d2e5206" specific-use="FOR"><bold>for</bold> each grid cell <bold>do</bold> <list>
    <list-item>
      <p id="d2e5217" specific-use="STATE">Fit a Gaussian probability density function <inline-formula><mml:math id="M131" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> for <inline-formula><mml:math id="M132" display="inline"><mml:mrow><mml:msubsup><mml:mover accent="true"><mml:mi>y</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mi>j</mml:mi><mml:mtext>ensemble</mml:mtext></mml:msubsup></mml:mrow></mml:math></inline-formula> using histograms and curve fitting algorithm.</p></list-item>
    <list-item>
      <p id="d2e5255" specific-use="STATE">Generate original distribution <inline-formula><mml:math id="M133" display="inline"><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> with mean <inline-formula><mml:math id="M134" display="inline"><mml:mover accent="true"><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover></mml:math></inline-formula> and standard deviation 0.05 (sampling error).</p></list-item>
    <list-item>
      <p id="d2e5291" specific-use="STATE">Calculate Kullback–Leibler divergence:</p>
      <p id="d2e5294" specific-use="STATE"><inline-formula><mml:math id="M135" display="inline"><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mtext>KL</mml:mtext></mml:msub><mml:mo>(</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>‖</mml:mo><mml:msub><mml:mi>Q</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mo>∑</mml:mo><mml:mi>i</mml:mi></mml:msub><mml:msub><mml:mi>P</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mi>log⁡</mml:mi><mml:mfenced close=")" open="("><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mfrac></mml:mstyle></mml:mfenced></mml:mrow></mml:math></inline-formula></p></list-item></list></p>
          </list-item>

    <list-item>

      <p id="d2e5383" specific-use="ENDFOR"><bold>end</bold> <bold>for</bold></p>
          </list-item>
        </list></disp-quote></boxed-text>
</app>
  </app-group><notes notes-type="codeavailability"><title>Code availability</title>

      <p id="d2e5396">The repository of the code to run the different models and analyse the outputs is available at <uri>https://github.com/paramnav/nn-toc</uri> <xref ref-type="bibr" rid="bib1.bibx46" id="paren.145"/> and <ext-link xlink:href="https://doi.org/10.5281/zenodo.12206145" ext-link-type="DOI">10.5281/zenodo.12206145</ext-link> <xref ref-type="bibr" rid="bib1.bibx47" id="paren.146"/>.</p>
  </notes><notes notes-type="dataavailability"><title>Data availability</title>

      <p id="d2e5414">The raw features and labels as well as the model outputs are available at <ext-link xlink:href="https://doi.org/10.5281/zenodo.11186224" ext-link-type="DOI">10.5281/zenodo.11186224</ext-link> <xref ref-type="bibr" rid="bib1.bibx47" id="paren.147"/>.</p>
  </notes><notes notes-type="authorcontribution"><title>Author contributions</title>

      <p id="d2e5426">NP: data curation, formal analysis, investigation, methodology, software, validation, visualization, writing – original draft preparation. EG: conceptualization, data curation, formal analysis, investigation, methodology, software, supervision, visualization, writing – original draft preparation. EBG: data curation, investigation, methodology, supervision, validation, writing – original draft preparation. MB: funding acquisition, project administration, supervision, validation, writing – original draft preparation. KW: conceptualization, funding acquisition, methodology, project administration, supervision, validation, writing – original draft preparation.</p>
  </notes><notes notes-type="competinginterests"><title>Competing interests</title>

      <p id="d2e5432">The contact author has declared that none of the authors has any competing interests.</p>
  </notes><notes notes-type="disclaimer"><title>Disclaimer</title>

      <p id="d2e5438">Publisher’s note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. While Copernicus Publications makes every effort to include appropriate place names, the final responsibility lies with the authors.</p>
  </notes><ack><title>Acknowledgements</title><p id="d2e5444">The first author wishes to thank the Helmholtz School for Marine Data Science (MarDATA) for its direct financial support.</p><p id="d2e5446">The authors would like to thank the reviewers and the editorial team, for their insightful reviews, which led to further improvement of the work.</p><p id="d2e5448">AI tools were used to correct the manuscript and optimize the code. The authors are grateful for the high-performance computing sources at Kiel University.</p></ack><notes notes-type="financialsupport"><title>Financial support</title>

      <p id="d2e5453">This work was partially funded by the Cluster of Excellence “The Ocean Floor – Earth’s Uncharted Interface” (EXC 2077) and the Deutsche Forschungsgemeinschaft (DFG) (project no. 390741603) hosted by the Research Faculty MARUM – Center for Marine Environmental Sciences, University of Bremen, Germany.The article processing charges for this open-access publication were covered by the GEOMAR Helmholtz Centre for Ocean Research Kiel.</p>
  </notes><notes notes-type="reviewstatement"><title>Review statement</title>

      <p id="d2e5462">This paper was edited by Sandra Arndt and reviewed by Taylor Lee and Sarah Paradis.</p>
  </notes><ref-list>
    <title>References</title>

      <ref id="bib1.bibx1"><label>Arndt et al.(2013)Arndt, Jørgensen, LaRowe, Middelburg, Pancost, and Regnier</label><mixed-citation>Arndt, S., Jørgensen, B., LaRowe, D., Middelburg, J., Pancost, R., and Regnier, P.: Quantifying the degradation of organic matter in marine sediments: A review and synthesis, Earth-Sci. Rev., 123, 53–86, <ext-link xlink:href="https://doi.org/10.1016/j.earscirev.2013.02.008" ext-link-type="DOI">10.1016/j.earscirev.2013.02.008</ext-link>, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx2"><label>Atwood et al.(2020)Atwood, Witt, Mayorga, Hammill, and Sala</label><mixed-citation>Atwood, T. B., Witt, A., Mayorga, J., Hammill, E., and Sala, E.: Global Patterns in Marine Sediment Carbon Stocks, Frontiers in Marine Science, 7, 165, <ext-link xlink:href="https://doi.org/10.3389/fmars.2020.00165" ext-link-type="DOI">10.3389/fmars.2020.00165</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx3"><label>Baturin(2007)</label><mixed-citation>Baturin, G. N.: Issue of the relationship between primary productivity of organic carbon in ocean and phosphate accumulation (Holocene-Late Jurassic), Lith. Miner. Resour., 42, 318–348, <ext-link xlink:href="https://doi.org/10.1134/S0024490207040025" ext-link-type="DOI">10.1134/S0024490207040025</ext-link>, 2007.</mixed-citation></ref>
      <ref id="bib1.bibx4"><label>Beazley(2003)</label><mixed-citation>Beazley, M. J.: The significance of organic carbon and sediment surface area to the benthic biogeochemistry of the slope and deep water environments of the northern Gulf of Mexico, Master's thesis, Texas A&amp;M University, <uri>http://hdl.handle.net/1969.1/534</uri> (last access: 3 February 2024), 2003.</mixed-citation></ref>
      <ref id="bib1.bibx5"><label>Becker et al.(2014)Becker, Wood, and Martin</label><mixed-citation>Becker, J. J., Wood, W. T., and Martin, K. M.: Global Crustal Heat Flow Using Random Decision Forest Prediction, in: AGU Fall Meeting Abstracts, vol. 2014,  NG31A–3788,  <uri>https://ui.adsabs.harvard.edu/abs/2014AGUFMNG31A3788B/abstract</uri> (last access: 3 February 2024), 2014.</mixed-citation></ref>
      <ref id="bib1.bibx6"><label>Berner(1982)</label><mixed-citation>Berner, R. A.: Burial of organic carbon and pyrite sulfur in the modern ocean: its geochemical and environmental significance, Am. J. Sci., 282, 451–473, <ext-link xlink:href="https://doi.org/10.2475/ajs.282.4.451" ext-link-type="DOI">10.2475/ajs.282.4.451</ext-link>, 1982.</mixed-citation></ref>
      <ref id="bib1.bibx7"><label>Berner(2004)</label><mixed-citation>Berner, R. A.: Processes of the Long-Term Carbon Cycle: Organic Matter and Carbonate Burial and Weathering, in: The Phanerozoic Carbon Cycle: CO<sub>2</sub> and O<sub>2</sub>, Oxford University Press, ISBN 9780195173338, <ext-link xlink:href="https://doi.org/10.1093/oso/9780195173338.003.0005" ext-link-type="DOI">10.1093/oso/9780195173338.003.0005</ext-link>, 2004.</mixed-citation></ref>
      <ref id="bib1.bibx8"><label>Boudreau(1997)</label><mixed-citation>Boudreau, B. P.: Diagenetic models and their implementation, vol. 505, Springer Berlin, ISBN 978-3-642-64399-6, <ext-link xlink:href="https://doi.org/10.1007/978-3-642-60421-8" ext-link-type="DOI">10.1007/978-3-642-60421-8</ext-link>, 1997.</mixed-citation></ref>
      <ref id="bib1.bibx9"><label>Boyer et al.(2013)Boyer, Antonov, Baranova, Coleman, Garcia, and Grodsky</label><mixed-citation>Boyer, T. P., Antonov, J. I., Baranova, O. K., Coleman, C., Garcia, H. E., and Grodsky, A.: World Ocean Database 2013, NOAA Atlas NESDIS 72, Technical Ed. Silver Spring, MD, <ext-link xlink:href="https://doi.org/10.7289/V5NZ85MT" ext-link-type="DOI">10.7289/V5NZ85MT</ext-link>, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx10"><label>Bradley and Arndt(2022)</label><mixed-citation>Bradley, J. A., Hülse, D., LaRowe, D. E., and Arndt, S.: Transfer efficiency of organic carbon in marine sediments, Nat. Commun., 13, 7297, <ext-link xlink:href="https://doi.org/10.1038/s41467-022-35112-9" ext-link-type="DOI">10.1038/s41467-022-35112-9</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx11"><label>Burdige(2005)</label><mixed-citation>Burdige, D. J.: Burial of terrestrial organic matter in marine sediments: A re-assessment, Global Biogeochem. Cy., 19,  GB4011, <ext-link xlink:href="https://doi.org/10.1029/2004GB002368" ext-link-type="DOI">10.1029/2004GB002368</ext-link>, 2005.</mixed-citation></ref>
      <ref id="bib1.bibx12"><label>Burdige(2007)</label><mixed-citation>Burdige, D. J.: Preservation of Organic Matter in Marine Sediments – Controls, Mechanisms, and an Imbalance in Sediment Organic Carbon Budgets?, Chem. Rev., 107, 467–485, <ext-link xlink:href="https://doi.org/10.1021/cr050347q" ext-link-type="DOI">10.1021/cr050347q</ext-link>, 2007.</mixed-citation></ref>
      <ref id="bib1.bibx13"><label>Crameri(2023)</label><mixed-citation>Crameri, F.: Scientific colour maps, Zenodo [data set], <ext-link xlink:href="https://doi.org/10.5281/zenodo.8409685" ext-link-type="DOI">10.5281/zenodo.8409685</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx14"><label>De Veaux and Ungar(1994)</label><mixed-citation> De Veaux, R. D. and Ungar, L. H.: Multicollinearity: A tale of two nonparametric regressions, in: Selecting Models from Data, edited by: Cheeseman, P. and Oldford, R. W., Springer New York, New York, NY, 393–402,  ISBN 978-1-4612-2660-4, 1994.</mixed-citation></ref>
      <ref id="bib1.bibx15"><label>Diesing et al.(2021)Diesing, Thorsnes, and Bjarnadóttir</label><mixed-citation>Diesing, M., Thorsnes, T., and Bjarnadóttir, L. R.: Organic carbon densities and accumulation rates in surface sediments of the North Sea and Skagerrak, Biogeosciences, 18, 2139–2160, <ext-link xlink:href="https://doi.org/10.5194/bg-18-2139-2021" ext-link-type="DOI">10.5194/bg-18-2139-2021</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx16"><label>Emerson and Hedges(1988)</label><mixed-citation>Emerson, S. and Hedges, J. I.: Processes controlling the organic carbon content of open ocean sediments, Paleoceanography, 3, 621–634, <ext-link xlink:href="https://doi.org/10.1029/PA003i005p00621" ext-link-type="DOI">10.1029/PA003i005p00621</ext-link>, 1988.</mixed-citation></ref>
      <ref id="bib1.bibx17"><label>Emery(1968)</label><mixed-citation> Emery, K. O.: Relict sediments on continental shelves of the world, Am. Assoc. Petr. Geol. B., 52, 445–464, 1968.</mixed-citation></ref>
      <ref id="bib1.bibx18"><label>Flanders Marine Institute(2021)</label><mixed-citation>Flanders Marine Institute: Global Oceans and Seas, version 1, <ext-link xlink:href="https://doi.org/10.14284/542" ext-link-type="DOI">10.14284/542</ext-link>,   <uri>https://www.marineregions.org/</uri> (last access: 25 August 2023), 2021.</mixed-citation></ref>
      <ref id="bib1.bibx19"><label>Gal and Ghahramani(2016)</label><mixed-citation>Gal, Y. and Ghahramani, Z.: Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning, in: Proceedings of The 33rd International Conference on Machine Learning, New York, New York, USA,  edited by: Balcan, M. F. and Weinberger, K. Q., Proceedings of Machine Learning Research,  vol. 48, PMLR,  19 June 2016, 1050–1059, <uri>https://proceedings.mlr.press/v48/gal16.html</uri> (last access: 24 February 2022), 2016.</mixed-citation></ref>
      <ref id="bib1.bibx20"><label>Garlan et al.(2018)Garlan, Gabelotaud, Lucas, and Marchès</label><mixed-citation> Garlan, T., Gabelotaud, I., Lucas, S., and Marchès, E.: A World Map of Seabed Sediment Based on 50 Years of Knowledge, World Academy of Science, Engineering and Technology, International Journal of Geological and Environmental Engineering, 12, 409–419, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx21"><label>Goyet et al.(2000)Goyet, Healy, and Ryan</label><mixed-citation>Goyet, C., Healy, R., and Ryan, J.: Global Distribution of Total Inorganic Carbon and Total Alkalinity Below the Deepest Winter Mixed Layer Depths, ORNIJCDIAC-127 NDP-076,  <ext-link xlink:href="https://doi.org/10.3334/zCDIAC/otg.ndp076" ext-link-type="DOI">10.3334/zCDIAC/otg.ndp076</ext-link>, 2000.</mixed-citation></ref>
      <ref id="bib1.bibx22"><label>Hall(2002)</label><mixed-citation>Hall, S. J.: The continental shelf benthic ecosystem: current status, agents for change and future prospects, Environ. Conserv., 29, 350–374, <uri>http://www.jstor.org/stable/44520615</uri> (last access: 26 June 2022), 2002.</mixed-citation></ref>
      <ref id="bib1.bibx23"><label>Hart-Davis et al.(2021)Hart-Davis, Piccioni, Dettmering, Schwatke, Passaro, and Seitz</label><mixed-citation>Hart-Davis, M. G., Piccioni, G., Dettmering, D., Schwatke, C., Passaro, M., and Seitz, F.: EOT20: a global ocean tide model from multi-mission satellite altimetry, Earth Syst. Sci. Data, 13, 3869–3884, <ext-link xlink:href="https://doi.org/10.5194/essd-13-3869-2021" ext-link-type="DOI">10.5194/essd-13-3869-2021</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx24"><label>He et al.(2015)He, Zhang, Ren, and Sun</label><mixed-citation> He, K., Zhang, X., Ren, S., and Sun, J.: Delving deep into rectifiers: Surpassing human-level performance on imagenet classification, in: Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile,  1026–1034, 7–13 December 2015.</mixed-citation></ref>
      <ref id="bib1.bibx25"><label>Hedges and Keil(1995)</label><mixed-citation>Hedges, J. I. and Keil, R. G.: Sedimentary organic matter preservation: an assessment and speculative synthesis, Mar. Chem., 49, 81–115, <ext-link xlink:href="https://doi.org/10.1016/0304-4203(95)00008-F" ext-link-type="DOI">10.1016/0304-4203(95)00008-F</ext-link>, 1995.</mixed-citation></ref>
      <ref id="bib1.bibx26"><label>Joint Panel on Oceanographic Tables and Standards(1991)</label><mixed-citation> Joint Panel on Oceanographic Tables and Standards: Processing of Oceanographic Station Data, UNESCO, Paris, ISBN 978-92-3-102756-7, 1991.</mixed-citation></ref>
      <ref id="bib1.bibx27"><label>Jørgensen et al.(2022)Jørgensen, Wenzhöfer, Egger, and Glud</label><mixed-citation>Jørgensen, B. B., Wenzhöfer, F., Egger, M., and Glud, R. N.: Sediment oxygen consumption: Role in the global marine carbon cycle, Earth-Sci. Rev., 228, 103987, <ext-link xlink:href="https://doi.org/10.1016/j.earscirev.2022.103987" ext-link-type="DOI">10.1016/j.earscirev.2022.103987</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx28"><label>Kim and Wessel(2011)</label><mixed-citation>Kim, S. S. and Wessel, P.: New global seamount census from the altimetry-derived gravity data, Geophys. J. Int., 186, 615–631, <ext-link xlink:href="https://doi.org/10.1111/j.1365-246X.2011.05076.x" ext-link-type="DOI">10.1111/j.1365-246X.2011.05076.x</ext-link>, 2011.</mixed-citation></ref>
      <ref id="bib1.bibx29"><label>Kullback and Leibler(1951)</label><mixed-citation>Kullback, S. and Leibler, R. A.: On Information and Sufficiency, Ann. Math. Stat., 22, 79–86, <uri>http://www.jstor.org/stable/2236703</uri> (last access: 19 April 2023), 1951.</mixed-citation></ref>
      <ref id="bib1.bibx30"><label>LaRowe et al.(2020a)LaRowe, Arndt, Bradley, Estes, Hoarfrost, Lang, Lloyd, Mahmoudi, Orsi, Shah Walter, Steen, and Zhao</label><mixed-citation>LaRowe, D., Arndt, S., Bradley, J., Estes, E., Hoarfrost, A., Lang, S., Lloyd, K., Mahmoudi, N., Orsi, W., Shah Walter, S., Steen, A., and Zhao, R.: The fate of organic carbon in marine sediments – New insights from recent data and analysis, Earth-Sci. Rev., 204, 103146, <ext-link xlink:href="https://doi.org/10.1016/j.earscirev.2020.103146" ext-link-type="DOI">10.1016/j.earscirev.2020.103146</ext-link>, 2020a.</mixed-citation></ref>
      <ref id="bib1.bibx31"><label>LaRowe et al.(2020b)LaRowe, Arndt, Bradley, Burwicz, Dale, and Amend</label><mixed-citation>LaRowe, D. E., Arndt, S., Bradley, J. A., Burwicz, E., Dale, A. W., and Amend, J. P.: Organic carbon and microbial activity in marine sediments on a global scale throughout the Quaternary, Geochim. Cosmochim. Ac., 286, 227–247, <ext-link xlink:href="https://doi.org/10.1016/j.gca.2020.07.017" ext-link-type="DOI">10.1016/j.gca.2020.07.017</ext-link>, 2020b.</mixed-citation></ref>
      <ref id="bib1.bibx32"><label>LeCun et al.(2015)LeCun, Bengio, and Hinton</label><mixed-citation>LeCun, Y., Bengio, Y., and Hinton, G.: Deep learning, Nature, 521, 436–444, <ext-link xlink:href="https://doi.org/10.1038/nature14539" ext-link-type="DOI">10.1038/nature14539</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx33"><label>Lee et al.(2019)Lee, Wood, and Phrampus</label><mixed-citation>Lee, T. R., Wood, W. T., and Phrampus, B. J.: A Machine Learning (kNN) Approach to Predicting Global Seafloor Total Organic Carbon, Global Biogeochem. Cy., 33, 37–46, <ext-link xlink:href="https://doi.org/10.1029/2018GB005992" ext-link-type="DOI">10.1029/2018GB005992</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx34"><label>Lee et al.(2020)Lee, Wood, Skarke, Phrampus, and Obelcz</label><mixed-citation>Lee, T. R., Wood, W. T., Skarke, A., Phrampus, B. J., and Obelcz, J.: Data files associated with the <inline-formula><mml:math id="M138" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>-nearest neighbor global prediction of isopachs for present to middle Miocene, Zenodo [data set], <ext-link xlink:href="https://doi.org/10.5281/zenodo.3675364" ext-link-type="DOI">10.5281/zenodo.3675364</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx35"><label>Legge et al.(2020)Legge, Johnson, Hicks, Jickells, Diesing, Aldridge, Andrews, Artioli, Bakker, Burrows, Carr, Cripps, Felgate, Fernand, Greenwood, Hartman, Kröger, Lessin, Mahaffey, Mayor, Parker, Queirós, Shutler, Silva, Stahl, Tinker, Underwood, Van Der Molen, Wakelin, Weston, and Williamson</label><mixed-citation>Legge, O., Johnson, M., Hicks, N., Jickells, T., Diesing, M., Aldridge, J., Andrews, J., Artioli, Y., Bakker, D. C. E., Burrows, M. T., Carr, N., Cripps, G., Felgate, S. L., Fernand, L., Greenwood, N., Hartman, S., Kröger, S., Lessin, G., Mahaffey, C., Mayor, D. J., Parker, R., Queirós, A. M., Shutler, J. D., Silva, T., Stahl, H., Tinker, J., Underwood, G. J. C., Van Der Molen, J., Wakelin, S., Weston, K., and Williamson, P.: Carbon on the Northwest European Shelf: Contemporary Budget and Future Influences, Frontiers in Marine Science, 7, 143, <ext-link xlink:href="https://doi.org/10.3389/fmars.2020.00143" ext-link-type="DOI">10.3389/fmars.2020.00143</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx36"><label>Lucas and Giles(2016)</label><mixed-citation>Lucas, M. and Giles, H.: Quantifying Uncertainty in Random Forests via Confidence Intervals and Hypothesis Tests, J. Mach. Learn. Res., 17, 1–41, <uri>http://jmlr.org/papers/v17/14-168.html</uri> (last access: 4 August 2023), 2016.</mixed-citation></ref>
      <ref id="bib1.bibx37"><label>Ludwig et al.(2011)Ludwig, Amiotte-Suchet, and Probst</label><mixed-citation>Ludwig, W., Amiotte-Suchet, P., and Probst, J. L.: ISLSCP II Global River Fluxes of Carbon and Sediments to the Oceans,  ORNL Distributed Active Archive Center [data set], <ext-link xlink:href="https://doi.org/10.3334/ORNLDAAC/1028" ext-link-type="DOI">10.3334/ORNLDAAC/1028</ext-link>, 2011.</mixed-citation></ref>
      <ref id="bib1.bibx38"><label>Lundberg and Lee(2017)</label><mixed-citation> Lundberg, S. M. and Lee, S.-I.: A unified approach to interpreting model predictions, in: Proceedings of the 31st International Conference on Neural Information Processing Systems, NIPS'17,  Curran Associates Inc., Red Hook, NY, USA, 4–9 December 2017, 4768–4777, ISBN 9781510860964, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx39"><label>Martin et al.(2015)Martin, Wood, and Becker</label><mixed-citation>Martin, K. M., Wood, W. T., and Becker, J. J.: A global prediction of seafloor sediment porosity using machine learning, Geophys. Res. Lett., 42, 10640–10646, <ext-link xlink:href="https://doi.org/10.1002/2015GL065279" ext-link-type="DOI">10.1002/2015GL065279</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx40"><label>NASA(2011)</label><mixed-citation>NASA: Announcement of Aquarius Level 2 Data Availability, Physical Oceanography Distributed Active Archive Center (PODAAC), <uri>https://aquarius.oceansciences.org/cgi/gal_density.htm</uri> (last access: 23 October 2023), 2011.</mixed-citation></ref>
      <ref id="bib1.bibx41"><label>NASA(2014)</label><mixed-citation>NASA: MODIS-Aqua Ocean Color Data, Goddard Space Flight Center, Ocean Ecology Laboratory, Ocean Biology Processing Group, <ext-link xlink:href="https://doi.org/10.5067/AQUA/MODIS_OC.2014.0" ext-link-type="DOI">10.5067/AQUA/MODIS_OC.2014.0</ext-link>, 2014.</mixed-citation></ref>
      <ref id="bib1.bibx42"><label>National Geophysical Data Center(1976)</label><mixed-citation>National Geophysical Data Center: The NGDC Seafloor Sediment Grain Size Database, first Version, NOAA National Centers for Environmental Information, <ext-link xlink:href="https://doi.org/10.7289/V5G44N6W" ext-link-type="DOI">10.7289/V5G44N6W</ext-link>, 1976.</mixed-citation></ref>
      <ref id="bib1.bibx43"><label>National Geophysical Data Center(2006)</label><mixed-citation>National Geophysical Data Center: 2-minute Gridded Global Relief Data (ETOPO2) v2, NCEI [data set], <ext-link xlink:href="https://doi.org/10.7289/V5J1012Q" ext-link-type="DOI">10.7289/V5J1012Q</ext-link>,  2006.</mixed-citation></ref>
      <ref id="bib1.bibx44"><label>Pape et al.(2020)Pape, Bünz, Hong, Torres, Riedel, Panieri, Lepland, Hsu, Wintersteller, Wallmann, Schmidt, Yao, and Bohrmann</label><mixed-citation>Pape, T., Bünz, S., Hong, W.-L., Torres, M. E., Riedel, M., Panieri, G., Lepland, A., Hsu, C.-W., Wintersteller, P., Wallmann, K., Schmidt, C., Yao, H., and Bohrmann, G.: Origin and Transformation of Light Hydrocarbons Ascending at an Active Pockmark on Vestnesa Ridge, Arctic Ocean, J. Geophys. Res.-Sol. Ea., 125, e2018JB016679, <ext-link xlink:href="https://doi.org/10.1029/2018JB016679" ext-link-type="DOI">10.1029/2018JB016679</ext-link>,  2020.</mixed-citation></ref>
      <ref id="bib1.bibx45"><label>Paradis et al.(2023)Paradis, Nakajima, Van der Voort, Gies, Wildberger, Blattmann, Bröder, and Eglinton</label><mixed-citation>Paradis, S., Nakajima, K., Van der Voort, T. S., Gies, H., Wildberger, A., Blattmann, T. M., Bröder, L., and Eglinton, T. I.: The Modern Ocean Sediment Archive and Inventory of Carbon (MOSAIC): version 2.0, Earth Syst. Sci. Data, 15, 4105–4125, <ext-link xlink:href="https://doi.org/10.5194/essd-15-4105-2023" ext-link-type="DOI">10.5194/essd-15-4105-2023</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx46"><label>Parameswaran et al.(2024a)Parameswaran, González, Burwicz-Galerne, Braack, and Wallmann</label><mixed-citation>Parameswaran, N., González, E., Burwicz-Galerne, E., Braack, M., and Wallmann, K.: Code for Global Prediction Of Total Organic Carbon In Marine Sediments Using Deep Neural Networks (nn-toc), GitHub [code], <uri>https://github.com/paramnav/nn-toc</uri> (last access: 13 November 2024), 2024a.</mixed-citation></ref>
      <ref id="bib1.bibx47"><label>Parameswaran et al.(2024b)Parameswaran, González, Burwicz-Galerne, Braack, and Wallmann</label><mixed-citation>Parameswaran, N., González, E., Burwicz-Galerne, E., Braack, M., and Wallmann, K.: Dataset for the Global Prediction Of Total Organic Carbon In Marine Sediments Using Deep Neural Networks (nn-toc), Zenodo [data set], <ext-link xlink:href="https://doi.org/10.5281/zenodo.11186224" ext-link-type="DOI">10.5281/zenodo.11186224</ext-link>, 2024b.</mixed-citation></ref>
      <ref id="bib1.bibx48"><label>Pavlis et al.(2008)Pavlis, Holmes, Kenyon, and Factor</label><mixed-citation>Pavlis, N. K., Holmes, S. A., Kenyon, S. C., and Factor, J. K.: The EGM2008 Global Gravitational Model, abstract 2008AGUFM.G22A..01P, 2008 General Assembly of the European Geosciences Union, Vienna, Austria, <uri>https://ui.adsabs.harvard.edu/abs/2008AGUFM.G22A..01P</uri> (last access: 7 October 2014), 2008.</mixed-citation></ref>
      <ref id="bib1.bibx49"><label>Phrampus et al.(2019)Phrampus, Lee, and Wood</label><mixed-citation>Phrampus, B. J., Lee, T. R., and Wood, W. T.: Predictor Grids for “A Global Probabilistic Prediction of Cold Seeps and Associated Seafloor Fluid Expulsion Anomalies (SEAFLEAs)”, Zenodo [data set], <ext-link xlink:href="https://doi.org/10.5281/zenodo.3459805" ext-link-type="DOI">10.5281/zenodo.3459805</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx50"><label>Rényi(1961)</label><mixed-citation> Rényi, A.: On measures of entropy and information, in: Proceedings of the fourth Berkeley symposium on mathematical statistics and probability, volume 1: contributions to the theory of statistics, Fourth Berkley Symposum on Mathematical Statistics and Probablity June 20-July 30, 1960, Statistical Laboratory of the University of California vol. 4, University of California Press, 547–562,   1961.</mixed-citation></ref>
      <ref id="bib1.bibx51"><label>Restreppo et al.(2020)Restreppo, Wood, and Phrampus</label><mixed-citation>Restreppo, G. A., Wood, W. T., and Phrampus, B. J.: Oceanic sediment accumulation rates predicted via machine learning algorithm: towards sediment characterization on a global scale, Geo-Mar. Lett., 40, 755–763, <ext-link xlink:href="https://doi.org/10.1007/s00367-020-00669-1" ext-link-type="DOI">10.1007/s00367-020-00669-1</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx52"><label>Restreppo et al.(2021)Restreppo, Wood, and Phrampus</label><mixed-citation>Restreppo, G. A., Wood, W. T., and Phrampus, B. J.: A machine-learning derived model of seafloor sediment accumulation, Mar. Geol., 440, 106577, <ext-link xlink:href="https://doi.org/10.1016/j.margeo.2021.106577" ext-link-type="DOI">10.1016/j.margeo.2021.106577</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx53"><label>Romankevich et al.(2009)Romankevich, Vetrov, and Peresypkin</label><mixed-citation>Romankevich, E., Vetrov, A., and Peresypkin, V.: Organic matter of the World Ocean, Russ. Geol. Geophys., 50, 299–307, <ext-link xlink:href="https://doi.org/10.1016/j.rgg.2009.03.013" ext-link-type="DOI">10.1016/j.rgg.2009.03.013</ext-link>, 2009.</mixed-citation></ref>
      <ref id="bib1.bibx54"><label>Sala et al.(2021)Sala, Mayorga, Bradley, Cabral, Atwood, Auber, Cheung, Costello, Ferretti, Friedl, Gaines, Garilao, Goodell, Halpern, Hinson, Kaschner, Kesner-Reyes, Leprieur, McGowan, Morgan, Mouillot, Palacios-Abrantes, Possingham, Rechberger, Worm, and Lubchenco</label><mixed-citation>Sala, E., Mayorga, J., Bradley, D., Cabral, R. B., Atwood, T. B., Auber, A., Cheung, W., Costello, C., Ferretti, F., Friedl, er, A. M., Gaines, S. D., Garilao, C., Goodell, W., Halpern, B. S., Hinson, A., Kaschner, K., Kesner-Reyes, K., Leprieur, F., McGowan, J., Morgan, L. E., Mouillot, D., Palacios-Abrantes, J., Possingham, H. P., Rechberger, K. D., Worm, B., and Lubchenco, J.: Protecting the global ocean for biodiversity, food and climate, Nature, 592, 397–402, <ext-link xlink:href="https://doi.org/10.1038/s41586-021-03371-z" ext-link-type="DOI">10.1038/s41586-021-03371-z</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx55"><label>Seiter et al.(2004)Seiter, Hensen, Schröter, and Zabel</label><mixed-citation>Seiter, K., Hensen, C., Schröter, J., and Zabel, M.: Organic carbon content in surface sediments – defining regional provinces, Deep-Sea Res. Pt. I, 51, 2001–2026, <ext-link xlink:href="https://doi.org/10.1016/j.dsr.2004.06.014" ext-link-type="DOI">10.1016/j.dsr.2004.06.014</ext-link>, 2004.</mixed-citation></ref>
      <ref id="bib1.bibx56"><label>Shannon(1948)</label><mixed-citation>Shannon, C. E.: A mathematical theory of communication, The Bell System Technical Journal, 27, 379–423, <ext-link xlink:href="https://doi.org/10.1002/j.1538-7305.1948.tb01338.x" ext-link-type="DOI">10.1002/j.1538-7305.1948.tb01338.x</ext-link>, 1948.</mixed-citation></ref>
      <ref id="bib1.bibx57"><label>Song et al.(2022)Song, Santos, Yu, Wang, Burnett, Bianchi, Dong, Lian, Zhao, Mayer, Yao, Yu, and Xu</label><mixed-citation>Song, S., Santos, I. R., Yu, H., Wang, F., Burnett, W. C., Bianchi, T. S., Dong, J., Lian, E., Zhao, B., Mayer, L., Yao, Q., Yu, Z., and Xu, B.: A global assessment of the mixed layer in coastal sediments and implications for carbon storage, Nat. Commun., 13, 4903, <ext-link xlink:href="https://doi.org/10.1038/s41467-022-32650-0" ext-link-type="DOI">10.1038/s41467-022-32650-0</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx58"><label>Song et al.(2023)Song, Pang, Hou, Xu, Xue, Sun, and Meng</label><mixed-citation>Song, T., Pang, C., Hou, B., Xu, G., Xue, J., Sun, H., and Meng, F.: A review of artificial intelligence in marine science, Front. Earth Sci., 11, <ext-link xlink:href="https://doi.org/10.3389/feart.2023.1090185" ext-link-type="DOI">10.3389/feart.2023.1090185</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx59"><label>The HYCOM+NCODA Ocean Reanalysis(2014)</label><mixed-citation>The HYCOM+NCODA Ocean Reanalysis: 1/12 deg global HYCOM+NCODA Ocean Reanalysis, funded by: U.S. Navy and the Modeling and Simulation Coordination Office, <uri>https://www.hycom.org/data/glbu0pt08/expt-19pt1</uri> (last access: 19 March 2014), 2014. </mixed-citation></ref>
      <ref id="bib1.bibx60"><label>Thyng et al.(2016)Thyng, Greene, Hetland, Zimmerle, and DiMarco</label><mixed-citation>Thyng, K. M., Greene, C. A., Hetland, R. D., Zimmerle, H. M., and DiMarco, S. F.: True Colors of Oceanography, Oceanography, 29, 9–13, <ext-link xlink:href="https://doi.org/10.5670/oceanog.2016.66" ext-link-type="DOI">10.5670/oceanog.2016.66</ext-link>, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx61"><label>van der Voort et al.(2021)van der Voort, Blattmann, Usman, Montluçon, Loeffler, Tavagna, Gruber, and Eglinton</label><mixed-citation>van der Voort, T. S., Blattmann, T. M., Usman, M., Montluçon, D., Loeffler, T., Tavagna, M. L., Gruber, N., and Eglinton, T. I.: MOSAIC (Modern Ocean Sediment Archive and Inventory of Carbon): a (radio)carbon-centric database for seafloor surficial sediments, Earth Syst. Sci. Data, 13, 2135–2146, <ext-link xlink:href="https://doi.org/10.5194/essd-13-2135-2021" ext-link-type="DOI">10.5194/essd-13-2135-2021</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx62"><label>Wei et al.(2010)Wei, Rowe, Escobar-Briones, Boetius, Soltwedel, and Caley</label><mixed-citation>Wei, C.-L., Rowe, G. T., Escobar-Briones, E., Boetius, A., Soltwedel, T., and Caley, M. J.: Global patterns and predictions of seafloor biomass using random forests, PLoS ONE, 5, e15323, <ext-link xlink:href="https://doi.org/10.1371/journal.pone.0015323" ext-link-type="DOI">10.1371/journal.pone.0015323</ext-link>, 2010.</mixed-citation></ref>
      <ref id="bib1.bibx63"><label>Whittaker et al.(2013)Whittaker, Goncharov, Williams, Müller, and Leitchenkov</label><mixed-citation>Whittaker, J., Goncharov, A., Williams, S., Müller, R. D., and Leitchenkov, G.: Global sediment thickness dataset updated for the Australian-Antarctic Southern Ocean, Geochem. Geophy. Geosy., 14, 3297–3305, <ext-link xlink:href="https://doi.org/10.1002/ggge.20181" ext-link-type="DOI">10.1002/ggge.20181</ext-link>, 2013.</mixed-citation></ref>

  </ref-list></back>
    <!--<article-title-html>NN-TOC v1: global prediction of total organic carbon in marine sediments using deep neural networks</article-title-html>
<abstract-html/>
<ref-html id="bib1.bib1"><label>Arndt et al.(2013)Arndt, Jørgensen, LaRowe, Middelburg, Pancost, and
Regnier</label><mixed-citation>
      
Arndt, S., Jørgensen, B., LaRowe, D., Middelburg, J., Pancost, R., and
Regnier, P.: Quantifying the degradation of organic matter in marine
sediments: A review and synthesis, Earth-Sci. Rev., 123, 53–86,
<a href="https://doi.org/10.1016/j.earscirev.2013.02.008" target="_blank">https://doi.org/10.1016/j.earscirev.2013.02.008</a>, 2013.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib2"><label>Atwood et al.(2020)Atwood, Witt, Mayorga, Hammill, and
Sala</label><mixed-citation>
      
Atwood, T. B., Witt, A., Mayorga, J., Hammill, E., and Sala, E.: Global
Patterns in Marine Sediment Carbon Stocks, Frontiers in Marine Science, 7, 165,
<a href="https://doi.org/10.3389/fmars.2020.00165" target="_blank">https://doi.org/10.3389/fmars.2020.00165</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib3"><label>Baturin(2007)</label><mixed-citation>
      
Baturin, G. N.: Issue of the relationship between primary productivity of
organic carbon in ocean and phosphate accumulation (Holocene-Late Jurassic),
Lith. Miner. Resour., 42, 318–348,
<a href="https://doi.org/10.1134/S0024490207040025" target="_blank">https://doi.org/10.1134/S0024490207040025</a>, 2007.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib4"><label>Beazley(2003)</label><mixed-citation>
      
Beazley, M. J.: The significance of organic carbon and sediment surface area to
the benthic biogeochemistry of the slope and deep water environments of the
northern Gulf of Mexico, Master's thesis, Texas A&amp;M University,
<a href="http://hdl.handle.net/1969.1/534" target="_blank"/> (last access: 3 February 2024), 2003.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib5"><label>Becker et al.(2014)Becker, Wood, and Martin</label><mixed-citation>
      
Becker, J. J., Wood, W. T., and Martin, K. M.: Global Crustal Heat Flow
Using Random Decision Forest Prediction, in: AGU Fall Meeting Abstracts,
vol. 2014,  NG31A–3788,  <a href="https://ui.adsabs.harvard.edu/abs/2014AGUFMNG31A3788B/abstract" target="_blank"/> (last access: 3 February 2024), 2014.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib6"><label>Berner(1982)</label><mixed-citation>
      
Berner, R. A.: Burial of organic carbon and pyrite sulfur in the modern ocean:
its geochemical and environmental significance, Am. J. Sci., 282, 451–473,
<a href="https://doi.org/10.2475/ajs.282.4.451" target="_blank">https://doi.org/10.2475/ajs.282.4.451</a>, 1982.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib7"><label>Berner(2004)</label><mixed-citation>
      
Berner, R. A.: Processes of the Long-Term Carbon Cycle: Organic Matter and
Carbonate Burial and Weathering, in: The Phanerozoic Carbon Cycle: CO<sub>2</sub>
and O<sub>2</sub>, Oxford University Press, ISBN 9780195173338,
<a href="https://doi.org/10.1093/oso/9780195173338.003.0005" target="_blank">https://doi.org/10.1093/oso/9780195173338.003.0005</a>, 2004.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib8"><label>Boudreau(1997)</label><mixed-citation>
      
Boudreau, B. P.: Diagenetic models and their implementation, vol. 505, Springer
Berlin, ISBN 978-3-642-64399-6, <a href="https://doi.org/10.1007/978-3-642-60421-8" target="_blank">https://doi.org/10.1007/978-3-642-60421-8</a>, 1997.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib9"><label>Boyer et al.(2013)Boyer, Antonov, Baranova, Coleman, Garcia, and
Grodsky</label><mixed-citation>
      
Boyer, T. P., Antonov, J. I., Baranova, O. K., Coleman, C., Garcia, H. E., and
Grodsky, A.: World Ocean Database 2013, NOAA Atlas NESDIS 72, Technical Ed.
Silver Spring, MD, <a href="https://doi.org/10.7289/V5NZ85MT" target="_blank">https://doi.org/10.7289/V5NZ85MT</a>, 2013.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib10"><label>Bradley and Arndt(2022)</label><mixed-citation>
      
Bradley, J. A., Hülse, D., LaRowe, D. E., and Arndt, S.: Transfer efficiency of organic
carbon in marine sediments, Nat. Commun., 13, 7297,
<a href="https://doi.org/10.1038/s41467-022-35112-9" target="_blank">https://doi.org/10.1038/s41467-022-35112-9</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib11"><label>Burdige(2005)</label><mixed-citation>
      
Burdige, D. J.: Burial of terrestrial organic matter in marine sediments: A
re-assessment, Global Biogeochem. Cy., 19,  GB4011,
<a href="https://doi.org/10.1029/2004GB002368" target="_blank">https://doi.org/10.1029/2004GB002368</a>, 2005.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib12"><label>Burdige(2007)</label><mixed-citation>
      
Burdige, D. J.: Preservation of Organic Matter in Marine Sediments – Controls,
Mechanisms, and an Imbalance in Sediment Organic Carbon Budgets?, Chem.
Rev., 107, 467–485, <a href="https://doi.org/10.1021/cr050347q" target="_blank">https://doi.org/10.1021/cr050347q</a>, 2007.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib13"><label>Crameri(2023)</label><mixed-citation>
      
Crameri, F.: Scientific colour maps, Zenodo [data set], <a href="https://doi.org/10.5281/zenodo.8409685" target="_blank">https://doi.org/10.5281/zenodo.8409685</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib14"><label>De Veaux and Ungar(1994)</label><mixed-citation>
      
De Veaux, R. D. and Ungar, L. H.: Multicollinearity: A tale of two
nonparametric regressions, in: Selecting Models from Data, edited by:
Cheeseman, P. and Oldford, R. W., Springer New York, New York,
NY, 393–402,  ISBN 978-1-4612-2660-4, 1994.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib15"><label>Diesing et al.(2021)Diesing, Thorsnes, and
Bjarnadóttir</label><mixed-citation>
      
Diesing, M., Thorsnes, T., and Bjarnadóttir, L. R.: Organic carbon densities and accumulation rates in surface sediments of the North Sea and Skagerrak, Biogeosciences, 18, 2139–2160, <a href="https://doi.org/10.5194/bg-18-2139-2021" target="_blank">https://doi.org/10.5194/bg-18-2139-2021</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib16"><label>Emerson and Hedges(1988)</label><mixed-citation>
      
Emerson, S. and Hedges, J. I.: Processes controlling the organic carbon content
of open ocean sediments, Paleoceanography, 3, 621–634,
<a href="https://doi.org/10.1029/PA003i005p00621" target="_blank">https://doi.org/10.1029/PA003i005p00621</a>, 1988.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib17"><label>Emery(1968)</label><mixed-citation>
      
Emery, K. O.: Relict sediments on continental shelves of the world, Am. Assoc.
Petr. Geol. B., 52, 445–464, 1968.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib18"><label>Flanders Marine Institute(2021)</label><mixed-citation>
      
Flanders Marine Institute: Global Oceans and Seas, version 1,
<a href="https://doi.org/10.14284/542" target="_blank">https://doi.org/10.14284/542</a>,   <a href="https://www.marineregions.org/" target="_blank"/> (last access: 25 August 2023), 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib19"><label>Gal and Ghahramani(2016)</label><mixed-citation>
      
Gal, Y. and Ghahramani, Z.: Dropout as a Bayesian Approximation: Representing
Model Uncertainty in Deep Learning, in: Proceedings of The 33rd International
Conference on Machine Learning, New York, New York, USA,  edited by: Balcan, M. F. and Weinberger,
K. Q., Proceedings of Machine Learning Research,  vol. 48,
PMLR,  19 June 2016, 1050–1059,
<a href="https://proceedings.mlr.press/v48/gal16.html" target="_blank"/> (last access: 24 February 2022), 2016.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib20"><label>Garlan et al.(2018)Garlan, Gabelotaud, Lucas, and
Marchès</label><mixed-citation>
      
Garlan, T., Gabelotaud, I., Lucas, S., and Marchès, E.: A World Map of Seabed
Sediment Based on 50 Years of Knowledge, World Academy of Science,
Engineering and Technology, International Journal of Geological and
Environmental Engineering, 12, 409–419, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib21"><label>Goyet et al.(2000)Goyet, Healy, and Ryan</label><mixed-citation>
      
Goyet, C., Healy, R., and Ryan, J.: Global Distribution of Total Inorganic
Carbon and Total Alkalinity Below the Deepest Winter Mixed Layer Depths,
ORNIJCDIAC-127 NDP-076,  <a href="https://doi.org/10.3334/zCDIAC/otg.ndp076" target="_blank">https://doi.org/10.3334/zCDIAC/otg.ndp076</a>, 2000.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib22"><label>Hall(2002)</label><mixed-citation>
      
Hall, S. J.: The continental shelf benthic ecosystem: current status, agents
for change and future prospects, Environ. Conserv., 29, 350–374,
<a href="http://www.jstor.org/stable/44520615" target="_blank"/> (last access: 26 June 2022), 2002.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib23"><label>Hart-Davis et al.(2021)Hart-Davis, Piccioni, Dettmering, Schwatke,
Passaro, and Seitz</label><mixed-citation>
      
Hart-Davis, M. G., Piccioni, G., Dettmering, D., Schwatke, C., Passaro, M., and Seitz, F.: EOT20: a global ocean tide model from multi-mission satellite altimetry, Earth Syst. Sci. Data, 13, 3869–3884, <a href="https://doi.org/10.5194/essd-13-3869-2021" target="_blank">https://doi.org/10.5194/essd-13-3869-2021</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib24"><label>He et al.(2015)He, Zhang, Ren, and Sun</label><mixed-citation>
      
He, K., Zhang, X., Ren, S., and Sun, J.: Delving deep into rectifiers: Surpassing human-level performance on imagenet classification, in: Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile,  1026–1034, 7–13 December 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib25"><label>Hedges and Keil(1995)</label><mixed-citation>
      
Hedges, J. I. and Keil, R. G.: Sedimentary organic matter preservation: an
assessment and speculative synthesis, Mar. Chem., 49, 81–115,
<a href="https://doi.org/10.1016/0304-4203(95)00008-F" target="_blank">https://doi.org/10.1016/0304-4203(95)00008-F</a>, 1995.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib26"><label>Joint Panel on Oceanographic Tables and
Standards(1991)</label><mixed-citation>
      
Joint Panel on Oceanographic Tables and Standards: Processing of
Oceanographic Station Data, UNESCO, Paris, ISBN 978-92-3-102756-7,
1991.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib27"><label>Jørgensen et al.(2022)Jørgensen, Wenzhöfer, Egger, and
Glud</label><mixed-citation>
      
Jørgensen, B. B., Wenzhöfer, F., Egger, M., and Glud, R. N.: Sediment oxygen
consumption: Role in the global marine carbon cycle, Earth-Sci. Rev.,
228, 103987, <a href="https://doi.org/10.1016/j.earscirev.2022.103987" target="_blank">https://doi.org/10.1016/j.earscirev.2022.103987</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib28"><label>Kim and Wessel(2011)</label><mixed-citation>
      
Kim, S. S. and Wessel, P.: New global seamount census from the
altimetry-derived gravity data, Geophys. J. Int., 186,
615–631, <a href="https://doi.org/10.1111/j.1365-246X.2011.05076.x" target="_blank">https://doi.org/10.1111/j.1365-246X.2011.05076.x</a>,
2011.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib29"><label>Kullback and Leibler(1951)</label><mixed-citation>
      
Kullback, S. and Leibler, R. A.: On Information and Sufficiency, Ann. Math.
Stat., 22, 79–86, <a href="http://www.jstor.org/stable/2236703" target="_blank"/> (last access: 19 April 2023), 1951.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib30"><label>LaRowe et al.(2020a)LaRowe, Arndt, Bradley, Estes,
Hoarfrost, Lang, Lloyd, Mahmoudi, Orsi, Shah Walter, Steen, and
Zhao</label><mixed-citation>
      
LaRowe, D., Arndt, S., Bradley, J., Estes, E., Hoarfrost, A., Lang, S., Lloyd,
K., Mahmoudi, N., Orsi, W., Shah Walter, S., Steen, A., and Zhao, R.: The
fate of organic carbon in marine sediments – New insights from recent data
and analysis, Earth-Sci. Rev., 204, 103146,
<a href="https://doi.org/10.1016/j.earscirev.2020.103146" target="_blank">https://doi.org/10.1016/j.earscirev.2020.103146</a>, 2020a.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib31"><label>LaRowe et al.(2020b)LaRowe, Arndt, Bradley, Burwicz,
Dale, and Amend</label><mixed-citation>
      
LaRowe, D. E., Arndt, S., Bradley, J. A., Burwicz, E., Dale, A. W., and Amend,
J. P.: Organic carbon and microbial activity in marine sediments on a global
scale throughout the Quaternary, Geochim. Cosmochim. Ac., 286,
227–247, <a href="https://doi.org/10.1016/j.gca.2020.07.017" target="_blank">https://doi.org/10.1016/j.gca.2020.07.017</a>,
2020b.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib32"><label>LeCun et al.(2015)LeCun, Bengio, and Hinton</label><mixed-citation>
      
LeCun, Y., Bengio, Y., and Hinton, G.: Deep learning, Nature, 521, 436–444,
<a href="https://doi.org/10.1038/nature14539" target="_blank">https://doi.org/10.1038/nature14539</a>, 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib33"><label>Lee et al.(2019)Lee, Wood, and Phrampus</label><mixed-citation>
      
Lee, T. R., Wood, W. T., and Phrampus, B. J.: A Machine Learning (kNN) Approach
to Predicting Global Seafloor Total Organic Carbon, Global Biogeochem. Cy., 33, 37–46, <a href="https://doi.org/10.1029/2018GB005992" target="_blank">https://doi.org/10.1029/2018GB005992</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib34"><label>Lee et al.(2020)Lee, Wood, Skarke, Phrampus, and
Obelcz</label><mixed-citation>
      
Lee, T. R., Wood, W. T., Skarke, A., Phrampus, B. J., and Obelcz, J.: Data
files associated with the <i>k</i>-nearest neighbor global prediction of isopachs
for present to middle Miocene, Zenodo [data set], <a href="https://doi.org/10.5281/zenodo.3675364" target="_blank">https://doi.org/10.5281/zenodo.3675364</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib35"><label>Legge et al.(2020)Legge, Johnson, Hicks, Jickells, Diesing, Aldridge,
Andrews, Artioli, Bakker, Burrows, Carr, Cripps, Felgate, Fernand, Greenwood,
Hartman, Kröger, Lessin, Mahaffey, Mayor, Parker, Queirós, Shutler, Silva,
Stahl, Tinker, Underwood, Van Der Molen, Wakelin, Weston, and
Williamson</label><mixed-citation>
      
Legge, O., Johnson, M., Hicks, N., Jickells, T., Diesing, M., Aldridge, J.,
Andrews, J., Artioli, Y., Bakker, D. C. E., Burrows, M. T., Carr, N., Cripps,
G., Felgate, S. L., Fernand, L., Greenwood, N., Hartman, S., Kröger, S.,
Lessin, G., Mahaffey, C., Mayor, D. J., Parker, R., Queirós, A. M., Shutler,
J. D., Silva, T., Stahl, H., Tinker, J., Underwood, G. J. C., Van Der Molen,
J., Wakelin, S., Weston, K., and Williamson, P.: Carbon on the Northwest
European Shelf: Contemporary Budget and Future Influences, Frontiers in
Marine Science, 7, 143, <a href="https://doi.org/10.3389/fmars.2020.00143" target="_blank">https://doi.org/10.3389/fmars.2020.00143</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib36"><label>Lucas and Giles(2016)</label><mixed-citation>
      
Lucas, M. and Giles, H.: Quantifying Uncertainty in Random Forests via
Confidence Intervals and Hypothesis Tests, J. Mach. Learn.
Res., 17, 1–41, <a href="http://jmlr.org/papers/v17/14-168.html" target="_blank"/> (last access: 4 August 2023),
2016.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib37"><label>Ludwig et al.(2011)Ludwig, Amiotte-Suchet, and Probst</label><mixed-citation>
      
Ludwig, W., Amiotte-Suchet, P., and Probst, J. L.: ISLSCP II Global River
Fluxes of Carbon and Sediments to the Oceans,  ORNL Distributed Active Archive Center [data set], <a href="https://doi.org/10.3334/ORNLDAAC/1028" target="_blank">https://doi.org/10.3334/ORNLDAAC/1028</a>,
2011.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib38"><label>Lundberg and Lee(2017)</label><mixed-citation>
      
Lundberg, S. M. and Lee, S.-I.: A unified approach to interpreting model
predictions, in: Proceedings of the 31st International Conference on Neural
Information Processing Systems, NIPS'17,  Curran Associates
Inc., Red Hook, NY, USA, 4–9 December 2017, 4768–4777, ISBN 9781510860964, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib39"><label>Martin et al.(2015)Martin, Wood, and Becker</label><mixed-citation>
      
Martin, K. M., Wood, W. T., and Becker, J. J.: A global prediction of seafloor
sediment porosity using machine learning, Geophys. Res. Lett., 42,
10640–10646, <a href="https://doi.org/10.1002/2015GL065279" target="_blank">https://doi.org/10.1002/2015GL065279</a>, 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib40"><label>NASA(2011)</label><mixed-citation>
      
NASA: Announcement of Aquarius Level 2 Data Availability, Physical
Oceanography Distributed Active Archive Center (PODAAC),
<a href="https://aquarius.oceansciences.org/cgi/gal_density.htm" target="_blank"/> (last access: 23 October 2023), 2011.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib41"><label>NASA(2014)</label><mixed-citation>
      
NASA: MODIS-Aqua Ocean Color Data, Goddard Space Flight Center, Ocean
Ecology Laboratory, Ocean Biology Processing Group,
<a href="https://doi.org/10.5067/AQUA/MODIS_OC.2014.0" target="_blank">https://doi.org/10.5067/AQUA/MODIS_OC.2014.0</a>, 2014.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib42"><label>National Geophysical Data Center(1976)</label><mixed-citation>
      
National Geophysical Data Center: The NGDC Seafloor Sediment Grain Size
Database, first Version, NOAA National Centers for Environmental Information,
<a href="https://doi.org/10.7289/V5G44N6W" target="_blank">https://doi.org/10.7289/V5G44N6W</a>, 1976.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib43"><label>National Geophysical Data Center(2006)</label><mixed-citation>
      
National Geophysical Data Center: 2-minute Gridded Global Relief Data
(ETOPO2) v2, NCEI [data set], <a href="https://doi.org/10.7289/V5J1012Q" target="_blank">https://doi.org/10.7289/V5J1012Q</a>,  2006.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib44"><label>Pape et al.(2020)Pape, Bünz, Hong, Torres, Riedel, Panieri, Lepland,
Hsu, Wintersteller, Wallmann, Schmidt, Yao, and Bohrmann</label><mixed-citation>
      
Pape, T., Bünz, S., Hong, W.-L., Torres, M. E., Riedel, M., Panieri, G.,
Lepland, A., Hsu, C.-W., Wintersteller, P., Wallmann, K., Schmidt, C., Yao,
H., and Bohrmann, G.: Origin and Transformation of Light Hydrocarbons
Ascending at an Active Pockmark on Vestnesa Ridge, Arctic Ocean, J.
Geophys. Res.-Sol. Ea., 125, e2018JB016679,
<a href="https://doi.org/10.1029/2018JB016679" target="_blank">https://doi.org/10.1029/2018JB016679</a>,  2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib45"><label>Paradis et al.(2023)Paradis, Nakajima, Van der Voort, Gies,
Wildberger, Blattmann, Bröder, and Eglinton</label><mixed-citation>
      
Paradis, S., Nakajima, K., Van der Voort, T. S., Gies, H., Wildberger, A., Blattmann, T. M., Bröder, L., and Eglinton, T. I.: The Modern Ocean Sediment Archive and Inventory of Carbon (MOSAIC): version 2.0, Earth Syst. Sci. Data, 15, 4105–4125, <a href="https://doi.org/10.5194/essd-15-4105-2023" target="_blank">https://doi.org/10.5194/essd-15-4105-2023</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib46"><label>Parameswaran et al.(2024a)Parameswaran, González,
Burwicz-Galerne, Braack, and Wallmann</label><mixed-citation>
      
Parameswaran, N., González, E., Burwicz-Galerne, E., Braack, M., and Wallmann,
K.: Code for Global Prediction Of Total Organic Carbon In Marine Sediments
Using Deep Neural Networks (nn-toc), GitHub [code],
<a href="https://github.com/paramnav/nn-toc" target="_blank"/> (last access: 13 November 2024), 2024a.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib47"><label>Parameswaran et al.(2024b)Parameswaran, González,
Burwicz-Galerne, Braack, and Wallmann</label><mixed-citation>
      
Parameswaran, N., González, E., Burwicz-Galerne, E., Braack, M., and Wallmann,
K.: Dataset for the Global Prediction Of Total Organic Carbon In Marine
Sediments Using Deep Neural Networks (nn-toc), Zenodo [data set], <a href="https://doi.org/10.5281/zenodo.11186224" target="_blank">https://doi.org/10.5281/zenodo.11186224</a>,
2024b.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib48"><label>Pavlis et al.(2008)Pavlis, Holmes, Kenyon, and
Factor</label><mixed-citation>
      
Pavlis, N. K., Holmes, S. A., Kenyon, S. C., and Factor, J. K.: The EGM2008
Global Gravitational Model, abstract 2008AGUFM.G22A..01P,
2008 General Assembly of the European Geosciences Union, Vienna, Austria,
<a href="https://ui.adsabs.harvard.edu/abs/2008AGUFM.G22A..01P" target="_blank"/> (last access: 7 October 2014), 2008.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib49"><label>Phrampus et al.(2019)Phrampus, Lee, and
Wood</label><mixed-citation>
      
Phrampus, B. J., Lee, T. R., and Wood, W. T.: Predictor Grids for “A Global
Probabilistic Prediction of Cold Seeps and Associated Seafloor Fluid
Expulsion Anomalies (SEAFLEAs)”, Zenodo [data set], <a href="https://doi.org/10.5281/zenodo.3459805" target="_blank">https://doi.org/10.5281/zenodo.3459805</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib50"><label>Rényi(1961)</label><mixed-citation>
      
Rényi, A.: On measures of entropy and information, in: Proceedings of the
fourth Berkeley symposium on mathematical statistics and probability, volume
1: contributions to the theory of statistics, Fourth Berkley Symposum on Mathematical Statistics and Probablity June 20-July 30, 1960, Statistical Laboratory of the University of California vol. 4,
University of California Press, 547–562,   1961.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib51"><label>Restreppo et al.(2020)Restreppo, Wood, and Phrampus</label><mixed-citation>
      
Restreppo, G. A., Wood, W. T., and Phrampus, B. J.: Oceanic sediment
accumulation rates predicted via machine learning algorithm: towards sediment
characterization on a global scale, Geo-Mar. Lett., 40, 755–763,
<a href="https://doi.org/10.1007/s00367-020-00669-1" target="_blank">https://doi.org/10.1007/s00367-020-00669-1</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib52"><label>Restreppo et al.(2021)Restreppo, Wood, and Phrampus</label><mixed-citation>
      
Restreppo, G. A., Wood, W. T., and Phrampus, B. J.: A machine-learning derived
model of seafloor sediment accumulation, Mar. Geol., 440, 106577,
<a href="https://doi.org/10.1016/j.margeo.2021.106577" target="_blank">https://doi.org/10.1016/j.margeo.2021.106577</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib53"><label>Romankevich et al.(2009)Romankevich, Vetrov, and
Peresypkin</label><mixed-citation>
      
Romankevich, E., Vetrov, A., and Peresypkin, V.: Organic matter of the World
Ocean, Russ. Geol. Geophys., 50, 299–307,
<a href="https://doi.org/10.1016/j.rgg.2009.03.013" target="_blank">https://doi.org/10.1016/j.rgg.2009.03.013</a>, 2009.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib54"><label>Sala et al.(2021)Sala, Mayorga, Bradley, Cabral, Atwood, Auber,
Cheung, Costello, Ferretti, Friedl, Gaines, Garilao, Goodell, Halpern,
Hinson, Kaschner, Kesner-Reyes, Leprieur, McGowan, Morgan, Mouillot,
Palacios-Abrantes, Possingham, Rechberger, Worm, and Lubchenco</label><mixed-citation>
      
Sala, E., Mayorga, J., Bradley, D., Cabral, R. B., Atwood, T. B., Auber, A.,
Cheung, W., Costello, C., Ferretti, F., Friedl, er, A. M., Gaines, S. D.,
Garilao, C., Goodell, W., Halpern, B. S., Hinson, A., Kaschner, K.,
Kesner-Reyes, K., Leprieur, F., McGowan, J., Morgan, L. E., Mouillot, D.,
Palacios-Abrantes, J., Possingham, H. P., Rechberger, K. D., Worm, B., and
Lubchenco, J.: Protecting the global ocean for biodiversity, food and
climate, Nature, 592, 397–402, <a href="https://doi.org/10.1038/s41586-021-03371-z" target="_blank">https://doi.org/10.1038/s41586-021-03371-z</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib55"><label>Seiter et al.(2004)Seiter, Hensen, Schröter, and
Zabel</label><mixed-citation>
      
Seiter, K., Hensen, C., Schröter, J., and Zabel, M.: Organic carbon content in
surface sediments – defining regional provinces, Deep-Sea Res. Pt. I, 51, 2001–2026,
<a href="https://doi.org/10.1016/j.dsr.2004.06.014" target="_blank">https://doi.org/10.1016/j.dsr.2004.06.014</a>, 2004.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib56"><label>Shannon(1948)</label><mixed-citation>
      
Shannon, C. E.: A mathematical theory of communication, The Bell System
Technical Journal, 27, 379–423, <a href="https://doi.org/10.1002/j.1538-7305.1948.tb01338.x" target="_blank">https://doi.org/10.1002/j.1538-7305.1948.tb01338.x</a>,
1948.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib57"><label>Song et al.(2022)Song, Santos, Yu, Wang, Burnett, Bianchi, Dong,
Lian, Zhao, Mayer, Yao, Yu, and Xu</label><mixed-citation>
      
Song, S., Santos, I. R., Yu, H., Wang, F., Burnett, W. C., Bianchi, T. S.,
Dong, J., Lian, E., Zhao, B., Mayer, L., Yao, Q., Yu, Z., and Xu, B.: A
global assessment of the mixed layer in coastal sediments and implications
for carbon storage, Nat. Commun., 13, 4903,
<a href="https://doi.org/10.1038/s41467-022-32650-0" target="_blank">https://doi.org/10.1038/s41467-022-32650-0</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib58"><label>Song et al.(2023)Song, Pang, Hou, Xu, Xue, Sun, and
Meng</label><mixed-citation>
      
Song, T., Pang, C., Hou, B., Xu, G., Xue, J., Sun, H., and Meng, F.: A review
of artificial intelligence in marine science, Front. Earth Sci., 11,
<a href="https://doi.org/10.3389/feart.2023.1090185" target="_blank">https://doi.org/10.3389/feart.2023.1090185</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib59"><label>The HYCOM+NCODA Ocean Reanalysis(2014)</label><mixed-citation>
      
The HYCOM+NCODA Ocean Reanalysis: 1/12 deg global HYCOM+NCODA Ocean
Reanalysis, funded
by: U.S. Navy and the Modeling and Simulation Coordination Office,
<a href="https://www.hycom.org/data/glbu0pt08/expt-19pt1" target="_blank"/> (last
access: 19 March 2014), 2014.


    </mixed-citation></ref-html>
<ref-html id="bib1.bib60"><label>Thyng et al.(2016)Thyng, Greene, Hetland, Zimmerle, and
DiMarco</label><mixed-citation>
      
Thyng, K. M., Greene, C. A., Hetland, R. D., Zimmerle, H. M., and DiMarco,
S. F.: True Colors of Oceanography, Oceanography, 29, 9–13, <a href="https://doi.org/10.5670/oceanog.2016.66" target="_blank">https://doi.org/10.5670/oceanog.2016.66</a>,
2016.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib61"><label>van der Voort et al.(2021)van der Voort, Blattmann, Usman,
Montluçon, Loeffler, Tavagna, Gruber, and Eglinton</label><mixed-citation>
      
van der Voort, T. S., Blattmann, T. M., Usman, M., Montluçon, D., Loeffler, T., Tavagna, M. L., Gruber, N., and Eglinton, T. I.: MOSAIC (Modern Ocean Sediment Archive and Inventory of Carbon): a (radio)carbon-centric database for seafloor surficial sediments, Earth Syst. Sci. Data, 13, 2135–2146, <a href="https://doi.org/10.5194/essd-13-2135-2021" target="_blank">https://doi.org/10.5194/essd-13-2135-2021</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib62"><label>Wei et al.(2010)Wei, Rowe, Escobar-Briones, Boetius, Soltwedel, and
Caley</label><mixed-citation>
      
Wei, C.-L., Rowe, G. T., Escobar-Briones, E., Boetius, A., Soltwedel, T., and
Caley, M. J.: Global patterns and predictions of seafloor biomass using
random forests, PLoS ONE, 5, e15323, <a href="https://doi.org/10.1371/journal.pone.0015323" target="_blank">https://doi.org/10.1371/journal.pone.0015323</a>,
2010.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib63"><label>Whittaker et al.(2013)Whittaker, Goncharov, Williams, Müller, and
Leitchenkov</label><mixed-citation>
      
Whittaker, J., Goncharov, A., Williams, S., Müller, R. D., and Leitchenkov,
G.: Global sediment thickness dataset updated for the Australian-Antarctic
Southern Ocean, Geochem. Geophy. Geosy., 14, 3297–3305,
<a href="https://doi.org/10.1002/ggge.20181" target="_blank">https://doi.org/10.1002/ggge.20181</a>, 2013.

    </mixed-citation></ref-html>--></article>
