<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing with OASIS Tables v3.0 20080202//EN" "https://jats.nlm.nih.gov/nlm-dtd/publishing/3.0/journalpub-oasis3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:oasis="http://docs.oasis-open.org/ns/oasis-exchange/table" xml:lang="en" dtd-version="3.0" article-type="research-article">
  <front>
    <journal-meta><journal-id journal-id-type="publisher">GMD</journal-id><journal-title-group>
    <journal-title>Geoscientific Model Development</journal-title>
    <abbrev-journal-title abbrev-type="publisher">GMD</abbrev-journal-title><abbrev-journal-title abbrev-type="nlm-ta">Geosci. Model Dev.</abbrev-journal-title>
  </journal-title-group><issn pub-type="epub">1991-9603</issn><publisher>
    <publisher-name>Copernicus Publications</publisher-name>
    <publisher-loc>Göttingen, Germany</publisher-loc>
  </publisher></journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.5194/gmd-19-7767-2026</article-id><title-group><article-title>Optimizing Gaussian process emulation and generalized additive model fitting for rapid, reproducible earth system model analysis</article-title><alt-title>Optimizing GP emulation and GAM fitting</alt-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author" corresp="yes" rid="aff1">
          <name><surname>Ghosh</surname><given-names>Kunal</given-names></name>
          <email>k.ghosh@leeds.ac.uk</email>
        <ext-link>https://orcid.org/0000-0002-3179-6844</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1 aff2 aff3">
          <name><surname>Regayre</surname><given-names>Leighton A.</given-names></name>
          
        <ext-link>https://orcid.org/0000-0003-2699-929X</ext-link></contrib>
        <aff id="aff1"><label>1</label><institution>School of Earth and Environment, University of Leeds, Leeds, LS2 9JT, UK</institution>
        </aff>
        <aff id="aff2"><label>2</label><institution>Met Office Hadley Centre, Exeter, Fitzroy Road, Exeter, Devon, EX1 3PB, UK</institution>
        </aff>
        <aff id="aff3"><label>3</label><institution>Centre for Environmental Modelling and Computation, University of Leeds, Leeds, LS2 9JT, UK</institution>
        </aff>
      </contrib-group>
      <author-notes><corresp id="corr1">Kunal Ghosh (k.ghosh@leeds.ac.uk)</corresp></author-notes><pub-date><day>21</day><month>August</month><year>2026</year></pub-date>
      
      <volume>19</volume>
      <issue>16</issue>
      <fpage>7767</fpage><lpage>7785</lpage>
      <history>
        <date date-type="received"><day>7</day><month>November</month><year>2025</year></date>
           <date date-type="rev-request"><day>21</day><month>December</month><year>2025</year></date>
           <date date-type="rev-recd"><day>12</day><month>July</month><year>2026</year></date>
           <date date-type="accepted"><day>12</day><month>August</month><year>2026</year></date>
      </history>
      <permissions>
        <copyright-statement>Copyright: © 2026 Kunal Ghosh</copyright-statement>
        <copyright-year>2026</copyright-year>
      <license license-type="open-access"><license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p></license></permissions><self-uri xlink:href="https://gmd.copernicus.org/articles/19/7767/2026/gmd-19-7767-2026.html">This article is available from https://gmd.copernicus.org/articles/19/7767/2026/gmd-19-7767-2026.html</self-uri><self-uri xlink:href="https://gmd.copernicus.org/articles/19/7767/2026/gmd-19-7767-2026.pdf">The full text article is available as a PDF file from https://gmd.copernicus.org/articles/19/7767/2026/gmd-19-7767-2026.pdf</self-uri>
      <abstract><title>Abstract</title>

      <p id="d2e103">Causes of model uncertainty in complex modeling systems can be identified using large perturbed-parameter ensembles (PPEs), combined with statistical emulators to increase sample size and enable variance-based sensitivity analyses and observational constraint. In global climate models such as the UK Earth System Model (UKESM), these approaches are typically applied at the global or regional mean scales for a limited set of variables. Accelerating progress in understanding the multi-faceted causes of climate model uncertainty, requires implementing such workflows at the model grid box scale, to enable analyses across variables that reveal how uncertainties propagate and interact spatially. However, this approach requires training millions of Gaussian process (GP) emulators and fitting an equal number of generalized additive models (GAMs) – a major computational bottleneck. We present a high-performance, open-source pipeline that introduces optimisations for this workflow. For GP emulation, we implement task-level parallelism and streamlined data handling on high-performance computing systems. For GAM fitting, we integrate a parallelized <monospace>pyGAM</monospace> interface with R's <monospace>mgcv::bam()</monospace> back end, using fast fREML estimation with discrete smoothing, memory-efficient batching, and improved input–output routines. These changes reduce GP training time by 97.5 % (6177 s <inline-formula><mml:math id="M1" display="inline"><mml:mo>→</mml:mo></mml:math></inline-formula> 154 s) and GAM fitting time by 95.2 % (10 623 s <inline-formula><mml:math id="M2" display="inline"><mml:mo>→</mml:mo></mml:math></inline-formula> 511 s), yielding a <inline-formula><mml:math id="M3" display="inline"><mml:mo>∼</mml:mo></mml:math></inline-formula> 25 <inline-formula><mml:math id="M4" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> faster end-to-end workflow (96 % total runtime reduction) and cutting peak memory use by a factor of 12. Outputs are numerically identical to the baseline implementation (Pearson correlation <inline-formula><mml:math id="M5" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 1.00 for both GP and GAM predictions). We demonstrate the approach using a UKESM PPE comprising 221 members scaled up to 1 million using GP emulators,  and GAM fits applied to output for a single target variable, to show that the improved performance enables multi-variable, higher-resolution, and potentially multi-model analyses that were previously impractical. These improvements pave the way for PPE studies to scale in scope without compromising statistical fidelity, enabling more comprehensive exploration of model parameter uncertainty within feasible HPC budgets.</p>
  </abstract>
    
<funding-group>
<award-group id="gs1">
<funding-source>Natural Environment Research Council</funding-source>
<award-id>NE/P013406/1</award-id>
<award-id>NE/X013901/1</award-id>
<award-id>NE/G006148/1</award-id>
</award-group>
<award-group id="gs2">
<funding-source>Horizon 2020</funding-source>
<award-id>821205</award-id>
</award-group>
<award-group id="gs3">
<funding-source>Department for Science, Innovation and Technology</funding-source>
<award-id>Met Office Hadley Centre Climate Programme</award-id>
</award-group>
</funding-group>
</article-meta>
  </front>
<body>
      

<sec id="Ch1.S1" sec-type="intro">
  <label>1</label><title>Introduction</title>
      <p id="d2e157">Perturbed-parameter ensemble (PPE) studies are increasingly recognised as a vital tool in climate and Earth system modelling because they provide information needed to characterise model uncertainty and identify its causes. Well designed PPEs enable systematic identification of structural model deficiencies, guide targeted model developments, and provide a basis for rigorous statistical calibration against observations <xref ref-type="bibr" rid="bib1.bibx3" id="paren.1"/>, unlike multi-model ensembles <xref ref-type="bibr" rid="bib1.bibx32 bib1.bibx4" id="paren.2"/>, which conflate structural and parametric uncertainties in uncontrolled ways. Community assessments consistently argue that PPEs should be prioritised alongside increases in model complexity and resolution, as they offer one of the best opportunities to improve model reliability and provide robust uncertainty quantification for downstream impacts and decision making <xref ref-type="bibr" rid="bib1.bibx10 bib1.bibx3" id="paren.3"/>.</p>
      <p id="d2e169">Complex Earth system models are computationally expensive, which typically limits PPEs to hundreds of members. To enable robust statistical analysis of parameter sensitivities and uncertainty, these ensembles are often augmented using model emulation, allowing exploration of millions of parameter combinations at relatively low computational cost <xref ref-type="bibr" rid="bib1.bibx29" id="paren.4"/>. While this expansion makes statistical analyses possible, it introduces additional computational demands during analysis, particularly for variance decomposition and sensitivity analysis across the enlarged parameter space.</p>
      <p id="d2e175">Recent progress harnessing the power of climate model PPEs has utilised statistical emulators, such as Gaussian Process (GP) surrogates <xref ref-type="bibr" rid="bib1.bibx16 bib1.bibx17" id="paren.5"/>, which enable fast approximations of expensive model components, and flexible smoothers such as generalised additive models (GAMs) <xref ref-type="bibr" rid="bib1.bibx36 bib1.bibx27" id="paren.6"/> to decompose variance, identify dominant sources of uncertainty, and provide interpretable functional relationships <xref ref-type="bibr" rid="bib1.bibx23 bib1.bibx18" id="paren.7"><named-content content-type="pre">e.g.</named-content></xref>. Ideally, these techniques would be applied at existing model temporal and spatial resolution. However, the lack of scalability in constructing GP emulators, sampling millions of parameter combinations, and subsequently fitting and evaluating GAMs severely limits the scope of climate model PPE analyses.</p>
      <p id="d2e189">The major computational bottleneck in the widespread implementation of statistical emulation of PPEs and subsequent variance decomposition arises from the cost of evaluating multiple model variables within each model grid box, a necessary step to establish interpretable functional relationships. For each model variable, applying these methods at moderate temporal and spatial resolution requires training and evaluating large numbers of GP emulators and GAMs, resulting in runtimes of tens of hours and memory requirements of hundreds of gigabytes <xref ref-type="bibr" rid="bib1.bibx39 bib1.bibx7 bib1.bibx12 bib1.bibx8" id="paren.8"/>. As a result, prior applications have often been restricted to single variables <xref ref-type="bibr" rid="bib1.bibx23" id="paren.9"/>, individual months or seasons <xref ref-type="bibr" rid="bib1.bibx18" id="paren.10"/>, or coarse regional aggregation <xref ref-type="bibr" rid="bib1.bibx21 bib1.bibx9 bib1.bibx22" id="paren.11"/>. This limited scope constrains the ability to fully exploit PPEs for multi-variable, multi-region, and multi-model analyses.</p>
      <p id="d2e205">Several frameworks have advanced the practical use of statistical emulators in climate science. <xref ref-type="bibr" rid="bib1.bibx34" id="text.12"/> introduced the Earth System Emulator (ESEm), a general framework for building GP surrogates across a range of model components, while <xref ref-type="bibr" rid="bib1.bibx38" id="text.13"/> proposed an additive GP approach to improve interpretability of parameter–output relationships. Multi-scale GP methods <xref ref-type="bibr" rid="bib1.bibx33" id="paren.14"/> and multi-fidelity emulators <xref ref-type="bibr" rid="bib1.bibx28" id="paren.15"/> have also addressed scaling challenges for remote sensing and radiative-transfer models. However, these approaches rely on statistical approximations or structural modifications of the emulator, which may reduce accuracy or reproducibility compared to directly leveraging high-performance computing resources. Similar challenges have been reported for scaling GAMs: <xref ref-type="bibr" rid="bib1.bibx37" id="text.16"/> showed that fitting spatiotemporal GAMs to gigadata from the UK Black Smoke monitoring network (almost 10 million daily observations) was computationally prohibitive using conventional methods, with model matrices alone requiring hundreds of gigabytes of memory. Their solution combined discretisation of covariates with parallelised matrix operations to achieve order-of-magnitude speed-ups, reducing runtimes from months to hours, though at the cost of some flexibility and potential loss of fine-scale precision. Together, these examples underline both the importance and the difficulty of deploying GP and GAM methods at scale, and highlight the urgent need for further innovation when such approaches are applied within PPE workflows involving millions of model evaluations.</p>
      <p id="d2e223">In this study, we introduce an open-source pipeline designed to remove the computational bottlenecks of GP emulation and GAM fitting in PPE workflows. For the GP emulation stage, we implement task-level parallelism and streamlined data handling to enable efficient use of high-performance computing resources. For the GAM stage, we likewise employ task-level parallelism together with R's <monospace>mgcv::bam()</monospace> function (fast fREML estimation with <monospace>discrete=TRUE</monospace>), hereafter “rBAM”, which combines discrete smoothing, batching, and optimised input–output operations. Together, these enhancements accelerate the combined workflow by a factor of about 25 (a 96 % reduction in total runtime, from <inline-formula><mml:math id="M6" display="inline"><mml:mo>∼</mml:mo></mml:math></inline-formula> 16 800 to 665 s). Additionally, the enhancements reduce peak memory demand by more than an order of magnitude, while reproducing baseline outputs to a high degree of numerical precision. We demonstrate performance gains using a single-variable example from a climate model PPE comprising 221 simulations. We analyse 833 grid boxes per month across the European domain. In each grid box, we train a GP emulator on the 221-member PPE, then evaluate 1 000 000 parameter combinations via the emulator to generate the extended ensemble, and fit a GAM to those 1 000 000 outputs for that grid box. This is done independently for each month (833 grid boxes <inline-formula><mml:math id="M7" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 12 months <inline-formula><mml:math id="M8" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 9996 emulators and 9996 GAM fits). Note, for ease of interpretation of our approach we include a lexicon of high-performance and statistical terms in Appendix <xref ref-type="sec" rid="App1.Ch1.S1"/>. By removing long-standing scaling limitations, the pipeline enables broader application of PPE methods such as detecting structural model deficiencies and constraining parametric uncertainty, making high-resolution, multi-variable, and even multi-model analyses tractable. Although demonstrated here for climate model PPEs, the pipeline is broadly applicable to other large-scale emulation and regression problems where statistical fidelity and computational efficiency are critical.</p>
</sec>
<sec id="Ch1.S2">
  <label>2</label><title>Exemplar problem</title>
      <p id="d2e264">To demonstrate and benchmark our approach, we focus on a PPE generated with the UK Earth System Model version 1 (UKESM1-A; <xref ref-type="bibr" rid="bib1.bibx26" id="altparen.17"/>), as developed and described in <xref ref-type="bibr" rid="bib1.bibx22" id="text.18"/>. This ensemble is a suitable exemplar because it combines scientific relevance, the PPE was designed to constrain aerosol–cloud interaction radiative forcing uncertainty – one of the largest sources of climate model uncertainty <xref ref-type="bibr" rid="bib1.bibx1" id="paren.19"/>, with the computational challenges of handling very high-dimensional model output. The dataset comprises 37 perturbed parameter combinations within 221 PPE members expanded by statistical GP emulation to one million parameter combinations in each model grid box. GAMs are subsequently fitted to decompose variance and assess parameter sensitivities, again at the model grid box scale. In total, to quantify causes of aerosol–cloud interaction uncertainty across 833 grid boxes per month (9996 grid boxes for 12 months), the workflow involves training one GP emulator for each grid box on the 221-member PPE, sampling one million parameter combinations from each emulator, and fitting a GAM to those emulated outputs. This configuration is therefore representative of the scale required for multi-variable, multi-regional, or multi-model studies.</p>
      <p id="d2e276">This configuration stresses both stages of the workflow. GP training must handle high-dimensional input space and repeated tasks at scale, while GAM fitting must operate efficiently on high-resolution spatiotemporal fields without resorting to coarse aggregation. Initial baseline implementations of these steps proved computationally prohibitive on HPC systems, motivating the optimisations described in this paper.</p>

      <fig id="F1"><label>Figure 1</label><caption><p id="d2e281">UKESM1 PPE domain using one example target variable: mean H<sub>2</sub>SO<sub>4</sub> concentration (<inline-formula><mml:math id="M11" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">µ</mml:mi></mml:mrow></mml:math></inline-formula>g m<sup>−3</sup>) for January 2017. The map shows the European domain with grid points centers marked by white circles. Orange circles highlight the first 100 grid points selected for in-depth analysis in this work.</p></caption>
        <graphic xlink:href="https://gmd.copernicus.org/articles/19/7767/2026/gmd-19-7767-2026-f01.png"/>

      </fig>

      <p id="d2e329">Figure <xref ref-type="fig" rid="F1"/> shows the UKESM1 PPE domain using one example target variable, the monthly mean H<sub>2</sub>SO<sub>4</sub> concentration for January 2017. This illustrates both the spatial coverage of the European domain and the distribution of grid points required for emulator training and variance analysis. A subset of 100 model grid boxes was used to validate workflow performance before scaling to the full set of 9996 grid points for the optimised workflow.</p>
<sec id="Ch1.S2.SS1">
  <label>2.1</label><title>Benchmarking setup</title>
      <p id="d2e359">All performance tests were carried out on the JASMIN “LOTUS” high-performance computing (HPC) cluster hosted at the UK Centre for Environmental Data Analysis (CEDA) <xref ref-type="bibr" rid="bib1.bibx11" id="paren.20"/>. The LOTUS cluster is a heterogeneous system with several thousand CPU cores available across multiple partitions. For this study, we used standard CPU nodes, each equipped with dual Intel Xeon processors (2.1–2.6 GHz) and 192 GB RAM, interconnected by high-speed InfiniBand. Jobs were scheduled using Slurm (v22.05), and all experiments were submitted through Slurm job arrays with task-level resource allocation. Unless otherwise noted, each task was executed on a single CPU core, and task submission was controlled using the <monospace>#SBATCH - -array</monospace> directive with a Slurm-array throttle (e.g. <monospace>- -array=0-N%200</monospace>). The value of 200 used throughout our scaling benchmarks was a deliberate production configuration choice rather than a hard limit of Slurm, the hardware, or the workflow. This throttle permits a maximum of 200 array tasks to run concurrently. Approximately 200 tasks generally executed concurrently under the benchmark configuration, although the number could occasionally be lower depending on resource availability and system load. The 10–16 GB value denotes the maximum memory requested per task and provides headroom for variability among grid cells and cases; actual memory consumption was lower and varied between tasks.</p>
      <p id="d2e371">Original dataset sizes mirrored those used in the scientific analysis such as those reported in <xref ref-type="bibr" rid="bib1.bibx23 bib1.bibx22" id="text.21"/>, <xref ref-type="bibr" rid="bib1.bibx13 bib1.bibx14" id="text.22"/>, and <xref ref-type="bibr" rid="bib1.bibx9" id="text.23"/>, though analyses were restricted in those studies. For benchmarking purposes, we measured end-to-end walltime, peak memory usage, and accuracy equivalence across both the currently used (baseline) and newly developed (optimised) workflows. Accuracy was quantified using Pearson correlation and root mean squared error (RMSE) between baseline and optimised outputs, while scalability was assessed as a function of task count and concurrency level. Default walltime limits were set to 24 h for all experiments, with memory limits varied between 10 and 32 GB per task depending on workload.</p>
      <p id="d2e383">All software were used in stable production releases available at the time of analysis. Python (v3.10) was used for pipeline orchestration, plotting, and baseline GAM fitting with <monospace>pyGAM</monospace> (v0.9.0) <xref ref-type="bibr" rid="bib1.bibx27" id="paren.24"/>. R (v4.2.2) <xref ref-type="bibr" rid="bib1.bibx20" id="paren.25"/> was used for GP emulation with <monospace>DiceKriging</monospace> (v1.6.0) <xref ref-type="bibr" rid="bib1.bibx24" id="paren.26"/> and for optimised GAM fitting via <monospace>mgcv</monospace> (v1.8-41) <xref ref-type="bibr" rid="bib1.bibx37" id="paren.27"/>. Additional R packages included <monospace>readr</monospace> <xref ref-type="bibr" rid="bib1.bibx35" id="paren.28"/>, <monospace>lhs</monospace> <xref ref-type="bibr" rid="bib1.bibx2" id="paren.29"/>, <monospace>sensitivity</monospace> <xref ref-type="bibr" rid="bib1.bibx19" id="paren.30"/>, <monospace>trapezoid</monospace> <xref ref-type="bibr" rid="bib1.bibx31" id="paren.31"/>, and <monospace>truncnorm</monospace> <xref ref-type="bibr" rid="bib1.bibx15" id="paren.32"/>. Parallelisation was supported by Slurm job arrays on the cluster side, and OpenMP (v4.5) within <monospace>mgcv::bam()</monospace> for multithreaded GAM fitting.</p>
</sec>
<sec id="Ch1.S2.SS2">
  <label>2.2</label><title>Baseline workflow and bottlenecks</title>
<sec id="Ch1.S2.SS2.SSS1">
  <label>2.2.1</label><title>Baseline GP emulation</title>
      <p id="d2e460">The baseline workflow combined GP emulation and GAM fitting using widely available Python libraries. GP emulators were trained with the <monospace>scikit-learn</monospace> <monospace>GaussianProcessRegressor</monospace> using a squared-exponential kernel with automatic relevance determination. Training was performed sequentially, with one emulator fitted per parameter setting, and relied on dense kernel inversions that scale cubically with the number of training points. Although individual emulators were modest in size, the aggregate cost of fitting tens to hundreds of thousands of models became prohibitive. Intermediate emulator outputs were written to disk for subsequent analysis, introducing additional input/output (I/O) overhead.</p>
</sec>
<sec id="Ch1.S2.SS2.SSS2">
  <label>2.2.2</label><title>Baseline GAM fitting</title>
      <p id="d2e477">GAM fitting was carried out entirely in Python using the <monospace>pyGAM</monospace> library with cubic regression splines. Fits were implemented serially at the grid box level, with each requiring repeated construction of large design matrices and independent smoothing‐parameter optimisation. This approach was straightforward to implement and completely reproducible, but computationally inefficient as no batching or parallelism was possible, and memory was not shared across tasks.</p>

      <fig id="F2"><label>Figure 2</label><caption><p id="d2e485">Measured runtime distribution of the baseline workflow over 100 operations.</p></caption>
            <graphic xlink:href="https://gmd.copernicus.org/articles/19/7767/2026/gmd-19-7767-2026-f02.png"/>

          </fig>

      <p id="d2e494">To clarify the distinct role of the GAM within the overall workflow, we briefly describe its relationship to the GP emulator. The GP emulator is first used to approximate the model response and generate a dense sampling of the 37-parameter space based on output from the original 221-member PPE. This dense sampling provides a smoother and more complete representation of the response surface. At the grid-box scale considered here, model responses are often highly non-linear and spatially heterogeneous, so fitting a GAM directly to the sparse PPE would provide only a coarse and potentially noisy estimate of parameter effects. The emulator-expanded sample therefore provides a more robust basis for variance decomposition and sensitivity analysis.</p>
      <p id="d2e498">While a GAM fitted directly to the original PPE is computationally inexpensive and can be performed without HPC resources, it remains constrained by the limited sampling density of the parameter space. In contrast, the GP emulator enables efficient generation of a much larger ensemble (order <inline-formula><mml:math id="M15" display="inline"><mml:mrow><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mn mathvariant="normal">6</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>), providing a more complete representation of the response surface. Fitting a GAM to this extended dataset becomes both memory-intensive and computationally demanding, motivating the need for the parallel HPC-based workflow developed in this study.</p>
      <p id="d2e512">Figure <xref ref-type="fig" rid="F2"/> summarises the runtime profile of the baseline workflow and shows that GAM fitting for one million combinations across 37 parameter dimensions dominated total runtime. GAM fitting was also the main contributor to peak memory use. In the baseline implementation, which uses <monospace>pyGAM</monospace>, memory usage within a single process regularly exceeded 80 GB during spline matrix assembly, reflecting the cost of constructing large design matrices in memory. In contrast, the optimised workflow employs rBAM, which is substantially more memory-efficient and enables block-wise computation and distributed execution across independent tasks. This allows each task to operate within the typical 10–16 GB memory allocation, making the approach well suited for large-scale parallelisation.</p>
      <p id="d2e520">Across 100 operations, GAM fitting required <inline-formula><mml:math id="M16" display="inline"><mml:mo>∼</mml:mo></mml:math></inline-formula> 8500 s (58.6 % of total wall time) and GP emulation <inline-formula><mml:math id="M17" display="inline"><mml:mo>∼</mml:mo></mml:math></inline-formula> 6000 s (41.4 %), giving a total of <inline-formula><mml:math id="M18" display="inline"><mml:mo>∼</mml:mo></mml:math></inline-formula> 14 500 s. Disk I/O overhead was negligible at this scale but becomes significant at larger ensemble sizes due to the churn of intermediate files.</p>
</sec>
<sec id="Ch1.S2.SS2.SSS3">
  <label>2.2.3</label><title>Baseline (median-hold) variance computation</title>
      <p id="d2e552">In the baseline implementation, the marginal importance of each parameter was obtained by a “median-hold” procedure. For each parameter <inline-formula><mml:math id="M19" display="inline"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, the multi-dimensional GAM model predictions were recomputed while varying <inline-formula><mml:math id="M20" display="inline"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> across the extended ensemble and holding all other parameters fixed at the median values used to train the GAM.</p>
      <p id="d2e577">The variance of the resulting one-dimensional prediction series quantifies the sensitivity of the model response to changes in <inline-formula><mml:math id="M21" display="inline"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. This quantity provides a measure of first-order marginal importance for each parameter. It is analogous in interpretation to an unnormalised first-order Sobol-type sensitivity measure under an additive (no-interaction) assumption, as each parameter is varied independently while all others are held fixed <xref ref-type="bibr" rid="bib1.bibx30 bib1.bibx25" id="paren.33"/>. It is not a formal Sobol index, as it is not normalised by the total variance and does not include higher-order interaction effects.</p>
      <p id="d2e594">Repeating this process for all 37 parameters yields a set of per-parameter variances <inline-formula><mml:math id="M22" display="inline"><mml:mrow><mml:mi mathvariant="normal">Var</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and gradient signs given by <inline-formula><mml:math id="M23" display="inline"><mml:mrow><mml:mi mathvariant="normal">sign</mml:mi><mml:mo>[</mml:mo><mml:mi mathvariant="normal">Cov</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>, which can be compared to quantify the relative importance of model parameters in driving output variance.</p>
      <p id="d2e645">Although this approach is conceptually straightforward, it requires rebuilding the prediction matrix <inline-formula><mml:math id="M24" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula> times (once per parameter) and is therefore computationally expensive. The baseline workflow had three bottlenecks. First, GP emulation ran as a loop-based wrapper with dense-kernel inversions executed sequentially, leaving the embarrassingly parallel PPE analysis structure under-utilised. Second, GAM fitting in <monospace>pyGAM</monospace> was serial, using cubic splines with cross-validation smoothing, which scaled poorly with the number of grid boxes and caused repeated disk access. Third, the stages communicated via per-grid <italic>text</italic> <monospace>.dat</monospace> files, creating tens of thousands of small files and heavy metadata traffic on the shared file system. On the JASMIN super-data-cluster, where jobs are limited to 36 h and 64–128 GB per node, these bottlenecks led to 20–30 h runtimes for a single large-scale experiment and frequent out-of-memory failures.</p>
</sec>
</sec>
</sec>
<sec id="Ch1.S3">
  <label>3</label><title>Optimised workflow</title>
      <p id="d2e674">The optimised pipeline targets the dominant computational costs identified in the baseline workflow: GP emulation and GAM fitting. The key design principle is to preserve the statistical formulation while restructuring execution and data handling to enable high concurrency, bounded memory usage, and minimal I/O.</p>

      <fig id="F3" specific-use="star"><label>Figure 3</label><caption><p id="d2e679">Schematic comparison of the baseline and optimised workflows. Grey boxes show the main workflow stages; orange boxes indicate baseline implementations; blue boxes indicate optimised implementations. The optimised workflow differs from the baseline in task execution, data handling, and prediction matrix construction. The GP emulation stage trains a Gaussian Process emulator for each grid box and generates one million parameter samples per emulator. The GAM fitting stage applies generalised additive models to the emulated outputs to decompose variance and identify parameter sensitivities.</p></caption>
        <graphic xlink:href="https://gmd.copernicus.org/articles/19/7767/2026/gmd-19-7767-2026-f03.png"/>

      </fig>

      <p id="d2e688">The optimisation is not limited to replacing a sequential loop with Slurm job arrays. Rather, it involves a restructuring of the workflow, including changes to task execution, data flow, and prediction matrix construction. Job arrays serve as an enabling mechanism for parallel execution, but the primary gains arise from eliminating competing processes, reducing I/O overhead, and avoiding repeated recomputation of large intermediate data structures.</p>
      <p id="d2e692">Across the full workflow, these changes yield an overall speed-up of <inline-formula><mml:math id="M25" display="inline"><mml:mo>∼</mml:mo></mml:math></inline-formula> 25 <inline-formula><mml:math id="M26" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> (96 % total runtime reduction) and reduce peak memory by a factor of 12, with outputs matching the baseline to numerical tolerance. Figure <xref ref-type="fig" rid="F3"/> summarises the end-to-end process; implementation details are given below and a baseline–optimised comparison is provided in Table <xref ref-type="table" rid="T1"/>.</p>
<sec id="Ch1.S3.SS1">
  <label>3.1</label><title>Optimised Gaussian process emulation stage</title>
      <p id="d2e720">In the baseline workflow, GP surrogates were trained and evaluated using serial loops that spawned multiple competing <monospace>Rscript</monospace> processes on a single node. This resulted in inefficient CPU utilisation, repeated construction of prediction matrices, heavy disk-based I/O through large text intermediates, and straggler tasks that delayed completion.</p>
      <p id="d2e726">The optimised implementation retains the same emulator formulation (<monospace>DiceKriging::km</monospace>) and large-sample prediction approach. However, it restructures execution and data handling. A task table is constructed such that each (month, latitude, longitude) combination maps to a single independent task executed via a Slurm job array. This ensures deterministic task assignment, isolates failures for straightforward re-submission, and enables cluster-wide concurrency.</p>
      <p id="d2e733">This optimisation is not limited to replacing a loop with Slurm job arrays; rather, it involves a restructuring of execution, data flow, and prediction matrix construction. In addition to task-level parallelisation, several key changes are introduced. First, competing <monospace>Rscript</monospace> processes are eliminated by assigning one emulator per task, avoiding contention within a node. Second, large intermediate files are avoided by replacing text-based I/O with in-memory data transfer or compact binary output, substantially reducing filesystem overhead. Third, prediction matrix construction is reorganised to minimise repeated recomputation, reducing both runtime and memory pressure. Together, these changes convert the GP emulation stage from a sequential, I/O-bound workflow into a scalable, task-parallel process with bounded per-task memory usage.</p>
      <p id="d2e739">Validation tests at representative grid points confirmed that the baseline and optimised workflows produced effectively identical emulator outputs (Appendix B, Fig. <xref ref-type="fig" rid="FB1"/>), demonstrating that the optimisation preserves statistical fidelity while substantially improving computational performance.</p>

      <fig id="F4" specific-use="star"><label>Figure 4</label><caption><p id="d2e747">Scaling of the GP emulation stage. <bold>(a)</bold> Walltime as a function of the number of tasks (log–log scale) for the baseline (serial) and optimised (parallel) workflows. <bold>(b)</bold> Measured speedup relative to the baseline, compared with the theoretical maximum defined as <inline-formula><mml:math id="M27" display="inline"><mml:mrow><mml:mo>min⁡</mml:mo><mml:mo>(</mml:mo><mml:mi>N</mml:mi><mml:mo>,</mml:mo><mml:mn mathvariant="normal">200</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, where <inline-formula><mml:math id="M28" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula> is the number of tasks and 200 is the array-task throttle. <bold>(c)</bold> Parallel efficiency, defined as the ratio of measured speedup to the theoretical maximum. The vertical dashed line indicates the 200-task throttle, beyond which tasks are processed in successive groups.</p></caption>
          <graphic xlink:href="https://gmd.copernicus.org/articles/19/7767/2026/gmd-19-7767-2026-f04.png"/>

        </fig>

      <p id="d2e790">Performance scaling of the GP emulation stage is shown in Fig. <xref ref-type="fig" rid="F4"/>. In the baseline workflow, runtime increases approximately linearly with the number of tasks, exceeding 6000 s for GP emulation in 100 grid boxes, consistent with a serial implementation based on a looped batch wrapper.</p>
      <p id="d2e795">In the optimised workflow, one emulator is launched per Slurm array task, with up to 200 array tasks permitted to run concurrently. Task counts above 200 were executed in successive groups under this production configuration. The 200-task concurrency throttle was selected following preliminary testing as a stable and reproducible operating point on the shared LOTUS system. Approximately 200 tasks generally ran concurrently, whereas higher settings occasionally increased resource and filesystem contention and caused individual tasks to fail during periods of high system load. For task counts below this throttle, walltime remains approximately constant, indicating near-ideal parallel scaling; for larger task counts, walltime increases as tasks are processed in successive groups. To quantify the scaling behaviour, we define the theoretical maximum speedup as the minimum of the number of tasks and the 200-task throttle. Measured speedup is calculated as the ratio of baseline to optimised walltime, and parallel efficiency is defined as measured speedup divided by the theoretical maximum. At higher task counts, we performed additional baseline experiments at 512 and 1000 tasks to extend the comparison beyond the throttle. These results confirm the expected deviation from ideal scaling due to batching and scheduling overheads. At low task counts, the measured speedup approaches the theoretical maximum, reaching approximately 130 times faster than the baseline. Parallel efficiency decreases beyond the throttle and exhibits non-monotonic behaviour at large task counts, reflecting variability in walltime associated with HPC scheduling and I/O effects. Overall, the optimised workflow removes the sequential bottleneck in the GP stage, transforming the emulator workflow into a scalable, cluster-parallel process and enabling efficient utilisation of available HPC resources.</p>
</sec>
<sec id="Ch1.S3.SS2">
  <label>3.2</label><title>Optimised GAM fitting stage</title>
      <p id="d2e806">In the baseline workflow, independent GAMs were fitted using <monospace>pyGAM</monospace> with cubic regression splines at each grid box. This required repeated construction of large design matrices and re-optimisation of smoothing parameters for every fit, resulting in high computational cost and memory usage for large ensembles. The optimised approach replaces this with an R-based Big Additive Model (rBAM), which is designed for efficient estimation on large datasets. In simple terms, the key change is that all parameter effects are evaluated simultaneously in a single pass, rather than recomputing model predictions separately for each parameter as in the baseline approach. This is achieved through fast restricted maximum likelihood (fREML) estimation, discrete smoothing, and multithreaded execution. In addition, the workflow is restructured to avoid repeated prediction and design-matrix reconstruction, allowing parameter contributions to be computed in a single batched (vectorised) operation. Together, these changes substantially reduce both runtime and memory requirements while preserving the statistical formulation of the model.</p>
<sec id="Ch1.S3.SS2.SSS1">
  <label>3.2.1</label><title>Vectorised variance and sign computation</title>
      <p id="d2e819">After fitting the additive model

              <disp-formula id="Ch1.Ex1"><mml:math id="M29" display="block"><mml:mrow><mml:mi mathvariant="italic">η</mml:mi><mml:mo>=</mml:mo><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>P</mml:mi></mml:munderover><mml:msub><mml:mi>f</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

            we evaluate all partial effects simultaneously across the full dataset. In practice, this means that for each parameter <inline-formula><mml:math id="M30" display="inline"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, the corresponding smooth contribution <inline-formula><mml:math id="M31" display="inline"><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is evaluated at all <inline-formula><mml:math id="M32" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula> samples.</p>
      <p id="d2e898">This produces a matrix

              <disp-formula id="Ch1.Ex2"><mml:math id="M33" display="block"><mml:mrow><mml:mi mathvariant="bold">T</mml:mi><mml:mo>=</mml:mo><mml:mo>[</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">t</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>P</mml:mi></mml:msub><mml:mo>]</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

            where each column <inline-formula><mml:math id="M34" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">t</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> represents the contribution of parameter <inline-formula><mml:math id="M35" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> across all samples.</p>
      <p id="d2e955">The marginal contribution of each parameter is then quantified using simple column-wise statistics:

              <disp-formula id="Ch1.Ex3"><mml:math id="M36" display="block"><mml:mrow><mml:mi mathvariant="normal">Var</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:mspace linebreak="nobreak" width="1em"/><mml:mi mathvariant="normal">sign</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="normal">Cov</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>)</mml:mo><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula></p>
      <p id="d2e1004">Here, <inline-formula><mml:math id="M37" display="inline"><mml:mrow><mml:mi mathvariant="normal">Var</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> measures the strength of the parameter's influence on the model response, while <inline-formula><mml:math id="M38" display="inline"><mml:mrow><mml:mi mathvariant="normal">sign</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="normal">Cov</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> indicates the direction of that influence.</p>
      <p id="d2e1055">This vectorised formulation computes all parameter effects in a single pass, rather than recomputing predictions separately for each parameter as in the baseline “median-hold” method. It therefore avoids repeated construction of design matrices and reduces computational cost from <inline-formula><mml:math id="M39" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula> prediction passes to one.</p>
      <p id="d2e1065">These quantities are invariant to additive constants in the partial contributions, since

              <disp-formula id="Ch1.Ex4"><mml:math id="M40" display="block"><mml:mrow><mml:mi mathvariant="normal">Var</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:mi>c</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mi mathvariant="normal">Var</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:mspace linebreak="nobreak" width="1em"/><mml:mi mathvariant="normal">Cov</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:mi>c</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mi mathvariant="normal">Cov</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula></p>
</sec>
<sec id="Ch1.S3.SS2.SSS2">
  <label>3.2.2</label><title>Implementation details</title>
      <p id="d2e1158">In practice, the matrix <inline-formula><mml:math id="M41" display="inline"><mml:mi mathvariant="bold">T</mml:mi></mml:math></inline-formula> is obtained using the <monospace>predict(..., type="terms")</monospace> functionality in <monospace>mgcv</monospace>, which returns the contribution of each smooth term evaluated across all samples. Computation is performed in blocks (via <monospace>block.size</monospace>) to control memory usage and combined with <monospace>bam(..., discrete=TRUE)</monospace> to enable efficient large-scale fitting. Parallelism is achieved both within each fit (via OpenMP multithreading) and across grid boxes (via Slurm job arrays).</p>
</sec>
<sec id="Ch1.S3.SS2.SSS3">
  <label>3.2.3</label><title>Interpretation of the variance term</title>
      <p id="d2e1188">The variance term <inline-formula><mml:math id="M42" display="inline"><mml:mrow><mml:mi mathvariant="normal">Var</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> represents the marginal contribution of parameter <inline-formula><mml:math id="M43" display="inline"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> within the full additive model. Each smooth function <inline-formula><mml:math id="M44" display="inline"><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is estimated jointly with all others, so <inline-formula><mml:math id="M45" display="inline"><mml:mrow><mml:mi mathvariant="normal">Var</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> reflects variability after accounting for the influence of the remaining parameters.</p>
      <p id="d2e1256">These quantities correspond to first-order marginal sensitivities and are directly comparable to the baseline median-hold approach. They should not be interpreted as total variance contributions, as interactions are not explicitly represented in the additive model. In particular,

              <disp-formula id="Ch1.Ex5"><mml:math id="M46" display="block"><mml:mrow><mml:mi mathvariant="normal">Var</mml:mi><mml:mspace width="-0.125em" linebreak="nobreak"/><mml:mfenced close=")" open="("><mml:mrow><mml:munder><mml:mo movablelimits="false">∑</mml:mo><mml:mi>i</mml:mi></mml:munder><mml:msub><mml:mi>f</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mfenced><mml:mo>=</mml:mo><mml:munder><mml:mo movablelimits="false">∑</mml:mo><mml:mi>i</mml:mi></mml:munder><mml:mi mathvariant="normal">Var</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>+</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:munder><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>&lt;</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:munder><mml:mi mathvariant="normal">Cov</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

            so the sum of individual variances does not equal the total output variance.</p>
      <p id="d2e1342">This vectorised formulation reproduces the baseline results exactly while substantially reducing computational cost.</p>
      <p id="d2e1345">To verify that the optimised formulation preserves the statistical fidelity of the baseline workflow, we evaluate variance contributions across a broad spatial sample.</p>

      <fig id="F5" specific-use="star"><label>Figure 5</label><caption><p id="d2e1351">Comparison of fractional variance contributions from GAM fitting across 100 grid boxes (3700 points in total). Each point represents one parameter at one grid cell. The dashed line denotes the 1 : 1 relationship.</p></caption>
            <graphic xlink:href="https://gmd.copernicus.org/articles/19/7767/2026/gmd-19-7767-2026-f05.png"/>

          </fig>

      <p id="d2e1360">Figure <xref ref-type="fig" rid="F5"/> shows that variance contributions from the optimised rBAM workflow closely match those from the baseline <monospace>pyGAM</monospace> implementation across all grid boxes and parameters. The points lie tightly along the 1 : 1 line, with an overall Pearson correlation of <inline-formula><mml:math id="M47" display="inline"><mml:mrow><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.999</mml:mn></mml:mrow></mml:math></inline-formula>, demonstrating that the optimised approach reproduces the baseline variance decomposition with high fidelity across the domain.</p>
      <p id="d2e1380">Small deviations from the 1 : 1 relationship are visible for a limited number of parameters and grid boxes. These differences arise from numerical and algorithmic distinctions between the two implementations rather than any change in the underlying statistical formulation. In particular, <monospace>pyGAM</monospace> and <monospace>mgcv::bam</monospace> differ in spline basis construction, smoothing parameter estimation (cross-validation versus fREML), and numerical optimisation strategies. In addition, the optimised workflow employs discrete smoothing and block-wise evaluation of the prediction matrix, which can introduce minor numerical approximations at large sample sizes (<inline-formula><mml:math id="M48" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mn mathvariant="normal">6</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>). These effects lead to small local differences in estimated smooth functions, which propagate into slight variations in the derived variance contributions. However, these differences are small in magnitude and do not affect the overall ranking or interpretation of parameter importance. To provide additional context, the spatial distribution of correlation coefficients across all analysed grid boxes is included in the Appendix <xref ref-type="sec" rid="App1.Ch1.S3"/> (Fig. <xref ref-type="fig" rid="FC1"/>).</p>
      <p id="d2e1406">To quantify the computational benefits of this restructuring, we examine the scaling behaviour of the GAM fitting stage (Fig. <xref ref-type="fig" rid="F6"/>). Panel (a) shows walltime, (b) the measured speed-up relative to the baseline, and (c) the corresponding parallel efficiency, together with the theoretical maximum scaling and the 200-task throttle.</p>

      <fig id="F6" specific-use="star"><label>Figure 6</label><caption><p id="d2e1413">Scaling of the GAM fitting stage. <bold>(a)</bold> Walltime for baseline and optimised workflows, <bold>(b)</bold> measured speed-up compared to the baseline alongside the theoretical maximum (limited by the 200-task throttle), and <bold>(c)</bold> parallel efficiency defined relative to the effective concurrency. The vertical dashed line indicates the 200-task throttle. Deviations from ideal scaling at higher task counts arise from processing in successive groups beyond the throttle and from variability in resource availability on the shared system.</p></caption>
            <graphic xlink:href="https://gmd.copernicus.org/articles/19/7767/2026/gmd-19-7767-2026-f06.png"/>

          </fig>

      <p id="d2e1432">In the baseline workflow, runtimes increase superlinearly with the number of tasks, quickly exceeding practical limits beyond a few hundred grid boxes, while peak memory usage frequently approaches or exceeds node capacity.</p>
      <p id="d2e1435">In the optimised workflow, rBAM substantially reduces both runtime and memory requirements, achieving more than 20 times faster performance at approximately <inline-formula><mml:math id="M49" display="inline"><mml:mrow><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mn mathvariant="normal">3</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> tasks. However, the scaling deviates from the theoretical maximum beyond moderate task counts. The 10–16 GB value represents the maximum memory requested per task and provides operational headroom for variation among grid cells and cases; it does not represent the memory continuously consumed by every task. Actual memory use was lower and varied between tasks, while approximately 200 tasks generally ran concurrently under the production configuration. Variability in resource availability and filesystem load on the shared LOTUS system nevertheless contributed to the observed flattening and fluctuations in speed-up and efficiency beyond <inline-formula><mml:math id="M50" display="inline"><mml:mo>∼</mml:mo></mml:math></inline-formula> 100 tasks.</p>
      <p id="d2e1456">Despite these practical constraints, the optimised workflow maintains substantially lower walltimes and improved efficiency compared to the baseline, enabling scalable application of GAM-based variance decomposition at resolutions and ensemble sizes that were previously infeasible.</p>
</sec>
</sec>
<sec id="Ch1.S3.SS3">
  <label>3.3</label><title>Pipeline integration and reproducibility</title>
      <p id="d2e1468">The pipeline implements and automates the workflow for large-scale statistical emulation. In this context, the workflow refers to the ordered sequence of tasks required to generate, emulate, and analyse model output, specifically, GP emulation, GAM fitting, and downstream analysis. The pipeline, by contrast, is the software and organisational framework that executes this workflow reproducibly and efficiently on high-performance computing systems. It manages task scheduling, data handling, and parallel execution, ensuring that the workflow can be scaled and repeated without manual intervention.</p>
      <p id="d2e1471">The pipeline organises the optimised workflow around a task-centred folder layout with separate stages for GP emulation, GAM fitting, and downstream analysis. Configuration is provided via plain-text (.YAML/.JSON) files (month, grid, variable), enabling deterministic re-runs that reproduce identical results from the same inputs. A consistent file-naming scheme based on month, latitude and longitude indices supports automatic discovery of missing or failed tasks. In previous implementations, a single task failure, often caused by node timeouts or memory overuse, could halt or invalidate large batch jobs, requiring manual log inspection and reruns of the entire workflow. In the optimised pipeline, each task writes an individual log and output file identified by its coordinates, allowing a lightweight post-check script to detect missing or incomplete outputs and automatically resubmit only the affected tasks. This design prevents error propagation, eliminates manual recovery steps, and greatly improves overall robustness and throughput on HPC systems. Slurm job arrays orchestrate parallel execution for a given output variable, with each array index corresponding to a single model grid box for one month (one task). The array structure can span any number of grid boxes, regional or global, depending on available HPC resources. Inputs are read once per task; outputs are kept in memory or written in binary to limit I/O overhead. Per-task logs (<monospace>slurm_output_%A_%a.out</monospace>) record program outputs and error messages, and a lightweight post-check script scans for failures and automatically resubmits only the affected tasks.</p>
      <p id="d2e1478">Software environments are fully version-controlled to ensure reproducibility. All Python dependencies are specified in the Conda environment file <monospace>envs/environment.yml</monospace>, while all R packages (including <monospace>mgcv</monospace> and <monospace>DiceKriging</monospace>) are declared in the manifest script <monospace>envs/install_R_deps.R</monospace>. This script automatically installs the required packages from CRAN and records the complete R session information (<monospace>sessionInfo()</monospace>), providing an explicit record of the package versions used in the analysis. The workflow is designed to run efficiently on high-performance computing (HPC) systems (e.g., JASMIN, ARCHER) and can also be executed locally on Windows, Linux, or macOS. The repository includes a fully self-contained demonstration using four grid points (H<sub>2</sub>SO<sub>4</sub>, January 2017) that allows users to test the complete GP–GAM pipeline without requiring HPC access. Software environments are fully reproducible using the provided Conda specification (<monospace>envs/environment.yml</monospace>) for Python and the manifest script (<monospace>envs/install_R_deps.R</monospace>) for R.</p>
      <p id="d2e1521">The individual improvements to GP emulation, GAM fitting, and workflow integration are brought together in Table <xref ref-type="table" rid="T1"/>, which contrasts key features of the baseline and optimised workflows.</p>

<table-wrap id="T1" specific-use="star"><label>Table 1</label><caption><p id="d2e1530">Baseline versus optimised workflows at the GP and GAM stages (key changes highlighted in bold).</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="3">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="justify" colwidth="7.5cm"/>
     <oasis:colspec colnum="3" colname="col3" align="justify" colwidth="7cm"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Aspect</oasis:entry>
         <oasis:entry colname="col2" align="left">Baseline</oasis:entry>
         <oasis:entry colname="col3" align="left">Optimised</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row rowsep="1">
         <oasis:entry namest="col1" nameend="col3"><bold>GP emulation stage</bold></oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Software</oasis:entry>
         <oasis:entry colname="col2" align="left">R <monospace>DiceKriging</monospace> (<monospace>km</monospace>/<monospace>predict.km</monospace>) via looped scripts</oasis:entry>
         <oasis:entry colname="col3" align="left">Same emulator/predictor, orchestrated with <bold>Slurm arrays</bold></oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Algorithm</oasis:entry>
         <oasis:entry colname="col2" align="left">Serial training; large-sample prediction; text intermediates</oasis:entry>
         <oasis:entry colname="col3" align="left">Same algorithms; <bold>task-table orchestration</bold> and <bold>binary/in-memory I/O</bold></oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Parallelisation</oasis:entry>
         <oasis:entry colname="col2" align="left">Few competing <monospace>Rscript</monospace> jobs on one node</oasis:entry>
         <oasis:entry colname="col3" align="left"><bold>Cluster-wide concurrency</bold>; one task per (month, lat, lon)</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Memory</oasis:entry>
         <oasis:entry colname="col2" align="left">Multiple processes share node RAM; high peaks</oasis:entry>
         <oasis:entry colname="col3" align="left"><bold>Bounded per-task RAM</bold>; minimal footprint</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">I/O</oasis:entry>
         <oasis:entry colname="col2" align="left">Shared/interleaved logs; many small files</oasis:entry>
         <oasis:entry colname="col3" align="left"><bold>Per-task logs</bold>; deterministic filenames</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry namest="col1" nameend="col3"><bold>GAM fitting stage</bold></oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Software</oasis:entry>
         <oasis:entry colname="col2" align="left">Python <monospace>pyGAM</monospace></oasis:entry>
         <oasis:entry colname="col3" align="left">Python–R bridge to <bold><monospace>mgcv::bam()</monospace></bold> (rBAM)</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Algorithm</oasis:entry>
         <oasis:entry colname="col2" align="left">Cubic splines; CV-based smoothing; rebuild design matrices per fit</oasis:entry>
         <oasis:entry colname="col3" align="left"><bold>fREML</bold> with <monospace>discrete=TRUE</monospace>; <bold>batched</bold> design matrices</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Parallelisation</oasis:entry>
         <oasis:entry colname="col2" align="left">Serial fits; limited threading</oasis:entry>
         <oasis:entry colname="col3" align="left"><bold>OpenMP within fits</bold> <inline-formula><mml:math id="M53" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula> <bold>Slurm arrays</bold> across tasks</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Memory</oasis:entry>
         <oasis:entry colname="col2" align="left">Large, repeatedly built matrices; peaks <inline-formula><mml:math id="M54" display="inline"><mml:mo>&gt;</mml:mo></mml:math></inline-formula> 80 GB</oasis:entry>
         <oasis:entry colname="col3" align="left"><bold>Sparse/batched</bold> design; substantially lower peaks</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">I/O</oasis:entry>
         <oasis:entry colname="col2" align="left">Filenames without standardisation; repeated re-reads</oasis:entry>
         <oasis:entry colname="col3" align="left"><bold>Deterministic outputs</bold>; minimal re-reads</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

</sec>
</sec>
<sec id="Ch1.S4">
  <label>4</label><title>End-to-end pipeline performance</title>
      <p id="d2e1769">The optimised workflow delivers large and systematic improvements in both runtime and memory at every stage of the pipeline (Fig. <xref ref-type="fig" rid="F7"/> and Table <xref ref-type="table" rid="TD1"/> in Appendix D). Median per-task walltime for the GP emulation falls from 6177 to 154 s (a <inline-formula><mml:math id="M55" display="inline"><mml:mrow><mml:mn mathvariant="normal">40.1</mml:mn><mml:mo>×</mml:mo></mml:mrow></mml:math></inline-formula> speed-up), while GAM fitting drops from 10 623 to 511 s (<inline-formula><mml:math id="M56" display="inline"><mml:mrow><mml:mn mathvariant="normal">20.8</mml:mn><mml:mo>×</mml:mo></mml:mrow></mml:math></inline-formula>). Taken together, the end-to-end time reduces from 16 800 to 665 s – about <inline-formula><mml:math id="M57" display="inline"><mml:mrow><mml:mn mathvariant="normal">25.3</mml:mn><mml:mo>×</mml:mo></mml:mrow></mml:math></inline-formula> faster overall (from <inline-formula><mml:math id="M58" display="inline"><mml:mo>∼</mml:mo></mml:math></inline-formula> 4 h 40 min to <inline-formula><mml:math id="M59" display="inline"><mml:mo>∼</mml:mo></mml:math></inline-formula> 11 min per model grid box), for the 100-task benchmark case considered here.</p>

      <fig id="F7"><label>Figure 7</label><caption><p id="d2e1823">Bar chart of speed-up factors showing GAM fitting, GP emulation, and total pipeline side-by-side for quick comparison. Baseline and optimised runtimes (left axis) are shown alongside the speed-up (right axis). GP achieves the largest gain (<inline-formula><mml:math id="M60" display="inline"><mml:mrow><mml:mn mathvariant="normal">40.1</mml:mn><mml:mo>×</mml:mo></mml:mrow></mml:math></inline-formula>; 6177 s <inline-formula><mml:math id="M61" display="inline"><mml:mo>→</mml:mo></mml:math></inline-formula> 154 s), GAM is <inline-formula><mml:math id="M62" display="inline"><mml:mrow><mml:mn mathvariant="normal">20.8</mml:mn><mml:mo>×</mml:mo></mml:mrow></mml:math></inline-formula> faster (10 623 s <inline-formula><mml:math id="M63" display="inline"><mml:mo>→</mml:mo></mml:math></inline-formula> 511 s), and the full pipeline is reduced by <inline-formula><mml:math id="M64" display="inline"><mml:mrow><mml:mn mathvariant="normal">25.3</mml:mn><mml:mo>×</mml:mo></mml:mrow></mml:math></inline-formula> overall (16 800 s <inline-formula><mml:math id="M65" display="inline"><mml:mo>→</mml:mo></mml:math></inline-formula> 665 s).</p></caption>
        <graphic xlink:href="https://gmd.copernicus.org/articles/19/7767/2026/gmd-19-7767-2026-f07.png"/>

      </fig>

      <p id="d2e1884">Peak memory usage shows comparable gains. GAM fitting, which dominated the baseline memory footprint, decreases from <inline-formula><mml:math id="M66" display="inline"><mml:mo>∼</mml:mo></mml:math></inline-formula> 100 to <inline-formula><mml:math id="M67" display="inline"><mml:mo>∼</mml:mo></mml:math></inline-formula> 10 GB (<inline-formula><mml:math id="M68" display="inline"><mml:mrow><mml:mn mathvariant="normal">10.0</mml:mn><mml:mo>×</mml:mo></mml:mrow></mml:math></inline-formula> reduction), and the GP emulation stage falls from <inline-formula><mml:math id="M69" display="inline"><mml:mo>∼</mml:mo></mml:math></inline-formula> 50 to <inline-formula><mml:math id="M70" display="inline"><mml:mo>∼</mml:mo></mml:math></inline-formula> 6.5 GB (<inline-formula><mml:math id="M71" display="inline"><mml:mrow><mml:mn mathvariant="normal">7.7</mml:mn><mml:mo>×</mml:mo></mml:mrow></mml:math></inline-formula>). As a result, the end-to-end peak memory requirement contracts by around <inline-formula><mml:math id="M72" display="inline"><mml:mrow><mml:mn mathvariant="normal">10.0</mml:mn><mml:mo>×</mml:mo></mml:mrow></mml:math></inline-formula>, enabling much denser job arrays on the same hardware and reducing the frequency of errors caused by memory limitations.</p>
      <p id="d2e1947">Practically, these improvements come from replacing slow, memory-intensive methods and I/O patterns with batched and streaming operations, using more efficient GAM fitting (e.g. discrete smoothing fits with <monospace>bam</monospace>) and lean GP emulation evaluation paths, while preserving the statistical specification of the original techniques. Figure <xref ref-type="fig" rid="F7"/> provides a side-by-side comparison of the GAM fitting, GP emulation, and total pipeline speed-ups to highlight where the largest gains are realised.</p>
</sec>
<sec id="Ch1.S5">
  <label>5</label><title>Application potential</title>
      <p id="d2e1964">The optimised pipeline substantially expands what is feasible in PPE-based climate model analysis. By reducing end-to-end runtime by more than an order of magnitude and lowering memory demand twelve-fold, our new workflow opens the door to in-depth analyses that were previously unachievable due to computational bottlenecks. In particular, four scientific areas of exploration are now enabled: (i) simultaneous treatment of multiple target variables, such as aerosol optical depth, cloud droplet number, and liquid water path, rather than single-variable case studies; (ii) finer temporal resolution, moving from annual mean to monthly, daily or even sub-daily outputs; (iii) extension to larger spatial domains, including global analyses without the need for coarse aggregation of model data; and (iv) cross-model comparisons in which multiple PPEs can be emulated and analysed under a unified statistical framework.</p>
      <p id="d2e1967">Example application scenarios include quantifying shared causes of aerosol radiative forcing uncertainty across multiple climate models, constraining interacting sources of uncertainty across model components and evaluating constraints across models to identify key structural model deficiencies in the current generation of climate models. The workflow is also readily applicable to non-climate large-ensemble settings where emulation and variance decomposition are required, such as hydrological forecasting, air-quality assessments or Earth system model evaluations.</p>
      <p id="d2e1970">Some opportunities for further refinement remain. We implemented discrete smoothing using rBAM, which significantly accelerates GAM fitting but may require parameter tuning when applied to extremely noisy or non-stationary predictors (not yet experienced in our chosen climate model outputs). Choice of kernel in GP emulator creation affects emulator fidelity – automated selection procedures could be implemented to trade some of the computation gains with improved emulator skill, which will be important as analyses scale. Finally, although per-task memory ceilings have been reduced substantially, very large multi-variable or multi-model applications, where distinct parameters are perturbed over model-specific ranges, will need more complex job-array design and, in some cases, access to distributed or GPU-enabled resources.</p>
</sec>
<sec id="Ch1.S6" sec-type="conclusions">
  <label>6</label><title>Conclusions</title>
      <p id="d2e1981">We have introduced and benchmarked a high-performance, open-source pipeline for large-scale statistical emulation in Earth system science. By restructuring workflow execution around task-level parallelism, streamlined I/O, and memory-efficient smoothing, we reduced GP training time by 97.5 % and GAM fitting time by 95.2 %, yielding a <inline-formula><mml:math id="M73" display="inline"><mml:mo>∼</mml:mo></mml:math></inline-formula> 25-fold end-to-end speed-up and a 12-fold reduction in peak memory demand. Crucially, these efficiency gains were achieved without loss of statistical fidelity, where emulator predictions and variance decompositions are numerically identical to the baseline implementation. The reduction in per-task memory is critical for enabling large-scale parallel execution across HPC systems, transforming PPE analysis from a memory-limited to a throughput-limited problem.</p>
      <p id="d2e1991">The workflow is designed for reproducibility and ease of adoption. All components are open-source, orchestrated via job arrays or containerised environments, and follow a deterministic file structure that simplifies reruns and failure recovery. This makes the pipeline directly reusable in other modelling contexts where emulators and smoothers are applied to high-dimensional outputs.</p>
      <p id="d2e1994">Future work will extend the framework in three directions: automatic kernel and smoothing-parameter selection to reduce the need for manual tuning, and distributed inference across heterogeneous HPC and cloud platforms. By combining statistical fidelity with computational tractability, the pipeline provides a scalable foundation for next-generation PPE studies and for broader applications of GP emulation and fractional variance decomposition methods in environmental science.</p>
</sec>

      
      </body>
    <back><app-group>

<app id="App1.Ch1.S1">
  <label>Appendix A</label><title>Lexicon of computational and workflow terms</title>

<table-wrap id="TA1a"><label>Table A1</label><caption><p id="d2e2012">This lexicon defines key computational and statistical terms used throughout the paper to aid readers who may be less familiar with high-performance computing and emulation workflow terminology.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="2">
     <oasis:colspec colnum="1" colname="col1" align="justify" colwidth="3cm"/>
     <oasis:colspec colnum="2" colname="col2" align="justify" colwidth="14cm"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Term/Acronym</oasis:entry>
         <oasis:entry colname="col2" align="left">Definition and Explanation</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Accuracy equivalence</oasis:entry>
         <oasis:entry colname="col2" align="left">Verification that optimised and baseline outputs are statistically identical, confirming no (or minimal to a specified level of accuracy) scientific change.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Array directive (Slurm)</oasis:entry>
         <oasis:entry colname="col2" align="left">A scheduler instruction that defines how many tasks to launch in a job array and optionally limits concurrent execution, for example <monospace>#SBATCH --array=0--999%500</monospace>. Used to implement throttling between 200 and 500 simultaneous tasks to balance load and filesystem performance.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Automatic resubmission</oasis:entry>
         <oasis:entry colname="col2" align="left">Recovery process in which failed tasks are detected and rerun automatically without manual intervention.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Baseline workflow</oasis:entry>
         <oasis:entry colname="col2" align="left">The original serial implementation of the GP–GAM procedure that used sequential loops, text-based I/O, and no parallelisation. It provided reference results for benchmarking.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Batching/batched design matrices</oasis:entry>
         <oasis:entry colname="col2" align="left">Processing large data in smaller chunks to reduce memory usage (e.g. In our GP stage, batching is used for kernel inversions; in our GAM stage, it limits design-matrix size during prediction).</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Benchmarking</oasis:entry>
         <oasis:entry colname="col2" align="left">Systematic measurement of runtime, memory usage, and accuracy to quantify performance improvements between workflows.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Binary I/O and intermediates</oasis:entry>
         <oasis:entry colname="col2" align="left">Storing intermediate data in compact binary format rather than text, drastically reducing file size and read/write time.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Bounded per-task RAM</oasis:entry>
         <oasis:entry colname="col2" align="left">Explicitly limiting the memory available to each task so that many tasks can run concurrently without exceeding node capacity.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Concurrency</oasis:entry>
         <oasis:entry colname="col2" align="left">The number of tasks executed simultaneously on available computing resources.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Conda</oasis:entry>
         <oasis:entry colname="col2" align="left">An open source package and environment management system for Python that helps users install, update and manage software packages and dependencies.</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1" align="left">Conda environment file (e.g. <monospace>environment.yml</monospace>)</oasis:entry>
         <oasis:entry colname="col2" align="left">Specification file listing all Python dependencies and versions for reproducible installation through Conda.</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

<table-wrap id="TA1b"><label>Table A1</label><caption><p id="d2e2150">Continued.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="2">
     <oasis:colspec colnum="1" colname="col1" align="justify" colwidth="3cm"/>
     <oasis:colspec colnum="2" colname="col2" align="justify" colwidth="14cm"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Term/Acronym</oasis:entry>
         <oasis:entry colname="col2" align="left">Definition and Explanation</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Container/ containerisation</oasis:entry>
         <oasis:entry colname="col2" align="left">A portable computing environment (e.g. Docker, Singularity) that packages software (across programming languages) and dependencies to ensure identical execution across systems.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Container recipe</oasis:entry>
         <oasis:entry colname="col2" align="left">Configuration file (e.g. Dockerfile or Singularity definition) that builds the container environment to match the HPC setup.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">CPU core</oasis:entry>
         <oasis:entry colname="col2" align="left">The smallest independent processor unit that executes individual tasks.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Cross-validation</oasis:entry>
         <oasis:entry colname="col2" align="left">A GAM evaluation technique used to select optimal smoothing by estimating how well the GAM generalizes to unseen data.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Deterministic file naming</oasis:entry>
         <oasis:entry colname="col2" align="left">A systematic naming convention (e.g. <monospace>latXXpXXX_lonXXpXXX</monospace>) that encodes coordinates and aids automatic tracking and recovery of task outputs.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Deterministic re-run</oasis:entry>
         <oasis:entry colname="col2" align="left">Re-execution of the workflow that reproduces identical outputs when given the same inputs and configuration.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Discrete smoothing</oasis:entry>
         <oasis:entry colname="col2" align="left">An optimisation in <monospace>mgcv::bam()</monospace> where continuous covariate values are discretised into bins before fitting. This reduces computation and memory demand with minimal loss of accuracy.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Error propagation</oasis:entry>
         <oasis:entry colname="col2" align="left">Situation where a single task failure could halt or invalidate entire job batches. Task isolation prevents error propagation.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Extended ensemble</oasis:entry>
         <oasis:entry colname="col2" align="left">A very large set of model outputs generated by sampling from a GP emulator (e.g. one million parameter combinations) beyond the original (221) PPE members.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">fREML (Fast Restricted Maximum Likelihood)</oasis:entry>
         <oasis:entry colname="col2" align="left">An efficient algorithm for estimating smoothing parameters in GAMs that avoids repeated cross-validation and scales well to large data sets.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Filesystem churn/metadata traffic</oasis:entry>
         <oasis:entry colname="col2" align="left">Performance loss caused by creation of many small files on a shared filesystem; Can be reduced via pipeline optimisation.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Filesystem contention</oasis:entry>
         <oasis:entry colname="col2" align="left">Performance degradation that occurs when too many processes simultaneously read from or write to the same storage system.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">GAM (Generalised Additive Model)</oasis:entry>
         <oasis:entry colname="col2" align="left">A regression framework that represents linear or nonlinear relationships between parameter values and/or variables using smooth functions.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Gigadata/high-dimensional output</oasis:entry>
         <oasis:entry colname="col2" align="left">Extremely large datasets with millions of observations or many predictor variables, typical of high-resolution Earth system models and data, or model ensembles.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">GP (Gaussian Process)</oasis:entry>
         <oasis:entry colname="col2" align="left">A non-parametric Bayesian statistical model used to emulate output from complex numerical simulations.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">GP emulator</oasis:entry>
         <oasis:entry colname="col2" align="left">The trained Gaussian Process model used as a statistical surrogate for expensive model runs. Once trained (e.g. on PPE output), GP emulators can rapidly predict millions of output values (mean and variance) at unseen parameter combinations.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Grid box/grid point</oasis:entry>
         <oasis:entry colname="col2" align="left">A spatial unit of climate model domain defined by latitude and longitude coordinates.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">HPC (High-Performance Computing)</oasis:entry>
         <oasis:entry colname="col2" align="left">Large computing systems consisting of many interconnected processors and nodes that allow parallel execution of computationally intensive workloads.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">I/O (Input–Output)</oasis:entry>
         <oasis:entry colname="col2" align="left">Reading from or writing data to disk or memory. Efficient I/O handling is essential for performance at scale.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">JASMIN/ARCHER</oasis:entry>
         <oasis:entry colname="col2" align="left">UK national HPC infrastructures used in this study; JASMIN hosts the LOTUS cluster at the Centre for Environmental Data Analysis (CEDA), and ARCHER is the UK's national supercomputing service.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Job array index</oasis:entry>
         <oasis:entry colname="col2" align="left">The unique identifier for a single task within a Slurm array, allowing direct mapping to a specific grid box and month.</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1" align="left">Loop-based wrapper/serial loop</oasis:entry>
         <oasis:entry colname="col2" align="left">A sequential control structure in which each task is executed one after another within a shell or Python loop.</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

<table-wrap id="TA1c"><label>Table A1</label><caption><p id="d2e2390">Continued.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="2">
     <oasis:colspec colnum="1" colname="col1" align="justify" colwidth="3cm"/>
     <oasis:colspec colnum="2" colname="col2" align="justify" colwidth="14cm"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Term/Acronym</oasis:entry>
         <oasis:entry colname="col2" align="left">Definition and Explanation</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">LOTUS cluster</oasis:entry>
         <oasis:entry colname="col2" align="left">HPC cluster within JASMIN used for all benchmarking experiments; provides thousands of CPU cores and high-speed interconnect.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Manifest script (<monospace>install_R_deps.R</monospace>)</oasis:entry>
         <oasis:entry colname="col2" align="left">The R script that installs required R packages and records <monospace>sessionInfo()</monospace>, and documents exact package versions used in the analysis.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Marginal importance/contribution</oasis:entry>
         <oasis:entry colname="col2" align="left">The result of variance decomposition analysis, where the marginal importance of each parameter <inline-formula><mml:math id="M74" display="inline"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the contribution to the modelled response variance after accounting for the influence of all other parameters in the full additive model. Represents the sensitivity of <inline-formula><mml:math id="M75" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula> to <inline-formula><mml:math id="M76" display="inline"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> within the multivariate GAM framework, not the effect of perturbing <inline-formula><mml:math id="M77" display="inline"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> alone.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Matrix operations/linear algebra kernels</oasis:entry>
         <oasis:entry colname="col2" align="left">Core computational routines (e.g. matrix multiplications, inversions) that dominate the cost of GP and GAM calculations. Parallelising or batching them greatly improves performance.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Median-hold method</oasis:entry>
         <oasis:entry colname="col2" align="left">The baseline approach for estimating parameter importance in GAMs. Each parameter <inline-formula><mml:math id="M78" display="inline"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is varied across the ensemble while all others are held fixed at their median values. The variance of the resulting one-dimensional prediction series quantifies the isolated influence of <inline-formula><mml:math id="M79" display="inline"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. Conceptually simple but computationally expensive, as it requires <inline-formula><mml:math id="M80" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula> (number of parameters in the PPE) independent prediction passes.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Memory footprint/peak memory</oasis:entry>
         <oasis:entry colname="col2" align="left">The amount of RAM used by a process; <italic>peak memory</italic> is the maximum observed usage during execution.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Memory reduction factor</oasis:entry>
         <oasis:entry colname="col2" align="left">Ratio of baseline to optimised peak memory usage.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Multithreading</oasis:entry>
         <oasis:entry colname="col2" align="left">Running multiple threads (lightweight tasks) simultaneously within a single process on a compute node, allowing different parts of a program to execute concurrently on separate cores.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Node</oasis:entry>
         <oasis:entry colname="col2" align="left">A physical compute unit within an HPC cluster that contains CPUs (and sometimes GPUs) and memory. Multiple tasks may run concurrently on one node.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Node-level contention</oasis:entry>
         <oasis:entry colname="col2" align="left">Competition among tasks running on the same HPC node for shared resources such as memory bandwidth, cache, or I/O channels. This contention can degrade performance and lead to superlinear scaling behaviour.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Numerical fidelity</oasis:entry>
         <oasis:entry colname="col2" align="left">The degree to which outputs from an optimised workflow exactly match those from a baseline implementation (e.g. Pearson <inline-formula><mml:math id="M81" display="inline"><mml:mrow><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1.000</mml:mn></mml:mrow></mml:math></inline-formula>).</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">OpenMP (v4.5)</oasis:entry>
         <oasis:entry colname="col2" align="left">A shared-memory parallel programming interface that allows multiple threads to run parts of a program simultaneously within one node.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Optimised workflow/pipeline</oasis:entry>
         <oasis:entry colname="col2" align="left">A redesigned workflow introducing task-level parallelism, efficient I/O, bounded memory, and reproducible automation to achieve large performance gains.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Parameter sensitivity</oasis:entry>
         <oasis:entry colname="col2" align="left">The strength or influence of each input parameter on model output (quantified here using GAM variance decomposition).</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Parameter space</oasis:entry>
         <oasis:entry colname="col2" align="left">The multidimensional range of all uncertain model parameter combinations systematically explored by the PPE and emulated by the GP.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Partial smooth term (<inline-formula><mml:math id="M82" display="inline"><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>)</oasis:entry>
         <oasis:entry colname="col2" align="left">A nonlinear smooth function in a GAM describing how model output changes with parameter <inline-formula><mml:math id="M83" display="inline"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. Each smooth <inline-formula><mml:math id="M84" display="inline"><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is fitted jointly with the others and contributes a term <inline-formula><mml:math id="M85" display="inline"><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> to the overall additive predictor <inline-formula><mml:math id="M86" display="inline"><mml:mrow><mml:mi mathvariant="italic">η</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mo>∑</mml:mo><mml:mi>i</mml:mi></mml:msub><mml:msub><mml:mi>f</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Pearson correlation</oasis:entry>
         <oasis:entry colname="col2" align="left">Statistical measure of linear correlation between two sets of values; used here to verify identical outputs between baseline and optimised workflows.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Per-task log</oasis:entry>
         <oasis:entry colname="col2" align="left">An individual log file (e.g. <monospace>slurm_output_%A_%a.out</monospace>) generated by each task, recording program outputs and error messages.</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1" align="left">Pipeline vs. workflow</oasis:entry>
         <oasis:entry colname="col2" align="left"><italic>Workflow</italic>: the scientific sequence of operations (GP emulation <inline-formula><mml:math id="M87" display="inline"><mml:mo>→</mml:mo></mml:math></inline-formula> GAM fitting <inline-formula><mml:math id="M88" display="inline"><mml:mo>→</mml:mo></mml:math></inline-formula> analysis). <italic>Pipeline</italic>: the software system that automates and manages workflow reproducibly on HPC resources.</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

<table-wrap id="TA1d"><label>Table A1</label><caption><p id="d2e2806">Continued.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="2">
     <oasis:colspec colnum="1" colname="col1" align="justify" colwidth="3cm"/>
     <oasis:colspec colnum="2" colname="col2" align="justify" colwidth="14cm"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Term/Acronym</oasis:entry>
         <oasis:entry colname="col2" align="left">Definition and Explanation</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">PPE (Perturbed-Parameter Ensemble)</oasis:entry>
         <oasis:entry colname="col2" align="left">A collection of model simulations designed to efficiently span high-dimensional parameter space.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Post-check script</oasis:entry>
         <oasis:entry colname="col2" align="left">A lightweight diagnostic script that scans task logs to identify failed or incomplete tasks and automatically resubmits only those tasks.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">rBAM (R Big Additive Model)</oasis:entry>
         <oasis:entry colname="col2" align="left">Implementation of GAM fitting using the R package <monospace>mgcv::bam()</monospace>, designed specifically for large data sets. It uses fast restricted maximum likelihood (fREML) estimation, discrete smoothing, and multithreading for efficiency.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Robustness</oasis:entry>
         <oasis:entry colname="col2" align="left">The ability of the pipeline to complete large numbers of tasks without failure or manual intervention.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Scalable design</oasis:entry>
         <oasis:entry colname="col2" align="left">Architecture that allows the workflow to expand from small test cases to full global or multi-model analyses with minimal code change.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Scalability/scaling test</oasis:entry>
         <oasis:entry colname="col2" align="left">Measurement of how runtime and memory usage change as the number of tasks or data size increases. Good scalability means near-linear increases in wall time as workload grows.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Slurm</oasis:entry>
         <oasis:entry colname="col2" align="left">An open-source workload manager used on many current HPC systems to schedule and monitor jobs across compute nodes. Version 22.05 was used here.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Slurm job array</oasis:entry>
         <oasis:entry colname="col2" align="left">A mechanism in Slurm that submits and manages large numbers of related tasks as a single batch job. Each array index runs one task (one grid box <inline-formula><mml:math id="M89" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> month).</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Sparse design matrices</oasis:entry>
         <oasis:entry colname="col2" align="left">Memory-efficient representation of matrices that store only non-zero elements, reducing memory load.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Speed-up factor</oasis:entry>
         <oasis:entry colname="col2" align="left">The ratio of baseline to optimised runtimes (e.g. 25 <inline-formula><mml:math id="M90" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> faster). Indicates performance improvement.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Superlinear scaling</oasis:entry>
         <oasis:entry colname="col2" align="left">A performance regime in which total runtime or memory increases faster than linearly with the number of concurrent tasks, often caused by node-level contention, shared resource limits, or excessive inter-process communication.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Task</oasis:entry>
         <oasis:entry colname="col2" align="left">The smallest independent computational unit in a pipeline (e.g. analysis in one model grid box for one month of data).</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Task-centred folder layout</oasis:entry>
         <oasis:entry colname="col2" align="left">Directory structure in which each grid box/month task has its own subfolder containing configuration, logs, and outputs, facilitating parallel execution and recovery.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Task-level parallelism</oasis:entry>
         <oasis:entry colname="col2" align="left">Running many small, independent tasks concurrently to maximise HPC utilisation.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Task table</oasis:entry>
         <oasis:entry colname="col2" align="left">A structured list or index that maps each task (month, latitude, longitude combination) to its associated Slurm array index for organised execution and recovery.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Task-table orchestration</oasis:entry>
         <oasis:entry colname="col2" align="left">Replacement for loop-based execution in which a structured table of tasks is distributed automatically to many CPUs or nodes.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Throttling</oasis:entry>
         <oasis:entry colname="col2" align="left">Limiting the number of concurrent tasks in a job array to avoid overloading the file system or exceeding resource limits.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Throttling limit</oasis:entry>
         <oasis:entry colname="col2" align="left">The maximum number of simultaneous tasks allowed in a job array to prevent filesystem contention (typically 200–500 tasks).</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Throughput</oasis:entry>
         <oasis:entry colname="col2" align="left">The total amount of computational work completed per unit time; improved here through concurrency and task isolation.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Variance decomposition</oasis:entry>
         <oasis:entry colname="col2" align="left">Statistical breakdown of the total output variance into contributions from individual model parameters and interactions between them.</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1" align="left">Vectorised terms-based computation</oasis:entry>
         <oasis:entry colname="col2" align="left">The optimised method for GAM variances decomposition using <monospace>predict(..., type="terms")</monospace> from <monospace>mgcv::bam()</monospace>. It evaluates all smooth terms simultaneously, obtaining per-parameter contributions <inline-formula><mml:math id="M91" display="inline"><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and computing <inline-formula><mml:math id="M92" display="inline"><mml:mrow><mml:mi mathvariant="normal">Var</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M93" display="inline"><mml:mrow><mml:mi mathvariant="normal">sign</mml:mi><mml:mo>[</mml:mo><mml:mi mathvariant="normal">Cov</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> in one batched pass. Numerically equivalent to the baseline median-hold approach but <inline-formula><mml:math id="M94" display="inline"><mml:mo>≥</mml:mo></mml:math></inline-formula> 20 <inline-formula><mml:math id="M95" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> faster.</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

<table-wrap id="TA1e"><label>Table A1</label><caption><p id="d2e3142">Continued.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="2">
     <oasis:colspec colnum="1" colname="col1" align="justify" colwidth="3cm"/>
     <oasis:colspec colnum="2" colname="col2" align="justify" colwidth="14cm"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Term/Acronym</oasis:entry>
         <oasis:entry colname="col2" align="left">Definition and Explanation</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Wall time</oasis:entry>
         <oasis:entry colname="col2" align="left">The total elapsed real time taken for a job or task to complete, including computation and waiting time on shared resources.</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1" align="left">YAML/JSON configuration files</oasis:entry>
         <oasis:entry colname="col2" align="left">Human-readable text files used to specify metadata (e.g. month, grid location, and variable name); useful in ensuring a reproducible configuration across runs.</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

</app>

<app id="App1.Ch1.S2">
  <label>Appendix B</label><title>Validation of GP emulator workflow optimisation</title>
      <p id="d2e3195">To confirm that the workflow optimisation did not alter emulator predictions, we compared probability density functions of mean H<sub>2</sub>SO<sub>4</sub> at representative grid points for January 2017.</p>

      <fig id="FB1"><label>Figure B1</label><caption><p id="d2e3218">Probability density functions of mean H<sub>2</sub>SO<sub>4</sub> for January 2017, obtained from one million parameter combinations using GP emulators. Panels <bold>(a)</bold>–<bold>(d)</bold> show results for four randomly selected grid points.</p></caption>
        
        <graphic xlink:href="https://gmd.copernicus.org/articles/19/7767/2026/gmd-19-7767-2026-f08.png"/>

      </fig>


</app>

<app id="App1.Ch1.S3">
  <label>Appendix C</label><title>Spatial consistency of GAM variance decomposition</title>
      <p id="d2e3263">Figure <xref ref-type="fig" rid="FC1"/> shows the spatial distribution of agreement between the baseline and optimised GAM variance decompositions across the analysed domain. For each grid box, the Pearson correlation coefficient is computed between the full set of parameter variance contributions derived from the two implementations. The results indicate consistently high agreement across the domain, with the majority of grid boxes exhibiting correlation coefficients close to unity. Out of the 100 grid boxes analysed, only four show a reduced correlation (<inline-formula><mml:math id="M100" display="inline"><mml:mrow><mml:mi>r</mml:mi><mml:mo>&lt;</mml:mo><mml:mn mathvariant="normal">0.99</mml:mn></mml:mrow></mml:math></inline-formula>), highlighted in Fig. <xref ref-type="fig" rid="FC1"/>.</p>
      <p id="d2e3282">These localised deviations are consistent with the small differences observed in Fig. <xref ref-type="fig" rid="F5"/> and are attributable to numerical and algorithmic differences between the implementations, including spline representation, smoothing parameter estimation, and discretisation in the optimised workflow. The magnitude and spatial extent of these differences are limited, and they do not affect the overall interpretation of parameter importance. This spatial analysis supports the conclusion that the optimised workflow preserves the statistical behaviour of the baseline method across the domain.</p>

      <fig id="FC1"><label>Figure C1</label><caption><p id="d2e3289">Spatial distribution of Pearson correlation coefficients (<inline-formula><mml:math id="M101" display="inline"><mml:mi>r</mml:mi></mml:math></inline-formula>) between baseline (<monospace>pyGAM</monospace>) and optimised (rBAM) variance contributions across 100 grid boxes for January. Each point represents one grid cell. Grid boxes with <inline-formula><mml:math id="M102" display="inline"><mml:mrow><mml:mi>r</mml:mi><mml:mo>&lt;</mml:mo><mml:mn mathvariant="normal">0.99</mml:mn></mml:mrow></mml:math></inline-formula> are highlighted with red circles.</p></caption>
        
        <graphic xlink:href="https://gmd.copernicus.org/articles/19/7767/2026/gmd-19-7767-2026-f09.png"/>

      </fig>


</app>

<app id="App1.Ch1.S4">
  <label>Appendix D</label><title>Runtime and peak memory usage</title>

<table-wrap id="TD1"><label>Table D1</label><caption><p id="d2e3336">Runtime and peak memory usage for baseline and optimised pipeline, by stage and end-to-end. Speed-up and memory-reduction factors are calculated as baseline/optimised.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="4">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Metric</oasis:entry>
         <oasis:entry colname="col2">GAM fitting</oasis:entry>
         <oasis:entry colname="col3">GP emulation</oasis:entry>
         <oasis:entry colname="col4">End-to-end</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">Baseline time (s)</oasis:entry>
         <oasis:entry colname="col2">10 623</oasis:entry>
         <oasis:entry colname="col3">6177</oasis:entry>
         <oasis:entry colname="col4">16 800</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Optimised time (s)</oasis:entry>
         <oasis:entry colname="col2">511</oasis:entry>
         <oasis:entry colname="col3">154</oasis:entry>
         <oasis:entry colname="col4">665</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Speed-up (<inline-formula><mml:math id="M103" display="inline"><mml:mo lspace="0mm">×</mml:mo></mml:math></inline-formula>)</oasis:entry>
         <oasis:entry colname="col2">20.8</oasis:entry>
         <oasis:entry colname="col3">40.1</oasis:entry>
         <oasis:entry colname="col4">25.3</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Baseline peak mem. (GB)</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M104" display="inline"><mml:mo>∼</mml:mo></mml:math></inline-formula> 100</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M105" display="inline"><mml:mo>∼</mml:mo></mml:math></inline-formula> 50</oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M106" display="inline"><mml:mo>∼</mml:mo></mml:math></inline-formula> 100</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Optimised peak mem. (GB)</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M107" display="inline"><mml:mo>∼</mml:mo></mml:math></inline-formula> 10</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M108" display="inline"><mml:mo>∼</mml:mo></mml:math></inline-formula> 6.5</oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M109" display="inline"><mml:mo>∼</mml:mo></mml:math></inline-formula> 10</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Mem. reduction (<inline-formula><mml:math id="M110" display="inline"><mml:mo lspace="0mm">×</mml:mo></mml:math></inline-formula>)</oasis:entry>
         <oasis:entry colname="col2">10.0</oasis:entry>
         <oasis:entry colname="col3">7.7</oasis:entry>
         <oasis:entry colname="col4">10.0</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table><table-wrap-foot><p id="d2e3339">Notes: Times are median walltime across tasks; memory values are approximate peaks per task.</p></table-wrap-foot></table-wrap>

</app>
  </app-group><notes notes-type="codedataavailability"><title>Code and data availability</title>

      <p id="d2e3522">All source code developed for this study, including the pipeline scripts, configuration files, and figure-generation notebooks, is publicly available through the GitHub repository <uri>https://github.com/Kunal198/gp-gam-optimisation-pipeline</uri> (last access: 18 August 2026). The tagged version corresponding to this paper is permanently archived on Zenodo (<ext-link xlink:href="https://doi.org/10.5281/zenodo.17543623" ext-link-type="DOI">10.5281/zenodo.17543623</ext-link>, <xref ref-type="bibr" rid="bib1.bibx5" id="altparen.34"/>). The repository includes a top-level <monospace>README.md</monospace> file providing step-by-step instructions and the commands required to reproduce all benchmarks, figures, and tables.</p>

      <p id="d2e3537">All simulation outputs required to reproduce the figures and tables in this paper are available as an accompanying dataset on Zenodo (<ext-link xlink:href="https://doi.org/10.5281/zenodo.17544324" ext-link-type="DOI">10.5281/zenodo.17544324</ext-link>, <xref ref-type="bibr" rid="bib1.bibx6" id="altparen.35"/>). The dataset includes emulator predictions, GAM-fitting outputs, and benchmarking measurements for the baseline and optimised workflows, together with metadata describing the input PPE configurations.</p>
  </notes><notes notes-type="authorcontribution"><title>Author contributions</title>

      <p id="d2e3549">KG designed and implemented the optimised pipeline, executed the experiments, performed the analysis, and prepared all figures and tables. LAR developed the baseline workflow scripts and provided methodological input. KG and LAR jointly interpreted the results and wrote the manuscript. Both authors reviewed and approved the final version.</p>
  </notes><notes notes-type="competinginterests"><title>Competing interests</title>

      <p id="d2e3555">The contact author has declared that neither of the authors has any competing interests.</p>
  </notes><notes notes-type="disclaimer"><title>Disclaimer</title>

      <p id="d2e3561">Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. The authors bear the ultimate responsibility for providing appropriate place names. Views expressed in the text are those of the authors and do not necessarily reflect the views of the publisher.</p>
  </notes><ack><title>Acknowledgements</title><p id="d2e3570">We thank Prof. Ken Carslaw for invaluable discussions and guidance throughout the development of this work. We are also grateful to Jill Johnson, Jonathan Owen, Léa Prévost, Iain Webb, and Jeremy Oakley for detailed feedback on early versions of the pipeline and manuscript. We additionally thank Lindsay Lee and Jill Johnson for sharing legacy code that was incorporated into the baseline workflow.</p><p id="d2e3572">We acknowledge the use of the JASMIN super-data-cluster, managed by the UK Centre for Environmental Data Analysis (CEDA), which hosted the simulations and benchmarking experiments presented in this paper. The PPE used in this research was created using the ARCHER UK National Supercomputing Service (<uri>http://www.archer.ac.uk</uri>, last access: 21 January 2021) under project allocation <monospace>n02-NEP013406</monospace>.</p></ack><notes notes-type="financialsupport"><title>Financial support</title>

      <p id="d2e3583">We acknowledge funding from the UK Natural Environment Research Council (NERC) under grants A-CURE (NE/P013406/1) and Aerosol-MFR (NE/X013901/1). Earlier related collaborations that informed this work were supported by NERC grant NE/G006148/1 (AEROS) and the FORCES project under the European Union's Horizon 2020 research programme with grant agreement 821205. LR was supported by the Met Office Hadley Centre Climate Programme funded by DSIT.</p>
  </notes><notes notes-type="reviewstatement"><title>Review statement</title>

      <p id="d2e3589">This paper was edited by Dan Lu and reviewed by two anonymous referees.</p>
  </notes><ref-list>
    <title>References</title>

      <ref id="bib1.bibx1"><label>Bellouin et al.(2020)</label><mixed-citation>Bellouin, N., Quaas, J., Gryspeerdt, E., Kinne, S., Stier, P., Watson-Parris, D., Boucher, O., Carslaw, K. S., Christensen, M., Daniau, A.-L., Dufresne, J.-L., Feingold, G., Fiedler, S., Forster, P., Gettelman, A., Haywood, J. M., Lohmann, U., Malavelle, F., Mauritsen, T., McCoy, D. T., Myhre, G., Mülmenstädt, J., Neubauer, D., Possner, A., Rugenstein, M., Sato, Y., Schulz, M., Schwartz, S. E., Sourdeval, O., Storelvmo, T., Toll, V., Winker, D., and Stevens, B.: Bounding Global Aerosol Radiative Forcing of Climate Change, Rev. Geophys., 58, e2019RG000660, <ext-link xlink:href="https://doi.org/10.1029/2019RG000660" ext-link-type="DOI">10.1029/2019RG000660</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx2"><label>Carnell(2022)</label><mixed-citation>Carnell, R.: lhs: Latin Hypercube Samples, r package version 1.1.6, <uri>https://CRAN.R-project.org/package=lhs</uri> (last access: 18 August 2026), 2022.</mixed-citation></ref>
      <ref id="bib1.bibx3"><label>Carslaw et al.(2026)Carslaw, Regayre, Proske, Gettelman, Sexton, Qian, Marshall, Wild, van Lier-Walqui, Oertel, Peatier, Yang, Johnson, Li, McCoy, Sanderson, Williamson, Elsaesser, Yamazaki, and Booth</label><mixed-citation>Carslaw, K. S., Regayre, L. A., Proske, U., Gettelman, A., Sexton, D. M. H., Qian, Y., Marshall, L. R., Wild, O., van Lier-Walqui, M., Oertel, A., Peatier, S., Yang, B., Johnson, J. S., Li, S., McCoy, D. T., Sanderson, B. M., Williamson, C. J., Elsaesser, G. S., Yamazaki, K., and Booth, B. B. B.: Opinion: The importance and future development of perturbed parameter ensembles in climate and atmospheric science, Atmos. Chem. Phys., 26, 4651–4667, <ext-link xlink:href="https://doi.org/10.5194/acp-26-4651-2026" ext-link-type="DOI">10.5194/acp-26-4651-2026</ext-link>, 2026.</mixed-citation></ref>
      <ref id="bib1.bibx4"><label>Durack et al.(2025)Durack, Taylor, Gleckler, Meehl, Lawrence, Covey, Stouffer, Levavasseur, Ben-Nasser, Denvil, Stockhause, Gregory, Juckes, Ames, Antonio, Bader, Dunne, Ellis, Eyring, Fiore, Joussaume, Kershaw, Lamarque, Lautenschlager, Lee, Mauzey, Mizielinski, Nassisi, Nuzzo, O'Rourke, Painter, Potter, Rodriguez, and Williams</label><mixed-citation>Durack, P. J., Taylor, K. E., Gleckler, P. J., Meehl, G. A., Lawrence, B. N., Covey, C., Stouffer, R. J., Levavasseur, G., Ben-Nasser, A., Denvil, S., Stockhause, M., Gregory, J. M., Juckes, M., Ames, S. K., Antonio, F., Bader, D. C., Dunne, J. P., Ellis, D., Eyring, V., Fiore, S. L., Joussaume, S., Kershaw, P., Lamarque, J.-F., Lautenschlager, M., Lee, J., Mauzey, C. F., Mizielinski, M., Nassisi, P., Nuzzo, A., O’Rourke, E., Painter, J., Potter, G. L., Rodriguez, S., and Williams, D. N.: The Coupled Model Intercomparison Project (CMIP): Reviewing project history, evolution, infrastructure and implementation, EGUsphere [preprint], <ext-link xlink:href="https://doi.org/10.5194/egusphere-2024-3729" ext-link-type="DOI">10.5194/egusphere-2024-3729</ext-link>, 2025.</mixed-citation></ref>
      <ref id="bib1.bibx5"><label>Ghosh and Regayre(2025a)</label><mixed-citation>Ghosh, K. and Regayre, L. A.: gp-gam-optimisation-pipeline: High-performance workflow for Gaussian Process emulation and GAM fitting, Zenodo [code], <ext-link xlink:href="https://doi.org/10.5281/zenodo.17543623" ext-link-type="DOI">10.5281/zenodo.17543623</ext-link>, 2025a.</mixed-citation></ref>
      <ref id="bib1.bibx6"><label>Ghosh and Regayre(2025b)</label><mixed-citation>Ghosh, K. and Regayre, L. A.: gp-gam-optimisation-dataset: Example input and output data for Gaussian Process and GAM workflow (H<sub>2</sub>SO<sub>4</sub>, January 2017), Zenodo [data set], <ext-link xlink:href="https://doi.org/10.5281/zenodo.17544324" ext-link-type="DOI">10.5281/zenodo.17544324</ext-link>, 2025b.</mixed-citation></ref>
      <ref id="bib1.bibx7"><label>Johnson et al.(2015)Johnson, Cui, Lee, Gosling, Blyth, and Carslaw</label><mixed-citation> Johnson, J., Cui, Z., Lee, L., Gosling, J., Blyth, A., and Carslaw, K.: Evaluating uncertainty in convective cloud microphysics using statistical emulation, J. Adv. Model. Earth Sy., 7, 162–187, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx8"><label>Johnson et al.(2018)Johnson, Regayre, Yoshioka, Pringle, Lee, Sexton, Rostron, Booth, and Carslaw</label><mixed-citation>Johnson, J. S., Regayre, L. A., Yoshioka, M., Pringle, K. J., Lee, L. A., Sexton, D. M. H., Rostron, J. W., Booth, B. B. B., and Carslaw, K. S.: The importance of comprehensive parameter sampling and multiple observations for robust constraint of aerosol radiative forcing, Atmos. Chem. Phys., 18, 13031–13053, <ext-link xlink:href="https://doi.org/10.5194/acp-18-13031-2018" ext-link-type="DOI">10.5194/acp-18-13031-2018</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx9"><label>Johnson et al.(2020)Johnson, Regayre, Yoshioka, Pringle, Turnock, Browse, Sexton, Rostron, Schutgens, Partridge, Liu, Allan, Coe, Ding, Cohen, Atanacio, Vakkari, Asmi, and Carslaw</label><mixed-citation>Johnson, J. S., Regayre, L. A., Yoshioka, M., Pringle, K. J., Turnock, S. T., Browse, J., Sexton, D. M. H., Rostron, J. W., Schutgens, N. A. J., Partridge, D. G., Liu, D., Allan, J. D., Coe, H., Ding, A., Cohen, D. D., Atanacio, A., Vakkari, V., Asmi, E., and Carslaw, K. S.: Robust observational constraint of uncertain aerosol processes and emissions in a climate model and the effect on aerosol radiative forcing, Atmos. Chem. Phys., 20, 9491–9524, <ext-link xlink:href="https://doi.org/10.5194/acp-20-9491-2020" ext-link-type="DOI">10.5194/acp-20-9491-2020</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx10"><label>Knutti(2008)</label><mixed-citation> Knutti, R.: Should we believe model predictions of future climate change?, Philos. T. Roy. Soc. A, 366, 4647–4664, 2008.</mixed-citation></ref>
      <ref id="bib1.bibx11"><label>Lawrence et al.(2013)Lawrence, Bennett, Churchill, Juckes, Kershaw, Pascoe, Pepler, Pritchard, and Stephens</label><mixed-citation> Lawrence, B. N., Bennett, V. L., Churchill, J., Juckes, M., Kershaw, P., Pascoe, S., Pepler, S., Pritchard, M., and Stephens, A.: Storing and manipulating environmental big data with JASMIN, in: 2013 IEEE international conference on big data, 68–75, IEEE, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx12"><label>Lee et al.(2011)Lee, Carslaw, Pringle, Mann, and Spracklen</label><mixed-citation>Lee, L. A., Carslaw, K. S., Pringle, K. J., Mann, G. W., and Spracklen, D. V.: Emulation of a complex global aerosol model to quantify sensitivity to uncertain parameters, Atmos. Chem. Phys., 11, 12253–12273, <ext-link xlink:href="https://doi.org/10.5194/acp-11-12253-2011" ext-link-type="DOI">10.5194/acp-11-12253-2011</ext-link>, 2011.</mixed-citation></ref>
      <ref id="bib1.bibx13"><label>Lee et al.(2012)Lee, Carslaw, Pringle, and Mann</label><mixed-citation>Lee, L. A., Carslaw, K. S., Pringle, K. J., and Mann, G. W.: Mapping the uncertainty in global CCN using emulation, Atmos. Chem. Phys., 12, 9739–9751, <ext-link xlink:href="https://doi.org/10.5194/acp-12-9739-2012" ext-link-type="DOI">10.5194/acp-12-9739-2012</ext-link>, 2012.</mixed-citation></ref>
      <ref id="bib1.bibx14"><label>Lee et al.(2016)Lee, Reddington, and Carslaw</label><mixed-citation>Lee, L. A., Reddington, C. L., and Carslaw, K. S.: On the relationship between aerosol model uncertainty and radiative forcing uncertainty, P. Natl. Acad. Sci., 113, 5820–5827, <ext-link xlink:href="https://doi.org/10.1073/pnas.1507050113" ext-link-type="DOI">10.1073/pnas.1507050113</ext-link>, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx15"><label>Mersmann(2022)</label><mixed-citation>Mersmann, O.: truncnorm: Truncated Normal Distribution, r package version 1.0-9, <uri>https://CRAN.R-project.org/package=truncnorm</uri> (last access: 18 August 2026), 2022.</mixed-citation></ref>
      <ref id="bib1.bibx16"><label>Oakley and O'Hagan(2002)</label><mixed-citation> Oakley, J. and O'Hagan, A.: Bayesian inference for the uncertainty distribution of computer model outputs, Biometrika, 89, 769–784, 2002.</mixed-citation></ref>
      <ref id="bib1.bibx17"><label>O'Hagan(2006)</label><mixed-citation>O'Hagan, A.: Bayesian analysis of computer code outputs: A tutorial, Reliab. Eng. Syst. Safe., 91, 1290–1300, <ext-link xlink:href="https://doi.org/10.1016/j.ress.2005.11.025" ext-link-type="DOI">10.1016/j.ress.2005.11.025</ext-link>, 2006.</mixed-citation></ref>
      <ref id="bib1.bibx18"><label>Prévost et al.(2026)Prévost, Regayre, Johnson, McNeall, Milton, and Carslaw</label><mixed-citation>Prévost, L. M. C., Regayre, L. A., Johnson, J. S., McNeall, D., Milton, S., and Carslaw, K. S.: Detection of potential structural deficiencies in a global aerosol model using a perturbed parameter ensemble, Atmos. Chem. Phys., 26, 2487–2530, <ext-link xlink:href="https://doi.org/10.5194/acp-26-2487-2026" ext-link-type="DOI">10.5194/acp-26-2487-2026</ext-link>, 2026.</mixed-citation></ref>
      <ref id="bib1.bibx19"><label>Pujol et al.(2017)Pujol, Iooss, and Janon</label><mixed-citation>Pujol, G., Iooss, B., and Janon, A.: sensitivity: Global Sensitivity Analysis of Model Outputs, r package version 1.18.1, <uri>https://CRAN.R-project.org/package=sensitivity</uri> (last access: 18 August 2026), 2017.</mixed-citation></ref>
      <ref id="bib1.bibx20"><label>R Core Team(2022)</label><mixed-citation>R Core Team: R: A Language and Environment for Statistical Computing, R Foundation for Statistical Computing, Vienna, Austria, <uri>https://www.R-project.org/</uri> (last access: 18 August 2026), 2022.</mixed-citation></ref>
      <ref id="bib1.bibx21"><label>Regayre et al.(2014)Regayre, Pringle, Booth, Lee, Mann, Browse, Woodhouse, Rap, Reddington, and Carslaw</label><mixed-citation>Regayre, L. A., Pringle, K. J., Booth, B. B. B., Lee, L. A., Mann, G. W., Browse, J., Woodhouse, M. T., Rap, A., Reddington, C. L., and Carslaw, K. S.: Uncertainty in the magnitude of aerosol-cloud radiative forcing over recent decades, Geophys. Res. Lett., 41, 9040–9049, <ext-link xlink:href="https://doi.org/10.1002/2014GL062029" ext-link-type="DOI">10.1002/2014GL062029</ext-link>, 2014.</mixed-citation></ref>
      <ref id="bib1.bibx22"><label>Regayre et al.(2023)Regayre, Deaconu, Grosvenor, Sexton, Symonds, Langton, Watson-Paris, Mulcahy, Pringle, Richardson, Johnson, Rostron, Gordon, Lister, Stier, and Carslaw</label><mixed-citation>Regayre, L. A., Deaconu, L., Grosvenor, D. P., Sexton, D. M. H., Symonds, C., Langton, T., Watson-Paris, D., Mulcahy, J. P., Pringle, K. J., Richardson, M., Johnson, J. S., Rostron, J. W., Gordon, H., Lister, G., Stier, P., and Carslaw, K. S.: Identifying climate model structural inconsistencies allows for tight constraint of aerosol radiative forcing, Atmos. Chem. Phys., 23, 8749–8768, <ext-link xlink:href="https://doi.org/10.5194/acp-23-8749-2023" ext-link-type="DOI">10.5194/acp-23-8749-2023</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx23"><label>Regayre et al.(2026)Regayre, Prévost, Ghosh, Johnson, Oakley, Owen, Webb, and Carslaw</label><mixed-citation>Regayre, L. A., Prévost, L. M. C., Ghosh, K., Johnson, J. S., Oakley, J. E., Owen, J., Webb, I., and Carslaw, K. S.: Remaining aerosol forcing uncertainty after observational constraint and the processes that cause it, Atmos. Chem. Phys., 26, 2293–2317, <ext-link xlink:href="https://doi.org/10.5194/acp-26-2293-2026" ext-link-type="DOI">10.5194/acp-26-2293-2026</ext-link>, 2026.</mixed-citation></ref>
      <ref id="bib1.bibx24"><label>Roustant et al.(2012)Roustant, Ginsbourger, and Deville</label><mixed-citation>Roustant, O., Ginsbourger, D., and Deville, Y.: DiceKriging, DiceOptim: Two R Packages for the Analysis of Computer Experiments by Kriging-Based Metamodeling and Optimization, J. Stat. Softw., 51, 1–55, <ext-link xlink:href="https://doi.org/10.18637/jss.v051.i01" ext-link-type="DOI">10.18637/jss.v051.i01</ext-link>, 2012.</mixed-citation></ref>
      <ref id="bib1.bibx25"><label>Saltelli et al.(2008)Saltelli, Ratto, Andres, Campolongo, Cariboni, Gatelli, Saisana, and Tarantola</label><mixed-citation>Saltelli, A., Ratto, M., Andres, T., Campolongo, F., Cariboni, J., Gatelli, D., Saisana, M., and Tarantola, S.: Global Sensitivity Analysis: The Primer, Wiley, <ext-link xlink:href="https://doi.org/10.1002/9780470725184" ext-link-type="DOI">10.1002/9780470725184</ext-link>, 2008.</mixed-citation></ref>
      <ref id="bib1.bibx26"><label>Sellar et al.(2019)</label><mixed-citation>Sellar, A. A., Jones, C. G., Mulcahy, J. P., Tang, Y., Yool, A., Wiltshire, A., O'Connor, F. M., Stringer, M., Hill, R., Palmieri, J., Woodward, S., de Mora, L., Kuhlbrodt, T., Rumbold, S. T., Kelley, D. I., Ellis, R., Johnson, C. E., Walton, J., Abraham, N. L., Andrews, M. B., Andrews, T., Archibald, A. T., Berthou, S., Burke, E., Blockley, E., Carslaw, K., Dalvi, M., Edwards, J., Folberth, G. A., Gedney, N., Griffiths, P. T., Harper, A. B., Hendry, M. A., Hewitt, A. J., Johnson, B., Jones, A., Jones, C. D., Keeble, J., Liddicoat, S., Morgenstern, O., Parker, R. J., Predoi, V., Robertson, E., Siahaan, A., Smith, R. S., Swaminathan, R., Woodhouse, M. T., Zeng, G., and Zerroukat, M.: UKESM1: Description and Evaluation of the U.K. Earth System Model, J. Adv. Model. Earth Sy., 11, 4513–4558, <ext-link xlink:href="https://doi.org/10.1029/2019MS001739" ext-link-type="DOI">10.1029/2019MS001739</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx27"><label>Servén et al.(2018)Servén, Brummitt, Abedi, and Hlink</label><mixed-citation>Servén, D., Brummitt, C., Abedi, H., and Hlink: dswah/pyGAM: v0.8.0, Zenodo, <ext-link xlink:href="https://doi.org/10.5281/zenodo.1208723" ext-link-type="DOI">10.5281/zenodo.1208723</ext-link>,  2018.</mixed-citation></ref>
      <ref id="bib1.bibx28"><label>Servera et al.(2023)Servera, Martino, Verrelst, and Camps-Valls</label><mixed-citation> Servera, J. V., Martino, L., Verrelst, J., and Camps-Valls, G.: Multifidelity Gaussian process emulation for atmospheric radiative transfer models, IEEE T. Geosci. Remote Sens., 61, 1–10, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx29"><label>Sexton et al.(2012)Sexton, Murphy, Collins, and Webb</label><mixed-citation> Sexton, D. M., Murphy, J. M., Collins, M., and Webb, M. J.: Multivariate probabilistic projections using imperfect climate models part I: outline of methodology, Clim. Dynam., 38, 2513–2542, 2012.</mixed-citation></ref>
      <ref id="bib1.bibx30"><label>Sobol'(2001)</label><mixed-citation> Sobol', I. M.: Global sensitivity indices for nonlinear mathematical models and their Monte Carlo estimates, Math. Comput. Simulat., 55, 271–280, 2001.</mixed-citation></ref>
      <ref id="bib1.bibx31"><label>Sottile and Gohel(2022)</label><mixed-citation>Sottile, G. and Gohel, D.: trapezoid: Trapezoidal Distribution, r package version 0.6, <uri>https://CRAN.R-project.org/package=trapezoid</uri> (last access: 18 August 2026), 2022.</mixed-citation></ref>
      <ref id="bib1.bibx32"><label>Stouffer et al.(2017)Stouffer, Eyring, Meehl, Bony, Senior, Stevens, and Taylor</label><mixed-citation>Stouffer, R. J., Eyring, V., Meehl, G. A., Bony, S., Senior, C., Stevens, B., and Taylor, K.: CMIP5 scientific gaps and recommendations for CMIP6, B. Am. Meteorol. Soc., 98, 95–105, 2017.  </mixed-citation></ref>
      <ref id="bib1.bibx33"><label>Susiluoto et al.(2020)Susiluoto, Spantini, Haario, Härkönen, and Marzouk</label><mixed-citation>Susiluoto, J., Spantini, A., Haario, H., Härkönen, T., and Marzouk, Y.: Efficient multi-scale Gaussian process regression for massive remote sensing data with satGP v0.1.2, Geosci. Model Dev., 13, 3439–3463, <ext-link xlink:href="https://doi.org/10.5194/gmd-13-3439-2020" ext-link-type="DOI">10.5194/gmd-13-3439-2020</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx34"><label>Watson-Parris et al.(2021)Watson-Parris, Williams, Deaconu, and Stier</label><mixed-citation>Watson-Parris, D., Williams, A., Deaconu, L., and Stier, P.: Model calibration using ESEm v1.1.0 – an open, scalable Earth system emulator, Geosci. Model Dev., 14, 7659–7672, <ext-link xlink:href="https://doi.org/10.5194/gmd-14-7659-2021" ext-link-type="DOI">10.5194/gmd-14-7659-2021</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx35"><label>Wickham and Bryan(2023)</label><mixed-citation>Wickham, H. and Bryan, J.: readr: Read Rectangular Text Data, r package version 2.1.4, <uri>https://CRAN.R-project.org/package=readr</uri> (last access: 18 August 2026), 2023.</mixed-citation></ref>
      <ref id="bib1.bibx36"><label>Wood(2020)</label><mixed-citation> Wood, S. N.: Inference and computation with generalized additive models and their extensions, Test, 29, 307–339, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx37"><label>Wood et al.(2015)Wood, Goude, and Shaw</label><mixed-citation> Wood, S. N., Goude, Y., and Shaw, S.: Generalized additive models for large data sets, J. Roy. Stat. Soc. C, 64, 139–155, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx38"><label>Yang et al.(2025)Yang, Elsaesser, van Lier-Walqui, and Eidhammer</label><mixed-citation>Yang, Q., Elsaesser, G. S., van Lier-Walqui, M., and Eidhammer, T.: A simple emulator that enables interpretation of parameter-output relationships, applied to two climate model PPEs, J. Adv. Model. Earth Sy., 17, e2024MS004766, <ext-link xlink:href="https://doi.org/10.1029/2024MS004766" ext-link-type="DOI">10.1029/2024MS004766</ext-link>, 2025.</mixed-citation></ref>
      <ref id="bib1.bibx39"><label>Yoshioka et al.(2019)Yoshioka, Regayre, Pringle, Johnson, Mann, Partridge, Sexton, Lister, Schutgens, Stier, Kipling, Bellouin, Browse, Booth, Johnson, Johnson, Mollard, Lee, and Carslaw</label><mixed-citation>Yoshioka, M., Regayre, L. A., Pringle, K. J., Johnson, J. S., Mann, G. W., Partridge, D. G., Sexton, D. M. H., Lister, G. M. S., Schutgens, N., Stier, P., Kipling, Z., Bellouin, N., Browse, J., Booth, B. B. B., Johnson, C. E., Johnson, B., Mollard, J. D. P., Lee, L., and Carslaw, K. S.: Ensembles of Global Climate Model Variants Designed for the Quantification and Constraint of Uncertainty in Aerosols and Their Radiative Forcing, J. Adv. Model. Earth Sy., 11, 3728–3754, <ext-link xlink:href="https://doi.org/10.1029/2019MS001628" ext-link-type="DOI">10.1029/2019MS001628</ext-link>, 2019.</mixed-citation></ref>

  </ref-list></back>
    <!--<article-title-html>Optimizing Gaussian process emulation and generalized additive model fitting for rapid, reproducible earth system model analysis</article-title-html>
<abstract-html/>
<ref-html id="bib1.bib1"><label>Bellouin et al.(2020)</label><mixed-citation>
      
Bellouin, N., Quaas, J., Gryspeerdt, E., Kinne, S., Stier, P., Watson-Parris,
D., Boucher, O., Carslaw, K. S., Christensen, M., Daniau, A.-L., Dufresne,
J.-L., Feingold, G., Fiedler, S., Forster, P., Gettelman, A., Haywood, J. M.,
Lohmann, U., Malavelle, F., Mauritsen, T., McCoy, D. T., Myhre, G.,
Mülmenstädt, J., Neubauer, D., Possner, A., Rugenstein, M., Sato, Y.,
Schulz, M., Schwartz, S. E., Sourdeval, O., Storelvmo, T., Toll, V., Winker,
D., and Stevens, B.: Bounding Global Aerosol Radiative Forcing of Climate
Change, Rev. Geophys., 58, e2019RG000660,
<a href="https://doi.org/10.1029/2019RG000660" target="_blank">https://doi.org/10.1029/2019RG000660</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib2"><label>Carnell(2022)</label><mixed-citation>
      
Carnell, R.: lhs: Latin Hypercube Samples, r package version
1.1.6,
<a href="https://CRAN.R-project.org/package=lhs" target="_blank"/> (last access: 18 August 2026), 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib3"><label>Carslaw et al.(2026)Carslaw, Regayre, Proske, Gettelman, Sexton,
Qian, Marshall, Wild, van Lier-Walqui, Oertel, Peatier, Yang, Johnson, Li,
McCoy, Sanderson, Williamson, Elsaesser, Yamazaki, and
Booth</label><mixed-citation>
      
Carslaw, K. S., Regayre, L. A., Proske, U., Gettelman, A., Sexton, D. M. H., Qian, Y., Marshall, L. R., Wild, O., van Lier-Walqui, M., Oertel, A., Peatier, S., Yang, B., Johnson, J. S., Li, S., McCoy, D. T., Sanderson, B. M., Williamson, C. J., Elsaesser, G. S., Yamazaki, K., and Booth, B. B. B.: Opinion: The importance and future development of perturbed parameter ensembles in climate and atmospheric science, Atmos. Chem. Phys., 26, 4651–4667, <a href="https://doi.org/10.5194/acp-26-4651-2026" target="_blank">https://doi.org/10.5194/acp-26-4651-2026</a>, 2026.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib4"><label>Durack et al.(2025)Durack, Taylor, Gleckler, Meehl, Lawrence, Covey,
Stouffer, Levavasseur, Ben-Nasser, Denvil, Stockhause, Gregory, Juckes, Ames,
Antonio, Bader, Dunne, Ellis, Eyring, Fiore, Joussaume, Kershaw, Lamarque,
Lautenschlager, Lee, Mauzey, Mizielinski, Nassisi, Nuzzo, O'Rourke, Painter,
Potter, Rodriguez, and Williams</label><mixed-citation>
      
Durack, P. J., Taylor, K. E., Gleckler, P. J., Meehl, G. A., Lawrence, B. N., Covey, C., Stouffer, R. J., Levavasseur, G., Ben-Nasser, A., Denvil, S., Stockhause, M., Gregory, J. M., Juckes, M., Ames, S. K., Antonio, F., Bader, D. C., Dunne, J. P., Ellis, D., Eyring, V., Fiore, S. L., Joussaume, S., Kershaw, P., Lamarque, J.-F., Lautenschlager, M., Lee, J., Mauzey, C. F., Mizielinski, M., Nassisi, P., Nuzzo, A., O’Rourke, E., Painter, J., Potter, G. L., Rodriguez, S., and Williams, D. N.: The Coupled Model Intercomparison Project (CMIP): Reviewing project history, evolution, infrastructure and implementation, EGUsphere [preprint], <a href="https://doi.org/10.5194/egusphere-2024-3729" target="_blank">https://doi.org/10.5194/egusphere-2024-3729</a>, 2025.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib5"><label>Ghosh and Regayre(2025a)</label><mixed-citation>
      
Ghosh, K. and Regayre, L. A.: gp-gam-optimisation-pipeline: High-performance
workflow for Gaussian Process emulation and GAM fitting, Zenodo [code],
<a href="https://doi.org/10.5281/zenodo.17543623" target="_blank">https://doi.org/10.5281/zenodo.17543623</a>, 2025a.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib6"><label>Ghosh and Regayre(2025b)</label><mixed-citation>
      
Ghosh, K. and Regayre, L. A.: gp-gam-optimisation-dataset: Example input and
output data for Gaussian Process and GAM workflow (H<sub>2</sub>SO<sub>4</sub>, January
2017), Zenodo [data set], <a href="https://doi.org/10.5281/zenodo.17544324" target="_blank">https://doi.org/10.5281/zenodo.17544324</a>, 2025b.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib7"><label>Johnson et al.(2015)Johnson, Cui, Lee, Gosling, Blyth, and
Carslaw</label><mixed-citation>
      
Johnson, J., Cui, Z., Lee, L., Gosling, J., Blyth, A., and Carslaw, K.:
Evaluating uncertainty in convective cloud microphysics using statistical
emulation, J. Adv. Model. Earth Sy., 7, 162–187, 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib8"><label>Johnson et al.(2018)Johnson, Regayre, Yoshioka, Pringle, Lee, Sexton,
Rostron, Booth, and Carslaw</label><mixed-citation>
      
Johnson, J. S., Regayre, L. A., Yoshioka, M., Pringle, K. J., Lee, L. A., Sexton, D. M. H., Rostron, J. W., Booth, B. B. B., and Carslaw, K. S.: The importance of comprehensive parameter sampling and multiple observations for robust constraint of aerosol radiative forcing, Atmos. Chem. Phys., 18, 13031–13053, <a href="https://doi.org/10.5194/acp-18-13031-2018" target="_blank">https://doi.org/10.5194/acp-18-13031-2018</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib9"><label>Johnson et al.(2020)Johnson, Regayre, Yoshioka, Pringle, Turnock,
Browse, Sexton, Rostron, Schutgens, Partridge, Liu, Allan, Coe, Ding, Cohen,
Atanacio, Vakkari, Asmi, and Carslaw</label><mixed-citation>
      
Johnson, J. S., Regayre, L. A., Yoshioka, M., Pringle, K. J., Turnock, S. T., Browse, J., Sexton, D. M. H., Rostron, J. W., Schutgens, N. A. J., Partridge, D. G., Liu, D., Allan, J. D., Coe, H., Ding, A., Cohen, D. D., Atanacio, A., Vakkari, V., Asmi, E., and Carslaw, K. S.: Robust observational constraint of uncertain aerosol processes and emissions in a climate model and the effect on aerosol radiative forcing, Atmos. Chem. Phys., 20, 9491–9524, <a href="https://doi.org/10.5194/acp-20-9491-2020" target="_blank">https://doi.org/10.5194/acp-20-9491-2020</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib10"><label>Knutti(2008)</label><mixed-citation>
      
Knutti, R.: Should we believe model predictions of future climate change?,
Philos. T. Roy. Soc. A, 366, 4647–4664, 2008.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib11"><label>Lawrence et al.(2013)Lawrence, Bennett, Churchill, Juckes, Kershaw,
Pascoe, Pepler, Pritchard, and Stephens</label><mixed-citation>
      
Lawrence, B. N., Bennett, V. L., Churchill, J., Juckes, M., Kershaw, P.,
Pascoe, S., Pepler, S., Pritchard, M., and Stephens, A.: Storing and
manipulating environmental big data with JASMIN, in: 2013 IEEE international
conference on big data, 68–75, IEEE, 2013.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib12"><label>Lee et al.(2011)Lee, Carslaw, Pringle, Mann, and
Spracklen</label><mixed-citation>
      
Lee, L. A., Carslaw, K. S., Pringle, K. J., Mann, G. W., and Spracklen, D. V.: Emulation of a complex global aerosol model to quantify sensitivity to uncertain parameters, Atmos. Chem. Phys., 11, 12253–12273, <a href="https://doi.org/10.5194/acp-11-12253-2011" target="_blank">https://doi.org/10.5194/acp-11-12253-2011</a>, 2011.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib13"><label>Lee et al.(2012)Lee, Carslaw, Pringle, and Mann</label><mixed-citation>
      
Lee, L. A., Carslaw, K. S., Pringle, K. J., and Mann, G. W.: Mapping the uncertainty in global CCN using emulation, Atmos. Chem. Phys., 12, 9739–9751, <a href="https://doi.org/10.5194/acp-12-9739-2012" target="_blank">https://doi.org/10.5194/acp-12-9739-2012</a>, 2012.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib14"><label>Lee et al.(2016)Lee, Reddington, and Carslaw</label><mixed-citation>
      
Lee, L. A., Reddington, C. L., and Carslaw, K. S.: On the relationship between
aerosol model uncertainty and radiative forcing uncertainty, P. Natl. Acad. Sci., 113, 5820–5827,
<a href="https://doi.org/10.1073/pnas.1507050113" target="_blank">https://doi.org/10.1073/pnas.1507050113</a>, 2016.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib15"><label>Mersmann(2022)</label><mixed-citation>
      
Mersmann, O.: truncnorm: Truncated Normal Distribution, r package
version 1.0-9,
<a href="https://CRAN.R-project.org/package=truncnorm" target="_blank"/> (last access: 18 August 2026), 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib16"><label>Oakley and O'Hagan(2002)</label><mixed-citation>
      
Oakley, J. and O'Hagan, A.: Bayesian inference for the uncertainty distribution
of computer model outputs, Biometrika, 89, 769–784, 2002.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib17"><label>O'Hagan(2006)</label><mixed-citation>
      
O'Hagan, A.: Bayesian analysis of computer code outputs: A tutorial,
Reliab. Eng. Syst. Safe., 91, 1290–1300,
<a href="https://doi.org/10.1016/j.ress.2005.11.025" target="_blank">https://doi.org/10.1016/j.ress.2005.11.025</a>, 2006.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib18"><label>Prévost et al.(2026)Prévost, Regayre, Johnson, McNeall, Milton,
and Carslaw</label><mixed-citation>
      
Prévost, L. M. C., Regayre, L. A., Johnson, J. S., McNeall, D., Milton, S., and Carslaw, K. S.: Detection of potential structural deficiencies in a global aerosol model using a perturbed parameter ensemble, Atmos. Chem. Phys., 26, 2487–2530, <a href="https://doi.org/10.5194/acp-26-2487-2026" target="_blank">https://doi.org/10.5194/acp-26-2487-2026</a>, 2026.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib19"><label>Pujol et al.(2017)Pujol, Iooss, and Janon</label><mixed-citation>
      
Pujol, G., Iooss, B., and Janon, A.: sensitivity: Global Sensitivity Analysis
of Model Outputs, r package
version 1.18.1,
<a href="https://CRAN.R-project.org/package=sensitivity" target="_blank"/> (last access: 18 August 2026), 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib20"><label>R Core Team(2022)</label><mixed-citation>
      
R Core Team: R: A Language and Environment for Statistical Computing, R
Foundation for Statistical Computing, Vienna, Austria,
<a href="https://www.R-project.org/" target="_blank"/> (last access: 18 August 2026), 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib21"><label>Regayre et al.(2014)Regayre, Pringle, Booth, Lee, Mann, Browse,
Woodhouse, Rap, Reddington, and Carslaw</label><mixed-citation>
      
Regayre, L. A., Pringle, K. J., Booth, B. B. B., Lee, L. A., Mann, G. W.,
Browse, J., Woodhouse, M. T., Rap, A., Reddington, C. L., and Carslaw, K. S.:
Uncertainty in the magnitude of aerosol-cloud radiative forcing over recent
decades, Geophys. Res. Lett., 41, 9040–9049,
<a href="https://doi.org/10.1002/2014GL062029" target="_blank">https://doi.org/10.1002/2014GL062029</a>, 2014.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib22"><label>Regayre et al.(2023)Regayre, Deaconu, Grosvenor, Sexton, Symonds,
Langton, Watson-Paris, Mulcahy, Pringle, Richardson, Johnson, Rostron,
Gordon, Lister, Stier, and Carslaw</label><mixed-citation>
      
Regayre, L. A., Deaconu, L., Grosvenor, D. P., Sexton, D. M. H., Symonds, C., Langton, T., Watson-Paris, D., Mulcahy, J. P., Pringle, K. J., Richardson, M., Johnson, J. S., Rostron, J. W., Gordon, H., Lister, G., Stier, P., and Carslaw, K. S.: Identifying climate model structural inconsistencies allows for tight constraint of aerosol radiative forcing, Atmos. Chem. Phys., 23, 8749–8768, <a href="https://doi.org/10.5194/acp-23-8749-2023" target="_blank">https://doi.org/10.5194/acp-23-8749-2023</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib23"><label>Regayre et al.(2026)Regayre, Prévost, Ghosh, Johnson, Oakley, Owen,
Webb, and Carslaw</label><mixed-citation>
      
Regayre, L. A., Prévost, L. M. C., Ghosh, K., Johnson, J. S., Oakley, J. E., Owen, J., Webb, I., and Carslaw, K. S.: Remaining aerosol forcing uncertainty after observational constraint and the processes that cause it, Atmos. Chem. Phys., 26, 2293–2317, <a href="https://doi.org/10.5194/acp-26-2293-2026" target="_blank">https://doi.org/10.5194/acp-26-2293-2026</a>, 2026.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib24"><label>Roustant et al.(2012)Roustant, Ginsbourger, and
Deville</label><mixed-citation>
      
Roustant, O., Ginsbourger, D., and Deville, Y.: DiceKriging, DiceOptim: Two R
Packages for the Analysis of Computer Experiments by Kriging-Based
Metamodeling and Optimization, J. Stat. Softw., 51, 1–55,
<a href="https://doi.org/10.18637/jss.v051.i01" target="_blank">https://doi.org/10.18637/jss.v051.i01</a>, 2012.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib25"><label>Saltelli et al.(2008)Saltelli, Ratto, Andres, Campolongo, Cariboni,
Gatelli, Saisana, and Tarantola</label><mixed-citation>
      
Saltelli, A., Ratto, M., Andres, T., Campolongo, F., Cariboni, J., Gatelli, D.,
Saisana, M., and Tarantola, S.: Global Sensitivity Analysis: The Primer,
Wiley, <a href="https://doi.org/10.1002/9780470725184" target="_blank">https://doi.org/10.1002/9780470725184</a>, 2008.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib26"><label>Sellar et al.(2019)</label><mixed-citation>
      
Sellar, A. A., Jones, C. G., Mulcahy, J. P., Tang, Y., Yool, A., Wiltshire, A.,
O'Connor, F. M., Stringer, M., Hill, R., Palmieri, J., Woodward, S., de Mora,
L., Kuhlbrodt, T., Rumbold, S. T., Kelley, D. I., Ellis, R., Johnson, C. E.,
Walton, J., Abraham, N. L., Andrews, M. B., Andrews, T., Archibald, A. T.,
Berthou, S., Burke, E., Blockley, E., Carslaw, K., Dalvi, M., Edwards, J.,
Folberth, G. A., Gedney, N., Griffiths, P. T., Harper, A. B., Hendry, M. A.,
Hewitt, A. J., Johnson, B., Jones, A., Jones, C. D., Keeble, J., Liddicoat,
S., Morgenstern, O., Parker, R. J., Predoi, V., Robertson, E., Siahaan, A.,
Smith, R. S., Swaminathan, R., Woodhouse, M. T., Zeng, G., and Zerroukat, M.:
UKESM1: Description and Evaluation of the U.K. Earth System Model, J.
Adv. Model. Earth Sy., 11, 4513–4558,
<a href="https://doi.org/10.1029/2019MS001739" target="_blank">https://doi.org/10.1029/2019MS001739</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib27"><label>Servén et al.(2018)Servén, Brummitt, Abedi, and
Hlink</label><mixed-citation>
      
Servén, D., Brummitt, C., Abedi, H., and Hlink: dswah/pyGAM: v0.8.0,
Zenodo, <a href="https://doi.org/10.5281/zenodo.1208723" target="_blank">https://doi.org/10.5281/zenodo.1208723</a>,  2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib28"><label>Servera et al.(2023)Servera, Martino, Verrelst, and
Camps-Valls</label><mixed-citation>
      
Servera, J. V., Martino, L., Verrelst, J., and Camps-Valls, G.: Multifidelity
Gaussian process emulation for atmospheric radiative transfer models, IEEE
T. Geosci. Remote Sens., 61, 1–10, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib29"><label>Sexton et al.(2012)Sexton, Murphy, Collins, and
Webb</label><mixed-citation>
      
Sexton, D. M., Murphy, J. M., Collins, M., and Webb, M. J.: Multivariate
probabilistic projections using imperfect climate models part I: outline of
methodology, Clim. Dynam., 38, 2513–2542, 2012.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib30"><label>Sobol'(2001)</label><mixed-citation>
      
Sobol', I. M.: Global sensitivity indices for nonlinear mathematical models and
their Monte Carlo estimates, Math. Comput. Simulat., 55,
271–280, 2001.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib31"><label>Sottile and Gohel(2022)</label><mixed-citation>
      
Sottile, G. and Gohel, D.: trapezoid: Trapezoidal Distribution, r package
version 0.6,
<a href="https://CRAN.R-project.org/package=trapezoid" target="_blank"/> (last access: 18 August 2026), 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib32"><label>Stouffer et al.(2017)Stouffer, Eyring, Meehl, Bony, Senior, Stevens,
and Taylor</label><mixed-citation>
      
Stouffer, R. J., Eyring, V., Meehl, G. A., Bony, S., Senior, C., Stevens, B.,
and Taylor, K.: CMIP5 scientific gaps and recommendations for CMIP6, B. Am. Meteorol. Soc., 98, 95–105, 2017.


    </mixed-citation></ref-html>
<ref-html id="bib1.bib33"><label>Susiluoto et al.(2020)Susiluoto, Spantini, Haario, Härkönen,
and Marzouk</label><mixed-citation>
      
Susiluoto, J., Spantini, A., Haario, H., Härkönen, T., and Marzouk, Y.: Efficient multi-scale Gaussian process regression for massive remote sensing data with satGP v0.1.2, Geosci. Model Dev., 13, 3439–3463, <a href="https://doi.org/10.5194/gmd-13-3439-2020" target="_blank">https://doi.org/10.5194/gmd-13-3439-2020</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib34"><label>Watson-Parris et al.(2021)Watson-Parris, Williams, Deaconu, and
Stier</label><mixed-citation>
      
Watson-Parris, D., Williams, A., Deaconu, L., and Stier, P.: Model calibration using ESEm v1.1.0 – an open, scalable Earth system emulator, Geosci. Model Dev., 14, 7659–7672, <a href="https://doi.org/10.5194/gmd-14-7659-2021" target="_blank">https://doi.org/10.5194/gmd-14-7659-2021</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib35"><label>Wickham and Bryan(2023)</label><mixed-citation>
      
Wickham, H. and Bryan, J.: readr: Read Rectangular Text Data, r package version
2.1.4,
<a href="https://CRAN.R-project.org/package=readr" target="_blank"/> (last access: 18 August 2026), 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib36"><label>Wood(2020)</label><mixed-citation>
      
Wood, S. N.: Inference and computation with generalized additive models and
their extensions, Test, 29, 307–339, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib37"><label>Wood et al.(2015)Wood, Goude, and Shaw</label><mixed-citation>
      
Wood, S. N., Goude, Y., and Shaw, S.: Generalized additive models for large
data sets, J. Roy. Stat. Soc. C, 64, 139–155, 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib38"><label>Yang et al.(2025)Yang, Elsaesser, van Lier-Walqui, and
Eidhammer</label><mixed-citation>
      
Yang, Q., Elsaesser, G. S., van Lier-Walqui, M., and Eidhammer, T.: A simple
emulator that enables interpretation of parameter-output relationships,
applied to two climate model PPEs, J. Adv. Model. Earth
Sy., 17, e2024MS004766, <a href="https://doi.org/10.1029/2024MS004766" target="_blank">https://doi.org/10.1029/2024MS004766</a>, 2025.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib39"><label>Yoshioka et al.(2019)Yoshioka, Regayre, Pringle, Johnson, Mann,
Partridge, Sexton, Lister, Schutgens, Stier, Kipling, Bellouin, Browse,
Booth, Johnson, Johnson, Mollard, Lee, and Carslaw</label><mixed-citation>
      
Yoshioka, M., Regayre, L. A., Pringle, K. J., Johnson, J. S., Mann, G. W.,
Partridge, D. G., Sexton, D. M. H., Lister, G. M. S., Schutgens, N., Stier,
P., Kipling, Z., Bellouin, N., Browse, J., Booth, B. B. B., Johnson, C. E.,
Johnson, B., Mollard, J. D. P., Lee, L., and Carslaw, K. S.: Ensembles of
Global Climate Model Variants Designed for the Quantification and Constraint
of Uncertainty in Aerosols and Their Radiative Forcing, J. Adv.
Model. Earth Sy., 11, 3728–3754,
<a href="https://doi.org/10.1029/2019MS001628" target="_blank">https://doi.org/10.1029/2019MS001628</a>, 2019.

    </mixed-citation></ref-html>--></article>
