<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing with OASIS Tables v3.0 20080202//EN" "journalpub-oasis3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:oasis="http://docs.oasis-open.org/ns/oasis-exchange/table" xml:lang="en" dtd-version="3.0">
  <front>
    <journal-meta><journal-id journal-id-type="publisher">GMD</journal-id><journal-title-group>
    <journal-title>Geoscientific Model Development</journal-title>
    <abbrev-journal-title abbrev-type="publisher">GMD</abbrev-journal-title><abbrev-journal-title abbrev-type="nlm-ta">Geosci. Model Dev.</abbrev-journal-title>
  </journal-title-group><issn pub-type="epub">1991-9603</issn><publisher>
    <publisher-name>Copernicus Publications</publisher-name>
    <publisher-loc>Göttingen, Germany</publisher-loc>
  </publisher></journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.5194/gmd-11-697-2018</article-id><title-group><article-title>Nine time steps: ultra-fast statistical consistency testing
of the Community Earth System Model (pyCECT v3.0)</article-title><alt-title>UF-CAM-ECT</alt-title>
      </title-group><?xmltex \runningtitle{UF-CAM-ECT}?><?xmltex \runningauthor{D.~J.~Milroy et al.}?>
      <contrib-group>
        <contrib contrib-type="author" corresp="yes" rid="aff1">
          <name><surname>Milroy</surname><given-names>Daniel J.</given-names></name>
          <email>daniel.milroy@colorado.edu</email>
        <ext-link>https://orcid.org/0000-0001-6500-3227</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff2">
          <name><surname>Baker</surname><given-names>Allison H.</given-names></name>
          
        <ext-link>https://orcid.org/0000-0003-2436-7838</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff2">
          <name><surname>Hammerling</surname><given-names>Dorit M.</given-names></name>
          
        <ext-link>https://orcid.org/0000-0003-3583-3611</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1">
          <name><surname>Jessup</surname><given-names>Elizabeth R.</given-names></name>
          
        <ext-link>https://orcid.org/0000-0002-7740-9985</ext-link></contrib>
        <aff id="aff1"><label>1</label><institution>University of Colorado,
Boulder, CO, USA</institution>
        </aff>
        <aff id="aff2"><label>2</label><institution>The National Center for
Atmospheric Research, Boulder, CO, USA</institution>
        </aff>
      </contrib-group>
      <author-notes><corresp id="corr1">Daniel J. Milroy (daniel.milroy@colorado.edu)</corresp></author-notes><pub-date><day>26</day><month>February</month><year>2018</year></pub-date>
      
      <volume>11</volume>
      <issue>2</issue>
      <fpage>697</fpage><lpage>711</lpage>
      <history>
        <date date-type="received"><day>28</day><month>February</month><year>2017</year></date>
           <date date-type="rev-request"><day>27</day><month>April</month><year>2017</year></date>
           <date date-type="rev-recd"><day>21</day><month>December</month><year>2017</year></date>
           <date date-type="accepted"><day>5</day><month>January</month><year>2018</year></date>
      </history>
      <permissions>
        <copyright-statement>Copyright: © 2018 Daniel J. Milroy et al.</copyright-statement>
        <copyright-year>2018</copyright-year>
      <license license-type="open-access"><license-p>This work is licensed under the Creative Commons Attribution 3.0 Unported License. To view a copy of this licence, visit <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/3.0/">https://creativecommons.org/licenses/by/3.0/</ext-link></license-p></license></permissions><self-uri xlink:href="https://gmd.copernicus.org/articles/11/697/2018/gmd-11-697-2018.html">This article is available from https://gmd.copernicus.org/articles/11/697/2018/gmd-11-697-2018.html</self-uri><self-uri xlink:href="https://gmd.copernicus.org/articles/11/697/2018/gmd-11-697-2018.pdf">The full text article is available as a PDF file from https://gmd.copernicus.org/articles/11/697/2018/gmd-11-697-2018.pdf</self-uri>
      <abstract><title>Abstract</title>
    <p id="d1e113">The Community Earth System Model Ensemble Consistency Test (CESM-ECT) suite
was developed as an alternative to requiring bitwise identical output for
quality assurance. This objective test provides a statistical measurement of
consistency between an accepted ensemble created by small initial temperature
perturbations and a test set of CESM simulations. In this work, we extend the
CESM-ECT suite with an inexpensive and robust test for ensemble consistency
that is applied to Community Atmospheric Model (CAM) output after only nine
model time steps. We demonstrate that adequate ensemble variability is
achieved with instantaneous variable values at the ninth step, despite rapid
perturbation growth and heterogeneous variable spread. We refer to this new
test as the Ultra-Fast CAM Ensemble Consistency Test (UF-CAM-ECT) and
demonstrate its effectiveness in practice, including its ability to detect
small-scale events and its applicability to the Community Land Model (CLM).
The new ultra-fast test facilitates CESM development, porting, and
optimization efforts, particularly when used to complement information from
the original CESM-ECT suite of tools.</p>
  </abstract>
    </article-meta>
  </front>
<body>
      

<sec id="Ch1.S1" sec-type="intro">
  <label>1</label><title>Introduction</title>
      <p id="d1e125">Requiring bit-for-bit (BFB) identical output for quality assurance of climate
codes is restrictive. The codes are complex and constantly evolving,
necessitating an objective method for assuring quality without BFB
equivalence. Once the BFB requirement is relaxed, evaluating the possible
ways data sets can be distinct while still representing similar states is
nontrivial. <xref ref-type="bibr" rid="bib1.bibx2" id="text.1"/> address this challenge by considering
statistical distinguishability from an ensemble for the Community Earth
System Model (CESM; <xref ref-type="bibr" rid="bib1.bibx6" id="altparen.2"/>), an open-source Earth system model (ESM)
developed principally at the National Center for Atmospheric Research (NCAR).
<xref ref-type="bibr" rid="bib1.bibx2" id="text.3"/> developed the CESM ensemble consistency test (CESM-ECT) to
address the need for a simple method of determining whether non-BFB CESM
outputs are statistically consistent with the expected output. Substituting
statistical indistinguishability for BFB equivalence allows for more
aggressive code optimizations, implementation of more efficient algorithms,
and execution on heterogeneous computational environments.</p>
      <p id="d1e137">CESM-ECT is a suite of tools which measures statistical consistency by
focusing on 12-month output from two different component models within CESM:
the Community Atmospheric Model (CAM) and the Parallel Ocean Program (POP),
with ensemble consistency testing tools referred to as CAM-ECT and POP-ECT,
respectively. The key idea of CESM-ECT is to compare new non-BFB CESM outputs
(e.g., from a recently built machine or modified code) to an ensemble of
simulation outputs from an “accepted” configuration (e.g., a trusted
machine and software and hardware configuration), quantifying their
differences by an objective statistical metric. CESM-ECT returns a pass for
the new output if it is statistically indistinguishable from the distribution
of the ensemble, and a fail if the results are distinct. The selection of an
“accepted” ensemble is integral to CESM-ECT's pass or fail determination
for test simulations. The question of the ensemble composition and size for
CAM-ECT is addressed in <xref ref-type="bibr" rid="bib1.bibx9" id="text.4"/>, which concludes that ensembles
created by aggregating sources of variability from different compilers
improve the classification power<?pagebreak page698?> and accuracy of the test. At this time,
CESM-ECT is used by CESM software engineers and scientists for both port
verification and quality assurance for code modification and updates. In
light of the success of CAM-ECT, the question arose as to whether the test
could also be performed using a time period shorter than 1 year, and in
particular, just a small number of time steps.</p>
      <p id="d1e143">The effects of rounding, truncation and initial condition perturbation on
chaotic dynamical systems is a well-studied area of research with foundations
in climate science. The growth of initial condition perturbations on CAM has
been investigated since <xref ref-type="bibr" rid="bib1.bibx11" id="text.5"/>, whose work resulted in the
PerGro test. This test examined the rate of divergence of CAM variables at
initial time steps between simulations with different initial conditions. The
rates were used to compare the behavior of CAM under modification to that of
an established version of the model in the context of the growth of machine
roundoff error. With the advent of CAM5, PerGro became less useful for
classifying model behavior, as the new parameterizations in the model
resulted in much more rapid spread. Accordingly, it was commonly held that
using a small number of time steps was an untenable strategy due to the
belief that the model's initial variability (that far exceeded machine
roundoff) was too great to measure statistical difference effectively.
Indeed, even the prospect of using runs of 1 simulation year for CESM-ECT
was met with initial skepticism. Note that prior to the CESM-ECT approach,
CESM verification was a subjective process predicated on climate scientists'
expertise in analyzing multi-century simulation output. The success of
CESM-ECT's technique of using properties of yearly CAM means
<xref ref-type="bibr" rid="bib1.bibx2" id="paren.6"/> translated to significant cost savings for verifying the
model.</p>
      <p id="d1e152">Motivated by the success and cost improvement of CESM-ECT, we were curious as
to whether its general technique could be applied after a few initial time
steps, in analogy with <xref ref-type="bibr" rid="bib1.bibx11" id="text.7"/>. This strategy would represent
potential further cost savings by reducing the length of the ensemble and
test simulations. We were not dissuaded by the fact that the rapid growth of
roundoff order perturbations in CAM5 negatively impacted PerGro's ability to
detect changes due to its comparison with machine epsilon. In fact, we show
that examination of ensemble variability after several time steps permits
accurate pass and fail determinations and complements CAM-ECT in terms of
identifying potential problems. In this paper we present an ensemble-based
consistency test that evaluates statistical distinguishability at nine time
steps, hereafter designated the Ultra-Fast CAM Ensemble Consistency Test
(UF-CAM-ECT).</p>
      <p id="d1e159">A notable difference between CAM-ECT and UF-CAM-ECT is the type of data
considered. CAM-ECT spatially averages the yearly mean output to make the
ensemble more robust (effectively a double average). Therefore, a limitation
of CAM-ECT is that if a bug only produces a small-scale effect, then the
overall climate may not be altered in an average sense at 12 months, and the
change may go undetected. In this case a longer simulation time may be needed
for the bug to impact the average climate. An example of this issue is the
modification of the dynamics hyperviscosity parameter (NU) in
<xref ref-type="bibr" rid="bib1.bibx2" id="text.8"/>, which was not detected by CAM-ECT. In contrast, UF-CAM-ECT
takes the spatial means of instantaneous values very early in the model run,
which can facilitate the detection of smaller-scale modifications. In terms
of simulation length for UF-CAM-ECT, we were aware that we would need to
satisfy two constraints in choosing an adequate number of initial time steps:
some variables can suffer excessive spread while others remain relatively
constant, complicating pass/fail determinations. Balancing the run time and
ensemble variability (hence test classification power) also alters the types
of statistical differences the test can detect; exploring the complementarity
between CAM-ECT and UF-CAM-ECT is a focus of our work.</p>
      <p id="d1e165">In particular, we make four contributions: we demonstrate that adequate
ensemble variability can be achieved at the ninth CESM time step in spite of
the heterogeneous spread among the variables considered; we evaluate
UF-CAM-ECT with experiments from <xref ref-type="bibr" rid="bib1.bibx2" id="text.9"/>, code modifications from
<xref ref-type="bibr" rid="bib1.bibx9" id="text.10"/>, and several new CAM tests; we propose an effective
ensemble size; and we demonstrate that changes to the Community Land Model
(CLM) can be detected by both UF-CAM-ECT and CAM-ECT.</p>
      <p id="d1e174">In Sect. <xref ref-type="sec" rid="Ch1.S2"/>, we provide background for CESM-ECT,
describe our experimental design and applications, and quantify CESM
divergence by time step. In Sect. <xref ref-type="sec" rid="Ch1.S3"/>, we detail the
UF-CAM-ECT method. We demonstrate the results of our investigation into the
appropriate ensemble size in Sect. <xref ref-type="sec" rid="Ch1.S4"/>. We present
experimental results in Sect. <xref ref-type="sec" rid="Ch1.S5"/>, provide guidance for
the tools' usage in Sect. <xref ref-type="sec" rid="Ch1.S6"/>, and conclude with
Sect. <xref ref-type="sec" rid="Ch1.S7"/>.</p>
</sec>
<sec id="Ch1.S2">
  <label>2</label><title>Background and motivation</title>
<sec id="Ch1.S2.SS1">
  <label>2.1</label><title>CESM and ECT</title>
      <p id="d1e205">The CESM consists of several geophysical models (e.g., atmosphere, ocean,
land); the original CESM-ECT study in <xref ref-type="bibr" rid="bib1.bibx2" id="text.11"/> focuses on data
from the Community Atmosphere Model (CAM). Currently, CESM-ECT consists of
tests for CAM (CAM-ECT and UF-CAM-ECT) and POP (POP-ECT).</p>
      <?pagebreak page699?><p id="d1e211">For CAM-ECT, CAM output data containing 12-month averages at each grid point
for the atmosphere variables are written in time slices to netCDF history
files. The CAM-ECT ensemble consists of CAM output data from simulations (151
by default or up to 453 simulations in <xref ref-type="bibr" rid="bib1.bibx9" id="altparen.12"/>) of 12-month runs
generated on an “accepted” machine with a trusted version of the CESM and
software stack. In <xref ref-type="bibr" rid="bib1.bibx2" id="text.13"/>, the CAM-ECT ensemble is a collection of
CESM simulations with identical parameters but different perturbations to the
initial atmospheric temperature field. In <xref ref-type="bibr" rid="bib1.bibx9" id="text.14"/>, the size and
composition of the ensemble are further explored, resulting in a refined
ensemble comprised of equal numbers of runs from different compilers (Intel,
GNU and PGI), and a recommended size of 300 or 453.</p>
      <p id="d1e223">Through sensitive dependence on initial conditions, unique initial
temperature perturbations guarantee unique trajectories through the state
space. The unique trajectories give rise to an ensemble of output variables
which represents the model's natural variability. Both <xref ref-type="bibr" rid="bib1.bibx2" id="text.15"/> and
<xref ref-type="bibr" rid="bib1.bibx9" id="text.16"/> use <inline-formula><mml:math id="M1" display="inline"><mml:mrow><mml:mi mathvariant="script">O</mml:mi><mml:mo>(</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">14</mml:mn></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> perturbations to the initial
atmospheric temperature field to create the ensemble at 12 months
<xref ref-type="bibr" rid="bib1.bibx2" id="paren.17"/>.</p>
      <p id="d1e255">To compare the statistical characteristics of new runs, CAM-ECT begins by
quantifying the variability of the accepted ensemble. Principal component
analysis (PCA) is used to transform the original space of CAM variables into
a subspace of linear combinations of the standardized variables, which are
then uncorrelated. To begin, a set of new runs (three by default) is given to
the Python CESM Ensemble Consistency Tool (pyCECT), which returns a pass or
fail depending on the number of PC scores that fall outside a specified
confidence interval (typically 95 %; <xref ref-type="bibr" rid="bib1.bibx2" id="altparen.18"/>). The area weighted
global means of the test runs are projected into the PC space of the
ensemble, and the tool determines if the scores of the new runs are within
two standard deviations of the distribution of the scores of the accepted
ensemble, marking any PC score outside two standard deviations as a failure
<xref ref-type="bibr" rid="bib1.bibx2" id="paren.19"/>. For example, let <inline-formula><mml:math id="M2" display="inline"><mml:mrow><mml:mi>P</mml:mi><mml:mo>=</mml:mo><mml:mo mathvariant="italic">{</mml:mo><mml:mi>A</mml:mi><mml:mo>,</mml:mo><mml:mi>B</mml:mi><mml:mo>,</mml:mo><mml:mi>C</mml:mi><mml:mo mathvariant="italic">}</mml:mo></mml:mrow></mml:math></inline-formula> be the set of sets of
PCs marked as failures. To make a pass or fail determination CAM-ECT operates
as follows: <inline-formula><mml:math id="M3" display="inline"><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>A</mml:mi><mml:mi>B</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>A</mml:mi><mml:mo>∩</mml:mo><mml:mi>B</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>A</mml:mi><mml:mi>C</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>A</mml:mi><mml:mo>∩</mml:mo><mml:mi>C</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>B</mml:mi><mml:mi>C</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>B</mml:mi><mml:mo>∩</mml:mo><mml:mi>C</mml:mi><mml:mo>;</mml:mo><mml:mi>S</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>A</mml:mi><mml:mi>B</mml:mi></mml:mrow></mml:msub><mml:mo>∪</mml:mo><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>A</mml:mi><mml:mi>C</mml:mi></mml:mrow></mml:msub><mml:mo>∪</mml:mo><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>B</mml:mi><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> and returns a failure if <inline-formula><mml:math id="M4" display="inline"><mml:mrow><mml:mo>|</mml:mo><mml:mi>S</mml:mi><mml:mo>|</mml:mo><mml:mo>≥</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:math></inline-formula>
<xref ref-type="bibr" rid="bib1.bibx9" id="paren.20"/>. The default parameters specifying the pass/fail criteria
yield a false positive rate of 0.5 %. Note that the scope of our method is
flexible: we can simply add additional output variables (including diagnostic
variables) to the test as long as the number of ensemble members is greater
than the number of variables, and recompute the PCA. For a comprehensive
explanation of the test method, see <xref ref-type="bibr" rid="bib1.bibx9" id="text.21"/> and <xref ref-type="bibr" rid="bib1.bibx2" id="text.22"/>.</p>
      <p id="d1e409">More recently, another module of CESM-ECT was developed to determine
statistical consistency for the Parallel  Ocean Program (POP) model of
CESM and designated POP-ECT <xref ref-type="bibr" rid="bib1.bibx3" id="paren.23"/>.  While similar to CAM-ECT
in that an ensemble of trusted simulations is used to evaluate new
runs, the statistical consistency test itself is not the same.  In POP data,
the spatial variation and temporal scales are much different than in
CAM data, and there are many fewer variables. Hence the test does not
involve PCA or spatial averages,
instead making comparisons at each grid location. Note that in this work we
demonstrate that CAM-ECT and UF-CAM-ECT testing can be applied directly and
successfully to the CLM component of CESM,
because of the tight coupling between CAM  and CLM, indicating that a
separate ECT module for CLM is likely unnecessary.</p>
</sec>
<sec id="Ch1.S2.SS2">
  <label>2.2</label><title>Motivation: CAM divergence in initial time steps</title>
      <p id="d1e423">Applying an ensemble consistency test at nine time steps is sensible only if
there is an adequate amount of ensemble variability to correctly evaluate new
runs as has been shown for the 1-year runs. This issue is key: with too
much spread a bug cannot be detected, and without enough spread the test can
be too restrictive in its pass and fail determinations. Many studies consider
the effects of initial condition perturbations to ensemble members on the
predictability of an ESM, and the references in <xref ref-type="bibr" rid="bib1.bibx7" id="text.24"/> contain
several examples. In particular, <xref ref-type="bibr" rid="bib1.bibx5" id="text.25"/> study uncertainty arising
from climate model internal variability using an ensemble method, and an
earlier work <xref ref-type="bibr" rid="bib1.bibx4" id="paren.26"/> considers the predictability and forecast
range of a climate model by examining the separate effects and timescales of
initial conditions and forcings. <xref ref-type="bibr" rid="bib1.bibx4" id="text.27"/> also study ensemble
global means and spread, as well as undertaking an entropy analysis of
leading empirical orthogonal functions (comparable to PCA). These studies are
primarily concerned with model variability and predictability at the
timescale of several years or more. However, we note that concurrent to our
work, a new method that considers 1 s time steps has been developed in
<xref ref-type="bibr" rid="bib1.bibx12" id="text.28"/>. Their focus is on comparing the numerical error in time
integration between a new run and control runs.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F1" specific-use="star"><?xmltex \currentcnt{1}?><label>Figure 1</label><caption><p id="d1e443">Representation of effects of initial CAM temperature perturbation
over 11 time steps (including <inline-formula><mml:math id="M5" display="inline"><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula>). CAM variables are listed on the
vertical axis, and the horizontal axis records the simulation time step. The
color bar designates equality of the corresponding variables between the
unperturbed and perturbed simulations' area weighted global means after being
rounded to <inline-formula><mml:math id="M6" display="inline"><mml:mi>n</mml:mi></mml:math></inline-formula> significant digits (<inline-formula><mml:math id="M7" display="inline"><mml:mi>n</mml:mi></mml:math></inline-formula> is the color) at each time step. Time
steps where the corresponding variable was not computed (subcycled variables)
are colored black. White indicates equality of greater than nine significant
digits (i.e., 10–17). Red variable names are not used by UF-CAM-ECT.</p></caption>
          <?xmltex \igopts{width=469.470472pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/11/697/2018/gmd-11-697-2018-f01.pdf"/>

        </fig>

      <p id="d1e478">As mentioned previously, we were curious about the behavior of CESM in its
initial time steps in terms of whether we would be able to determine
statistical distinguishability. Figure <xref ref-type="fig" rid="Ch1.F1"/> represents our
initial inquiry into this behavior. To generate the data, we ran two
simulations of 11 time steps each: one with no initial condition perturbation
and one with a perturbation of <inline-formula><mml:math id="M8" display="inline"><mml:mrow><mml:mi mathvariant="script">O</mml:mi><mml:mo>(</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">14</mml:mn></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> to the initial
atmospheric temperature. The vertical axis labels designate CAM variables,
while the horizontal axis specifies the CESM time step. The color of each
step represents the number of significant figures in common between the
perturbed and unperturbed simulations' area weighted global means: a small
number of figures in common (darker red) indicates a large difference. Black
tiles specify time steps where the variable's value is not computed due to
model sub-cycling <xref ref-type="bibr" rid="bib1.bibx6" id="paren.29"/>. White tiles indicate between 10 and 17
significant figures in common (i.e., a small magnitude of difference). Most
CAM variables exhibit a difference from the unperturbed simulation at the
initial time step (0), and nearly all have diverged to only a few figures in
common by step 10. Figure <xref ref-type="fig" rid="Ch1.F1"/> demonstrates sensitive dependence
on initial conditions in CAM and suggests that choosing a small number of
time steps may provide sufficient variability to determine statistical
distinguishability resulting from significant changes. We further examine the
ninth time step as it is the last step on the plot where sub-cycled variables
are<?pagebreak page700?> calculated. Of the total 134 CAM variables output by default in our
version of CESM (see Sect. <xref ref-type="sec" rid="Ch1.S3"/>), 117 are utilized by CAM-ECT,
as 17 are either redundant or have zero variance. In all following analyses,
the first nine sub-cycled variables distinguished by red labels (AODDUST1,
AODDUST3, AODVIS, BURDENBC, BURDENDUST, BURDENPOM, BURDENSEASALT, BURDENSO4,
and BURDENSOA) are discarded, as they take constant values through time step
45. Thus we use 108 variables from Fig. <xref ref-type="fig" rid="Ch1.F1"/> in the UF-CAM-ECT
ensemble.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F2" specific-use="star"><?xmltex \currentcnt{2}?><label>Figure 2</label><caption><p id="d1e516">The vertical axis labels the difference between the unperturbed
simulation and the perturbed simulations' area weighted global means, divided
by the unperturbed simulation's area
weighted global mean for the indicated variable at each time step. The
horizontal axis is the CESM time step with intervals chosen as multiples of
nine. The left column <bold>(a)</bold> plots are time series representations of
the values of three CAM variables chosen as representatives of the entire set
(variables behave similarly to one of these three). The right column
<bold>(b)</bold> plots are statistical representations of the 30 values plotted
at each vertical grid line of the corresponding left column. More directly,
each box plot depicts the distribution of values of each variable for each
time step from 9 to 45 in multiples of 9.</p></caption>
          <?xmltex \igopts{width=441.017717pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/11/697/2018/gmd-11-697-2018-f02.pdf"/>

        </fig>

      <p id="d1e531">Next we examine the time series of each CAM variable by looking at the first
45 CESM time steps (<inline-formula><mml:math id="M9" display="inline"><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> through <inline-formula><mml:math id="M10" display="inline"><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mn mathvariant="normal">45</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>) for 30 simulations to select a
time step when all variables experience sufficient divergence from the values
of the reference unperturbed simulation. To illustrate ensemble variability
at initial time steps, Fig. <xref ref-type="fig" rid="Ch1.F2"/> depicts the time evolution
of three representative CAM variables from <inline-formula><mml:math id="M11" display="inline"><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> to <inline-formula><mml:math id="M12" display="inline"><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mn mathvariant="normal">45</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>. Most CAM
variables' behavior is analogous to one of the rows in this figure. The data
set was generated by running 31 simulations: one with no initial atmospheric
temperature perturbation, and 30 with different <inline-formula><mml:math id="M13" display="inline"><mml:mrow><mml:mi mathvariant="script">O</mml:mi><mml:mo>(</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">14</mml:mn></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> K
perturbations. The vertical axis labels the difference between the
unperturbed simulation and the perturbed simulations' area weighted global
means, divided by the unperturbed simulation's area
weighted global mean value for the indicated variable at each time step. The
right column visualizes the distributions of the data in<?pagebreak page701?> the left column.
Each box plot represents the values in the left column at nine time step
intervals from 9 to 45 (inclusive). In most cases, the variables attain
measurable but well-contained spread in approximately the first nine time
steps. From the standpoint of CAM-ECT, Fig. <xref ref-type="fig" rid="Ch1.F2"/> suggests
that an ensemble created at the ninth time step will likely contain
sufficient variability to categorize experimental sets correctly. Using
additional time steps is unlikely to be beneficial in terms of UF-CAM-ECT
sensitivity or classification accuracy, and choosing a smaller number of time
steps is advantageous from the standpoint of capturing the state of test
cases before feedback mechanisms take place (e.g.,
Sect. <xref ref-type="sec" rid="Ch1.S5.SS3.SSS3"/>). Note that we do not claim that 9 time steps
is optimal in terms of computational cost, but the difference in run time
between 9 and 45 time steps is negligible in comparison to the cost of
CAM-ECT 12-month simulations (and the majority of time for such short runs is
initialization and I/O). We further discuss ensemble generation and size in
Sect. <xref ref-type="sec" rid="Ch1.S4"/> with an investigation of the properties of
ensembles created from the ninth time step and compare their pass/fail
determinations of experimental simulations with that of CAM-ECT in
Sect. <xref ref-type="sec" rid="Ch1.S5"/>.</p>
</sec>
</sec>
<sec id="Ch1.S3">
  <label>3</label><title>UF-CAM-ECT approach</title>
      <p id="d1e618">UF-CAM-ECT employs the same essential test method as CAM-ECT described in
<xref ref-type="bibr" rid="bib1.bibx2" id="text.30"/>, but with a CESM simulation length of nine time steps
(which is approximately 5 simulation hours) using the default CAM time step
of 1800 s (30 min). By considering a specific time step, we are using
instantaneous values in contrast to CAM-ECT, which uses yearly average
values. UF-CAM-ECT inputs are spatially averaged, so averaged once, whereas
CAM-ECT inputs are averaged across the 12 simulation months and spatially
averaged, so averaged twice. As a consequence of using instantaneous values,
UF-CAM-ECT is more sensitive to localized phenomena (see
Sect. <xref ref-type="sec" rid="Ch1.S5.SS3.SSS3"/>). By virtue of the small number of
modifications required to transform<?pagebreak page702?> CAM-ECT into UF-CAM-ECT, we consider the
ECT framework to have surprisingly broad applicability. Substituting
instantaneous values for yearly averages permits the discernment of different
features and modifications – see Sect. <xref ref-type="sec" rid="Ch1.S5"/> and
<xref ref-type="sec" rid="Ch1.S5.SS2"/> for evidence of this assertion.</p>
      <p id="d1e630">As in <xref ref-type="bibr" rid="bib1.bibx2" id="text.31"/> and <xref ref-type="bibr" rid="bib1.bibx9" id="text.32"/>, we run CESM simulations on a
1<inline-formula><mml:math id="M14" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> global grid using the CAM5 model version described in
<xref ref-type="bibr" rid="bib1.bibx7" id="text.33"/>, and despite the rapid growth in perturbations in CAM5 with
the default time step of 1800 s, we can still characterize its variability.
We run simulations with 900 MPI processes and two OpenMP threads per MPI
process (unless otherwise noted) on the Yellowstone machine at NCAR.
Yellowstone is composed of 4536 compute nodes, with two Xeon Sandy Bridge
CPUs and 32 GB memory per node. The default compiler on Yellowstone for our
CESM version is Intel 13.1.2 with -O2 optimization. We also use the
CESM-supported compilers GNU 4.8.0 and PGI 13.0 in this study. With 900 MPI
processes and two OpenMP threads per process, a simulation of nine time steps
on Yellowstone is a factor of approximately 70 cheaper in terms of CPU time
than a 12-month CESM simulation.</p>
      <p id="d1e651">Either single- or double-precision output is suitable for UF-CAM-ECT. While
CAM can be instructed to write its history files in single- or double-precision floating-point form, its default is single precision which was used
for CAM-ECT in <xref ref-type="bibr" rid="bib1.bibx2" id="text.34"/> and <xref ref-type="bibr" rid="bib1.bibx9" id="text.35"/>. Similarly,
UF-CAM-ECT takes single-precision output by default. However, we chose to
generate double-precision output to facilitate the study represented by
Fig. <xref ref-type="fig" rid="Ch1.F1"/>; it would have been impossible to perform a
significance test of up to 17 digits otherwise. In the case of new runs
written in double precision, both CAM-ECT and UF-CAM-ECT compare ensemble
values promoted to double precision with the unmodified new outputs. We
determined that the effects of using double- or single-precision outputs for
ensemble generation and the evaluation of new runs did not impact statistical
distinguishability.</p>
</sec>
<sec id="Ch1.S4">
  <label>4</label><title>UF-CAM-ECT ensemble size</title>
      <p id="d1e670">In this section we consider the properties of the UF-CAM-ECT ensemble,
particularly focusing on  ensemble size.   Given the use of
instantaneous values at nine time steps in UF-CAM-ECT, our expectation
was that the size of the ensemble would differ from that of CAM-ECT.
We considered it plausible that a larger number would be required to
make proper pass and fail determinations.  The ensemble should contain
enough variability that UF-CAM-ECT classifies experiments expected to
be statistically indistinguishable as consistent with the ensemble.
Furthermore, for experiments that significantly alter the climate,
UF-CAM-ECT should classify them as statistically distinct from the
ensemble.  Accordingly, the ensemble itself is key, and examining its
size allows us to quantify the variability it contains as the number
of ensemble members increases.</p>
      <p id="d1e673">Our sets of experimental simulations (new runs) typically consist of 30
members, but by default pyCECT was written to do a single test on three runs.
Performing the full set of possible CAM-ECT tests from a given set of
experimental simulations allows us to make robust failure determinations as
opposed to a single binary pass/fail test. In this work we utilize the
Ensemble Exhaustive Test (EET) tool described in <xref ref-type="bibr" rid="bib1.bibx9" id="text.36"/> to
calculate an overall failure rate for sets of new runs larger than the pyCECT
default. A failure rate provides more detail on the statistical difference
between the ensemble and experimental set. To calculate the failure rate, EET
efficiently performs all possible tests which are equal in number to the ways
<inline-formula><mml:math id="M15" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">test</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> simulations can be chosen from all <inline-formula><mml:math id="M16" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">tot</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>
simulations (i.e., the binomial coefficient:
<inline-formula><mml:math id="M17" display="inline"><mml:mfenced open="(" close=")"><mml:mfrac linethickness="0"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">tot</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">test</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mfenced></mml:math></inline-formula>). For this work most experiments
consist of 30 simulations, so <inline-formula><mml:math id="M18" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">tot</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">30</mml:mn></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M19" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">test</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:math></inline-formula>
yields 4060 possible combinations. With this tool we can make a comparison
between the exhaustive test failure rate and the single-test CAM-ECT false
positive rate calibrated to be 0.5 %.</p>
      <p id="d1e751">To determine a desirable UF-CAM-ECT ensemble size, we gauge whether ensembles
of varying sizes contain sufficient variability by excluding sets of ensemble
simulations and performing exhaustive testing against ensembles formed from
the remaining elements. Since the test sets and ensemble members are
generated by the same type of initial condition perturbation, the
test sets should pass. We begin with
a set of 801 CESM simulations of nine time steps, differing by unique
perturbations to the initial atmospheric temperature field in
{<inline-formula><mml:math id="M20" display="inline"><mml:mrow><mml:mfenced close=")" open="["><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">9.99</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">14</mml:mn></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:mfenced></mml:mrow></mml:math></inline-formula>,
<inline-formula><mml:math id="M21" display="inline"><mml:mrow><mml:mfenced open="(" close="]"><mml:mrow><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">9.99</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">14</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mfenced></mml:mrow></mml:math></inline-formula>} K. The motivation for generating a
large number of outputs was our expectation that ensembles created from
instantaneous values would contain less variability. Moreover, since the runs
are comparatively cheap, it was easy to run many simulations for testing
purposes. In the following description, all draws are made without
replacement. We first randomly select a subset from the 801 simulations and
compute the PC loadings. From the remaining simulations, we choose 30 at
random and run EET against this experimental set. For each ensemble size, we
make 100 random draws to form an ensemble, and for each ensemble we make 100
random draws of experimental sets. This results in 10 000 EET results per
ensemble size. For example, to test the variability of the size 350 ensemble,
we choose 350 simulations at random from our set of 801 to form an ensemble.
From the remaining 451 simulations, we randomly choose 30 and exhaustively
test them against the generated ensemble with EET (4060 individual tests).
This is repeated 99 times for the ensemble. Then 99 more ensembles are
created in the same way, yielding 10 000 tests for size 350. As such, we
tested sizes from 100 through 750, and include a plot of the results in
Fig. <xref ref-type="fig" rid="Ch1.F3"/>. Since all 801 simulations are created by the
same type of perturbation, we expect EET to issue a pass for each
experimental set against<?pagebreak page703?> each ensemble. This study is essentially a
resampling method without replacement used jointly with cross validation to
ascertain the minimum ensemble size for stable PC calculations and pass/fail
determinations. With greater ensemble size the distribution of EET failure
rates should narrow, reflecting the increased stability of calculated PC
loadings that accompanies larger sample sizes. The EET failure rates will
never be uniformly zero due to the statistical nature of the test. The chosen
false positive rate of 0.5 % is reflected by the red horizontal line in
Fig. <xref ref-type="fig" rid="Ch1.F3"/>. We define an adequate ensemble size as one
whose median is less than 0.5 % and whose interquartile range (IQR) is
narrow. The IQR is defined as the difference between the upper and lower
quartiles of a distribution. For the remainder of this work we use the size
350 ensemble shown in Fig. <xref ref-type="fig" rid="Ch1.F3"/>, as it is the smallest
ensemble that meets our criteria of median below 0.5 % and narrow IQR. The
larger ensembles represent diminishing returns at greater computational
expense. Note that the relationship between model time step number and the
ensemble size necessary to optimize test accuracy is complex.
<xref ref-type="bibr" rid="bib1.bibx9" id="text.37"/> concludes that ensembles of size 300 or 453 are necessary
for accurate CAM-ECT test results, which bounds the 350 chosen for UF-CAM-ECT
above and below. Minimizing the cost of ensemble generation and test
evaluation is not a main consideration of this study, as UF-CAM-ECT is
already a sizable improvement over CAM-ECT.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F3"><?xmltex \currentcnt{3}?><label>Figure 3</label><caption><p id="d1e817">Box plot of EET failure rate distributions as a function of ensemble
size. The distributions are generated by randomly selecting a number of
simulations (ensemble size) from a set of 801 simulations to compute PC
loadings. From the remaining set, 30 simulations are chosen at random. These
simulations are projected into the PC space of the ensemble and evaluated via
EET. For each ensemble size, 100 ensembles are created and 100 experimental
sets are selected and evaluated. Thus each distribution contains 10 000 EET
results (40 600 000 total tests per distribution). The red horizontal line
indicates the chosen false positive rate of 0.5 %.</p></caption>
        <?xmltex \igopts{width=241.848425pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/11/697/2018/gmd-11-697-2018-f03.pdf"/>

      </fig>

</sec>
<sec id="Ch1.S5">
  <label>5</label><title>Results</title>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T1"><?xmltex \currentcnt{1}?><label>Table 1</label><caption><p id="d1e838">CAM-ECT and UF-CAM-ECT return the same result (pass or fail) for the
experiments listed. The CAM-ECT column is the result of a single test
on three runs. The UF-CAM-ECT column represents EET results from 30 runs.
Descriptions of the porting experiments (EDISON, CHEYENNE, and SUMMIT)
are listed in Sect. <xref ref-type="sec" rid="Ch1.S5.SS1"/>. The remaining experiments are
described in Appendices <xref ref-type="sec" rid="App1.Ch1.S1.SS1"/> and <xref ref-type="sec" rid="App1.Ch1.S1.SS2"/>.
</p></caption><oasis:table frame="topbot"><?xmltex \begin{scaleboxenv}{.85}[.85]?><oasis:tgroup cols="4">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="left"/>
     <oasis:colspec colnum="3" colname="col3" align="left"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:thead>
       <oasis:row>
         <oasis:entry colname="col1">Experiment</oasis:entry>
         <oasis:entry colname="col2">CAM-ECT </oasis:entry>
         <oasis:entry rowsep="1" namest="col3" nameend="col4" align="center">UF-CAM-ECT </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">Result</oasis:entry>
         <oasis:entry colname="col3">Result</oasis:entry>
         <oasis:entry colname="col4">EET failure %</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">INTEL-15</oasis:entry>
         <oasis:entry colname="col2">Pass</oasis:entry>
         <oasis:entry colname="col3">Pass</oasis:entry>
         <oasis:entry colname="col4">0.1 %</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">PGI</oasis:entry>
         <oasis:entry colname="col2">Pass</oasis:entry>
         <oasis:entry colname="col3">Pass</oasis:entry>
         <oasis:entry colname="col4">0.1 %</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">GNU</oasis:entry>
         <oasis:entry colname="col2">Pass</oasis:entry>
         <oasis:entry colname="col3">Pass</oasis:entry>
         <oasis:entry colname="col4">0.0 %</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">NO-OPT</oasis:entry>
         <oasis:entry colname="col2">Pass</oasis:entry>
         <oasis:entry colname="col3">Pass</oasis:entry>
         <oasis:entry colname="col4">0.0 %</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">NO-THRD</oasis:entry>
         <oasis:entry colname="col2">Pass</oasis:entry>
         <oasis:entry colname="col3">Pass</oasis:entry>
         <oasis:entry colname="col4">0.0 %</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">EDISON</oasis:entry>
         <oasis:entry colname="col2">Pass</oasis:entry>
         <oasis:entry colname="col3">Pass</oasis:entry>
         <oasis:entry colname="col4">0.1 %</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">CHEYENNE (AVX2 disabled)</oasis:entry>
         <oasis:entry colname="col2">Pass</oasis:entry>
         <oasis:entry colname="col3">Pass</oasis:entry>
         <oasis:entry colname="col4">2.1 %</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">SUMMIT (FMA disabled)</oasis:entry>
         <oasis:entry colname="col2">Pass</oasis:entry>
         <oasis:entry colname="col3">Pass</oasis:entry>
         <oasis:entry colname="col4">0.0 %</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">C</oasis:entry>
         <oasis:entry colname="col2">Pass</oasis:entry>
         <oasis:entry colname="col3">Pass</oasis:entry>
         <oasis:entry colname="col4">0.7 %</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">E</oasis:entry>
         <oasis:entry colname="col2">Pass</oasis:entry>
         <oasis:entry colname="col3">Pass</oasis:entry>
         <oasis:entry colname="col4">0.0 %</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">DM</oasis:entry>
         <oasis:entry colname="col2">Pass</oasis:entry>
         <oasis:entry colname="col3">Pass</oasis:entry>
         <oasis:entry colname="col4">0.0 %</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">UO</oasis:entry>
         <oasis:entry colname="col2">Pass</oasis:entry>
         <oasis:entry colname="col3">Pass</oasis:entry>
         <oasis:entry colname="col4">0.0 %</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">P</oasis:entry>
         <oasis:entry colname="col2">Pass</oasis:entry>
         <oasis:entry colname="col3">Pass</oasis:entry>
         <oasis:entry colname="col4">0.0 %</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">CHEYENNE (AVX2 enabled)</oasis:entry>
         <oasis:entry colname="col2">Fail</oasis:entry>
         <oasis:entry colname="col3">Fail</oasis:entry>
         <oasis:entry colname="col4">92.2 %</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">SUMMIT (FMA enabled)</oasis:entry>
         <oasis:entry colname="col2">Fail</oasis:entry>
         <oasis:entry colname="col3">Fail</oasis:entry>
         <oasis:entry colname="col4">77.2 %</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">DUST</oasis:entry>
         <oasis:entry colname="col2">Fail</oasis:entry>
         <oasis:entry colname="col3">Fail</oasis:entry>
         <oasis:entry colname="col4">100.0 %</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">FACTB</oasis:entry>
         <oasis:entry colname="col2">Fail</oasis:entry>
         <oasis:entry colname="col3">Fail</oasis:entry>
         <oasis:entry colname="col4">100.0 %</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">FACTIC</oasis:entry>
         <oasis:entry colname="col2">Fail</oasis:entry>
         <oasis:entry colname="col3">Fail</oasis:entry>
         <oasis:entry colname="col4">100.0 %</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">RH-MIN-LOW</oasis:entry>
         <oasis:entry colname="col2">Fail</oasis:entry>
         <oasis:entry colname="col3">Fail</oasis:entry>
         <oasis:entry colname="col4">100.0 %</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">RH-MIN-HIGH</oasis:entry>
         <oasis:entry colname="col2">Fail</oasis:entry>
         <oasis:entry colname="col3">Fail</oasis:entry>
         <oasis:entry colname="col4">100.0 %</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">CLDFRC-DP</oasis:entry>
         <oasis:entry colname="col2">Fail</oasis:entry>
         <oasis:entry colname="col3">Fail</oasis:entry>
         <oasis:entry colname="col4">100.0 %</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">UW-SH</oasis:entry>
         <oasis:entry colname="col2">Fail</oasis:entry>
         <oasis:entry colname="col3">Fail</oasis:entry>
         <oasis:entry colname="col4">100.0 %</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">CONV-LND</oasis:entry>
         <oasis:entry colname="col2">Fail</oasis:entry>
         <oasis:entry colname="col3">Fail</oasis:entry>
         <oasis:entry colname="col4">100.0 %</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">CONV-OCN</oasis:entry>
         <oasis:entry colname="col2">Fail</oasis:entry>
         <oasis:entry colname="col3">Fail</oasis:entry>
         <oasis:entry colname="col4">100.0 %</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">NU-P</oasis:entry>
         <oasis:entry colname="col2">Fail</oasis:entry>
         <oasis:entry colname="col3">Fail</oasis:entry>
         <oasis:entry colname="col4">100.0 %</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup><?xmltex \end{scaleboxenv}?></oasis:table></table-wrap>

      <?pagebreak page704?><p id="d1e1272">The UF-CAM-ECT must have properties comparable or complementary to CAM-ECT
including high classification accuracy. In particular, its response to
modifications known to produce statistically distinguishable output should be
a fail, and to changes not expected to result in statistically
distinguishable output, a pass. We verify its effectiveness by performing the
same tests as before with CAM-ECT: CAM namelist alterations and compiler
changes from <xref ref-type="bibr" rid="bib1.bibx2" id="text.38"/> as well as code modifications from
<xref ref-type="bibr" rid="bib1.bibx9" id="text.39"/>. We further explore UF-CAM-ECT properties with experiments
from CLM and several new CAM experiments. In the following sections,
UF-CAM-ECT experiments consist of 30 runs due to their low cost, allowing us
to do exhaustive testing. For CAM-ECT, we only run EET (which is far more
expensive due to the need for more than three 12-month runs) in
Sect. <xref ref-type="sec" rid="Ch1.S5.SS3"/>, where the expected experiment outcomes
are less certain. The UF-CAM-ECT ensemble selected for testing is size 350
(see Fig. <xref ref-type="fig" rid="Ch1.F3"/>), and the CAM-ECT ensemble is size 300,
comprised of 100 simulations built by Intel, GNU, and PGI compilers (the
smallest size recommended in <xref ref-type="bibr" rid="bib1.bibx9" id="altparen.40"/>). <?xmltex \hack{\newpage}?></p>
<sec id="Ch1.S5.SS1">
  <label>5.1</label><title>Matching expectation and result: where UF and CAM-ECT agree</title>
      <p id="d1e1296">UF-CAM-ECT should return a pass when run against experiments expected to be
statistically indistinguishable from the ensemble. A comparison between the
EET failures of UF-CAM-ECT and the single-test CAM-ECT for several types of
experiments that should all pass is presented in the upper section of
Table <xref ref-type="table" rid="Ch1.T1"/>. The first type of examples for this “should pass”
category includes building CESM with a different compiler or a different
value-safe optimization order (e.g., with no optimization: -O0), or running
CESM without OpenMP threading. These tests are labeled INTEL-15, PGI,
GNU, NO-OPT, and NO-THRD (see Appendix <xref ref-type="sec" rid="App1.Ch1.S1.SS1"/> for
further details).</p>
      <p id="d1e1303">A second type of should pass examples includes the important category of
port verification to other (i.e., not Yellowstone) CESM-supported machines.
<xref ref-type="bibr" rid="bib1.bibx9" id="text.41"/> determined that running CESM with fused multiply–add (FMA)
CPU instructions enabled resulted in statistically distinguishable output on
the Argonne National Laboratory Mira supercomputer. With the instructions
disabled, the machine passed CAM-ECT. We list results from the machines in
Table <xref ref-type="table" rid="Ch1.T1"/> with FMA enabled and disabled on SUMMIT (note that
the SUMMIT results can also be found in <xref ref-type="bibr" rid="bib1.bibx1" id="altparen.42"/>), and with
xCORE-AVX2 (a set of optimizations that activates FMA) enabled and disabled
on CHEYENNE.</p>
      <p id="d1e1314"><list list-type="bullet">
            <list-item>

      <p id="d1e1319"><italic>EDISON</italic> Cray XC30 with Xeon Ivy Bridge CPUs at the National
Energy Research Scientific Computing Center (NERSC; Intel compiler, no FMA capability)</p>
            </list-item>
            <list-item>

      <p id="d1e1327"><italic>CHEYENNE</italic> SGI ICE XA cluster with Xeon Broadwell CPUs at NCAR
(Intel compiler, FMA capable)</p>
            </list-item>
            <list-item>

      <p id="d1e1335"><italic>SUMMIT</italic> Dell C6320 cluster with Xeon Haswell CPUs at the
University of Colorado Boulder for
the Rocky Mountain Advanced Computing Consortium (RMACC; Intel compiler, FMA capable)</p>
            </list-item>
          </list></p>
      <p id="d1e1342">Finally, a third type of should pass experiments is the minimal code
modifications from Milroy et al. (2016) that were developed to test the
variability and classification power of CESM-ECT. These code modifications
that should pass UF-CAM-ECT include the following: Combine (C), Expand (E),
Division-to-Multiplication (DM), Unpack-Order (UO), and Precision (P; see Appendix <xref ref-type="sec" rid="App1.Ch1.S1.SS2"/> for full descriptions). Note that all
EET failure rates for the types of experiments that should pass (in the upper
section of Table <xref ref-type="table" rid="Ch1.T1"/>) are close to zero for UF-CAM-ECT,
indicating full agreement between CAM-ECT and UF-CAM-ECT.</p>
      <p id="d1e1350">Next we further exercise UF-CAM-ECT by performing tests that are expected to
fail, which are presented in the lower section of Table <xref ref-type="table" rid="Ch1.T1"/>. We
perform the following CAM namelist experiments from <xref ref-type="bibr" rid="bib1.bibx2" id="text.43"/>:
DUST, FACTB, FACTIC, RH-MIN-LOW, RH-MIN-HIGH, CLDFRC-DP, UW-SH,
CONV-LND, CONV-OCN, and NU-P (see Appendix <xref ref-type="sec" rid="App1.Ch1.S1.SS1"/> for
descriptions). For UF-CAM-ECT each “expected to fail” result in
Table <xref ref-type="table" rid="Ch1.T1"/> (lower portion) is identically a 100 % EET failure:
a clear indication of statistical distinctness from the size 350 UF-CAM-ECT
ensemble. Therefore, the CAM-ECT and UF-CAM-ECT tests are in agreement for
the entire list of examples presented in Table <xref ref-type="table" rid="Ch1.T1"/>, which is a
testament to the utility of UF-CAM-ECT.</p>
</sec>
<sec id="Ch1.S5.SS2">
  <label>5.2</label><title>CLM modifications</title>
      <p id="d1e1373">The CLM, the land model component of CESM,
was initially developed to study land surface processes and
land–atmosphere interactions, and was a
product of a merging of a community land model with the NCAR Land
Surface Model <xref ref-type="bibr" rid="bib1.bibx10" id="paren.44"/>.  More recent versions benefit from
the integration of far more sophisticated physical processes than in
the original code. Specifically, CLM 4.0 integrates models of
vegetation phenology,  surface albedos, radiative fluxes, soil and
snow temperatures, hydrology, photosynthesis,  river transport, urban
areas, carbon–nitrogen cycles, and dynamic global vegetation, among
many others <xref ref-type="bibr" rid="bib1.bibx10" id="paren.45"/>.  Moreover, the CLM receives state variables from
CAM and updates hydrology calculations, outputting the fields back to
CAM <xref ref-type="bibr" rid="bib1.bibx10" id="paren.46"/>. It is sensible to assume that since
information propagates between the land and atmosphere models, in
particular between CLM and CAM, CAM-ECT and UF-CAM-ECT should be
capable of detecting changes to CLM.</p>
      <p id="d1e1385">For our tests we use CLM version 4.0, which is the default for our CESM
version (see Sect. <xref ref-type="sec" rid="Ch1.S3"/>) and the same version used in all
experiments in this work. Our CLM experiments are described as follows:</p>
      <p id="d1e1390"><list list-type="bullet">
            <list-item>

      <p id="d1e1395"><italic>CLM_INIT</italic> changes from using the default land initial
condition file to using a cold restart.</p>
            </list-item>
            <list-item>

      <p id="d1e1403"><italic>CO2_PPMV_280</italic> reduces the CO<inline-formula><mml:math id="M22" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> type and concentration
from CLM_CO2_TYPE <inline-formula><mml:math id="M23" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> “diagnostic”
to CLM_CO2_TYPE <inline-formula><mml:math id="M24" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> “constant” and CCSM_CO2_PPMV <inline-formula><mml:math id="M25" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 280.0.</p>
            </list-item>
            <list-item>

      <p id="d1e1441"><italic>CLM_VEG</italic> activates CN mode (carbon–nitrogen cycle coupling).</p>
            </list-item>
            <list-item>

      <p id="d1e1449"><italic>CLM_URBAN</italic> disables urban air conditioning/heating
and the waste heat associated with these processes so that the internal
building temperature floats freely.</p>
            </list-item>
          </list>See Table <xref ref-type="table" rid="Ch1.T2"/> for the test results of the experiments. The
pass and fail results in this table reflect our high confidence in the
expected outcome: all test determinations are in agreement, the UF-CAM-ECT
passes represent EET failure rates <inline-formula><mml:math id="M26" display="inline"><mml:mrow><mml:mo>&lt;</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> %, and failing UF-CAM-ECT tests
are all 100 % EET failures. We expected failures for CLM_INIT because the
CLM and CAM coupling period is 30 simulation minutes, and such a substantial
change to the initial conditions should be detected immediately and persist
through 12 months. CLM_CO2_PPMV_280 is also a tremendous change as it
effectively resets the atmospheric CO<inline-formula><mml:math id="M27" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> concentration to a preindustrial
value, and changes which CO<inline-formula><mml:math id="M28" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> value the model uses. In particular, for
CLM_CO2_TYPE <inline-formula><mml:math id="M29" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> “diagnostic” CLM uses the value from the atmosphere
(367.0 ppmv), while CLM_CO2_TYPE <inline-formula><mml:math id="M30" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> “constant” instructs CLM to use the
value specified by CCSM_CO2_PPMV. Therefore both tests detect the large
reduction in CO<inline-formula><mml:math id="M31" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> concentration, generating failures at the ninth time step
and in the 12-month average. CLM_VEG was also expected to fail immediately,
given how quickly the CN coupling is expressed. Finally, the passing results
of both CAM-ECT and UF-CAM-ECT for CLM_URBAN is unsurprising as the urban
fraction is less than 1 % of the land surface, and heating and air
conditioning only occur over a fraction of this 1 % as well.</p>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T2"><?xmltex \currentcnt{2}?><label>Table 2</label><caption><p id="d1e1515">These CLM experiments show agreement between CAM-ECT and
UF-CAM-ECT as well as with the expected outcome.
The CAM-ECT column is the result of a single ECT test on
three runs.  The UF-CAM-ECT column represents EET failure rates
from 30 runs.      </p></caption><oasis:table frame="topbot"><?xmltex \begin{scaleboxenv}{.95}[.95]?><oasis:tgroup cols="4">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="left"/>
     <oasis:colspec colnum="3" colname="col3" align="left"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:thead>
       <oasis:row>
         <oasis:entry colname="col1">Experiment</oasis:entry>
         <oasis:entry colname="col2">CAM-ECT </oasis:entry>
         <oasis:entry rowsep="1" namest="col3" nameend="col4" align="center">UF-CAM-ECT </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">Result</oasis:entry>
         <oasis:entry colname="col3">Result</oasis:entry>
         <oasis:entry colname="col4">EET failure %</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">CLM_INIT</oasis:entry>
         <oasis:entry colname="col2">Fail</oasis:entry>
         <oasis:entry colname="col3">Fail</oasis:entry>
         <oasis:entry colname="col4">100.0 %</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">CLM_CO2_PPMV_280</oasis:entry>
         <oasis:entry colname="col2">Fail</oasis:entry>
         <oasis:entry colname="col3">Fail</oasis:entry>
         <oasis:entry colname="col4">100.0 %</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">CLM_VEG</oasis:entry>
         <oasis:entry colname="col2">Fail</oasis:entry>
         <oasis:entry colname="col3">Fail</oasis:entry>
         <oasis:entry colname="col4">100.0 %</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">CLM_URBAN</oasis:entry>
         <oasis:entry colname="col2">Pass</oasis:entry>
         <oasis:entry colname="col3">Pass</oasis:entry>
         <oasis:entry colname="col4">0.1 %</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup><?xmltex \end{scaleboxenv}?></oasis:table></table-wrap>

      <p id="d1e1624">Our experiments thus far indicate that both CAM-ECT and UF-CAM-ECT
will detect errors in CLM, and that a separate CESM-ECT module for CLM
(required for POP) is most likely not needed. While this finding may be
unsurprising given how tightly CAM and CLM are coupled, it represents a
significant broadening of the tools' applicability and utility.
Note that while we have not
generated CAM-ECT or UF-CAM-ECT ensembles with CN mode activated in
CLM (which is a common configuration for land modeling), we have no reason to
believe that statistical consistency testing of CN-related CLM code
changes would not be equally successful.  Consistency testing of
active CN mode code changes bears further
investigation and will be a subject of future work.</p>
</sec>
<?pagebreak page705?><sec id="Ch1.S5.SS3">
  <label>5.3</label><title>UF-CAM-ECT and CAM-ECT disagreement</title>
      <p id="d1e1635">In this section we test experiments that result in contradictory
determinations by UF-CAM-ECT and CAM-ECT. Due to the disagreement, all tests'
EET failure percentages are reported for 30 run experimental sets for both
UF-CAM-ECT and CAM-ECT. We present the results in Table <xref ref-type="table" rid="Ch1.T3"/>.
The modifications are described in the following list (note that NU and
RAND-MT can also be found in <xref ref-type="bibr" rid="bib1.bibx2" id="text.47"/> and <xref ref-type="bibr" rid="bib1.bibx8" id="text.48"/>,
respectively):</p>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T3"><?xmltex \currentcnt{3}?><label>Table 3</label><caption><p id="d1e1649">These experiments represent disagreement between
UF-CAM-ECT and fail CAM-ECT.  Shown are the EET
failure rates from 30 runs.
</p></caption><oasis:table frame="topbot"><?xmltex \begin{scaleboxenv}{.95}[.95]?><oasis:tgroup cols="3">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:thead>
       <oasis:row>
         <oasis:entry colname="col1">Experiment</oasis:entry>
         <oasis:entry colname="col2" align="left">CAM-ECT </oasis:entry>
         <oasis:entry colname="col3" align="left">UF-CAM-ECT </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">EET failure %</oasis:entry>
         <oasis:entry colname="col3">EET failure %</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">RAND-MT</oasis:entry>
         <oasis:entry colname="col2">4.7 %</oasis:entry>
         <oasis:entry colname="col3">99.4 %</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">TSTEP_TYPE</oasis:entry>
         <oasis:entry colname="col2">2.5 %</oasis:entry>
         <oasis:entry colname="col3">100 %</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">QSPLIT</oasis:entry>
         <oasis:entry colname="col2">1.8 %</oasis:entry>
         <oasis:entry colname="col3">100 %</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">CPL_BUG</oasis:entry>
         <oasis:entry colname="col2">41.6 %</oasis:entry>
         <oasis:entry colname="col3">0.1 %</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">CLM_HYDRO_BASEFLOW</oasis:entry>
         <oasis:entry colname="col2">30.7 %</oasis:entry>
         <oasis:entry colname="col3">0.1 %</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">NU</oasis:entry>
         <oasis:entry colname="col2">33.0 %</oasis:entry>
         <oasis:entry colname="col3">100 %</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">CLM_ALBICE_00</oasis:entry>
         <oasis:entry colname="col2">12.8 %</oasis:entry>
         <oasis:entry colname="col3">96.3 %</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup><?xmltex \end{scaleboxenv}?></oasis:table></table-wrap>

      <p id="d1e1777"><list list-type="bullet">
            <list-item>

      <p id="d1e1782"><italic>RAND-MT</italic> substitutes the Mersenne Twister pseudo-random number generator
(PRNG) for the default PRNG in radiation modules.</p>
            </list-item>
            <list-item>

      <p id="d1e1790"><italic>TSTEP_TYPE</italic> changes the time-stepping method for the spectral element dynamical core from
4 (Kinnmark &amp; Gray Runge–Kutta 4 stage) to 5 (Kinnmark &amp; Gray Runge–Kutta 5 stage).</p>
            </list-item>
            <list-item>

      <p id="d1e1798"><italic>QSPLIT</italic> alters how often tracer advection is done in terms of dynamics
time steps, the default
is one, and we increase it to nine.</p>
            </list-item>
            <list-item>

      <p id="d1e1806"><italic>CPL_BUG</italic> sets albedos to zero above 57<inline-formula><mml:math id="M32" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula>N latitude in the coupler.</p>
            </list-item>
            <list-item>

      <p id="d1e1823"><italic>CLM_HYDRO_BASEFLOW</italic> increases the soil hydrology baseflow rate
coefficient in CLM from <inline-formula><mml:math id="M33" display="inline"><mml:mrow><mml:mn mathvariant="normal">5.5</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> to <inline-formula><mml:math id="M34" display="inline"><mml:mn mathvariant="normal">55</mml:mn></mml:math></inline-formula>.</p>
            </list-item>
            <list-item>

      <p id="d1e1857"><italic>NU</italic> changes the dynamics hyperviscosity (horizontal diffusion) from
<inline-formula><mml:math id="M35" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mn mathvariant="normal">15</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> to <inline-formula><mml:math id="M36" display="inline"><mml:mrow><mml:mn mathvariant="normal">9</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mn mathvariant="normal">14</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>.</p>
            </list-item>
            <list-item>

      <p id="d1e1895"><italic>CLM_ALBICE_00</italic> changes the albedo of bare ice on glaciers
(visible and near-infrared albedos for glacier ice) from 0.80, 0.55
to 0.00, 0.00.</p>
            </list-item>
          </list></p>
<sec id="Ch1.S5.SS3.SSS1">
  <label>5.3.1</label><?xmltex \opttitle{Minor setting changes: RAND-MT, TSTEP\_TYPE, and QSPLIT}?><title>Minor setting changes: RAND-MT, TSTEP_TYPE, and QSPLIT</title>
      <p id="d1e1911">RAND-MT is a test of the response to substituting the CAM default
PRNG in the radiation module for a
different CESM-supported PRNG <xref ref-type="bibr" rid="bib1.bibx8" id="paren.49"/>.  Since the PRNG
affects radiation modules which compute cloud properties, it is
reasonable to conclude that the change alters the distributions of
cloud-related CAM variables (such as cloud covers).  Both CAM and its
PRNG are deterministic; the variability at nine time steps
exhibits different characteristics depending on the PRNG.
However, we would not expect (nor would we want) a change
to the PRNG to induce statistically distinguishable results over
a longer period such as a simulation year, and this expectation is
confirmed by CAM-ECT.</p>
      <?pagebreak page706?><p id="d1e1917">TSTEP_TYPE and QSPLIT are changes to attributes of the model dynamics:
TSTEP_TYPE alters the time-stepping method in the dynamical core, and QSPLIT
modifies the frequency of tracer advection computation relative to the
dynamics time step. It is well known that CAM is generally much more
sensitive to the physics time step than to the dynamics time step.
Time-stepping errors in CAM dynamics do not affect large-scale well-resolved
waves in the atmosphere but they do affect small-scale fast waves. While
short-term weather should be affected, the model climate is not expected to
be affected by time-stepping method or dynamics time-step. However, like the
RAND-MT example, UF-CAM-ECT registers the less “smoothed” instantaneous
global means as failures for both tests, while CAM-ECT finds the yearly
averaged global means to be statistically indistinguishable. Small grid-scale
waves are affected by choice of TSTEP_TYPE in short runs; however, the
long-term climate is not affected by time-stepping method. The results of
these experiments are shown in the top section of Table <xref ref-type="table" rid="Ch1.T3"/>,
and for experiments of this type, CAM-ECT yields anticipated results. This
group of experiments exemplifies the categories of experiments to which
UF-CAM-ECT may be sensitive: small-scale or minor changes to initial
conditions or settings which are irrelevant in the long term. Therefore,
while the UF-CAM-ECT results can be misleading in particular in these cases,
they may indicate a larger problem as will be seen in the examples in
Sect. <xref ref-type="sec" rid="Ch1.S5.SS3.SSS3"/>.</p>
</sec>
<sec id="Ch1.S5.SS3.SSS2">
  <label>5.3.2</label><?xmltex \opttitle{Contrived experiments: CPL\_BUG and CLM\_HYDRO\_BASEFLOW}?><title>Contrived experiments: CPL_BUG and CLM_HYDRO_BASEFLOW</title>
      <p id="d1e1933">Motivated by experiments which bifurcate the tests' findings, we seek the
reverse of the previous three experiments in
Sect. <xref ref-type="sec" rid="Ch1.S5.SS3.SSS1"/>: examples of a parameter change or
code modification that are distinguishable in the yearly global means, but
are undetectable in the first time steps. We consulted with climate
scientists and CESM software engineers, testing a large number of possible
modifications to find some that would pass UF-CAM-ECT and fail CAM-ECT. The
results in the center section of Table <xref ref-type="table" rid="Ch1.T3"/> represent a small
fraction of the tests performed, as examples that met the condition of
UF-CAM-ECT pass and CAM-ECT fail were exceedingly difficult to find. In fact,
CPL_BUG and CLM_HYDRO_BASEFLOW were devised specifically for that purpose.
That their failure rates are far from 100 % is an indication of the
challenge of finding an error that is not present at nine time steps, but
manifests clearly in the annual average.</p>
      <p id="d1e1940">CPL_BUG is devised to demonstrate that it is possible to construct an
example that does not yield substantial differences in output at nine time
steps, but does impact the yearly average. Selectively setting the albedos to
zero above 57<inline-formula><mml:math id="M37" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula>N latitude has little effect at nine time steps since
this region experiences almost zero solar radiation during the first 5 h of 1 January. The nonzero CAM-ECT result is a consequence of using
annual averages since for the Northern Hemisphere summer months this region
is exposed to nearly constant solar irradiance.</p>
      <p id="d1e1952">CLM_HYDRO_BASEFLOW is another manufactured example of a change
designed   to be undetectable at the ninth time step.  It is an
increase  in the exponent of the soil hydrology baseflow rate
coefficient, which controls the amount of water drained from the soil.
This substantial change (4 orders of magnitude) cannot be detected
by UF-CAM-ECT since the differences at nine time steps are confined to
deep layers of the soil.  However, through the year they propagate to and
eventually influence the atmosphere through changes in surface
fluxes, which is corroborated by
the much higher CAM-ECT failure rate.</p>
</sec>
<sec id="Ch1.S5.SS3.SSS3">
  <label>5.3.3</label><?xmltex \opttitle{Small but consequential changes: NU and CLM\_ALBICE\_00}?><title>Small but consequential changes: NU and CLM_ALBICE_00</title>
      <p id="d1e1964">CAM-ECT results in <xref ref-type="bibr" rid="bib1.bibx2" id="text.50"/> for the NU experiment are of particular
interest as climate scientists expected this experiment to fail. NU is an
extraordinary case, as <xref ref-type="bibr" rid="bib1.bibx2" id="text.51"/> acknowledge: “[b]ecause CESM-ECT
[currently CAM-ECT] looks at variable annual global means, the “pass”
result [for NU] is not entirely surprising as errors in small-scale behavior
are unlikely to be detected in a yearly global mean”. The change to NU
should be evident only where there are strong field gradients and small-scale
precipitation. We applied EET for CAM-ECT with 30 runs and determined the
failure rate to be 33.0 % against the reference ensemble. In terms of
CAM-ECT this experiment was borderline, as the probability that CAM-ECT will
classify three NU runs a “pass” is not much greater than a “fail”
outcome. In contrast UF-CAM-ECT is able to detect this difference much more
definitively in the instantaneous data at the ninth time step. We would also
expect this experiment to fail more definitively for simulations longer than
12 months, once the small but nevertheless consequential change in NU had
time to manifest.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F4" specific-use="star"><?xmltex \currentcnt{4}?><label>Figure 4</label><caption><p id="d1e1975">Each box plot represents the statistical distribution of the
difference between the global mean of each variable and the unperturbed,
ensemble global mean, then scaled by the unperturbed, ensemble global mean
for both the 30 ensemble members and 30 CLM_ALBICE_00 members. The plots on
the left <bold>(a)</bold> are generated from nine time step simulations, while
those on the right <bold>(b)</bold> are from one simulation year.</p></caption>
            <?xmltex \igopts{width=441.017717pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/11/697/2018/gmd-11-697-2018-f04.pdf"/>

          </fig>

      <p id="d1e1990">CLM_ALBICE_00 affects a very small percent of the land area. Furthermore,
of that land area, only regions where the fractional snow cover is <inline-formula><mml:math id="M38" display="inline"><mml:mo>&lt;</mml:mo></mml:math></inline-formula> 1 and
incoming solar radiation is present will be affected by the modification to
the bare ice albedo. Therefore, it was expected to pass both CAM-ECT and
UF-CAM-ECT, yet the EET failure rate for UF-CAM-ECT was 96.3 %. We consider
the CLM_ALBICE_00 experiment in greater detail to better understand the
differences between UF-CAM-ECT and CAM-ECT. Since the change is small and
localized, we need to discover the reason why UF-CAM-ECT detects a
statistical difference, particularly given that many northern regions with
glaciation receive little or no solar radiation at time step 9 (1 January).
To explain this unanticipated result, we created box plots of all 108 CAM
variables tested by UF-CAM-ECT to compare the distributions (at nine time
steps and at 1 year) of the ensemble versus the 30 CLM_ALBICE_00
simulations. Each plot was generated by subtracting the unperturbed ensemble
value from<?pagebreak page707?> each value, and then rescaling by the unperturbed ensemble value.
After analyzing the plots, we isolated four variables (FSDSC: clear-sky
downwelling solar flux at surface; FSNSC: clear-sky net solar flux at
surface;
FSNTC: clear-sky net solar flux at top of model; and FSNTOAC: clear-sky net
solar flux at top of atmosphere) that exhibited markedly different behaviors
between the ensemble and experimental outputs. Figure <xref ref-type="fig" rid="Ch1.F4"/>
displays the results. The left column represents distributions from the ninth
time step which demonstrate the distinction between the ensemble and
experiment: the top three variables' distributions have no overlap. For the
12-month runs, the ensemble and experiments are much less distinct. It is
sensible that the global mean net fluxes are increased by the albedo change,
as the incident solar radiation should be a constant, while the zero albedo
forces all radiation impinging on exposed ice to be absorbed. The absorption
reduces the negative radiation flux, making the net flux more positive. The
yearly mean distributions are not altered enough for CAM-ECT to robustly
detect a fail (12.8 % EET failure rate), which is due to feedback
mechanisms having taken hold and leading to spatially heterogeneous effects,
which are seen as such in the spatial and temporal 12-month average.</p>
      <?pagebreak page708?><p id="d1e2003">The percentage of grid cells affected by the CLM_ALBICE_00 modification is
0.36 % (calculated by counting the number of cells with nonzero FSNS,
fractional snow cover (FSNO in CLM) <inline-formula><mml:math id="M39" display="inline"><mml:mrow><mml:mo>∈</mml:mo><mml:mo>(</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, and PCT_GLACIER greater
than zero in the surface data set). Remarkably, despite such a small area
being affected by the CLM_ALBICE_00 change, UF-CAM-ECT flags these simulations as statistically
distinguishable. The results of CLM_ALBICE_00 taken together with NU
indicate that UF-CAM-ECT demonstrates the ability to detect small-scale
events, fulfilling the desired capability of CAM-ECT mentioned as future work
in <xref ref-type="bibr" rid="bib1.bibx2" id="text.52"/>.</p>
</sec>
</sec>
</sec>
<sec id="Ch1.S6">
  <label>6</label><title>Implications and ECT guidelines</title>
      <p id="d1e2037">In this section we summarize the lessons learned in Sect. <xref ref-type="sec" rid="Ch1.S5"/>
to provide both clarification and guidance on the use of the complementary
tools UF-CAM-ECT and CAM-ECT in practice. Our extensive experiments, a
representative subset of which are presented in Sect. <xref ref-type="sec" rid="Ch1.S5"/>,
indicate that UF-CAM-ECT and CAM-ECT typically return the same determination.
Indeed, finding counterexamples was non-trivial. Certainly the types of
modifications that occur frequently in the CESM development cycle (e.g.,
compiler upgrades, new CESM-supported platforms, minor code rearrangements
for optimization, and initial state changes) are all equally well classified
by both UF-CAM-ECT and CAM-ECT. Therefore, in practice we recommend the use
of the cheaper UF-CAM-ECT as a first step for port verification, code
optimization, and compiler flag changes, as well as other frequent CESM
quality assurance procedures. Moreover, the low cost of ensemble generation
provides researchers and software engineers with the ability to generate
ensembles rapidly for evaluation of new physics, chemistry, or other
modifications which could affect the climate.</p>
      <p id="d1e2044">CAM-ECT is used as a second step only when needed for complementary
information as follows. First, our experimentation indicates that if
UF-CAM-ECT issues a pass, it is very likely that CAM-ECT will also issue a
pass. While devising examples where UF-CAM-ECT issues a pass and CAM-ECT
issues a fail is conceptually straightforward (e.g., a seasonal or
slow-propagating effect), in practice none of the changes suggested by
climate scientists and software engineers resulted in a discrepancy between
CAM-ECT and UF-CAM-ECT. Hence, we constructed the two examples presented in
Sect. <xref ref-type="sec" rid="Ch1.S5.SS3.SSS2"/>, using changes which were formulated
specifically to be undetectable by UF-CAM-ECT, but flagged as statistically
distinguishable by CAM-ECT. It appears that if a change propagates so slowly
as not to be detected at the ninth time step, its later effects can be
smoothed by the annual averaging which includes the initial behavior.
Accordingly, the change may go undetected by CAM-ECT when used without EET
(e.g., failure rates for CAM-ECT in the lower third of
Table <xref ref-type="table" rid="Ch1.T3"/> are well below 100 %). A user may choose to run
both tests, but in practice applying CAM-ECT as a second step should only be
considered when UF-CAM-ECT issues a fail. In particular, because we have
shown that UF-CAM-ECT is quite sensitive to small-scale errors or alterations
(see CLM_ALBICE_00 in Sect. <xref ref-type="sec" rid="Ch1.S5.SS3.SSS3"/>, which impacted less
than 1 % of land area), by running CAM-ECT when UF-CAM-ECT fails, we can
further determine whether the change also impacted statistical consistency
during the first year. If CAM-ECT also fails, then the UF-CAM-ECT result is
confirmed. On the other hand, if CAM-ECT passes, the situation is more
nuanced. Either a small-scale change has occurred that is unimportant in
the long term for the mean climate (e.g., RAND-MT), or a small-scale change has
occurred that will require a longer timescale than 12 months to manifest
decisively (e.g., NU). In either case, the user must have an understanding of
the characteristics of the modification being tested to reconcile the results
at this point. Future work will include investigation of ensembles at longer
timescales, which will aid in the overall determination of the relevance of
the error.</p>
</sec>
<sec id="Ch1.S7" sec-type="conclusions">
  <label>7</label><title>Conclusions</title>
      <p id="d1e2061">We developed a new Ultra-Fast CAM Ensemble Consistency test from
output at the ninth time step of the CESM. Conceived largely out of
curiosity, it proved  to possess surprisingly wide
applicability in part  due to its use of instantaneous values rather
than annual means.  The short simulation time translated to a cost
savings of a factor of approximately 70 over a simulation of 12 months,
considerably reducing the expense  of ensemble and test run creation.
Through methodical testing, we selected a UF-CAM-ECT ensemble  size
(350) that balances the variability contained in the  ensemble (hence
its ability to classify new runs) with the cost of generation.   We
performed extensive experimentation to test which modifications known
to produce statistically distinguishable and indistinguishable results
would be classified as such by UF-CAM-ECT. These experiments yielded
clear pass/fail results that were in agreement between the two tests,
allowing us to more confidently prescribe use cases for UF-CAM-ECT.
Due to the established feedback mechanisms between the CLM component of CESM and CAM, we extended CESM-ECT testing to
CLM. Our determination that both CAM-ECT and UF-CAM-ECT are capable of
identifying  statistical distinguishability resulting from alterations
to CLM indicates that a separate ECT module for CLM is likely
unnecessary.  By studying experiments where CAM-ECT and UF-CAM-ECT
arrived at different findings we concluded that UF-CAM-ECT is capable
of detecting small-scale changes, a feature that facilitates root
cause analysis for test failures in conjunction with CAM-ECT.</p>
      <p id="d1e2064">UF-CAM-ECT will be an asset to CESM model developers, software engineers, and
climate scientists. The ultra-fast test is cheap and quick, and further
testing is not<?pagebreak page709?> required when a passing result indicating statistical
consistency is issued. Ultimately the two tests can be used in concert to
provide richer feedback to software engineers, hardware experts, and climate
scientists: combining the results from the ninth time step and 12 months
enhances understanding of the timescales on which changes become operative
and influential. We intend to refine our understanding of both UF-CAM-ECT and
CAM-ECT via an upcoming study on decadal simulations. We hope to determine
whether the tests are capable of identifying statistical consistency (or lack
thereof) of modifications that may take many years to manifest fully. Another
potential application of the tests is the detection of hardware or software
issues during the initial evaluation and routine operation of a
supercomputer. <?xmltex \hack{\newpage}?></p>
</sec>

      
      </body>
    <back><notes notes-type="codedataavailability"><title>Code and data availability</title>

      <p id="d1e2072">The current version (v3.0.0) of Python tools can be
obtained directly from <uri>https://github.com/NCAR/PyCECT</uri>. In addition,
CESM-ECT is available as part of the Common Infrastructure for Modeling Earth
(CIME) at <uri>https://github.com/ESMCI/cime</uri>. CESM-ECT will be included in
the CESM public releases beginning with the version 2.0 series.</p>
  </notes><?xmltex \hack{\clearpage}?><app-group>

<?pagebreak page710?><app id="App1.Ch1.S1">
  <label>Appendix A</label><title>Referenced experiments</title>
<sec id="App1.Ch1.S1.SS1">
  <label>A1</label><title>Experiments from Baker et al. (2015)</title>
      <p id="d1e2099"><list list-type="bullet">
            <list-item>

      <p id="d1e2104">NO-OPT: changing the Intel compiler flag to remove optimization</p>
            </list-item>
            <list-item>

      <p id="d1e2110">INTEL-15: changing the Intel compiler version to 15.0.0</p>
            </list-item>
            <list-item>

      <p id="d1e2116">NO-THRD: compiling CAM without threading (MPI-only)</p>
            </list-item>
            <list-item>

      <p id="d1e2122">PGI: using the CESM-supported PGI compiler (13.0)</p>
            </list-item>
            <list-item>

      <p id="d1e2128">GNU: using the CESM-supported GNU compiler (4.8.0)</p>
            </list-item>
            <list-item>

      <p id="d1e2135">EDISON:  National  Energy  Research  Scientific  Computing Center (Cray XC30, Intel)</p>
            </list-item>
            <list-item>

      <p id="d1e2141">DUST: dust emissions; dust_emis_fact <inline-formula><mml:math id="M40" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 0.45 (original default 0.55)</p>
            </list-item>
            <list-item>

      <p id="d1e2154">FACTB: wet deposition of aerosols convection factor; sol_factb_interstitial <inline-formula><mml:math id="M41" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 1.0 (original default 0.1)</p>
            </list-item>
            <list-item>

      <p id="d1e2167">FACTIC: wet deposition of aerosols convection factor; sol_factic_interstitial <inline-formula><mml:math id="M42" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 1.0 (original default 0.4)</p>
            </list-item>
            <list-item>

      <p id="d1e2180">RH-MIN-LOW: min. relative humidity for low clouds; cldfrc_rhminl <inline-formula><mml:math id="M43" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 0.85 (original default 0.8975)</p>
            </list-item>
            <list-item>

      <p id="d1e2193">RH-MIN-HIGH: min. relative humidity for high clouds; cldfrc_rhminh <inline-formula><mml:math id="M44" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 0.9 (original default 0.8)</p>
            </list-item>
            <list-item>

      <p id="d1e2207">CLDFRC-DP:  deep  convection  cloud  fraction; cld_frc_dp1 <inline-formula><mml:math id="M45" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 0.14 (original default 0.10)</p>
            </list-item>
            <list-item>

      <p id="d1e2220">UW-SH: penetrative entrainment efficiency – shallow; uwschu_rpen <inline-formula><mml:math id="M46" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 10.0 (original default 5.0)</p>
            </list-item>
            <list-item>

      <p id="d1e2233">CONV-LND: autoconversion over land in deep convection; zmconv_c0_lnd <inline-formula><mml:math id="M47" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 0.0035 (original default 0.0059)</p>
            </list-item>
            <list-item>

      <p id="d1e2246">CONV-OCN: autoconversion over ocean in deep convection; zmconv_c0_ocn <inline-formula><mml:math id="M48" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 0.0035 (original default 0.045)</p>
            </list-item>
            <list-item>

      <p id="d1e2259">NU-P:  hyperviscosity  for  layer  thickness  (vertical  lagrangian dynamics); nu_p <inline-formula><mml:math id="M49" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M50" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mn mathvariant="normal">14</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> (original default <inline-formula><mml:math id="M51" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mn mathvariant="normal">15</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>)</p>
            </list-item>
            <list-item>

      <p id="d1e2302">NU: dynamics hyperviscosity (horizontal diffusion); nu <inline-formula><mml:math id="M52" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M53" display="inline"><mml:mrow><mml:mn mathvariant="normal">9</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mn mathvariant="normal">14</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> (original default <inline-formula><mml:math id="M54" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mn mathvariant="normal">15</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>)</p>
            </list-item>
          </list></p>
</sec>
<sec id="App1.Ch1.S1.SS2">
  <label>A2</label><title>Experiments from Milroy et al. (2016)</title>
      <p id="d1e2352"><list list-type="bullet">
            <list-item>

      <p id="d1e2357"><italic>Combine</italic> (C) is a single-line code change to the <italic>preq_omega_ps</italic> subroutine.</p>
            </list-item>
            <list-item>

      <p id="d1e2368"><italic>Expand</italic> (E) is a modification to the <italic>preq_hydrostatic</italic>
subroutine.  We expand the calculation of the variable <monospace>phi</monospace>.</p>
            </list-item>
            <list-item>

      <p id="d1e2382"><italic>Division-to-Multiplication</italic> (DM): the original version of the
<italic>euler_step</italic> subroutine of the primitive trace advection module
(<italic>prim_advection_mod.F90</italic>) includes an operation that divides by a
spherical mass matrix <monospace>spheremp</monospace>. The modification to this kernel
consists of declaring a temporary variable (<monospace>tmpsphere</monospace>) defined as the
inverse of <monospace>spheremp</monospace>, and substituting a multiplication for the more
expensive division operation.</p>
            </list-item>
            <list-item>

      <p id="d1e2405"><italic>Unpack-Order</italic> (UO) changes the order that an MPI
receive buffer is unpacked in the <italic>edgeVunpack</italic> subroutine of
<italic>edge_mod.F90</italic>.  Changing the order of buffer unpacking has implications
for performance, as traversing the buffer sub-optimally can prevent cache
prefetching.</p>
            </list-item>
            <list-item>

      <p id="d1e2419"><italic>Precision</italic> (P) is a performance-oriented
modification to the water vapor saturation module (<italic>wv_sat_methods.F90</italic>).
From a performance perspective this could be extremely
advantageous and could present an opportunity for co-processor acceleration due
to superior single-precision computation speed.  We modify
the elemental function that computes saturation vapor pressure by substituting <monospace>r4</monospace>
for <monospace>r8</monospace> and casting to single precision in the original.</p>
            </list-item>
          </list></p><?xmltex \hack{\clearpage}?>
</sec>
</app>
  </app-group><notes notes-type="competinginterests"><title>Competing interests</title>

      <p id="d1e2441">The authors declare that they have no conflict of
interest.</p>
  </notes><ack><title>Acknowledgements</title><p id="d1e2447">Many thanks are due to Haiying Xu for her help modifying pyCECT for
comparison of CESM-ECT single- and double-precision floating-point
statistical consistency results.
We wish to thank Michael Levy for his suggestion of altering
albedos above a certain northern latitude (CPL_BUG), Peter Lauritzen
for his fruitful discussion of CAM physics and dynamics, and
Keith Oleson for his considerable assistance interpreting and
visualizing CLM experiments and output.</p><p id="d1e2449">We would like to acknowledge high-performance computing support from
Yellowstone (ark:/85065/d7wd3xhc) and Cheyenne
(<ext-link xlink:href="http://dx.doi.org/10.5065/D6RX99HX">doi:10.5065/D6RX99HX</ext-link>) provided by
NCAR's Computational and Information Systems Laboratory, sponsored by the
National Science Foundation. This work utilized the RMACC Summit
supercomputer, which is supported by the National Science Foundation (awards
ACI-1532235 and ACI-1532236), the University of Colorado Boulder, and
Colorado State University. The Summit supercomputer is a joint effort of the
University of Colorado Boulder and Colorado State University. This research
also used computing resources provided by the National Energy Research
Scientific Computing Center, a DOE Office of Science User Facility supported
by the Office of Science of the US Department of Energy under contract no.
DEAC02-05CH11231. This work was funded in part by the Intel Parallel
Computing Center for Weather and Climate Simulation
(<ext-link xlink:href="https://software.intel.com/en-us/articles/intel-parallel-computing-center-at-the-university-of-colorado-boulder-and-the-national">https://software.intel.com/en-us/articles/intel-parallel-computing-center-at-the-university-of-colorado-boulder-and-the-national</ext-link>).<?xmltex \hack{\newline}?><?xmltex \hack{\newline}?>Edited
by: Olivier Marti <?xmltex \hack{\newline}?>Reviewed by: two anonymous referees</p></ack><ref-list>
    <title>References</title>

      <ref id="bib1.bibx1"><label>Anderson et al.(2017)Anderson, Burns, Milroy, Ruprecht, Hauser, and
Siegel</label><mixed-citation>Anderson, J., Burns, P. J., Milroy, D., Ruprecht, P., Hauser, T., and Siegel,
H. J.: Deploying RMACC Summit: An HPC Resource for the Rocky Mountain Region,
in: Proceedings of the Practice and Experience in Advanced Research Computing
2017 on Sustainability, Success and Impact, PEARC17, 8:1–8:7, ACM, New
York, NY, USA, <ext-link xlink:href="https://doi.org/10.1145/3093338.3093379" ext-link-type="DOI">10.1145/3093338.3093379</ext-link>,
2017.</mixed-citation></ref>
      <ref id="bib1.bibx2"><label>Baker et al.(2015)Baker, Hammerling, Levy, Xu, Dennis, Eaton,
Edwards, Hannay, Mickelson, Neale, Nychka, Shollenberger, Tribbia,
Vertenstein, and Williamson</label><mixed-citation>Baker, A. H., Hammerling, D. M., Levy, M. N., Xu, H., Dennis, J. M., Eaton,
B. E., Edwards, J., Hannay, C., Mickelson, S. A., Neale, R. B., Nychka, D.,
Shollenberger, J., Tribbia, J., Vertenstein, M., and Williamson, D.: A new
ensemble-based consistency test for the Community Earth System Model (pyCECT
v1.0), Geosci. Model Dev., 8, 2829–2840,
<ext-link xlink:href="https://doi.org/10.5194/gmd-8-2829-2015" ext-link-type="DOI">10.5194/gmd-8-2829-2015</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx3"><label>Baker et al.(2016)Baker, Hu, Hammerling, Tseng, Xu, Huang, Bryan, and
Yang</label><mixed-citation>Baker, A. H., Hu, Y., Hammerling, D. M., Tseng, Y.-H., Xu, H., Huang, X.,
Bryan, F. O., and Yang, G.: Evaluating statistical consistency in the ocean
model component of the Community Earth System Model (pyCECT v2.0), Geosci.
Model Dev., 9, 2391–2406, <ext-link xlink:href="https://doi.org/10.5194/gmd-9-2391-2016" ext-link-type="DOI">10.5194/gmd-9-2391-2016</ext-link>, 2016.
</mixed-citation></ref><?xmltex \hack{\newpage}?>
      <ref id="bib1.bibx4"><label>Branstator and Teng(2010)</label><mixed-citation>Branstator, G. and Teng, H.: Two Limits of Initial-Value Decadal Predictability
in a CGCM, J. Climate, 23, 6292–6311, <ext-link xlink:href="https://doi.org/10.1175/2010JCLI3678.1" ext-link-type="DOI">10.1175/2010JCLI3678.1</ext-link>,
2010.</mixed-citation></ref>
      <ref id="bib1.bibx5"><label>Deser et al.(2012)Deser, Phillips, Bourdette, and Teng</label><mixed-citation>Deser, C., Phillips, A., Bourdette, V., and Teng, H.: Uncertainty in climate
change projections: the role of internal variability, Clim. Dynam., 38,
527–546, <ext-link xlink:href="https://doi.org/10.1007/s00382-010-0977-x" ext-link-type="DOI">10.1007/s00382-010-0977-x</ext-link>,
2012.</mixed-citation></ref>
      <ref id="bib1.bibx6"><label>Hurrell et al.(2013)Hurrell, Holland, Gent, Ghan, Kay, Kushner,
Lamarque, Large, Lawrence, Lindsay, Lipscomb, Long, Mahowald, Marsh, Neale,
Rasch, Vavrus, Vertenstein, Bader, Collins, Hack, Kiehl, and
Marshall</label><mixed-citation>Hurrell, J., Holland, M., Gent, P., Ghan, S., Kay, J., Kushner, P., Lamarque,
J.-F., Large, W., Lawrence, D., Lindsay, K., Lipscomb, W., Long, M.,
Mahowald, N., Marsh, D., Neale, R., Rasch, P., Vavrus, S., Vertenstein, M.,
Bader, D., Collins, W., Hack, J., Kiehl, J., and Marshall, S.: The
Community Earth System Model: A Framework for Collaborative Research,
B. Am. Meteorol. Soc., 94, 1339–1360,
<ext-link xlink:href="https://doi.org/10.1175/BAMS-D-12-00121.1" ext-link-type="DOI">10.1175/BAMS-D-12-00121.1</ext-link>, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx7"><label>Kay et al.(2015)Kay, Deser, Phillips, Mai, Hannay, Strand, Arblaster,
Bates, Danabasoglu, Edwards, Holland, Kushner, Lamarque, Lawrence, Lindsay,
Middleton, Munoz, Neale, Oleson, Polvani, and Vertenstein</label><mixed-citation>Kay, J. E., Deser, C., Phillips, A., Mai, A., Hannay, C., Strand, G.,
Arblaster, J. M., Bates, S. C., Danabasoglu, G., Edwards, J., Holland, M.,
Kushner, P., Lamarque, J.-F., Lawrence, D., Lindsay, K., Middleton, A.,
Munoz, E., Neale, R., Oleson, K., Polvani, L., and Vertenstein, M.: The
Community Earth System Model (CESM) Large Ensemble Project: A Community
Resource for Studying Climate Change in the Presence of Internal Climate
Variability, B. Am. Meteorol. Soc., 96, 1333–1349,
<ext-link xlink:href="https://doi.org/10.1175/BAMS-D-13-00255.1" ext-link-type="DOI">10.1175/BAMS-D-13-00255.1</ext-link>,
2015.</mixed-citation></ref>
      <ref id="bib1.bibx8"><label>Milroy(2015)</label><mixed-citation>Milroy, D. J.: Refining the composition of CESM-ECT ensembles, ProQuest
Dissertations and Theses, p. 60, available at:
<uri>http://search.proquest.com/docview/1760171657?accountid=28174</uri> (last
access: 10 January 2017),
2015.</mixed-citation></ref>
      <ref id="bib1.bibx9"><label>Milroy et al.(2016)Milroy, Baker, Hammerling, Dennis, Mickelson, and
Jessup</label><mixed-citation>Milroy, D. J., Baker, A. H., Hammerling, D. M., Dennis, J. M., Mickelson,
S. A., and Jessup, E. R.: Towards Characterizing the Variability of
Statistically Consistent Community Earth System Model Simulations, Procedia
Computer Science, 80, 1589–1600,
<ext-link xlink:href="https://doi.org/10.1016/j.procs.2016.05.489" ext-link-type="DOI">10.1016/j.procs.2016.05.489</ext-link>,
International Conference on Computational Science 2016, ICCS 2016,
6–8 June 2016, San Diego, California, USA, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx10"><label>Oleson et al.(2010)Oleson, Lawrence, B, Flanner, Kluzek, J, Levis,
Swenson, Thornton, Feddema, Heald, Lamarque, Niu, Qian, Running, Sakaguchi,
Yang, Zeng, and Zeng</label><mixed-citation>
Oleson, K. W., Lawrence, D. M., B, G., Flanner, M. G., Kluzek, E., J, P.,
Levis, S., Swenson, S. C., Thornton, E., Feddema, J., Heald, C. L., Lamarque,
J., Niu, G., Qian, T., Running, S., Sakaguchi, K., Yang, L., Zeng, X., and
Zeng, X.: Technical description of version 4.0 of the Community Land Model, NCAR Tech. Note NCAR/TN-478+STR, 257.</mixed-citation></ref>
      <ref id="bib1.bibx11"><label>Rosinski and Williamson(1997)</label><mixed-citation>Rosinski, J. M. and Williamson, D. L.: The Accumulation of Rounding Errors and
Port Validation for Global Atmospheric Models, SIAM J. Sci. Comput., 18,
552–564, <ext-link xlink:href="https://doi.org/10.1137/S1064827594275534" ext-link-type="DOI">10.1137/S1064827594275534</ext-link>,
1997.</mixed-citation></ref>
      <ref id="bib1.bibx12"><label>Wan et al.(2017)Wan, Zhang, Rasch, Singh, Chen, and
Edwards</label><mixed-citation>Wan, H., Zhang, K., Rasch, P. J., Singh, B., Chen, X., and Edwards, J.: A new
and inexpensive non-bit-for-bit solution reproducibility test based on time
step convergence (TSC1.0), Geosci. Model Dev., 10, 537–552,
<ext-link xlink:href="https://doi.org/10.5194/gmd-10-537-2017" ext-link-type="DOI">10.5194/gmd-10-537-2017</ext-link>, 2017.</mixed-citation></ref>

  </ref-list></back>
    <!--<article-title-html>Nine time steps: ultra-fast statistical consistency testing of the Community Earth System Model (pyCECT v3.0)</article-title-html>
<abstract-html><p>The Community Earth System Model Ensemble Consistency Test (CESM-ECT) suite
was developed as an alternative to requiring bitwise identical output for
quality assurance. This objective test provides a statistical measurement of
consistency between an accepted ensemble created by small initial temperature
perturbations and a test set of CESM simulations. In this work, we extend the
CESM-ECT suite with an inexpensive and robust test for ensemble consistency
that is applied to Community Atmospheric Model (CAM) output after only nine
model time steps. We demonstrate that adequate ensemble variability is
achieved with instantaneous variable values at the ninth step, despite rapid
perturbation growth and heterogeneous variable spread. We refer to this new
test as the Ultra-Fast CAM Ensemble Consistency Test (UF-CAM-ECT) and
demonstrate its effectiveness in practice, including its ability to detect
small-scale events and its applicability to the Community Land Model (CLM).
The new ultra-fast test facilitates CESM development, porting, and
optimization efforts, particularly when used to complement information from
the original CESM-ECT suite of tools.</p></abstract-html>
<ref-html id="bib1.bib1"><label>Anderson et al.(2017)Anderson, Burns, Milroy, Ruprecht, Hauser, and
Siegel</label><mixed-citation>
Anderson, J., Burns, P. J., Milroy, D., Ruprecht, P., Hauser, T., and Siegel,
H. J.: Deploying RMACC Summit: An HPC Resource for the Rocky Mountain Region,
in: Proceedings of the Practice and Experience in Advanced Research Computing
2017 on Sustainability, Success and Impact, PEARC17, 8:1–8:7, ACM, New
York, NY, USA, <a href="https://doi.org/10.1145/3093338.3093379" target="_blank">https://doi.org/10.1145/3093338.3093379</a>,
2017.
</mixed-citation></ref-html>
<ref-html id="bib1.bib2"><label>Baker et al.(2015)Baker, Hammerling, Levy, Xu, Dennis, Eaton,
Edwards, Hannay, Mickelson, Neale, Nychka, Shollenberger, Tribbia,
Vertenstein, and Williamson</label><mixed-citation>
Baker, A. H., Hammerling, D. M., Levy, M. N., Xu, H., Dennis, J. M., Eaton,
B. E., Edwards, J., Hannay, C., Mickelson, S. A., Neale, R. B., Nychka, D.,
Shollenberger, J., Tribbia, J., Vertenstein, M., and Williamson, D.: A new
ensemble-based consistency test for the Community Earth System Model (pyCECT
v1.0), Geosci. Model Dev., 8, 2829–2840,
<a href="https://doi.org/10.5194/gmd-8-2829-2015" target="_blank">https://doi.org/10.5194/gmd-8-2829-2015</a>, 2015.
</mixed-citation></ref-html>
<ref-html id="bib1.bib3"><label>Baker et al.(2016)Baker, Hu, Hammerling, Tseng, Xu, Huang, Bryan, and
Yang</label><mixed-citation>
Baker, A. H., Hu, Y., Hammerling, D. M., Tseng, Y.-H., Xu, H., Huang, X.,
Bryan, F. O., and Yang, G.: Evaluating statistical consistency in the ocean
model component of the Community Earth System Model (pyCECT v2.0), Geosci.
Model Dev., 9, 2391–2406, <a href="https://doi.org/10.5194/gmd-9-2391-2016" target="_blank">https://doi.org/10.5194/gmd-9-2391-2016</a>, 2016.

</mixed-citation></ref-html>
<ref-html id="bib1.bib4"><label>Branstator and Teng(2010)</label><mixed-citation>
Branstator, G. and Teng, H.: Two Limits of Initial-Value Decadal Predictability
in a CGCM, J. Climate, 23, 6292–6311, <a href="https://doi.org/10.1175/2010JCLI3678.1" target="_blank">https://doi.org/10.1175/2010JCLI3678.1</a>,
2010.
</mixed-citation></ref-html>
<ref-html id="bib1.bib5"><label>Deser et al.(2012)Deser, Phillips, Bourdette, and Teng</label><mixed-citation>
Deser, C., Phillips, A., Bourdette, V., and Teng, H.: Uncertainty in climate
change projections: the role of internal variability, Clim. Dynam., 38,
527–546, <a href="https://doi.org/10.1007/s00382-010-0977-x" target="_blank">https://doi.org/10.1007/s00382-010-0977-x</a>,
2012.
</mixed-citation></ref-html>
<ref-html id="bib1.bib6"><label>Hurrell et al.(2013)Hurrell, Holland, Gent, Ghan, Kay, Kushner,
Lamarque, Large, Lawrence, Lindsay, Lipscomb, Long, Mahowald, Marsh, Neale,
Rasch, Vavrus, Vertenstein, Bader, Collins, Hack, Kiehl, and
Marshall</label><mixed-citation>
Hurrell, J., Holland, M., Gent, P., Ghan, S., Kay, J., Kushner, P., Lamarque,
J.-F., Large, W., Lawrence, D., Lindsay, K., Lipscomb, W., Long, M.,
Mahowald, N., Marsh, D., Neale, R., Rasch, P., Vavrus, S., Vertenstein, M.,
Bader, D., Collins, W., Hack, J., Kiehl, J., and Marshall, S.: The
Community Earth System Model: A Framework for Collaborative Research,
B. Am. Meteorol. Soc., 94, 1339–1360,
<a href="https://doi.org/10.1175/BAMS-D-12-00121.1" target="_blank">https://doi.org/10.1175/BAMS-D-12-00121.1</a>, 2013.
</mixed-citation></ref-html>
<ref-html id="bib1.bib7"><label>Kay et al.(2015)Kay, Deser, Phillips, Mai, Hannay, Strand, Arblaster,
Bates, Danabasoglu, Edwards, Holland, Kushner, Lamarque, Lawrence, Lindsay,
Middleton, Munoz, Neale, Oleson, Polvani, and Vertenstein</label><mixed-citation>
Kay, J. E., Deser, C., Phillips, A., Mai, A., Hannay, C., Strand, G.,
Arblaster, J. M., Bates, S. C., Danabasoglu, G., Edwards, J., Holland, M.,
Kushner, P., Lamarque, J.-F., Lawrence, D., Lindsay, K., Middleton, A.,
Munoz, E., Neale, R., Oleson, K., Polvani, L., and Vertenstein, M.: The
Community Earth System Model (CESM) Large Ensemble Project: A Community
Resource for Studying Climate Change in the Presence of Internal Climate
Variability, B. Am. Meteorol. Soc., 96, 1333–1349,
<a href="https://doi.org/10.1175/BAMS-D-13-00255.1" target="_blank">https://doi.org/10.1175/BAMS-D-13-00255.1</a>,
2015.
</mixed-citation></ref-html>
<ref-html id="bib1.bib8"><label>Milroy(2015)</label><mixed-citation>
Milroy, D. J.: Refining the composition of CESM-ECT ensembles, ProQuest
Dissertations and Theses, p. 60, available at:
<a href="http://search.proquest.com/docview/1760171657?accountid=28174" target="_blank">http://search.proquest.com/docview/1760171657?accountid=28174</a> (last
access: 10 January 2017),
2015.
</mixed-citation></ref-html>
<ref-html id="bib1.bib9"><label>Milroy et al.(2016)Milroy, Baker, Hammerling, Dennis, Mickelson, and
Jessup</label><mixed-citation>
Milroy, D. J., Baker, A. H., Hammerling, D. M., Dennis, J. M., Mickelson,
S. A., and Jessup, E. R.: Towards Characterizing the Variability of
Statistically Consistent Community Earth System Model Simulations, Procedia
Computer Science, 80, 1589–1600,
<a href="https://doi.org/10.1016/j.procs.2016.05.489" target="_blank">https://doi.org/10.1016/j.procs.2016.05.489</a>,
International Conference on Computational Science 2016, ICCS 2016,
6–8 June 2016, San Diego, California, USA, 2016.
</mixed-citation></ref-html>
<ref-html id="bib1.bib10"><label>Oleson et al.(2010)Oleson, Lawrence, B, Flanner, Kluzek, J, Levis,
Swenson, Thornton, Feddema, Heald, Lamarque, Niu, Qian, Running, Sakaguchi,
Yang, Zeng, and Zeng</label><mixed-citation>
Oleson, K. W., Lawrence, D. M., B, G., Flanner, M. G., Kluzek, E., J, P.,
Levis, S., Swenson, S. C., Thornton, E., Feddema, J., Heald, C. L., Lamarque,
J., Niu, G., Qian, T., Running, S., Sakaguchi, K., Yang, L., Zeng, X., and
Zeng, X.: Technical description of version 4.0 of the Community Land Model, NCAR Tech. Note NCAR/TN-478+STR, 257.
</mixed-citation></ref-html>
<ref-html id="bib1.bib11"><label>Rosinski and Williamson(1997)</label><mixed-citation>
Rosinski, J. M. and Williamson, D. L.: The Accumulation of Rounding Errors and
Port Validation for Global Atmospheric Models, SIAM J. Sci. Comput., 18,
552–564, <a href="https://doi.org/10.1137/S1064827594275534" target="_blank">https://doi.org/10.1137/S1064827594275534</a>,
1997.
</mixed-citation></ref-html>
<ref-html id="bib1.bib12"><label>Wan et al.(2017)Wan, Zhang, Rasch, Singh, Chen, and
Edwards</label><mixed-citation>
Wan, H., Zhang, K., Rasch, P. J., Singh, B., Chen, X., and Edwards, J.: A new
and inexpensive non-bit-for-bit solution reproducibility test based on time
step convergence (TSC1.0), Geosci. Model Dev., 10, 537–552,
<a href="https://doi.org/10.5194/gmd-10-537-2017" target="_blank">https://doi.org/10.5194/gmd-10-537-2017</a>, 2017.
</mixed-citation></ref-html>--></article>
