<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing with OASIS Tables v3.0 20080202//EN" "journalpub-oasis3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:oasis="http://docs.oasis-open.org/ns/oasis-exchange/table" xml:lang="en" dtd-version="3.0">
  <front>
    <journal-meta><journal-id journal-id-type="publisher">GMD</journal-id><journal-title-group>
    <journal-title>Geoscientific Model Development</journal-title>
    <abbrev-journal-title abbrev-type="publisher">GMD</abbrev-journal-title><abbrev-journal-title abbrev-type="nlm-ta">Geosci. Model Dev.</abbrev-journal-title>
  </journal-title-group><issn pub-type="epub">1991-9603</issn><publisher>
    <publisher-name>Copernicus Publications</publisher-name>
    <publisher-loc>Göttingen, Germany</publisher-loc>
  </publisher></journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.5194/gmd-13-2355-2020</article-id><title-group><article-title>Improving climate model coupling through a complete mesh
representation: a case study with E3SM (v1) and MOAB (v5.x)</article-title><alt-title>Improving climate model coupling through a complete mesh representation</alt-title>
      </title-group><?xmltex \runningtitle{Improving climate model coupling through a complete mesh representation}?><?xmltex \runningauthor{V. S. Mahadevan et al.}?>
      <contrib-group>
        <contrib contrib-type="author" corresp="yes">
          <name><surname>Mahadevan</surname><given-names>Vijay S.</given-names></name>
          <email>mahadevan@anl.gov</email>
        </contrib>
        <contrib contrib-type="author" corresp="no">
          <name><surname>Grindeanu</surname><given-names>Iulian</given-names></name>
          
        </contrib>
        <contrib contrib-type="author" corresp="no">
          <name><surname>Jacob</surname><given-names>Robert</given-names></name>
          
        </contrib>
        <contrib contrib-type="author" corresp="no">
          <name><surname>Sarich</surname><given-names>Jason</given-names></name>
          
        </contrib>
        <aff id="aff1"><institution>Argonne National Laboratory, 9700 S. Cass Avenue, Lemont, IL, USA</institution>
        </aff>
      </contrib-group>
      <author-notes><corresp id="corr1">Vijay S. Mahadevan (mahadevan@anl.gov)</corresp></author-notes><pub-date><day>26</day><month>May</month><year>2020</year></pub-date>
      
      <volume>13</volume>
      <issue>5</issue>
      <fpage>2355</fpage><lpage>2377</lpage>
      <history>
        <date date-type="received"><day>3</day><month>November</month><year>2018</year></date>
           <date date-type="rev-request"><day>20</day><month>November</month><year>2018</year></date>
           <date date-type="rev-recd"><day>24</day><month>April</month><year>2019</year></date>
           <date date-type="accepted"><day>13</day><month>May</month><year>2019</year></date>
      </history>
      <permissions>
        <copyright-statement>Copyright: © 2020 Vijay S. Mahadevan et al.</copyright-statement>
        <copyright-year>2020</copyright-year>
      <license license-type="open-access"><license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p></license></permissions><self-uri xlink:href="https://gmd.copernicus.org/articles/13/2355/2020/gmd-13-2355-2020.html">This article is available from https://gmd.copernicus.org/articles/13/2355/2020/gmd-13-2355-2020.html</self-uri><self-uri xlink:href="https://gmd.copernicus.org/articles/13/2355/2020/gmd-13-2355-2020.pdf">The full text article is available as a PDF file from https://gmd.copernicus.org/articles/13/2355/2020/gmd-13-2355-2020.pdf</self-uri>
      <abstract><title>Abstract</title>
    <p id="d1e104">One of the fundamental factors contributing to the spatiotemporal inaccuracy in climate modeling is the
mapping of solution field data between different discretizations and numerical grids used in the coupled component
models.
The typical climate computational workflow involves evaluation and serialization of the remapping weights during
the preprocessing step, which is then consumed by the coupled driver infrastructure during simulation to compute
field projections.
Tools like Earth System Modeling Framework (ESMF) <xref ref-type="bibr" rid="bib1.bibx28" id="paren.1"/> and TempestRemap <xref ref-type="bibr" rid="bib1.bibx73" id="paren.2"/> offer
capability to generate conservative remapping weights, while the Model Coupling Toolkit (MCT) <xref ref-type="bibr" rid="bib1.bibx38" id="paren.3"/>
that is utilized in many production climate models exposes functionality to make use of the operators to solve the
coupled problem. However, such multistep processes present several hurdles in terms of the scientific workflow and
impede research productivity. In order to overcome these limitations, we present a fully integrated infrastructure
based on the Mesh Oriented datABase (MOAB) <xref ref-type="bibr" rid="bib1.bibx66 bib1.bibx47" id="paren.4"/> library, which allows for a
complete description of the numerical grids and solution data used in each submodel.
Through a scalable advancing-front intersection algorithm, the supermesh of the source and target grids are computed,
which is then used to assemble the high-order, conservative, and monotonicity-preserving remapping weights between
discretization specifications. The Fortran-compatible interfaces in MOAB are utilized to directly link the submodels
in the Energy Exascale Earth System Model (E3SM) to enable online remapping strategies in order to simplify the
coupled workflow process.
We demonstrate the superior computational efficiency of the remapping algorithms in comparison with other state-of-the-science
tools and present strong scaling results on large-scale machines for computing remapping weights between the spectral element
atmosphere and finite volume discretizations on the polygonal ocean grids.</p>
  </abstract>
    </article-meta>
  <notes notes-type="copyrightstatement">
  
      <p id="d1e126">The submitted manuscript has been created by UChicago Argonne, LLC,
Operator of Argonne National Laboratory (“Argonne”). Argonne, a U.S.
Department of Energy Office of Science laboratory, is operated under
Contract No. DE-AC02-06CH11357. The U.S. Government retains for itself,
and others acting on its behalf, a paid-up nonexclusive, irrevocable
worldwide license in said article to reproduce, prepare derivative
works, distribute copies to the public, and perform publicly and display
publicly, by or on behalf of the Government. The Department of Energy
will provide public access to these results of federally sponsored
research in accordance with the DOE Public Access Plan
(<uri>http://energy.gov/downloads/doe-public-access-plan</uri>, last access: 16 May 2020).</p>
</notes></front>
<body>
      


<sec id="Ch1.S1" sec-type="intro">
  <label>1</label><title>Introduction</title>
      <p id="d1e140">Understanding Earth's climate evolution through robust and accurate modeling
of the intrinsically complex, coupled ocean–atmosphere–land–ice–biosphere models
requires extreme-scale computational power <xref ref-type="bibr" rid="bib1.bibx79" id="paren.5"/>.
In such coupled applications, the different component models may employ
unstructured spatial meshes that are specifically generated to resolve
problem-dependent solution variations, which introduces several challenges in
performing a consistent solution coupling. It is known that operator decomposition
and unresolved coupling errors in partitioned atmosphere and ocean model simulations <xref ref-type="bibr" rid="bib1.bibx2" id="paren.6"/>,
or physics and dynamics components<?pagebreak page2356?> of an atmosphere, can lead to large
approximation errors that cause severe numerical stability issues. In this context, one
factor contributing to the spatiotemporal accuracy is the mapping between different discretizations of the sphere used in the
components of a coupled climate model.  Accurate remapping strategies in
such multi-mesh problems are critical to preserve higher-order resolution
but are in general computationally expensive given the disparate spatial scales
across which conservative projections are calculated. Since the primal solution
or auxiliary-derived data defined on a donor physics component mesh (source model) need to
be transferred to their coupled dependent physics mesh (target model), robust
numerical algorithms are necessary to preserve discretization accuracy during
these operations <xref ref-type="bibr" rid="bib1.bibx24 bib1.bibx12" id="paren.7"/>, in addition to conservation and monotonicity
properties in the field profile.</p>
      <p id="d1e152">An important consideration is that in addition to maintaining the overall discretization
accuracy of the solution during remapping, global conservation
and sometimes local element-wise conservation for critical quantities <xref ref-type="bibr" rid="bib1.bibx33" id="paren.8"/>
need to be imposed during the workflow. Such stringent requirements on
key flux fields that couple components along boundary interfaces are necessary in
order to mitigate any numerical deviations in coupled climate simulations. Note that these
physics meshes are usually never embedded or are not linked by any trivial linear transformations, which
render existence of exact projection or interpolation operators unfeasible, even if the
same continuous geometric topology is discretized in the models.
Additionally, the unique domain decomposition used for each of
the component physics meshes complicates the communication pattern during
intra-physics transfer, since aggregation of point location requests needs to be handled
efficiently in order to reduce overheads during the remapping workflow <xref ref-type="bibr" rid="bib1.bibx54 bib1.bibx65" id="paren.9"/>.</p>
      <p id="d1e161">Adaptive block-structured cubed-sphere or unstructured refinement of icosahedral/polygonal
meshes <xref ref-type="bibr" rid="bib1.bibx64" id="paren.10"/> are often used to resolve the complex fluid dynamics behavior in
atmosphere and ocean models efficiently.
In such models, conservative, local flux-preserving remapping schemes are critically important
<xref ref-type="bibr" rid="bib1.bibx4" id="paren.11"/> to effectively reduce multi-mesh errors, especially during computation of tracer
advection such as water vapor or <inline-formula><mml:math id="M1" display="inline"><mml:mrow class="chem"><mml:msub><mml:mi mathvariant="normal">CO</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> <xref ref-type="bibr" rid="bib1.bibx39" id="paren.12"/>.
This is also an issue in atmosphere models where physics and dynamics
are computed on non-embedded grids <xref ref-type="bibr" rid="bib1.bibx13" id="paren.13"/>, and the improper spatial coupling between
these multiscale models could introduce numerical artifacts.
Hence, the availability of different consistent and accurate
remapping schemes under one flexible climate simulation framework is vital to better
understand the pros and cons of the adaptive multi-resolution choices <xref ref-type="bibr" rid="bib1.bibx57" id="paren.14"/>.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F1" specific-use="star"><?xmltex \currentcnt{1}?><label>Figure 1</label><caption><p id="d1e194">E3SM coupled climate solver: <bold>(a)</bold> current model; <bold>(b)</bold> newer MOAB-based coupler.</p></caption>
        <?xmltex \igopts{width=312.980315pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/13/2355/2020/gmd-13-2355-2020-f01.png"/>

      </fig>

<sec id="Ch1.S1.SS1">
  <label>1.1</label><title>Hub-and-spoke vs. distributed coupling workflow</title>
      <p id="d1e216">The hub-and-spoke centralized model as shown in Fig. <xref ref-type="fig" rid="Ch1.F1"/>a is used
in the current Exascale Earth System Model (E3SM; <uri>https://www.e3sm.org</uri>, last access: 16 May 2020) driver and relies on several
tools and libraries that have been developed to simplify the regridding workflow within the
climate community. Most of the current tools used in E3SM and the Community Earth System Model (CESM)
<xref ref-type="bibr" rid="bib1.bibx30" id="paren.15"/> are included in a single package called the Common Infrastructure for Modeling the
Earth (CIME; <uri>http://esmci.github.io/cime</uri>, last access: 16 May 2020),
which builds on previous couplers used in CESM <xref ref-type="bibr" rid="bib1.bibx10 bib1.bibx11" id="paren.16"/>.
These modeling tools approach the problem in a two-step computational process:</p>
      <p id="d1e233"><list list-type="order">
            <list-item>

      <p id="d1e238">Compute the projection or remapping weights for a solution field
from a source component physics to a target component physics as an “offline
process”.</p>
            </list-item>
            <list-item>

      <p id="d1e244">During runtime, the CIME coupled solver loads the remapping weights from a file and handles the
partition-aware  “communication and weight matrix” application to project coupled fields between
components.</p>
            </list-item>
          </list></p>
      <p id="d1e249">The first task in this workflow is currently accomplished through a variety of standard state-of-the-science
tools such as the Earth Science Modeling Framework (ESMF) <xref ref-type="bibr" rid="bib1.bibx28" id="paren.17"/>, Spherical Coordinate
Remapping and Interpolation Package (SCRIP) <xref ref-type="bibr" rid="bib1.bibx34" id="paren.18"/>, and TempestRemap <xref ref-type="bibr" rid="bib1.bibx73 bib1.bibx71" id="paren.19"/>.
The Model Coupling Toolkit (MCT) <xref ref-type="bibr" rid="bib1.bibx38 bib1.bibx32" id="paren.20"/> used in the CIME solver provides
data structures for the second part of the workflow.
Traditionally, the first workflow phase is executed decoupled from the simulation driver during a
preprocessing step, and hence any updates to the field discretization or the underlying mesh
resolution immediately necessitate recomputation of the remapping weight generation workflow
with updated inputs. This process flow also prohibits the component solvers from performing
any runtime spatial adaptivity, since the remapping weights have to be recomputed dynamically
after any changes in grid positions. To overcome such deficiencies, and to accelerate the
current coupling workflow, recent efforts have been undertaken to implement a fully
integrated remapping weight generation process within E3SM using a scalable infrastructure
provided by the topology, decomposition, and data-aware Mesh Oriented datABase (MOAB;
<uri>http://sigma.mcs.anl.gov/moab-library</uri>, last access: 16 May 2020)
<xref ref-type="bibr" rid="bib1.bibx66 bib1.bibx47" id="paren.21"/> and TempestRemap <xref ref-type="bibr" rid="bib1.bibx73" id="paren.22"/>
software libraries as shown in Fig. <xref ref-type="fig" rid="Ch1.F1"/>b. Note that regardless of whether a
hub-and-spoke or distributed coupling model is used to drive the simulation, a minimal layer of
driver logic is necessary to compute weighted combination of fluxes, validation metrics, and other diagnostic outputs.</p>
      <?pagebreak page2357?><p id="d1e276">The paper is organized as follows. In Sect. <xref ref-type="sec" rid="Ch1.S2"/>, we present the necessary background
and motivations to develop an online remapping workflow implementation in E3SM. Section <xref ref-type="sec" rid="Ch1.S3"/>
covers details on the scalable, mesh- and partition-aware, conservative remapping algorithmic implementation
to improve
scientific productivity of the climate scientists, and to simplify the overall computational workflow for
complex problem simulations. Then, the performance of these algorithms is first evaluated in serial for
various grid combinations, and the parallel scalability of the workflow is demonstrated on large-scale
machines in Sect. <xref ref-type="sec" rid="Ch1.S4"/>.</p>
</sec>
</sec>
<sec id="Ch1.S2">
  <label>2</label><title>Background</title>
      <p id="d1e294">Conservative remapping of nonlinearly coupled solution fields is a critical task to
ensure consistency and accuracy in climate and numerical weather prediction simulations
<xref ref-type="bibr" rid="bib1.bibx64" id="paren.23"/>. While there are various ways to compute a projection of a solution
defined on a source grid <inline-formula><mml:math id="M2" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">Ω</mml:mi><mml:mi mathvariant="normal">S</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> to a target grid <inline-formula><mml:math id="M3" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">Ω</mml:mi><mml:mi mathvariant="normal">T</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, the requirements
related to global or local conservation in the remapped solution reduce the number of
potential algorithms that can be employed for such problems.</p>
      <p id="d1e322">Depending on whether (global or local) conservation is important, and if higher-order, monotone interpolators
are required, there are several consistent algorithmic options that can be used <xref ref-type="bibr" rid="bib1.bibx12" id="paren.24"/>. All of these
different remapping schemes usually have one of these characteristic traits: non-conservative (NC),
globally conservative (GC), and locally conservative (LC). Note that strong local-conservation
prescriptions also guarantee global conservation for the remapped fields.</p>
      <p id="d1e328"><list list-type="order">
          <list-item>

      <p id="d1e333">NC/GC: solution interpolation approximations:
<list list-type="bullet"><list-item>
      <p id="d1e338">NC: (approximate or exact) nearest-neighbor
interpolation;</p></list-item><list-item>
      <p id="d1e342">NC/GC: radial basis function (RBF) <xref ref-type="bibr" rid="bib1.bibx19" id="paren.25"/>
interpolators  and patch-based least-squares reconstructions <xref ref-type="bibr" rid="bib1.bibx81 bib1.bibx18" id="paren.26"/>; and</p></list-item><list-item>
      <p id="d1e352">GC: consistent finite element (FE) interpolation (bilinear, biquadratic, etc.) with area
renormalization.</p></list-item></list></p>
          </list-item>
          <list-item>

      <p id="d1e358">LC: mass- (<inline-formula><mml:math id="M4" display="inline"><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>) and gradient-preserving (<inline-formula><mml:math id="M5" display="inline"><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>)
projections:
<list list-type="bullet"><list-item>
      <p id="d1e385">embedded finite element (FE), finite difference (FD), and finite volume (FV) meshes in adaptive
computations;</p></list-item><list-item>
      <p id="d1e389">intersection-based field integrators with consistent higher-order discretization
<xref ref-type="bibr" rid="bib1.bibx34" id="paren.27"/>; and</p></list-item><list-item>
      <p id="d1e396">constrained projections to ensure conservation <xref ref-type="bibr" rid="bib1.bibx4 bib1.bibx1" id="paren.28"/> and monotonicity
<xref ref-type="bibr" rid="bib1.bibx56" id="paren.29"/>.</p></list-item></list></p>
          </list-item>
        </list></p>
      <p id="d1e407">Typically, in climate applications, flux fields are interpolated using first-order (locally)
conservative interpolation, while other scalar fields use non-conservative but higher-order
interpolators (e.g., bilinear or biquadratic).
For scalar solutions that do not need to be conserved, consistent FE interpolation,
patch-wise reconstruction schemes <xref ref-type="bibr" rid="bib1.bibx20" id="paren.30"/>, or even nearest-neighbor
interpolation <xref ref-type="bibr" rid="bib1.bibx5" id="paren.31"/> can be performed efficiently using
Kd-tree-based search-and-locate point infrastructure.
Vector fields like velocities or wind stresses are interpolated
using these same routines by separately tackling each Cartesian-decomposed
component of the field.
However, conservative remapping of flux fields requires computation of a supermesh
<xref ref-type="bibr" rid="bib1.bibx17" id="paren.32"/>, or a global intersection mesh that can be viewed as <inline-formula><mml:math id="M6" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">Ω</mml:mi><mml:mi mathvariant="normal">S</mml:mi></mml:msub><mml:mo>⋃</mml:mo><mml:msub><mml:mi mathvariant="normal">Ω</mml:mi><mml:mi mathvariant="normal">T</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, which is then used to compute projection weights that contain additional
conservation and monotonicity constraints embedded in them.</p>
      <p id="d1e438">In general, remapping implementations have three distinct steps
to accomplish the projection of solution fields from a source to a target grid. First, the target points of interest
are identified and located in the source grid, such that the target cells are a subset
of the covering (source) mesh. Next,<?pagebreak page2358?> an intersection between this covering (source) mesh
and the target mesh is performed, in order to calculate the individual weight contribution to
each target cell, without approximations to the component field discretizations that can be defined with
arbitrary-order FV or FE basis.
Finally, application of the weight matrix yields the projection required to
conservatively transfer the data onto the target grid.</p>
      <p id="d1e441">To illustrate some key differences between some NC to GC or LC schemes,
we show a 1-D Gaussian hill solution, projected onto
a coarse grid through linear basis interpolation and weighted least-squares (<inline-formula><mml:math id="M7" display="inline"><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>) minimization,
as shown in Fig. <xref ref-type="fig" rid="Ch1.F2"/>. While the point-wise linear interpolator is
computationally efficient, and second-order accurate (Fig. <xref ref-type="fig" rid="Ch1.F2"/>a)
for smooth profiles, it does not preserve the exact area under the curve. In contrast, the
<inline-formula><mml:math id="M8" display="inline"><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> minimizer conserves the global integral area but can exhibit spurious oscillatory modes, as
shown in Fig. <xref ref-type="fig" rid="Ch1.F2"/>b, when
dealing with solutions with strong gradients (Gibbs phenomena; <xref ref-type="bibr" rid="bib1.bibx23" id="altparen.33"/>). This
demonstration confirms that even for the simple 1-D example, a conservative and monotonic projector is necessary to preserve both
stability and accuracy for repeated remapping operator applications, in order to accurately transfer
fields between grids with very different resolutions. These requirements are magnified manifold when
dealing with real-world climate simulation data.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F2" specific-use="star"><?xmltex \currentcnt{2}?><label>Figure 2</label><caption><p id="d1e478">An illustration comparing point interpolation vs. <inline-formula><mml:math id="M9" display="inline"><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> minimization; impact
on conservation and monotonicity properties.</p></caption>
        <?xmltex \igopts{width=469.470472pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/13/2355/2020/gmd-13-2355-2020-f02.png"/>

      </fig>

      <p id="d1e498">While there is a delicate balance in optimizing the computational efficiency of these operations without sacrificing the
numerical accuracy or consistency of the procedure, several researchers have implemented
algorithms that are useful for a variety of problem domains.
In recent years, the growing interest to rigorously tackle coupled multiphysics applications
has led to research efforts focused on developing new regridding algorithms. The Data Transfer Kit
(DTK; <uri>https://github.com/ORNL-CEES/DataTransferKit</uri>, last access: 16 May 2020)
<xref ref-type="bibr" rid="bib1.bibx62" id="paren.34"/> from
Oak Ridge National Labs was originally developed for nuclear engineering applications but has
been extended for other problem domains through custom adaptors for meshes. DTK is more
suited for non-conservative interpolation of scalar variables with either mesh-aware (using
consistent discretization bases) or RBF-based meshless (point-cloud) representations
<xref ref-type="bibr" rid="bib1.bibx63" id="paren.35"/> that can be extended to model transport schemes on a sphere <xref ref-type="bibr" rid="bib1.bibx19" id="paren.36"/>.
The Portage (<uri>https://github.com/laristra/portage</uri>, last access: 16 May 2020)
library <xref ref-type="bibr" rid="bib1.bibx27" id="paren.37"/> from Los Alamos National
Laboratory also provides several key capabilities that are useful for geology and geophysics modeling applications
including porous flow and seismology systems. Using advanced clipping algorithms to compute the intersection of
axis-aligned squares/cubes against faces of a triangle/tetrahedron in 2-D and 3-D, respectively, general intersections
of arbitrary convex polyhedral domains can be computed efficiently <xref ref-type="bibr" rid="bib1.bibx55" id="paren.38"/>.
Support for conservative solution transfer between grids and bound preservation (to ensure monotonicity)
<xref ref-type="bibr" rid="bib1.bibx7" id="paren.39"/> has also been recently added.
While Portage does support hybrid-level parallelism (MPI <inline-formula><mml:math id="M10" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula> OpenMP), demonstrations on large-scale
machines to compute remapping weights for climate science applications have not been pursued previously.
Based on the software package documentation, support for remapping of vector fields with conservation
constraints in DTK and Portage is not directly available for use in climate workflows. Additionally,
unavailability of native support for projection of high-order spectral element data on a sphere onto a
target mesh restricts the use of these tools for certain component models in E3SM.</p>
      <p id="d1e533">In Earth science applications, the state-of-the-science regridding tool that is often used by many
researchers is the ESMF (<uri>https://www.earthsystemcog.org/projects/esmf</uri>, last access: 16 May 2020)
library, and the set of utility tools that are distributed along with it <xref ref-type="bibr" rid="bib1.bibx8 bib1.bibx15" id="paren.40"/>, to
simplify the traditional offline–online computational workflow as described in Sect. <xref ref-type="sec" rid="Ch1.S1.SS1"/>.
ESMF is implemented in a component architecture <xref ref-type="bibr" rid="bib1.bibx80" id="paren.41"/> and
provides capabilities to generate the remapping weights for different discretization
combinations on the source and target grids in serial and parallel. ESMF provides a stand-alone tool,
ESMF_RegridWeightGen (<uri>https://www.earthsystemcog.org/projects/regridweightgen</uri>, last access: 16 May 2020), to generate offline weights
that can be consumed by climate applications such as E3SM. ESMF also exposes interfaces that enable
drivers to directly invoke the remapping algorithms in order to enable the fully online workflow as well.</p>
      <p id="d1e551">Currently, the E3SM components are integrated together in a hub-and-spoke model
(Fig. <xref ref-type="fig" rid="Ch1.F1"/>a), with the inter-model communication being
handled by the Model Coupling Toolkit (MCT) <xref ref-type="bibr" rid="bib1.bibx38 bib1.bibx32" id="paren.42"/> in CIME.
The MCT library consumes the offline weights generated with ESMF or similar tools and
provides the functionality to interface with models, decompose the field data, and apply
the remapping weights loaded from a file during the setup phase. Hence, MCT serves
to abstract the communication of data in the E3SM ecosystem. However, without the offline
remapping weight generation phase for fixed grid resolutions and model combinations, the
workflow in Fig. <xref ref-type="fig" rid="Ch1.F1"/>a is incomplete.</p>
      <p id="d1e561">Similar to the CIME-MCT driver used by E3SM, OASIS3-MCT <xref ref-type="bibr" rid="bib1.bibx75 bib1.bibx9" id="paren.43"/> is a
coupler used by many European climate models, where the interpolation weights
can be generated offline through SCRIP (included as part of OASIS3-MCT).
An option to call SCRIP in an online mode is also available.
The OASIS team has recently parallelized SCRIP to speed up its calculation time <xref ref-type="bibr" rid="bib1.bibx76" id="paren.44"/>.
OASIS3-MCT also supports application of global conservation operations after interpolation and
does not require a strict hub-and-spoke coupler.  Similar to the coupler in CIME, OASIS3-MCT
utilizes MCT to perform both the communication of fields between components and<?pagebreak page2359?> the
application of the precomputed interpolation weights in parallel.</p>
      <p id="d1e570">ESMF and SCRIP traditionally handle only cell-centered data that target finite volume
discretizations (FV-to-FV projections), with first- or second-order conservation constraints.
Hence, generating remapping weights for
atmosphere–ocean grids with a spectral element (SE) source grid definition requires
generation of an intermediate and spectrally equivalent “dual” grid, which
matches the areas of the polygons to the weight of each Gauss–Lobatto–Legendre (GLL) nodes
<xref ref-type="bibr" rid="bib1.bibx52" id="paren.45"/>. Such procedures add more steps to the offline process and can
degrade the accuracy in the remapped solution since the original spectral order is neglected
(transformation from <inline-formula><mml:math id="M11" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula> order to first order). These procedures may also introduce numerical
uncertainty in the coupled solution that could produce high solution dispersion <xref ref-type="bibr" rid="bib1.bibx74" id="paren.46"/>.</p>
      <p id="d1e586">To calculate remapping weights directly for high-order spectral element grids, E3SM uses the
TempestRemap C++ library (<uri>https://github.com/ClimateGlobalChange/tempestremap</uri>, last access: 16 May 2020) <xref ref-type="bibr" rid="bib1.bibx73" id="paren.47"/>.
TempestRemap is a uni-process tool focused on the mathematically rigorous implementations of the
remapping algorithms <xref ref-type="bibr" rid="bib1.bibx71 bib1.bibx74" id="paren.48"/> and provides higher-order conservative and
monotonicity-preserving interpolators with different discretization basis such as FV, the spectrally
equivalent continuous Galerkin FE with GLL basis (cGLL), and discontinuous Galerkin FE with GLL basis
(dGLL). This library was developed as part of the effort to fill the gap in generating consistent
remapping operators for non-FV discretizations without a need for intermediate dual meshes. Computation
of conservative interpolators between any combination of these discretizations (FV, cGLL, dGLL) and grid
definitions is supported by TempestRemap library. However, since this regridding tool can only be executed
in serial, the usage of TempestRemap prior to the work presented here has been restricted primarily to
generating the required mapping weights in the offline stage.</p>
      <p id="d1e598">Even though ESMF and OASIS3-MCT have been used in online remapping studies, weight generation as part
of a preprocessing step currently remains the preferred workflow for many production climate models.
While this decoupling provides flexibility in terms of choice of remapping tools, the data management
of the mapping files for different discretizations, field constraints, and grids can render provenance,
reproducibility, and experimentation a difficult task.  It also precludes the ability to handle moving
or dynamically adaptive meshes in coupled simulations. However, it should be noted that the shift of
the remapping computation process from a preprocessing stage in the workflow, to the simulation
stage, imposes additional onus on the users to better understand the underlying component grid
properties, their decompositions, the solution fields being transferred, and the preferred
options for computing the weights. This also raises interesting workflow modifications to ensure
verification of the online weights such that consistency, conservation, and dissipation of key
fields are within user-specified constraints. In the implementation discussed here, the online
remapping computation uses the exact same input grids and specifications like the offline
workflow, along with ability to write the weights to file, which can be used to run detailed verification studies as needed.</p>
      <p id="d1e601">There are several challenges in scalably computing the regridding operators in parallel, since
it is imperative to have both a mesh- and partition-aware data structure to handle this part of the regridding workflow.
A few climate models have begun to calculate weights online as part of their regular operation.
The ICON GCM <xref ref-type="bibr" rid="bib1.bibx78" id="paren.49"/> uses YAC <xref ref-type="bibr" rid="bib1.bibx26" id="paren.50"/> and FGOALS <xref ref-type="bibr" rid="bib1.bibx41" id="paren.51"/> uses
the C-Coupler <xref ref-type="bibr" rid="bib1.bibx43 bib1.bibx44" id="paren.52"/> framework. These
codes expose both offline and online remapping capabilities with parallel decomposition management
similar to the ongoing effort presented in the current work for E3SM. Both of these packages provide
algorithmic options to<?pagebreak page2360?> perform in-memory search-and-locate operations, interpolation of field data
between meshes with first-order conservative remapping, higher-order patch-recovery <xref ref-type="bibr" rid="bib1.bibx81" id="paren.53"/>
and RBF schemes and the NC nearest-neighbor queries. The use of non-blocking communication for field
data in these packages aligns closely with scalable strategies implemented in MCT <xref ref-type="bibr" rid="bib1.bibx32" id="paren.54"/>.
While these capabilities are used routinely in production runs for their respective models, the motivation
for the work presented here is to tackle coupled high-resolution runs on next-generation architectures
with scalable algorithms (the high-resolution E3SM coupler routinely runs on 13 000 MPI tasks), without sacrificing
numerical accuracy for all discretization descriptions (FV, cGLL, dGLL) on unstructured grids.</p>
      <p id="d1e624">In the E3SM workflow supported by CIME, the ESMF regridder understands the component
grid definitions and generates the weight matrices (offline).
The CIME driver loads these operators at runtime and places them in MCT data types,
which treat them as discrete operators to compute
the interpolation or projection of data on the target grids. Additional changes in conservation
requirements or monotonicity of the field data cannot be imposed as a runtime or post-processing step in
such a workflow. In the current work, we present a new infrastructure with scalable
algorithms implemented using the MOAB mesh library and TempestRemap package to replace the
ESMF-E3SM-MCT remapper/coupler workflow. A detailed review of the algorithmic approach used
in the MOAB-TempestRemap (MBTR) workflow, along with the software
interfaces exposed to E3SM, is presented next.</p>
</sec>
<sec id="Ch1.S3">
  <label>3</label><title>Algorithmic approach</title>
      <p id="d1e635">Efficient, conservative, and accurate multi-mesh solution transfer workflows
<xref ref-type="bibr" rid="bib1.bibx32 bib1.bibx65" id="paren.55"/> are a complex process. This is
due to the fact that in order to ensure conservation of critical quantities in a given norm,
exact cell intersections between the source and target grids have to be computed. This is
complicated in a parallel setting since the domain decompositions between the source and
target grids may not have any overlaps, making it a potentially all-to-all collective
communication problem. Hence, efficient implementations of regridding operators need to
be mesh, resolution, field, and decomposition aware in order to provide optimal performance in emerging architectures.</p>
      <p id="d1e641">Fully online remapping capability within a complex ecosystem such as E3SM requires
a flexible infrastructure to generate the projection weights. In order to fulfill
these needs, we utilize the MOAB mesh data structure combined with the TempestRemap
libraries in order to provide an in-memory remapping layer to dynamically compute
the weight matrices during the setup phase of the simulations for static source–target
grid combinations. For dynamically adaptive and moving grids, the remapping operator
can be recomputed at runtime as needed. The introduction of such a software stack allows
higher-order conservation of fields while being able to transfer and maintain field
relations in parallel, within the context of the fully decomposed mesh view. This is
an improvement to the E3SM workflow where MCT is oblivious to the underlying mesh
data structure in the component models. Having a fully mesh-aware implementation
with element connectivity and adjacency information, along with parallel ghosting
and decomposition information, also provides opportunities to implement dynamic load-balancing
algorithms to gain optimal performance on large-scale machines. Without the mesh topology,
MCT is limited to performing trivial decompositions based on global ID spaces during mesh
migration operations. YAC interpolator <xref ref-type="bibr" rid="bib1.bibx26" id="paren.56"/> and the multidimensional Common
Remapping software (CoR) in C-Coupler2 <xref ref-type="bibr" rid="bib1.bibx44" id="paren.57"/> provide similar capabilities
to perform a parallel tree-based search for point location and interpolation through various supported numerical schemes.</p>
      <p id="d1e650">MOAB is a fully distributed, compact, array-based mesh data structure, and the local
entity lists are stored in ranges along with connectivity and ownership information,
rather than explicit lists, thereby leading to a high degree of memory compression.
The memory constraints per process scale well in parallel <xref ref-type="bibr" rid="bib1.bibx65" id="paren.58"/>
and are only proportional to the number of entities in the local partition, which reduces
as the number of processes increases (strong scaling limit). This is similar to the Global
Segment Map (GSMap) in MCT, which in contrast is stored in every processor, leading to
<inline-formula><mml:math id="M12" display="inline"><mml:mrow><mml:mi>O</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mi>x</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> memory requirements. The parallel communication infrastructure in MOAB is
heavily leveraged <xref ref-type="bibr" rid="bib1.bibx67" id="paren.59"/> to utilize the scalable crystal router algorithm
<xref ref-type="bibr" rid="bib1.bibx21 bib1.bibx59" id="paren.60"/> in order to scalably communicate the covering cells to
different processors. This parallel mesh infrastructure in MOAB provides the necessary
algorithmic tools for optimally executing online remapping strategies, so that MCT in E3SM
can be replaced with a MOAB-based coupler.</p>
      <p id="d1e679">In order to illustrate the online remapping algorithm implemented with the MOAB-TempestRemap
infrastructure, we define the following terms. Let <inline-formula><mml:math id="M13" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi mathvariant="normal">c</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="normal">S</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> be the component processes
for source mesh, <inline-formula><mml:math id="M14" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi mathvariant="normal">c</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="normal">T</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> be the component processes for target mesh, and <inline-formula><mml:math id="M15" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> be the
coupler processes where the remapping operator is computed. More generally, the problem
statement can be defined as transferring a solution field <inline-formula><mml:math id="M16" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula> defined on the domain
<inline-formula><mml:math id="M17" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">Ω</mml:mi><mml:mi mathvariant="normal">S</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>  and processes <inline-formula><mml:math id="M18" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi mathvariant="normal">c</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="normal">S</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> to the domain <inline-formula><mml:math id="M19" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">Ω</mml:mi><mml:mi mathvariant="normal">T</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and
processes <inline-formula><mml:math id="M20" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi mathvariant="normal">c</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="normal">T</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> through a centralized coupler with domain information
<inline-formula><mml:math id="M21" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">Ω</mml:mi><mml:mi mathvariant="normal">S</mml:mi></mml:msub><mml:mo>⋃</mml:mo><mml:msub><mml:mi mathvariant="normal">Ω</mml:mi><mml:mi mathvariant="normal">T</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> defined on <inline-formula><mml:math id="M22" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> processes.
Such a complex online remapping workflow for projecting the field data from a source to
target mesh follows the algorithm shown in Algorithm 1.</p>
      <p id="d1e817"><?xmltex \hack{\begin{figure*}[t]}?><?xmltex \igopts{width=497.923228pt}?><inline-graphic xlink:href="https://gmd.copernicus.org/articles/13/2355/2020/gmd-13-2355-2020-g01.png"/><?xmltex \hack{\end{figure*}}?></p>
      <p id="d1e825">In the following sections, the new E3SM online remapping interface implemented with a combination of
the MOAB and TempestRemap libraries is explained. Details regarding the algorithmic aspects to compute
conservative, high-order<?pagebreak page2361?> remapping weights in parallel, without sacrificing discretization accuracy on
next-generation hardware, are presented.</p>
<sec id="Ch1.S3.SS1">
  <label>3.1</label><title>Interfacing to component models in E3SM</title>
      <p id="d1e835">Within the E3SM simulation ecosystem, there are multiple component models (atmosphere–ocean–land–ice–runoff)
that are coupled to each other. While the MCT infrastructure primarily manages the global degree-of-freedom
(DoF) partitions without a notion of the underlying mesh, the new MOAB-based coupler infrastructure provides
the ability to natively interface to the component mesh and intricately understand the field DoF data layout
associated with each model.  MOAB can recognize the difference between values on a cell center and values on a
cell edge or corner. In the current work, the MOAB mesh database has been used to create the relevant
integration abstraction for the High-Order Methods Modeling Environment (HOMME)  atmosphere model <xref ref-type="bibr" rid="bib1.bibx69 bib1.bibx68" id="paren.61"/>
(cubed-sphere SE grid) and the Model for Prediction Across Scales (MPAS) ocean model <xref ref-type="bibr" rid="bib1.bibx58 bib1.bibx53" id="paren.62"/>
(polygonal meshes with holes representing land and ice regions). Since details of the mesh are not
available at the level of the coupler interface, additional MOAB (Fortran) calls<?pagebreak page2362?> via the <monospace>iMOAB</monospace>
interface are added to HOMME and MPAS component models to describe the details of the unstructured mesh
to MOAB with explicit vertex and element connectivity information, in contrast to MCT coupler that is
oblivious to the underlying grid. The atmosphere–ocean coupling requires the largest computational
effort in the coupler (since they cover about 70 % of the coupled domain), and hence the bulk of
discussions in the current work will focus on remapping and coupling between these two component models.</p>
      <p id="d1e847">MOAB can handle the finite element zoo of elements on a sphere (triangles, quadrangles, and
polygons), making it an appropriate layer to store both the mesh layout (vertices, elements,
connectivity, adjacencies) and the parallel decomposition for the component models along with
information on shared and ghosted entities. While having a uniform partitioning methodology
across components may be advantageous for improving the efficiency of coupled climate simulations,
the parallel partition of the meshes are chosen according to the requirements in individual component
solvers. Figure <xref ref-type="fig" rid="Ch1.F3"/> shows examples of partitioned SE and MPAS meshes, visualized
through the native MOAB plugin for VisIt (<uri>https://visit.llnl.gov</uri>, last access: 16 May 2020) <xref ref-type="bibr" rid="bib1.bibx77" id="paren.63"/>.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F3"><?xmltex \currentcnt{3}?><label>Figure 3</label><caption><p id="d1e860">MOAB representation of partitioned component meshes.</p></caption>
          <?xmltex \igopts{width=241.848425pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/13/2355/2020/gmd-13-2355-2020-f03.jpg"/>

        </fig>

      <p id="d1e870">The coupled field data that are to be remapped from the source grid to the target grid also need
to be serialized as part of the MOAB mesh database in terms of an internally contiguous MOAB
data storage structure named a “tag” <xref ref-type="bibr" rid="bib1.bibx66" id="paren.64"/>.
For E3SM, we use element-based tags to store the partitioned field data that are required to be
remapped between components. Typically, the number of DoFs per element (nDoF<inline-formula><mml:math id="M23" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:math></inline-formula>) is
determined based on the underlying discretization; nDoF<inline-formula><mml:math id="M24" display="inline"><mml:mrow><mml:msub><mml:mi/><mml:mi mathvariant="normal">e</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msup><mml:mi>p</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> values in HOMME, where
<inline-formula><mml:math id="M25" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula> is the order of SE discretization,
and nDoF<inline-formula><mml:math id="M26" display="inline"><mml:mrow><mml:msub><mml:mi/><mml:mi mathvariant="normal">e</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> for the FV discretization in MPAS ocean. With this complete description
of the mesh and associated data for each component model,
MOAB contains the necessary information to proceed with the remapping workflow.<?xmltex \hack{\newpage}?></p>
</sec>
<sec id="Ch1.S3.SS2">
  <label>3.2</label><title>Migration of component mesh to coupler</title>
      <p id="d1e932">E3SM's driver supports multiple modes of partitioning the various components in the global processor space. This is
usually fine tuned based on the estimated computational load in each physics, according to the problem case definition.
A sample process-execution (PE) layout for a E3SM run on 9000 processes with atmosphere (ATM)
on 5400 and ocean (OCN) on 3600 processes is
shown in Fig. <xref ref-type="fig" rid="Ch1.F4"/>.
In the case shown in the schematic, <inline-formula><mml:math id="M27" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi mathvariant="normal">c</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="normal">ATM</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">5400</mml:mn></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M28" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi mathvariant="normal">c</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="normal">OCN</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">3600</mml:mn></mml:mrow></mml:math></inline-formula>, and <inline-formula><mml:math id="M29" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi>x</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">4800</mml:mn></mml:mrow></mml:math></inline-formula>. In such a PE layout,
the atmosphere component mesh from HOMME, distributed on <inline-formula><mml:math id="M30" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi mathvariant="normal">c</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="normal">ATM</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> (5400) processes, needs to be migrated and
redistributed on <inline-formula><mml:math id="M31" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> (4800) processes and similarly from <inline-formula><mml:math id="M32" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi mathvariant="normal">c</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="normal">OCN</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> (3600) to <inline-formula><mml:math id="M33" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> (4800) processes for the
MPAS ocean mesh.
In the hub-and-spoke coupling model, as shown in Fig. <xref ref-type="fig" rid="Ch1.F1"/>, the remapping computation is performed
only in the coupler processors. Hence, inference of a communication pattern becomes necessary to ensure scalable data
transfers between the components and the coupler. In the existing implementation, MCT handles such communication,
which is being replaced by point-to-point communication kernels in MOAB to transfer mesh and data between different
components or component–coupler PEs. Note that in a distributed coupler, source and target components can communicate
directly, without any intermediate transfers (through the coupler). Under the unified infrastructure provided by MOAB,
minimal changes are required to enable either the hub-and-spoke or the distributed coupler for E3SM runs, which offers
opportunities to minimize time to solution without any changes in spatial coupling behavior.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F4"><?xmltex \currentcnt{4}?><label>Figure 4</label><caption><p id="d1e1051">Example E3SM process execution layout for a problem case.</p></caption>
          <?xmltex \igopts{width=241.848425pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/13/2355/2020/gmd-13-2355-2020-f04.png"/>

        </fig>

      <p id="d1e1060">For illustration, let <inline-formula><mml:math id="M34" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">c</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> be the number of component process elements and <inline-formula><mml:math id="M35" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> be the number of
coupler process elements. In order to migrate the mesh and associated data from <inline-formula><mml:math id="M36" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">c</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> to <inline-formula><mml:math id="M37" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>,
we first compute a trivial partition of elements that map directly in the partition space, the
same<?pagebreak page2363?> partitioning as used in the CIME-MCT coupler. In MOAB, we have exposed parallel graph and geometric
repartitioning schemes through interfaces to
Zoltan <xref ref-type="bibr" rid="bib1.bibx14" id="paren.65"/> or ParMetis <xref ref-type="bibr" rid="bib1.bibx36" id="paren.66"/> in order to evaluate optimized migration
patterns to minimize the volume of data communicated between component and coupler. We intend to
analyze the impact of different migration schemes on the scalability of the remapping operation in
Sect. <xref ref-type="sec" rid="Ch1.S4"/>. These optimizations have the potential to minimize data movement in the
MOAB-based remapper and to make it a competitive data broker to replace the current MCT <xref ref-type="bibr" rid="bib1.bibx32" id="paren.67"/> coupler in E3SM.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F5" specific-use="star"><?xmltex \currentcnt{5}?><label>Figure 5</label><caption><p id="d1e1122">Migration strategies to
repartition from <inline-formula><mml:math id="M38" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">c</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> to <inline-formula><mml:math id="M39" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>.</p></caption>
          <?xmltex \igopts{width=441.017717pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/13/2355/2020/gmd-13-2355-2020-f05.jpg"/>

        </fig>

      <p id="d1e1153">We show an example of a decomposed ocean mesh (polygonal MPAS mesh) that is replicated in a E3SM problem case run
on two processes in Fig. <xref ref-type="fig" rid="Ch1.F5"/>. Figure <xref ref-type="fig" rid="Ch1.F5"/>a is the
original decomposed mesh on two processes <inline-formula><mml:math id="M40" display="inline"><mml:mrow><mml:mo>∈</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">c</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, while Fig. <xref ref-type="fig" rid="Ch1.F5"/>b and c
show the impact of migrating a mesh from <inline-formula><mml:math id="M41" display="inline"><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">c</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> processes to four processes <inline-formula><mml:math id="M42" display="inline"><mml:mrow><mml:mo>∈</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> with a trivial linear
partitioner and a Zoltan-based geometric online partitioner. The decomposition in Fig. <xref ref-type="fig" rid="Ch1.F5"/>b
shows that the element-ID-based linear partitioner can produce bad data locality, which may require a large number of
nearest-neighbor communications when computing a source coverage mesh. The resulting communication pattern can also make the migration
and coverage computation process non-scalable on larger core counts. In contrast, in Fig. <xref ref-type="fig" rid="Ch1.F5"/>c,
the Zoltan partitioners produce much better load-balanced decompositions with hypergraph (PHG), recursive coordinate bisection
(RCB), or recursive inertial bisection (RIB) algorithms to reduce communication overheads in the remapping workflow. In order to
better understand the impact of online decomposition strategies on the overall remapping operation, we need to better
understand the impact of the repartitioner on two communication-heavy steps:
<list list-type="order"><list-item>
      <p id="d1e1208">mesh migration from component to coupler involving communication between <inline-formula><mml:math id="M43" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi mathvariant="normal">c</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="normal">s</mml:mi><mml:mo>/</mml:mo><mml:mi mathvariant="normal">t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M44" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>; and</p></list-item><list-item>
      <p id="d1e1243">computing the coverage mesh requiring gather/scatter of source mesh elements to cover local target elements.</p></list-item></list>
In a hub-and-spoke model with online remapping, the best coupler strategy will require a simultaneous
partition optimization for all grids such that mesh migration includes constraints on geometric
coordinates of component pairs. While such extensions can be implemented within the infrastructure
presented here, the performance discussions in Sect. <xref ref-type="sec" rid="Ch1.S4"/> will only focus on the trivial
and Zoltan-based partitioners. It is also worth noting that in a distributed coupler, pair-wise migration
optimizations can be performed seamlessly using a master (target)–slave (source)
strategy to maximize partition overlaps.</p>
</sec>
<sec id="Ch1.S3.SS3">
  <label>3.3</label><title>Computing the regridding operator</title>
      <p id="d1e1257">Standard approaches to compute the intersection of two convex polygonal meshes involve the creation of a
Kd-tree <xref ref-type="bibr" rid="bib1.bibx29" id="paren.68"/> or bounding volume hierarchy (BVH)-tree data structure <xref ref-type="bibr" rid="bib1.bibx31" id="paren.69"/> to enable fast element
location of relevant target points. In general, each target point of interest is located on the source
mesh by querying the tree data structure, and the corresponding (source) element is then marked as a
contributor to the remapping weight computation of the target DoF. This process is repeated to form a
list of source elements that interact directly according to the consistent discretization basis.
TempestRemap, ESMF, and YAC use variations of this search-and-clip strategy tailored to their underlying mesh representations.</p>
<sec id="Ch1.S3.SS3.SSS1">
  <label>3.3.1</label><title>Advancing-front intersection – a linear complexity algorithm</title>
      <p id="d1e1273">The intersection algorithm used in this paper follows the ideas from <xref ref-type="bibr" rid="bib1.bibx46" id="text.70"/> and <xref ref-type="bibr" rid="bib1.bibx22" id="text.71"/>,
in which two meshes are covering the same domain. At the core is an advancing-front method that aims
to traverse through the source and target meshes to compute a union (super) mesh. First, two convex
cells from the source coverage mesh and the target meshes that intersect are identified by using an
adaptive Kd-tree search tree data structure. This process also includes determination of the seed
(the starting cell) for the advancing front in each of the partitions independently. Advancing in
both meshes using face adjacency information, incrementally all possible intersections are computed
<xref ref-type="bibr" rid="bib1.bibx6" id="paren.72"/> accurately to a user defined tolerance (default <inline-formula><mml:math id="M45" display="inline"><mml:mrow><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mi>e</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">15</mml:mn></mml:mrow></mml:math></inline-formula>) in linear time.</p>
      <p id="d1e1301">While the advancing-front algorithm is not restricted to convex cells, the intersection computation
is simpler if they are strictly convex. If concave polygons exist in the initial source or target
meshes, they are recursively decomposed into simpler convex polygons, by splitting along interior
diagonals. Note that the intersection between two convex polygons results in a strictly convex
polygon. Hence, the underlying intersection algorithm remains robust to resolve even arbitrary
non-convex meshes covering the same domain space.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F6"><?xmltex \currentcnt{6}?><label>Figure 6</label><caption><p id="d1e1306">Illustration of the advancing-front intersection algorithm.</p></caption>
            <?xmltex \igopts{width=213.395669pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/13/2355/2020/gmd-13-2355-2020-f06.jpg"/>

          </fig>

      <?pagebreak page2364?><p id="d1e1316">Figure <xref ref-type="fig" rid="Ch1.F6"/> illustrates how the algorithm advances.
In each local partition of the coupler PEs, a pair of source (blue) and target
(red) cells that intersect is found (Fig. <xref ref-type="fig" rid="Ch1.F6"/>a). Using
face adjacency queries for the source mesh, as shown in Fig. <xref ref-type="fig" rid="Ch1.F6"/>b,
all source cells  that intersect the current target cell are found.  New possible pairs
between cells adjacent to the current target cell and other source cells are added to a
local queue. After the current target cell is resolved, a pair from the queue is considered next.
Figure <xref ref-type="fig" rid="Ch1.F6"/>c shows the resolving of a second target cell, which is
intersecting here with three source cells. If both meshes are contiguous, this algorithm guarantees
to compute all possible intersections between source and target cells. Figure <xref ref-type="fig" rid="Ch1.F6"/>d
shows the color map representation of the progression in which the intersection polygons
were found, from blue (low count) towards red. This advance/progression is also illustrated
through videos in both the serial <xref ref-type="bibr" rid="bib1.bibx48" id="paren.73"/> and parallel <xref ref-type="bibr" rid="bib1.bibx49" id="paren.74"/> contexts on partitioned meshes.</p>
      <p id="d1e1336">This flooding-like advancing front needs a stable and robust methodology of intersecting
edges/segments in two cells that belong to different meshes. Any pair of segments that
intersect can appear in four different pairs of cells. A list of intersection points
is maintained on each target edge, so that the intersection points are unique. Also,
a geometric tolerance is used to merge intersection points that are close to each
other or if they are proximal to the original vertices in both meshes. Decisions
regarding whether points are inside, outside, or at the boundary of a convex enclosure
are handled separately. If necessary, more robust techniques such as adaptive
precision arithmetic procedures used in Triangle <xref ref-type="bibr" rid="bib1.bibx60" id="paren.75"/>
can be employed to resolve the fronts more accurately.
Note that the advancing-front strategy can be employed for meshes with topological
holes (e.g., ocean meshes in which the continents are excluded) without any further
modifications by using a new pair for each disconnected region in the target mesh.</p>
</sec>
<sec id="Ch1.S3.SS3.SSSx1" specific-use="unnumbered">
  <title>Note on gnomonic projection for spherical geometry</title>
      <p id="d1e1348">Meshes that appear in climate applications are often on a sphere. Cell edges are
considered to be great circle arcs. A simple gnomonic projection is used to
project the edges on one of the six planes parallel to the coordinate axis and
tangent to the sphere <xref ref-type="bibr" rid="bib1.bibx73" id="paren.76"/>. With this projection, all curvilinear
cells on the sphere are transformed to linear polygons on a gnomonic plane, which
simplifies the computation of intersection between multiple grids. Once the
intersection points and cells are computed on the gnomonic plane, these are
projected back onto the original spherical domain without approximations.
This is possible due to the fact that intersection can be computed to machine
precision as the edges become straight lines in a gnomonic plane (projected
from great circle arcs on a sphere). If curves on a sphere are not great
circle arcs (splines, for example), the intersections between those curves
have to be computed using some nonlinear iterative procedures such as Newton–Raphson
(depending on the representation of the curves).</p>
</sec>
</sec>
<sec id="Ch1.S3.SS4">
  <label>3.4</label><title>Parallel implementation considerations</title>
      <?pagebreak page2365?><p id="d1e1363">Existing infrastructure from MOAB <xref ref-type="bibr" rid="bib1.bibx66" id="paren.77"/> was used to extend the
advancing-front algorithm in parallel. The expensive intersection computation can
be carried out independently, in parallel, once we redistribute the source mesh
to envelope the target mesh areas fully, in a step we refer to as “source coverage mesh” computation.<?xmltex \hack{\newpage}?></p>
<sec id="Ch1.S3.SS4.SSS1">
  <label>3.4.1</label><title>Computation of a source coverage mesh</title>
      <p id="d1e1377">We select the target mesh as the driver for redistribution of the source mesh.
On each task, we first compute the bounding box of the local target mesh. This
information is then gathered and communicated to all coupler PEs and used for redistributing the local source mesh.
Cells that intersect the bounding boxes of other processors are sent to the
corresponding owner task using the aggregating crystal router algorithm
that is particularly efficient in performing all-to-all strategies with <inline-formula><mml:math id="M46" display="inline"><mml:mrow><mml:mi>O</mml:mi><mml:mo>(</mml:mo><mml:mi>log⁡</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mi>x</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>
complexity. This graph is computed once during the setup phase to establish point-to-point
communication patterns, which are then used to pack and send/receive mesh elements or data at runtime.</p>
      <p id="d1e1403">This workflow guarantees that the target mesh on each processor is completely
enveloped by the covering mesh repartitioned from its original source mesh decomposition,
as shown in Fig. <xref ref-type="fig" rid="Ch1.F7"/>. In other words, the covering mesh fully
encompasses and bounds the target mesh in each task. It is important to note that some
source coverage cells might be sent to multiple processors during this step, depending
on the target mesh resolution and decomposition.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F7"><?xmltex \currentcnt{7}?><label>Figure 7</label><caption><p id="d1e1410">Source coverage mesh fully covers local target mesh;  local
intersection proceeds between the source atmosphere (quadrangle) and the target ocean (polygonal) grids.</p></caption>
            <?xmltex \igopts{width=241.848425pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/13/2355/2020/gmd-13-2355-2020-f07.jpg"/>

          </fig>

      <p id="d1e1420">Once the relevant covering mesh is accumulated locally on each process, the
intersection computation can be carried out in parallel, completely independently,
using the advancing-front algorithm (Sect. <xref ref-type="sec" rid="Ch1.S3.SS3.SSS1"/>). After computation
of the local intersection polygons, the vertices on the shared edges between processes
are communicated to avoid duplication. In order to ensure consistent local conservation
constraints in the weight matrix in the parallel setting, there might be a need for
additional communication of ghost intersection elements to nearest neighbors. This extra
communication step is only required for computing interpolators for flux variables
and can generally be avoided when transferring scalar fields with non-conservative bilinear
or higher-order interpolations.  Note that this ghost exchange on the intersection mesh only
requires nearest-neighbor communications within the coupler PEs, since the communication graph has been established a priori.</p>
      <p id="d1e1425">The parallel advancing-front algorithm presented here to globally compute the intersection supermesh can
be extended to expose finer-grained parallelism using hybrid-threaded (OpenMP) programming or a task-based
execution model, where each task handles a unique front in the computation queue. Such task or
hybrid-threaded parallelism can be employed in combination with the MPI-based mesh decompositions. Using local
partitions computed with Metis <xref ref-type="bibr" rid="bib1.bibx35" id="paren.78"/>
and through standard coloring approaches, each thread or task can then
proceed to compute the intersection elements until the front collides with another and until all the
overlap elements have been computed in each process. Such a parallel hybrid algorithm has the potential
to scale well even on heterogeneous architectures and provides options to improve the computational
throughput of the regridding process <xref ref-type="bibr" rid="bib1.bibx45" id="paren.79"/>.</p>
</sec>
</sec>
<sec id="Ch1.S3.SS5">
  <label>3.5</label><title>Computation of remapping operator with TempestRemap</title>
      <p id="d1e1444">For illustration, consider a scalar field <inline-formula><mml:math id="M47" display="inline"><mml:mi mathvariant="bold-italic">U</mml:mi></mml:math></inline-formula>
discretized with standard Galerkin finite element method (FEM) on source <inline-formula><mml:math id="M48" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">Ω</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>
and target <inline-formula><mml:math id="M49" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">Ω</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> meshes with different resolutions. The projection of the scalar field on
the target grid is in general given as follows:
            <disp-formula id="Ch1.E1" content-type="numbered"><label>1</label><mml:math id="M50" display="block"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">U</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold">Ω</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:msubsup><mml:mi mathvariant="bold">Π</mml:mi><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:msub><mml:mi mathvariant="bold-italic">U</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold">Ω</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
          where <inline-formula><mml:math id="M51" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="bold">Π</mml:mi><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:math></inline-formula> is the discrete solution interpolator of <inline-formula><mml:math id="M52" display="inline"><mml:mi mathvariant="bold-italic">U</mml:mi></mml:math></inline-formula> defined on <inline-formula><mml:math id="M53" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">Ω</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> to <inline-formula><mml:math id="M54" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">Ω</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>.
This interpolator <inline-formula><mml:math id="M55" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="bold">Π</mml:mi><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:math></inline-formula> in Eq. (<xref ref-type="disp-formula" rid="Ch1.E1"/>) is often referred to as the
remapping operator, which is precomputed in the coupled climate workflows using ESMF and TempestRemap.
For embedded meshes, the remapping operator can be calculated exactly as a restriction or
prolongation from the source to target grid. However, for general unstructured meshes and
in cases where the source and target meshes are topologically different, the numerical
integration to assemble <inline-formula><mml:math id="M56" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="bold">Π</mml:mi><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:math></inline-formula> needs to be carried out on the supermesh <xref ref-type="bibr" rid="bib1.bibx71" id="paren.80"/>.
Since a unique source and target parent element exist for every intersection element belonging to the
supermesh <inline-formula><mml:math id="M57" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">Ω</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>⋃</mml:mo><mml:msub><mml:mi mathvariant="bold">Ω</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M58" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="bold">Π</mml:mi><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:math></inline-formula> is assembled as the sum of local
mass matrix contributions on the intersection elements, by using the consistent discretization basis
for the source and target field descriptions <xref ref-type="bibr" rid="bib1.bibx74" id="paren.81"/>. The intersection mesh typically
contains arbitrary convex polygons, and hence subsequent triangulation may be necessary before
evaluating the integration. This global linear operator directly couples<?pagebreak page2366?> source and target DoFs
based on the participating intersection element parents <xref ref-type="bibr" rid="bib1.bibx72" id="paren.82"/>.</p>
      <p id="d1e1633">MOAB supports point-wise FEM interpolation (bilinear and higher-order spectral) with
local or global subset normalization <xref ref-type="bibr" rid="bib1.bibx65" id="paren.83"/>, in addition to a
conservative first-order remapping scheme. However, higher-order conservative monotone
weight computations are currently unsupported natively. To fill this gap for climate
applications, and to leverage existing developments in rigorous numerical algorithms
to compute the conservative weights, interfaces to TempestRemap in MOAB were added to
scalably compute the remap operator in parallel, without sacrificing field discretization accuracy.
The MOAB interface to the E3SM component models provides access to the underlying type
and order of field discretization, along with the global partitioning for the DoF
numbering. Hence, the projection or the weight matrix can be assembled in parallel by
traversing through the intersection elements and associating the appropriate source
and target DoF parent to columns and rows, respectively. The MOAB implementation uses
a sparse matrix representation using the Eigen3 (<uri>https://eigen.tuxfamily.org</uri>, last access: 16 May 2020) library <xref ref-type="bibr" rid="bib1.bibx25" id="paren.84"/> to store the local weight matrix. Except
for the particular case of projection onto a target grid with cGLL description, the matrix
rows do not share any contributions from the same source DoFs. This implies that for FV and
dGLL target field descriptions, the application of the weight matrix does not require
global collective operations and sparse matrix vector (SpMV) applications scale ideally
(still memory bandwidth limited). In the cGLL case, we perform a reduction of the parallel
vector along the shared DoFs to accumulate contributions exactly. However, it is non-trivial
to ensure full bit-for-bit (BFB) reproducibility during such reductions, and currently, the MBTR
workflow does not support exact reproducibility. Requirements for rigorous bit-wise reproduction
for online remapping need careful implementation to enforce that the advancing-front
intersection and weight matrix is computed in the exact same global element order, in addition
to ensuring that the parallel SpMV products are reduced identically, independent of the parallel mesh decompositions.</p>
      <p id="d1e1645">It is also possible to use the transpose of the remapping operator computed between
a particular source and target component combination to project the solution back
to the original source grid. Such an operation has the advantage of preserving the
consistency and conservation metrics originally imposed in finding the remapping
operator and reduces computation cost by avoiding recomputation of the weight
matrix for the new directional pair. For example, when computing the remap
operator between atmosphere and ocean models (with holes), it is advantageous
to use the atmosphere model as the source grid, since the advancing-front seed
computation may require multiple trials if the initial front begins within a hole in the source mesh.
Given that the seed or the initial cell determination on the target mesh is
chosen at random, the corresponding intersecting cell on the source mesh found
through a linear search could be contained within a hole in the source mesh. In
such a case, a new target cell is then chosen and the source cell search is repeated.
Hence, multiple trials may be required for the advancing-front algorithm to start
propagating, depending on the mesh topology and decomposition. Note that the linear
search in the source mesh can easily be replaced with a Kd-tree data structure to
provide better computational complexity for cases where both source and target meshes have many holes.
Additionally, such transpose vector applications can also make the global coupling symmetric,
which may have favorable implications when pursuing implicit temporal integration schemes.</p>
</sec>
<sec id="Ch1.S3.SS6">
  <label>3.6</label><title>Note on MBTR remapper implementation</title>
      <p id="d1e1656">The remapping algorithms presented in the previous section are exposed through a
combination of implementations in MOAB and TempestRemap libraries. Since both
libraries are written in C++, direct inheritance of key data structures such as the
GridElements (mesh) and OfflineMap (projection weights) are available to minimize
data movement between the libraries. Additionally, Fortran codes such as E3SM
can invoke computations of the intersection mesh and the remapping weights through
specialized language-agnostic interfaces in MOAB: <monospace>iMOAB</monospace> <xref ref-type="bibr" rid="bib1.bibx47" id="paren.85"/>.
These interfaces offer the flexibility to query, manipulate, and transfer the mesh between
groups of processes that represent the component and coupler processing elements.</p>
      <p id="d1e1665">Using the <monospace>iMOAB</monospace> interfaces, the E3SM coupler can coordinate the online remapping workflow during
the setup phase of the simulation and compute the projection operators for component and
scalar or vector coupled field combinations. For each pair of coupled components, the
following sequence of steps are then executed to consistently compute the remapping
operator and transfer the solution fields in parallel.</p>
      <p id="d1e1671"><list list-type="order">
            <list-item>

      <p id="d1e1676"><monospace>iMOAB_SendMesh</monospace> and <monospace>iMOAB_ReceiveMesh</monospace>:  send the component mesh
(defined on <inline-formula><mml:math id="M59" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi mathvariant="normal">c</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="normal">l</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> processes) and receive the complete unstructured mesh copy
in the coupler processes (<inline-formula><mml:math id="M60" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>). This mesh migration undergoes an online mesh repartition
either through a trivial decomposition scheme or with advanced Zoltan algorithms (geometric or graph
partitioners).</p>
            </list-item>
            <list-item>

      <p id="d1e1714"><monospace>iMOAB_ComputeMeshIntersectionOnSphere</monospace>: the advancing-front intersection scheme
is invoked to compute the overlap mesh in the coupler processes.</p>
            </list-item>
            <list-item>

      <p id="d1e1722"><monospace>iMOAB_CoverageGraph</monospace>: update the parallel communication graph based on the
(source) coverage mesh association in each process.</p>
            </list-item>
            <list-item>

      <p id="d1e1730"><monospace>iMOAB_ComputeScalarProjectionWeights</monospace>: the remapping weight operator is
computed and assembled with discretization-specific (FV, SE) calls to<?pagebreak page2367?> TempestRemap and stored in Eigen3 SparseMatrix
object.</p>
            </list-item>
          </list></p>
      <p id="d1e1737">Once the remapping operator is serialized in memory for each coupled scalar and flux
field, this operator is then used at every time step to compute the actual projection of the data.</p>
      <p id="d1e1741"><list list-type="order">
            <list-item>

      <p id="d1e1746"><monospace>iMOAB_SendElementTag</monospace> and <monospace>iMOAB_ReceiveElementTag</monospace>: using the coverage
graph computed previously, direct one-to-one communication of the field data is enabled between
<inline-formula><mml:math id="M61" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi mathvariant="normal">c</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="normal">l</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M62" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, before and after application of the weight
operator.</p>
            </list-item>
            <list-item>

      <p id="d1e1784"><monospace>iMOAB_ApplyScalarProjectionWeights</monospace>: in order to compute the field
interpolation or projection from the source component to the target component, a
matvec product of the weight matrix and the field vector defined on the source grid
is performed. The source field vector is received from source processes <inline-formula><mml:math id="M63" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi mathvariant="normal">c</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="normal">s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>
and after weight application, the target field vector is sent to target processes
<inline-formula><mml:math id="M64" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi mathvariant="normal">c</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="normal">l</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>.</p>
            </list-item>
          </list></p>
      <p id="d1e1823">Additionally, to facilitate offline generation of projection weights, a MOAB-based parallel
tool <monospace>mbtempest</monospace> has been written in C++, similar to ESMF and TempestRemap (serial)
stand-alone tools. <monospace>mbtempest</monospace> can load the source and target meshes from files, in
parallel, and compute the intersection and remapping weights through TempestRemap. The weights
can then be written back to a SCRIP-compatible file format, for any of the supported field
discretization combinations in source and destination components. Added capability to apply
the weight matrix onto the source solution field vectors, and native visualization plugins
in VisIt for MOAB, simplify the verification of conservation and monotonicity for complex
remapping workflows. This workflow allows users to validate the underlying assumptions for
remapping solution fields across unstructured grids and can be executed in both a serial and a parallel setting.</p>
</sec>
</sec>
<sec id="Ch1.S4">
  <label>4</label><title>Results</title>
      <p id="d1e1842">Evaluating the performance of the in-memory MBTR remapping infrastructure
requires recursive profiling and optimization to ensure scalability for large-scale simulations.
In order to showcase the advantage of using the mesh-aware MOAB data structure as the MCT coupler
replacement, we need to understand the per task performance of the regridder in addition to the
parallel point locator scalability and overall time for remapping weight computation. Note that
except for the weight application for each solution field from a source grid to a target grid, the
in-memory copy of the component meshes, migration to coupler PEs, computation of intersection
elements, and remapping weights is done only once during the setup phase in E3SM, per coupled component model pair.</p>
<sec id="Ch1.S4.SS1">
  <label>4.1</label><title>Serial performance</title>
      <p id="d1e1852">We compare the total cost for computing the supermesh and the remapping weights for
several source and target grid combinations through three different methods to determine the serial computational complexity.</p>
      <p id="d1e1855"><list list-type="order">
            <list-item>

      <p id="d1e1860">ESMF: Kd-tree-based regridder and weight generation for first-/second-order
FV<inline-formula><mml:math id="M65" display="inline"><mml:mo>→</mml:mo></mml:math></inline-formula>FV conservative remapping.</p>
            </list-item>
            <list-item>

      <p id="d1e1873">TempestRemap: Kd-tree-based supermesh generation and conservative, monotonic,
high-order remap operator for FV<inline-formula><mml:math id="M66" display="inline"><mml:mo>→</mml:mo></mml:math></inline-formula>FV, SE<inline-formula><mml:math id="M67" display="inline"><mml:mo>→</mml:mo></mml:math></inline-formula>FV, SE<inline-formula><mml:math id="M68" display="inline"><mml:mo>→</mml:mo></mml:math></inline-formula>SE
projection.</p>
            </list-item>
            <list-item>

      <p id="d1e1900">MBTempest: advancing-front intersection with MOAB and conservative weight
generation with TempestRemap interfaces.</p>
            </list-item>
          </list></p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F8" specific-use="star"><?xmltex \currentcnt{8}?><label>Figure 8</label><caption><p id="d1e1907">Comparison of serial regridding computation (supermesh and projection weight
generation) between ESMF, TempestRemap, and MBTempest.</p></caption>
          <?xmltex \igopts{width=384.112205pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/13/2355/2020/gmd-13-2355-2020-f08.png"/>

        </fig>

      <p id="d1e1917">Figure <xref ref-type="fig" rid="Ch1.F8"/> shows the serial performance
of the remappers for computing the conservative interpolator from cubed-sphere (CS)
grids to polygonal MPAS grids of different resolutions for a FV<inline-formula><mml:math id="M69" display="inline"><mml:mo>→</mml:mo></mml:math></inline-formula>FV field
transfer. This total time includes the computation of intersection mesh or supermesh,
in addition to the remapping weights with field conservation specifications. These
serial runs were executed on a machine with 8<inline-formula><mml:math id="M70" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> Intel Xeon(R) CPU E7-4820 at 2.00 GHz
(total of 64 cores) and 1.47 TB of RAM.
As the source grid resolution increases, the advancing-front intersection with linear
complexity outperforms the Kd-tree intersection algorithms used by TempestRemap and ESMF.
The time spent in the remapping task, including the overlap mesh generation, provides an
overall metric on the single task performance when memory bandwidth or communication
concerns do not dominate in a parallel run. In this comparison with three remapping
software libraries, the total computational time in the fine-resolution limit as
<inline-formula><mml:math id="M71" display="inline"><mml:mrow><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mtext>nele(source)</mml:mtext><mml:mtext>nele(target)</mml:mtext></mml:mfrac></mml:mstyle><mml:mo>≈</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> consistently increases
(going diagonally from left to right in Fig. <xref ref-type="fig" rid="Ch1.F8"/>).
We note that the serial version of TempestRemap is comparable to ESMF and can even provide
better timing on the highly refined cases, while the MBTempest remapper consistently
outperforms both tools, with a 2<inline-formula><mml:math id="M72" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> speedup on average. The relatively better
performance in MBTempest is accomplished through the linear complexity advancing-front
algorithm, which further offers avenues to incorporate finer-grain task or thread-level
parallelism to accelerate the on-node performance on multicore and general purpose graphical processing unit (GPGPU) architectures.</p>
</sec>
<sec id="Ch1.S4.SS2">
  <label>4.2</label><title>Scalability of the MOAB Kd-tree point locator</title>
      <?pagebreak page2368?><p id="d1e1970">In addition to being able to compute the supermesh between <inline-formula><mml:math id="M73" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">Ω</mml:mi><mml:mi mathvariant="normal">S</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and
<inline-formula><mml:math id="M74" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">Ω</mml:mi><mml:mi mathvariant="normal">T</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, MOAB also offers data structures to query source elements containing
points that correspond to the target DoFs locations. This operation is critical in
evaluating bilinear and biquadratic interpolator approximations for scalar variables
when conservative projection is not required by the underlying coupled model. The
solution interpolation for the multi-mesh case involves two distinct phases.</p>
      <p id="d1e1995"><list list-type="order">
            <list-item>

      <p id="d1e2000">Setup phase: use Kd tree to build the search data structure to locate
points corresponding to vertices in the target mesh on the source mesh.</p>
            </list-item>
            <list-item>

      <p id="d1e2006">Run phase: use the elements containing the located points to compute
consistent interpolation onto target mesh vertices.</p>
            </list-item>
          </list></p>
      <p id="d1e2011">Studies were performed to evaluate the strong and weak scalability of the parallel
Kd-tree point search implementation in MOAB. The scalability results were generated
with the CIAN2  (<uri>https://github.com/tpeterka/cian2</uri>, last access: 16 May 2020) coupling mini-app <xref ref-type="bibr" rid="bib1.bibx51" id="paren.86"/>,
which links to MOAB to handle traversal of the unstructured grids and transfer of
solution fields between the grids. For this case, a series of hexahedral and
tetrahedral meshes was used to interpolate an analytical solution. By changing
the basis interpolation order, and mesh resolutions, the convergence of the
interpolator was verified to provide theoretical accuracy orders of convergence in the asymptotic fine limit.</p>
      <p id="d1e2020">The performance tests were executed on the IBM BlueGene/Q Mira at 16 MPI ranks
per node, with 2 GB RAM per MPI rank, at up to 500 K MPI processes.  The
strong scaling results and error convergence were computed with a grid size of <inline-formula><mml:math id="M75" display="inline"><mml:mrow><mml:msup><mml:mn mathvariant="normal">1024</mml:mn><mml:mn mathvariant="normal">3</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>. The solution
interpolation on varying mesh resolutions was performed by projecting
an analytical solution from a tetrahedral to a hexahedral to a tetrahedral
grid, with total number of points/rank varied between [2 K, 32 K] in the study.
Note that the total number of DoFs in this study is much larger than that in typical climate production
runs, and hence we use these experiments to showcase the strong scaling of the bilinear
interpolation operation at the high-resolution limit.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F9" specific-use="star"><?xmltex \currentcnt{9}?><label>Figure 9</label><caption><p id="d1e2037">MOAB 3-D Kd-tree point location: strong scaling on Mira (BG/Q).</p></caption>
          <?xmltex \igopts{width=412.564961pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/13/2355/2020/gmd-13-2355-2020-f09.png"/>

        </fig>

      <?pagebreak page2369?><p id="d1e2046">First, the root-mean-square (rms) error was measured in the bilinearly interpolated solution
against the analytical solution and plotted for different source and target mesh resolutions.
Figure <xref ref-type="fig" rid="Ch1.F9"/>a demonstrates that the error convergence of the interpolants
matches the expected theoretical second-order rates, and that the error constant is proportional to
ratio of source-to-target mesh resolution. Next, Fig. <xref ref-type="fig" rid="Ch1.F9"/>b shows
the strong scaling efficiency of around 50 % is achieved on a maximum of 512 K cores (66 % of Mira).
We note that the computational complexity of the Kd-tree data structure scales as <inline-formula><mml:math id="M76" display="inline"><mml:mrow><mml:mi>O</mml:mi><mml:mo>(</mml:mo><mml:mi>n</mml:mi><mml:mi>log⁡</mml:mi><mml:mo>(</mml:mo><mml:mi>n</mml:mi><mml:mo>)</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>
asymptotically, and the point location phase during initial search setup dominates the total
cost on higher core counts. This is evident in the timing breakdown for each phase shown in
Fig. <xref ref-type="fig" rid="Ch1.F9"/>c. Since the point location is performed only once during
simulation startup, while the interpolation is performed multiple times per time step during
the run, we expect the total cost of the projection for scalar variables to be amortized over
transient climate simulations with fixed grids. Further investigations with optimal BVH-tree
<xref ref-type="bibr" rid="bib1.bibx37" id="paren.87"/> or R-tree implementations for these interpolation cases could help reduce the overall cost.</p>
      <p id="d1e2080">The full 3-D point location and interpolation operations provided by MOAB are comparable
to the implementation in the common remapping component used in the C-Coupler <xref ref-type="bibr" rid="bib1.bibx42" id="paren.88"/>
and provide relatively much stronger scalability on larger core counts <xref ref-type="bibr" rid="bib1.bibx43" id="paren.89"/>
for the remapping operation. Such higher-order interpolators for multicomponent physics
variables can provide better performance in atmospheric chemistry calculations.
Additionally, as component mesh resolutions are increased to sub-kilometer regimes, the expectations from remapping libraries
such as MOAB to provide scalable search and location of points become important.
Currently, only the NC bilinear or biquadratic interpolation of scalar fields with
subset normalization <xref ref-type="bibr" rid="bib1.bibx65" id="paren.90"/> is supported directly in MOAB
(via Kd-tree point location and interpolation), and the advancing-front intersection
algorithm does not make use of these data structures. In contrast, TempestRemap
and ESMF use a Kd-tree search to not only compute the location of points but
also to evaluate the supermesh <inline-formula><mml:math id="M77" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">Ω</mml:mi><mml:mi mathvariant="normal">S</mml:mi></mml:msub><mml:mo>⋃</mml:mo><mml:msub><mml:mi mathvariant="normal">Ω</mml:mi><mml:mi mathvariant="normal">T</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>,
and hence the computational complexity for the intersection mesh determination
scales as <inline-formula><mml:math id="M78" display="inline"><mml:mrow><mml:mi>O</mml:mi><mml:mo>(</mml:mo><mml:mi>n</mml:mi><mml:mi>log⁡</mml:mi><mml:mo>(</mml:mo><mml:mi>n</mml:mi><mml:mo>)</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, in contrast to the linear complexity (<inline-formula><mml:math id="M79" display="inline"><mml:mrow><mml:mi>O</mml:mi><mml:mo>(</mml:mo><mml:mi>n</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>) of the
advancing-front intersection algorithm implemented in MOAB.</p>
</sec>
<?pagebreak page2370?><sec id="Ch1.S4.SS3">
  <label>4.3</label><title>The parallel MBTR remapping algorithm</title>
      <p id="d1e2155">The MBTR online weight generation workflow within E3SM was employed to verify and test
the projection of real simulation data generated during the coupled atmosphere–ocean model runs.
A choice was made to use the model-computed temperature on the lowest level of the atmosphere,
since the heat fluxes that nonlinearly couple the atmosphere and ocean models are directly
proportional to this interface temperature field.  By convention, the fluxes are computed
on the ocean mesh, and hence the atmosphere temperature must be interpolated onto MPAS
polygonal mesh. We use this scenario as a test case for demonstrating the strong scalability results in this section.</p>
      <p id="d1e2158">The atmosphere run with approximately 4<inline-formula><mml:math id="M80" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> grid size and 11 elements per edge on a
cubed sphere (NE11) in E3SM, and the projection of its lowest level temperature onto two
different MPAS meshes (with approximate grid size of 240 km) are shown in Fig. <xref ref-type="fig" rid="Ch1.F10"/>.
The conservative field projection from SE to FV on a mesh with holes corresponding
to land regions is given in Fig. <xref ref-type="fig" rid="Ch1.F10"/>b, where the continents are
shown in transparent shading. To contrast, we also present the remapped field on an MPAS
mesh without holes (Fig. <xref ref-type="fig" rid="Ch1.F10"/>c) to show the differences in the
remapped solutions as a function of mesh topology.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F10" specific-use="star"><?xmltex \currentcnt{10}?><label>Figure 10</label><caption><p id="d1e2178">Projection of the NE11 SE bottom atmospheric temperature field onto the MPAS ocean grid.</p></caption>
          <?xmltex \igopts{width=412.564961pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/13/2355/2020/gmd-13-2355-2020-f10.jpg"/>

        </fig>

<sec id="Ch1.S4.SS3.SSS1">
  <label>4.3.1</label><?xmltex \opttitle{Scaling comparison of conservative remappers (FV$\rightarrow$FV)}?><title>Scaling comparison of conservative remappers (FV<inline-formula><mml:math id="M81" display="inline"><mml:mo>→</mml:mo></mml:math></inline-formula>FV)</title>
      <p id="d1e2203">The strong scaling studies for computation of remapping weights to project a FV
solution field between CS grids of varying resolutions were performed on the Blues
large-scale cluster (with 16 Sandy Bridge Xeon E5-2670 2.6 GHz cores and 32 GB RAM
per node) at ANL and the Cori supercomputer at NERSC (with 64 Haswell Xeon E5-2698v3
2.3 GHz cores and 128 GB RAM per node). Figure <xref ref-type="fig" rid="Ch1.F11"/> shows
that the MBTR workflow consistently outperforms ESMF on both machines as the
number of processes used by the coupler is increased. The timing shown here
represents the total remapping time, i.e., cumulative computational time for
generating the super mesh and the (conservative) remapping weights.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F11"><?xmltex \currentcnt{11}?><label>Figure 11</label><caption><p id="d1e2210">CS (<inline-formula><mml:math id="M82" display="inline"><mml:mrow><mml:mi>E</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">614</mml:mn><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mn mathvariant="normal">400</mml:mn></mml:mrow></mml:math></inline-formula> quads) <inline-formula><mml:math id="M83" display="inline"><mml:mo>→</mml:mo></mml:math></inline-formula> CS (<inline-formula><mml:math id="M84" display="inline"><mml:mrow><mml:mi>E</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">153</mml:mn><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">600</mml:mn></mml:mrow></mml:math></inline-formula> quads) remapping (-m conserve)
on LCRC/ALCF and NERSC machines.</p></caption>
            <?xmltex \igopts{width=241.848425pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/13/2355/2020/gmd-13-2355-2020-f11.png"/>

          </fig>

      <p id="d1e2256">The relatively better scaling for MOAB on the Blues cluster is due to faster
hardware and memory bandwidth compared to the Cori machine. The strong scaling
efficiency approaches a plateau on Cori Haswell nodes as communication costs
for the coverage mesh computation start dominating the overall remapping processes,
especially in the limit of <inline-formula><mml:math id="M85" display="inline"><mml:mrow><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mtext>nele</mml:mtext><mml:mtext>process</mml:mtext></mml:mfrac></mml:mstyle><mml:mo>→</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> at large node counts.</p>
</sec>
<sec id="Ch1.S4.SS3.SSS2">
  <label>4.3.2</label><?xmltex \opttitle{Strong scalability of spectral projection (SE$\rightarrow$FV)}?><title>Strong scalability of spectral projection (SE<inline-formula><mml:math id="M86" display="inline"><mml:mo>→</mml:mo></mml:math></inline-formula>FV)</title>
      <p id="d1e2291">To further evaluate the characteristics of in-memory remapping computation, along with
cost of application of the weights during a transient simulation, a series of further
studies was executed on the NERSC Cori system to determine the spectral projection of
a real dataset between atmosphere and ocean components in E3SM.
The source mesh contains fourth-order spectral element temperature data defined on
Gauss–Lobatto quadrature nodes (cGLL discretization) of the CS mesh, and the
projection is performed on a MPAS polygonal mesh with holes (FV discretization).
A direct comparison to ESMF was unfeasible in this study since the traditional
workflow requires the computation of a dual mesh transformation of the spectral
grid. Hence, only timing for MBTR workflow is shown here.</p>
      <p id="d1e2294">Two specific cases were considered for this SE<inline-formula><mml:math id="M87" display="inline"><mml:mo>→</mml:mo></mml:math></inline-formula>FV strong scaling
study with conservation and monotonicity constraints.
<list list-type="order"><list-item>
      <p id="d1e2306">Case A (NE30): 1<inline-formula><mml:math id="M88" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> CS (30 edges per side) SE mesh (nele<inline-formula><mml:math id="M89" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula>5400 quads)
with <inline-formula><mml:math id="M90" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">4</mml:mn></mml:mrow></mml:math></inline-formula> to MPAS mesh (nele<inline-formula><mml:math id="M91" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula>235 160 polygons).</p></list-item><list-item>
      <p id="d1e2345">Case B (NE120): <inline-formula><mml:math id="M92" display="inline"><mml:mrow><mml:msup><mml:mn mathvariant="normal">0.25</mml:mn><mml:mo>∘</mml:mo></mml:msup></mml:mrow></mml:math></inline-formula> CS (120 edges per side) SE mesh (nele<inline-formula><mml:math id="M93" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula>86 400 quads)
with <inline-formula><mml:math id="M94" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">4</mml:mn></mml:mrow></mml:math></inline-formula> to MPAS mesh (nele<inline-formula><mml:math id="M95" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula>3 693 225 polygons).</p></list-item></list></p>
      <p id="d1e2385">The performance tests for each of these cases were launched with three different
process execution layouts for the atmosphere, ocean components, and the coupler.
<list list-type="custom"><list-item><label>a.</label>
      <p id="d1e2390">Fully colocated PE layout: <inline-formula><mml:math id="M96" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">atm</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and
<inline-formula><mml:math id="M97" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">ocn</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>.</p></list-item><list-item><label>b.</label>
      <p id="d1e2430">Disjoint-ATM model PE layout: <inline-formula><mml:math id="M98" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">atm</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mi>x</mml:mi></mml:msub><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula> and
<inline-formula><mml:math id="M99" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">ocn</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>.</p></list-item><list-item><label>c.</label>
      <p id="d1e2474">Disjoint-OCN model PE layout: <inline-formula><mml:math id="M100" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">atm</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and
<inline-formula><mml:math id="M101" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">ocn</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mi>x</mml:mi></mml:msub><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula>.</p></list-item></list></p>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T1" specific-use="star"><?xmltex \currentcnt{1}?><label>Table 1</label><caption><p id="d1e2521">Strong scaling on Cori for SE<inline-formula><mml:math id="M102" display="inline"><mml:mo>→</mml:mo></mml:math></inline-formula>FV projection with two different resolutions on a fully colocated PE layout.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="5">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right" colsep="1"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:colspec colnum="5" colname="col5" align="right"/>
     <oasis:thead>
       <oasis:row>
         <oasis:entry colname="col1">Number of</oasis:entry>
         <oasis:entry namest="col2" nameend="col3" align="center" colsep="1">Case A  </oasis:entry>
         <oasis:entry namest="col4" nameend="col5" align="center">Case B   </oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">processors</oasis:entry>
         <oasis:entry rowsep="1" namest="col2" nameend="col3" align="center" colsep="1">(NE30) </oasis:entry>
         <oasis:entry rowsep="1" namest="col4" nameend="col5" align="center">(NE120) </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">Intersection (s)</oasis:entry>
         <oasis:entry colname="col3">Compute weights (s)</oasis:entry>
         <oasis:entry colname="col4">Intersection (s)</oasis:entry>
         <oasis:entry colname="col5">Compute weights (s)</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">16</oasis:entry>
         <oasis:entry colname="col2">0.936846</oasis:entry>
         <oasis:entry colname="col3">0.64983</oasis:entry>
         <oasis:entry colname="col4">145.623</oasis:entry>
         <oasis:entry colname="col5">9.732</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">32</oasis:entry>
         <oasis:entry colname="col2">0.449022</oasis:entry>
         <oasis:entry colname="col3">0.429028</oasis:entry>
         <oasis:entry colname="col4">53.1244</oasis:entry>
         <oasis:entry colname="col5">5.78093</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">64</oasis:entry>
         <oasis:entry colname="col2">0.377767</oasis:entry>
         <oasis:entry colname="col3">0.373476</oasis:entry>
         <oasis:entry colname="col4">22.7167</oasis:entry>
         <oasis:entry colname="col5">4.92151</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">128</oasis:entry>
         <oasis:entry colname="col2">0.255154</oasis:entry>
         <oasis:entry colname="col3">0.270574</oasis:entry>
         <oasis:entry colname="col4">6.70485</oasis:entry>
         <oasis:entry colname="col5">2.79397</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">256</oasis:entry>
         <oasis:entry colname="col2">0.180136</oasis:entry>
         <oasis:entry colname="col3">0.18272</oasis:entry>
         <oasis:entry colname="col4">2.26435</oasis:entry>
         <oasis:entry colname="col5">1.71835</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">512</oasis:entry>
         <oasis:entry colname="col2">0.162388</oasis:entry>
         <oasis:entry colname="col3">0.104737</oasis:entry>
         <oasis:entry colname="col4">1.25471</oasis:entry>
         <oasis:entry colname="col5">0.928622</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">1024</oasis:entry>
         <oasis:entry colname="col2">0.203354</oasis:entry>
         <oasis:entry colname="col3">0.0932475</oasis:entry>
         <oasis:entry colname="col4">0.680122</oasis:entry>
         <oasis:entry colname="col5">0.618943</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d1e2721">A breakdown of computational time for key tasks on Cori with up to 1024
processes for both cases is tabulated in Table <xref ref-type="table" rid="Ch1.T1"/>
on a fully colocated decomposition; i.e., <inline-formula><mml:math id="M103" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">ocn</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">atm</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. It is
clear that the computation of parallel intersection mesh strong scales well for these
production cases, especially for larger mesh resolutions (Case B). For the smaller
source and target mesh resolution (Case A), we notice that the intersection time
hits a lower bound that is dominated by the computation of the coverage mesh to
enclose the target mesh in each task. It is important to stress that this
one-time setup call to compute remap operator, per component pair, is relatively much
cheaper compared to individual component and solver initializations and gets amortized
over longer transient simulations. It is also worth noting that as the I/O bandwidth
in emerging architectures is not scaling in line with the compute throughput, such
an online workflow can generally be faster than parallel I/O for reading the weights
from file at scale. The MBTR implementation is also flexible to allow loading the
weights from file directly in order to preserve the existing coupler process with MCT.
In comparison to the computation of the intersection mesh, the time to assemble the
remapping weight operator in parallel is generally<?pagebreak page2371?> smaller. Even though both of these
operations are performed only once during the setup phase of the E3SM simulation,
the weight operator computation involves several validation checks that utilize
collective MPI operations, which do destroy the embarrassingly parallel nature
of the calculation, once appropriate coverage mesh is determined in each task.</p>
      <p id="d1e2751">The component-wise breakdown for the advancing-front intersection mesh, the parallel
communication graph for sending and receiving data between component and coupler,
and finally, the remapping weight generation for the SE<inline-formula><mml:math id="M104" display="inline"><mml:mo>→</mml:mo></mml:math></inline-formula>FV setup for NE30
and NE120 cases is shown in Fig. <xref ref-type="fig" rid="Ch1.F12"/>. The cumulative time for
this remapping process is shown to scale linearly for the NE120 case, even if the parallel
efficiency decreases significantly in the NE30 case, as expected based on the results in
Table <xref ref-type="table" rid="Ch1.T1"/>. Note that the MBTR workflow provides a unique
capability to consistently and accurately compute SE<inline-formula><mml:math id="M105" display="inline"><mml:mo>→</mml:mo></mml:math></inline-formula>FV projection weights in
parallel, without any need for an external preprocessing step to compute the dual mesh
(as required by ESMF) or running the entire remapping process in serial (TempestRemap).</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F12" specific-use="star"><?xmltex \currentcnt{12}?><label>Figure 12</label><caption><p id="d1e2774">Strong scaling study for the NE30 and NE120 cases for spectral projection with Zoltan repartitioner on Cori.</p></caption>
            <?xmltex \igopts{width=384.112205pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/13/2355/2020/gmd-13-2355-2020-f12.png"/>

          </fig>

      <?xmltex \floatpos{t}?><fig id="Ch1.F13" specific-use="star"><?xmltex \currentcnt{13}?><label>Figure 13</label><caption><p id="d1e2785">Scaling of the communication kernels driven with the parallel graph computed with a
trivial redistribution and the Zoltan geometric (RCB) repartitioner for the NE120 case with
<inline-formula><mml:math id="M106" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">ocn</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M107" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">atm</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mi>x</mml:mi></mml:msub><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula> on Cori.</p></caption>
            <?xmltex \igopts{width=384.112205pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/13/2355/2020/gmd-13-2355-2020-f13.png"/>

          </fig>

      <p id="d1e2835">Another key aspect of the results shown in Fig. <xref ref-type="fig" rid="Ch1.F12"/> is the relative indifference
in performance of the algorithms to the type of PE layout used to partition the component process space.
Theoretically, we expect the fully disjoint case to perform the
worst, and a full colocated case with maximal overlap to perform the best, since the layout
directly affects the total amount of data communicated for both the mesh and field data.
However, in practice, with online repartitioning<?pagebreak page2372?> strategies exposed through Zoltan (PHG and RCB),
overall scaling of the remapping algorithm is nearly independent of the PE layout for the
simulation. This is especially evident from the timing for the coverage mesh computation
for the NE120 case for all three PE layouts.</p>
</sec>
</sec>
<sec id="Ch1.S4.SS4">
  <label>4.4</label><title>Effect of partitioning strategy</title>
      <p id="d1e2850">In order to determine the effect of partitioning strategies described in
Fig. <xref ref-type="fig" rid="Ch1.F5"/>, the NE120 case with the trivial
decomposition and Zoltan geometric partitioner (RCB) was tested in parallel.
Figure <xref ref-type="fig" rid="Ch1.F13"/> compares the two strategies for
optimizing the mesh migration from the component to coupler. These strategies
play a critical role in task mapping and data locality for the source coverage
mesh computation, in addition to determining the communication graph complexity
between the components and the coupler.
This comparison highlights that the coverage mesh cost reduces uniformly at scale,
while the trivial partitioning scheme behaves better on lower core counts as shown
in Fig. <xref ref-type="fig" rid="Ch1.F13"/>a. The communication of field data
between the atmosphere component and the coupler resulting from the partitioning
strategy is a critical operation during the transient simulation and generally
stays within network latency limits in Cori (shown in Fig.<?pagebreak page2373?> <xref ref-type="fig" rid="Ch1.F13"/>b).
Even though the communication kernel does not show ideal scaling on increasing node counts, the
relative cost of the operation should be insignificant in comparison to total time spent in
individual component solvers. Note that production climate model solvers require multiple
data fields to be remapped at every rendezvous time step, and hence the size of the packed
messages may be larger for such simulations. We also note that there is a factor of 3
increase in the communication time to send and receive data, which occurs after the 64
process counts on Cori in Fig. <xref ref-type="fig" rid="Ch1.F13"/>b. This is
an artifact of the additional communication latency due to the transition from an
intra-node (each Haswell node in Cori accommodates 64 processes) to inter-node
nearest-neighbor data transfer when using multiple nodes. In this strong scaling study, the net
data size being transferred reduces with increasing core counts, and hence the
point-to-point communication beyond 128 processes is primarily dominated by network
latency and task placement. As part of future extensions, we will further explore
task mapping strategies with the Zoltan2 <xref ref-type="bibr" rid="bib1.bibx40" id="paren.91"/> library, in addition to
online partition rebalancing to maximize geometric overlap and minimize communication time during remapping.</p>
</sec>
<sec id="Ch1.S4.SS5">
  <label>4.5</label><title>Note on application of weights</title>
      <p id="d1e2875">Generally, operations involving SpMV products are memory bandwidth limited
<xref ref-type="bibr" rid="bib1.bibx3" id="paren.92"/> and occur during the application of remapping weights operator onto the source
solution field vector, in order to compute the field projection onto the target grid.
In addition to the communication of field data shown in Fig. <xref ref-type="fig" rid="Ch1.F13"/>b,
the cost of remapping weight application in parallel (presented in Fig. <xref ref-type="fig" rid="Ch1.F14"/>)
determines the total cost of the remapping operation during runtime. Except for the case of cGLL target
discretizations, the parallel SpMV operation during the weight application does not involve any global
collective reductions. In the current E3SM and OASIS3-MCT workflow, these operations are handled by the MCT library.
In high-resolution simulations of E3SM, the total time for the remapping operation in MCT is primarily
dominated by the communication costs based on the communication graph, similar to the MBTR workflow.
However, a direct comparison of the communication kernels in these two workflows is not yet possible,
since the offline maps for MCT that are generated with ESMF use the “dual” grid, while the online
maps generated with MBTR utilize the original spectral grid with no approximations, which results
in very different communication graph and non-zero pattern in the remap weight matrices.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F14"><?xmltex \currentcnt{14}?><label>Figure 14</label><caption><p id="d1e2887">SE<inline-formula><mml:math id="M108" display="inline"><mml:mo>→</mml:mo></mml:math></inline-formula>FV remapping weight operator application for the NE120 case on Cori.</p></caption>
          <?xmltex \igopts{width=227.622047pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/13/2355/2020/gmd-13-2355-2020-f14.png"/>

        </fig>

<?xmltex \hack{\newpage}?>
</sec>
</sec>
<?pagebreak page2374?><sec id="Ch1.S5" sec-type="conclusions">
  <label>5</label><title>Conclusions</title>
      <p id="d1e2914">Understanding and controlling primary sources of
errors in a coupled system dynamically will be key to achieving
predictable and verifiable climate simulations on emerging architectures.
Traditionally, the computational workflow for coupled climate simulations has involved two distinct
steps, with an offline preprocessing phase using remapping tools to generate solution field projection
weights (ESMF, TempestRemap, SCRIP), which are then consumed by the coupler to transfer field data between the component grids.</p>
      <p id="d1e2917">The offline steps include generating grid description files and running the offline tools with the
problem-specific options. Additionally many of state-of-the-science tools such as ESMF and SCRIP
require additional steps to specially
handle interpolators from SE grids. Such workflows create bottlenecks that do not scale and
can inhibit scientific research productivity.
When experimenting with refined grids, a goal for E3SM, this tool chain has to exercised repeatedly.
Additionally, when component meshes are dynamically modified, either through mesh adaptivity or
dynamical mesh movement to track moving boundaries, the underlying remapping weights must be recomputed on the fly.</p>
      <p id="d1e2920">To overcome some of these limitations, we have presented scalable algorithms and software interfaces
to create a direct component coupling with online regridding and weight generation tools. The
remapping algorithms utilize the numerics exposed by TempestRemap and leverage the parallel
mesh handling infrastructure in MOAB to create a scalable in-memory remapping infrastructure
that can be integrated with existing coupled climate solvers. Such a methodology invalidates
the need for dual grids, preserves higher-order spectral accuracy, and locally conserves the
field data, in addition to monotonicity constraints, when transferring solutions between grids with non-matching resolutions.</p>
      <p id="d1e2923">The serial and parallel performance of the MOAB advancing-front intersection algorithm with
linear complexity (<inline-formula><mml:math id="M109" display="inline"><mml:mrow><mml:mi>O</mml:mi><mml:mo>(</mml:mo><mml:mi>n</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>) was demonstrated for a variety of source and target mesh resolution
combinations and compared with the current state-of-the-science regridding tools such as ESMF
(serial/parallel) and TempestRemap (serial) that have a <inline-formula><mml:math id="M110" display="inline"><mml:mrow><mml:mi>O</mml:mi><mml:mo>(</mml:mo><mml:mi>n</mml:mi><mml:mi>log⁡</mml:mi><mml:mo>(</mml:mo><mml:mi>n</mml:mi><mml:mo>)</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> complexity using the
Kd-tree data structure. The MOAB-TempestRemap (MBTR) software infrastructure yields a balance
of the scalable performance on emerging architectures without sacrificing discretization
accuracy for component field interpolators. There are also several optimizations in the MBTR
algorithms that can be implemented to improve finer-grained parallelism on heterogeneous
architectures and to minimize data movement with better partitioning in combination with load
rebalancing strategies. Such a software infrastructure provides a foundation to build a new
coupler to replace the current offline–online, hub-and-spoke MCT-based coupler in E3SM and
offers extensions to enable a fully distributed coupling paradigm (without the need for a centralized
coupler) to minimize computational bottlenecks in a task-based workflow.</p>
</sec>

      
      </body>
    <back><notes notes-type="codeavailability"><title>Code availability</title>

      <p id="d1e2966">Information on the availability of source code
for the algorithmic infrastructure and models featured in this paper is tabulated below.</p>

      <p id="d1e2969"><list list-type="bullet">
        <list-item>

      <p id="d1e2974">E3SM  <xref ref-type="bibr" rid="bib1.bibx16" id="paren.93"/> is under active development funded by the US
Department of Energy. E3SM version 1.1 has been publicly released under an open-source three-clause
BSD license in August 2018 and is available on GitHub (<uri>https://github.com/E3SM-Project/E3SM</uri>, last access: 16 May 2020, E3SM Project, 2018).</p>
        </list-item>
        <list-item>

      <p id="d1e2986">MOAB <xref ref-type="bibr" rid="bib1.bibx66" id="paren.94"/> is an open-source library under the
umbrella of the <xref ref-type="bibr" rid="bib1.bibx61" id="text.95"/> <xref ref-type="bibr" rid="bib1.bibx47" id="paren.96"/> and is publicly available under the Lesser GNU Public
License (v3) on BitBucket <uri>https://bitbucket.org/fathomteam/moab</uri> (last access: 16 May 2020, Tautges et al., 2004). Version 5.1.0 was
released on 7 January 2019 and available at <uri>http://ftp.mcs.anl.gov/pub/fathom/moab-5.1.0.tar.gz</uri> (last access: 16 May 2020). DOI: <ext-link xlink:href="https://doi.org/10.5281/zenodo.2584863" ext-link-type="DOI">10.5281/zenodo.2584863</ext-link> <xref ref-type="bibr" rid="bib1.bibx50" id="paren.97"/>.</p>
        </list-item>
        <list-item>

      <p id="d1e3014">TempestRemap <xref ref-type="bibr" rid="bib1.bibx71 bib1.bibx74" id="paren.98"/> source code is
available under a BSD open-source license and hosted on GitHub
(<uri>https://github.com/ClimateGlobalChange/tempestremap</uri>, last access: 16 May 2020, <xref ref-type="bibr" rid="bib1.bibx70" id="altparen.99"/>). Version 2.0.2 was
released on 19 December 2018 and available at <uri>https://github.com/ClimateGlobalChange/tempestremap/archive/v2.0.2.tar.gz</uri> (last access: 16 May 2020).</p>
        </list-item>
      </list></p>
  </notes><notes notes-type="videosupplement"><title>Video supplement</title>

      <p id="d1e3034">The video supplements for the serial and parallel advancing-front mesh intersection
algorithms to compute the supermesh (<inline-formula><mml:math id="M111" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">Ω</mml:mi><mml:mi mathvariant="normal">S</mml:mi></mml:msub><mml:mo>⋃</mml:mo><mml:msub><mml:mi mathvariant="bold">Ω</mml:mi><mml:mi mathvariant="normal">T</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) of a
source (<inline-formula><mml:math id="M112" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">Ω</mml:mi><mml:mi mathvariant="normal">S</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) and target (<inline-formula><mml:math id="M113" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">Ω</mml:mi><mml:mi mathvariant="normal">T</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) grid are demonstrated.
<list list-type="bullet"><list-item>
      <p id="d1e3079">Serial advancing-front mesh intersection:  intersection between CS and MPAS grids on a single
task is illustrated; <ext-link xlink:href="https://doi.org/10.6084/m9.figshare.7294901.v1" ext-link-type="DOI">10.6084/m9.figshare.7294901.v1</ext-link> <xref ref-type="bibr" rid="bib1.bibx48" id="paren.100"/>.</p></list-item><list-item>
      <p id="d1e3089">Parallel advancing-front mesh intersection: simultaneous parallel intersections between CS and
MPAS grids on two different tasks are illustrated side by side; <ext-link xlink:href="https://doi.org/10.6084/m9.figshare.7294919.v2" ext-link-type="DOI">10.6084/m9.figshare.7294919.v2</ext-link>
<xref ref-type="bibr" rid="bib1.bibx49" id="paren.101"/>.</p></list-item></list></p>
  </notes><notes notes-type="authorcontribution"><title>Author contributions</title>

      <p id="d1e3101">VSM and RJ wrote the paper (with comments
from IG and JS). VSM and IG designed and implemented the MOAB integration with TempestRemap
library, along with exposing the necessary infrastructure for online remapping through iMOAB
interfaces. IG and JS configured the MOAB-TempestRemap remapper within E3SM and verified weight
generation to transfer solution fields between atmosphere and ocean component models. VSM
conducted numerical verification studies and executed both the serial and parallel scalability
studies on Blues<?pagebreak page2375?> and Cori LCF machines to quantify performance characteristics of the remapping
algorithms. The broader project idea was conceived by Andy Salinger (SNL), RJ, VSM, and IG.</p>
  </notes><notes notes-type="competinginterests"><title>Competing interests</title>

      <p id="d1e3107">The authors declare that they have no conflict of
interest.</p>
  </notes><ack><title>Acknowledgements</title><p id="d1e3113">This research used resources of the Argonne Leadership Computing Facility at Argonne National Laboratory, which is supported by the Office of Science of the US Department of Energy under contract DEAC02-06CH11357 and resources of the National Energy Research Scientific Computing Center, a DOE Office of Science User Facility supported by the Office of Science of the US Department of Energy under contract no. DEAC02-05CH11231. We gratefully acknowledge the computing resources provided on Blues, a high-performance computing cluster operated by the Laboratory Computing Resource Center at Argonne National Laboratory. We would also like to thank Dr. Paul Ullrich at University of California, Davis, for several helpful discussions regarding remapping schemes and implementations in TempestRemap.</p></ack><notes notes-type="financialsupport"><title>Financial support</title>

      <p id="d1e3118">This research has been supported by the Department of Energy, Labor and Economic Growth (grant no. FWP 66368). Funding for this work was provided by the Climate Model Development and Validation – Software Modernization project, and partially by the SciDAC Coupling Approaches for Next Generation Architectures (CANGA) project, which is funded by the US Department of Energy (DOE) and Office of Science Biological and Environmental Research programs. CANGA is also funded by the DOE Office of Advanced Scientific Computing Research (ASCR).</p>
  </notes><notes notes-type="reviewstatement"><title>Review statement</title>

      <p id="d1e3124">This paper was edited by Sophie Valcke and reviewed by two anonymous referees.</p>
  </notes><ref-list>
    <title>References</title>

      <ref id="bib1.bibx1"><?xmltex \def\ref@label{{Aguerre et~al.(2017)Aguerre, Dami{\'{a}}n, Gimenez, and
Nigro}}?><label>Aguerre et al.(2017)Aguerre, Damián, Gimenez, and
Nigro</label><?label aguerre2017?><mixed-citation>
Aguerre, H. J., Damián, S. M., Gimenez, J. M., and Nigro, N. M.:
Conservative handling of arbitrary non-conformal interfaces using an
efficient supermesh, J. Comput. Phys., 335, 21–49, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx2"><?xmltex \def\ref@label{{Beljaars et~al.(2017)Beljaars, Dutra, Balsamo, and
Lemari{\'{e}}}}?><label>Beljaars et al.(2017)Beljaars, Dutra, Balsamo, and
Lemarié</label><?label beljaars_2017?><mixed-citation>Beljaars, A., Dutra, E., Balsamo, G., and Lemarié, F.: On the numerical stability of surface–atmosphere coupling in weather and climate models, Geosci. Model Dev., 10, 977–989, <ext-link xlink:href="https://doi.org/10.5194/gmd-10-977-2017" ext-link-type="DOI">10.5194/gmd-10-977-2017</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx3"><label>Bell and Garland(2009)</label><?label bell2009?><mixed-citation>
Bell, N. and Garland, M.: Implementing sparse matrix-vector multiplication on
throughput-oriented processors, in: Proceedings of the conference on high
performance computing networking, storage and analysis, p. 18, ACM, 2009.</mixed-citation></ref>
      <ref id="bib1.bibx4"><label>Berger(1987)</label><?label berger1987?><mixed-citation>
Berger, M. J.: On conservation at grid interfaces, SIAM J. Numer.
Anal., 24, 967–984, 1987.</mixed-citation></ref>
      <ref id="bib1.bibx5"><label>Blanco and Rai(2014)</label><?label blanco2014nanoflann?><mixed-citation>Blanco, J. L. and Rai, P. K.: nanoflann: a C++ header-only fork of FLANN, a
library for Nearest Neighbor (NN) wih KD-trees, available at:
<uri>https://github.com/jlblancoc/nanoflann</uri> (last access: 16 May 2020), 2014.</mixed-citation></ref>
      <ref id="bib1.bibx6"><?xmltex \def\ref@label{{B\v{r}ezina and Exner(2017)}}?><label>Březina and Exner(2017)</label><?label Brezina2017?><mixed-citation>Březina, J. and Exner, P.: Fast algorithms for intersection of non-matching
grids using Plücker coordinates, Comput. Math. Appl.,
74, 174–187, <ext-link xlink:href="https://doi.org/10.1016/j.camwa.2017.01.028" ext-link-type="DOI">10.1016/j.camwa.2017.01.028</ext-link>,
2017.</mixed-citation></ref>
      <ref id="bib1.bibx7"><label>Certik et al.(2017)Certik, Ferenbaugh, Garimella, Herring, Jean,
Malone, and Sewell</label><?label certik_2017?><mixed-citation>Certik, O., Ferenbaugh, C., Garimella, R., Herring, A., Jean, B., Malone, C.,
and Sewell, C.: A Flexible Conservative Remapping Framework for Exascale
Computing, available at:
<uri>https://permalink.lanl.gov/object/tr?what=info:lanl-repo/lareport/LA-UR-17-21749</uri> (last access: 16 May 2020),
sIAM Conference on Computational Science &amp; Engineering, February 2017.</mixed-citation></ref>
      <ref id="bib1.bibx8"><label>Collins et al.(2005)Collins, Theurich, Deluca, Suarez, Trayanov,
Balaji, Li, Yang, Hill, and Da Silva</label><?label collins2005?><mixed-citation>
Collins, N., Theurich, G., Deluca, C., Suarez, M., Trayanov, A., Balaji, V.,
Li, P., Yang, W., Hill, C., and Da Silva, A.: Design and implementation of
components in the Earth System Modeling Framework,  Int. J.
High Perform. C., 19, 341–350, 2005.</mixed-citation></ref>
      <ref id="bib1.bibx9"><label>Craig et al.(2017)Craig, Valcke, and Coquart</label><?label craig2017?><mixed-citation>Craig, A., Valcke, S., and Coquart, L.: Development and performance of a new version of the OASIS coupler, OASIS3-MCT_3.0, Geosci. Model Dev., 10, 3297–3308, <ext-link xlink:href="https://doi.org/10.5194/gmd-10-3297-2017" ext-link-type="DOI">10.5194/gmd-10-3297-2017</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx10"><label>Craig et al.(2005)Craig, Jacob, Kauffman, Bettge, Larson, Ong, Ding,
and He</label><?label craig2005?><mixed-citation>Craig, A. P., Jacob, R., Kauffman, B., Bettge, T., Larson, J., Ong, E., Ding,
C., and He, Y.: CPL6: The New Extensible, High Performance Parallel Coupler
for the Community Climate System Model, Int. J. High
Perform. C., 19, 309–327,
<ext-link xlink:href="https://doi.org/10.1177/1094342005056117" ext-link-type="DOI">10.1177/1094342005056117</ext-link>,
2005.</mixed-citation></ref>
      <ref id="bib1.bibx11"><label>Craig et al.(2012)Craig, Vertenstein, and Jacob</label><?label craig2012?><mixed-citation>Craig, A. P., Vertenstein, M., and Jacob, R.: A new flexible coupler for earth
system modeling developed for CCSM4 and CESM1, Int. J. High
Perform. C., 26, 31–42,
<ext-link xlink:href="https://doi.org/10.1177/1094342011428141" ext-link-type="DOI">10.1177/1094342011428141</ext-link>,
2012.</mixed-citation></ref>
      <ref id="bib1.bibx12"><label>de Boer et al.(2008)de Boer, van Zuijlen, and Bijl</label><?label DeBoer2008?><mixed-citation>de Boer, A., van Zuijlen, A., and Bijl, H.: Comparison of conservative and
consistent approaches for the coupling of non-matching meshes, Comput.
Method. Appl. M., 197, 4284–4297,
<ext-link xlink:href="https://doi.org/10.1016/j.cma.2008.05.001" ext-link-type="DOI">10.1016/j.cma.2008.05.001</ext-link>,
2008.</mixed-citation></ref>
      <ref id="bib1.bibx13"><label>Dennis et al.(2012)Dennis, Edwards, Evans, Guba, Lauritzen, Mirin,
St-Cyr, Taylor, and Worley</label><?label dennis_2012?><mixed-citation>
Dennis, J. M., Edwards, J., Evans, K. J., Guba, O., Lauritzen, P. H., Mirin,
A. A., St-Cyr, A., Taylor, M. A., and Worley, P. H.: CAM-SE: A scalable
spectral element dynamical core for the Community Atmosphere Model,
Int. J.  High Perform. C., 26, 74–89,
2012.</mixed-citation></ref>
      <ref id="bib1.bibx14"><label>Devine et al.(2002)Devine, Boman, Heaphy, Hendrickson, and
Vaughan</label><?label zoltan2002?><mixed-citation>
Devine, K., Boman, E., Heaphy, R., Hendrickson, B., and Vaughan, C.: Zoltan
Data Management Services for Parallel Dynamic Applications, Comput.
Sci. Eng., 4, 90–97, 2002.</mixed-citation></ref>
      <ref id="bib1.bibx15"><label>Dunlap et al.(2013)Dunlap, Rugaber, and Mark</label><?label dunlap2013?><mixed-citation>
Dunlap, R., Rugaber, S., and Mark, L.: A feature model of coupling technologies
for Earth System Models, Comput. Geosci., 53, 13–20, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx16"><label>E3SM Project(2018)</label><?label e3sm-model?><mixed-citation>E3SM Project: Energy Exascale Earth System Model (E3SM), [Computer
Software],
<ext-link xlink:href="https://doi.org/10.11578/E3SM/dc.20180418.36" ext-link-type="DOI">10.11578/E3SM/dc.20180418.36</ext-link>,
2018.</mixed-citation></ref>
      <ref id="bib1.bibx17"><label>Farrell and Maddison(2011)</label><?label farrell_2011?><mixed-citation>Farrell, P. and Maddison, J.: Conservative interpolation between volume meshes
by local Galerkin projection, Comput. Method. Appl. M., 200, 89–100,
<ext-link xlink:href="https://doi.org/10.1016/j.cma.2010.07.015" ext-link-type="DOI">10.1016/j.cma.2010.07.015</ext-link>,
2011.</mixed-citation></ref>
      <ref id="bib1.bibx18"><label>Fleishman et al.(2005)Fleishman, Cohen-Or, and Silva</label><?label fleishman2005?><mixed-citation>
Fleishman, S., Cohen-Or, D., and Silva, C. T.: Robust moving least-squares
fitting with sharp features, ACM T. Graphic, 24,
544–552, 2005.</mixed-citation></ref>
      <ref id="bib1.bibx19"><label>Flyer and Wright(2007)</label><?label flyer2007?><mixed-citation>
Flyer, N. and Wright, G. B.: Transport schemes on a sphere using radial basis
functions, J. Comput. Phys., 226, 1059–1084, 2007.</mixed-citation></ref>
      <?pagebreak page2376?><ref id="bib1.bibx20"><label>Fornberg and Piret(2008)</label><?label fornberg2008?><mixed-citation>
Fornberg, B. and Piret, C.: On choosing a radial basis function and a shape
parameter when solving a convective PDE on a sphere, J. Comput.
Phys., 227, 2758–2780, 2008.</mixed-citation></ref>
      <ref id="bib1.bibx21"><label>Fox et al.(1989)Fox, Johnson, Lyzenga, Otto, Salmon, Walker, and
White</label><?label fox1989?><mixed-citation>
Fox, G., Johnson, M., Lyzenga, G., Otto, S., Salmon, J., Walker, D., and White,
R. L.: Solving Problems On Concurrent Processors Vol. 1: General Techniques
and Regular Problems, Comput. Phys., 3, 83–84, 1989.</mixed-citation></ref>
      <ref id="bib1.bibx22"><label>Gander and Japhet(2013)</label><?label Gander2013?><mixed-citation>Gander, M. J. and Japhet, C.: Algorithm 932: PANG: software for nonmatching
grid projections in 2D and 3D with linear complexity, ACM T.
Math. Software, 40, 1–25, <ext-link xlink:href="https://doi.org/10.1145/2513109.2513115" ext-link-type="DOI">10.1145/2513109.2513115</ext-link>, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx23"><label>Gottlieb and Shu(1997)</label><?label gottlieb1997?><mixed-citation>
Gottlieb, D. and Shu, C.-W.: On the Gibbs phenomenon and its resolution, SIAM
Rev., 39, 644–668, 1997.</mixed-citation></ref>
      <ref id="bib1.bibx24"><label>Grandy(1999)</label><?label grandy_1999?><mixed-citation>
Grandy, J.: Conservative remapping and region overlays by intersecting
arbitrary polyhedra, J. Comput. Phys., 148, 433–466, 1999.</mixed-citation></ref>
      <ref id="bib1.bibx25"><label>Guennebaud et al.(2010)Guennebaud, Jacob et al.</label><?label eigen3web?><mixed-citation>Guennebaud, G., Jacob, B., et al.: Eigen v3, available at: <uri>http://eigen.tuxfamily.org</uri> (last access: 16 May 2020), 2010.</mixed-citation></ref>
      <ref id="bib1.bibx26"><label>Hanke et al.(2016)Hanke, Redler, Holfeld, and Yastremsky</label><?label yac_2016?><mixed-citation>Hanke, M., Redler, R., Holfeld, T., and Yastremsky, M.: YAC 1.2.0: new aspects for coupling software in Earth system modelling, Geosci. Model Dev., 9, 2755–2769, <ext-link xlink:href="https://doi.org/10.5194/gmd-9-2755-2016" ext-link-type="DOI">10.5194/gmd-9-2755-2016</ext-link>, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx27"><label>Herring et al.(2017)Herring, Certik, Ferenbaugh, Garimella, Jean,
Malone, and Sewell</label><?label herring_2017?><mixed-citation>Herring, A. M., Certik, O., Ferenbaugh, C. R., Garimella, R. V., Jean, B. A.,
Malone, C. M., and Sewell, C. M.: (U) Introduction to Portage, Tech. rep.,
Los Alamos National Lab.(LANL), Los Alamos, NM (United States), available at:
<uri>https://permalink.lanl.gov/object/tr?what=info:lanl-repo/lareport/LA-UR-17-20831</uri> (last access: 16 May 2020),
2017.</mixed-citation></ref>
      <ref id="bib1.bibx28"><label>Hill et al.(2004)Hill, DeLuca, Balaji, Suarez, and
Silva</label><?label esmf_hill_2004?><mixed-citation>
Hill, C., DeLuca, C., Balaji, Suarez, M., and Silva, A. D.: The architecture of
the earth system modeling framework, Comput. Sci.  Eng., 6,
18–28, 2004.</mixed-citation></ref>
      <ref id="bib1.bibx29"><label>Hunt et al.(2006)Hunt, Mark, and Stoll</label><?label kdtree_hunt2006?><mixed-citation>
Hunt, W., Mark, W. R., and Stoll, G.: Fast kd-tree construction with an
adaptive error-bounded heuristic, in: Interactive Ray Tracing 2006,
81–88,  2006.</mixed-citation></ref>
      <ref id="bib1.bibx30"><label>Hurrell et al.(2013)Hurrell, Holland, Gent, Ghan, Kay, Kushner,
Lamarque, Large, Lawrence, Lindsay et al.</label><?label hurrell2013?><mixed-citation>
Hurrell, J. W., Holland, M. M., Gent, P. R., Ghan, S., Kay, J. E., Kushner,
P. J., Lamarque, J.-F., Large, W. G., Lawrence, D., Lindsay, K.,
Lipscomb, W. H., Long, M. C., Mahowald, N., Marsh, D. R., Neale, R. B., Rasch, P., Vavrus, S., Vertenstein, M., Bader, D., Collins, W. D., Hack, J. J., Kiehl, J., and Marshall, S. L.: The
community earth system model: a framework for collaborative research,
B. Am. Meteorol. Soc., 94, 1339–1360, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx31"><label>Ize et al.(2007)Ize, Wald, and Parker</label><?label bvhtree_ize2007?><mixed-citation>
Ize, T., Wald, I., and Parker, S. G.: Asynchronous BVH construction for ray
tracing dynamic scenes on parallel multi-core architectures, in: Proceedings
of the 7th Eurographics conference on Parallel Graphics and Visualization,
Eurographics Association,  101–108, 2007.</mixed-citation></ref>
      <ref id="bib1.bibx32"><label>Jacob et al.(2005)Jacob, Larson, and Ong</label><?label jacob2005m?><mixed-citation>Jacob, R., Larson, J., and Ong, E.: M<inline-formula><mml:math id="M114" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula>N communication and parallel
interpolation in Community Climate System Model Version 3 using the model
coupling toolkit,  Int. J. High Perform. C., 19, 293–307, 2005.</mixed-citation></ref>
      <ref id="bib1.bibx33"><label>Jiao and Heath(2004)</label><?label jiao_heath_2004?><mixed-citation>
Jiao, X. and Heath, M. T.: Common-refinement-based data transfer between
non-matching meshes in multiphysics simulations, Int. J.
Numer. Meth. Eng., 61, 2402–2427, 2004.</mixed-citation></ref>
      <ref id="bib1.bibx34"><label>Jones(1999)</label><?label Jones1999?><mixed-citation>
Jones, P. W.: First-and second-order conservative remapping schemes for grids
in spherical coordinates, Mon. Weather Rev., 127, 2204–2210, 1999.</mixed-citation></ref>
      <ref id="bib1.bibx35"><label>Karypis and Kumar(1998)</label><?label karypis1998fast?><mixed-citation>
Karypis, G. and Kumar, V.: A fast and high quality multilevel scheme for partitioning irregular graphs, SIAM J. Sci. Comput., 20, 359–392, 1998.</mixed-citation></ref>
      <ref id="bib1.bibx36"><label>Karypis et al.(1997)Karypis, Schloegel, and Kumar</label><?label parmetis1997?><mixed-citation>
Karypis, G., Schloegel, K., and Kumar, V.: Parmetis: Parallel graph
partitioning and sparse matrix ordering library, Version 1.0, Dept. of
Computer Science, University of Minnesota, 22 pp., 1997.</mixed-citation></ref>
      <ref id="bib1.bibx37"><label>Larsen et al.(1999)Larsen, Gottschalk, Lin, and Manocha</label><?label larsen1999?><mixed-citation>
Larsen, E., Gottschalk, S., Lin, M. C., and Manocha, D.: Fast proximity queries
with swept sphere volumes, Tech. rep., Technical Report TR99-018, Department
of Computer Science, University of North Carolina, 1999.</mixed-citation></ref>
      <ref id="bib1.bibx38"><label>Larson et al.(2001)Larson, Jacob, Foster, and Guo</label><?label mct_larson2001?><mixed-citation>
Larson, J. W., Jacob, R. L., Foster, I., and Guo, J.: The model coupling
toolkit, in: International Conference on Computational Science,
Springer, 185–194, 2001.</mixed-citation></ref>
      <ref id="bib1.bibx39"><label>Lauritzen et al.(2010)Lauritzen, Nair, and Ullrich</label><?label lauritzen_2010?><mixed-citation>Lauritzen, P. H., Nair, R. D., and Ullrich, P. A.: A conservative
semi-Lagrangian multi-tracer transport scheme (CSLAM) on the cubed-sphere
grid, J. Comput. Phys., 229, 1401–1424,
<ext-link xlink:href="https://doi.org/10.1016/j.jcp.2009.10.036" ext-link-type="DOI">10.1016/j.jcp.2009.10.036</ext-link>,
2010.</mixed-citation></ref>
      <ref id="bib1.bibx40"><label>Leung et al.(2014)Leung, Rajamanickam, Pedretti, Olivier, Devine,
Deveci, Catalyurek, and Bunde</label><?label zoltan2?><mixed-citation>
Leung, V. J., Rajamanickam, S., Pedretti, K., Olivier, S. L., Devine, K. D.,
Deveci, M., Catalyurek, U., and Bunde, D. P.: Zoltan2: Exploiting Geometric
Partitioning in Task Mapping for Parallel Computers, Tech. rep., Sandia
National Lab.(SNL-NM), Albuquerque, NM (United States), 2014.</mixed-citation></ref>
      <ref id="bib1.bibx41"><label>Li et al.(2013)</label><?label Li2013?><mixed-citation>Li, L., Lin, P., Yu, Y., Wang, B., Zhou, T., Liu, L., Liu, J., Bao, Q., Xu, S.,
Huang, W., Xia, K., Pu, Y., Dong, L., Shen, S., Liu, Y., Hu, N., Liu, M.,
Sun, W., Shi, X., Zheng, W., Wu, B., Song, M., Liu, H., Zhang, X., Wu, G.,
Xue, W., Huang, X., Yang, G., Song, Z., and Qiao, F.: The flexible global
ocean-atmosphere-land system model, Grid-point Version 2: FGOALS-g2, Adv.
Atmos. Sci., 30, 543–560, <ext-link xlink:href="https://doi.org/10.1007/s00376-012-2140-6" ext-link-type="DOI">10.1007/s00376-012-2140-6</ext-link>,
2013.</mixed-citation></ref>
      <ref id="bib1.bibx42"><label>Liu et al.(2013)Liu, Yang, and Wang</label><?label liu2013?><mixed-citation>Liu, L., Yang, G., and Wang, B.: CoR: a multi-dimensional common remapping
software for Earth System Models, in: The Second Workshop on Coupling
Technologies for Earth System Models (CW2013), available at:
<uri>https://wiki.cc.gatech.edu/CW2013/index.php/Program</uri> (last access: 8 May
2014), 2013.</mixed-citation></ref>
      <ref id="bib1.bibx43"><label>Liu et al.(2014)Liu, Yang, Wang, Zhang, Li, Zhang, Ji, and
Wang</label><?label ccoupler1_2014?><mixed-citation>Liu, L., Yang, G., Wang, B., Zhang, C., Li, R., Zhang, Z., Ji, Y., and Wang, L.: C-Coupler1: a Chinese community coupler for Earth system modeling, Geosci. Model Dev., 7, 2281–2302, <ext-link xlink:href="https://doi.org/10.5194/gmd-7-2281-2014" ext-link-type="DOI">10.5194/gmd-7-2281-2014</ext-link>, 2014.</mixed-citation></ref>
      <ref id="bib1.bibx44"><label>Liu et al.(2018)Liu, Zhang, Li, Wang, and Yang</label><?label ccoupler2_2018?><mixed-citation>Liu, L., Zhang, C., Li, R., Wang, B., and Yang, G.: C-Coupler2: a flexible and user-friendly community coupler for model coupling and nesting, Geosci. Model Dev., 11, 3557–3586, <ext-link xlink:href="https://doi.org/10.5194/gmd-11-3557-2018" ext-link-type="DOI">10.5194/gmd-11-3557-2018</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx45"><?xmltex \def\ref@label{{L{\"{o}}hner(2014)}}?><label>Löhner(2014)</label><?label Lohner2014?><mixed-citation>
Löhner, R.: Recent advances in parallel advancing front grid generation,
Arch. Comput. Meth. E., 21, 127–140, 2014.</mixed-citation></ref>
      <ref id="bib1.bibx46"><?xmltex \def\ref@label{{L{\"{o}}hner and Parikh(1988)}}?><label>Löhner and Parikh(1988)</label><?label Lohner_1988?><mixed-citation>
Löhner, R. and Parikh, P.: Generation of three-dimensional unstructured
grids by the advancing-front method, Int. J. Numer.
Meth. Fl., 8, 1135–1149, 1988.</mixed-citation></ref>
      <ref id="bib1.bibx47"><label>Mahadevan et al.(2015)Mahadevan, Grindeanu, Ray, Jain, and
Wu</label><?label sigma_v12_osti_1224985?><mixed-citation>Mahadevan, V., Grindeanu, I. R., Ray, N., Jain, R., and Wu, D.: SIGMA Release
v1.2 – Capabilities, Enhancements and Fixes, <ext-link xlink:href="https://doi.org/10.2172/1224985" ext-link-type="DOI">10.2172/1224985</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx48"><?xmltex \def\ref@label{{Mahadevan et~al.(2018{\natexlab{a}})Mahadevan, Grindeanu, Jacob, and
Sarich}}?><label>Mahadevan et al.(2018a)Mahadevan, Grindeanu, Jacob, and
Sarich</label><?label Mahadevan2018-af1?><mixed-citation>Mahadevan, V., Grindeanu, I., Jacob, R., and Sarich, J.: MOAB: Serial
Advancing Front Intersection Computation,
<ext-link xlink:href="https://doi.org/10.6084/m9.figshare.7294901.v1" ext-link-type="DOI">10.6084/m9.figshare.7294901.v1</ext-link>,
2018a.</mixed-citation></ref>
      <ref id="bib1.bibx49"><?xmltex \def\ref@label{{Mahadevan et~al.(2018{\natexlab{b}})Mahadevan, Grindeanu, Jacob, and
Sarich}}?><label>Mahadevan et al.(2018b)Mahadevan, Grindeanu, Jacob, and
Sarich</label><?label Mahadevan2018-af2?><mixed-citation>Mahadevan, V., Grindeanu, I., Jacob, R., and Sarich, J.: MOAB: Parallel
Advancing Front Mesh Intersection Algorithm,
<ext-link xlink:href="https://doi.org/10.6084/m9.figshare.7294919.v2" ext-link-type="DOI">10.6084/m9.figshare.7294919.v2</ext-link>,
2018b.</mixed-citation></ref>
      <ref id="bib1.bibx50"><label>Mahadevan er al.(2020)</label><?label Mahadevan2020?><mixed-citation>Mahadevan, V., Grindeanu, I., Jain, R., Shriwise, P., and Wilson, P.: MOAB v5.1.0, <ext-link xlink:href="https://doi.org/10.5281/zenodo.2584863" ext-link-type="DOI">10.5281/zenodo.2584863</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx51"><label>Morozov and Peterka(2016)</label><?label morozov_ldav16?><mixed-citation>
Morozov, D. and Peterka, T.: Block-Parallel Data Analysis with DIY2, 2016 IEEE 6th Symposium on Large Data Analysis and Visualization (LDAV), 29–36, 2016.</mixed-citation></ref>
      <?pagebreak page2377?><ref id="bib1.bibx52"><label>Mundt et al.(2016)Mundt, Boslough, Taylor, and
Roesler</label><?label gll-dual-mesh?><mixed-citation>
Mundt, M., Boslough, M., Taylor, M., and Roesler, E.: A Modification to the
remapping of Gauss-Lobatto nodes to the cubed sphere, Tech. rep., Technical
Report SAND2016-0830R, Sandia National Laboratory, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx53"><label>Petersen et al.(2015)Petersen, Jacobsen, Ringler, Hecht, and
Maltrud</label><?label petersen2015?><mixed-citation>
Petersen, M. R., Jacobsen, D. W., Ringler, T. D., Hecht, M. W., and Maltrud,
M. E.: Evaluation of the arbitrary Lagrangian–Eulerian vertical coordinate
method in the MPAS-Ocean model, Ocean Model., 86, 93–113, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx54"><label>Plimpton et al.(2004)Plimpton, Hendrickson, and
Stewart</label><?label Plimpton2004?><mixed-citation>
Plimpton, S. J., Hendrickson, B., and Stewart, J. R.: A parallel rendezvous
algorithm for interpolation between multiple grids, J. Parallel
Distr. Com., 64, 266–276, 2004.</mixed-citation></ref>
      <ref id="bib1.bibx55"><label>Powell and Abel(2015)</label><?label r3d_powell_2015?><mixed-citation>Powell, D. and Abel, T.: An exact general remeshing scheme applied to
physically conservative voxelization, J. Comput. Phys., 297,
340–356, <ext-link xlink:href="https://doi.org/10.1016/j.jcp.2015.05.022" ext-link-type="DOI">10.1016/j.jcp.2015.05.022</ext-link>,
2015.</mixed-citation></ref>
      <ref id="bib1.bibx56"><?xmltex \def\ref@label{{Ran{\v{c}}i{\'{c}}(1995)}}?><label>Rančić(1995)</label><?label ranvcic1995?><mixed-citation>
Rančić, M.: An efficient, conservative, monotonic remapping for
semi-Lagrangian transport algorithms, Mon. Weather Rev., 123,
1213–1217, 1995.</mixed-citation></ref>
      <ref id="bib1.bibx57"><label>Reichler and Kim(2008)</label><?label reichler_2008?><mixed-citation>
Reichler, T. and Kim, J.: How well do coupled models simulate today's climate?,
B. Am. Meteorol. Soc., 89, 303–312, 2008.</mixed-citation></ref>
      <ref id="bib1.bibx58"><label>Ringler et al.(2013)Ringler, Petersen, Higdon, Jacobsen, Jones, and
Maltrud</label><?label ringler2013?><mixed-citation>
Ringler, T., Petersen, M., Higdon, R. L., Jacobsen, D., Jones, P. W., and
Maltrud, M.: A multi-resolution approach to global ocean modeling, Ocean
Model., 69, 211–232, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx59"><label>Schliephake and Laure(2015)</label><?label schliephake-2015?><mixed-citation>
Schliephake, M. and Laure, E.: Performance Analysis of Irregular Collective
Communication with the Crystal Router Algorithm, in: Solving Software
Challenges for Exascale, edited by: Markidis, S. and Laure, E.,
Springer International Publishing, Cham, 130–140, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx60"><label>Shewchuk(1996)</label><?label shewchuk1996?><mixed-citation>
Shewchuk, J. R.: Triangle: Engineering a 2D quality mesh generator and Delaunay
triangulator, in: Applied computational geometry towards geometric
engineering,  Springer, 203–222, 1996.</mixed-citation></ref>
      <ref id="bib1.bibx61"><label>SIGMA toolkit(2014)</label><?label sigma_website?><mixed-citation>SIGMA toolkit: Scalable Interfaces for Geometry and Mesh based Applications
(SIGMA) toolkit, available at: <uri>http://sigma.mcs.anl.gov</uri> (last access: 16 May 2020),
2014.</mixed-citation></ref>
      <ref id="bib1.bibx62"><label>Slattery et al.(2013)Slattery, Wilson, and
Pawlowski</label><?label dtk_slattery_2013?><mixed-citation>
Slattery, S., Wilson, P., and Pawlowski, R.: The data transfer kit: a geometric
rendezvous-based tool for multiphysics data transfer, in: International
conference on mathematics &amp; computational methods applied to nuclear science
&amp; engineering (M&amp;C 2013),   5–9, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx63"><label>Slattery(2016)</label><?label dtk_slattery_2016?><mixed-citation>
Slattery, S. R.: Mesh-free data transfer algorithms for partitioned
multiphysics problems: Conservation, accuracy, and parallelism, J.
Comput. Phys., 307, 164–188, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx64"><label>Slingo et al.(2009)Slingo, Bates, Nikiforakis, Piggott, Roberts,
Shaffrey, Stevens, Vidale, and Weller</label><?label slingo_2009?><mixed-citation>
Slingo, J., Bates, K., Nikiforakis, N., Piggott, M., Roberts, M., Shaffrey, L.,
Stevens, I., Vidale, P. L., and Weller, H.: Developing the next-generation
climate system models: challenges and achievements, Philos.
T. R. Soc. S-A, 367, 815–831, 2009.</mixed-citation></ref>
      <ref id="bib1.bibx65"><label>Tautges and Caceres(2009)</label><?label tautges_scalable_2009?><mixed-citation>Tautges, T. J. and Caceres, A.: Scalable parallel solution coupling for
multiphysics reactor simulation, J.  Phys. Conf. Ser., 180, 012017, <ext-link xlink:href="https://doi.org/10.1088/1742-6596/180/1/012017" ext-link-type="DOI">10.1088/1742-6596/180/1/012017</ext-link>, 2009.</mixed-citation></ref>
      <ref id="bib1.bibx66"><label>Tautges et al.(2004)Tautges, Meyers, Merkley, Stimpson, and
Ernst</label><?label moab_tautges_2004?><mixed-citation>
Tautges, T. J., Meyers, R., Merkley, K., Stimpson, C., and Ernst, C.: MOAB: A
Mesh-Oriented datABase, SAND2004-1592, Sandia National Laboratories,
2004.</mixed-citation></ref>
      <ref id="bib1.bibx67"><label>Tautges et al.(2012)Tautges, Kraftcheck, Bertram, Sachdeva, and
Magerlein</label><?label tautges_jason_2012?><mixed-citation>Tautges, T. J., Kraftcheck, J., Bertram, N., Sachdeva, V., and Magerlein, J.:
Mesh Interface Resolution and Ghost Exchange in a Parallel Mesh
Representation, IEEE, Shanghai, China, 2012.
 </mixed-citation></ref><?xmltex \hack{\newpage}?>
      <ref id="bib1.bibx68"><label>Taylor et al.(2007)Taylor, Edwards, Thomas, and
Nair</label><?label homme_taylor_2007?><mixed-citation>Taylor, M., Edwards, J., Thomas, S., and Nair, R.: A mass and energy conserving
spectral element atmospheric dynamical core on the cubed-sphere grid,
J. Phys. Conf. Ser.,  78, p. 012074, <ext-link xlink:href="https://doi.org/10.1088/1742-6596/78/1/012074" ext-link-type="DOI">10.1088/1742-6596/78/1/012074</ext-link>, 2007.</mixed-citation></ref>
      <ref id="bib1.bibx69"><label>Thomas and Loft(2005)</label><?label thomas2005?><mixed-citation>
Thomas, S. J. and Loft, R. D.: The NCAR spectral element climate dynamical
core: Semi-implicit Eulerian formulation, J. Sci. Comput.,
25, 307–322, 2005.</mixed-citation></ref>
      <ref id="bib1.bibx70"><label>Ullrich and Mahadevan(2020)</label><?label UllrichMahadevan?><mixed-citation>Ullrich, P. and Mahadevan, V.: TempestRemap v2.0.2, available at:
<uri>https://github.com/ClimateGlobalChange/tempestremap</uri>, last access: 16 May 2020.</mixed-citation></ref>
      <ref id="bib1.bibx71"><label>Ullrich and Taylor(2015)</label><?label ullrich2015?><mixed-citation>
Ullrich, P. A. and Taylor, M. A.: Arbitrary-order conservative and consistent
remapping and a theory of linear maps: Part I, Mon. Weather Rev., 143,
2419–2440, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx72"><label>Ullrich et al.(2009)</label><?label ullrich2009?><mixed-citation>
Ullrich, P. A., Lauritzen, P. H., and Jablonowski, C.: Geometrically Exact
Conservative Remapping (GECoRe): Regular latitude–longitude and
cubed-sphere grids, Mon. Weather Rev., 137, 1721–1741, 2009.</mixed-citation></ref>
      <ref id="bib1.bibx73"><label>Ullrich et al.(2013)</label><?label ullrich_2013?><mixed-citation>
Ullrich, P. A., Lauritzen, P. H., and Jablonowski, C.: Some considerations for
high-order “incremental remap”-based transport schemes: edges,
reconstructions, and area integration, Int. J. Nume.
Meth. Fl., 71, 1131–1151, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx74"><label>Ullrich et al.(2016)</label><?label ullrich2016?><mixed-citation>
Ullrich, P. A., Devendran, D., and Johansen, H.: Arbitrary-order conservative
and consistent remapping and a theory of linear maps: Part II, Mon.
Weather Rev., 144, 1529–1549, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx75"><label>Valcke(2013)</label><?label valcke2013oasis3?><mixed-citation>Valcke, S.: The OASIS3 coupler: a European climate modelling community software, Geosci. Model Dev., 6, 373–388, <ext-link xlink:href="https://doi.org/10.5194/gmd-6-373-2013" ext-link-type="DOI">10.5194/gmd-6-373-2013</ext-link>, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx76"><label>Valcke et al.(2018)</label><?label oasis-mct?><mixed-citation>Valcke, S., Craig, A., and Coquart, L.: OASIS3-MCT User Guide, OASIS3-MCT4.0,
Tech. rep., CECI, Université de Toulouse, CNRS, CERFACS – TR-CMGC-18-77,
Toulouse, France,
available at: <uri>https://cerfacs.fr/wp-content/uploads/2018/07/GLOBC-TR-oasis3mct_UserGuide4.0_30062018.pdf</uri> (last access: 19 May 2020),
2018.</mixed-citation></ref>
      <ref id="bib1.bibx77"><label>VisIt(2005)</label><?label visit_2005?><mixed-citation>
VisIt: VisIt User's Guide, Tech. Rep. UCRL-SM-220449, Lawrence Livermore
National Laboratory, 2005.</mixed-citation></ref>
      <ref id="bib1.bibx78"><label>Wan et al.(2013)</label><?label Wan2013?><mixed-citation>Wan, H., Giorgetta, M. A., Zängl, G., Restelli, M., Majewski, D., Bonaventura, L., Fröhlich, K., Reinert, D., Rípodas, P., Kornblueh, L., and Förstner, J.: The ICON-1.2 hydrostatic atmospheric dynamical core on triangular grids – Part 1: Formulation and performance of the baseline version, Geosci. Model Dev., 6, 735–763, <ext-link xlink:href="https://doi.org/10.5194/gmd-6-735-2013" ext-link-type="DOI">10.5194/gmd-6-735-2013</ext-link>, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx79"><label>Washington et al.(2008)Washington, Bader, and
Collins</label><?label washington_2008?><mixed-citation>
Washington, W., Bader, D., and Collins, B. E. A.: Challenges in climate change
science and the role of computing at the extreme scale, in: Proc. of the
Workshop on Climate Science, 2008.</mixed-citation></ref>
      <ref id="bib1.bibx80"><label>Zhou(2006)</label><?label zhou2006?><mixed-citation>
Zhou, S.: Coupling climate models with the earth system modeling framework and
the common component architecture, Concurr. Comp.-Pract. E., 18, 203–213, 2006.</mixed-citation></ref>
      <ref id="bib1.bibx81"><label>Zienkiewicz and Zhu(1992)</label><?label zienkiewicz1992?><mixed-citation>
Zienkiewicz, O. C. and Zhu, J. Z.: The superconvergent patch recovery and a
posteriori error estimates. Part 1: The recovery technique, Int.
J. Numer. Meth. Eng., 33, 1331–1364, 1992.</mixed-citation></ref>

  </ref-list></back>
    <!--<article-title-html>Improving climate model coupling through a complete mesh representation: a case study with E3SM (v1) and MOAB (v5.x)</article-title-html>
<abstract-html><p>One of the fundamental factors contributing to the spatiotemporal inaccuracy in climate modeling is the
mapping of solution field data between different discretizations and numerical grids used in the coupled component
models.
The typical climate computational workflow involves evaluation and serialization of the remapping weights during
the preprocessing step, which is then consumed by the coupled driver infrastructure during simulation to compute
field projections.
Tools like Earth System Modeling Framework (ESMF) (Hill et al., 2004) and TempestRemap (Ullrich et al., 2013) offer
capability to generate conservative remapping weights, while the Model Coupling Toolkit (MCT) (Larson et al., 2001)
that is utilized in many production climate models exposes functionality to make use of the operators to solve the
coupled problem. However, such multistep processes present several hurdles in terms of the scientific workflow and
impede research productivity. In order to overcome these limitations, we present a fully integrated infrastructure
based on the Mesh Oriented datABase (MOAB) (Mahadevan et al., 2015; Tautges et al., 2004) library, which allows for a
complete description of the numerical grids and solution data used in each submodel.
Through a scalable advancing-front intersection algorithm, the supermesh of the source and target grids are computed,
which is then used to assemble the high-order, conservative, and monotonicity-preserving remapping weights between
discretization specifications. The Fortran-compatible interfaces in MOAB are utilized to directly link the submodels
in the Energy Exascale Earth System Model (E3SM) to enable online remapping strategies in order to simplify the
coupled workflow process.
We demonstrate the superior computational efficiency of the remapping algorithms in comparison with other state-of-the-science
tools and present strong scaling results on large-scale machines for computing remapping weights between the spectral element
atmosphere and finite volume discretizations on the polygonal ocean grids.</p></abstract-html>
<ref-html id="bib1.bib1"><label>Aguerre et al.(2017)Aguerre, Damián, Gimenez, and
Nigro</label><mixed-citation>
Aguerre, H. J., Damián, S. M., Gimenez, J. M., and Nigro, N. M.:
Conservative handling of arbitrary non-conformal interfaces using an
efficient supermesh, J. Comput. Phys., 335, 21–49, 2017.
</mixed-citation></ref-html>
<ref-html id="bib1.bib2"><label>Beljaars et al.(2017)Beljaars, Dutra, Balsamo, and
Lemarié</label><mixed-citation>
Beljaars, A., Dutra, E., Balsamo, G., and Lemarié, F.: On the numerical stability of surface–atmosphere coupling in weather and climate models, Geosci. Model Dev., 10, 977–989, <a href="https://doi.org/10.5194/gmd-10-977-2017" target="_blank">https://doi.org/10.5194/gmd-10-977-2017</a>, 2017.
</mixed-citation></ref-html>
<ref-html id="bib1.bib3"><label>Bell and Garland(2009)</label><mixed-citation>
Bell, N. and Garland, M.: Implementing sparse matrix-vector multiplication on
throughput-oriented processors, in: Proceedings of the conference on high
performance computing networking, storage and analysis, p. 18, ACM, 2009.
</mixed-citation></ref-html>
<ref-html id="bib1.bib4"><label>Berger(1987)</label><mixed-citation>
Berger, M. J.: On conservation at grid interfaces, SIAM J. Numer.
Anal., 24, 967–984, 1987.
</mixed-citation></ref-html>
<ref-html id="bib1.bib5"><label>Blanco and Rai(2014)</label><mixed-citation>
Blanco, J. L. and Rai, P. K.: nanoflann: a C++ header-only fork of FLANN, a
library for Nearest Neighbor (NN) wih KD-trees, available at:
<a href="https://github.com/jlblancoc/nanoflann" target="_blank"/> (last access: 16 May 2020), 2014.
</mixed-citation></ref-html>
<ref-html id="bib1.bib6"><label>Březina and Exner(2017)</label><mixed-citation>
Březina, J. and Exner, P.: Fast algorithms for intersection of non-matching
grids using Plücker coordinates, Comput. Math. Appl.,
74, 174–187, <a href="https://doi.org/10.1016/j.camwa.2017.01.028" target="_blank">https://doi.org/10.1016/j.camwa.2017.01.028</a>,
2017.
</mixed-citation></ref-html>
<ref-html id="bib1.bib7"><label>Certik et al.(2017)Certik, Ferenbaugh, Garimella, Herring, Jean,
Malone, and Sewell</label><mixed-citation>
Certik, O., Ferenbaugh, C., Garimella, R., Herring, A., Jean, B., Malone, C.,
and Sewell, C.: A Flexible Conservative Remapping Framework for Exascale
Computing, available at:
<a href="https://permalink.lanl.gov/object/tr?what=info:lanl-repo/lareport/LA-UR-17-21749" target="_blank"/> (last access: 16 May 2020),
sIAM Conference on Computational Science &amp; Engineering, February 2017.
</mixed-citation></ref-html>
<ref-html id="bib1.bib8"><label>Collins et al.(2005)Collins, Theurich, Deluca, Suarez, Trayanov,
Balaji, Li, Yang, Hill, and Da Silva</label><mixed-citation>
Collins, N., Theurich, G., Deluca, C., Suarez, M., Trayanov, A., Balaji, V.,
Li, P., Yang, W., Hill, C., and Da Silva, A.: Design and implementation of
components in the Earth System Modeling Framework,  Int. J.
High Perform. C., 19, 341–350, 2005.
</mixed-citation></ref-html>
<ref-html id="bib1.bib9"><label>Craig et al.(2017)Craig, Valcke, and Coquart</label><mixed-citation>
Craig, A., Valcke, S., and Coquart, L.: Development and performance of a new version of the OASIS coupler, OASIS3-MCT_3.0, Geosci. Model Dev., 10, 3297–3308, <a href="https://doi.org/10.5194/gmd-10-3297-2017" target="_blank">https://doi.org/10.5194/gmd-10-3297-2017</a>, 2017.
</mixed-citation></ref-html>
<ref-html id="bib1.bib10"><label>Craig et al.(2005)Craig, Jacob, Kauffman, Bettge, Larson, Ong, Ding,
and He</label><mixed-citation>
Craig, A. P., Jacob, R., Kauffman, B., Bettge, T., Larson, J., Ong, E., Ding,
C., and He, Y.: CPL6: The New Extensible, High Performance Parallel Coupler
for the Community Climate System Model, Int. J. High
Perform. C., 19, 309–327,
<a href="https://doi.org/10.1177/1094342005056117" target="_blank">https://doi.org/10.1177/1094342005056117</a>,
2005.
</mixed-citation></ref-html>
<ref-html id="bib1.bib11"><label>Craig et al.(2012)Craig, Vertenstein, and Jacob</label><mixed-citation>
Craig, A. P., Vertenstein, M., and Jacob, R.: A new flexible coupler for earth
system modeling developed for CCSM4 and CESM1, Int. J. High
Perform. C., 26, 31–42,
<a href="https://doi.org/10.1177/1094342011428141" target="_blank">https://doi.org/10.1177/1094342011428141</a>,
2012.
</mixed-citation></ref-html>
<ref-html id="bib1.bib12"><label>de Boer et al.(2008)de Boer, van Zuijlen, and Bijl</label><mixed-citation>
de Boer, A., van Zuijlen, A., and Bijl, H.: Comparison of conservative and
consistent approaches for the coupling of non-matching meshes, Comput.
Method. Appl. M., 197, 4284–4297,
<a href="https://doi.org/10.1016/j.cma.2008.05.001" target="_blank">https://doi.org/10.1016/j.cma.2008.05.001</a>,
2008.
</mixed-citation></ref-html>
<ref-html id="bib1.bib13"><label>Dennis et al.(2012)Dennis, Edwards, Evans, Guba, Lauritzen, Mirin,
St-Cyr, Taylor, and Worley</label><mixed-citation>
Dennis, J. M., Edwards, J., Evans, K. J., Guba, O., Lauritzen, P. H., Mirin,
A. A., St-Cyr, A., Taylor, M. A., and Worley, P. H.: CAM-SE: A scalable
spectral element dynamical core for the Community Atmosphere Model,
Int. J.  High Perform. C., 26, 74–89,
2012.
</mixed-citation></ref-html>
<ref-html id="bib1.bib14"><label>Devine et al.(2002)Devine, Boman, Heaphy, Hendrickson, and
Vaughan</label><mixed-citation>
Devine, K., Boman, E., Heaphy, R., Hendrickson, B., and Vaughan, C.: Zoltan
Data Management Services for Parallel Dynamic Applications, Comput.
Sci. Eng., 4, 90–97, 2002.
</mixed-citation></ref-html>
<ref-html id="bib1.bib15"><label>Dunlap et al.(2013)Dunlap, Rugaber, and Mark</label><mixed-citation>
Dunlap, R., Rugaber, S., and Mark, L.: A feature model of coupling technologies
for Earth System Models, Comput. Geosci., 53, 13–20, 2013.
</mixed-citation></ref-html>
<ref-html id="bib1.bib16"><label>E3SM Project(2018)</label><mixed-citation>
E3SM Project: Energy Exascale Earth System Model (E3SM), [Computer
Software],
<a href="https://doi.org/10.11578/E3SM/dc.20180418.36" target="_blank">https://doi.org/10.11578/E3SM/dc.20180418.36</a>,
2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib17"><label>Farrell and Maddison(2011)</label><mixed-citation>
Farrell, P. and Maddison, J.: Conservative interpolation between volume meshes
by local Galerkin projection, Comput. Method. Appl. M., 200, 89–100,
<a href="https://doi.org/10.1016/j.cma.2010.07.015" target="_blank">https://doi.org/10.1016/j.cma.2010.07.015</a>,
2011.
</mixed-citation></ref-html>
<ref-html id="bib1.bib18"><label>Fleishman et al.(2005)Fleishman, Cohen-Or, and Silva</label><mixed-citation>
Fleishman, S., Cohen-Or, D., and Silva, C. T.: Robust moving least-squares
fitting with sharp features, ACM T. Graphic, 24,
544–552, 2005.
</mixed-citation></ref-html>
<ref-html id="bib1.bib19"><label>Flyer and Wright(2007)</label><mixed-citation>
Flyer, N. and Wright, G. B.: Transport schemes on a sphere using radial basis
functions, J. Comput. Phys., 226, 1059–1084, 2007.
</mixed-citation></ref-html>
<ref-html id="bib1.bib20"><label>Fornberg and Piret(2008)</label><mixed-citation>
Fornberg, B. and Piret, C.: On choosing a radial basis function and a shape
parameter when solving a convective PDE on a sphere, J. Comput.
Phys., 227, 2758–2780, 2008.
</mixed-citation></ref-html>
<ref-html id="bib1.bib21"><label>Fox et al.(1989)Fox, Johnson, Lyzenga, Otto, Salmon, Walker, and
White</label><mixed-citation>
Fox, G., Johnson, M., Lyzenga, G., Otto, S., Salmon, J., Walker, D., and White,
R. L.: Solving Problems On Concurrent Processors Vol. 1: General Techniques
and Regular Problems, Comput. Phys., 3, 83–84, 1989.
</mixed-citation></ref-html>
<ref-html id="bib1.bib22"><label>Gander and Japhet(2013)</label><mixed-citation>
Gander, M. J. and Japhet, C.: Algorithm 932: PANG: software for nonmatching
grid projections in 2D and 3D with linear complexity, ACM T.
Math. Software, 40, 1–25, <a href="https://doi.org/10.1145/2513109.2513115" target="_blank">https://doi.org/10.1145/2513109.2513115</a>, 2013.
</mixed-citation></ref-html>
<ref-html id="bib1.bib23"><label>Gottlieb and Shu(1997)</label><mixed-citation>
Gottlieb, D. and Shu, C.-W.: On the Gibbs phenomenon and its resolution, SIAM
Rev., 39, 644–668, 1997.
</mixed-citation></ref-html>
<ref-html id="bib1.bib24"><label>Grandy(1999)</label><mixed-citation>
Grandy, J.: Conservative remapping and region overlays by intersecting
arbitrary polyhedra, J. Comput. Phys., 148, 433–466, 1999.
</mixed-citation></ref-html>
<ref-html id="bib1.bib25"><label>Guennebaud et al.(2010)Guennebaud, Jacob et al.</label><mixed-citation>
Guennebaud, G., Jacob, B., et al.: Eigen v3, available at: <a href="http://eigen.tuxfamily.org" target="_blank"/> (last access: 16 May 2020), 2010.
</mixed-citation></ref-html>
<ref-html id="bib1.bib26"><label>Hanke et al.(2016)Hanke, Redler, Holfeld, and Yastremsky</label><mixed-citation>
Hanke, M., Redler, R., Holfeld, T., and Yastremsky, M.: YAC 1.2.0: new aspects for coupling software in Earth system modelling, Geosci. Model Dev., 9, 2755–2769, <a href="https://doi.org/10.5194/gmd-9-2755-2016" target="_blank">https://doi.org/10.5194/gmd-9-2755-2016</a>, 2016.
</mixed-citation></ref-html>
<ref-html id="bib1.bib27"><label>Herring et al.(2017)Herring, Certik, Ferenbaugh, Garimella, Jean,
Malone, and Sewell</label><mixed-citation>
Herring, A. M., Certik, O., Ferenbaugh, C. R., Garimella, R. V., Jean, B. A.,
Malone, C. M., and Sewell, C. M.: (U) Introduction to Portage, Tech. rep.,
Los Alamos National Lab.(LANL), Los Alamos, NM (United States), available at:
<a href="https://permalink.lanl.gov/object/tr?what=info:lanl-repo/lareport/LA-UR-17-20831" target="_blank"/> (last access: 16 May 2020),
2017.
</mixed-citation></ref-html>
<ref-html id="bib1.bib28"><label>Hill et al.(2004)Hill, DeLuca, Balaji, Suarez, and
Silva</label><mixed-citation>
Hill, C., DeLuca, C., Balaji, Suarez, M., and Silva, A. D.: The architecture of
the earth system modeling framework, Comput. Sci.  Eng., 6,
18–28, 2004.
</mixed-citation></ref-html>
<ref-html id="bib1.bib29"><label>Hunt et al.(2006)Hunt, Mark, and Stoll</label><mixed-citation>
Hunt, W., Mark, W. R., and Stoll, G.: Fast kd-tree construction with an
adaptive error-bounded heuristic, in: Interactive Ray Tracing 2006,
81–88,  2006.
</mixed-citation></ref-html>
<ref-html id="bib1.bib30"><label>Hurrell et al.(2013)Hurrell, Holland, Gent, Ghan, Kay, Kushner,
Lamarque, Large, Lawrence, Lindsay et al.</label><mixed-citation>
Hurrell, J. W., Holland, M. M., Gent, P. R., Ghan, S., Kay, J. E., Kushner,
P. J., Lamarque, J.-F., Large, W. G., Lawrence, D., Lindsay, K.,
Lipscomb, W. H., Long, M. C., Mahowald, N., Marsh, D. R., Neale, R. B., Rasch, P., Vavrus, S., Vertenstein, M., Bader, D., Collins, W. D., Hack, J. J., Kiehl, J., and Marshall, S. L.: The
community earth system model: a framework for collaborative research,
B. Am. Meteorol. Soc., 94, 1339–1360, 2013.
</mixed-citation></ref-html>
<ref-html id="bib1.bib31"><label>Ize et al.(2007)Ize, Wald, and Parker</label><mixed-citation>
Ize, T., Wald, I., and Parker, S. G.: Asynchronous BVH construction for ray
tracing dynamic scenes on parallel multi-core architectures, in: Proceedings
of the 7th Eurographics conference on Parallel Graphics and Visualization,
Eurographics Association,  101–108, 2007.
</mixed-citation></ref-html>
<ref-html id="bib1.bib32"><label>Jacob et al.(2005)Jacob, Larson, and Ong</label><mixed-citation>
Jacob, R., Larson, J., and Ong, E.: M × N communication and parallel
interpolation in Community Climate System Model Version 3 using the model
coupling toolkit,  Int. J. High Perform. C., 19, 293–307, 2005.
</mixed-citation></ref-html>
<ref-html id="bib1.bib33"><label>Jiao and Heath(2004)</label><mixed-citation>
Jiao, X. and Heath, M. T.: Common-refinement-based data transfer between
non-matching meshes in multiphysics simulations, Int. J.
Numer. Meth. Eng., 61, 2402–2427, 2004.
</mixed-citation></ref-html>
<ref-html id="bib1.bib34"><label>Jones(1999)</label><mixed-citation>
Jones, P. W.: First-and second-order conservative remapping schemes for grids
in spherical coordinates, Mon. Weather Rev., 127, 2204–2210, 1999.
</mixed-citation></ref-html>
<ref-html id="bib1.bib35"><label>Karypis and Kumar(1998)</label><mixed-citation>
Karypis, G. and Kumar, V.: A fast and high quality multilevel scheme for partitioning irregular graphs, SIAM J. Sci. Comput., 20, 359–392, 1998.
</mixed-citation></ref-html>
<ref-html id="bib1.bib36"><label>Karypis et al.(1997)Karypis, Schloegel, and Kumar</label><mixed-citation>
Karypis, G., Schloegel, K., and Kumar, V.: Parmetis: Parallel graph
partitioning and sparse matrix ordering library, Version 1.0, Dept. of
Computer Science, University of Minnesota, 22 pp., 1997.
</mixed-citation></ref-html>
<ref-html id="bib1.bib37"><label>Larsen et al.(1999)Larsen, Gottschalk, Lin, and Manocha</label><mixed-citation>
Larsen, E., Gottschalk, S., Lin, M. C., and Manocha, D.: Fast proximity queries
with swept sphere volumes, Tech. rep., Technical Report TR99-018, Department
of Computer Science, University of North Carolina, 1999.
</mixed-citation></ref-html>
<ref-html id="bib1.bib38"><label>Larson et al.(2001)Larson, Jacob, Foster, and Guo</label><mixed-citation>
Larson, J. W., Jacob, R. L., Foster, I., and Guo, J.: The model coupling
toolkit, in: International Conference on Computational Science,
Springer, 185–194, 2001.
</mixed-citation></ref-html>
<ref-html id="bib1.bib39"><label>Lauritzen et al.(2010)Lauritzen, Nair, and Ullrich</label><mixed-citation>
Lauritzen, P. H., Nair, R. D., and Ullrich, P. A.: A conservative
semi-Lagrangian multi-tracer transport scheme (CSLAM) on the cubed-sphere
grid, J. Comput. Phys., 229, 1401–1424,
<a href="https://doi.org/10.1016/j.jcp.2009.10.036" target="_blank">https://doi.org/10.1016/j.jcp.2009.10.036</a>,
2010.
</mixed-citation></ref-html>
<ref-html id="bib1.bib40"><label>Leung et al.(2014)Leung, Rajamanickam, Pedretti, Olivier, Devine,
Deveci, Catalyurek, and Bunde</label><mixed-citation>
Leung, V. J., Rajamanickam, S., Pedretti, K., Olivier, S. L., Devine, K. D.,
Deveci, M., Catalyurek, U., and Bunde, D. P.: Zoltan2: Exploiting Geometric
Partitioning in Task Mapping for Parallel Computers, Tech. rep., Sandia
National Lab.(SNL-NM), Albuquerque, NM (United States), 2014.
</mixed-citation></ref-html>
<ref-html id="bib1.bib41"><label>Li et al.(2013)</label><mixed-citation>
Li, L., Lin, P., Yu, Y., Wang, B., Zhou, T., Liu, L., Liu, J., Bao, Q., Xu, S.,
Huang, W., Xia, K., Pu, Y., Dong, L., Shen, S., Liu, Y., Hu, N., Liu, M.,
Sun, W., Shi, X., Zheng, W., Wu, B., Song, M., Liu, H., Zhang, X., Wu, G.,
Xue, W., Huang, X., Yang, G., Song, Z., and Qiao, F.: The flexible global
ocean-atmosphere-land system model, Grid-point Version 2: FGOALS-g2, Adv.
Atmos. Sci., 30, 543–560, <a href="https://doi.org/10.1007/s00376-012-2140-6" target="_blank">https://doi.org/10.1007/s00376-012-2140-6</a>,
2013.
</mixed-citation></ref-html>
<ref-html id="bib1.bib42"><label>Liu et al.(2013)Liu, Yang, and Wang</label><mixed-citation>
Liu, L., Yang, G., and Wang, B.: CoR: a multi-dimensional common remapping
software for Earth System Models, in: The Second Workshop on Coupling
Technologies for Earth System Models (CW2013), available at:
<a href="https://wiki.cc.gatech.edu/CW2013/index.php/Program" target="_blank"/> (last access: 8 May
2014), 2013.
</mixed-citation></ref-html>
<ref-html id="bib1.bib43"><label>Liu et al.(2014)Liu, Yang, Wang, Zhang, Li, Zhang, Ji, and
Wang</label><mixed-citation>
Liu, L., Yang, G., Wang, B., Zhang, C., Li, R., Zhang, Z., Ji, Y., and Wang, L.: C-Coupler1: a Chinese community coupler for Earth system modeling, Geosci. Model Dev., 7, 2281–2302, <a href="https://doi.org/10.5194/gmd-7-2281-2014" target="_blank">https://doi.org/10.5194/gmd-7-2281-2014</a>, 2014.
</mixed-citation></ref-html>
<ref-html id="bib1.bib44"><label>Liu et al.(2018)Liu, Zhang, Li, Wang, and Yang</label><mixed-citation>
Liu, L., Zhang, C., Li, R., Wang, B., and Yang, G.: C-Coupler2: a flexible and user-friendly community coupler for model coupling and nesting, Geosci. Model Dev., 11, 3557–3586, <a href="https://doi.org/10.5194/gmd-11-3557-2018" target="_blank">https://doi.org/10.5194/gmd-11-3557-2018</a>, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib45"><label>Löhner(2014)</label><mixed-citation>
Löhner, R.: Recent advances in parallel advancing front grid generation,
Arch. Comput. Meth. E., 21, 127–140, 2014.
</mixed-citation></ref-html>
<ref-html id="bib1.bib46"><label>Löhner and Parikh(1988)</label><mixed-citation>
Löhner, R. and Parikh, P.: Generation of three-dimensional unstructured
grids by the advancing-front method, Int. J. Numer.
Meth. Fl., 8, 1135–1149, 1988.
</mixed-citation></ref-html>
<ref-html id="bib1.bib47"><label>Mahadevan et al.(2015)Mahadevan, Grindeanu, Ray, Jain, and
Wu</label><mixed-citation>
Mahadevan, V., Grindeanu, I. R., Ray, N., Jain, R., and Wu, D.: SIGMA Release
v1.2 – Capabilities, Enhancements and Fixes, <a href="https://doi.org/10.2172/1224985" target="_blank">https://doi.org/10.2172/1224985</a>, 2015.
</mixed-citation></ref-html>
<ref-html id="bib1.bib48"><label>Mahadevan et al.(2018a)Mahadevan, Grindeanu, Jacob, and
Sarich</label><mixed-citation>
Mahadevan, V., Grindeanu, I., Jacob, R., and Sarich, J.: MOAB: Serial
Advancing Front Intersection Computation,
<a href="https://doi.org/10.6084/m9.figshare.7294901.v1" target="_blank">https://doi.org/10.6084/m9.figshare.7294901.v1</a>,
2018a.
</mixed-citation></ref-html>
<ref-html id="bib1.bib49"><label>Mahadevan et al.(2018b)Mahadevan, Grindeanu, Jacob, and
Sarich</label><mixed-citation>
Mahadevan, V., Grindeanu, I., Jacob, R., and Sarich, J.: MOAB: Parallel
Advancing Front Mesh Intersection Algorithm,
<a href="https://doi.org/10.6084/m9.figshare.7294919.v2" target="_blank">https://doi.org/10.6084/m9.figshare.7294919.v2</a>,
2018b.
</mixed-citation></ref-html>
<ref-html id="bib1.bib50"><label>Mahadevan er al.(2020)</label><mixed-citation>
Mahadevan, V., Grindeanu, I., Jain, R., Shriwise, P., and Wilson, P.: MOAB v5.1.0, <a href="https://doi.org/10.5281/zenodo.2584863" target="_blank">https://doi.org/10.5281/zenodo.2584863</a>, 2020.
</mixed-citation></ref-html>
<ref-html id="bib1.bib51"><label>Morozov and Peterka(2016)</label><mixed-citation>
Morozov, D. and Peterka, T.: Block-Parallel Data Analysis with DIY2, 2016 IEEE 6th Symposium on Large Data Analysis and Visualization (LDAV), 29–36, 2016.
</mixed-citation></ref-html>
<ref-html id="bib1.bib52"><label>Mundt et al.(2016)Mundt, Boslough, Taylor, and
Roesler</label><mixed-citation>
Mundt, M., Boslough, M., Taylor, M., and Roesler, E.: A Modification to the
remapping of Gauss-Lobatto nodes to the cubed sphere, Tech. rep., Technical
Report SAND2016-0830R, Sandia National Laboratory, 2016.
</mixed-citation></ref-html>
<ref-html id="bib1.bib53"><label>Petersen et al.(2015)Petersen, Jacobsen, Ringler, Hecht, and
Maltrud</label><mixed-citation>
Petersen, M. R., Jacobsen, D. W., Ringler, T. D., Hecht, M. W., and Maltrud,
M. E.: Evaluation of the arbitrary Lagrangian–Eulerian vertical coordinate
method in the MPAS-Ocean model, Ocean Model., 86, 93–113, 2015.
</mixed-citation></ref-html>
<ref-html id="bib1.bib54"><label>Plimpton et al.(2004)Plimpton, Hendrickson, and
Stewart</label><mixed-citation>
Plimpton, S. J., Hendrickson, B., and Stewart, J. R.: A parallel rendezvous
algorithm for interpolation between multiple grids, J. Parallel
Distr. Com., 64, 266–276, 2004.
</mixed-citation></ref-html>
<ref-html id="bib1.bib55"><label>Powell and Abel(2015)</label><mixed-citation>
Powell, D. and Abel, T.: An exact general remeshing scheme applied to
physically conservative voxelization, J. Comput. Phys., 297,
340–356, <a href="https://doi.org/10.1016/j.jcp.2015.05.022" target="_blank">https://doi.org/10.1016/j.jcp.2015.05.022</a>,
2015.
</mixed-citation></ref-html>
<ref-html id="bib1.bib56"><label>Rančić(1995)</label><mixed-citation>
Rančić, M.: An efficient, conservative, monotonic remapping for
semi-Lagrangian transport algorithms, Mon. Weather Rev., 123,
1213–1217, 1995.
</mixed-citation></ref-html>
<ref-html id="bib1.bib57"><label>Reichler and Kim(2008)</label><mixed-citation>
Reichler, T. and Kim, J.: How well do coupled models simulate today's climate?,
B. Am. Meteorol. Soc., 89, 303–312, 2008.
</mixed-citation></ref-html>
<ref-html id="bib1.bib58"><label>Ringler et al.(2013)Ringler, Petersen, Higdon, Jacobsen, Jones, and
Maltrud</label><mixed-citation>
Ringler, T., Petersen, M., Higdon, R. L., Jacobsen, D., Jones, P. W., and
Maltrud, M.: A multi-resolution approach to global ocean modeling, Ocean
Model., 69, 211–232, 2013.
</mixed-citation></ref-html>
<ref-html id="bib1.bib59"><label>Schliephake and Laure(2015)</label><mixed-citation>
Schliephake, M. and Laure, E.: Performance Analysis of Irregular Collective
Communication with the Crystal Router Algorithm, in: Solving Software
Challenges for Exascale, edited by: Markidis, S. and Laure, E.,
Springer International Publishing, Cham, 130–140, 2015.
</mixed-citation></ref-html>
<ref-html id="bib1.bib60"><label>Shewchuk(1996)</label><mixed-citation>
Shewchuk, J. R.: Triangle: Engineering a 2D quality mesh generator and Delaunay
triangulator, in: Applied computational geometry towards geometric
engineering,  Springer, 203–222, 1996.
</mixed-citation></ref-html>
<ref-html id="bib1.bib61"><label>SIGMA toolkit(2014)</label><mixed-citation>
SIGMA toolkit: Scalable Interfaces for Geometry and Mesh based Applications
(SIGMA) toolkit, available at: <a href="http://sigma.mcs.anl.gov" target="_blank"/> (last access: 16 May 2020),
2014.
</mixed-citation></ref-html>
<ref-html id="bib1.bib62"><label>Slattery et al.(2013)Slattery, Wilson, and
Pawlowski</label><mixed-citation>
Slattery, S., Wilson, P., and Pawlowski, R.: The data transfer kit: a geometric
rendezvous-based tool for multiphysics data transfer, in: International
conference on mathematics &amp; computational methods applied to nuclear science
&amp; engineering (M&amp;C 2013),   5–9, 2013.
</mixed-citation></ref-html>
<ref-html id="bib1.bib63"><label>Slattery(2016)</label><mixed-citation>
Slattery, S. R.: Mesh-free data transfer algorithms for partitioned
multiphysics problems: Conservation, accuracy, and parallelism, J.
Comput. Phys., 307, 164–188, 2016.
</mixed-citation></ref-html>
<ref-html id="bib1.bib64"><label>Slingo et al.(2009)Slingo, Bates, Nikiforakis, Piggott, Roberts,
Shaffrey, Stevens, Vidale, and Weller</label><mixed-citation>
Slingo, J., Bates, K., Nikiforakis, N., Piggott, M., Roberts, M., Shaffrey, L.,
Stevens, I., Vidale, P. L., and Weller, H.: Developing the next-generation
climate system models: challenges and achievements, Philos.
T. R. Soc. S-A, 367, 815–831, 2009.
</mixed-citation></ref-html>
<ref-html id="bib1.bib65"><label>Tautges and Caceres(2009)</label><mixed-citation>
Tautges, T. J. and Caceres, A.: Scalable parallel solution coupling for
multiphysics reactor simulation, J.  Phys. Conf. Ser., 180, 012017, <a href="https://doi.org/10.1088/1742-6596/180/1/012017" target="_blank">https://doi.org/10.1088/1742-6596/180/1/012017</a>, 2009.
</mixed-citation></ref-html>
<ref-html id="bib1.bib66"><label>Tautges et al.(2004)Tautges, Meyers, Merkley, Stimpson, and
Ernst</label><mixed-citation>
Tautges, T. J., Meyers, R., Merkley, K., Stimpson, C., and Ernst, C.: MOAB: A
Mesh-Oriented datABase, SAND2004-1592, Sandia National Laboratories,
2004.
</mixed-citation></ref-html>
<ref-html id="bib1.bib67"><label>Tautges et al.(2012)Tautges, Kraftcheck, Bertram, Sachdeva, and
Magerlein</label><mixed-citation>
Tautges, T. J., Kraftcheck, J., Bertram, N., Sachdeva, V., and Magerlein, J.:
Mesh Interface Resolution and Ghost Exchange in a Parallel Mesh
Representation, IEEE, Shanghai, China, 2012.

</mixed-citation></ref-html>
<ref-html id="bib1.bib68"><label>Taylor et al.(2007)Taylor, Edwards, Thomas, and
Nair</label><mixed-citation>
Taylor, M., Edwards, J., Thomas, S., and Nair, R.: A mass and energy conserving
spectral element atmospheric dynamical core on the cubed-sphere grid,
J. Phys. Conf. Ser.,  78, p. 012074, <a href="https://doi.org/10.1088/1742-6596/78/1/012074" target="_blank">https://doi.org/10.1088/1742-6596/78/1/012074</a>, 2007.
</mixed-citation></ref-html>
<ref-html id="bib1.bib69"><label>Thomas and Loft(2005)</label><mixed-citation>
Thomas, S. J. and Loft, R. D.: The NCAR spectral element climate dynamical
core: Semi-implicit Eulerian formulation, J. Sci. Comput.,
25, 307–322, 2005.
</mixed-citation></ref-html>
<ref-html id="bib1.bib70"><label>Ullrich and Mahadevan(2020)</label><mixed-citation>
Ullrich, P. and Mahadevan, V.: TempestRemap v2.0.2, available at:
<a href="https://github.com/ClimateGlobalChange/tempestremap" target="_blank"/>, last access: 16 May 2020.
</mixed-citation></ref-html>
<ref-html id="bib1.bib71"><label>Ullrich and Taylor(2015)</label><mixed-citation>
Ullrich, P. A. and Taylor, M. A.: Arbitrary-order conservative and consistent
remapping and a theory of linear maps: Part I, Mon. Weather Rev., 143,
2419–2440, 2015.
</mixed-citation></ref-html>
<ref-html id="bib1.bib72"><label>Ullrich et al.(2009)</label><mixed-citation>
Ullrich, P. A., Lauritzen, P. H., and Jablonowski, C.: Geometrically Exact
Conservative Remapping (GECoRe): Regular latitude–longitude and
cubed-sphere grids, Mon. Weather Rev., 137, 1721–1741, 2009.
</mixed-citation></ref-html>
<ref-html id="bib1.bib73"><label>Ullrich et al.(2013)</label><mixed-citation>
Ullrich, P. A., Lauritzen, P. H., and Jablonowski, C.: Some considerations for
high-order “incremental remap”-based transport schemes: edges,
reconstructions, and area integration, Int. J. Nume.
Meth. Fl., 71, 1131–1151, 2013.
</mixed-citation></ref-html>
<ref-html id="bib1.bib74"><label>Ullrich et al.(2016)</label><mixed-citation>
Ullrich, P. A., Devendran, D., and Johansen, H.: Arbitrary-order conservative
and consistent remapping and a theory of linear maps: Part II, Mon.
Weather Rev., 144, 1529–1549, 2016.
</mixed-citation></ref-html>
<ref-html id="bib1.bib75"><label>Valcke(2013)</label><mixed-citation>
Valcke, S.: The OASIS3 coupler: a European climate modelling community software, Geosci. Model Dev., 6, 373–388, <a href="https://doi.org/10.5194/gmd-6-373-2013" target="_blank">https://doi.org/10.5194/gmd-6-373-2013</a>, 2013.
</mixed-citation></ref-html>
<ref-html id="bib1.bib76"><label>Valcke et al.(2018)</label><mixed-citation>
Valcke, S., Craig, A., and Coquart, L.: OASIS3-MCT User Guide, OASIS3-MCT4.0,
Tech. rep., CECI, Université de Toulouse, CNRS, CERFACS – TR-CMGC-18-77,
Toulouse, France,
available at: <a href="https://cerfacs.fr/wp-content/uploads/2018/07/GLOBC-TR-oasis3mct_UserGuide4.0_30062018.pdf" target="_blank"/> (last access: 19 May 2020),
2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib77"><label>VisIt(2005)</label><mixed-citation>
VisIt: VisIt User's Guide, Tech. Rep. UCRL-SM-220449, Lawrence Livermore
National Laboratory, 2005.
</mixed-citation></ref-html>
<ref-html id="bib1.bib78"><label>Wan et al.(2013)</label><mixed-citation>
Wan, H., Giorgetta, M. A., Zängl, G., Restelli, M., Majewski, D., Bonaventura, L., Fröhlich, K., Reinert, D., Rípodas, P., Kornblueh, L., and Förstner, J.: The ICON-1.2 hydrostatic atmospheric dynamical core on triangular grids – Part 1: Formulation and performance of the baseline version, Geosci. Model Dev., 6, 735–763, <a href="https://doi.org/10.5194/gmd-6-735-2013" target="_blank">https://doi.org/10.5194/gmd-6-735-2013</a>, 2013.
</mixed-citation></ref-html>
<ref-html id="bib1.bib79"><label>Washington et al.(2008)Washington, Bader, and
Collins</label><mixed-citation>
Washington, W., Bader, D., and Collins, B. E. A.: Challenges in climate change
science and the role of computing at the extreme scale, in: Proc. of the
Workshop on Climate Science, 2008.
</mixed-citation></ref-html>
<ref-html id="bib1.bib80"><label>Zhou(2006)</label><mixed-citation>
Zhou, S.: Coupling climate models with the earth system modeling framework and
the common component architecture, Concurr. Comp.-Pract. E., 18, 203–213, 2006.
</mixed-citation></ref-html>
<ref-html id="bib1.bib81"><label>Zienkiewicz and Zhu(1992)</label><mixed-citation>
Zienkiewicz, O. C. and Zhu, J. Z.: The superconvergent patch recovery and a
posteriori error estimates. Part 1: The recovery technique, Int.
J. Numer. Meth. Eng., 33, 1331–1364, 1992.
</mixed-citation></ref-html>--></article>
