<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing with OASIS Tables v3.0 20080202//EN" "https://jats.nlm.nih.gov/nlm-dtd/publishing/3.0/journalpub-oasis3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:oasis="http://docs.oasis-open.org/ns/oasis-exchange/table" xml:lang="en" dtd-version="3.0" article-type="research-article"><?xmltex \bartext{Development and technical paper}?>
  <front>
    <journal-meta><journal-id journal-id-type="publisher">GMD</journal-id><journal-title-group>
    <journal-title>Geoscientific Model Development</journal-title>
    <abbrev-journal-title abbrev-type="publisher">GMD</abbrev-journal-title><abbrev-journal-title abbrev-type="nlm-ta">Geosci. Model Dev.</abbrev-journal-title>
  </journal-title-group><issn pub-type="epub">1991-9603</issn><publisher>
    <publisher-name>Copernicus Publications</publisher-name>
    <publisher-loc>Göttingen, Germany</publisher-loc>
  </publisher></journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.5194/gmd-17-899-2024</article-id><title-group><article-title>Graphics-processing-unit-accelerated ice flow solver for unstructured meshes using the Shallow-Shelf Approximation (FastIceFlo v1.0.1)</article-title><alt-title>FastIceFlo v2.0</alt-title>
      </title-group><?xmltex \runningtitle{FastIceFlo v2.0}?><?xmltex \runningauthor{A. Sandip et al.}?>
      <contrib-group>
        <contrib contrib-type="author" corresp="yes" rid="aff1">
          <name><surname>Sandip</surname><given-names>Anjali</given-names></name>
          <email>anjali.sandip@und.edu</email>
        <ext-link>https://orcid.org/0000-0001-9221-1910</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff2 aff3 aff5">
          <name><surname>Räss</surname><given-names>Ludovic</given-names></name>
          
        <ext-link>https://orcid.org/0000-0002-1136-899X</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff4">
          <name><surname>Morlighem</surname><given-names>Mathieu</given-names></name>
          
        <ext-link>https://orcid.org/0000-0001-5219-1310</ext-link></contrib>
        <aff id="aff1"><label>1</label><institution>Department of Mechanical Engineering, University of North Dakota, North Dakota, USA</institution>
        </aff>
        <aff id="aff2"><label>2</label><institution>Laboratory of Hydraulics, Hydrology and Glaciology (VAW), ETH Zurich, Zurich, Switzerland</institution>
        </aff>
        <aff id="aff3"><label>3</label><institution>Swiss Federal Institute for Forest, Snow and Landscape Research (WSL), Birmensdorf, Switzerland</institution>
        </aff>
        <aff id="aff4"><label>4</label><institution>Department of Earth Sciences, Dartmouth College, New Hampshire, USA</institution>
        </aff>
        <aff id="aff5"><label>a</label><institution>now at: Swiss Geocomputing Centre, Faculty of Geosciences and Environment, <?xmltex \hack{\break}?>University of Lausanne, Lausanne, Switzerland</institution>
        </aff>
      </contrib-group>
      <author-notes><corresp id="corr1">Anjali Sandip (anjali.sandip@und.edu)</corresp></author-notes><pub-date><day>2</day><month>February</month><year>2024</year></pub-date>
      
      <volume>17</volume>
      <issue>2</issue>
      <fpage>899</fpage><lpage>909</lpage>
      <history>
        <date date-type="received"><day>3</day><month>March</month><year>2023</year></date>
           <date date-type="rev-request"><day>10</day><month>May</month><year>2023</year></date>
           <date date-type="rev-recd"><day>23</day><month>November</month><year>2023</year></date>
           <date date-type="accepted"><day>19</day><month>December</month><year>2023</year></date>
      </history>
      <permissions>
        <copyright-statement>Copyright: © 2024 Anjali Sandip et al.</copyright-statement>
        <copyright-year>2024</copyright-year>
      <license license-type="open-access"><license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p></license></permissions><self-uri xlink:href="https://gmd.copernicus.org/articles/17/899/2024/gmd-17-899-2024.html">This article is available from https://gmd.copernicus.org/articles/17/899/2024/gmd-17-899-2024.html</self-uri><self-uri xlink:href="https://gmd.copernicus.org/articles/17/899/2024/gmd-17-899-2024.pdf">The full text article is available as a PDF file from https://gmd.copernicus.org/articles/17/899/2024/gmd-17-899-2024.pdf</self-uri>
      <abstract><title>Abstract</title>

      <p id="d1e129">Ice-sheet flow models capable of accurately projecting their future mass balance constitute tools to improve flood risk assessment and assist sea-level rise mitigation associated with enhanced ice discharge. Some processes that need to be captured, such as grounding-line migration, require high spatial resolution (under the kilometer scale). Conventional ice flow models mainly execute on central processing units (CPUs), which feature limited parallel processing capabilities and peak memory bandwidth. This may hinder model scalability and result in long run times, requiring significant computational resources. As an alternative, graphics processing units (GPUs) are ideally suited for high spatial resolution, as the calculations can be performed concurrently by thousands of threads, processing most of the computational domain simultaneously. In this study, we combine a GPU-based approach with the pseudo-transient (PT) method, an accelerated iterative and matrix-free solution strategy, and investigate its performance for finite elements and unstructured meshes with application to two-dimensional (2-D) models of real glaciers at a regional scale. For both the Jakobshavn and Pine Island glacier models, the number of nonlinear PT iterations required to converge a given number of vertices (<inline-formula><mml:math id="M1" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula>) scales in the order of <inline-formula><mml:math id="M2" display="inline"><mml:mrow><mml:mi mathvariant="script">O</mml:mi><mml:mo>(</mml:mo><mml:msup><mml:mi>N</mml:mi><mml:mn mathvariant="normal">1.2</mml:mn></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> or better. We further compare the performance of the PT CUDA C implementation with a standard finite-element CPU-based implementation using the price-to-performance metric. The price of a single Tesla V100 GPU is 1.5 times that of two Intel Xeon Gold 6140 CPUs. We expect a minimum speedup of at least 1.5 times to justify the Tesla V100 GPU price to performance. Our developments result in a GPU-based implementation that achieves this goal with a speedup beyond 1.5 times. This study represents a first step toward leveraging GPU processing power, enabling more accurate polar ice discharge predictions. The insights gained will benefit efforts to diminish spatial resolution constraints at higher computing performance. The higher computing performance will allow for ensembles of ice-sheet flow simulations to be run at the continental scale and higher resolution, a previously challenging task. The advances will further enable the quantification of model sensitivity to changes in upcoming climate forcings. These findings will significantly benefit process-oriented sea-level-projection studies over the coming decades.</p>
  </abstract>
    </article-meta>
  </front>
<body>
      

<sec id="Ch1.S1" sec-type="intro">
  <label>1</label><title>Introduction</title>
      <?pagebreak page900?><p id="d1e165">Global mean sea level is rising at an average rate of 3.7 mm yr<inline-formula><mml:math id="M3" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>, posing a significant threat to coastal communities and global ecosystems <xref ref-type="bibr" rid="bib1.bibx13 bib1.bibx17" id="paren.1"/>. The increase in ice discharge from the Greenland and Antarctic ice sheets significantly contributes to sea-level rise. However, their dynamic response to climate change remains a fundamental uncertainty in future projection <xref ref-type="bibr" rid="bib1.bibx28 bib1.bibx4 bib1.bibx14" id="paren.2"/>. While much progress has been made over the last decades, several critical physical processes, such as calving and ice-sheet basal sliding, remain poorly understood <xref ref-type="bibr" rid="bib1.bibx23" id="paren.3"/>. Existing computational resources limit the spatial resolution and simulation time on which continental-scale ice-sheet models can run. Some processes, such as grounding-line migration or ice front dynamics, require spatial resolutions in the order of 1 km or smaller <xref ref-type="bibr" rid="bib1.bibx18 bib1.bibx1 bib1.bibx3" id="paren.4"/>.</p>
      <p id="d1e192">Most numerical models use a solution strategy designed to target central processing units (CPUs) and shared memory parallelization. CPUs' parallel processing capabilities, peak memory bandwidth, and power consumption remain limiting factors. It remains to be seen whether high-resolution modeling will become feasible at the continental scale (or ice-sheet scale). Specifically, complex flow models, such as full-Stokes models, may remain challenging to employ beyond the regional scale. Trying to overcome the technical limitations tied to CPU-based computing, graphics processing units (GPUs) feature interesting capabilities and have been booming over the past decade <xref ref-type="bibr" rid="bib1.bibx2 bib1.bibx12" id="paren.5"/>. Developing algorithms and solvers to leverage GPU computing capabilities has become essential and has resulted in active development within scientific computing and high-performance computing (HPC) communities.</p>
      <p id="d1e198">The traditional way of solving the partial differential equations governing ice-sheet flow, employing, e.g., finite-element analysis, may represent a challenge to leverage GPU acceleration efficiently. Handling unstructured grid geometries and having global-to-local indexing patterns may significantly hinder efficient memory transfers and optimal bandwidth utilization. <xref ref-type="bibr" rid="bib1.bibx26" id="text.6"/> proposed an alternative approach by reformulating the flow equations in the form of pseudo-transient (PT) updates. The PT method augments the time-independent governing ice-sheet flow equation by physically motivated pseudo-time-dependent terms. The added pseudo-time <inline-formula><mml:math id="M4" display="inline"><mml:mi mathvariant="italic">τ</mml:mi></mml:math></inline-formula> terms turn the initial time-independent elliptic equations into a parabolic form, allowing for an explicit iterative pseudo-time integration to reach a steady state and, thus, the solution of the initial elliptic problem. The explicit pseudo-time integration scheme eliminates the need for the expensive direct–iterative type of solvers, making the proposed approach matrix-free and attractive for various parallel computing approaches <xref ref-type="bibr" rid="bib1.bibx7 bib1.bibx24 bib1.bibx16" id="paren.7"/>. <xref ref-type="bibr" rid="bib1.bibx26" id="text.8"/> introduced this method specifically targeting GPU computing to enable the development of high-spatial-resolution full-Stokes ice-sheet flow solvers in two dimensions (2-D) and three dimensions (3-D), respectively, on uniform grids <xref ref-type="bibr" rid="bib1.bibx26" id="paren.9"/>. The approach unveils a promising solution strategy, but the finite-difference discretization on uniform and structured grids and the idealized test cases represent actual limitations.</p>
      <p id="d1e220">Here, we build upon work from previous studies <xref ref-type="bibr" rid="bib1.bibx25 bib1.bibx26" id="paren.10"/> on the accelerated PT method for finite-difference discretization on uniform, structured grids and extend it to finite-element discretization and unstructured meshes. We developed a CUDA C implementation of the PT depth-integrated Shallow-Shelf Approximation (SSA) and applied it to regional-scale glaciers, Pine Island Glacier and Jakobshavn Isbræ, in West Antarctica and Greenland, respectively. We compare the PT CUDA C implementation with a more standard finite-element CPU-based implementation available within the Ice-sheet and Sea-level System Model (ISSM). Our comparison uses the same mesh, model equations, and boundary conditions. In Sect. <xref ref-type="sec" rid="Ch1.S2"/>, we present the mathematical reformulation of the 2-D SSA momentum balance equations to incorporate the additional pseudo-transient terms needed for the PT method. We provide the weak formulation and discuss the spatial discretization. Section <xref ref-type="sec" rid="Ch1.S3"/> describes the numerical experiments conducted, chosen glacier model configurations, hardware implementation, and performance assessment metrics. In Sects. <xref ref-type="sec" rid="Ch1.S4"/> and <xref ref-type="sec" rid="Ch1.S5"/>, we illustrate the method's performance and conclude on future research directions.</p>
</sec>
<sec id="Ch1.S2">
  <label>2</label><title>Methods</title>
<sec id="Ch1.S2.SS1">
  <label>2.1</label><title>Mathematical formulation of the 2-D SSA model</title>
      <p id="d1e249">We employ the SSA <xref ref-type="bibr" rid="bib1.bibx19" id="paren.11"/> formulation to solve the momentum balance equation:
            <disp-formula id="Ch1.E1" content-type="numbered"><label>1</label><mml:math id="M5" display="block"><mml:mrow><mml:mi mathvariant="normal">∇</mml:mi><mml:mo>⋅</mml:mo><mml:mfenced close=")" open="("><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mi>H</mml:mi><mml:mi mathvariant="italic">μ</mml:mi><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">ε</mml:mi><mml:mo mathvariant="normal">˙</mml:mo></mml:mover><mml:mi mathvariant="normal">SSA</mml:mi></mml:msub></mml:mrow></mml:mfenced><mml:mo>=</mml:mo><mml:mi mathvariant="italic">ρ</mml:mi><mml:mi>g</mml:mi><mml:mi>H</mml:mi><mml:mi mathvariant="normal">∇</mml:mi><mml:mi mathvariant="bold-italic">s</mml:mi><mml:mo>+</mml:mo><mml:msup><mml:mi mathvariant="italic">α</mml:mi><mml:mn mathvariant="bold">2</mml:mn></mml:msup><mml:mi mathvariant="bold-italic">v</mml:mi><mml:mspace linebreak="nobreak" width="0.25em"/><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
          where the 2-D SSA strain rate <inline-formula><mml:math id="M6" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">ε</mml:mi><mml:mo mathvariant="normal">˙</mml:mo></mml:mover><mml:mi mathvariant="normal">SSA</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is defined as
            <disp-formula id="Ch1.E2" content-type="numbered"><label>2</label><mml:math id="M7" display="block"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">ε</mml:mi><mml:mo mathvariant="normal">˙</mml:mo></mml:mover><mml:mi mathvariant="normal">SSA</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mfenced close=")" open="("><mml:mtable class="array" columnalign="center center"><mml:mtr><mml:mtd><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mstyle displaystyle="false"><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mrow><mml:mo>∂</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>∂</mml:mo><mml:mi>x</mml:mi></mml:mrow></mml:mfrac></mml:mstyle></mml:mstyle><mml:mo>+</mml:mo><mml:mstyle displaystyle="false"><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mrow><mml:mo>∂</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>∂</mml:mo><mml:mi>y</mml:mi></mml:mrow></mml:mfrac></mml:mstyle></mml:mstyle></mml:mrow></mml:mtd><mml:mtd><mml:mstyle displaystyle="false"><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mrow><mml:mo>∂</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>∂</mml:mo><mml:mi>y</mml:mi></mml:mrow></mml:mfrac></mml:mstyle></mml:mstyle></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd/></mml:mtr><mml:mtr><mml:mtd><mml:mstyle displaystyle="false"><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mrow><mml:mo>∂</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>∂</mml:mo><mml:mi>y</mml:mi></mml:mrow></mml:mfrac></mml:mstyle></mml:mstyle></mml:mtd><mml:mtd><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mstyle displaystyle="false"><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mrow><mml:mo>∂</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>∂</mml:mo><mml:mi>y</mml:mi></mml:mrow></mml:mfrac></mml:mstyle></mml:mstyle><mml:mo>+</mml:mo><mml:mstyle displaystyle="false"><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mrow><mml:mo>∂</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>∂</mml:mo><mml:mi>x</mml:mi></mml:mrow></mml:mfrac></mml:mstyle></mml:mstyle></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mfenced><mml:mspace width="0.25em" linebreak="nobreak"/><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>
          The terms <inline-formula><mml:math id="M8" display="inline"><mml:mrow><mml:msub><mml:mi>v</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M9" display="inline"><mml:mrow><mml:msub><mml:mi>v</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> represent the respective <inline-formula><mml:math id="M10" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M11" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula> ice velocity components, <inline-formula><mml:math id="M12" display="inline"><mml:mi>H</mml:mi></mml:math></inline-formula> is the ice thickness, <inline-formula><mml:math id="M13" display="inline"><mml:mi mathvariant="italic">ρ</mml:mi></mml:math></inline-formula> is the ice density, <inline-formula><mml:math id="M14" display="inline"><mml:mi>g</mml:mi></mml:math></inline-formula> is the gravitational acceleration, <inline-formula><mml:math id="M15" display="inline"><mml:mi>s</mml:mi></mml:math></inline-formula> is the glacier's upper surface <inline-formula><mml:math id="M16" display="inline"><mml:mi>z</mml:mi></mml:math></inline-formula> coordinate, and <inline-formula><mml:math id="M17" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="italic">α</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mi mathvariant="bold-italic">v</mml:mi></mml:mrow></mml:math></inline-formula> is the basal friction term. The ice viscosity <inline-formula><mml:math id="M18" display="inline"><mml:mi mathvariant="italic">μ</mml:mi></mml:math></inline-formula> follows Glen's flow law <xref ref-type="bibr" rid="bib1.bibx9" id="paren.12"/>:
            <disp-formula id="Ch1.E3" content-type="numbered"><label>3</label><mml:math id="M19" display="block"><mml:mrow><mml:mi mathvariant="italic">μ</mml:mi><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mi>B</mml:mi><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mspace width="0.25em" linebreak="nobreak"/><mml:msubsup><mml:mover accent="true"><mml:mi mathvariant="italic">ε</mml:mi><mml:mo mathvariant="normal">˙</mml:mo></mml:mover><mml:mi mathvariant="normal">e</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>n</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>)</mml:mo><mml:mo>/</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:mfrac></mml:mstyle><mml:mspace linebreak="nobreak" width="0.25em"/><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
          where <inline-formula><mml:math id="M20" display="inline"><mml:mi>B</mml:mi></mml:math></inline-formula> is ice rigidity,  <inline-formula><mml:math id="M21" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi mathvariant="italic">ε</mml:mi><mml:mo mathvariant="normal">˙</mml:mo></mml:mover><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the effective strain rate, and <inline-formula><mml:math id="M22" display="inline"><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:math></inline-formula> is Glen's power-law exponent. We regularize the strain-rate-dependent viscosity formulation in the numerical implementation by capping it at <inline-formula><mml:math id="M23" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mn mathvariant="normal">5</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> to address the singularity arising in regions of the computational domain where the strain rate tends toward zero.</p>
      <p id="d1e645">As boundary conditions, we apply water pressure at the ice front <inline-formula><mml:math id="M24" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">Γ</mml:mi><mml:mi mathvariant="italic">σ</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and nonhomogeneous Dirichlet boundary conditions <inline-formula><mml:math id="M25" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">Γ</mml:mi><mml:mi>u</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> on the other boundaries (based on observed velocity).</p>
</sec>
<?pagebreak page901?><sec id="Ch1.S2.SS2">
  <label>2.2</label><title>Mathematical reformulation of the 2-D SSA model to incorporate the PT method</title>
      <p id="d1e678">The solution to the SSA ice flow problem is commonly achieved by discretizing Eq. (<xref ref-type="disp-formula" rid="Ch1.E1"/>) using the finite-element or finite-difference method. The discretized problem can be solved using a direct, direct–iterative, or iterative approach. Robust matrix-based direct-type solvers exhibit significant scaling limitations restricting their applicability when considering high-resolution or 3-D configurations. Iterative solving approaches allow one to circumvent most scaling limitations. However, they may encounter convergence issues for suboptimally conditioned, stiff, or highly nonlinear problems, resulting in the non-tractable growth of the iteration count. Thus, one challenge is to prevent the iteration count from growing exponentially. We propose the accelerated pseudo-transient (PT) method as an alternative approach. The method augments the steady-state viscous flow (Eq. <xref ref-type="disp-formula" rid="Ch1.E1"/>) by adding the usually ignored transient term, which can be further used to integrate the equations in pseudo-time <inline-formula><mml:math id="M26" display="inline"><mml:mi mathvariant="italic">τ</mml:mi></mml:math></inline-formula>, seeking an implicit solution once the steady state is reached, i.e., when <inline-formula><mml:math id="M27" display="inline"><mml:mrow><mml:mi mathvariant="italic">τ</mml:mi><mml:mo>→</mml:mo><mml:mi mathvariant="normal">∞</mml:mi></mml:mrow></mml:math></inline-formula>.</p>
      <p id="d1e704">Building upon work from previous studies <xref ref-type="bibr" rid="bib1.bibx22 bib1.bibx6 bib1.bibx25 bib1.bibx26 bib1.bibx27" id="paren.13"/>, we reformulate the 2-D SSA steady-state momentum balance equations into a transient diffusion-like formulation for flow velocities <inline-formula><mml:math id="M28" display="inline"><mml:mrow><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> by incorporating the usually omitted time derivative.
            <disp-formula id="Ch1.E4" content-type="numbered"><label>4</label><mml:math id="M29" display="block"><mml:mrow><mml:mi mathvariant="normal">∇</mml:mi><mml:mo>⋅</mml:mo><mml:mfenced open="(" close=")"><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mi>H</mml:mi><mml:mi mathvariant="italic">μ</mml:mi><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">ε</mml:mi><mml:mo mathvariant="normal">˙</mml:mo></mml:mover><mml:mi mathvariant="normal">SSA</mml:mi></mml:msub></mml:mrow></mml:mfenced><mml:mo>-</mml:mo><mml:mi mathvariant="italic">ρ</mml:mi><mml:mi>g</mml:mi><mml:mi>H</mml:mi><mml:mi mathvariant="normal">∇</mml:mi><mml:mi>s</mml:mi><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="italic">α</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mi mathvariant="bold-italic">v</mml:mi><mml:mo>=</mml:mo><mml:mi mathvariant="italic">ρ</mml:mi><mml:mi>H</mml:mi><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mo>∂</mml:mo><mml:mi mathvariant="bold-italic">v</mml:mi></mml:mrow><mml:mrow><mml:mo>∂</mml:mo><mml:mi mathvariant="italic">τ</mml:mi></mml:mrow></mml:mfrac></mml:mstyle><mml:mspace linebreak="nobreak" width="0.25em"/></mml:mrow></mml:math></disp-formula>
          The velocity–time derivatives represent physically motivated expressions that we can further use to iteratively reach a steady state, which provides the solution of the original time-independent equations. As we are only interested in the steady-state here, transient processes evolve in numerical time or pseudo-time <inline-formula><mml:math id="M30" display="inline"><mml:mi mathvariant="italic">τ</mml:mi></mml:math></inline-formula>:
            <disp-formula id="Ch1.E5" content-type="numbered"><label>5</label><mml:math id="M31" display="block"><mml:mtable rowspacing="0.2ex" class="split" displaystyle="true" columnalign="right"><mml:mtr><mml:mtd><mml:mrow><mml:mi mathvariant="italic">ρ</mml:mi><mml:mi>H</mml:mi><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mo>∂</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>∂</mml:mo><mml:mi mathvariant="italic">τ</mml:mi></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>=</mml:mo><mml:msub><mml:mi>R</mml:mi><mml:mi>x</mml:mi></mml:msub><mml:mspace linebreak="nobreak" width="0.25em"/><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mi mathvariant="italic">ρ</mml:mi><mml:mi>H</mml:mi><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mo>∂</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>∂</mml:mo><mml:mi mathvariant="italic">τ</mml:mi></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>=</mml:mo><mml:msub><mml:mi>R</mml:mi><mml:mi>y</mml:mi></mml:msub><mml:mspace linebreak="nobreak" width="0.25em"/><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
          where <inline-formula><mml:math id="M32" display="inline"><mml:mrow><mml:msub><mml:mi>R</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M33" display="inline"><mml:mrow><mml:msub><mml:mi>R</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> correspond to the right-hand-side expressions of Eq. (<xref ref-type="disp-formula" rid="Ch1.E1"/>) and define the residuals of the original SSA equations for which we are seeking a solution. We define the transient <italic>pseudo</italic>-time step <inline-formula><mml:math id="M34" display="inline"><mml:mrow><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi mathvariant="italic">τ</mml:mi></mml:mrow></mml:math></inline-formula> as a field variable (that is spatially variable) chosen to minimize the number of nonlinear PT iterations.</p>
</sec>
<sec id="Ch1.S2.SS3">
  <label>2.3</label><title>Pseudo-time-stepping method</title>
      <p id="d1e916">Here, we advance in numerical pseudo-time using a forward Euler pseudo-time-stepping method. We choose our time derivative by approximating the transient diffusive system for both <inline-formula><mml:math id="M35" display="inline"><mml:mrow><mml:msub><mml:mi>v</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M36" display="inline"><mml:mrow><mml:msub><mml:mi>v</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>:
            <disp-formula id="Ch1.E6" content-type="numbered"><label>6</label><mml:math id="M37" display="block"><mml:mtable class="split" rowspacing="0.2ex" displaystyle="true" columnalign="right left"><mml:mtr><mml:mtd><mml:mrow><mml:mi mathvariant="italic">ρ</mml:mi><mml:mi>H</mml:mi><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mo>∂</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>∂</mml:mo><mml:mi mathvariant="italic">τ</mml:mi></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mo>∂</mml:mo><mml:mrow><mml:mo>∂</mml:mo><mml:mi>x</mml:mi></mml:mrow></mml:mfrac></mml:mstyle><mml:mfenced open="(" close=")"><mml:mrow><mml:mn mathvariant="normal">4</mml:mn><mml:mi>H</mml:mi><mml:mi mathvariant="italic">μ</mml:mi><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mo>∂</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>∂</mml:mo><mml:mi>x</mml:mi></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:mfenced><mml:mspace linebreak="nobreak" width="0.25em"/><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mi mathvariant="italic">ρ</mml:mi><mml:mi>H</mml:mi><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mo>∂</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>∂</mml:mo><mml:mi mathvariant="italic">τ</mml:mi></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mo>∂</mml:mo><mml:mrow><mml:mo>∂</mml:mo><mml:mi>y</mml:mi></mml:mrow></mml:mfrac></mml:mstyle><mml:mfenced close=")" open="("><mml:mrow><mml:mn mathvariant="normal">4</mml:mn><mml:mi>H</mml:mi><mml:mi mathvariant="italic">μ</mml:mi><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mo>∂</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>∂</mml:mo><mml:mi>y</mml:mi></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:mfenced><mml:mspace linebreak="nobreak" width="0.25em"/><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
          where one recognizes the diffusive variables <inline-formula><mml:math id="M38" display="inline"><mml:mrow><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> and the effective dynamic viscosity <inline-formula><mml:math id="M39" display="inline"><mml:mrow><mml:mn mathvariant="normal">4</mml:mn><mml:mi mathvariant="italic">μ</mml:mi><mml:mo>/</mml:mo><mml:mi mathvariant="italic">ρ</mml:mi></mml:mrow></mml:math></inline-formula> as a diffusion coefficient. Using the analogy of a diffusive process, we can define a CFL (Courant–Friedrichs–Lewy)-like stability criterion for the PT iterative scheme. The explicit CFL-stability-based time step for viscous flow is given by the following:
            <disp-formula id="Ch1.E7" content-type="numbered"><label>7</label><mml:math id="M40" display="block"><mml:mrow><mml:mi mathvariant="normal">Δ</mml:mi><mml:msub><mml:mi mathvariant="italic">τ</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mi mathvariant="italic">ρ</mml:mi><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mi mathvariant="normal">Δ</mml:mi><mml:msup><mml:mi>x</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:mn mathvariant="normal">4</mml:mn><mml:mi mathvariant="italic">μ</mml:mi><mml:mo>(</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>+</mml:mo><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi mathvariant="normal">b</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>×</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mi mathvariant="normal">dim</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:mspace width="0.25em" linebreak="nobreak"/><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
          where <inline-formula><mml:math id="M41" display="inline"><mml:mrow><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:math></inline-formula> represents the grid spacing, <inline-formula><mml:math id="M42" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi mathvariant="normal">b</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the numerical bulk ice viscosity, and <inline-formula><mml:math id="M43" display="inline"><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi mathvariant="normal">dim</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">2.1</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">4.1</mml:mn></mml:mrow></mml:math></inline-formula>, and 6.1 in 1-D, 2-D, and 3-D, respectively.</p>
</sec>
<sec id="Ch1.S2.SS4">
  <label>2.4</label><title>Viscosity continuation</title>
      <p id="d1e1203">We implement a continuation on the nonlinear strain-rate-dependent effective viscosity <inline-formula><mml:math id="M44" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi mathvariant="normal">eff</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> to avoid the iterative solution process diverging, as strain-rate values may not satisfy  the momentum balance at the beginning of the iterative process and may, thus, be far from equilibrium. At every pseudo-time step, the effective viscosity <inline-formula><mml:math id="M45" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi mathvariant="normal">eff</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is updated in the logarithmic space:
            <disp-formula id="Ch1.E8" content-type="numbered"><label>8</label><mml:math id="M46" display="block"><mml:mrow><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi mathvariant="normal">eff</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mi>exp⁡</mml:mi><mml:mfenced close=")" open="("><mml:mrow><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mi mathvariant="italic">μ</mml:mi></mml:msub><mml:mi>log⁡</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">μ</mml:mi><mml:mo>)</mml:mo><mml:mo>+</mml:mo><mml:mo>(</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mi mathvariant="italic">μ</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mi>log⁡</mml:mi><mml:mo>(</mml:mo><mml:msubsup><mml:mi mathvariant="italic">μ</mml:mi><mml:mi mathvariant="normal">eff</mml:mi><mml:mi mathvariant="normal">old</mml:mi></mml:msubsup></mml:mrow></mml:mfenced><mml:mspace linebreak="nobreak" width="0.25em"/><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
          where the scalar <inline-formula><mml:math id="M47" display="inline"><mml:mrow><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:msup><mml:mo>&lt;</mml:mo><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mi mathvariant="italic">μ</mml:mi></mml:msub><mml:mo>&lt;</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> is selected such that we provide sufficient time to relax the nonlinear viscosity at the start of the pseudo-iterative loop.</p>
</sec>
<sec id="Ch1.S2.SS5">
  <label>2.5</label><title>Acceleration owing to damping</title>
      <p id="d1e1322">The major limitation of this simple first-order, or Picard-type, iterative approach resides in the poor iteration count scaling with increased numerical resolution. The number of iterations needed to converge for a given problem for <inline-formula><mml:math id="M48" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula> number of grid points involved in the computation scales in the order of <inline-formula><mml:math id="M49" display="inline"><mml:mrow><mml:mi mathvariant="script">O</mml:mi><mml:mo>(</mml:mo><mml:msup><mml:mi>N</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>.</p>
      <?pagebreak page902?><p id="d1e1349">To address this limitation, we consider a second-order method, referred to as the second-order Richardson method, as introduced by <xref ref-type="bibr" rid="bib1.bibx7" id="text.14"/>. This approach allows us to aggressively reduce the number of iterations to the number of grid points, making the method scale as <inline-formula><mml:math id="M50" display="inline"><mml:mrow><mml:mo>≈</mml:mo><mml:mi mathvariant="script">O</mml:mi><mml:mo>(</mml:mo><mml:msup><mml:mi>N</mml:mi><mml:mn mathvariant="normal">1.2</mml:mn></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. Optimal scaling can be achieved by realizing that the PT framework's diffusion type of updates readily provided can be divided into two wave-like update steps. Transitioning from diffusion to wave-like pseudo-physics exhibits two main advantages: (i) the wave-like time step limit is a function of <inline-formula><mml:math id="M51" display="inline"><mml:mrow><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:math></inline-formula> instead of <inline-formula><mml:math id="M52" display="inline"><mml:mrow><mml:mi mathvariant="normal">Δ</mml:mi><mml:msup><mml:mi>x</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>; (ii) it is possible to turn the wave equation into a damped wave equation. The latter permits one to find optimal tuning parameters to achieve optimal damping, resulting in fast convergence. Let us assume the following diffusion-like update step, reported here for the <inline-formula><mml:math id="M53" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> direction only:

                <disp-formula id="Ch1.E9" content-type="numbered"><label>9</label><mml:math id="M54" display="block"><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mi mathvariant="italic">ρ</mml:mi><mml:mi>H</mml:mi><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mo>∂</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>∂</mml:mo><mml:mi mathvariant="italic">τ</mml:mi></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mo>∂</mml:mo><mml:mrow><mml:mo>∂</mml:mo><mml:mi>x</mml:mi></mml:mrow></mml:mfrac></mml:mstyle><mml:mn mathvariant="normal">4</mml:mn><mml:mi>H</mml:mi><mml:mi mathvariant="italic">μ</mml:mi><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mo>∂</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>∂</mml:mo><mml:mi>x</mml:mi></mml:mrow></mml:mfrac></mml:mstyle><mml:mspace width="0.25em" linebreak="nobreak"/><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>

          The above expression results in the following update rule:

                <disp-formula id="Ch1.E10" content-type="numbered"><label>10</label><mml:math id="M55" display="block"><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:msub><mml:mi>v</mml:mi><mml:mi>x</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msubsup><mml:mi>v</mml:mi><mml:mi>x</mml:mi><mml:mi mathvariant="normal">old</mml:mi></mml:msubsup><mml:mo>+</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mi mathvariant="normal">Δ</mml:mi><mml:msub><mml:mi mathvariant="italic">τ</mml:mi><mml:mi mathvariant="normal">D</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mi mathvariant="italic">ρ</mml:mi><mml:mi>H</mml:mi></mml:mrow></mml:mfrac></mml:mstyle><mml:mfenced open="(" close=")"><mml:mrow><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mo>∂</mml:mo><mml:mrow><mml:mo>∂</mml:mo><mml:mi>x</mml:mi></mml:mrow></mml:mfrac></mml:mstyle><mml:mn mathvariant="normal">4</mml:mn><mml:mi>H</mml:mi><mml:mi mathvariant="italic">μ</mml:mi><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mo>∂</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>∂</mml:mo><mml:mi>x</mml:mi></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:mfenced><mml:mspace width="0.25em" linebreak="nobreak"/><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

          where <inline-formula><mml:math id="M56" display="inline"><mml:mrow><mml:mi mathvariant="normal">Δ</mml:mi><mml:msub><mml:mi mathvariant="italic">τ</mml:mi><mml:mi mathvariant="normal">D</mml:mi></mml:msub><mml:mo>≈</mml:mo><mml:mi mathvariant="normal">Δ</mml:mi><mml:msup><mml:mi>x</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>/</mml:mo><mml:mo>(</mml:mo><mml:mn mathvariant="normal">4</mml:mn><mml:mi mathvariant="italic">μ</mml:mi><mml:mo>/</mml:mo><mml:mi mathvariant="italic">ρ</mml:mi><mml:mo>)</mml:mo><mml:mo>/</mml:mo><mml:mn mathvariant="normal">4.1</mml:mn></mml:mrow></mml:math></inline-formula> is the diffusion-like time step limit. This system can be separated into a residual assignment <inline-formula><mml:math id="M57" display="inline"><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and the velocity update <inline-formula><mml:math id="M58" display="inline"><mml:mrow><mml:msub><mml:mi>v</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>:

                <disp-formula specific-use="align" content-type="numbered"><mml:math id="M59" display="block"><mml:mtable rowspacing="5pt" displaystyle="true"><mml:mlabeledtr id="Ch1.E11"><mml:mtd><mml:mtext>11</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:msub><mml:mi>A</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:mi mathvariant="italic">ρ</mml:mi><mml:mi>H</mml:mi></mml:mrow></mml:mfrac></mml:mstyle><mml:mfenced close=")" open="("><mml:mrow><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mo>∂</mml:mo><mml:mrow><mml:mo>∂</mml:mo><mml:mi>x</mml:mi></mml:mrow></mml:mfrac></mml:mstyle><mml:mn mathvariant="normal">4</mml:mn><mml:mi>H</mml:mi><mml:mi mathvariant="italic">μ</mml:mi><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mo>∂</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>∂</mml:mo><mml:mi>x</mml:mi></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:mfenced><mml:mspace width="0.25em" linebreak="nobreak"/><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr><mml:mlabeledtr id="Ch1.E12"><mml:mtd><mml:mtext>12</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:msub><mml:mi>v</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mo>=</mml:mo><mml:msubsup><mml:mi>v</mml:mi><mml:mi>x</mml:mi><mml:mi mathvariant="normal">old</mml:mi></mml:msubsup><mml:mo>+</mml:mo><mml:mi mathvariant="normal">Δ</mml:mi><mml:msub><mml:mi mathvariant="italic">τ</mml:mi><mml:mi mathvariant="normal">D</mml:mi></mml:msub><mml:msub><mml:mi>A</mml:mi><mml:mi>x</mml:mi></mml:msub><mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula>

            Converting Eq. (<xref ref-type="disp-formula" rid="Ch1.E11"/>) into an update rule using a step size of <inline-formula><mml:math id="M60" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:mi mathvariant="italic">γ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>,

                <disp-formula id="Ch1.E13" content-type="numbered"><label>13</label><mml:math id="M61" display="block"><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:msub><mml:mi>A</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mo>=</mml:mo><mml:msubsup><mml:mi>A</mml:mi><mml:mi>x</mml:mi><mml:mi mathvariant="normal">old</mml:mi></mml:msubsup><mml:mo>(</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:mi mathvariant="italic">γ</mml:mi><mml:mo>)</mml:mo><mml:mo>+</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:mi mathvariant="italic">ρ</mml:mi><mml:mi>H</mml:mi></mml:mrow></mml:mfrac></mml:mstyle><mml:mfenced open="(" close=")"><mml:mrow><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mo>∂</mml:mo><mml:mrow><mml:mo>∂</mml:mo><mml:mi>x</mml:mi></mml:mrow></mml:mfrac></mml:mstyle><mml:mn mathvariant="normal">4</mml:mn><mml:mi>H</mml:mi><mml:mi mathvariant="italic">μ</mml:mi><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mo>∂</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>∂</mml:mo><mml:mi>x</mml:mi></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:mfenced><mml:mspace linebreak="nobreak" width="0.25em"/><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

          turns the system composed of Eqs. (<xref ref-type="disp-formula" rid="Ch1.E12"/>) and (<xref ref-type="disp-formula" rid="Ch1.E13"/>) into a damped wave equation similar to what was suggested by <xref ref-type="bibr" rid="bib1.bibx7" id="text.15"/>. Ideal convergence can be reached upon selecting the appropriate damping parameter <inline-formula><mml:math id="M62" display="inline"><mml:mi mathvariant="italic">γ</mml:mi></mml:math></inline-formula>. To maintain solution stability, we include relaxation <inline-formula><mml:math id="M63" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mi>v</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>:

                <disp-formula id="Ch1.E14" content-type="numbered"><label>14</label><mml:math id="M64" display="block"><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:msub><mml:mi>v</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mo>=</mml:mo><mml:msubsup><mml:mi>v</mml:mi><mml:mi>x</mml:mi><mml:mi mathvariant="normal">old</mml:mi></mml:msubsup><mml:mo>+</mml:mo><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mi>v</mml:mi></mml:msub><mml:mi mathvariant="normal">Δ</mml:mi><mml:msub><mml:mi mathvariant="italic">τ</mml:mi><mml:mi mathvariant="normal">D</mml:mi></mml:msub><mml:msub><mml:mi>A</mml:mi><mml:mi>x</mml:mi></mml:msub><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

          where <inline-formula><mml:math id="M65" display="inline"><mml:mrow><mml:mn mathvariant="normal">0</mml:mn><mml:mo>&lt;</mml:mo><mml:mi mathvariant="italic">γ</mml:mi><mml:mo>&lt;</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M66" display="inline"><mml:mrow><mml:mn mathvariant="normal">0</mml:mn><mml:mo>&lt;</mml:mo><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mi>v</mml:mi></mml:msub><mml:mo>&lt;</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula>.</p>
      <p id="d1e1914">Alternative and complementary details about the PT acceleration can be found in <xref ref-type="bibr" rid="bib1.bibx25" id="text.16"/>, <xref ref-type="bibr" rid="bib1.bibx6" id="text.17"/>, and <xref ref-type="bibr" rid="bib1.bibx26" id="text.18"/>, while an in-depth analysis is provided in <xref ref-type="bibr" rid="bib1.bibx27" id="text.19"/>.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F1"><?xmltex \currentcnt{1}?><?xmltex \def\figurename{Figure}?><label>Figure 1</label><caption><p id="d1e1932">PT iterative algorithm for unstructured meshes applied to solve 2-D SSA momentum balance equations.</p></caption>
          <?xmltex \igopts{width=213.395669pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/17/899/2024/gmd-17-899-2024-f01.png"/>

        </fig>

</sec>
<sec id="Ch1.S2.SS6">
  <label>2.6</label><title>Weak formulation and finite-element discretization</title>
      <p id="d1e1949">Using the PT method, the equations to solve are as referenced in Eq. (<xref ref-type="disp-formula" rid="Ch1.E4"/>):
            <disp-formula id="Ch1.E15" content-type="numbered"><label>15</label><mml:math id="M67" display="block"><mml:mrow><mml:mi mathvariant="italic">ρ</mml:mi><mml:mi>H</mml:mi><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mo>∂</mml:mo><mml:mi mathvariant="bold-italic">v</mml:mi></mml:mrow><mml:mrow><mml:mo>∂</mml:mo><mml:mi mathvariant="italic">τ</mml:mi></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>=</mml:mo><mml:mi mathvariant="normal">∇</mml:mi><mml:mo>⋅</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:mi>H</mml:mi><mml:mi mathvariant="italic">μ</mml:mi><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">ε</mml:mi><mml:mo mathvariant="normal">˙</mml:mo></mml:mover><mml:mi mathvariant="normal">SSA</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mi mathvariant="italic">ρ</mml:mi><mml:mi>g</mml:mi><mml:mi>H</mml:mi><mml:mi mathvariant="normal">∇</mml:mi><mml:mi>s</mml:mi><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="italic">α</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mi mathvariant="bold-italic">v</mml:mi><mml:mspace width="0.25em" linebreak="nobreak"/><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula></p>
      <p id="d1e2020">The weak form of the equation, assuming homogeneous Dirichlet conditions along all model boundaries for simplicity, reads: <inline-formula><mml:math id="M68" display="inline"><mml:mrow><mml:mo>∀</mml:mo><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mo>∈</mml:mo><mml:msup><mml:mi mathvariant="script">H</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msup><mml:mfenced open="(" close=")"><mml:mi mathvariant="normal">Ω</mml:mi></mml:mfenced></mml:mrow></mml:math></inline-formula>,
            <disp-formula id="Ch1.E16" content-type="numbered"><label>16</label><mml:math id="M69" display="block"><mml:mtable rowspacing="0.2ex" class="split" displaystyle="true" columnalign="right left"><mml:mtr><mml:mtd><mml:mrow><mml:munder><mml:mo movablelimits="false">∫</mml:mo><mml:mi mathvariant="normal">Ω</mml:mi></mml:munder><mml:mi mathvariant="italic">ρ</mml:mi><mml:mi>H</mml:mi><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mo>∂</mml:mo><mml:mi mathvariant="bold-italic">v</mml:mi></mml:mrow><mml:mrow><mml:mo>∂</mml:mo><mml:mi mathvariant="italic">τ</mml:mi></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>⋅</mml:mo><mml:mi mathvariant="bold">w</mml:mi><mml:mi mathvariant="normal">d</mml:mi><mml:mi mathvariant="normal">Ω</mml:mi></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mo>+</mml:mo><mml:munder><mml:mo movablelimits="false">∫</mml:mo><mml:mi mathvariant="normal">Ω</mml:mi></mml:munder><mml:mn mathvariant="normal">2</mml:mn><mml:mi>H</mml:mi><mml:mi mathvariant="italic">μ</mml:mi><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">ε</mml:mi><mml:mo mathvariant="normal">˙</mml:mo></mml:mover><mml:mi mathvariant="normal">SSA</mml:mi></mml:msub><mml:mo>⋅</mml:mo><mml:mi mathvariant="normal">∇</mml:mi><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mspace linebreak="nobreak" width="0.25em"/><mml:mi mathvariant="normal">d</mml:mi><mml:mi mathvariant="normal">Ω</mml:mi><mml:mo>=</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:munder><mml:mo movablelimits="false">∫</mml:mo><mml:mi mathvariant="normal">Ω</mml:mi></mml:munder><mml:mo>-</mml:mo><mml:mi mathvariant="italic">ρ</mml:mi><mml:mi>g</mml:mi><mml:mi>H</mml:mi><mml:mi mathvariant="normal">∇</mml:mi><mml:mi>s</mml:mi><mml:mo>⋅</mml:mo><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="italic">α</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mi mathvariant="bold-italic">v</mml:mi><mml:mo>⋅</mml:mo><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mspace width="0.25em" linebreak="nobreak"/><mml:mi mathvariant="normal">d</mml:mi><mml:mi mathvariant="normal">Ω</mml:mi><mml:mspace linebreak="nobreak" width="0.25em"/><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
          where <inline-formula><mml:math id="M70" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="script">H</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msup><mml:mfenced open="(" close=")"><mml:mi mathvariant="normal">Ω</mml:mi></mml:mfenced></mml:mrow></mml:math></inline-formula> is the space of square-integrable functions whose first derivatives are also square integrable.</p>
      <p id="d1e2179">Once discretized using the finite-element method, the matrix system to solve reads:
            <disp-formula id="Ch1.E17" content-type="numbered"><label>17</label><mml:math id="M71" display="block"><mml:mrow><mml:mi mathvariant="bold">M</mml:mi><mml:mover accent="true"><mml:mi mathvariant="bold-italic">V</mml:mi><mml:mo mathvariant="normal">˙</mml:mo></mml:mover><mml:mo>+</mml:mo><mml:mi mathvariant="bold">K</mml:mi><mml:mi mathvariant="bold-italic">V</mml:mi><mml:mo>=</mml:mo><mml:mi mathvariant="bold-italic">F</mml:mi><mml:mspace width="0.25em" linebreak="nobreak"/><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
          where <inline-formula><mml:math id="M72" display="inline"><mml:mi mathvariant="bold-italic">M</mml:mi></mml:math></inline-formula> is the mass matrix, <inline-formula><mml:math id="M73" display="inline"><mml:mi mathvariant="bold">K</mml:mi></mml:math></inline-formula> is the stiffness matrix, <inline-formula><mml:math id="M74" display="inline"><mml:mi mathvariant="bold-italic">F</mml:mi></mml:math></inline-formula> is the right-hand-side or load vector, and <inline-formula><mml:math id="M75" display="inline"><mml:mi mathvariant="bold-italic">V</mml:mi></mml:math></inline-formula> is the vector of ice velocity.</p>
      <?pagebreak page903?><p id="d1e2236">We can compute <inline-formula><mml:math id="M76" display="inline"><mml:mover accent="true"><mml:mi mathvariant="bold">V</mml:mi><mml:mo mathvariant="normal">˙</mml:mo></mml:mover></mml:math></inline-formula> by solving
            <disp-formula id="Ch1.E18" content-type="numbered"><label>18</label><mml:math id="M77" display="block"><mml:mrow><mml:mover accent="true"><mml:mi mathvariant="bold">V</mml:mi><mml:mo mathvariant="normal">˙</mml:mo></mml:mover><mml:mo>≃</mml:mo><mml:msubsup><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">L</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msubsup><mml:mfenced close=")" open="("><mml:mrow><mml:mo>-</mml:mo><mml:mi mathvariant="bold">KV</mml:mi><mml:mo>+</mml:mo><mml:mi mathvariant="bold-italic">F</mml:mi></mml:mrow></mml:mfenced><mml:mspace width="0.25em" linebreak="nobreak"/><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
          where <inline-formula><mml:math id="M78" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">L</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> stands for the lumped mass matrix that permits one to avoid the resolution of a matrix system.</p>
      <p id="d1e2297">Hence, we have an explicit expression of the time derivative of the ice velocity for each vertex of the mesh:

                <disp-formula specific-use="gather" content-type="numbered"><mml:math id="M79" display="block"><mml:mtable displaystyle="true"><mml:mlabeledtr id="Ch1.E19"><mml:mtd><mml:mtext>19</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mtable rowspacing="0.2ex" class="split" displaystyle="true" columnalign="right left"><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi>v</mml:mi><mml:mo mathvariant="normal">˙</mml:mo></mml:mover><mml:mrow><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:mi mathvariant="italic">ρ</mml:mi><mml:mi>H</mml:mi><mml:msub><mml:mi>m</mml:mi><mml:mrow><mml:mi mathvariant="normal">L</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:mfenced close="" open="("><mml:mrow><mml:mo>-</mml:mo><mml:munder><mml:mo movablelimits="false">∫</mml:mo><mml:mi mathvariant="normal">Ω</mml:mi></mml:munder><mml:mfenced open="(" close=")"><mml:mrow><mml:mn mathvariant="normal">4</mml:mn><mml:mi>H</mml:mi><mml:mi mathvariant="italic">μ</mml:mi><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mo>∂</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>∂</mml:mo><mml:mi>x</mml:mi></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>+</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:mi>H</mml:mi><mml:mi mathvariant="italic">μ</mml:mi><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mo>∂</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>∂</mml:mo><mml:mi>y</mml:mi></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:mfenced><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mo>∂</mml:mo><mml:msub><mml:mi mathvariant="italic">φ</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>∂</mml:mo><mml:mi>x</mml:mi></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:mfenced></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mo>+</mml:mo><mml:mfenced close=")" open="("><mml:mrow><mml:mi>H</mml:mi><mml:mi mathvariant="italic">μ</mml:mi><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mo>∂</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>∂</mml:mo><mml:mi>y</mml:mi></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>+</mml:mo><mml:mi>H</mml:mi><mml:mi mathvariant="italic">μ</mml:mi><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mo>∂</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>∂</mml:mo><mml:mi>x</mml:mi></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:mfenced><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mo>∂</mml:mo><mml:msub><mml:mi mathvariant="italic">φ</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>∂</mml:mo><mml:mi>y</mml:mi></mml:mrow></mml:mfrac></mml:mstyle><mml:mi mathvariant="normal">d</mml:mi><mml:mi mathvariant="normal">Ω</mml:mi></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mo>+</mml:mo><mml:mfenced close=")" open=""><mml:mrow><mml:munder><mml:mo movablelimits="false">∫</mml:mo><mml:mi mathvariant="normal">Ω</mml:mi></mml:munder><mml:mo>-</mml:mo><mml:mi mathvariant="italic">ρ</mml:mi><mml:mi>g</mml:mi><mml:mi>H</mml:mi><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mo>∂</mml:mo><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mo>∂</mml:mo><mml:mi>x</mml:mi></mml:mrow></mml:mfrac></mml:mstyle><mml:msub><mml:mi mathvariant="italic">φ</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="italic">α</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:msub><mml:mi>v</mml:mi><mml:mi>x</mml:mi></mml:msub><mml:msub><mml:mi mathvariant="italic">φ</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mspace linebreak="nobreak" width="0.25em"/><mml:mi mathvariant="normal">d</mml:mi><mml:mi mathvariant="normal">Ω</mml:mi></mml:mrow></mml:mfenced><mml:mspace linebreak="nobreak" width="0.25em"/><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mtd></mml:mlabeledtr><mml:mlabeledtr id="Ch1.E20"><mml:mtd><mml:mtext>20</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mtable rowspacing="0.2ex" class="split" displaystyle="true" columnalign="right left"><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi>v</mml:mi><mml:mo mathvariant="normal">˙</mml:mo></mml:mover><mml:mrow><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:mi mathvariant="italic">ρ</mml:mi><mml:mi>H</mml:mi><mml:msub><mml:mi>m</mml:mi><mml:mrow><mml:mi mathvariant="normal">L</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:mfenced close="" open="("><mml:mrow><mml:mo>-</mml:mo><mml:munder><mml:mo movablelimits="false">∫</mml:mo><mml:mi mathvariant="normal">Ω</mml:mi></mml:munder><mml:mfenced open="(" close=")"><mml:mrow><mml:mn mathvariant="normal">4</mml:mn><mml:mi>H</mml:mi><mml:mi mathvariant="italic">μ</mml:mi><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mo>∂</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>∂</mml:mo><mml:mi>y</mml:mi></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>+</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:mi>H</mml:mi><mml:mi mathvariant="italic">μ</mml:mi><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mo>∂</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>∂</mml:mo><mml:mi>x</mml:mi></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:mfenced><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mo>∂</mml:mo><mml:msub><mml:mi mathvariant="italic">φ</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>∂</mml:mo><mml:mi>y</mml:mi></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:mfenced></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mo>+</mml:mo><mml:mfenced close=")" open="("><mml:mrow><mml:mi>H</mml:mi><mml:mi mathvariant="italic">μ</mml:mi><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mo>∂</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>∂</mml:mo><mml:mi>y</mml:mi></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>+</mml:mo><mml:mi>H</mml:mi><mml:mi mathvariant="italic">μ</mml:mi><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mo>∂</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>∂</mml:mo><mml:mi>x</mml:mi></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:mfenced><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mo>∂</mml:mo><mml:msub><mml:mi mathvariant="italic">φ</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>∂</mml:mo><mml:mi>x</mml:mi></mml:mrow></mml:mfrac></mml:mstyle><mml:mi mathvariant="normal">d</mml:mi><mml:mi mathvariant="normal">Ω</mml:mi></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mfenced close=")" open=""><mml:mrow><mml:mo>+</mml:mo><mml:munder><mml:mo movablelimits="false">∫</mml:mo><mml:mi mathvariant="normal">Ω</mml:mi></mml:munder><mml:mo>-</mml:mo><mml:mi mathvariant="italic">ρ</mml:mi><mml:mi>g</mml:mi><mml:mi>H</mml:mi><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mo>∂</mml:mo><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mo>∂</mml:mo><mml:mi>y</mml:mi></mml:mrow></mml:mfrac></mml:mstyle><mml:msub><mml:mi mathvariant="italic">φ</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="italic">α</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:msub><mml:mi>v</mml:mi><mml:mi>y</mml:mi></mml:msub><mml:msub><mml:mi mathvariant="italic">φ</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mi mathvariant="normal">d</mml:mi><mml:mi mathvariant="normal">Ω</mml:mi></mml:mrow></mml:mfenced><mml:mspace linebreak="nobreak" width="0.25em"/><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula>

            where <inline-formula><mml:math id="M80" display="inline"><mml:mrow><mml:msub><mml:mi>m</mml:mi><mml:mrow><mml:mi mathvariant="normal">L</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> is the component number <inline-formula><mml:math id="M81" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula> along the diagonal of the lumped mass matrix <inline-formula><mml:math id="M82" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mi mathvariant="normal">L</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>.</p>
      <p id="d1e2825">For every nonlinear PT iteration,  we compute the rate of change in velocity <inline-formula><mml:math id="M83" display="inline"><mml:mover accent="true"><mml:mi mathvariant="bold-italic">v</mml:mi><mml:mo mathvariant="normal">˙</mml:mo></mml:mover></mml:math></inline-formula> and the explicit CFL-stability-based time step <inline-formula><mml:math id="M84" display="inline"><mml:mrow><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi mathvariant="italic">τ</mml:mi></mml:mrow></mml:math></inline-formula>. We then deploy the reformulated 2-D SSA momentum balance equations to update ice velocity <inline-formula><mml:math id="M85" display="inline"><mml:mi mathvariant="bold-italic">v</mml:mi></mml:math></inline-formula> followed by ice viscosity <inline-formula><mml:math id="M86" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi mathvariant="normal">eff</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. We iterate in pseudo-time until the stopping criterion is met (Fig. <xref ref-type="fig" rid="Ch1.F1"/>).</p>
</sec>
</sec>
<sec id="Ch1.S3">
  <label>3</label><title>Numerical experiments</title>
<sec id="Ch1.S3.SS1">
  <label>3.1</label><title>Glacier model configurations</title>
      <p id="d1e2884">To test the performance of the PT method beyond simple idealized geometries, we apply it to two regional-scale glaciers: Jakobshavn Isbræ, in Western Greenland, and Pine Island Glacier, in West Antarctica (Fig. <xref ref-type="fig" rid="Ch1.F2"/>). For Jakobshavn Isbræ, we rely on BedMachine Greenland v4 <xref ref-type="bibr" rid="bib1.bibx20" id="paren.20"/> and also invert for basal friction to infer the basal boundary conditions. Note that the inversion is run on the Ice-sheet and Sea-level System Model (ISSM), using a standard approach <xref ref-type="bibr" rid="bib1.bibx18" id="paren.21"/>. For Pine Island Glacier, we initialize the ice geometry using BedMachine Antarctica v2 <xref ref-type="bibr" rid="bib1.bibx21" id="paren.22"/> and infer the friction coefficient using surface velocities derived from satellite interferometry <xref ref-type="bibr" rid="bib1.bibx29" id="paren.23"/>.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F2"><?xmltex \currentcnt{2}?><?xmltex \def\figurename{Figure}?><label>Figure 2</label><caption><p id="d1e2903">Glacier model configurations: observed surface velocities (in m yr<inline-formula><mml:math id="M87" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>) interpolated on a uniform mesh. Panels <bold>(a)</bold> and <bold>(b)</bold>  correspond to Jakobshavn Isbræ and Pine Island Glacier, respectively.</p></caption>
          <?xmltex \igopts{width=241.848425pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/17/899/2024/gmd-17-899-2024-f02.png"/>

        </fig>

</sec>
<sec id="Ch1.S3.SS2">
  <label>3.2</label><title>Hardware implementation</title>
      <p id="d1e2938">We developed a CUDA C implementation to solve the SSA equations using the PT approach on unstructured meshes. We choose a stopping criterion of <inline-formula><mml:math id="M88" display="inline"><mml:mrow><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msup><mml:mi>v</mml:mi><mml:mi mathvariant="normal">old</mml:mi></mml:msup><mml:mo>-</mml:mo><mml:mi>v</mml:mi><mml:mo>|</mml:mo><mml:msub><mml:mo>|</mml:mo><mml:mi mathvariant="normal">∞</mml:mi></mml:msub><mml:mo>&lt;</mml:mo><mml:mn mathvariant="normal">10</mml:mn></mml:mrow></mml:math></inline-formula> m yr<inline-formula><mml:math id="M89" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>. The software solves the 2-D SSA momentum balance equations on a single GPU. We use an NVIDIA Tesla V100 SXM2 GPU with 16 GB (gigabytes) of device RAM and an NVIDIA A100 SXM4 with 80 GB of device RAM. We compare the PT implementation's results on a Tesla V100 GPU with ISSM's “standard” CPU implementation using a conjugate gradient (CG) iterative solver. We used a 64-bit 18-core Intel Xeon Gold 6140 processor for the CPU comparison, with 192 GB of RAM available. We perform multicore Message Passing Interface (MPI)-parallelized ice-sheet flow simulations on two CPUs, with all 36 cores enabled <xref ref-type="bibr" rid="bib1.bibx18 bib1.bibx11" id="paren.24"/>. All simulations use double-precision arithmetic computations.</p>
</sec>
<sec id="Ch1.S3.SS3">
  <label>3.3</label><title>Performance assessment metrics</title>
      <p id="d1e2994">To investigate the PT CUDA C implementation for unstructured meshes, we report the number of vertices (or grid size) and the corresponding number of nonlinear PT iterations needed to meet the stopping criterion. We employ the computational time required to reach convergence as a proxy to assess and compare the performance of the PT CUDA C with the ISSM CG CPU implementation. We make sure to exclude all pre- and post-processing steps from the timing. We quantify the relative performance of the CPU and GPU implementations as the speedup (<inline-formula><mml:math id="M90" display="inline"><mml:mi>S</mml:mi></mml:math></inline-formula>), given by the following:
            <disp-formula id="Ch1.E21" content-type="numbered"><label>21</label><mml:math id="M91" display="block"><mml:mrow><mml:mi>S</mml:mi><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mi mathvariant="normal">CPU</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mi mathvariant="normal">GPU</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:mspace width="0.25em" linebreak="nobreak"/><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula></p>
      <p id="d1e3030">The PT method employed to solve the nonlinear momentum balance equations results in a memory-bound algorithm <xref ref-type="bibr" rid="bib1.bibx26" id="paren.25"/>; therefore, the wall time depends on the memory throughput. In addition to the speedup, we employ the effective memory throughput metric to assess the<?pagebreak page904?> performance of the PT CUDA C implementation developed in this study  <xref ref-type="bibr" rid="bib1.bibx26 bib1.bibx27" id="paren.26"/>, which is defined as follows:

                <disp-formula id="Ch1.E22" content-type="numbered"><label>22</label><mml:math id="M92" display="block"><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:msub><mml:mi>T</mml:mi><mml:mi mathvariant="normal">eff</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi mathvariant="normal">n</mml:mi></mml:msub><mml:mspace linebreak="nobreak" width="0.25em"/><mml:msub><mml:mi>n</mml:mi><mml:mi mathvariant="normal">iter</mml:mi></mml:msub><mml:mspace linebreak="nobreak" width="0.25em"/><mml:msub><mml:mi>n</mml:mi><mml:mi mathvariant="normal">IO</mml:mi></mml:msub><mml:mspace linebreak="nobreak" width="0.25em"/><mml:msub><mml:mi>n</mml:mi><mml:mi mathvariant="normal">p</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:msup><mml:mn mathvariant="normal">1024</mml:mn><mml:mn mathvariant="normal">3</mml:mn></mml:msup><mml:mspace linebreak="nobreak" width="0.25em"/><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi mathvariant="normal">iter</mml:mi></mml:msub></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:mspace linebreak="nobreak" width="0.25em"/><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

          where <inline-formula><mml:math id="M93" display="inline"><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi mathvariant="normal">n</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> represents the total number of vertices,  <inline-formula><mml:math id="M94" display="inline"><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi mathvariant="normal">iter</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is a given number of PT iterations, <inline-formula><mml:math id="M95" display="inline"><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi mathvariant="normal">p</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the arithmetic precision, <inline-formula><mml:math id="M96" display="inline"><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi mathvariant="normal">iter</mml:mi></mml:msub></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> is the time taken to complete <inline-formula><mml:math id="M97" display="inline"><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi mathvariant="normal">iter</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> iterations, and <inline-formula><mml:math id="M98" display="inline"><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi mathvariant="normal">IO</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the minimal number of nonredundant memory accesses (read and write operations). The number of read and write operations needed for this study would be eight: update <inline-formula><mml:math id="M99" display="inline"><mml:mrow><mml:msub><mml:mi>v</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M100" display="inline"><mml:mrow><mml:msub><mml:mi>v</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, and nonlinear viscosity arrays for every PT iteration, in addition to reading the basal friction coefficient and the masks.</p>
</sec>
</sec>
<sec id="Ch1.S4">
  <label>4</label><title>Results and discussion</title>
      <p id="d1e3200">To investigate the performance of the PT CUDA C implementation on unstructured meshes, we report the number of vertices (or grid size) and the corresponding number of nonlinear PT iterations needed to meet the stopping criterion (Fig. <xref ref-type="fig" rid="Ch1.F3"/>). For both the Jakobshavn and Pine Island glacier models, the number of nonlinear PT iterations required to converge for a given number of vertices (<inline-formula><mml:math id="M101" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula>) scales in the order of <inline-formula><mml:math id="M102" display="inline"><mml:mrow><mml:mo>≈</mml:mo><mml:mi mathvariant="script">O</mml:mi><mml:mo>(</mml:mo><mml:msup><mml:mi>N</mml:mi><mml:mn mathvariant="normal">1.2</mml:mn></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> or better. We chose the damping parameter <inline-formula><mml:math id="M103" display="inline"><mml:mi mathvariant="italic">γ</mml:mi></mml:math></inline-formula>,  nonlinear viscosity relaxation scalar <inline-formula><mml:math id="M104" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mi mathvariant="italic">μ</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, and transient <italic>pseudo</italic>-time step <inline-formula><mml:math id="M105" display="inline"><mml:mrow><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi mathvariant="italic">τ</mml:mi></mml:mrow></mml:math></inline-formula> to maintain the linear scaling described above; optimal parameter values are listed in the Appendix (Table A1). We observed an exception at <inline-formula><mml:math id="M106" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">3</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mn mathvariant="normal">7</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> degrees of freedom (DoFs) for the Pine Island Glacier model; optimal parameter values are unidentifiable. We will further investigate the convergence for the Pine Island Glacier model at <inline-formula><mml:math id="M107" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">3</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mn mathvariant="normal">7</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> DoFs. Among the two glacier models chosen in this study, for a given number of vertices (<inline-formula><mml:math id="M108" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula>), Jakobshavn Isbræ resulted in faster convergence rates, which we attribute to differences in scale and bed topography and the nonlinearity of the problem (Fig. <xref ref-type="fig" rid="Ch1.F4"/>).</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F3"><?xmltex \currentcnt{3}?><?xmltex \def\figurename{Figure}?><label>Figure 3</label><caption><p id="d1e3308">Performance assessment of the PT CUDA C implementation for unstructured meshes.</p></caption>
        <?xmltex \igopts{width=241.848425pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/17/899/2024/gmd-17-899-2024-f03.png"/>

      </fig>

      <?xmltex \floatpos{t}?><fig id="Ch1.F4"><?xmltex \currentcnt{4}?><?xmltex \def\figurename{Figure}?><label>Figure 4</label><caption><p id="d1e3319">Residual error evolution of the PT CUDA C implementation for unstructured meshes.</p></caption>
        <?xmltex \igopts{width=241.848425pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/17/899/2024/gmd-17-899-2024-f04.png"/>

      </fig>

      <?xmltex \floatpos{t}?><fig id="Ch1.F5"><?xmltex \currentcnt{5}?><?xmltex \def\figurename{Figure}?><label>Figure 5</label><caption><p id="d1e3331">Performance comparison of the PT Tesla V100 implementation with the CPU implementation employing wall time (or computational time to reach convergence).  Note that wall time does not include pre- and post-processing steps.</p></caption>
        <?xmltex \igopts{width=241.848425pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/17/899/2024/gmd-17-899-2024-f05.png"/>

      </fig>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T1"><?xmltex \currentcnt{1}?><label>Table 1</label><caption><p id="d1e3343">Performance comparison of the PT Tesla V100 implementation with the CPU implementation employing speedup <inline-formula><mml:math id="M109" display="inline"><mml:mi>S</mml:mi></mml:math></inline-formula>.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="4">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:thead>
       <oasis:row>
         <oasis:entry colname="col1">Jakobshavn Isbræ</oasis:entry>
         <oasis:entry colname="col2">Speedup</oasis:entry>
         <oasis:entry colname="col3">Pine Island</oasis:entry>
         <oasis:entry colname="col4">Speedup</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">DoFs</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M110" display="inline"><mml:mi>S</mml:mi></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">Glacier   DoFs</oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M111" display="inline"><mml:mi>S</mml:mi></mml:math></inline-formula></oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">88 458</oasis:entry>
         <oasis:entry colname="col2">3.6</oasis:entry>
         <oasis:entry colname="col3">28 920</oasis:entry>
         <oasis:entry colname="col4">3.73</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">329 362</oasis:entry>
         <oasis:entry colname="col2">11.68</oasis:entry>
         <oasis:entry colname="col3">71 292</oasis:entry>
         <oasis:entry colname="col4">15.30</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">787 542</oasis:entry>
         <oasis:entry colname="col2">2.64</oasis:entry>
         <oasis:entry colname="col3">139 578</oasis:entry>
         <oasis:entry colname="col4">1.73</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">1 335 458</oasis:entry>
         <oasis:entry colname="col2">7.37</oasis:entry>
         <oasis:entry colname="col3">2 221 410</oasis:entry>
         <oasis:entry colname="col4">1.5</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">21 328 514</oasis:entry>
         <oasis:entry colname="col2">0.299</oasis:entry>
         <oasis:entry colname="col3"/>
         <oasis:entry colname="col4"/>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table><?xmltex \gdef\@currentlabel{1}?></table-wrap>

      <?xmltex \floatpos{t}?><fig id="Ch1.F6"><?xmltex \currentcnt{6}?><?xmltex \def\figurename{Figure}?><label>Figure 6</label><caption><p id="d1e3488">Performance assessment of the PT CUDA C implementation across GPU architectures employing effective memory throughput.</p></caption>
        <?xmltex \igopts{width=241.848425pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/17/899/2024/gmd-17-899-2024-f06.png"/>

      </fig>

      <p id="d1e3497">We further compare the performance of the PT CUDA C implementation with a standard finite-element CPU-based implementation using the price-to-performance metric. The price of a single Tesla V100 GPU is 1.5 times that of two Intel Xeon Gold 6140 CPU processors. <fn id="Ch1.Footn1"><p id="d1e3500">Intel Xeon Gold 6140 processor specification sheet: <uri>https://ark.intel.com/content/www/us/en/ark/products/120485/intel-xeon-gold-6140-processor-24-75m-cache-2-30-ghz.html</uri> (last access: 31 December 2023).</p></fn> We expect a minimum speedup of at least 1.5 times to justify the Tesla V100 GPU price to performance. We recorded the computational time to reach convergence for the ISSM CG CPU and the PT GPU solver implementations (Fig. <xref ref-type="fig" rid="Ch1.F5"/>) for up to <inline-formula><mml:math id="M112" display="inline"><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mn mathvariant="normal">7</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>  DoFs. Across glacier configurations, we report a speedup of <inline-formula><mml:math id="M113" display="inline"><mml:mo>&gt;</mml:mo></mml:math></inline-formula>1.5 times on the Tesla V100 GPU. We report a speedup of approximately 7 times at <inline-formula><mml:math id="M114" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mn mathvariant="normal">6</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> DoFs for the Jakobshavn glacier model. This larger speedup at <inline-formula><mml:math id="M115" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mn mathvariant="normal">6</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> DoFs indicates the PT GPU implementation's suitability to develop high-spatial-resolution ice-sheet flow models. We report an exception for the Jakobshavn glacier model at <inline-formula><mml:math id="M116" display="inline"><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mn mathvariant="normal">7</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> DoFs (speedup of 0.28 times). We suggest that readers compare the speedup results reported in this study (Table <xref ref-type="table" rid="Ch1.T1"/>) with other parallelization strategies.</p>
      <?pagebreak page905?><p id="d1e3583">The PT method applied to solve nonlinear momentum balance equations is a memory-bound algorithm, as described in Sect. <xref ref-type="sec" rid="Ch1.S3.SS3"/>. The profiling results on the Tesla V100 GPU indicate an up to 85 % increase in the device's available memory resources utilization with the increase in DoFs. This further confirms the memory-bounded nature of the implementation. To assess the performance of the memory-bound PT CUDA C implementation, we employ the effective memory throughput metric defined in Sect. <xref ref-type="sec" rid="Ch1.S3.SS3"/>. We report the effective memory throughput to DoFs for the PT CUDA C single-GPU implementation (Fig. <xref ref-type="fig" rid="Ch1.F6"/>). We observe a significant drop in effective memory throughput on both GPU architectures at DoFs <inline-formula><mml:math id="M117" display="inline"><mml:mrow><mml:mo>&gt;</mml:mo><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mn mathvariant="normal">7</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>, which explains the drop in the speedup. We attribute the drop partly to the nonoptimal global memory access patterns reported in the L1TEX and L2 cache. We identify excessive nonlocal data access patterns in the ice stiffness and strain-rate computations, which involve accessing element-to-vertex connectivities and vice versa. For optimal or fully coalesced global memory access patterns,  the threads in a warp must access the same relative address. We are investigating techniques to reduce the mesh non-localities and allow for coalesced global accesses.</p>
      <p id="d1e3608">We report a peak memory throughput for the NVIDIA Tesla V100 and A100 GPUs of 785 and 1536 GB s<inline-formula><mml:math id="M118" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>, respectively. The peak memory throughput reflects the maximal memory transfer rates for performing memory copy-only operations. It represents the hardware performance limit in a memory-bound regime. Across glacier model configurations for the DoFs chosen in this study, the PT CUDA C implementation achieves a maximum of 23 and 58 GB s<inline-formula><mml:math id="M119" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> for the NVIDIA Tesla V100 and A100 GPUs, respectively. The measured memory throughput is in the order of 500 GB s<inline-formula><mml:math id="M120" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>, as reported by the NVIDIA Nsight Compute profiling tool 2022.2 on the NVIDIA Tesla V100. The measured memory throughput values reflect that we efficiently saturate the memory bandwidth. In contrast, the lower effective memory throughput values indicate that some of the memory accesses are redundant and could be further optimized.</p>
      <p id="d1e3647">Minimizing the memory footprint is critical when assessing the performance of memory-bounded algorithms, further speedups, and increased ability to solve large-scale problems. Due to insufficient memory at <inline-formula><mml:math id="M121" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mn mathvariant="normal">8</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> DoFs for the Pine Island Glacier model configuration, we could neither execute the standard CPU solver on four 18-core Intel Xeon Gold 6140 processors and 3 TB of RAM nor the PT GPU implementation on the Tesla V100 GPU architecture. However, we could implement PT GPU on a single A100 SXM4 featuring 80 GB of device RAM. Thus, we could further confirm the necessity to keep the memory footprint minimal for models targeting high spatial resolution.</p>
      <p id="d1e3665">In this study, we tested up to an estimated <inline-formula><mml:math id="M122" display="inline"><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mn mathvariant="normal">7</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> DoFs needed to maintain a spatial resolution of <inline-formula><mml:math id="M123" display="inline"><mml:mo>∼</mml:mo></mml:math></inline-formula> 1 km or better in grounding-line regions for Antarctic and Greenland-wide ice flow models. Future studies may involve extending the PT CUDA C implementation from (i) the regional scale to the ice-sheet scale and from (ii) a 2-D SSA to a 3-D Blatter–Pattyn higher-order (HO) approximation. Extending the PT CUDA C implementation to the ice-sheet scale will require the user to carefully choose the damping parameter <inline-formula><mml:math id="M124" display="inline"><mml:mi mathvariant="italic">γ</mml:mi></mml:math></inline-formula>, nonlinear viscosity relaxation scalar <inline-formula><mml:math id="M125" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mi mathvariant="italic">μ</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, and transient <italic>pseudo</italic>-time step <inline-formula><mml:math id="M126" display="inline"><mml:mrow><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi mathvariant="italic">τ</mml:mi></mml:mrow></mml:math></inline-formula>. The shared elliptical nature of the 2-D SSA and 3-D HO formulations and corresponding partial differential-equation-based models <xref ref-type="bibr" rid="bib1.bibx8 bib1.bibx31" id="paren.27"/> suggests the PT method's ability to solve the 3-D HO momentum balance applied to unstructured meshes. The overarching goal is to diminish spatial resolution constraints at higher computing performance to improve predictions of ice-sheet evolution.</p>
</sec>
<sec id="Ch1.S5" sec-type="conclusions">
  <label>5</label><title>Conclusions</title>
      <p id="d1e3734">Recent studies have implemented techniques that keep computational resources manageable at the ice-sheet scale while increasing the spatial resolution dynamically in areas where the grounding lines migrate during prognostic simulations <xref ref-type="bibr" rid="bib1.bibx5 bib1.bibx10" id="paren.28"/>. In terms of computer memory footprint and execution time, the computational cost associated with solving the momentum balance equations to predict the ice velocity and pressure represents one of the primary bottlenecks <xref ref-type="bibr" rid="bib1.bibx15" id="paren.29"/>. This preliminary study introduces a PT solver, applied to unstructured meshes, that leverages the GPU computing power to alleviate this bottleneck. Coupling the GPU-based ice velocity and pressure simulations with CPU-based ice thickness<?pagebreak page906?> and temperature simulations can provide an enhanced balance between speed and predictive performance.</p>
      <p id="d1e3743">This study aimed to investigate the PT CUDA C implementation for unstructured meshes and its application to the 2-D SSA model formulation. For both the Jakobshavn and Pine Island glacier models, the number of nonlinear PT iterations required to converge for a given number of vertices (<inline-formula><mml:math id="M127" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula>) scales in the order of <inline-formula><mml:math id="M128" display="inline"><mml:mrow><mml:mo>≈</mml:mo><mml:mi mathvariant="script">O</mml:mi><mml:mo>(</mml:mo><mml:msup><mml:mi>N</mml:mi><mml:mn mathvariant="normal">1.2</mml:mn></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> or better. We observed an exception at <inline-formula><mml:math id="M129" display="inline"><mml:mrow><mml:mn mathvariant="normal">3</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mn mathvariant="normal">7</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> degrees of freedom (DoFs) for the Pine Island Glacier model; optimal solver parameters are unidentifiable. We further compare the performance of the PT CUDA C implementation with a standard finite-element CPU-based implementation using the price-to-performance metric. We justify the GPU implementation in the price-to-performance metric for up to million-grid-point spatial resolutions.</p>
      <p id="d1e3787">In addition to the price-to-performance metric, we preliminary investigated the power consumption. The power consumption of the PT GPU implementation was measured using the NVIDIA System Management Interface 460.32.03. For the range of DoFs tested, the power usage for both glacier configurations to meet the stopping criterion was <inline-formula><mml:math id="M130" display="inline"><mml:mrow><mml:mn mathvariant="normal">38</mml:mn><mml:mo>±</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> W.  The power consumption measurement for the CPU implementation was taken from the hardware specification sheet: thermal design power. For a 64-bit 18-core Intel Xeon Gold 6140 processor, the thermal design power is 140 W.<fn id="Ch1.Footn2"><p id="d1e3802">Intel Xeon Gold 6140 processor specification sheet: <uri>https://ark.intel.com/content/www/us/en/ark/products/120485/intel-xeon-gold-6140-processor-24-75m-cache-2-30-ghz.html</uri> (last access: 31 December 2023).</p></fn> We executed the CPU-based multicore MPI-parallelized ice-sheet flow simulations on two CPUs, with all 36 cores enabled, and we chose the power consumption to be 280 W. This is a first-order estimate. Thus, the power consumption of the PT GPU implementation was approximately one-seventh of the traditional CPU implementation for the test cases chosen in this study.  We will investigate this further.</p>
      <p id="d1e3809">This study represents a first step toward leveraging GPU processing power, enabling more accurate polar ice discharge predictions. The insights gained will benefit efforts to diminish spatial resolution constraints at higher computing performance. The higher computing performance will allow users to run ensembles of ice-sheet flow simulations at the continental scale and higher resolution, a previously challenging task. The advances will further enable the quantification of model sensitivity to changes in upcoming climate forcings. These findings will significantly benefit process-oriented sea-level-projection studies over the coming decades.</p><?xmltex \hack{\clearpage}?>
</sec>

      
      </body>
    <back><app-group>

<?pagebreak page907?><app id="App1.Ch1.S1">
  <?xmltex \currentcnt{A}?><label>Appendix A</label><title/>

<?xmltex \floatpos{h!}?><table-wrap id="App1.Ch1.S1.T2"><?xmltex \hack{\hsize\textwidth}?><?xmltex \currentcnt{A1}?><label>Table A1</label><caption><p id="d1e3827">Optimal combination of damping parameter <inline-formula><mml:math id="M131" display="inline"><mml:mi mathvariant="italic">γ</mml:mi></mml:math></inline-formula>,  nonlinear viscosity relaxation scalar <inline-formula><mml:math id="M132" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mi mathvariant="italic">μ</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, and relaxation <inline-formula><mml:math id="M133" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mi>v</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>  to maintain the linear scaling and solution stability for the glacier model configurations and DoFs listed.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="10">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:colspec colnum="5" colname="col5" align="right"/>
     <oasis:colspec colnum="6" colname="col6" align="right"/>
     <oasis:colspec colnum="7" colname="col7" align="right"/>
     <oasis:colspec colnum="8" colname="col8" align="right"/>
     <oasis:colspec colnum="9" colname="col9" align="left"/>
     <oasis:colspec colnum="10" colname="col10" align="left"/>
     <oasis:thead>
       <oasis:row>
         <oasis:entry colname="col1">Jakobshavn Isbræ</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M134" display="inline"><mml:mi mathvariant="italic">γ</mml:mi></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M135" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mi>v</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M136" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mi mathvariant="italic">μ</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col5">Pine Island</oasis:entry>
         <oasis:entry colname="col6"><inline-formula><mml:math id="M137" display="inline"><mml:mi mathvariant="italic">γ</mml:mi></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col7"><inline-formula><mml:math id="M138" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mi>v</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col8"><inline-formula><mml:math id="M139" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mi mathvariant="italic">μ</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col9"/>
         <oasis:entry colname="col10"/>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">DoFs</oasis:entry>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3"/>
         <oasis:entry colname="col4"/>
         <oasis:entry colname="col5">Glacier DoFs</oasis:entry>
         <oasis:entry colname="col6"/>
         <oasis:entry colname="col7"/>
         <oasis:entry colname="col8"/>
         <oasis:entry colname="col9"/>
         <oasis:entry colname="col10"/>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">88 458</oasis:entry>
         <oasis:entry colname="col2">0.98</oasis:entry>
         <oasis:entry colname="col3">0.99</oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M140" display="inline"><mml:mrow><mml:mn mathvariant="normal">3</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col5">28 920</oasis:entry>
         <oasis:entry colname="col6">0.98</oasis:entry>
         <oasis:entry colname="col7">0.6</oasis:entry>
         <oasis:entry colname="col8"><inline-formula><mml:math id="M141" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col9"/>
         <oasis:entry colname="col10"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">329 362</oasis:entry>
         <oasis:entry colname="col2">0.987</oasis:entry>
         <oasis:entry colname="col3">0.98</oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M142" display="inline"><mml:mrow><mml:mn mathvariant="normal">7</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col5">71 292</oasis:entry>
         <oasis:entry colname="col6">0.99</oasis:entry>
         <oasis:entry colname="col7">0.49</oasis:entry>
         <oasis:entry colname="col8"><inline-formula><mml:math id="M143" display="inline"><mml:mrow><mml:mn mathvariant="normal">8</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col9"/>
         <oasis:entry colname="col10"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">787 542</oasis:entry>
         <oasis:entry colname="col2">0.99</oasis:entry>
         <oasis:entry colname="col3">0.99</oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M144" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col5">139 578</oasis:entry>
         <oasis:entry colname="col6">0.991</oasis:entry>
         <oasis:entry colname="col7">0.99</oasis:entry>
         <oasis:entry colname="col8"><inline-formula><mml:math id="M145" display="inline"><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col9"/>
         <oasis:entry colname="col10"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">1 335 458</oasis:entry>
         <oasis:entry colname="col2">0.992</oasis:entry>
         <oasis:entry colname="col3">0.999</oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M146" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col5">2 221 410</oasis:entry>
         <oasis:entry colname="col6">0.998</oasis:entry>
         <oasis:entry colname="col7">0.995</oasis:entry>
         <oasis:entry colname="col8"><inline-formula><mml:math id="M147" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col9"/>
         <oasis:entry colname="col10"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">21 328 514</oasis:entry>
         <oasis:entry colname="col2">0.998</oasis:entry>
         <oasis:entry colname="col3">0.999</oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M148" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6"/>
         <oasis:entry colname="col7"/>
         <oasis:entry colname="col8"/>
         <oasis:entry colname="col9"/>
         <oasis:entry colname="col10"/>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table><?xmltex \gdef\@currentlabel{A1}?></table-wrap>

</app>
  </app-group><notes notes-type="codedataavailability"><title>Code and data availability</title>

      <p id="d1e4300">The current version of FastIceFlo is available for download from GitHub at <uri>https://github.com/AnjaliSandip/FastIceFlo</uri> (last access: 18 September 2023) under the MIT license. The exact version of the model used to produce the results used in this paper is archived on Zenodo (<ext-link xlink:href="https://doi.org/10.5281/zenodo.8356351" ext-link-type="DOI">10.5281/zenodo.8356351</ext-link>, <xref ref-type="bibr" rid="bib1.bibx30" id="altparen.30"/>) along with the input data and scripts to run the model and produce the plots for all of the simulations presented in this paper. The PT CUDA C implementation runs on a CUDA-capable GPU device. The research data are presented in the paper.</p>
  </notes><notes notes-type="authorcontribution"><title>Author contributions</title>

      <p id="d1e4315">AS developed the PT CUDA C implementation, conducted the performance assessment tests described in the paper, and was responsible for data analysis. LR provided guidance on the early stages of the mathematical reformulation of the 2-D SSA model to incorporate the PT method and supported the PT CUDA C implementation. MM reformulated the 2-D SSA model to incorporate the PT method, developed the weak formulation, and wrote the first versions of the code in MATLAB and then C. All authors participated in writing the manuscript.</p>
  </notes><notes notes-type="competinginterests"><title>Competing interests</title>

      <p id="d1e4321">At least one of the (co-)authors is a member of the editorial board of <italic>Geoscientific Model Development</italic>. The peer-review process was guided by an independent editor, and the authors also have no other competing interests to declare.</p>
  </notes><notes notes-type="disclaimer"><title>Disclaimer</title>

      <p id="d1e4330">Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. While Copernicus Publications makes every effort to include appropriate place names, the final responsibility lies with the authors.</p>
  </notes><?xmltex \hack{\newpage}?><?xmltex \hack{~\\[53mm]}?><ack><title>Acknowledgements</title><p id="d1e4338">We thank Dan Martin and the anonymous reviewer for their valuable feedback that enhanced the study. We acknowledge the University of North Dakota Computational Research Center for computing resources on the Talon cluster and are grateful to David Apostal and Aaron Bergstrom for technical support.  We thank the NVIDIA Applied Research Accelerator Program for hardware support. We are also grateful to the NVIDIA solution architects, Oded Green, Zoe Ryan, and Jonathan Dursi, for their thoughtful input. Ludovic Räss thanks Ivan Utkin for helpful discussions and acknowledges the Laboratory of Hydraulics, Hydrology, and Glaciology (VAW) at ETH Zurich for computing access to the Superzack GPU server. CPU and GPU hardware architectures were made available through the first author's access to the Talon (University of North Dakota) and Curiosity (NVIDIA) clusters.</p></ack><notes notes-type="reviewstatement"><title>Review statement</title>

      <p id="d1e4343">This paper was edited by Philippe Huybrechts and reviewed by Daniel Martin and one anonymous referee.</p>
  </notes><ref-list>
    <title>References</title>

      <ref id="bib1.bibx1"><?xmltex \def\ref@label{{Aschwanden et~al.(2021)Aschwanden, Bartholomaus, Brinkerhoff, and
Truffer}}?><label>Aschwanden et al.(2021)Aschwanden, Bartholomaus, Brinkerhoff, and Truffer</label><?label aschwanden2021brief?><mixed-citation>Aschwanden, A., Bartholomaus, T. C., Brinkerhoff, D. J., and Truffer, M.: Brief communication: A roadmap towards credible projections of ice sheet contribution to sea level, The Cryosphere, 15, 5705–5715, <ext-link xlink:href="https://doi.org/10.5194/tc-15-5705-2021" ext-link-type="DOI">10.5194/tc-15-5705-2021</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx2"><?xmltex \def\ref@label{{Br{\ae}dstrup et~al.(2014)Br{\ae}dstrup, Damsgaard, and
Egholm}}?><label>Brædstrup et al.(2014)Brædstrup, Damsgaard, and Egholm</label><?label braedstrup2014ice?><mixed-citation>Brædstrup, C. F., Damsgaard, A., and Egholm, D. L.: Ice-sheet modelling accelerated by graphics cards, Comput. Geosci., 72, 210–220, <ext-link xlink:href="https://doi.org/10.1016/j.cageo.2014.07.019" ext-link-type="DOI">10.1016/j.cageo.2014.07.019</ext-link>, 2014.</mixed-citation></ref>
      <ref id="bib1.bibx3"><?xmltex \def\ref@label{{Castleman et~al.(2022)Castleman, Schlegel, Caron, Larour, and
Khazendar}}?><label>Castleman et al.(2022)Castleman, Schlegel, Caron, Larour, and Khazendar</label><?label castleman2022derivation?><mixed-citation>Castleman, B. A., Schlegel, N.-J., Caron, L., Larour, E., and Khazendar, A.: Derivation of bedrock topography measurement requirements for the reduction of uncertainty in ice-sheet model projections of Thwaites Glacier, The Cryosphere, 16, 761–778, <ext-link xlink:href="https://doi.org/10.5194/tc-16-761-2022" ext-link-type="DOI">10.5194/tc-16-761-2022</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx4"><?xmltex \def\ref@label{{Chen et~al.(2017)Chen, Zhang, Church, Watson, King, Monselesan,
Legresy, and Harig}}?><label>Chen et al.(2017)Chen, Zhang, Church, Watson, King, Monselesan, Legresy, and Harig</label><?label chen2017increasing?><mixed-citation>Chen, X., Zhang, X., Church, J. A., Watson, C. S., King, M. A., Monselesan, D., Legresy, B., and Harig, C.: The increasing rate of global mean sea-level rise during 1993–2014, Nat. Clim. Change, 7, 492–495, <ext-link xlink:href="https://doi.org/10.1038/nclimate3325" ext-link-type="DOI">10.1038/nclimate3325</ext-link>, 2017.</mixed-citation></ref>
      <?pagebreak page908?><ref id="bib1.bibx5"><?xmltex \def\ref@label{{Cornford et~al.(2013)Cornford, Martin, Graves, Ranken, Le~Brocq,
Gladstone, Payne, Ng, and Lipscomb}}?><label>Cornford et al.(2013)Cornford, Martin, Graves, Ranken, Le Brocq, Gladstone, Payne, Ng, and Lipscomb</label><?label cornford2013adaptive?><mixed-citation>Cornford, S. L., Martin, D. F., Graves, D. T., Ranken, D. F., Le Brocq, A. M., Gladstone, R. M., Payne, A. J., Ng, E. G., and Lipscomb, W. H.: Adaptive mesh, finite volume modeling of marine ice sheets, J. Comput. Phys., 232, 529–549, <ext-link xlink:href="https://doi.org/10.1016/j.jcp.2012.08.037" ext-link-type="DOI">10.1016/j.jcp.2012.08.037</ext-link>, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx6"><?xmltex \def\ref@label{{Duretz et~al.(2019)Duretz, R{\"{a}}ss, Podladchikov, and
Schmalholz}}?><label>Duretz et al.(2019)Duretz, Räss, Podladchikov, and Schmalholz</label><?label duretz2019resolving?><mixed-citation>Duretz, T., Räss, L., Podladchikov, Y., and Schmalholz, S.: Resolving thermomechanical coupling in two and three dimensions: spontaneous strain localization owing to shear heating, Geophys. J. Int., 216, 365–379, <ext-link xlink:href="https://doi.org/10.1093/gji/ggy434" ext-link-type="DOI">10.1093/gji/ggy434</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx7"><?xmltex \def\ref@label{{Frankel(1950)}}?><label>Frankel(1950)</label><?label frankel1950convergence?><mixed-citation> Frankel, S. P.: Convergence rates of iterative treatments of partial differential equations, Math. Comput., 4, 65–75, 1950.</mixed-citation></ref>
      <ref id="bib1.bibx8"><?xmltex \def\ref@label{{Gilbarg and Trudinger(1977)}}?><label>Gilbarg and Trudinger(1977)</label><?label gilbarg1977elliptic?><mixed-citation>Gilbarg, D. and Trudinger, N. S.: Elliptic partial differential equations of second order, vol. 224, Springer, <ext-link xlink:href="https://doi.org/10.1007/978-3-642-61798-0" ext-link-type="DOI">10.1007/978-3-642-61798-0</ext-link>, 1977.</mixed-citation></ref>
      <ref id="bib1.bibx9"><?xmltex \def\ref@label{{Glen(1955)}}?><label>Glen(1955)</label><?label glen1955creep?><mixed-citation>Glen, J. W.: The creep of polycrystalline ice, P. Roy. Soc. Lond.-Ser. A, 228, 519–538, <ext-link xlink:href="https://doi.org/10.1098/rspa.1955.0066" ext-link-type="DOI">10.1098/rspa.1955.0066</ext-link>, 1955.</mixed-citation></ref>
      <ref id="bib1.bibx10"><?xmltex \def\ref@label{{Goelzer et~al.(2017)Goelzer, Robinson, Seroussi, and Van
De~Wal}}?><label>Goelzer et al.(2017)Goelzer, Robinson, Seroussi, and Van De Wal</label><?label goelzer2017recent?><mixed-citation>Goelzer, H., Robinson, A., Seroussi, H., and Van De Wal, R. S.: Recent progress in Greenland ice sheet modelling, Curr. Clim. Change Rep., 3, 291–302, <ext-link xlink:href="https://doi.org/10.1007/s40641-017-0073-y" ext-link-type="DOI">10.1007/s40641-017-0073-y</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx11"><?xmltex \def\ref@label{{Habbal et~al.(2017)Habbal, Larour, Morlighem, Seroussi, Borstad, and
Rignot}}?><label>Habbal et al.(2017)Habbal, Larour, Morlighem, Seroussi, Borstad, and Rignot</label><?label habbal2017optimal?><mixed-citation>Habbal, F., Larour, E., Morlighem, M., Seroussi, H., Borstad, C. P., and Rignot, E.: Optimal numerical solvers for transient simulations of ice flow using the Ice Sheet System Model (ISSM versions 4.2.5 and 4.11), Geosci. Model Dev., 10, 155–168, <ext-link xlink:href="https://doi.org/10.5194/gmd-10-155-2017" ext-link-type="DOI">10.5194/gmd-10-155-2017</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx12"><?xmltex \def\ref@label{{H{\"{a}}fner et~al.(2021)H{\"{a}}fner, Nuterman, and
Jochum}}?><label>Häfner et al.(2021)Häfner, Nuterman, and Jochum</label><?label hafner2021fast?><mixed-citation>Häfner, D., Nuterman, R., and Jochum, M.: Fast, cheap, and turbulent – Global ocean modeling with GPU acceleration in python, J. Adv. Model. Earth Sy., 13, e2021MS002717, <ext-link xlink:href="https://doi.org/10.1029/2021MS002717" ext-link-type="DOI">10.1029/2021MS002717</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx13"><?xmltex \def\ref@label{{Hinkel et~al.(2014)Hinkel, Lincke, Vafeidis, Perrette, Nicholls, Tol,
Marzeion, Fettweis, Ionescu, and Levermann}}?><label>Hinkel et al.(2014)Hinkel, Lincke, Vafeidis, Perrette, Nicholls, Tol, Marzeion, Fettweis, Ionescu, and Levermann</label><?label hinkel2014coastal?><mixed-citation>Hinkel, J., Lincke, D., Vafeidis, A. T., Perrette, M., Nicholls, R. J., Tol, R. S., Marzeion, B., Fettweis, X., Ionescu, C., and Levermann, A.: Coastal flood damage and adaptation costs under 21st century sea-level rise, P. Natl. Acad.  Sci. USA, 111, 3292–3297, <ext-link xlink:href="https://doi.org/10.1073/pnas.1222469111" ext-link-type="DOI">10.1073/pnas.1222469111</ext-link>, 2014.</mixed-citation></ref>
      <ref id="bib1.bibx14"><?xmltex \def\ref@label{{IPCC(2021)}}?><label>IPCC(2021)</label><?label ipcc2021climate?><mixed-citation> IPCC: Climate Change 2021 The Physical Science Basis, Working Group 1 (WG1) Contribution to the Sixth Assessment Report of the Intergovernmental Panel on Climate Change, Tech. rep., Intergovernmental Panel on Climate Change, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx15"><?xmltex \def\ref@label{{Jouvet et~al.(2022)Jouvet, Cordonnier, Kim, L{\"{u}}thi, Vieli, and
Aschwanden}}?><label>Jouvet et al.(2022)Jouvet, Cordonnier, Kim, Lüthi, Vieli, and Aschwanden</label><?label jouvet2022deep?><mixed-citation>Jouvet, G., Cordonnier, G., Kim, B., Lüthi, M., Vieli, A., and Aschwanden, A.: Deep learning speeds up ice flow modelling by several orders of magnitude, J. Glaciol., 68, 651–664, <ext-link xlink:href="https://doi.org/10.1017/jog.2021.120" ext-link-type="DOI">10.1017/jog.2021.120</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx16"><?xmltex \def\ref@label{{Kelley and Liao(2013)}}?><label>Kelley and Liao(2013)</label><?label kelley2013explicit?><mixed-citation> Kelley, C. and Liao, L.-Z.: Explicit pseudo-transient continuation, Computing, 15, 18, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx17"><?xmltex \def\ref@label{{Kopp et~al.(2016)Kopp, Kemp, Bittermann, Horton, Donnelly, Gehrels,
Hay, Mitrovica, Morrow, and Rahmstorf}}?><label>Kopp et al.(2016)Kopp, Kemp, Bittermann, Horton, Donnelly, Gehrels, Hay, Mitrovica, Morrow, and Rahmstorf</label><?label kopp2016temperature?><mixed-citation>Kopp, R. E., Kemp, A. C., Bittermann, K., Horton, B. P., Donnelly, J. P., Gehrels, W. R., Hay, C. C., Mitrovica, J. X., Morrow, E. D., and Rahmstorf, S.: Temperature-driven global sea-level variability in the Common Era, P. Natl. Acad. Sci. USA, 113, E1434–E1441, <ext-link xlink:href="https://doi.org/10.1073/pnas.1517056113" ext-link-type="DOI">10.1073/pnas.1517056113</ext-link>, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx18"><?xmltex \def\ref@label{{Larour et~al.(2012)Larour, Seroussi, Morlighem, and
Rignot}}?><label>Larour et al.(2012)Larour, Seroussi, Morlighem, and Rignot</label><?label larour2012continental?><mixed-citation>Larour, E., Seroussi, H., Morlighem, M., and Rignot, E.: Continental scale, high order, high spatial resolution, ice sheet modeling using the Ice Sheet System Model (ISSM), J. Geophys. Res.-Earth, 117, 22, <ext-link xlink:href="https://doi.org/10.1029/2011JF002140" ext-link-type="DOI">10.1029/2011JF002140</ext-link>, 2012.</mixed-citation></ref>
      <ref id="bib1.bibx19"><?xmltex \def\ref@label{{MacAyeal(1989)}}?><label>MacAyeal(1989)</label><?label macayeal1989large?><mixed-citation> MacAyeal, D. R.: Large-scale ice flow over a viscous basal sediment: Theory and application to ice stream B, Antarctica, J. Geophys. Res.-Sol. Ea., 94, 4071–4087, 1989.</mixed-citation></ref>
      <ref id="bib1.bibx20"><?xmltex \def\ref@label{{Morlighem et~al.(2017)Morlighem, Williams, Rignot, An, Arndt, Bamber,
Catania, Chauch{\'{e}}, Dowdeswell, Dorschel, Fenty, Hogan, Howat, Hubbard,
Jakobsson, Jordan, Kjeldsen, Millan, Mayer, Mouginot, Noel, O’Cofaigh,
Palmer, Rysgaard, Seroussi, Siegert, Slabon, Straneo, Van Den~Broeke,
Weinrebe, Wood, and Zinglersen}}?><label>Morlighem et al.(2017)Morlighem, Williams, Rignot, An, Arndt, Bamber, Catania, Chauché, Dowdeswell, Dorschel, Fenty, Hogan, Howat, Hubbard, Jakobsson, Jordan, Kjeldsen, Millan, Mayer, Mouginot, Noel, O’Cofaigh, Palmer, Rysgaard, Seroussi, Siegert, Slabon, Straneo, Van Den Broeke, Weinrebe, Wood, and Zinglersen</label><?label morlighem2017bedmachine?><mixed-citation>Morlighem, M., Williams, C. N., Rignot, E., An, L., Arndt, J. E., Bamber, J. L., Catania, G., Chauché, N., Dowdeswell, J. A., Dorschel, B., Fenty, I., Hogan, K., Howat, I., Hubbard, A., Jakobsson, M., Jordan, T. M., Kjeldsen, K. K., Millan, R., Mayer, L., Mouginot, J., Noel, B. P. Y., O’Cofaigh, C., Palmer, S., Rysgaard, S., Seroussi, H., Siegert, M. J., Slabon, P., Straneo, F., Van Den Broeke, M. R., Weinrebe, W., Wood, M., and Zinglersen, K. B.: BedMachine v3: Complete bed topography and ocean bathymetry mapping of Greenland from multibeam echo sounding combined with mass conservation, Geophys. Res. Lett., 44, 11–51, <ext-link xlink:href="https://doi.org/10.1002/2017GL074954" ext-link-type="DOI">10.1002/2017GL074954</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx21"><?xmltex \def\ref@label{{Morlighem et~al.(2020)Morlighem, Rignot, Binder, Blankenship, Drews,
Eagles, Eisen, Ferraccioli, Forsberg, Fretwell, Goel, Greenbaum, Gudmundsson,
Guo, Helm, Hofstede, Howat, Humbert, Jokat, Karlsson, Lee, Matsuoka, Millan,
Mouginot, Paden, Pattyn, Roberts, Rosier, Ruppel, Seroussi, Smith, Steinhage,
Sun, Van Den~Broeke, Van~Ommen, Wessem, and Young}}?><label>Morlighem et al.(2020)Morlighem, Rignot, Binder, Blankenship, Drews, Eagles, Eisen, Ferraccioli, Forsberg, Fretwell, Goel, Greenbaum, Gudmundsson, Guo, Helm, Hofstede, Howat, Humbert, Jokat, Karlsson, Lee, Matsuoka, Millan, Mouginot, Paden, Pattyn, Roberts, Rosier, Ruppel, Seroussi, Smith, Steinhage, Sun, Van Den Broeke, Van Ommen, Wessem, and Young</label><?label morlighem2020deep?><mixed-citation>Morlighem, M., Rignot, E., Binder, T., Blankenship, D., Drews, R., Eagles, G., Eisen, O., Ferraccioli, F., Forsberg, R., Fretwell, P., Goel, V., Greenbaum, J. S., Gudmundsson, H., Guo, J., Helm, V., Hofstede, C., Howat, I., Humbert, A., Jokat, W., Karlsson, N., Lee, W. S., Matsuoka, K., Millan, R., Mouginot, J., Paden, J., Pattyn, F., Roberts, J., Rosier, S., Ruppel, A., Seroussi, H., Smith, E. C., Steinhage, D., Sun, B., Van Den Broeke, M. R., Van Ommen, T. D., Wessem, M. V., and Young, D. A.: Deep glacial troughs and stabilizing ridges unveiled beneath the margins of the Antarctic ice sheet, Nat. Geosci., 13, 132–137, <ext-link xlink:href="https://doi.org/10.1038/s41561-019-0510-8" ext-link-type="DOI">10.1038/s41561-019-0510-8</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx22"><?xmltex \def\ref@label{{Omlin et~al.(2018)Omlin, R{\"{a}}ss, and
Podladchikov}}?><label>Omlin et al.(2018)Omlin, Räss, and Podladchikov</label><?label omlin2018simulation?><mixed-citation>Omlin, S., Räss, L., and Podladchikov, Y. Y.: Simulation of three-dimensional viscoelastic deformation coupled to porous fluid flow, Tectonophysics, 746, 695–701, <ext-link xlink:href="https://doi.org/10.1016/j.tecto.2017.08.012" ext-link-type="DOI">10.1016/j.tecto.2017.08.012</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx23"><?xmltex \def\ref@label{{Pattyn and Morlighem(2020)}}?><label>Pattyn and Morlighem(2020)</label><?label pattyn2020uncertain?><mixed-citation>Pattyn, F. and Morlighem, M.: The uncertain future of the Antarctic Ice Sheet, Science, 367, 1331–1335, <ext-link xlink:href="https://doi.org/10.1126/science.aaz5487" ext-link-type="DOI">10.1126/science.aaz5487</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx24"><?xmltex \def\ref@label{{Poliakov et~al.(1993)Poliakov, Cundall, Podladchikov, and
Lyakhovsky}}?><label>Poliakov et al.(1993)Poliakov, Cundall, Podladchikov, and Lyakhovsky</label><?label Poliakov1993?><mixed-citation>Poliakov, A. N., Cundall, P. A., Podladchikov, Y. Y., and Lyakhovsky, V. A.: An explicit inertial method for the simulation of viscoelastic flow: an evaluation of elastic effects on diapiric flow in two- and three- layers models, Flow and creep in the solar system: observations, modeling and theory,   in:  Flow and Creep in the Solar System: Observations, Modeling and Theory, edited by: Stone, D. B. and Runcorn, S. K., NATO ASI Series, vol. 391, Springer, Dordrecht, 175–195, <ext-link xlink:href="https://doi.org/10.1007/978-94-015-8206-3_12" ext-link-type="DOI">10.1007/978-94-015-8206-3_12</ext-link>, 1993.</mixed-citation></ref>
      <ref id="bib1.bibx25"><?xmltex \def\ref@label{{R{\"{a}}ss et~al.(2019)R{\"{a}}ss, Duretz, and
Podladchikov}}?><label>Räss et al.(2019)Räss, Duretz, and Podladchikov</label><?label rass2019resolving?><mixed-citation>Räss, L., Duretz, T., and Podladchikov, Y.: Resolving hydromechanical coupling in two and three dimensions: spontaneous channelling of porous fluids owing to decompaction weakening, Geophys. J. Int., 218, 1591–1616, <ext-link xlink:href="https://doi.org/10.1093/gji/ggz239" ext-link-type="DOI">10.1093/gji/ggz239</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx26"><?xmltex \def\ref@label{{R{\"{a}}ss et~al.(2020)R{\"{a}}ss, Licul, Herman, Podladchikov, and
Suckale}}?><label>Räss et al.(2020)Räss, Licul, Herman, Podladchikov, and Suckale</label><?label rass2020modelling?><mixed-citation>Räss, L., Licul, A., Herman, F., Podladchikov, Y. Y., and Suckale, J.: Modelling thermomechanical ice deformation using an implicit pseudo-transient method (FastICE v1.0) based on graphical processing units (GPUs), Geosci. Model Dev., 13, 955–976, <ext-link xlink:href="https://doi.org/10.5194/gmd-13-955-2020" ext-link-type="DOI">10.5194/gmd-13-955-2020</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx27"><?xmltex \def\ref@label{{R{\"{a}}ss et~al.(2022)R{\"{a}}ss, Utkin, Duretz, Omlin, and
Podladchikov}}?><label>Räss et al.(2022)Räss, Utkin, Duretz, Omlin, and Podladchikov</label><?label rass2022assessing?><mixed-citation>Räss, L., Utkin, I., Duretz, T., Omlin, S., and Podladchikov, Y. Y.: Assessing the robustness and scalability of the accelerated pseudo-transient method, Geosci. Model Dev., 15, 5757–5786, <ext-link xlink:href="https://doi.org/10.5194/gmd-15-5757-2022" ext-link-type="DOI">10.5194/gmd-15-5757-2022</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx28"><?xmltex \def\ref@label{{Rietbroek et~al.(2016)Rietbroek, Brunnabend, Kusche, Schr{\"{o}}ter,
and Dahle}}?><label>Rietbroek et al.(2016)Rietbroek, Brunnabend, Kusche, Schröter, and Dahle</label><?label rietbroek2016revisiting?><mixed-citation>Rietbroek, R., Brunnabend, S.-E., Kusche, J., Schröter, J., and Dahle, C.: Revisiting the contemporary sea-level budget on global and regional scales, P. Natl. Acad. Sci. USA, 113, 1504–1509, <ext-link xlink:href="https://doi.org/10.1073/pnas.1519132113" ext-link-type="DOI">10.1073/pnas.1519132113</ext-link>, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx29"><?xmltex \def\ref@label{{Rignot et~al.(2011)Rignot, Mouginot, and
Scheuchl}}?><label>Rignot et al.(2011)Rignot, Mouginot, and Scheuchl</label><?label rignot2011antarctic?><mixed-citation>Rignot, E., Mouginot, J., and Scheuchl, B.: Antarctic grounding line mapping from differential satellite radar interferometry, Geophys. Res. Lett., 38, 49, <ext-link xlink:href="https://doi.org/10.1029/2011GL047109" ext-link-type="DOI">10.1029/2011GL047109</ext-link>, 2011.</mixed-citation></ref>
      <ref id="bib1.bibx30"><?xmltex \def\ref@label{Sandip et al.(2023)}?><label>Sandip et al.(2023)</label><?label anjali_sandip_2023_8356351?><mixed-citation>Sandip, A., Morlighem, M., and  Räss, L.:  AnjaliSandip/FastIceFlo: FastIceFlo v1.0.1 (v1.0.1), Zenodo [code], <ext-link xlink:href="https://doi.org/10.5281/zenodo.8356351" ext-link-type="DOI">10.5281/zenodo.8356351</ext-link>, 2023.</mixed-citation></ref>
      <?pagebreak page909?><ref id="bib1.bibx31"><?xmltex \def\ref@label{{Tezaur et~al.(2015)Tezaur, Perego, Salinger, Tuminaro, and
Price}}?><label>Tezaur et al.(2015)Tezaur, Perego, Salinger, Tuminaro, and Price</label><?label tezaur2015albany?><mixed-citation>Tezaur, I. K., Perego, M., Salinger, A. G., Tuminaro, R. S., and Price, S. F.: Albany/FELIX: a parallel, scalable and robust, finite element, first-order Stokes approximation ice sheet solver built for advanced analysis, Geosci. Model Dev., 8, 1197–1220, <ext-link xlink:href="https://doi.org/10.5194/gmd-8-1197-2015" ext-link-type="DOI">10.5194/gmd-8-1197-2015</ext-link>, 2015.</mixed-citation></ref>

  </ref-list></back>
    <!--<article-title-html>Graphics-processing-unit-accelerated ice flow solver for unstructured meshes using the Shallow-Shelf Approximation (FastIceFlo v1.0.1)</article-title-html>
<abstract-html/>
<ref-html id="bib1.bib1"><label>Aschwanden et al.(2021)Aschwanden, Bartholomaus, Brinkerhoff, and
Truffer</label><mixed-citation>
      
Aschwanden, A., Bartholomaus, T. C., Brinkerhoff, D. J., and Truffer, M.: Brief communication: A roadmap towards credible projections of ice sheet contribution to sea level, The Cryosphere, 15, 5705–5715, <a href="https://doi.org/10.5194/tc-15-5705-2021" target="_blank">https://doi.org/10.5194/tc-15-5705-2021</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib2"><label>Brædstrup et al.(2014)Brædstrup, Damsgaard, and
Egholm</label><mixed-citation>
      
Brædstrup, C. F., Damsgaard, A., and Egholm, D. L.: Ice-sheet modelling
accelerated by graphics cards, Comput. Geosci., 72, 210–220,
<a href="https://doi.org/10.1016/j.cageo.2014.07.019" target="_blank">https://doi.org/10.1016/j.cageo.2014.07.019</a>, 2014.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib3"><label>Castleman et al.(2022)Castleman, Schlegel, Caron, Larour, and
Khazendar</label><mixed-citation>
      
Castleman, B. A., Schlegel, N.-J., Caron, L., Larour, E., and Khazendar, A.: Derivation of bedrock topography measurement requirements for the reduction of uncertainty in ice-sheet model projections of Thwaites Glacier, The Cryosphere, 16, 761–778, <a href="https://doi.org/10.5194/tc-16-761-2022" target="_blank">https://doi.org/10.5194/tc-16-761-2022</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib4"><label>Chen et al.(2017)Chen, Zhang, Church, Watson, King, Monselesan,
Legresy, and Harig</label><mixed-citation>
      
Chen, X., Zhang, X., Church, J. A., Watson, C. S., King, M. A., Monselesan, D.,
Legresy, B., and Harig, C.: The increasing rate of global mean sea-level rise
during 1993–2014, Nat. Clim. Change, 7, 492–495,
<a href="https://doi.org/10.1038/nclimate3325" target="_blank">https://doi.org/10.1038/nclimate3325</a>, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib5"><label>Cornford et al.(2013)Cornford, Martin, Graves, Ranken, Le Brocq,
Gladstone, Payne, Ng, and Lipscomb</label><mixed-citation>
      
Cornford, S. L., Martin, D. F., Graves, D. T., Ranken, D. F., Le Brocq, A. M.,
Gladstone, R. M., Payne, A. J., Ng, E. G., and Lipscomb, W. H.: Adaptive
mesh, finite volume modeling of marine ice sheets, J. Comput.
Phys., 232, 529–549, <a href="https://doi.org/10.1016/j.jcp.2012.08.037" target="_blank">https://doi.org/10.1016/j.jcp.2012.08.037</a>, 2013.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib6"><label>Duretz et al.(2019)Duretz, Räss, Podladchikov, and
Schmalholz</label><mixed-citation>
      
Duretz, T., Räss, L., Podladchikov, Y., and Schmalholz, S.: Resolving
thermomechanical coupling in two and three dimensions: spontaneous strain
localization owing to shear heating, Geophys. J. Int., 216,
365–379, <a href="https://doi.org/10.1093/gji/ggy434" target="_blank">https://doi.org/10.1093/gji/ggy434</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib7"><label>Frankel(1950)</label><mixed-citation>
      
Frankel, S. P.: Convergence rates of iterative treatments of partial
differential equations, Math. Comput., 4, 65–75, 1950.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib8"><label>Gilbarg and Trudinger(1977)</label><mixed-citation>
      
Gilbarg, D. and Trudinger, N. S.: Elliptic partial differential equations of
second order, vol. 224, Springer, <a href="https://doi.org/10.1007/978-3-642-61798-0" target="_blank">https://doi.org/10.1007/978-3-642-61798-0</a>, 1977.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib9"><label>Glen(1955)</label><mixed-citation>
      
Glen, J. W.: The creep of polycrystalline ice, P. Roy. Soc.
Lond.-Ser. A, 228, 519–538,
<a href="https://doi.org/10.1098/rspa.1955.0066" target="_blank">https://doi.org/10.1098/rspa.1955.0066</a>, 1955.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib10"><label>Goelzer et al.(2017)Goelzer, Robinson, Seroussi, and Van
De Wal</label><mixed-citation>
      
Goelzer, H., Robinson, A., Seroussi, H., and Van De Wal, R. S.: Recent progress
in Greenland ice sheet modelling, Curr. Clim. Change Rep., 3,
291–302, <a href="https://doi.org/10.1007/s40641-017-0073-y" target="_blank">https://doi.org/10.1007/s40641-017-0073-y</a>, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib11"><label>Habbal et al.(2017)Habbal, Larour, Morlighem, Seroussi, Borstad, and
Rignot</label><mixed-citation>
      
Habbal, F., Larour, E., Morlighem, M., Seroussi, H., Borstad, C. P., and Rignot, E.: Optimal numerical solvers for transient simulations of ice flow using the Ice Sheet System Model (ISSM versions 4.2.5 and 4.11), Geosci. Model Dev., 10, 155–168, <a href="https://doi.org/10.5194/gmd-10-155-2017" target="_blank">https://doi.org/10.5194/gmd-10-155-2017</a>, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib12"><label>Häfner et al.(2021)Häfner, Nuterman, and
Jochum</label><mixed-citation>
      
Häfner, D., Nuterman, R., and Jochum, M.: Fast, cheap, and
turbulent – Global ocean modeling with GPU acceleration in python, J.
Adv. Model. Earth Sy., 13, e2021MS002717,
<a href="https://doi.org/10.1029/2021MS002717" target="_blank">https://doi.org/10.1029/2021MS002717</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib13"><label>Hinkel et al.(2014)Hinkel, Lincke, Vafeidis, Perrette, Nicholls, Tol,
Marzeion, Fettweis, Ionescu, and Levermann</label><mixed-citation>
      
Hinkel, J., Lincke, D., Vafeidis, A. T., Perrette, M., Nicholls, R. J., Tol,
R. S., Marzeion, B., Fettweis, X., Ionescu, C., and Levermann, A.: Coastal
flood damage and adaptation costs under 21st century sea-level rise,
P. Natl. Acad.  Sci. USA, 111, 3292–3297,
<a href="https://doi.org/10.1073/pnas.1222469111" target="_blank">https://doi.org/10.1073/pnas.1222469111</a>, 2014.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib14"><label>IPCC(2021)</label><mixed-citation>
      
IPCC: Climate Change 2021 The Physical Science Basis, Working Group 1 (WG1)
Contribution to the Sixth Assessment Report of the Intergovernmental Panel on
Climate Change, Tech. rep., Intergovernmental Panel on Climate Change, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib15"><label>Jouvet et al.(2022)Jouvet, Cordonnier, Kim, Lüthi, Vieli, and
Aschwanden</label><mixed-citation>
      
Jouvet, G., Cordonnier, G., Kim, B., Lüthi, M., Vieli, A., and Aschwanden,
A.: Deep learning speeds up ice flow modelling by several orders of
magnitude, J. Glaciol., 68, 651–664, <a href="https://doi.org/10.1017/jog.2021.120" target="_blank">https://doi.org/10.1017/jog.2021.120</a>,
2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib16"><label>Kelley and Liao(2013)</label><mixed-citation>
      
Kelley, C. and Liao, L.-Z.: Explicit pseudo-transient continuation, Computing,
15, 18, 2013.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib17"><label>Kopp et al.(2016)Kopp, Kemp, Bittermann, Horton, Donnelly, Gehrels,
Hay, Mitrovica, Morrow, and Rahmstorf</label><mixed-citation>
      
Kopp, R. E., Kemp, A. C., Bittermann, K., Horton, B. P., Donnelly, J. P.,
Gehrels, W. R., Hay, C. C., Mitrovica, J. X., Morrow, E. D., and Rahmstorf,
S.: Temperature-driven global sea-level variability in the Common Era,
P. Natl. Acad. Sci. USA, 113, E1434–E1441,
<a href="https://doi.org/10.1073/pnas.1517056113" target="_blank">https://doi.org/10.1073/pnas.1517056113</a>, 2016.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib18"><label>Larour et al.(2012)Larour, Seroussi, Morlighem, and
Rignot</label><mixed-citation>
      
Larour, E., Seroussi, H., Morlighem, M., and Rignot, E.: Continental scale,
high order, high spatial resolution, ice sheet modeling using the Ice Sheet
System Model (ISSM), J. Geophys. Res.-Earth, 117, 22,
<a href="https://doi.org/10.1029/2011JF002140" target="_blank">https://doi.org/10.1029/2011JF002140</a>, 2012.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib19"><label>MacAyeal(1989)</label><mixed-citation>
      
MacAyeal, D. R.: Large-scale ice flow over a viscous basal sediment: Theory and
application to ice stream B, Antarctica, J. Geophys. Res.-Sol. Ea., 94, 4071–4087, 1989.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib20"><label>Morlighem et al.(2017)Morlighem, Williams, Rignot, An, Arndt, Bamber,
Catania, Chauché, Dowdeswell, Dorschel, Fenty, Hogan, Howat, Hubbard,
Jakobsson, Jordan, Kjeldsen, Millan, Mayer, Mouginot, Noel, O’Cofaigh,
Palmer, Rysgaard, Seroussi, Siegert, Slabon, Straneo, Van Den Broeke,
Weinrebe, Wood, and Zinglersen</label><mixed-citation>
      
Morlighem, M., Williams, C. N., Rignot, E., An, L., Arndt, J. E., Bamber,
J. L., Catania, G., Chauché, N., Dowdeswell, J. A., Dorschel, B., Fenty,
I., Hogan, K., Howat, I., Hubbard, A., Jakobsson, M., Jordan, T. M.,
Kjeldsen, K. K., Millan, R., Mayer, L., Mouginot, J., Noel, B. P. Y.,
O’Cofaigh, C., Palmer, S., Rysgaard, S., Seroussi, H., Siegert, M. J.,
Slabon, P., Straneo, F., Van Den Broeke, M. R., Weinrebe, W., Wood, M., and
Zinglersen, K. B.: BedMachine v3: Complete bed topography and ocean
bathymetry mapping of Greenland from multibeam echo sounding combined with
mass conservation, Geophys. Res. Lett., 44, 11–51,
<a href="https://doi.org/10.1002/2017GL074954" target="_blank">https://doi.org/10.1002/2017GL074954</a>, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib21"><label>Morlighem et al.(2020)Morlighem, Rignot, Binder, Blankenship, Drews,
Eagles, Eisen, Ferraccioli, Forsberg, Fretwell, Goel, Greenbaum, Gudmundsson,
Guo, Helm, Hofstede, Howat, Humbert, Jokat, Karlsson, Lee, Matsuoka, Millan,
Mouginot, Paden, Pattyn, Roberts, Rosier, Ruppel, Seroussi, Smith, Steinhage,
Sun, Van Den Broeke, Van Ommen, Wessem, and Young</label><mixed-citation>
      
Morlighem, M., Rignot, E., Binder, T., Blankenship, D., Drews, R., Eagles, G.,
Eisen, O., Ferraccioli, F., Forsberg, R., Fretwell, P., Goel, V., Greenbaum,
J. S., Gudmundsson, H., Guo, J., Helm, V., Hofstede, C., Howat, I., Humbert,
A., Jokat, W., Karlsson, N., Lee, W. S., Matsuoka, K., Millan, R., Mouginot,
J., Paden, J., Pattyn, F., Roberts, J., Rosier, S., Ruppel, A., Seroussi, H.,
Smith, E. C., Steinhage, D., Sun, B., Van Den Broeke, M. R., Van Ommen,
T. D., Wessem, M. V., and Young, D. A.: Deep glacial troughs and stabilizing
ridges unveiled beneath the margins of the Antarctic ice sheet, Nat.
Geosci., 13, 132–137, <a href="https://doi.org/10.1038/s41561-019-0510-8" target="_blank">https://doi.org/10.1038/s41561-019-0510-8</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib22"><label>Omlin et al.(2018)Omlin, Räss, and
Podladchikov</label><mixed-citation>
      
Omlin, S., Räss, L., and Podladchikov, Y. Y.: Simulation of
three-dimensional viscoelastic deformation coupled to porous fluid flow,
Tectonophysics, 746, 695–701, <a href="https://doi.org/10.1016/j.tecto.2017.08.012" target="_blank">https://doi.org/10.1016/j.tecto.2017.08.012</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib23"><label>Pattyn and Morlighem(2020)</label><mixed-citation>
      
Pattyn, F. and Morlighem, M.: The uncertain future of the Antarctic Ice Sheet,
Science, 367, 1331–1335, <a href="https://doi.org/10.1126/science.aaz5487" target="_blank">https://doi.org/10.1126/science.aaz5487</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib24"><label>Poliakov et al.(1993)Poliakov, Cundall, Podladchikov, and
Lyakhovsky</label><mixed-citation>
      
Poliakov, A. N., Cundall, P. A., Podladchikov, Y. Y., and Lyakhovsky, V. A.:
An explicit inertial method for the simulation of viscoelastic flow: an
evaluation of elastic effects on diapiric flow in two- and three- layers
models, Flow and creep in the solar system: observations, modeling and
theory,   in:  Flow and Creep in the Solar System: Observations, Modeling and Theory, edited by: Stone, D. B. and Runcorn, S. K., NATO ASI Series, vol. 391, Springer, Dordrecht, 175–195,
<a href="https://doi.org/10.1007/978-94-015-8206-3_12" target="_blank">https://doi.org/10.1007/978-94-015-8206-3_12</a>,
1993.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib25"><label>Räss et al.(2019)Räss, Duretz, and
Podladchikov</label><mixed-citation>
      
Räss, L., Duretz, T., and Podladchikov, Y.: Resolving hydromechanical
coupling in two and three dimensions: spontaneous channelling of porous
fluids owing to decompaction weakening, Geophys. J. Int.,
218, 1591–1616, <a href="https://doi.org/10.1093/gji/ggz239" target="_blank">https://doi.org/10.1093/gji/ggz239</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib26"><label>Räss et al.(2020)Räss, Licul, Herman, Podladchikov, and
Suckale</label><mixed-citation>
      
Räss, L., Licul, A., Herman, F., Podladchikov, Y. Y., and Suckale, J.: Modelling thermomechanical ice deformation using an implicit pseudo-transient method (FastICE v1.0) based on graphical processing units (GPUs), Geosci. Model Dev., 13, 955–976, <a href="https://doi.org/10.5194/gmd-13-955-2020" target="_blank">https://doi.org/10.5194/gmd-13-955-2020</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib27"><label>Räss et al.(2022)Räss, Utkin, Duretz, Omlin, and
Podladchikov</label><mixed-citation>
      
Räss, L., Utkin, I., Duretz, T., Omlin, S., and Podladchikov, Y. Y.: Assessing the robustness and scalability of the accelerated pseudo-transient method, Geosci. Model Dev., 15, 5757–5786, <a href="https://doi.org/10.5194/gmd-15-5757-2022" target="_blank">https://doi.org/10.5194/gmd-15-5757-2022</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib28"><label>Rietbroek et al.(2016)Rietbroek, Brunnabend, Kusche, Schröter,
and Dahle</label><mixed-citation>
      
Rietbroek, R., Brunnabend, S.-E., Kusche, J., Schröter, J., and Dahle, C.:
Revisiting the contemporary sea-level budget on global and regional scales,
P. Natl. Acad. Sci. USA, 113, 1504–1509,
<a href="https://doi.org/10.1073/pnas.1519132113" target="_blank">https://doi.org/10.1073/pnas.1519132113</a>, 2016.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib29"><label>Rignot et al.(2011)Rignot, Mouginot, and
Scheuchl</label><mixed-citation>
      
Rignot, E., Mouginot, J., and Scheuchl, B.: Antarctic grounding line mapping
from differential satellite radar interferometry, Geophys. Res.
Lett., 38, 49, <a href="https://doi.org/10.1029/2011GL047109" target="_blank">https://doi.org/10.1029/2011GL047109</a>, 2011.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib30"><label>Sandip et al.(2023)</label><mixed-citation>
      
Sandip, A., Morlighem, M., and  Räss, L.:  AnjaliSandip/FastIceFlo: FastIceFlo v1.0.1 (v1.0.1), Zenodo [code], <a href="https://doi.org/10.5281/zenodo.8356351" target="_blank">https://doi.org/10.5281/zenodo.8356351</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib31"><label>Tezaur et al.(2015)Tezaur, Perego, Salinger, Tuminaro, and
Price</label><mixed-citation>
      
Tezaur, I. K., Perego, M., Salinger, A. G., Tuminaro, R. S., and Price, S. F.: Albany/FELIX: a parallel, scalable and robust, finite element, first-order Stokes approximation ice sheet solver built for advanced analysis, Geosci. Model Dev., 8, 1197–1220, <a href="https://doi.org/10.5194/gmd-8-1197-2015" target="_blank">https://doi.org/10.5194/gmd-8-1197-2015</a>, 2015.

    </mixed-citation></ref-html>--></article>
