<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing with OASIS Tables v3.0 20080202//EN" "https://jats.nlm.nih.gov/nlm-dtd/publishing/3.0/journalpub-oasis3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:oasis="http://docs.oasis-open.org/ns/oasis-exchange/table" xml:lang="en" dtd-version="3.0" article-type="research-article">
  <front>
    <journal-meta><journal-id journal-id-type="publisher">GMD</journal-id><journal-title-group>
    <journal-title>Geoscientific Model Development</journal-title>
    <abbrev-journal-title abbrev-type="publisher">GMD</abbrev-journal-title><abbrev-journal-title abbrev-type="nlm-ta">Geosci. Model Dev.</abbrev-journal-title>
  </journal-title-group><issn pub-type="epub">1991-9603</issn><publisher>
    <publisher-name>Copernicus Publications</publisher-name>
    <publisher-loc>Göttingen, Germany</publisher-loc>
  </publisher></journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.5194/gmd-19-7525-2026</article-id><title-group><article-title>GPU-accelerated finite-element method for the three-dimensional unstructured mesh atmospheric dynamic framework</article-title><alt-title>GPU-accelerated FEM for atmospheric model</alt-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author" corresp="no" rid="aff1">
          <name><surname>Li</surname><given-names>Leisheng</given-names></name>
          
        <ext-link>https://orcid.org/0000-0002-5234-4150</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1 aff2">
          <name><surname>Fu</surname><given-names>Ximeng</given-names></name>
          
        <ext-link>https://orcid.org/0009-0000-5033-6133</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1 aff2">
          <name><surname>Zheng</surname><given-names>Xiyu</given-names></name>
          
        </contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1">
          <name><surname>Li</surname><given-names>Huiyuan</given-names></name>
          
        <ext-link>https://orcid.org/0000-0002-6326-9926</ext-link></contrib>
        <contrib contrib-type="author" corresp="yes" rid="aff3">
          <name><surname>Li</surname><given-names>Jinxi</given-names></name>
          <email>ljx2311@mail.iap.ac.cn</email>
        </contrib>
        <aff id="aff1"><label>1</label><institution>Institute of Software, Chinese Academy of Sciences, Beijing 100190, China</institution>
        </aff>
        <aff id="aff2"><label>2</label><institution>University of Chinese Academy of Sciences, Beijing 100190, China</institution>
        </aff>
        <aff id="aff3"><label>3</label><institution>State Key Laboratory of Atmospheric Environment and Extreme Meteorology, Institute of Atmospheric Physics,  Chinese Academy of Sciences, Beijing 100029, China</institution>
        </aff>
      </contrib-group>
      <author-notes><corresp id="corr1">Jinxi Li (ljx2311@mail.iap.ac.cn)</corresp></author-notes><pub-date><day>14</day><month>August</month><year>2026</year></pub-date>
      
      <volume>19</volume>
      <issue>15</issue>
      <fpage>7525</fpage><lpage>7544</lpage>
      <history>
        <date date-type="received"><day>5</day><month>February</month><year>2026</year></date>
           <date date-type="rev-request"><day>5</day><month>March</month><year>2026</year></date>
           <date date-type="rev-recd"><day>9</day><month>June</month><year>2026</year></date>
           <date date-type="accepted"><day>19</day><month>July</month><year>2026</year></date>
      </history>
      <permissions>
        <copyright-statement>Copyright: © 2026 Leisheng Li et al.</copyright-statement>
        <copyright-year>2026</copyright-year>
      <license license-type="open-access"><license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p></license></permissions><self-uri xlink:href="https://gmd.copernicus.org/articles/19/7525/2026/gmd-19-7525-2026.html">This article is available from https://gmd.copernicus.org/articles/19/7525/2026/gmd-19-7525-2026.html</self-uri><self-uri xlink:href="https://gmd.copernicus.org/articles/19/7525/2026/gmd-19-7525-2026.pdf">The full text article is available as a PDF file from https://gmd.copernicus.org/articles/19/7525/2026/gmd-19-7525-2026.pdf</self-uri>
      <abstract><title>Abstract</title>

      <p id="d2e133">The three-dimensional unstructured-mesh finite-element atmospheric dynamical framework is gaining significance owing to its flexibility in representing complex topography and capability for multi-scale simulations in high resolutions. However, this framework has substantial bottlenecks. Unlike structured-grid models, the unstructured finite element method (FEM) must frequently access irregular mesh connectivity among nodes, edges, and elements, causing indirect memory addressing, inadequate data locality, and substantial memory bandwidth bottlenecks on conventional CPU architectures. Consequently, element-wise computations and global assembly are among the primary contributors alongside the sparse linear solver to the runtime in high-resolution simulations.</p>

      <p id="d2e136">This study develops a GPU-parallel implementation of the Fluidity-Atmosphere dynamical core to address these challenges. The GPU-oriented data structures and optimized kernels are designed to efficiently leverage the computing power of GPUs. These kernels enable parallelized element integration and are efficient solvers for specific size matrices; a parallel assembly strategy enhances memory throughput during global sparse matrix construction. On the NVIDIA A100 GPU, the optimized kernels achieve speeds over 100<inline-formula><mml:math id="M1" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> compared to the single CPU core baseline for element-wise computations and up to 389.02 times for global matrix assembly, resulting in an overall acceleration of 8.57 times with four messages passing interface (MPI) processes. The proposed framework demonstrates that tailored GPU parallelization is effective in overcoming the computational bottleneck of unstructured FEM-based atmospheric models, facilitating high-resolution simulations on heterogeneous architectures.</p>
  </abstract>
    
<funding-group>
<award-group id="gs1">
<funding-source>National Key Research and Development Program of China</funding-source>
<award-id>2023YFC3705701</award-id>
</award-group>
</funding-group>
</article-meta>
  </front>
<body>
      

<sec id="Ch1.S1" sec-type="intro">
  <label>1</label><title>Introduction</title>
      <p id="d2e155">Atmospheric dynamic frameworks are the fundamental tools for weather forecasting and climate simulation. In recent years, growing demands for the accurate prediction and management of extreme pollution and weather events have driven atmospheric models toward higher spatial resolutions and more sophisticated physical parameterizations. These advances create formidable computational challenges: doubling the spatial resolution can lead to exponential growth in computational cost, substantially increasing the demand for computing power and storage resources <xref ref-type="bibr" rid="bib1.bibx26" id="paren.1"/>. For the discretization methods, the choice of scheme is primarily dependent on the topological structure of the underlying mesh. For high-resolution simulations, models are now able to resolve increasingly complex underlying surface features, such as terrains and urban structures. Owing to their geometric constraints, traditional terrain-following coordinate grids typically produce significant mesh distortions near steep terrain, introducing significant computational errors. Therefore, conventional atmospheric models based on structured “horizontal <inline-formula><mml:math id="M2" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula> vertical” grid configurations encounter inherent limitations in representing steep terrains, offering inadequate flexibility for localized adaptation.</p>
      <p id="d2e168">To address this challenge, fully three-dimensional (3D) unstructured grids have emerged as an effective solution. Replacing the conventional quasi-3D approach that separates horizontal and vertical discretization, unstructured grids provide superior geometric adaptability and more flexible mesh generation capabilities <xref ref-type="bibr" rid="bib1.bibx40 bib1.bibx43 bib1.bibx13 bib1.bibx45 bib1.bibx47" id="paren.2"/>. The finite element method (FEM) proves particularly well-suited to such grid architectures, offering inherent advantages in handling complex geometrical configurations. This synergistic combination has established FEM as an increasingly prominent methodology in atmospheric modeling <xref ref-type="bibr" rid="bib1.bibx7 bib1.bibx19 bib1.bibx25 bib1.bibx38" id="paren.3"/>. FEM not only can be formulated to satisfy discrete conservation properties and numerical stability on irregular geometries but also enables local mesh refinement in critical regions, thereby improving simulation accuracy <xref ref-type="bibr" rid="bib1.bibx44 bib1.bibx33 bib1.bibx28 bib1.bibx15 bib1.bibx17 bib1.bibx10" id="paren.4"/>. To leverage these advantages, the Institute of Atmospheric Physics, Chinese Academy of Sciences, and AMCG group at Imperial College London jointly developed Fluidity-Atmosphere, a 3D adaptive atmospheric model <xref ref-type="bibr" rid="bib1.bibx22 bib1.bibx11" id="paren.5"/>. Based on the unstructured FEM, Fluidity-Atmosphere tightly couples the Navier-Stokes equations, complete advection-diffusion dynamics, and anisotropic adaptive mesh algorithms, facilitating dynamic mesh optimization during simulations to capture multi-scale flow phenomena.</p>
      <p id="d2e183">However, while the FEM on unstructured meshes provides superior geometric flexibility, it also introduces inherent computational challenges that are absent in structured-grid models. In unstructured FEM frameworks such as Fluidity-Atmosphere, each element must frequently access irregularly connected nodes, edges, and faces during numerical integration and global sparse matrix assembly. This results in non-contiguous memory access, indirect addressing, and load imbalance across elements. These challenges are typically absent in structured grids, where data are stored in regular arrays with predictable access patterns. Moreover, the adaptive mesh refinement (AMR) feature of Fluidity-Atmosphere dynamically modifies the mesh topology during simulations, further complicating the memory layout. Consequently, the two most time-consuming components, namely, element-wise computations and global sparse matrix assembly, become dominated by irregular data movement rather than floating-point arithmetic, causing substantial memory bandwidth bottlenecks on CPU architectures <xref ref-type="bibr" rid="bib1.bibx42" id="paren.6"/>.</p>
      <p id="d2e189">Meanwhile, the slowdown of Moore's law and energy-efficiency bottlenecks of CPUs render CPU-only optimization inadequate for high-resolution simulations. In this context, GPUs have emerged as the widely used heterogeneous accelerators in high-performance computing. With massive parallelism and high-bandwidth device memory, GPUs have delivered orders-of-magnitude speedups in many numerical applications <xref ref-type="bibr" rid="bib1.bibx6" id="paren.7"/>. For atmospheric models, particularly those based on FEM, computational intensity primarily arises from solving large-scale linear systems, parameterizing physical processes, and performing massive element-level operations, each amenable to GPU acceleration. Reengineering these components to run efficiently on GPUs has therefore become a key research direction.</p>
      <p id="d2e196">Recent studies have reported significant advances in GPU acceleration of atmospheric models, including Navier-Stokes solvers <xref ref-type="bibr" rid="bib1.bibx14 bib1.bibx46" id="paren.8"/>, advection-diffusion equations <xref ref-type="bibr" rid="bib1.bibx39" id="paren.9"/>, and iterative solvers for implicit schemes <xref ref-type="bibr" rid="bib1.bibx29" id="paren.10"/>, typically with speedups of several orders of magnitude. Some Earth system models such as ocean model <xref ref-type="bibr" rid="bib1.bibx32" id="paren.11"/>, tsunami model <xref ref-type="bibr" rid="bib1.bibx5" id="paren.12"/>, cloud-resolving atmosphere model <xref ref-type="bibr" rid="bib1.bibx3" id="paren.13"/> and sea-ice model <xref ref-type="bibr" rid="bib1.bibx16" id="paren.14"/> also were accelerated using GPUs. FEM modules such as numerical integration <xref ref-type="bibr" rid="bib1.bibx24 bib1.bibx36" id="paren.15"/>, matrix assembly <xref ref-type="bibr" rid="bib1.bibx18 bib1.bibx37" id="paren.16"/>, and linear solvers <xref ref-type="bibr" rid="bib1.bibx34 bib1.bibx18" id="paren.17"/> have been explored, achieving speedups from tens to hundreds of times. Typically, the solution of large sparse linear systems is regarded as the dominant cost in FEM-based simulations and has been extensively studied, with numerous GPU-parallel implementations already achieving mature performance. However, the computation of elements and subsequent matrix assembly may pose significant challenges, primarily due to indirect memory addressing and low arithmetic intensity, particularly in large-scale unstructured atmospheric models. During each iteration, local element-wise matrices and right-hand sides must be computed and assembled into global sparse matrices. As the simulation domain expands, the number of mesh elements grows exponentially, leading to a proportional increase in computational cost. Without targeted optimization, the efficiency of the entire model can be substantially reduced.</p>
      <p id="d2e230">In practice, element-wise computations and global matrix assembly involve frequent access to unstructured mesh data, where the topological relationships among points, edges, faces, and cells are maintained through multiple interlinked index arrays. These indirect data access operations disrupt spatial locality and cause irregular memory traffic, resulting in low cache utilization and high latency. Therefore, despite substantial progress in GPU acceleration of structured or semi-structured atmospheric models, unstructured FEM-based frameworks–characterized by irregular data dependencies and adaptive mesh refinement–still lack systematic and efficient GPU solutions. To address this research gap, this study focuses on the GPU parallelization and optimization of the two most time-consuming components of Fluidity-Atmosphere: <italic>element-wise computations</italic> (e.g., evaluating local mass and stiffness matrices via numerical integration) and <italic>global matrix assembly</italic>. The primary contributions are as follows. <list list-type="order"><list-item>
      <p id="d2e241"><italic>GPU-based element-wise computations:</italic> Multiple optimization techniques are applied to accelerate compute-intensive kernels via GPU offloading, achieving more than 100<inline-formula><mml:math id="M3" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> speedups compared with CPU versions.</p></list-item><list-item>
      <p id="d2e254"><italic>GPU-based global matrix assembly:</italic> Parallel strategies are designed and implemented for element-wise assembly.</p></list-item><list-item>
      <p id="d2e260"><italic>Integrated GPU implementation:</italic> The GPU kernels are integrated into Fluidity-Atmosphere with pinned memory to optimize data transfer. The results show that GPU kernels facilitate speedups exceeding two orders of magnitude, with 4 MPI processes and one GPU, achieving an overall acceleration of 8.57 times compared to a single CPU process.</p></list-item></list></p>
      <p id="d2e265">The remainder of this paper is organized as follows. Section 2 introduces the program structure of Fluidity-Atmosphere, with emphasis on element-wise computations and global matrix assembly. Section 3 presents the GPU implementation and optimization of element-wise computations. Section 4 describes GPU-based global matrix assembly strategies. Section 5 evaluates the integrated GPU implementation and its performance. Section 6 concludes the paper, highlighting the future scope.</p>
</sec>
<sec id="Ch1.S2">
  <label>2</label><title>Target scientific application definition: fluidity-atmosphere</title>
      <p id="d2e276">Fluidity-Atmosphere employs a mixed finite element/finite volume framework that incorporates the Navier-Stokes momentum equations, advection-diffusion equations for potential temperature and water vapor, and the compressible continuity equation within an anisotropic adaptive mesh algorithm. This design facilitates dynamic co-optimization of the mesh and physical fields. Fluidity-Atmosphere solves a coupled system of nonlinear equations with (typically) time-varying solutions. The time-marching algorithm employed uses a nonlinear iteration scheme known as Picard iteration in which each equation is solved using the currently optimal solution for the other variables. The dynamical framework supports both the continuous Galerkin (CG) and discontinuous Galerkin (DG) finite element formulations. This study is focused exclusively on the CG scheme. Specifically, in the performance evaluations and benchmarks presented in this study, the dynamic framework predominantly utilizes linear Lagrange elements (<inline-formula><mml:math id="M4" display="inline"><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>) for the continuous Galerkin (CG) discretization. This choice inherently yields 4 degrees of freedom per tetrahedral element, resulting in the 4 <inline-formula><mml:math id="M5" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 4 local element matrices evaluated during the simulation.</p>

      <fig id="F1" specific-use="star"><label>Figure 1</label><caption><p id="d2e299">Time loop of Fluidity-Atmosphere.</p></caption>
        <graphic xlink:href="https://gmd.copernicus.org/articles/19/7525/2026/gmd-19-7525-2026-f01.png"/>

      </fig>

      <p id="d2e308">Figure 1 illustrates the general iteration loop of Fluidity-Atmosphere <xref ref-type="bibr" rid="bib1.bibx22" id="paren.18"/>. Fluidity-Atmosphere could invoke the adaptive algorithm at regular timesteps to ensure that the dynamics do not extend beyond the zone of adapted resolution. The adaptive algorithm has been parallelized using MPI processes, which run efficiently on CPUs. In our performance profiling, the nonlinear iterations were predominant. In each timestep iteration, the model solves governing equations such as the advection-diffusion and momentum equations. The processes of solving these equations in Fluidity-Atmosphere typically are divided into two major computational stages: <list list-type="custom"><list-item><label>1.</label>
      <p id="d2e316"><italic>Construct matrices:</italic> For each element, numerical integration is performed using Gaussian quadrature to evaluate local stiffness, mass matrices, and right-hand side (RHS) vectors. These local contributions are subsequently inserted into the global sparse matrix system.</p></list-item><list-item><label>2.</label>
      <p id="d2e322"><italic>Solution of linear systems:</italic> The assembled sparse matrices are solved using parallel numerical libraries such as PETSc, employing iterative solvers and preconditioners well-suited for large-scale sparse problems. The PETSc library has good performance and scalability. In the simulated case studies, its contribution to the overall runtime was a significant portion (37 %) in the CPU-only baseline, and as detailed later, it becomes the dominant bottleneck once the matrix assembly is accelerated on the GPU. In addition, PETSc already supports GPUs <xref ref-type="bibr" rid="bib1.bibx27" id="paren.19"/>. Hence, future studies will be focused on the integration of Fluidity-Atmosphere and PETSc GPU computing.</p></list-item></list></p>
      <p id="d2e331">The computation times of constructing matrices contributes significantly to the overall runtime. For instance, in the Mountainwave3D simulation, the number of mesh nodes and mesh elements are 103 635 and 554 394, respectively. Table 1 shows the computational times of constructing matrices in a timestep. In particular, the pressure diffusion matrix is used to calculate the stabilization term, which is only calculated at the first timestep or when the grid changes because of the adaptive process. In the Mountainwave3D simulation, the pressure diffusion matrix is computed every ten time steps. Without calculating the stabilization term, the total time of the timestep will decrease, but the proportion of time required to construct the matrices will slightly increase. As quantified in Table 1, the processes of constructing matrices account for approximately 60 % of the total timestep duration in both scenarios (with and without the stabilization term). This high proportion confirms that accelerating <italic>element-wise computations</italic> and <italic>global assembly</italic> is the most effective strategy for improving the overall performance of the Fluidity-Atmosphere model.</p>
      <p id="d2e340">The processes of constructing matrices share highly similar loop patterns. Hence, their general workflow can be summarized as follows: <list list-type="custom"><list-item><label>Step 1.</label>
      <p id="d2e346"><italic>Preparing Data:</italic> Initialize element mass matrices, right-hand sides, and local physical variables (such as velocity, density, and viscosity). </p></list-item><list-item><label>Step 2.</label>
      <p id="d2e353"><italic>Coordinate transformation:</italic> The <monospace>transform_</monospace><monospace>to_physical</monospace> function maps shape function gradients into physical space and computes Gaussian weights (<monospace>detwei</monospace>).</p></list-item><list-item><label>Step 3.</label>
      <p id="d2e368"><italic>Setup test function.</italic></p></list-item><list-item><label>Step 4.</label>
      <p id="d2e373"><italic>Contribution evaluation:</italic> Call physics modules to accumulate contributions into element matrices and right-hand sides. Several functions, each corresponding to a distinct physical process, such as mass, advection, diffusion, source, and viscosity terms, are called. In addition to the calculation of physical variables, these functions extensively call the integration calculation function on the elements such as <monospace>dshape_tensor_dshape</monospace>, <monospace>shape_dshape</monospace>, and <monospace>shape_shape</monospace>. In the process of assembling the pressure diffusion matrix, a special function named <monospace>get_edge_lengths</monospace> (in which the small matrix solver is called) is used to calculate the element length scale in the physical space. All these functions are well-suited for GPU parallel computing.</p></list-item><list-item><label>Step 5.</label>
      <p id="d2e391"><italic>Assembly:</italic> Local contributions are inserted into the global matrix and RHS. The subsequent irregular insertion operations into the global matrix makes it highly data-intensive and memory-bound. Owing to the massive throughput of GPU device memory, these functions can be accelerated on GPUs.</p></list-item></list></p>
      <p id="d2e396">Therefore, the following sections describe the implementation and performance optimization methods of the element-wise computations and global matrix assembly on the GPU.</p>
      <p id="d2e399">Although the current implementation is specifically optimized for the NVIDIA A100 GPU utilizing the CUDA programming model <xref ref-type="bibr" rid="bib1.bibx30" id="paren.20"/>, the proposed optimization strategies are largely hardware-agnostic at the algorithmic level. Specifically, the element-wise parallelization strategy (one-thread-per-element), the data layout design, and the sparse matrix assembly workflow are general and can be applied to other GPU architectures.</p>

<table-wrap id="T1" specific-use="star"><label>Table 1</label><caption><p id="d2e408">The computation times of constructing matrices. Bold values highlight the total computational time and percentage for the processes of constructing matrices.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="5">
     <oasis:colspec colnum="1" colname="col1" align="justify" colwidth="60mm"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right" colsep="1"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:colspec colnum="5" colname="col5" align="right"/>
     <oasis:thead>
       <oasis:row>
         <oasis:entry colname="col1" align="left">Functional module</oasis:entry>
         <oasis:entry rowsep="1" namest="col2" nameend="col3" align="center" colsep="1">With stabilization term </oasis:entry>
         <oasis:entry rowsep="1" namest="col4" nameend="col5" align="center">Without stabilization term </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left"/>
         <oasis:entry colname="col2">CPU time (<inline-formula><mml:math id="M6" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">ms</mml:mi></mml:mrow></mml:math></inline-formula>)</oasis:entry>
         <oasis:entry colname="col3">Percentage (%)</oasis:entry>
         <oasis:entry colname="col4">CPU time (<inline-formula><mml:math id="M7" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">ms</mml:mi></mml:mrow></mml:math></inline-formula>)</oasis:entry>
         <oasis:entry colname="col5">Percentage (%)</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1" align="left">The advection-diffusion matrix (Temperature)</oasis:entry>
         <oasis:entry colname="col2">2477.545</oasis:entry>
         <oasis:entry colname="col3">4.84</oasis:entry>
         <oasis:entry colname="col4">2466.594</oasis:entry>
         <oasis:entry colname="col5">5.95</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1" align="left">The advection-diffusion matrix (WaterVapor)</oasis:entry>
         <oasis:entry colname="col2">2477.194</oasis:entry>
         <oasis:entry colname="col3">4.82</oasis:entry>
         <oasis:entry colname="col4">2471.302</oasis:entry>
         <oasis:entry colname="col5">5.96</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1" align="left">The advection-diffusion matrix (CloudWater)</oasis:entry>
         <oasis:entry colname="col2">2466.853</oasis:entry>
         <oasis:entry colname="col3">4.82</oasis:entry>
         <oasis:entry colname="col4">2494.424</oasis:entry>
         <oasis:entry colname="col5">6.02</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1" align="left">The advection-diffusion matrix (RainWater)</oasis:entry>
         <oasis:entry colname="col2">1785.827</oasis:entry>
         <oasis:entry colname="col3">3.49</oasis:entry>
         <oasis:entry colname="col4">1799.104</oasis:entry>
         <oasis:entry colname="col5">4.34</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1" align="left">The momentum matrix</oasis:entry>
         <oasis:entry colname="col2">13 029.408</oasis:entry>
         <oasis:entry colname="col3">25.47</oasis:entry>
         <oasis:entry colname="col4">12 093.145</oasis:entry>
         <oasis:entry colname="col5">29.19</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1" align="left">The divergence matrix</oasis:entry>
         <oasis:entry colname="col2">3662.082</oasis:entry>
         <oasis:entry colname="col3">7.16</oasis:entry>
         <oasis:entry colname="col4">3634.111</oasis:entry>
         <oasis:entry colname="col5">8.77</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1" align="left">The pressure diffusion matrix</oasis:entry>
         <oasis:entry colname="col2">3505.105</oasis:entry>
         <oasis:entry colname="col3">6.85</oasis:entry>
         <oasis:entry colname="col4"/>
         <oasis:entry colname="col5"/>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">The projection matrix</oasis:entry>
         <oasis:entry colname="col2">1126.747</oasis:entry>
         <oasis:entry colname="col3">2.20</oasis:entry>
         <oasis:entry colname="col4">1125.822</oasis:entry>
         <oasis:entry colname="col5">2.72</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1" align="left">Total of constructing matrices</oasis:entry>
         <oasis:entry colname="col2"><bold>30 530.760</bold></oasis:entry>
         <oasis:entry colname="col3">59.68</oasis:entry>
         <oasis:entry colname="col4"><bold>26 084.502</bold></oasis:entry>
         <oasis:entry colname="col5"><bold>62.96</bold></oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Others</oasis:entry>
         <oasis:entry colname="col2">20 625.871</oasis:entry>
         <oasis:entry colname="col3">40.32</oasis:entry>
         <oasis:entry colname="col4">15 348.450</oasis:entry>
         <oasis:entry colname="col5">37.04</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1" align="left">Total timestep duration</oasis:entry>
         <oasis:entry colname="col2">51 156.631</oasis:entry>
         <oasis:entry colname="col3">100.00</oasis:entry>
         <oasis:entry colname="col4">41 432.952</oasis:entry>
         <oasis:entry colname="col5">100.00</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d2e679">However, certain implementation details rely strictly on CUDA-specific hardware features, such as double-precision atomic operations in the L2 cache, memory hierarchy optimization, and specific kernel launch configurations. Porting the implementation to other platforms, such as AMD GPUs, would primarily involve adapting CUDA-specific APIs to alternative frameworks like HIP <xref ref-type="bibr" rid="bib1.bibx2" id="paren.21"/>. Since HIP provides similar abstractions for thread hierarchy, memory management, and atomic operations, the overall syntax porting effort is expected to be moderate.</p>
      <p id="d2e685">Nevertheless, achieving true performance portability would require additional tuning to account for architectural differences in memory bandwidth, cache structure, and execution models. For instance, while NVIDIA V100 and A100 GPUs handle massive atomic updates highly efficiently, evaluating similar unstructured-grid applications on AMD CDNA architectures (e.g., MI100) has shown that standard atomic updates can incur higher penalties, sometimes necessitating alternative data restructuring or register-based aggregation to achieve optimal efficiency <xref ref-type="bibr" rid="bib1.bibx41" id="paren.22"/>. Overall, the proposed algorithmic approach is expected to maintain its fundamental effectiveness across heterogeneous architectures, although achieving peak optimal performance on different platforms will inevitably require platform-specific tuning.</p>
</sec>
<sec id="Ch1.S3">
  <label>3</label><title>GPU parallelization of element-wise computations</title>
      <p id="d2e699">In Fluidity-Atmosphere, element-level operations such as numerical integration and coordinate transformation constitute the core of unstructured finite element-wise computations. Profiling results indicate that these routines are invoked millions of times per timestep. Each nonlinear iteration involves re-evaluating element Jacobians, transforming basis functions, and accumulating physical contributions for all elements, which cumulatively dominate the computational workload. These observations identify numerical integration and coordinate transformation as the primary performance bottlenecks and thus the key focus of GPU parallelization in this study. The computation for each mesh element can be performed independently of others, making this step an ideal candidate for parallelization <xref ref-type="bibr" rid="bib1.bibx12" id="paren.23"/>. GPUs feature a massively parallel processor architecture, which enables the simultaneous launch of a large number of parallel threads. These threads can be associated with different mesh elements to execute element-wise computations. These results are stored for subsequent assembly into the global matrix. Building on the workflow analysis in Section 2, this section presents the GPU implementation and optimization strategies for key subroutines, followed by a summary of general acceleration techniques.</p>

      <fig id="F2" specific-use="star"><label>Figure 2</label><caption><p id="d2e707">Data layout of elements and variables.</p></caption>
        <graphic xlink:href="https://gmd.copernicus.org/articles/19/7525/2026/gmd-19-7525-2026-f02.png"/>

      </fig>

<sec id="Ch1.S3.SS1">
  <label>3.1</label><title>Data structures and parallelization strategy</title>
      <p id="d2e723">In unstructured meshes, element connectivity must be explicitly stored. Each tetrahedral element in Fluidity-Atmosphere records four vertex indices. Physical variables are stored in a unified Fields data structure organized as (number of vertices <inline-formula><mml:math id="M8" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> dimension) arrays: scalars (1D), vectors (3D), and tensors (9D). The same layout is adopted in GPU memory to ensure contiguous access and consistent indexing (Fig. 2). Common fields such as density and temperature are stored as scalars, positions and velocities as vectors, and viscosity terms as tensors.</p>
      <p id="d2e733">The element-wise computations in the FEM are intrinsically independent, making them particularly suitable for massively parallel execution on GPU architectures. Unlike spectral methods that involve global communication or high-order finite-difference schemes relying on extended stencils, finite-element computations are confined strictly within each element boundary. It is important to clarify that while Fluidity-Atmosphere utilizes a mixed CG/DG discretization, the GPU acceleration in this study focuses exclusively on the Continuous Galerkin (CG) operators (e.g., for the momentum and advection-diffusion equations). In the CG formulation, the numerical integration to evaluate local matrices remains mathematically entirely local. Unlike DG methods, which require calculating inter-element fluxes and jumps over facets, the CG element-wise computations do not introduce cross-element mathematical dependencies during the integration phase. The global coupling arising from shared mesh nodes is manifested exclusively during the global matrix assembly phase through indirect memory access.</p>
      <p id="d2e736">This strong locality offers two principal advantages: (1) all element-wise computations can proceed independently, eliminating inter-thread communication during the compute phase; and (2) the workload is naturally balanced across elements, minimizing synchronization overhead. Consequently, a “one-thread-per-element” parallelization strategy is adopted. Each GPU thread handles one element. In our implementation, thread block sizes of 128 and 256 were both utilized, with the specific parameter for each kernel determined empirically via NVIDIA Nsight Compute. For the target mesh comprising 554 394 elements, a block size of 128 yields 4332 blocks, while 256 yields 2166 blocks. Given that the NVIDIA A100 GPU features 108 Streaming Multiprocessors (SMs), each with a theoretical maximum of 2048 resident threads, both configurations generate sufficient numbers of thread blocks to provide approximately 2.5 execution waves across the GPU, which helps maintain high occupancy and effectively hide memory latency.</p>
</sec>
<sec id="Ch1.S3.SS2">
  <label>3.2</label><title>GPU parallel coordinate transformation</title>
      <p id="d2e748">With the data layout on GPUs established, the next step is to implement the core finite element operation: coordinate transformation. Using a reference element (in this study, a tetrahedron), shape functions and Gauss points are consistently defined. In Fluidity-Atmosphere, each physical element employs 11 Gauss points: four vertices, six edge midpoints, and the centroid (Fig. 3). For each Gauss point, the transformation routine evaluates the integrand and multiplies it by the predefined weight.</p>

      <fig id="F3"><label>Figure 3</label><caption><p id="d2e753">The reference element and Gauss points.</p></caption>
          <graphic xlink:href="https://gmd.copernicus.org/articles/19/7525/2026/gmd-19-7525-2026-f03.png"/>

        </fig>

      <p id="d2e762">Even though element-independent and theoretically parallelizable, this kernel is computationally intensive and one of the most frequently executed routines in the model. For every iteration and every element, it performs numerous small matrix operations: matrix multiplications, adjoint and determinant evaluations, and transformations of derivative matrices.</p>
      <p id="d2e766">Algorithm 1 outlines the <monospace>transform_to_physical</monospace> procedure, where input and output arrays are stored contiguously to maximize memory efficiency. Crucially, the algorithm incorporates a conditional optimization for linear (<inline-formula><mml:math id="M9" display="inline"><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>) elements: because standard affine mappings yield a constant Jacobian matrix, expensive operations (Jacobian construction, determinant, and inverse) are evaluated only at the first Gauss point (<inline-formula><mml:math id="M10" display="inline"><mml:mrow><mml:msub><mml:mi>g</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula>) and cached.</p><boxed-text content-type="algorithm" position="float" id="Ch1.Prog1"><label>Algorithm 1</label><caption><p id="d2e799">Transform_to_physical.</p></caption><disp-quote content-type="algorithmic" specific-use="numbering{0}"><list>

    <list-item>

      <p id="d2e806" specific-use="STATE"><bold>Input:</bold> Element nodal coordinates <inline-formula><mml:math id="M11" display="inline"><mml:mi>X</mml:mi></mml:math></inline-formula>, shape functions <inline-formula><mml:math id="M12" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula>, reference derivatives <inline-formula><mml:math id="M13" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">∇</mml:mi><mml:mi mathvariant="italic">ξ</mml:mi></mml:msub><mml:mi>N</mml:mi></mml:mrow></mml:math></inline-formula>, Gauss points <inline-formula><mml:math id="M14" display="inline"><mml:mrow><mml:msub><mml:mi>g</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula></p>
            </list-item>

    <list-item>

      <p id="d2e851" specific-use="STATE"><bold>Output:</bold> Physical derivatives <inline-formula><mml:math id="M15" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">∇</mml:mi><mml:mi>x</mml:mi></mml:msub><mml:mi>N</mml:mi></mml:mrow></mml:math></inline-formula>, weights detwei, <inline-formula><mml:math id="M16" display="inline"><mml:mi mathvariant="bold">J</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M17" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold">J</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>, det <inline-formula><mml:math id="M18" display="inline"><mml:mi mathvariant="bold">J</mml:mi></mml:math></inline-formula></p>
            </list-item>

    <list-item>

      <p id="d2e899" specific-use="FOR"><bold>for</bold> <inline-formula><mml:math id="M19" display="inline"><mml:mrow><mml:msub><mml:mi>g</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mtext> to </mml:mtext><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mi>g</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>  <bold>do</bold> <list>
    <list-item>
      <p id="d2e935" specific-use="STATE">compute <inline-formula><mml:math id="M20" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">∇</mml:mi><mml:mi mathvariant="italic">ξ</mml:mi></mml:msub><mml:mi mathvariant="bold-italic">N</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>g</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula></p></list-item>
    <list-item>
      <p id="d2e961" specific-use="IF"><bold>if</bold> mapping is nonlinear <bold>or</bold> <inline-formula><mml:math id="M21" display="inline"><mml:mrow><mml:msub><mml:mi>g</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> <bold>then</bold> <list>
    <list-item>
      <p id="d2e992" specific-use="STATE">construct Jacobian matrix <inline-formula><mml:math id="M22" display="inline"><mml:mrow><mml:mi mathvariant="bold">J</mml:mi><mml:mo>←</mml:mo><mml:mi mathvariant="bold">X</mml:mi><mml:mo>⋅</mml:mo><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold">∇</mml:mi><mml:mi mathvariant="italic">ξ</mml:mi></mml:msub><mml:mi mathvariant="bold-italic">N</mml:mi><mml:msup><mml:mo>)</mml:mo><mml:mi>T</mml:mi></mml:msup></mml:mrow></mml:math></inline-formula></p></list-item>
    <list-item>
      <p id="d2e1024" specific-use="STATE">compute <inline-formula><mml:math id="M23" display="inline"><mml:mrow><mml:mtext>adj</mml:mtext><mml:mo>(</mml:mo><mml:mi mathvariant="bold">J</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and det <inline-formula><mml:math id="M24" display="inline"><mml:mi mathvariant="bold">J</mml:mi></mml:math></inline-formula></p></list-item>
    <list-item>
      <p id="d2e1049" specific-use="STATE">normalize inverse matrix <inline-formula><mml:math id="M25" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold">J</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup><mml:mo>←</mml:mo><mml:mtext>adj</mml:mtext><mml:mo>(</mml:mo><mml:mi mathvariant="bold">J</mml:mi><mml:mo>)</mml:mo><mml:mo>/</mml:mo><mml:mi mathvariant="normal">det</mml:mi><mml:mi mathvariant="bold">J</mml:mi></mml:mrow></mml:math></inline-formula></p></list-item></list></p></list-item>
    <list-item>
      <p id="d2e1083" specific-use="ENDIF"><bold>end</bold> <bold>if</bold></p></list-item>
    <list-item>
      <p id="d2e1092" specific-use="STATE">transform to physical derivatives <inline-formula><mml:math id="M26" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">∇</mml:mi><mml:mi>x</mml:mi></mml:msub><mml:mi mathvariant="bold-italic">N</mml:mi><mml:mo>←</mml:mo><mml:msup><mml:mi mathvariant="bold">J</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup><mml:mo>⋅</mml:mo><mml:msub><mml:mi mathvariant="bold">∇</mml:mi><mml:mi mathvariant="italic">ξ</mml:mi></mml:msub><mml:mi mathvariant="bold-italic">N</mml:mi></mml:mrow></mml:math></inline-formula></p></list-item>
    <list-item>
      <p id="d2e1128" specific-use="STATE">compute quadrature weights <inline-formula><mml:math id="M27" display="inline"><mml:mrow><mml:mtext>detwei</mml:mtext><mml:mo>(</mml:mo><mml:msub><mml:mi>g</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>←</mml:mo><mml:mo>|</mml:mo><mml:mi mathvariant="normal">det</mml:mi><mml:mi mathvariant="bold">J</mml:mi><mml:mo>|</mml:mo><mml:mo>⋅</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>g</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula></p></list-item></list></p>
            </list-item>

    <list-item>

      <p id="d2e1170" specific-use="ENDFOR"><bold>end</bold> <bold>for</bold></p>
            </list-item>
          </list></disp-quote></boxed-text>
      <p id="d2e1179">The inner loop over Gauss points is deliberately retained to preserve framework generality for higher-order elements. Furthermore, rather than precomputing and storing these geometric quantities, which would exacerbate memory traffic in this heavily memory-bound framework, we employ a “compute-on-the-fly” strategy, trading abundant GPU arithmetic cycles to conserve precious memory bandwidth <xref ref-type="bibr" rid="bib1.bibx20 bib1.bibx4" id="paren.24"/>.</p>
</sec>
<sec id="Ch1.S3.SS3">
  <label>3.3</label><title>GPU parallelization of element-wise integration</title>
      <p id="d2e1194">Element-wise integration routines constitute the computational core of unstructured finite-element models, as they are responsible for evaluating the contributions of each element to the global system. Profiling of Fluidity-Atmosphere indicates that these kernels are among the most frequently executed routines. During each nonlinear iteration, they are called once per element per physical equation, resulting in millions of invocations per timestep. As each call involves multi-dimensional tensor operations and numerical integration over multiple Gauss points, the accumulated cost dominates both the element-computation phase and the overall runtime of the model. Therefore, optimizing element integration is critical to achieving substantial end-to-end performance gains. The principal integration functions are as follows: <list list-type="bullet"><list-item>
      <p id="d2e1199"><italic><monospace>dshape_tensor_dshape</monospace></italic> couples shape function gradients with tensor fields, supporting stabilization terms and physical field coupling.</p></list-item><list-item>
      <p id="d2e1206"><italic><monospace>shape_shape</monospace></italic> integrates shape functions to construct the mass matrix, required in momentum and energy conservation equations.</p></list-item><list-item>
      <p id="d2e1213"><italic><monospace>shape_dshape</monospace></italic> performs weighted integration of shape functions and their gradients at nodal points.</p></list-item></list></p>
      <p id="d2e1219">Among these functions, the <monospace>dshape_tensor_dshape</monospace> function is the most computationally intensive, as it involves both tensor-vector and vector-vector multiplications for every integration point. On the GPU, this procedure is parallelized at the thread level, with each thread processing one element. Beyond kernel-level parallelization, a series of GPU-specific performance tuning techniques are applied to completely exploit the hardware potential. These optimizations are guided by profiling by employing NVIDIA Nsight Compute, focusing on reducing memory latency, register pressure, and control overhead:</p>
      <p id="d2e1225"><list list-type="bullet">
            <list-item>

      <p id="d2e1231"><italic>Memory optimization:</italic> The <monospace>__restrict__</monospace> qualifier is applied to pointer variables to eliminate aliasing and enable more aggressive memory-access optimizations. Small <monospace>__device__</monospace> functions are annotated with <monospace>__forceinline__</monospace> to reduce function-call overhead.</p>
            </list-item>
            <list-item>

      <p id="d2e1248"><italic>Loop optimization:</italic> Loop-invariant expressions (particularly global memory accesses) are hoisted outside the loop body; manual loop unrolling or <monospace>#pragma unroll</monospace> directives are applied to innermost loops to reduce control overhead and increase instruction-level parallelism.</p>
            </list-item>
            <list-item>

      <p id="d2e1259"><italic>Template metaprogramming:</italic> As the dimensions of the different matrices and vectors are already known at compile time, the <monospace>__device__</monospace> functions such as matrix-matrix product and dot product are implemented as templates, facilitating the compiler to apply deeper optimizations for specific instantiations.</p>
            </list-item>
          </list></p>
      <p id="d2e1269">These measures are reflected in Nsight Compute reports as improved register utilization and reduced memory bottlenecks. Loop unrolling and templating yield substantial enhancements in small-scale computations. In combination with GPU-parallel element integration and small-matrix solvers, these optimizations form a cohesive strategy that maximizes performance while maintaining full consistency with the original CPU implementation.</p>
      <p id="d2e1273">It should be noted that the CPU baseline used for comparison represents the official, unmodified production code of Fluidity-Atmosphere, compiled with aggressive optimization flags <monospace>-O3 -ffast-math</monospace>. Under these flags, modern CPU compilers automatically apply substantial loop unrolling and vectorization. However, due to the warp-based execution model and strict register allocation of GPUs, the CUDA compiler frequently requires explicit manual directives (e.g., <monospace>#pragma unroll</monospace>) and compile-time size resolution (via templates) to achieve optimal instruction scheduling and register reuse. Thus, these GPU-specific manual optimizations are implemented to fully unlock the hardware's architectural potential rather than to introduce an algorithmic disparity. The optimized code is listed in Appendix A.</p>
</sec>
<sec id="Ch1.S3.SS4">
  <label>3.4</label><title>GPU parallelization of small-scale linear algebra solvers</title>
      <p id="d2e1290">In addition to numerical integration, the construction of stabilization terms involves numerous small-scale linear algebraic operations, such as solving 6 <inline-formula><mml:math id="M28" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 6 systems or computing eigen-decompositions of 3 <inline-formula><mml:math id="M29" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 3 matrices. Even though these operations are apparently lightweight, they are called repeatedly for every element and Gauss point, resulting in a substantial cumulative cost. On CPUs, relying on standard library routines such as LAPACK's <monospace>DSPEV</monospace> or Fortran's intrinsic solvers for these operations is highly inefficient. While a small 6 <inline-formula><mml:math id="M30" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 6 matrix easily fits into the L1 data cache, gathering these matrices from an unstructured mesh via indirect addressing disrupts spatial locality. Furthermore, these independent tiny matrices offer minimal opportunity for data reuse, limiting cache efficiency. From a software perspective, pre-compiled external library routines typically cannot be automatically inlined by the compiler within the tight element loop. Invoking them millions of times introduces profound function-call overhead, and their internal argument validation and branching dwarf the actual floating-point arithmetic.</p>
      <p id="d2e1317">To address this challenge, this study implements dedicated GPU-based small-matrix solvers and eigen-decomposition kernels as inline <monospace>__device__</monospace> operators. <list list-type="bullet"><list-item>
      <p id="d2e1325"><italic>Small linear systems (e.g., 6</italic> <inline-formula><mml:math id="M31" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> <italic>6):</italic> These are solved on GPUs using Gaussian elimination. The input is an augmented matrix <inline-formula><mml:math id="M32" display="inline"><mml:mi mathvariant="bold">A</mml:mi></mml:math></inline-formula> (in column-major layout, including the right-hand side) with dimensions m<inline-formula><mml:math id="M33" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula>n. Pivoting and row exchanges are applied to control numerical errors, whereas column-major storage and in-place operations reduce register pressure and facilitate the alignment with GPU memory characteristics.</p></list-item><list-item>
      <p id="d2e1355"><italic>Eigen-decomposition of 3</italic> <inline-formula><mml:math id="M34" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> <italic>3 real symmetric matrices:</italic> Both iterative and direct GPU solvers are implemented, based on the methods employed in prior research. To improve efficiency, the implementation employs template-based array classes and includes sorting functions for eigenvalues and eigenvectors. By exploiting symmetry, only six entries (the diagonal and upper triangular part) are stored, minimizing register usage and avoiding spills to local memory.</p></list-item></list></p>
      <p id="d2e1370">These GPU-implemented solvers are lightweight and reusable across kernels. The modular <monospace>__device__</monospace> design also promotes code reuse across physical modules and supports further fusion with integration kernels.</p>
</sec>
</sec>
<sec id="Ch1.S4">
  <label>4</label><title>GPU parallelization of matrix assembly</title>
      <p id="d2e1385">In the FEM, matrix assembly serves as the critical bridge between local element-wise computations and the solution of global linear systems. As this stage involves extensive operations on sparse data structures and frequent memory access, it typically hinders the performance of the entire program. Therefore, migrating the matrix assembly procedure to GPUs and applying targeted optimizations is crucial for enhancing the computational efficiency of atmospheric models. This section introduces the global matrix storage formats, the element-wise parallel assembly strategy, and the GPU implementation of RHS vector assembly.</p>
      <p id="d2e1388">Before detailing our assembly strategy, it is worth noting that a prominent trend in high-performance finite-element simulations is to discard global matrix assembly in favor of matrix-free (or partially assembled) operator evaluations <xref ref-type="bibr" rid="bib1.bibx20 bib1.bibx21 bib1.bibx35" id="paren.25"/>. Matrix-free algorithms were fundamentally developed to address the performance bottlenecks inherent in high-order finite elements. In high-order discretizations, full matrix assembly leads to an exponential explosion in memory footprint and floating-point operations per degree of freedom (DoF). Matrix-free approaches, typically coupled with sum factorization, suppress this complexity, thereby drastically increasing computational intensity and parallel efficiency <xref ref-type="bibr" rid="bib1.bibx8" id="paren.26"/>.</p>
      <p id="d2e1397">However, the dynamic framework evaluated in this study predominantly employs low-order linear elements (<inline-formula><mml:math id="M35" display="inline"><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>). For low-order FEM, the memory access per DoF is intrinsically small and remains comparable whether using fully assembled operators or on-the-fly matrix-free evaluations. Since low-order evaluations are heavily memory-bound, transitioning to matrix-free methods does not yield the profound bandwidth savings observed in high-order regimes. Furthermore, explicitly assembling the global sparse matrix provides a critical advantage: it enables the use of highly optimized, generic Sparse Matrix-Vector multiplication (SpMV) kernels and sophisticated algebraic preconditioners, such as Algebraic Multigrid (AMG) and Incomplete LU (ILU) provided by PETSc, which are indispensable for the stability and efficiency of the implicit time-stepping solvers in Fluidity-Atmosphere. Therefore, accelerating the global sparse matrix assembly remains the most rational and impactful optimization pathway for this class of low-order dynamical cores.</p>
<sec id="Ch1.S4.SS1">
  <label>4.1</label><title>Global matrix assembly</title>
      <p id="d2e1418">Following local integration, the algorithm accumulates (<inline-formula><mml:math id="M36" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold">A</mml:mi><mml:mi>e</mml:mi></mml:msup></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M37" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">F</mml:mi><mml:mi>e</mml:mi></mml:msup></mml:mrow></mml:math></inline-formula>) into the global matrix <inline-formula><mml:math id="M38" display="inline"><mml:mi mathvariant="bold">A</mml:mi></mml:math></inline-formula> and vector <inline-formula><mml:math id="M39" display="inline"><mml:mi mathvariant="bold-italic">F</mml:mi></mml:math></inline-formula> according to the connectivity information of the mesh. This procedure, widely recognized as the global assembly, transforms element-level contributions into a consistent global representation based on the mapping between local and global degrees of freedom, as shown in Algorithm 2. Here, <inline-formula><mml:math id="M40" display="inline"><mml:mi mathvariant="italic">ε</mml:mi></mml:math></inline-formula> is the set of elements and <inline-formula><mml:math id="M41" display="inline"><mml:mi>e</mml:mi></mml:math></inline-formula> is an element. The <inline-formula><mml:math id="M42" display="inline"><mml:mrow><mml:mi>e</mml:mi><mml:mi>l</mml:mi><mml:mi>e</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:math></inline-formula> function performs element-wise computations to get element matrices. The <inline-formula><mml:math id="M43" display="inline"><mml:mi>L</mml:mi></mml:math></inline-formula> function retrieves the node numbering.</p><boxed-text content-type="algorithm" position="float" id="Ch1.Prog2"><label>Algorithm 2</label><caption><p id="d2e1494">The Global Matrix Assembly.</p></caption><disp-quote content-type="algorithmic" specific-use="numbering{0}"><list>

    <list-item>

      <p id="d2e1501" specific-use="STATE"><bold>Output:</bold> Global matrices <inline-formula><mml:math id="M44" display="inline"><mml:mi mathvariant="bold">A</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M45" display="inline"><mml:mi mathvariant="bold-italic">F</mml:mi></mml:math></inline-formula></p>
            </list-item>

    <list-item>

      <p id="d2e1522" specific-use="STATE">Initialize <inline-formula><mml:math id="M46" display="inline"><mml:mi mathvariant="bold">A</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M47" display="inline"><mml:mi mathvariant="bold-italic">F</mml:mi></mml:math></inline-formula> to zero</p>
            </list-item>

    <list-item>

      <p id="d2e1542" specific-use="FORALL"><bold>for all</bold> <inline-formula><mml:math id="M48" display="inline"><mml:mrow><mml:mi>e</mml:mi><mml:mo>∈</mml:mo><mml:mi mathvariant="italic">ε</mml:mi></mml:mrow></mml:math></inline-formula> <bold>do</bold> <list>
    <list-item>
      <p id="d2e1565" specific-use="STATE">Compute <inline-formula><mml:math id="M49" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold">A</mml:mi><mml:mi>e</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">F</mml:mi><mml:mi>e</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mtext>elem</mml:mtext><mml:mo>(</mml:mo><mml:mi>e</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula></p></list-item>
    <list-item>
      <p id="d2e1601" specific-use="FORALL"><bold>for all</bold> local degrees of freedom <inline-formula><mml:math id="M50" display="inline"><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> of <inline-formula><mml:math id="M51" display="inline"><mml:mi>e</mml:mi></mml:math></inline-formula> <bold>do</bold> <list>
    <list-item>
      <p id="d2e1630" specific-use="STATE"><inline-formula><mml:math id="M52" display="inline"><mml:mrow><mml:mi mathvariant="bold-italic">F</mml:mi><mml:mo>(</mml:mo><mml:mi>L</mml:mi><mml:mo>(</mml:mo><mml:mi>e</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>)</mml:mo><mml:mo>)</mml:mo><mml:mo>+</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">F</mml:mi><mml:mi>e</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula></p></list-item>
    <list-item>
      <p id="d2e1679" specific-use="FORALL"><bold>for all</bold> local degrees of freedom <inline-formula><mml:math id="M53" display="inline"><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> of <inline-formula><mml:math id="M54" display="inline"><mml:mi>e</mml:mi></mml:math></inline-formula> <bold>do</bold> <list>
    <list-item>
      <p id="d2e1708" specific-use="STATE"><inline-formula><mml:math id="M55" display="inline"><mml:mrow><mml:mi mathvariant="bold">A</mml:mi><mml:mo>(</mml:mo><mml:mi>L</mml:mi><mml:mo>(</mml:mo><mml:mi>e</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:mi>L</mml:mi><mml:mo>(</mml:mo><mml:mi>e</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub><mml:mo>)</mml:mo><mml:mo>)</mml:mo><mml:mo>+</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="bold">A</mml:mi><mml:mi>e</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula></p></list-item></list></p></list-item>
    <list-item>
      <p id="d2e1782" specific-use="ENDFOR"><bold>end</bold> <bold>for</bold></p></list-item></list></p></list-item>
    <list-item>
      <p id="d2e1791" specific-use="ENDFOR"><bold>end</bold> <bold>for</bold></p></list-item></list></p>
            </list-item>

    <list-item>

      <p id="d2e1801" specific-use="ENDFOR"><bold>end</bold> <bold>for</bold></p>
            </list-item>
          </list></disp-quote></boxed-text>
      <p id="d2e1810">In contrast to structured-grid models, where mesh nodes are arranged in a regular <inline-formula><mml:math id="M56" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> indexing order and memory access can be performed sequentially and predictably, unstructured meshes lack geometric regularity. The connectivity of each element varies depending on its neighbours, and the mapping between local and global indices must be retrieved indirectly through lookup tables. Consequently, the assembly phase in unstructured FEM involves frequent non-contiguous global memory reads and writes, with irregular strides between consecutive memory locations. This destroys spatial locality and drastically increases cache miss rates, thereby amplifying the ratio of memory access time to floating-point computation.</p>
      <p id="d2e1834">On GPU architectures, these irregular access patterns exacerbate performance degradation. Concurrent threads typically need to update overlapping entries of the global sparse matrix, requiring synchronization or the use of atomic operations to ensure correctness. The combined effects of irregular global memory traffic, low cache reuse, and write conflicts make matrix assembly a highly memory-bound process.</p>
      <p id="d2e1838">Consequently, global matrix assembly presents two major challenges: (1) the imbalance between memory traffic and arithmetic intensity, caused by the dominance of irregular global data movement over computation; (2) the difficulty of ensuring data consistency during concurrent updates. Cumulatively, these factors make matrix assembly another critical performance bottleneck, and thus a central focus of GPU parallel optimization in this work.</p>
</sec>
<sec id="Ch1.S4.SS2">
  <label>4.2</label><title>Sparse matrix storage formats</title>
      <p id="d2e1850">Fluidity-Atmosphere employs different sparse storage formats for different types of physical fields. For scalar field matrices, the <italic>compressed sparse row (CSR)</italic> format is used. For vector field matrices such as three-component velocity, the <italic>block-CSR (BCSR)</italic> format is adopted.</p>
      <p id="d2e1859">For a scalar field, the local stiffness matrix can be expressed as Eq. (1):

            <disp-formula id="Ch1.E1" content-type="numbered"><label>1</label><mml:math id="M57" display="block"><mml:mrow><mml:msub><mml:mi>K</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:munder><mml:mo movablelimits="false">∫</mml:mo><mml:mrow><mml:msub><mml:mi mathvariant="normal">Ω</mml:mi><mml:mi>e</mml:mi></mml:msub></mml:mrow></mml:munder><mml:mo>(</mml:mo><mml:mi mathvariant="bold">∇</mml:mi><mml:msub><mml:mi mathvariant="bold-italic">N</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msup><mml:mo>)</mml:mo><mml:mi>T</mml:mi></mml:msup><mml:mi mathvariant="bold">C</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold">∇</mml:mi><mml:msub><mml:mi mathvariant="bold-italic">N</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mi mathvariant="normal">d</mml:mi><mml:mi mathvariant="normal">Ω</mml:mi></mml:mrow></mml:math></disp-formula>

          where <inline-formula><mml:math id="M58" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M59" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> are basis functions, <inline-formula><mml:math id="M60" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">Ω</mml:mi><mml:mi>e</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the integration domain of the element, and <inline-formula><mml:math id="M61" display="inline"><mml:mi mathvariant="bold">C</mml:mi></mml:math></inline-formula> is the associated tensor. For scalar fields, <inline-formula><mml:math id="M62" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is a scalar and <inline-formula><mml:math id="M63" display="inline"><mml:mrow><mml:msub><mml:mi>K</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> is a scalar integral, resulting in a 4 <inline-formula><mml:math id="M64" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 4 local element matrix. By obtaining the global node indices <inline-formula><mml:math id="M65" display="inline"><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M66" display="inline"><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> corresponding to the local nodes <inline-formula><mml:math id="M67" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M68" display="inline"><mml:mi>j</mml:mi></mml:math></inline-formula>, the result <inline-formula><mml:math id="M69" display="inline"><mml:mrow><mml:msub><mml:mi>K</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> is updated in the global matrix entryy<inline-formula><mml:math id="M70" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>.</p>
      <p id="d2e2065">The global matrix is typically stored in the CSR format, which consists of three arrays: <italic>values</italic> (non-zero entries), <italic>col_index</italic> (column indices), and <italic>row_pointers</italic> (row offsets) <xref ref-type="bibr" rid="bib1.bibx31" id="paren.27"/>. Within a row range, the target column can be located using binary search.</p>
      <p id="d2e2080">For vector fields, each node contains three degrees of freedom (DOFs). Thus, while assembling the local stiffness matrix, all interactions among the nodes and their DOFs must be considered. This results in a set of 4 <inline-formula><mml:math id="M71" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 4 blocks, each of size 3 <inline-formula><mml:math id="M72" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 3:

            <disp-formula id="Ch1.E2" content-type="numbered"><label>2</label><mml:math id="M73" display="block"><mml:mrow><mml:mi mathvariant="bold">K</mml:mi><mml:mo>=</mml:mo><mml:mfenced open="[" close="]"><mml:mtable class="matrix" columnalign="center center center center" framespacing="0em"><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mi mathvariant="bold">K</mml:mi><mml:mn mathvariant="normal">11</mml:mn></mml:msub></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:msub><mml:mi mathvariant="bold">K</mml:mi><mml:mn mathvariant="normal">12</mml:mn></mml:msub></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:msub><mml:mi mathvariant="bold">K</mml:mi><mml:mn mathvariant="normal">13</mml:mn></mml:msub></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:msub><mml:mi mathvariant="bold">K</mml:mi><mml:mn mathvariant="normal">14</mml:mn></mml:msub></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mi mathvariant="bold">K</mml:mi><mml:mn mathvariant="normal">21</mml:mn></mml:msub></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:msub><mml:mi mathvariant="bold">K</mml:mi><mml:mn mathvariant="normal">22</mml:mn></mml:msub></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:msub><mml:mi mathvariant="bold">K</mml:mi><mml:mn mathvariant="normal">23</mml:mn></mml:msub></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:msub><mml:mi mathvariant="bold">K</mml:mi><mml:mn mathvariant="normal">24</mml:mn></mml:msub></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mi mathvariant="bold">K</mml:mi><mml:mn mathvariant="normal">31</mml:mn></mml:msub></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:msub><mml:mi mathvariant="bold">K</mml:mi><mml:mn mathvariant="normal">32</mml:mn></mml:msub></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:msub><mml:mi mathvariant="bold">K</mml:mi><mml:mn mathvariant="normal">33</mml:mn></mml:msub></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:msub><mml:mi mathvariant="bold">K</mml:mi><mml:mn mathvariant="normal">34</mml:mn></mml:msub></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mi mathvariant="bold">K</mml:mi><mml:mn mathvariant="normal">41</mml:mn></mml:msub></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:msub><mml:mi mathvariant="bold">K</mml:mi><mml:mn mathvariant="normal">42</mml:mn></mml:msub></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:msub><mml:mi mathvariant="bold">K</mml:mi><mml:mn mathvariant="normal">43</mml:mn></mml:msub></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:msub><mml:mi mathvariant="bold">K</mml:mi><mml:mn mathvariant="normal">44</mml:mn></mml:msub></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mfenced><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula></p>
      <p id="d2e2227">Each element of the element matrix <inline-formula><mml:math id="M74" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">K</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> is actually a 3 <inline-formula><mml:math id="M75" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 3 matrix.

            <disp-formula id="Ch1.E3" content-type="numbered"><label>3</label><mml:math id="M76" display="block"><mml:mrow><mml:msub><mml:mi mathvariant="bold">K</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfenced open="[" close="]"><mml:mtable class="matrix" columnalign="center center center" framespacing="0em"><mml:mtr><mml:mtd><mml:mi>a</mml:mi></mml:mtd><mml:mtd><mml:mi>b</mml:mi></mml:mtd><mml:mtd><mml:mi>c</mml:mi></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>d</mml:mi></mml:mtd><mml:mtd><mml:mi>e</mml:mi></mml:mtd><mml:mtd><mml:mi>f</mml:mi></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>g</mml:mi></mml:mtd><mml:mtd><mml:mi>h</mml:mi></mml:mtd><mml:mtd><mml:mi>i</mml:mi></mml:mtd></mml:mtr></mml:mtable></mml:mfenced><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula></p>
      <p id="d2e2300">Here, Fluidity adopts the BCSR format, which is structurally similar to CSR, but stores continuous block data in the value array. This block-based representation better supports sparse operations for vector fields.</p>
</sec>
<sec id="Ch1.S4.SS3">
  <label>4.3</label><title>GPU parallel element-wise assembly</title>
      <p id="d2e2311">In our GPU parallelization, the contributions of each mesh element are computed and stored in parallel, with each GPU thread responsible for one element. A dedicated kernel is launched to assemble element matrices.</p>
      <p id="d2e2314">In the implementation, each thread iterates over the entries <inline-formula><mml:math id="M77" display="inline"><mml:mrow><mml:msub><mml:mi>K</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> of the local element matrix and looks up the corresponding global non-zero indices. Given the global node indices <inline-formula><mml:math id="M78" display="inline"><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M79" display="inline"><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, the <italic>row_pointers</italic> array is first used to locate the row range [row_pointers(<inline-formula><mml:math id="M80" display="inline"><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>), row_pointers(<inline-formula><mml:math id="M81" display="inline"><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula>)). Within this range, the column index <inline-formula><mml:math id="M82" display="inline"><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is searched using binary search. To enhance the efficiency, the implemented GPU assembly function consistently adopts binary search.</p>
      <p id="d2e2394">A significant challenge arises during parallel write-back, where race conditions may occur if multiple threads attempt to update the same non-zero entry simultaneously. An atomic operation is indivisible and guarantees consistency by preventing interference from other threads. Modern GPUs provide hardware-level support for atomic addition, enabling conflict-free concurrent updates without explicit locks. In this work, the atomic add approach is adopted for its efficiency and simplicity. The CUDA parallel assembly code is listed in Appendix B.</p>
      <p id="d2e2397">To summarize the complete execution pipeline, the global matrix assembly occurs entirely within the GPU device memory. Following the element-wise computations, the local element matrices are accumulated into the global sparse matrix arrays (i.e., the <monospace>values</monospace> array in CSR/BCSR format), which are pre-allocated on the GPU. This is achieved using the aforementioned parallel assembly kernel equipped with hardware-level <monospace>atomicAdd</monospace> operations. Upon the completion of the assembly kernel, the fully constructed global matrix arrays are transferred back to the CPU host memory via PCIe. The assembled matrix is then handed over to the CPU-bound PETSc library to perform the subsequent large-scale linear system solve.</p>
</sec>
<sec id="Ch1.S4.SS4">
  <label>4.4</label><title>GPU parallel assembly of Right-Hand-Side (RHS) vectors</title>
      <p id="d2e2415">Similar to the global matrix, the RHS vector is assembled using an element-wise parallel strategy. Each GPU thread computes the nodal contributions of one element and accumulates them into the corresponding global vector entries.</p>
      <p id="d2e2418">The implementation first determines the indices of the nodes belonging to each element, then loads the nodal values, and finally applies <monospace>atomicAdd</monospace> operations to ensure correctness during concurrent accumulation. For cases where multiple variables are associated with a node, appropriate offsets are added to the corresponding positions. The Appendix C illustrates the procedure of RHS assembly, which achieves efficient large-scale parallel accumulation while preserving correctness.</p>
</sec>
</sec>
<sec id="Ch1.S5">
  <label>5</label><title>Results and evaluation</title>
<sec id="Ch1.S5.SS1">
  <label>5.1</label><title>Experimental environment</title>
      <p id="d2e2444">This section presents a systematic analysis of the performance of the proposed GPU parallelization strategies. Experiments were conducted to first validate correctness on both CPU and GPU platforms, and then to evaluate performance in four aspects: element-wise computation, global matrix assembly, data transfer with computation overlap, and overall model performance. The GPU and CPU experimental environments are shown in Table 2. All CPU baseline evaluations were executed using the specified number of CPU cores (MPI processes) on a single physical AMD EPYC 7713 processor.</p>

<table-wrap id="T2"><label>Table 2</label><caption><p id="d2e2451">GPU and CPU and compiler information.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="2">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="left"/>
     <oasis:thead>
       <oasis:row>
         <oasis:entry colname="col1">GPU NVIDIA</oasis:entry>
         <oasis:entry colname="col2">A100</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Peak perf. (FP64)</oasis:entry>
         <oasis:entry colname="col2">9.7 TFLOPS</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Max Clock Freq</oasis:entry>
         <oasis:entry colname="col2">1410 <inline-formula><mml:math id="M83" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">MHz</mml:mi></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">L1 Cache</oasis:entry>
         <oasis:entry colname="col2">192 <inline-formula><mml:math id="M84" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">KiB</mml:mi></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">L2 Cache</oasis:entry>
         <oasis:entry colname="col2">40 960 <inline-formula><mml:math id="M85" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">KiB</mml:mi></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">SM per GPU</oasis:entry>
         <oasis:entry colname="col2">108</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Register file size per SM</oasis:entry>
         <oasis:entry colname="col2">256 <inline-formula><mml:math id="M86" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">KiB</mml:mi></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Device Memory Bandwidth</oasis:entry>
         <oasis:entry colname="col2">1555 <inline-formula><mml:math id="M87" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">GB</mml:mi><mml:mspace linebreak="nobreak" width="0.125em"/><mml:msup><mml:mi mathvariant="normal">s</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">PCIe Bandwidth</oasis:entry>
         <oasis:entry colname="col2">32 <inline-formula><mml:math id="M88" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">GB</mml:mi><mml:mspace width="0.125em" linebreak="nobreak"/><mml:msup><mml:mi mathvariant="normal">s</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">nvcc</oasis:entry>
         <oasis:entry colname="col2">13.0</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">CPU</oasis:entry>
         <oasis:entry colname="col2">AMD EPYC 7713 64-Core</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">Processor</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">ISA</oasis:entry>
         <oasis:entry colname="col2">X86_64</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Max Clock Freq</oasis:entry>
         <oasis:entry colname="col2">3720 <inline-formula><mml:math id="M89" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">Mhz</mml:mi></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">L1 Cache</oasis:entry>
         <oasis:entry colname="col2">8 <inline-formula><mml:math id="M90" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">MiB</mml:mi></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">L2 Cache</oasis:entry>
         <oasis:entry colname="col2">64 <inline-formula><mml:math id="M91" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">MiB</mml:mi></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Cores</oasis:entry>
         <oasis:entry colname="col2">64</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Memory Bandwidth</oasis:entry>
         <oasis:entry colname="col2">204.8 <inline-formula><mml:math id="M92" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">GB</mml:mi><mml:mspace width="0.125em" linebreak="nobreak"/><mml:msup><mml:mi mathvariant="normal">s</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">gcc</oasis:entry>
         <oasis:entry colname="col2">13.3</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">CMake</oasis:entry>
         <oasis:entry colname="col2">3.22.1</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">gfortran</oasis:entry>
         <oasis:entry colname="col2">13.3</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">mpicc</oasis:entry>
         <oasis:entry colname="col2">MPICH 3.4.2</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Optimization</oasis:entry>
         <oasis:entry colname="col2">-O3 -ffast-math</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d2e2775">GPU acceleration was applied to the momentum equation, advection equation, and stabilization term construction of Fluidity-Atmosphere. Then, the GPU-accelerated version was validated for correctness and evaluated for performance to ensure both accuracy and efficiency.</p>
</sec>
<sec id="Ch1.S5.SS2">
  <label>5.2</label><title>Validation methodology</title>
      <p id="d2e2787">For validation of correctness, the 3D idealized mountain wave test case <xref ref-type="bibr" rid="bib1.bibx22" id="paren.28"/> was used. The mountain wave test serves as a crucial validation procedure for assessing the dynamic framework model of the Fluidity-Atmosphere model in simulating orography-induced airflow. By comparing numerical results against theoretical solutions of idealized orographically forced flows, this test effectively evaluates the capability of the model to capture mountain wave generation and propagation phenomena under complex topographic conditions. The computational domain spans 60 <inline-formula><mml:math id="M93" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula> in both horizontal dimensions with a vertical extent of 16 <inline-formula><mml:math id="M94" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula>. Adaptive mesh resolution dynamically ranges from 125 <inline-formula><mml:math id="M95" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">m</mml:mi></mml:mrow></mml:math></inline-formula> to 10 <inline-formula><mml:math id="M96" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula> throughout the simulation domain, with mesh refinement actively guided by variations in velocity magnitude and potential temperature. The mesh consists of 554 394 elements and 103 635 nodes. A 3D bell-shaped mountain profile is mathematically described as follows:

            <disp-formula id="Ch1.E4" content-type="numbered"><label>4</label><mml:math id="M97" display="block"><mml:mrow><mml:mi>h</mml:mi><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msub><mml:mi>h</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow><mml:mrow><mml:msup><mml:mfenced open="(" close=")"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>+</mml:mo><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mrow><mml:msup><mml:mi>x</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mi>y</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:msup><mml:mi>a</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:mfenced><mml:mfrac><mml:mn mathvariant="normal">3</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:mfrac></mml:msup></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:math></disp-formula>

          where <inline-formula><mml:math id="M98" display="inline"><mml:mrow><mml:msub><mml:mi>h</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M99" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 400 <inline-formula><mml:math id="M100" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">m</mml:mi></mml:mrow></mml:math></inline-formula> represents the peak elevation and <inline-formula><mml:math id="M101" display="inline"><mml:mi>a</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M102" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 1000 <inline-formula><mml:math id="M103" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">m</mml:mi></mml:mrow></mml:math></inline-formula> denotes the characteristic half-width parameter. The stratified atmospheric background is characterized by <inline-formula><mml:math id="M104" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M105" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 0.01 <inline-formula><mml:math id="M106" display="inline"><mml:mrow class="unit"><mml:msup><mml:mi mathvariant="normal">s</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>, while the surface potential temperature initializes at <inline-formula><mml:math id="M107" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M108" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 293.15 <inline-formula><mml:math id="M109" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">K</mml:mi></mml:mrow></mml:math></inline-formula>. The incoming flow maintains a constant velocity of <inline-formula><mml:math id="M110" display="inline"><mml:mi mathvariant="bold-italic">u</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M111" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M112" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mn mathvariant="normal">10</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:msup><mml:mo>)</mml:mo><mml:mi>T</mml:mi></mml:msup></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M113" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">m</mml:mi><mml:mspace linebreak="nobreak" width="0.125em"/><mml:msup><mml:mi mathvariant="normal">s</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>. For numerical stability, an absorbing layer is implemented in the upper atmospheric region (10–16 <inline-formula><mml:math id="M114" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula>-altitude), with additional dissipative layers extending for 10 <inline-formula><mml:math id="M115" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula> inward from all lateral boundaries (comprehensive specifications available in Li et al., 2021).</p>
      <p id="d2e3064">In scientific applications demanding high-precision floating-point operations, microscopic rounding differences between CPU and GPU architectures are inevitable. In a fully dynamic simulation, these bit-wise differences can cause the Adaptive Mesh Refinement (AMR) algorithm to split elements differently over time. Comparing results across divergent meshes necessitates spatial interpolation, which introduces algorithmic noise and obscures the underlying arithmetic consistency. To eliminate structural variations and rigorously validate the GPU implementation, we conduct evaluations under two complementary fixed-mesh scenarios.</p>
      <p id="d2e3067">First, to validate the core operators during the non-stationary initial stage, we extract the refined unstructured mesh generated at <inline-formula><mml:math id="M116" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M117" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 3000 <inline-formula><mml:math id="M118" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">s</mml:mi></mml:mrow></mml:math></inline-formula> from an adaptive run and employ it as a fixed, static grid right from the beginning of the simulation (<inline-formula><mml:math id="M119" display="inline"><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula>). Both the CPU and GPU versions simulate the transient phase up to <inline-formula><mml:math id="M120" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M121" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 3000 <inline-formula><mml:math id="M122" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">s</mml:mi></mml:mrow></mml:math></inline-formula> with the AMR module deactivated. While this static mesh is not dynamically optimized for the early stages, it ensures an identical degree-of-freedom layout from <inline-formula><mml:math id="M123" display="inline"><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula> and completely avoids spatial interpolation errors. Second, to verify long-term stability in the steady-state regime, the simulation is run up to <inline-formula><mml:math id="M124" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M125" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 3000 <inline-formula><mml:math id="M126" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">s</mml:mi></mml:mrow></mml:math></inline-formula> under the standard CPU-AMR configuration until the wave profile stabilizes. The mesh is then fixed, and both versions advance the physical time independently from <inline-formula><mml:math id="M127" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M128" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 3000 <inline-formula><mml:math id="M129" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">s</mml:mi></mml:mrow></mml:math></inline-formula> to <inline-formula><mml:math id="M130" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M131" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 5000 <inline-formula><mml:math id="M132" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">s</mml:mi></mml:mrow></mml:math></inline-formula>. The computational results for both scenarios are presented in Fig. 4. Figure 4a illustrates the 3D wireframe and a 2D cross-sectional slice of the mesh, demonstrating the localized high-resolution refinement adapted to the mountain wave. For the transient fixed-mesh run (Fig. 4b and c), the absolute differences between the CPU and GPU fields at <inline-formula><mml:math id="M133" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M134" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 3000 <inline-formula><mml:math id="M135" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">s</mml:mi></mml:mrow></mml:math></inline-formula> peak at 1.44 <inline-formula><mml:math id="M136" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 10<sup>−14</sup> for the vertical velocity (<inline-formula><mml:math id="M138" display="inline"><mml:mi>w</mml:mi></mml:math></inline-formula>) and 4.73 <inline-formula><mml:math id="M139" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 10<sup>−15</sup> for the potential temperature perturbation (<inline-formula><mml:math id="M141" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>′</mml:mo></mml:msup></mml:mrow></mml:math></inline-formula>). For the steady-state continuation run (Fig. 4d and e), the differences remain bounded within a similar magnitude at <inline-formula><mml:math id="M142" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M143" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 5000 <inline-formula><mml:math id="M144" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">s</mml:mi></mml:mrow></mml:math></inline-formula>. These minuscule variations strictly correspond to the double-precision machine epsilon, confirming that the GPU-accelerated operators are fundamentally and mathematically consistent with the original CPU baseline across all simulation phases.</p>

      <fig id="F4" specific-use="star"><label>Figure 4</label><caption><p id="d2e3311">Comparison between the CPU baseline and the GPU-accelerated computation results under fixed-mesh configurations. <bold>(a)</bold> The unstructured mesh layout visualized via a 3D wireframe and a 2D cross-sectional slice (origin: <inline-formula><mml:math id="M145" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mn mathvariant="normal">30</mml:mn><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mn mathvariant="normal">000</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">30</mml:mn><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">000</mml:mn><mml:mo>,</mml:mo><mml:mo>-</mml:mo><mml:mn mathvariant="normal">8000</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, normal: <inline-formula><mml:math id="M146" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>) rendered in ParaView; <bold>(b)</bold> Simulated vertical velocity (<inline-formula><mml:math id="M147" display="inline"><mml:mi>w</mml:mi></mml:math></inline-formula>) and potential temperature perturbation (<inline-formula><mml:math id="M148" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>′</mml:mo></mml:msup></mml:mrow></mml:math></inline-formula>) fields at <inline-formula><mml:math id="M149" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M150" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 3000 <inline-formula><mml:math id="M151" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">s</mml:mi></mml:mrow></mml:math></inline-formula> from the transient validation run (simulated continuously from <inline-formula><mml:math id="M152" display="inline"><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula> using the fixed-mesh); <bold>(c)</bold> Absolute differences between CPU and GPU results at <inline-formula><mml:math id="M153" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M154" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 3000 <inline-formula><mml:math id="M155" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">s</mml:mi></mml:mrow></mml:math></inline-formula> for the transient fixed-mesh run; <bold>(d)</bold> Simulated fields at <inline-formula><mml:math id="M156" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M157" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 5000 <inline-formula><mml:math id="M158" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">s</mml:mi></mml:mrow></mml:math></inline-formula> from the steady-state stability validation run (advanced from <inline-formula><mml:math id="M159" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M160" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 3000 <inline-formula><mml:math id="M161" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">s</mml:mi></mml:mrow></mml:math></inline-formula> on the fixed-mesh); <bold>(e)</bold> Absolute differences between CPU and GPU results at <inline-formula><mml:math id="M162" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M163" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 5000 <inline-formula><mml:math id="M164" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">s</mml:mi></mml:mrow></mml:math></inline-formula> for the steady-state stability run.</p></caption>
          <graphic xlink:href="https://gmd.copernicus.org/articles/19/7525/2026/gmd-19-7525-2026-f04.png"/>

        </fig>

</sec>
<sec id="Ch1.S5.SS3">
  <label>5.3</label><title>Results and discussion of performance optimization</title>
<sec id="Ch1.S5.SS3.SSS1">
  <label>5.3.1</label><title>GPU element-wise computation performance</title>
      <p id="d2e3543">In unstructured meshes, element node values are typically stored non-contiguously in memory. Thus, the performance of element-wise computation is primarily limited by memory access efficiency. While CPUs incur high memory overhead for non-contiguous access, GPUs leverage high bandwidth and massive thread parallelism to alleviate this bottleneck. Table 3 compares GPU and CPU performance in scalar, vector, and tensor data access. Specifically, functions such as <monospace>ele_val_scalar</monospace> serve as data-gathering routines that collect scattered node-based field values from global arrays into contiguous element-local arrays for each element prior to numerical integration. The results demonstrate the GPU's efficiency in accelerating these heavily memory-bound “gather” operations. This is primarily because arrays storing scalar fields have a simpler, lower-dimensional layout, facilitating GPUs to completely leverage high-bandwidth access. All the performance profiling data were collected with the mesh in the Mountainwave3D simulation, in which the number of mesh nodes and elements are 103 635 and 554 394, respectively.</p>

<table-wrap id="T3"><label>Table 3</label><caption><p id="d2e3552">GPU Performance of GPU parallel element data access (Mesh: 103 635 nodes, 554 394 elements; timings are averaged over 1000 invocations).</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="4">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="center"/>
     <oasis:colspec colnum="4" colname="col4" align="center"/>
     <oasis:thead>
       <oasis:row>
         <oasis:entry colname="col1">Function</oasis:entry>
         <oasis:entry colname="col2">CPU time</oasis:entry>
         <oasis:entry colname="col3">GPU time</oasis:entry>
         <oasis:entry colname="col4">Speedup</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">(<inline-formula><mml:math id="M165" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">ms</mml:mi></mml:mrow></mml:math></inline-formula>)</oasis:entry>
         <oasis:entry colname="col3">(<inline-formula><mml:math id="M166" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">ms</mml:mi></mml:mrow></mml:math></inline-formula>)</oasis:entry>
         <oasis:entry colname="col4"/>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">ele_val_scalar</oasis:entry>
         <oasis:entry colname="col2">6.628</oasis:entry>
         <oasis:entry colname="col3">0.0276</oasis:entry>
         <oasis:entry colname="col4">240.14</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">ele_val_vector</oasis:entry>
         <oasis:entry colname="col2">13.633</oasis:entry>
         <oasis:entry colname="col3">0.0768</oasis:entry>
         <oasis:entry colname="col4">177.51</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">ele_val_tensor</oasis:entry>
         <oasis:entry colname="col2">32.345</oasis:entry>
         <oasis:entry colname="col3">0.1976</oasis:entry>
         <oasis:entry colname="col4">163.69</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

<table-wrap id="T4"><label>Table 4</label><caption><p id="d2e3667">GPU kernel performance in element-wise computation (Mesh: 103 635 nodes, 554 394 elements).</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="4">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="center"/>
     <oasis:colspec colnum="4" colname="col4" align="center"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Function</oasis:entry>
         <oasis:entry colname="col2">CPU time  (<inline-formula><mml:math id="M167" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">ms</mml:mi></mml:mrow></mml:math></inline-formula>)</oasis:entry>
         <oasis:entry colname="col3">GPU time  (<inline-formula><mml:math id="M168" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">ms</mml:mi></mml:mrow></mml:math></inline-formula>)</oasis:entry>
         <oasis:entry colname="col4">Speedup</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">transform_to_physical</oasis:entry>
         <oasis:entry colname="col2">525.575</oasis:entry>
         <oasis:entry colname="col3">3.964</oasis:entry>
         <oasis:entry colname="col4">132.59</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">dshape_tensor_dshape</oasis:entry>
         <oasis:entry colname="col2">1865.290</oasis:entry>
         <oasis:entry colname="col3">1.969</oasis:entry>
         <oasis:entry colname="col4">947.33</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">shape_shape</oasis:entry>
         <oasis:entry colname="col2">104.221</oasis:entry>
         <oasis:entry colname="col3">1.012</oasis:entry>
         <oasis:entry colname="col4">103.01</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">shape_dshape</oasis:entry>
         <oasis:entry colname="col2">246.047</oasis:entry>
         <oasis:entry colname="col3">1.371</oasis:entry>
         <oasis:entry colname="col4">179.41</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">get_edge_lengths</oasis:entry>
         <oasis:entry colname="col2">1369.171</oasis:entry>
         <oasis:entry colname="col3">7.476</oasis:entry>
         <oasis:entry colname="col4">183.14</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d2e3797">Furthermore, Table 4 presents the GPU performance of numerous core functions in Fluidity-Atmosphere (coordinate transformation, integration, and auxiliary routines). Results indicate that parallelizing dense element-wise computations on GPUs significantly accelerates the performance. All listed functions achieved more than <italic>103</italic><inline-formula><mml:math id="M169" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> speedups compared with CPU versions. These functions are computationally intensive, involving repeated local linear system and eigenvalue/eigenvector solves for each element in single-core CPU execution. In contrast, the GPU can execute such highly independent tasks concurrently across thousands of threads, comprehensively leveraging parallelism.</p>

<table-wrap id="T5"><label>Table 5</label><caption><p id="d2e3812">Performance of dshape_tensor_dshape with optimizations (Mesh: 554 394 elements).</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="3">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:thead>
       <oasis:row>
         <oasis:entry colname="col1">Version</oasis:entry>
         <oasis:entry colname="col2">dshape1 and dshape2</oasis:entry>
         <oasis:entry colname="col3">dshape1 and dshape2</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">are the same array.</oasis:entry>
         <oasis:entry colname="col3">are different arrays.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">Time  (<inline-formula><mml:math id="M170" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">ms</mml:mi></mml:mrow></mml:math></inline-formula>)</oasis:entry>
         <oasis:entry colname="col3">Time  (<inline-formula><mml:math id="M171" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">ms</mml:mi></mml:mrow></mml:math></inline-formula>)</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">Baseline</oasis:entry>
         <oasis:entry colname="col2">35.799</oasis:entry>
         <oasis:entry colname="col3">36.314</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Unroll</oasis:entry>
         <oasis:entry colname="col2">20.847</oasis:entry>
         <oasis:entry colname="col3">26.890</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Templated matrix product</oasis:entry>
         <oasis:entry colname="col2">1.969</oasis:entry>
         <oasis:entry colname="col3">3.689</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">and dot product</oasis:entry>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3"/>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d2e3928">To analyse the effects of GPU-oriented optimizations, we take the <monospace>dshape_tensor_dshape</monospace> function as an example. This function is a hotspot in finite-element computation, involving frequent access to gradients of shape functions at Gauss points and dense matrix multiplications/dot products. Table 5 shows performance enhancement due to loop unrolling and optimization via the use of template functions. Notably, Fluidity-Atmosphere typically calls this routine with identical <monospace>dshape1</monospace> and <monospace>dshape2</monospace> arrays. However, for generality we also tested cases with two different <monospace>dshape1</monospace> and <monospace>dshape2</monospace> arrays. While the same array case performs better, both scenarios show consistent performance gains.</p>
      <p id="d2e3946">Nsight Compute profiling further indicates that templated matrix product and dot product increased memory throughput from 70 % to 82.35 %, slightly improved compute throughput, and enhanced L1 cache hit rates. SASS code inspection revealed that baseline <monospace>__device__</monospace> functions were only locally reordered, thereby limiting the performance. After templating, the compiler applied deeper optimizations, interleaving SASS instructions between <monospace>__device__</monospace> and <monospace>__global__</monospace> functions, yielding significant performance enhancements.</p>
      <p id="d2e3958">These results confirm that GPU parallelization achieves order-of-magnitude acceleration in element-wise computations, while targeted optimizations further unlock the hardware potential.</p>
</sec>
<sec id="Ch1.S5.SS3.SSS2">
  <label>5.3.2</label><title>Performance of global matrix assembly optimization</title>
      <p id="d2e3969">Table 6 shows assembly performance for the advection-diffusion matrix (Temperature) and the right-hand sides. Compared with CPU execution, GPU parallel assembly achieved speedups of up to 389.02<inline-formula><mml:math id="M172" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> and 415.64<inline-formula><mml:math id="M173" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula>, respectively, highlighting the advantages of massive threading and ultra-high device memory bandwidth. Compared with existing GPU studies on sparse matrix assembly, a substantially higher acceleration is achieved in the proposed method, demonstrating the superior performance of our kernel designs and data structures.</p>

<table-wrap id="T6"><label>Table 6</label><caption><p id="d2e3989">Global matrix assembly performance (Mesh: 103 635 nodes, 554 394 elements).</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="4">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="center"/>
     <oasis:colspec colnum="4" colname="col4" align="center"/>
     <oasis:thead>
       <oasis:row>
         <oasis:entry colname="col1">Task</oasis:entry>
         <oasis:entry colname="col2">CPU</oasis:entry>
         <oasis:entry colname="col3">GPU</oasis:entry>
         <oasis:entry colname="col4">Speedup</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">assembly</oasis:entry>
         <oasis:entry colname="col3">assembly</oasis:entry>
         <oasis:entry colname="col4"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">time</oasis:entry>
         <oasis:entry colname="col3">time</oasis:entry>
         <oasis:entry colname="col4"/>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">(<inline-formula><mml:math id="M174" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">ms</mml:mi></mml:mrow></mml:math></inline-formula>)</oasis:entry>
         <oasis:entry colname="col3">(<inline-formula><mml:math id="M175" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">ms</mml:mi></mml:mrow></mml:math></inline-formula>)</oasis:entry>
         <oasis:entry colname="col4"/>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">Matrix</oasis:entry>
         <oasis:entry colname="col2">327.538</oasis:entry>
         <oasis:entry colname="col3">0.842</oasis:entry>
         <oasis:entry colname="col4">389.02</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">RHS</oasis:entry>
         <oasis:entry colname="col2">48.254</oasis:entry>
         <oasis:entry colname="col3">0.116</oasis:entry>
         <oasis:entry colname="col4">415.64</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d2e4112">Regarding the parallel assembly, it is worth noting that multiple elements inevitably share nodes in unstructured meshes, which can theoretically introduce memory contention when using atomic operations. To address this, we evaluated alternative algorithms that avoid atomic operations. While this contention-free approach further reduced the sparse matrix update time (e.g., from approximately 5 to 3 <inline-formula><mml:math id="M176" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">ms</mml:mi></mml:mrow></mml:math></inline-formula> in specific test stages), it required an explicit mesh preprocessing step. This preprocessing consumed several seconds, introducing an overhead that far outweighed the modest kernel-level savings. Therefore, the direct atomic operation algorithm was selected for its simplicity and overall efficiency. This design choice is further supported by recent evaluations of unstructured-grid CFD applications on GPUs, which demonstrate that NVIDIA V100 and A100 architectures inherently deliver exceptional performance on kernels dominated by double-precision atomic updates, often outperforming complex register-based aggregation or data restructuring methods <xref ref-type="bibr" rid="bib1.bibx41" id="paren.29"/>.</p>
</sec>
<sec id="Ch1.S5.SS3.SSS3">
  <label>5.3.3</label><title>Hardware utilization and bottleneck analysis</title>
      <p id="d2e4134">To further comprehend the GPU execution efficiency, we analyzed the hardware utilization of our core kernels using NVIDIA Nsight Compute. The theoretical FP64 peak performance of the A100 GPU is 9.7 <inline-formula><mml:math id="M177" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">TFLOP</mml:mi><mml:mspace linebreak="nobreak" width="0.125em"/><mml:msup><mml:mi mathvariant="normal">s</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>. However, in unstructured FEM frameworks, performance is predominantly bounded by memory bandwidth rather than compute capabilities.</p>
      <p id="d2e4154">For the element-wise integration kernel <monospace>dshape_</monospace><monospace>tensor_dshape</monospace>, the Compute (SM) Throughput reached approximately 8.50 % of the theoretical peak. While this kernel involves intensive small-matrix multiplications, its arithmetic intensity is inherently constrained by the necessity to fetch scattered nodal variables and coordinates from the global memory for each element. This is strongly evidenced by its Memory Throughput, which achieved an outstanding 90.44 % of the device's peak memory bandwidth, indicating that the kernel is optimally saturating the hardware's memory subsystem.</p>
      <p id="d2e4163">Similarly, the global matrix assembly <monospace>assemble_</monospace><monospace>csr_matrix</monospace> and RHS construction <monospace>rhs_addto_</monospace><monospace>kernel</monospace> kernels are archetypal memory-bound operations. Their Compute Throughputs are limited to 7.93 % and 18.29 %, respectively, primarily because they execute sparse floating-point operations (such as atomic additions) amidst heavy indirect memory accesses. Instead, their execution efficiency is profoundly reflected in their Memory Throughputs, which achieved 79.48 % and 82.98 %, respectively. Given the highly irregular access patterns dictated by unstructured mesh connectivity (e.g., indirect addressing via <monospace>row_pointers</monospace> and <monospace>ndglno</monospace>), maintaining such high memory bandwidth utilization demonstrates that our data layout and atomic-based assembly strategies effectively saturate the hardware limits for these memory-intensive workloads.</p>

<table-wrap id="T7" specific-use="star"><label>Table 7</label><caption><p id="d2e4189">Pinned vs non-pinned memory transfer performance.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="6">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="center"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:colspec colnum="5" colname="col5" align="center"/>
     <oasis:colspec colnum="6" colname="col6" align="center"/>
     <oasis:thead>
       <oasis:row>
         <oasis:entry colname="col1">Data</oasis:entry>
         <oasis:entry colname="col2">Size (Byte)</oasis:entry>
         <oasis:entry colname="col3">Non-pinned</oasis:entry>
         <oasis:entry colname="col4">Non-pinned</oasis:entry>
         <oasis:entry colname="col5">Pinned</oasis:entry>
         <oasis:entry colname="col6">Pinned</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3">time (<inline-formula><mml:math id="M178" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">ms</mml:mi></mml:mrow></mml:math></inline-formula>)</oasis:entry>
         <oasis:entry colname="col4">Bandwidth</oasis:entry>
         <oasis:entry colname="col5">time</oasis:entry>
         <oasis:entry colname="col6">BW (<inline-formula><mml:math id="M179" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">GB</mml:mi><mml:mspace linebreak="nobreak" width="0.125em"/><mml:msup><mml:mi mathvariant="normal">s</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>)</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3"/>
         <oasis:entry colname="col4">(<inline-formula><mml:math id="M180" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">GB</mml:mi><mml:mspace linebreak="nobreak" width="0.125em"/><mml:msup><mml:mi mathvariant="normal">s</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>)</oasis:entry>
         <oasis:entry colname="col5">(<inline-formula><mml:math id="M181" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">ms</mml:mi></mml:mrow></mml:math></inline-formula>)</oasis:entry>
         <oasis:entry colname="col6"/>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">ndglno array</oasis:entry>
         <oasis:entry colname="col2">8 870 304</oasis:entry>
         <oasis:entry colname="col3">0.686</oasis:entry>
         <oasis:entry colname="col4">12.94</oasis:entry>
         <oasis:entry colname="col5">0.348</oasis:entry>
         <oasis:entry colname="col6">25.47</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">scalar field array</oasis:entry>
         <oasis:entry colname="col2">829 080</oasis:entry>
         <oasis:entry colname="col3">0.086</oasis:entry>
         <oasis:entry colname="col4">9.67</oasis:entry>
         <oasis:entry colname="col5">0.047</oasis:entry>
         <oasis:entry colname="col6">17.60</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">vector field array</oasis:entry>
         <oasis:entry colname="col2">2 487 240</oasis:entry>
         <oasis:entry colname="col3">0.204</oasis:entry>
         <oasis:entry colname="col4">12.18</oasis:entry>
         <oasis:entry colname="col5">0.105</oasis:entry>
         <oasis:entry colname="col6">23.60</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">tensor field array</oasis:entry>
         <oasis:entry colname="col2">7461 720</oasis:entry>
         <oasis:entry colname="col3">0.586</oasis:entry>
         <oasis:entry colname="col4">12.73</oasis:entry>
         <oasis:entry colname="col5">0.332</oasis:entry>
         <oasis:entry colname="col6">22.44</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

</sec>
<sec id="Ch1.S5.SS3.SSS4">
  <label>5.3.4</label><title>CPU-GPU data transfer and overlap with computation</title>
      <p id="d2e4422">In hotspot acceleration, CPU-GPU data transfer is another critical performance factor. In our system, CPU and GPU are connected via PCIe 4.0 <inline-formula><mml:math id="M182" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 16, with a bidirectional bandwidth of 32 <inline-formula><mml:math id="M183" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">GB</mml:mi><mml:mspace width="0.125em" linebreak="nobreak"/><mml:msup><mml:mi mathvariant="normal">s</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>. When the CPU has pinned memory, transfer bandwidth utilization is significantly higher than that with pageable memory. During Fluidity-Atmosphere computation, the type of data transferred frequently between the CPU and GPU includes scalar, vector, and tensor field arrays, the element-node connectivity array (<italic>ndglno</italic>), as well as the fully assembled global sparse matrix arrays returning from the GPU device memory to the host for PETSc solvers. Table 7 compares pinned and non-pinned transfer performance, displaying an evidently superior bandwidth with pinned memory.</p>
      <p id="d2e4452">In addition, CUDA provides stream concurrency and asynchronous APIs such as cudaMemcpyAsync, facilitating simultaneous data transfer and kernel execution. The <monospace>transform_to_physical</monospace> kernels can run concurrently with transfers of subsequent physical field data. The transfer-computation overlap effectively mitigates communication bottlenecks and further leverages HPC system performance.</p>

<table-wrap id="T8" specific-use="star"><label>Table 8</label><caption><p id="d2e4461">GPU acceleration performance of major modules.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="4">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="center"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Functional module</oasis:entry>
         <oasis:entry colname="col2">CPU time  (<inline-formula><mml:math id="M184" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">ms</mml:mi></mml:mrow></mml:math></inline-formula>)</oasis:entry>
         <oasis:entry colname="col3">GPU time  (<inline-formula><mml:math id="M185" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">ms</mml:mi></mml:mrow></mml:math></inline-formula>)</oasis:entry>
         <oasis:entry colname="col4">Speedup</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">the advection-diffusion matrix (Temperature)</oasis:entry>
         <oasis:entry colname="col2">2563.158</oasis:entry>
         <oasis:entry colname="col3">11.566</oasis:entry>
         <oasis:entry colname="col4">221.60</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">the advection-diffusion matrix (WaterVapor)</oasis:entry>
         <oasis:entry colname="col2">2542.717</oasis:entry>
         <oasis:entry colname="col3">11.251</oasis:entry>
         <oasis:entry colname="col4">226.00</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">the advection–diffusion matrix (CloudWater)</oasis:entry>
         <oasis:entry colname="col2">2566.037</oasis:entry>
         <oasis:entry colname="col3">11.296</oasis:entry>
         <oasis:entry colname="col4">227.17</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">the advection-diffusion matrix (RainWater)</oasis:entry>
         <oasis:entry colname="col2">1877.409</oasis:entry>
         <oasis:entry colname="col3">9.481</oasis:entry>
         <oasis:entry colname="col4">198.02</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">the momentum matrix</oasis:entry>
         <oasis:entry colname="col2">12 641.197</oasis:entry>
         <oasis:entry colname="col3">46.363</oasis:entry>
         <oasis:entry colname="col4">272.66</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">the divergence matrix</oasis:entry>
         <oasis:entry colname="col2">3657.007</oasis:entry>
         <oasis:entry colname="col3">21.497</oasis:entry>
         <oasis:entry colname="col4">170.12</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">the pressure diffusion matrix</oasis:entry>
         <oasis:entry colname="col2">3505.105</oasis:entry>
         <oasis:entry colname="col3">28.507</oasis:entry>
         <oasis:entry colname="col4">122.96</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">the projection matrix</oasis:entry>
         <oasis:entry colname="col2">1126.354</oasis:entry>
         <oasis:entry colname="col3">10.944</oasis:entry>
         <oasis:entry colname="col4">102.92</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>


</sec>
<sec id="Ch1.S5.SS3.SSS5">
  <label>5.3.5</label><title>Performance of functional modules</title>
      <p id="d2e4647">After embedding GPU-accelerated element-wise computation and global assembly strategies into Fluidity-Atmosphere, we evaluated the overall performance at three levels: module, single timestep, and multi-process parallelism. In the restarting Mountainwave3D simulation, the pressure diffusion matrix was calculated only once, and the others were the average times of simulation timesteps which ranges from 3000 to 5000 <inline-formula><mml:math id="M186" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">s</mml:mi></mml:mrow></mml:math></inline-formula>. Table 8 summarizes performance across major functional modules.</p>
      <p id="d2e4658">As the global sparse linear solver was not GPU-parallelized, the overall acceleration for the single-process execution remained limited. However, Fluidity-Atmosphere already supports MPI, enabling MPI+GPU hybrid execution for maximal performance. Table 9 compares the one timestep execution performance using different numbers of CPU cores (MPI processes) and GPU-enabled configurations.</p>

<table-wrap id="T9"><label>Table 9</label><caption><p id="d2e4664">The performance of GPU accelerated Fluidity-Atmosphere.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="3">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="center"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Version</oasis:entry>
         <oasis:entry colname="col2">Time (<inline-formula><mml:math id="M187" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">s</mml:mi></mml:mrow></mml:math></inline-formula>)</oasis:entry>
         <oasis:entry colname="col3">Speedup (vs 1 CPU)</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">1 CPU core (Baseline)</oasis:entry>
         <oasis:entry colname="col2">41.838</oasis:entry>
         <oasis:entry colname="col3">1.00</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">1 CPU core <inline-formula><mml:math id="M188" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula> 1 GPU</oasis:entry>
         <oasis:entry colname="col2">17.147</oasis:entry>
         <oasis:entry colname="col3">2.44</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">4 CPU cores</oasis:entry>
         <oasis:entry colname="col2">13.376</oasis:entry>
         <oasis:entry colname="col3">3.13</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">4 CPU cores <inline-formula><mml:math id="M189" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula> 1 GPU</oasis:entry>
         <oasis:entry colname="col2">4.880</oasis:entry>
         <oasis:entry colname="col3">8.57</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d2e4767">The results show that GPU acceleration achieved a <italic>2.44</italic><inline-formula><mml:math id="M190" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> <italic>speedup</italic> for single-process execution. With 4 MPI processes (each handling approximately 25 000 nodes and 140 000 elements), the hybrid MPI <inline-formula><mml:math id="M191" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula> GPU version achieved <italic>8.57</italic><inline-formula><mml:math id="M192" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> <italic>speedup</italic> compared with single CPU core execution. Table 9 presents the average times of simulation timesteps which ranges from 3000 to 5000 <inline-formula><mml:math id="M193" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">s</mml:mi></mml:mrow></mml:math></inline-formula>.</p>

<table-wrap id="T10" specific-use="star"><label>Table 10</label><caption><p id="d2e4813">Comparison of computational performance and runtime distribution across different configurations (Unit: Time in seconds, Percentage of total timestep).</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="7">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right" colsep="1"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:colspec colnum="5" colname="col5" align="right" colsep="1"/>
     <oasis:colspec colnum="6" colname="col6" align="right"/>
     <oasis:colspec colnum="7" colname="col7" align="right"/>
     <oasis:thead>
       <oasis:row>
         <oasis:entry colname="col1">Functional module</oasis:entry>
         <oasis:entry rowsep="1" namest="col2" nameend="col3" align="center" colsep="1">Baseline (1 CPU core) </oasis:entry>
         <oasis:entry rowsep="1" namest="col4" nameend="col5" align="center" colsep="1">1 CPU core <inline-formula><mml:math id="M194" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula> 1 GPU </oasis:entry>
         <oasis:entry rowsep="1" namest="col6" nameend="col7" align="center">4 CPU cores <inline-formula><mml:math id="M195" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula> 1 GPU </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">Time (<inline-formula><mml:math id="M196" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">s</mml:mi></mml:mrow></mml:math></inline-formula>)</oasis:entry>
         <oasis:entry colname="col3">Percentage (%)</oasis:entry>
         <oasis:entry colname="col4">Time (<inline-formula><mml:math id="M197" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">s</mml:mi></mml:mrow></mml:math></inline-formula>)</oasis:entry>
         <oasis:entry colname="col5">Percentage (%)</oasis:entry>
         <oasis:entry colname="col6">Time (<inline-formula><mml:math id="M198" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">s</mml:mi></mml:mrow></mml:math></inline-formula>)</oasis:entry>
         <oasis:entry colname="col7">Percentage (%)</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">Temperature Matrix</oasis:entry>
         <oasis:entry colname="col2">2.467</oasis:entry>
         <oasis:entry colname="col3">5.95</oasis:entry>
         <oasis:entry colname="col4">0.0226</oasis:entry>
         <oasis:entry colname="col5">0.13</oasis:entry>
         <oasis:entry colname="col6">0.0120</oasis:entry>
         <oasis:entry colname="col7">0.25</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">WaterVapor Matrix</oasis:entry>
         <oasis:entry colname="col2">2.471</oasis:entry>
         <oasis:entry colname="col3">5.96</oasis:entry>
         <oasis:entry colname="col4">0.0137</oasis:entry>
         <oasis:entry colname="col5">0.08</oasis:entry>
         <oasis:entry colname="col6">0.0041</oasis:entry>
         <oasis:entry colname="col7">0.08</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">CloudWater Matrix</oasis:entry>
         <oasis:entry colname="col2">2.494</oasis:entry>
         <oasis:entry colname="col3">6.02</oasis:entry>
         <oasis:entry colname="col4">0.0139</oasis:entry>
         <oasis:entry colname="col5">0.08</oasis:entry>
         <oasis:entry colname="col6">0.0038</oasis:entry>
         <oasis:entry colname="col7">0.08</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">RainWater Matrix</oasis:entry>
         <oasis:entry colname="col2">1.799</oasis:entry>
         <oasis:entry colname="col3">4.34</oasis:entry>
         <oasis:entry colname="col4">0.0104</oasis:entry>
         <oasis:entry colname="col5">0.06</oasis:entry>
         <oasis:entry colname="col6">0.0029</oasis:entry>
         <oasis:entry colname="col7">0.06</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Momentum Matrix</oasis:entry>
         <oasis:entry colname="col2">12.093</oasis:entry>
         <oasis:entry colname="col3">29.19</oasis:entry>
         <oasis:entry colname="col4">0.0313</oasis:entry>
         <oasis:entry colname="col5">0.18</oasis:entry>
         <oasis:entry colname="col6">0.0384</oasis:entry>
         <oasis:entry colname="col7">0.79</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Pressure Correction Matrix</oasis:entry>
         <oasis:entry colname="col2">4.760</oasis:entry>
         <oasis:entry colname="col3">11.49</oasis:entry>
         <oasis:entry colname="col4">0.0165</oasis:entry>
         <oasis:entry colname="col5">0.10</oasis:entry>
         <oasis:entry colname="col6">0.0092</oasis:entry>
         <oasis:entry colname="col7">0.19</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Total PETSc (Setup <inline-formula><mml:math id="M199" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula> Solver)</oasis:entry>
         <oasis:entry colname="col2">5.212</oasis:entry>
         <oasis:entry colname="col3">12.58</oasis:entry>
         <oasis:entry colname="col4">5.6045</oasis:entry>
         <oasis:entry colname="col5">32.68</oasis:entry>
         <oasis:entry colname="col6">2.1921</oasis:entry>
         <oasis:entry colname="col7">44.92</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Other unaccelerated modules</oasis:entry>
         <oasis:entry colname="col2">10.136</oasis:entry>
         <oasis:entry colname="col3">24.46</oasis:entry>
         <oasis:entry colname="col4">11.4341</oasis:entry>
         <oasis:entry colname="col5">66.68</oasis:entry>
         <oasis:entry colname="col6">2.6175</oasis:entry>
         <oasis:entry colname="col7">53.63</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Total timestep duration</oasis:entry>
         <oasis:entry colname="col2">41.433</oasis:entry>
         <oasis:entry colname="col3">100.00</oasis:entry>
         <oasis:entry colname="col4">17.147</oasis:entry>
         <oasis:entry colname="col5">100.00</oasis:entry>
         <oasis:entry colname="col6">4.880</oasis:entry>
         <oasis:entry colname="col7">100.00</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d2e5152">Notably, the 4 CPU cores <inline-formula><mml:math id="M200" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula> 1 GPU configuration achieves an 8.57<inline-formula><mml:math id="M201" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> speedup, which is more than triple the speedup of the 1 CPU core <inline-formula><mml:math id="M202" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula> 1 GPU configuration (2.44<inline-formula><mml:math id="M203" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula>). This synergistic effect is attributed to two factors: first, the domain decomposition in multi-process execution allows the remaining CPU-bound linear solvers (PETSc) to benefit from improved cache locality on smaller sub-domains; second, multiple MPI processes can concurrently issue kernels to the GPU via CUDA streams, leading to higher hardware occupancy on the NVIDIA A100. As detailed in Table 10, this configuration effectively addresses the “bottleneck shift” dictated by Amdahl's Law. In the 1 CPU core <inline-formula><mml:math id="M204" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula> 1 GPU setup, GPU matrix assembly is compressed to less than 1 % of the total time, leaving the PETSc solver and other unaccelerated modules as the dominant bottlenecks. However, by leveraging 4 MPI processes, the absolute execution time of the PETSc solver is dramatically reduced from 5.60 to 2.19 <inline-formula><mml:math id="M205" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">s</mml:mi></mml:mrow></mml:math></inline-formula>, and the time for other unaccelerated modules drops from 11.43 to 2.62 <inline-formula><mml:math id="M206" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">s</mml:mi></mml:mrow></mml:math></inline-formula>. This multi-core mitigation of the remaining CPU bottlenecks, combined with sustained GPU efficiency, strictly validates the 8.57<inline-formula><mml:math id="M207" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> overall performance gain.</p>
</sec>
</sec>
</sec>
<sec id="Ch1.S6" sec-type="conclusions">
  <label>6</label><title>Conclusions</title>
      <p id="d2e5224">This study focuses on the unstructured-mesh finite-element atmospheric model named Fluidity-Atmosphere, thereby addressing the computational bottlenecks in element-wise computations and matrix assembly. Leveraging the architectural features of the NVIDIA A100 GPU and the CUDA programming model, we designed and implemented GPU-oriented parallel optimization strategies. The proposed high-performance GPU kernels substantially enhanced the computational efficiency while solving the momentum equations, advection equations, and stabilization term construction, thereby accelerating critical processes in atmospheric numerical simulations. For element-wise computations, template-based kernel design and deep optimizations achieved speedups ranging from tens to several hundred times. With asynchronous data transfer and MPI parallelization, an overall acceleration that is 8.57 times greater than that of the CPU version was achieved using four processes while ensuring correctness and stability in multi-time step simulations. The proposed parallelization methods substantially enhance the performance of atmospheric models on heterogeneous systems, remarkably extending the support for high-performance implementations of unstructured-mesh models. Future work will focus on further overcoming performance bottlenecks in the overall simulation, particularly by migrating the sparse linear solver (which is still CPU-dependent) to GPUs, thereby facilitating end-to-end acceleration.</p>
</sec>

      
      </body>
    <back><app-group>

<app id="App1.Ch1.S1">
  <label>Appendix A</label><title>Code</title>
      <p id="d2e5239"><preformat><![CDATA[template<int T_M>
__device__ double dot_product_t(double*__restrict__ a,
double* __restrict__ b){
    double r=0.0;
    #pragma unroll
    for(int i=0;i<T_M;i++){
        r += a[i]*b[i];
    }
    return r;
}

// T_M, T_N, T_K: Matrix dimensions for generic small matrix multiplication
template<int T_M, int T_N, int T_K>
__device__ void __forceinline__  matmul_t(const double* a, const double* b, double* C){
    for (int mi = 0; mi < T_M; ++mi) {
        for (int ki = 0; ki < T_K; ++ki) {
            #pragma unroll
            for (int ni = 0; ni < T_N; ++ni) {
               // Column-major layout stride; efficient due to thread-local
               // register/L1 caching
               C[ni* T_M + mi] +=  a[ki * T_M + mi] * b[T_K * ni + ki];
            }
        }
    }
}

// T_dim: Spatial dimension (e.g., 3 for 3D)
// T_loc1, T_loc2: Number of local nodes per element
// T_ngi: Number of Gauss integration points
template<int T_dim, int T_loc1, int T_loc2, int T_ngi>
__device__ void dshape_tensor_dshape_unroll_t(const double *dshape1,
const double *tensor,const double *dshape2, const double *detwei, double *r){
    double tmp[T_dim]={0.0},mat_mul[T_dim]={0.0},dotproduct=0.0;
    for(int gi=0; gi< T_ngi;++gi){
        double gidetwei=detwei[gi];
        for(int iloc=0;iloc<T_loc1;iloc++){
            #pragma unroll
            for(int n=0;n<T_dim;++n){
                tmp[n]=dshape1[n*T_ngi*T_loc1+gi*T_loc1+iloc];
                mat_mul[n] = 0.0;
            }
            // <1,3,3> literals explicitly used to unroll 3D spatial
            // vector-tensor interaction
            matmul_t<1,3,3>(tmp,&tensor[gi*T_dim*T_dim],mat_mul);
            for(int jloc=0;jloc<T_loc2;jloc++){
                #pragma unroll
                for(int n=0;n<T_dim;++n){
                    tmp[n]=dshape2[n*T_ngi*T_loc2+gi*T_loc2+jloc];
                }
                dotproduct = dot_product_t<T_dim>(mat_mul,tmp);
                r[jloc*T_loc1+iloc]=r[jloc*T_loc1+iloc]+dotproduct*gidetwei;
            }
        }
    }
}]]></preformat></p>
</app>

<app id="App1.Ch1.S2">
  <label>Appendix B</label><title>Code</title>
      <p id="d2e5254"><preformat><![CDATA[template<int T_dim, int T_loc, int T_ngi>
__global__ void assemble_csr_matrix(double *result, double* values,
int *row_pointers, int *col_index, int *ndglno, const int ele_num){
    int element_index = blockIdx.x*blockDim.x + threadIdx.x;
    if(element_index>=ele_num) return;
    int loc1=T_loc;
    int loc2=T_loc;
    int mpos=0,base=0,upper_j=0,upper_pos=0,lower_j=0,lower_pos=0,
        this_pos=0,this_j=0;
    int inode[T_loc]={0};
    int j=0;
    double *element_result=result + element_index *T_loc * T_loc;
    int *row=nullptr;
    int iloc=0, jloc=0,  n=0;
    //ndglno stores the connectivity for neighbouring elements.
    for(int i = 0; i < T_loc; ++i){
      inode[i]= ndglno[element_index * loc1 + i];
    }
    for(iloc=0;iloc<loc1;iloc++){
        //the index in Fluidity-Atmosphere Fortran code is from 1.
        base = row_pointers[inode[iloc]-1]-1;
        n = row_pointers[inode[iloc]] -1 - base;
        for(jloc=0;jloc<loc2;jloc++){
            if(element_result[jloc*loc1+iloc]==0)continue;
            row = col_index+base;
            upper_pos=n-1;
            upper_j=row[n-1]-1;
            lower_pos=0;
            lower_j=row[0]-1;
            mpos=-2;
            j = inode[jloc]-1;
            if (upper_j<j){
                mpos=-1;
            }else if (upper_j==j){
                mpos=upper_pos+base;
            }else if (lower_j>j){
                mpos=-1;
            }else if(lower_j==j){
                mpos=lower_pos+base;
            }

            while(((upper_pos-lower_pos)>1)&&(mpos==-2)){
                this_pos=(upper_pos+lower_pos)/2;
                this_j=row[this_pos]-1;
                if(this_j == (inode[jloc]-1)){
                    mpos=this_pos+base;
                }
                else if(this_j > (inode[jloc]-1)){
                    upper_j=this_j;
                    upper_pos=this_pos;
                }
                else{
                    lower_j=this_j;
                    lower_pos=this_pos;
                }
            }
            if(mpos<0){
            }else{
                atomicAdd(values+mpos,element_result[jloc*loc1+iloc]);
            }
        }
    }
}]]></preformat></p>
</app>

<app id="App1.Ch1.S3">
  <label>Appendix C</label><title>Code</title>
      <p id="d2e5269"><preformat><![CDATA[template<int T_loc>
__global__ void rhs_addto_kernel(GpuScalarField *field, const int ele_num,
double *gpu_ele_val){
    int  element_index = blockDim.x * blockIdx.x + threadIdx.x;
    if(element_index >=ele_num) return;
    //element_index will be used to find the node indexes starting from 1.
    element_index = element_index + 1;
    int nodes[T_loc];
    //get the node indexes of current element.
    gpu_ele_nodes_scalar(*field, element_index,nodes);
    double *ptr = gpu_ele_val + (element_index -1)*T_loc;
    #pragma unroll
    for(int i = 0; i < T_loc;++i){
        atomicAdd(&(*field).val[(nodes[i] -1)],*(ptr+i));
    }
}]]></preformat></p>
</app>
  </app-group><notes notes-type="codedataavailability"><title>Code and data availability</title>

      <p id="d2e5277">The original Fluidity model is distributed free of charge under the GNU Lesser General Public License (LGPL). The source code is publicly available from its official repository at <uri>https://fluidityproject.github.io/get-fluidity.html</uri> <xref ref-type="bibr" rid="bib1.bibx1" id="paren.30"/>. The version used in this study corresponds to Fluidity release 2025.12.</p>

      <p id="d2e5286">The GPU-accelerated extension developed in this work is permanently archived at Zenodo. The version used to generate all results presented in this paper corresponds to release v2 and is available at: <ext-link xlink:href="https://doi.org/10.5281/zenodo.18799735" ext-link-type="DOI">10.5281/zenodo.18799735</ext-link> <xref ref-type="bibr" rid="bib1.bibx9" id="paren.31"/>. All simulation data produced in this study are publicly available at <ext-link xlink:href="https://doi.org/10.5281/zenodo.17824052" ext-link-type="DOI">10.5281/zenodo.17824052</ext-link> <xref ref-type="bibr" rid="bib1.bibx23" id="paren.32"/>.</p>

      <p id="d2e5301">The Zenodo archive includes: (1) the complete GPU-modified source files, (2) CUDA kernels and GPU interface implementation, (3) compilation scripts, (4) a detailed README file providing step-by-step instructions for reproducing the numerical experiments and figures reported in this paper.</p>
  </notes><notes notes-type="authorcontribution"><title>Author contributions</title>

      <p id="d2e5307">Conceptualization and Methodology: LL, LJ and LH. provided technical guidance and supervision. Software and Investigation: LL, ZX and FX developed the software and performed the testing. Writing – Original Draft: FX prepared the manuscript. Writing – Review and Editing: LL and LJ revised the manuscript. All co-authors reviewed and approved the final manuscript.</p>
  </notes><notes notes-type="competinginterests"><title>Competing interests</title>

      <p id="d2e5313">The contact author has declared that none of the authors has any competing interests.</p>
  </notes><notes notes-type="disclaimer"><title>Disclaimer</title>

      <p id="d2e5319">Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. The authors bear the ultimate responsibility for providing appropriate place names. Views expressed in the text are those of the authors and do not necessarily reflect the views of the publisher.</p>
  </notes><ack><title>Acknowledgements</title><p id="d2e5329">This work is jointly supported by the State Key Laboratory of Atmospheric Environment and Extreme Meteorology (2024ZD04). Jinxi Li thanks for the technical support of the National Large Scientific and Technological Infrastructure “Earth System Numerical Simulation Facility” (Grant 2025-EL-PT-000886, <uri>https://cstr.cn/31134.02.EL</uri>, last access: 10 August 2026).</p></ack><notes notes-type="financialsupport"><title>Financial support</title>

      <p id="d2e5337">This research has been supported by the National Key Research and Development Program of China (grant no. 2023YFC3705701).</p>
  </notes><notes notes-type="reviewstatement"><title>Review statement</title>

      <p id="d2e5344">This paper was edited by Lele Shu and reviewed by three anonymous referees.</p>
  </notes><ref-list>
    <title>References</title>

      <ref id="bib1.bibx1"><label>AMCG(2014)</label><mixed-citation>AMCG: Fluidity manual, GitHub [code], <uri>https://fluidityproject.github.io/get-fluidity.html</uri> (last access: 10 August 2026), 2014.</mixed-citation></ref>
      <ref id="bib1.bibx2"><label>AMD(2025)</label><mixed-citation>AMD: CUDA to HIP API Function Comparison, <uri>https://rocm.docs.amd.com/projects/HIP/en/latest/reference/api_syntax.html</uri> (last access: 10 August 2026), 2025.</mixed-citation></ref>
      <ref id="bib1.bibx3"><label>Bogenschutz et al.(2025)Bogenschutz, Clevenger, Bradley, Caldwell, Beydoun, Mahfouz, Keen, Guba, Bertagna, Foucar, Zhang, and Donahue</label><mixed-citation>Bogenschutz, P. A., Clevenger, T. C., Bradley, A. M., Caldwell, P. M., Beydoun, H., Mahfouz, N., Keen, N. D., Guba, O., Bertagna, L., Foucar, J., Zhang, J., and Donahue, A. S.: High Performance, High Fidelity: A GPU-Accelerated Doubly-Periodic Configuration of the Simple Cloud-Resolving E3SM Atmosphere Model Version 1 (DP-SCREAMv1), J. Adv. Model. Earth Sy., 17, e2025MS005127, <ext-link xlink:href="https://doi.org/10.1029/2025MS005127" ext-link-type="DOI">10.1029/2025MS005127</ext-link>, 2025. </mixed-citation></ref>
      <ref id="bib1.bibx4"><label>Böhm et al.(2025)Böhm, Bauer, Kohl, Alappat, Thönnes, Mohr, Köstler, and Rüde</label><mixed-citation>Böhm, F., Bauer, D., Kohl, N., Alappat, C. L., Thönnes, D., Mohr, M., Köstler, H., and Rüde, U.: Code Generation and Performance Engineering for Matrix-Free Finite Element Methods on Hybrid Tetrahedral Grids, SIAM J. Sci. Comput., 47, B131–B159, <ext-link xlink:href="https://doi.org/10.1137/24M1653756" ext-link-type="DOI">10.1137/24M1653756</ext-link>, 2025.</mixed-citation></ref>
      <ref id="bib1.bibx5"><label>Conde et al.(2025)Conde, Ferreira, Canelas, Ricardo, and Mendes</label><mixed-citation>Conde, D. A. S., Ferreira, R. M. L., Canelas, R., Ricardo, A. M., and Mendes, L.: A Distributed-Heterogeneous Design for Explicit Hyperbolic Solvers. Application to Tsunami Urban Run-Up Modelling, J. Adv. Model. Earth Sy., 17, e2024MS004602, <ext-link xlink:href="https://doi.org/10.1029/2024MS004602" ext-link-type="DOI">10.1029/2024MS004602</ext-link>, 2025.</mixed-citation></ref>
      <ref id="bib1.bibx6"><label>Czarnul et al.(2020)Czarnul, Proficz, and Drypczewski</label><mixed-citation>Czarnul, P., Proficz, J., and Drypczewski, K.: Survey of Methodologies, Approaches, and Challenges in Parallel Programming Using High-Performance Computing Systems, Sci. Programming-Neth., 2020, 4176794, <ext-link xlink:href="https://doi.org/10.1155/2020/4176794" ext-link-type="DOI">10.1155/2020/4176794</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx7"><label>Farrell et al.(2009)Farrell, Piggott, Pain, Gorman, and Wilson</label><mixed-citation>Farrell, P. E., Piggott, M. D., Pain, C. C., Gorman, G. J., and Wilson, C. R.: Conservative interpolation between unstructured meshes via supermesh construction, Comput. Method. Appl. M., 198, 2632–2642, <ext-link xlink:href="https://doi.org/10.1016/j.cma.2009.03.004" ext-link-type="DOI">10.1016/j.cma.2009.03.004</ext-link>, 2009.</mixed-citation></ref>
      <ref id="bib1.bibx8"><label>Fischer et al.(2020)Fischer, Min, Rathnayake, Dutta, Kolev, Dobrev, Camier, Kronbichler, Warburton, Swirydowicz, and Brown</label><mixed-citation>Fischer, P., Min, M., Rathnayake, T., Dutta, S., Kolev, T., Dobrev, V., Camier, J.-S., Kronbichler, M., Warburton, T., Swirydowicz, K., and Brown, J.: Scalability of High-Performance PDE Solvers, Int. J. High Perform. C., 34, 562–586, <ext-link xlink:href="https://doi.org/10.1177/1094342020915762" ext-link-type="DOI">10.1177/1094342020915762</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx9"><label>Fu(2026)</label><mixed-citation>Fu, X.: GPU-Accelerated Implementation of the Fluidity-Atmosphere Dynamical Core, Zenodo [code], <ext-link xlink:href="https://doi.org/10.5281/zenodo.18799735" ext-link-type="DOI">10.5281/zenodo.18799735</ext-link>, 2026.</mixed-citation></ref>
      <ref id="bib1.bibx10"><label>Gan et al.(2026a)Gan, Li, Fang, Wu, Zhu, Wang, Zhu, and Zou</label><mixed-citation>Gan, P., Li, J., Fang, F., Wu, X., Zhu, J., Wang, Z., Zhu, M., and Zou, X.: HyMeshAI: Deep learning enabled three-dimensional adaptive mesh generator for high-resolution atmospheric simulations, J. Comput. Phys., 554, 114760, <ext-link xlink:href="https://doi.org/10.1016/j.jcp.2026.114760" ext-link-type="DOI">10.1016/j.jcp.2026.114760</ext-link>, 2026a.</mixed-citation></ref>
      <ref id="bib1.bibx11"><label>Gan et al.(2026b)Gan, Li, Wu, Wu, Zou, Wang, Zhu, Yuan, Xie, Tang, Li, and Fang</label><mixed-citation>Gan, P., Li, J., Wu, X., Wu, Q., Zou, X., Wang, Z., Zhu, J., Yuan, H., Xie, F., Tang, X., Li, L., and Fang, F.: Tibetan Plateau Mountain Wave Simulation Using AI-Driven 3D Adaptive Mesh Refinement, J. Geophys. Res.-Atmos., 131, <ext-link xlink:href="https://doi.org/10.1029/2025JD045585" ext-link-type="DOI">10.1029/2025JD045585</ext-link>, 2026b.</mixed-citation></ref>
      <ref id="bib1.bibx12"><label>Georgescu et al.(2013)Georgescu, Chow, and Okuda</label><mixed-citation>Georgescu, S., Chow, P., and Okuda, H.: GPU Acceleration for FEM-Based Structural Analysis, Arch. Comput. Method. E., 20, 111–121, <ext-link xlink:href="https://doi.org/10.1007/s11831-013-9082-8" ext-link-type="DOI">10.1007/s11831-013-9082-8</ext-link>, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx13"><label>Giraldo(2020)</label><mixed-citation>Giraldo, F. X.: An Introduction to Element-Based Galerkin Methods, Texts in Computational Science and Engineering, 1 edn., Springer, <ext-link xlink:href="https://doi.org/10.1007/978-3-030-55069-1" ext-link-type="DOI">10.1007/978-3-030-55069-1</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx14"><label>Goddeke et al.(2009)Goddeke, Buijssen, Wobker, and Turek</label><mixed-citation>Goddeke, D., Buijssen, S., Wobker, H., and Turek, S.: GPU acceleration of an unmodified parallel finite element Navier-Stokes solver, in: Proceedings of the 2009 International Conference on High Performance Computing and Simulation, HPCS 2009, pp. 12 – 21, <ext-link xlink:href="https://doi.org/10.1109/HPCSIM.2009.5191718" ext-link-type="DOI">10.1109/HPCSIM.2009.5191718</ext-link>, 2009.</mixed-citation></ref>
      <ref id="bib1.bibx15"><label>Huang et al.(2022)Huang, Chen, Li, Shen, and Xiao</label><mixed-citation>Huang, P., Chen, C., Li, X., Shen, X., and Xiao, F.: An Adaptive Nonhydrostatic Atmospheric Dynamical Core Using a Multi-Moment Constrained Finite Volume Method, Adv. Atmos. Sci., 39, 487–501, <ext-link xlink:href="https://doi.org/10.1007/s00376-021-1185-9" ext-link-type="DOI">10.1007/s00376-021-1185-9</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx16"><label>Jendersie et al.(2025)Jendersie, Lessig, and Richter</label><mixed-citation>Jendersie, R., Lessig, C., and Richter, T.: A GPU parallelization of the neXtSIM-DG dynamical core (v0.3.1), Geosci. Model Dev., 18, 3017–3040, <ext-link xlink:href="https://doi.org/10.5194/gmd-18-3017-2025" ext-link-type="DOI">10.5194/gmd-18-3017-2025</ext-link>, 2025.</mixed-citation></ref>
      <ref id="bib1.bibx17"><label>Kam et al.(2025)Kam, Tam, Sze, Cheung, Ng, and Lee</label><mixed-citation>Kam, P.-H., Tam, C.-Y., Sze, W.-P., Cheung, C.-C., Ng, K.-K., and Lee, S.-H.: Dynamically Adapting Mesh Refinement in an Unstructured Grid Global Model for Numerical Weather Prediction, Weather Forecast., 40, 1445 – 1462, <ext-link xlink:href="https://doi.org/10.1175/WAF-D-24-0239.1" ext-link-type="DOI">10.1175/WAF-D-24-0239.1</ext-link>, 2025.</mixed-citation></ref>
      <ref id="bib1.bibx18"><label>Kiran et al.(2020)Kiran, Gautam, and Sharma</label><mixed-citation>Kiran, U., Gautam, S. S., and Sharma, D.: GPU-based matrix-free finite element solver exploiting symmetry of elemental matrices, Computing, 102, 1941–1965, <ext-link xlink:href="https://doi.org/10.1007/s00607-020-00827-4" ext-link-type="DOI">10.1007/s00607-020-00827-4</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx19"><label>Kopera and Giraldo(2014)</label><mixed-citation>Kopera, M. A. and Giraldo, F. X.: Analysis of adaptive mesh refinement for IMEX discontinuous Galerkin solutions of the compressible Euler equations with application to atmospheric simulations, J. Comput. Phys., 275, 92–117, <ext-link xlink:href="https://doi.org/10.1016/j.jcp.2014.06.026" ext-link-type="DOI">10.1016/j.jcp.2014.06.026</ext-link>, 2014.</mixed-citation></ref>
      <ref id="bib1.bibx20"><label>Kronbichler and Kormann(2012)</label><mixed-citation>Kronbichler, M. and Kormann, K.: A generic interface for parallel cell-based finite element operator application, Comput. Fluids, 63, 135–147, <ext-link xlink:href="https://doi.org/10.1016/j.compfluid.2012.04.012" ext-link-type="DOI">10.1016/j.compfluid.2012.04.012</ext-link>, 2012.</mixed-citation></ref>
      <ref id="bib1.bibx21"><label>Kronbichler and Kormann(2019)</label><mixed-citation>Kronbichler, M. and Kormann, K.: Fast Matrix-Free Evaluation of Discontinuous Galerkin Finite Element Operators, ACM T. Math. Software, 45, 1–40, <ext-link xlink:href="https://doi.org/10.1145/3325864" ext-link-type="DOI">10.1145/3325864</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx22"><label>Li et al.(2021)Li, Fang, Steppeler, Zhu, Cheng, and Wu</label><mixed-citation>Li, J., Fang, F., Steppeler, J., Zhu, J., Cheng, Y., and Wu, X.: Demonstration of a three-dimensional dynamically adaptive atmospheric dynamic framework for the simulation of mountain waves, Meteorol. Atmos. Phys., 133, 1627–1645, <ext-link xlink:href="https://doi.org/10.1007/s00703-021-00828-8" ext-link-type="DOI">10.1007/s00703-021-00828-8</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx23"><label>Li et al.(2025)Li, Fu, Zheng, Li, and Li</label><mixed-citation>Li, L., Fu, X., Zheng, X., Li, H., and Li, J.: Atmospheric Mountain Wave Simulation Dataset (Fluidity-Atmosphere), Zenodo [data set], <ext-link xlink:href="https://doi.org/10.5281/zenodo.17824052" ext-link-type="DOI">10.5281/zenodo.17824052</ext-link>, 2025.</mixed-citation></ref>
      <ref id="bib1.bibx24"><label>Macioł et al.(2010)Macioł, Płaszewski, and Banaś</label><mixed-citation>Macioł, P., Płaszewski, P., and Banaś, K.: 3D finite element numerical integration on GPUs, Procedia Comput. Sci., 1, 1093–1100, <ext-link xlink:href="https://doi.org/10.1016/j.procs.2010.04.121" ext-link-type="DOI">10.1016/j.procs.2010.04.121</ext-link>, 2010.</mixed-citation></ref>
      <ref id="bib1.bibx25"><label>Marras et al.(2015)Marras, Kelly, Moragues Ginard, Müller, Kopera, Vázquez, Giraldo, Houzeaux, and Jorba</label><mixed-citation>Marras, S., Kelly, J., Moragues Ginard, M., Müller, A., Kopera, M., Vázquez, M., Giraldo, F., Houzeaux, G., and Jorba, O.: A Review of Element-Based Galerkin Methods for Numerical Weather Prediction: Finite Elements, Spectral Elements, and Discontinuous Galerkin, Arch. Comput. Method. E., 23, 673–722, <ext-link xlink:href="https://doi.org/10.1007/s11831-015-9152-1" ext-link-type="DOI">10.1007/s11831-015-9152-1</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx26"><label>Michalakes(2020)</label><mixed-citation>Michalakes, J.: HPC for Weather Forecasting, in: Parallel Algorithms in Computational Science and Engineering, edited by: Grama, A. and Sameh, A. H., Springer International Publishing, Cham, pp. 297–323, <ext-link xlink:href="https://doi.org/10.1007/978-3-030-43736-7_10" ext-link-type="DOI">10.1007/978-3-030-43736-7_10</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx27"><label>Mills et al.(2021)Mills, Adams, Balay, Brown, Dener, Knepley, Kruger, Morgan, Munson, Rupp, Smith, Zampini, Zhang, and Zhang</label><mixed-citation>Mills, R. T., Adams, M. F., Balay, S., Brown, J., Dener, A., Knepley, M., Kruger, S. E., Morgan, H., Munson, T., Rupp, K., Smith, B. F., Zampini, S., Zhang, H., and Zhang, J.: Toward performance-portable PETSc for GPU-based exascale systems, Parallel Comput., 108, 102831, <ext-link xlink:href="https://doi.org/10.1016/j.parco.2021.102831" ext-link-type="DOI">10.1016/j.parco.2021.102831</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx28"><label>Müller et al.(2013)Müller, Behrens, Giraldo, and Wirth</label><mixed-citation>Müller, A., Behrens, J., Giraldo, F. X., and Wirth, V.: Comparison between adaptive and uniform discontinuous Galerkin simulations in dry 2D bubble experiments, J. Comput. Phys., 235, 371–393, <ext-link xlink:href="https://doi.org/10.1016/j.jcp.2012.10.038" ext-link-type="DOI">10.1016/j.jcp.2012.10.038</ext-link>, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx29"><label>Müller et al.(2015)Müller, Scheichl, and Vainikko</label><mixed-citation>Müller, E. H., Scheichl, R., and Vainikko, E.: Petascale solvers for anisotropic PDEs in atmospheric modelling on GPU clusters, Parallel Comput., 50, 53–69, <ext-link xlink:href="https://doi.org/10.1016/j.parco.2015.10.007" ext-link-type="DOI">10.1016/j.parco.2015.10.007</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx30"><label>NVIDIA(2025a)</label><mixed-citation>NVIDIA: CUDA C<inline-formula><mml:math id="M208" display="inline"><mml:mrow><mml:mo>+</mml:mo><mml:mo>+</mml:mo></mml:mrow></mml:math></inline-formula> Programming Guide, <uri>https://docs.nvidia.com/cuda/cuda-c-programming-guide/index.html</uri> (last access: 10 August 2026), 2025a.</mixed-citation></ref>
      <ref id="bib1.bibx31"><label>NVIDIA(2025b)</label><mixed-citation>NVIDIA: Sparse Matrix Formats, <uri>https://docs.nvidia.com/nvpl/latest/sparse/storage_format/sparse_matrix.html</uri> (last access: 10 August 2026), 2025b.</mixed-citation></ref>
      <ref id="bib1.bibx32"><label>Petersen et al.(2026)Petersen, Asay-Davis, Barthel, Begeman, Bishnu, Brus, Jones, Kang, Kim, Mametjanov, O'Neill, Ringel, Smith, Sreepathi, Van Roekel, and Waruszewski</label><mixed-citation>Petersen, M. R., Asay-Davis, X. S., Barthel, A. M., Begeman, C. B., Bishnu, S., Brus, S. R., Jones, P. W., Kang, H.-G., Kim, Y., Mametjanov, A., O'Neill, B. J., Overfelt, J. R., Ringel, K. K., Smith, K. M., Sreepathi, S., Van Roekel, L. P., and Waruszewski, M.: The ocean model for E3SM global applications: Omega version 0.1.0 – a new high-performance computing code for exascale architectures, Geosci. Model Dev., 19, 3569–3594, <ext-link xlink:href="https://doi.org/10.5194/gmd-19-3569-2026" ext-link-type="DOI">10.5194/gmd-19-3569-2026</ext-link>, 2026.</mixed-citation></ref>
      <ref id="bib1.bibx33"><label>Piggott et al.(2008)Piggott, Gorman, Pain, Allison, Candy, Martin, and Wells</label><mixed-citation>Piggott, M. D., Gorman, G. J., Pain, C. C., Allison, P. A., Candy, A. S., Martin, B. T., and Wells, M. R.: A new computational framework for multi-scale ocean modelling based on adapting unstructured meshes, Int. J. Numer. Meth. Fl., 56, 1003–1015, <ext-link xlink:href="https://doi.org/10.1002/fld.1663" ext-link-type="DOI">10.1002/fld.1663</ext-link>, 2008.</mixed-citation></ref>
      <ref id="bib1.bibx34"><label>Ratnakar et al.(2021)Ratnakar, Sanfui, and Sharma</label><mixed-citation>Ratnakar, S. K., Sanfui, S., and Sharma, D.: Graphics Processing Unit-Based Element-by-Element Strategies for Accelerating Topology Optimization of Three-Dimensional Continuum Structures Using Unstructured All-Hexahedral Mesh, J. Comput. Inf. Sci. Eng., 22, 021013, <ext-link xlink:href="https://doi.org/10.1115/1.4052892" ext-link-type="DOI">10.1115/1.4052892</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx35"><label>Rudi et al.(2015)Rudi, Malossi, Isaac, Stadler, Gurnis, Staar, Ineichen, Bekas, Curioni, and Ghattas</label><mixed-citation>Rudi, J., Malossi, A. C. I., Isaac, T., Stadler, G., Gurnis, M., Staar, P. W. J., Ineichen, Y., Bekas, C., Curioni, A., and Ghattas, O.: An extreme-scale implicit solver for complex PDEs: highly heterogeneous flow in earth's mantle, in: SC '15: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, pp. 1–12, <ext-link xlink:href="https://doi.org/10.1145/2807591.2807675" ext-link-type="DOI">10.1145/2807591.2807675</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx36"><label>Sanfui and Sharma(2020)</label><mixed-citation>Sanfui, S. and Sharma, D.: A three-stage graphics processing unit-based finite element analyses matrix generation strategy for unstructured meshes, Int. J. Numer. Meth. Eng., 121, 3824–3848, <ext-link xlink:href="https://doi.org/10.1002/nme.6383" ext-link-type="DOI">10.1002/nme.6383</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx37"><label>Sanfui and Sharma(2021)</label><mixed-citation>Sanfui, S. and Sharma, D.: Symbolic and Numeric Kernel Division for Graphics Processing Unit-Based Finite Element Analysis Assembly of Regular Meshes With Modified Sparse Storage Formats, J. Comput. Inf. Sci. Eng., 22, <ext-link xlink:href="https://doi.org/10.1115/1.4051123" ext-link-type="DOI">10.1115/1.4051123</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx38"><label>Savre et al.(2016)Savre, Percival, Herzog, and Pain</label><mixed-citation>Savre, J., Percival, J., Herzog, M., and Pain, C.: Two-Dimensional Evaluation of ATHAM-Fluidity, a Nonhydrostatic Atmospheric Model Using Mixed Continuous/Discontinuous Finite Elements and Anisotropic Grid Optimization, Mon. Weather Rev., 144, 4349 – 4372, <ext-link xlink:href="https://doi.org/10.1175/MWR-D-15-0398.1" ext-link-type="DOI">10.1175/MWR-D-15-0398.1</ext-link>, 2016. </mixed-citation></ref>
      <ref id="bib1.bibx39"><label>Simek et al.(2009)Simek, Dvorak, Zboril, and Kunovsky</label><mixed-citation>Simek, V., Dvorak, R., Zboril, F., and Kunovsky, J.: Towards Accelerated Computation of Atmospheric Equations Using CUDA, in: 2009 11th International Conference on Computer Modelling and Simulation, pp. 449–454, <ext-link xlink:href="https://doi.org/10.1109/UKSIM.2009.25" ext-link-type="DOI">10.1109/UKSIM.2009.25</ext-link>, 2009.</mixed-citation></ref>
      <ref id="bib1.bibx40"><label>Smolarkiewicz et al.(2013)Smolarkiewicz, Szmelter, and Wyszogrodzki</label><mixed-citation>Smolarkiewicz, P. K., Szmelter, J., and Wyszogrodzki, A. A.: An unstructured-mesh atmospheric model for nonhydrostatic dynamics, J. Comput. Phys., 254, 184–199, <ext-link xlink:href="https://doi.org/10.1016/j.jcp.2013.07.027" ext-link-type="DOI">10.1016/j.jcp.2013.07.027</ext-link>, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx41"><label>Stone et al.(2021)Stone, Walden, Zubair, and Nielsen</label><mixed-citation>Stone, C. P., Walden, A., Zubair, M., and Nielsen, E. J.: Accelerating unstructured-grid CFD algorithms on NVIDIA and AMD GPUs, in: 2021 IEEE/ACM 11th Workshop on Irregular Applications: Architectures and Algorithms (IA3), pp. 19–26, <ext-link xlink:href="https://doi.org/10.1109/IA354616.2021.00010" ext-link-type="DOI">10.1109/IA354616.2021.00010</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx42"><label>Sulyok et al.(2019)Sulyok, Balogh, Reguly, and Mudalige</label><mixed-citation>Sulyok, A. A., Balogh, G. D., Reguly, I. Z., and Mudalige, G. R.: Locality optimized unstructured mesh algorithms on GPUs, J. Parallel Distr. Com., 134, 50–64, <ext-link xlink:href="https://doi.org/10.1016/j.jpdc.2019.07.011" ext-link-type="DOI">10.1016/j.jpdc.2019.07.011</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx43"><label>Szmelter et al.(2015)Szmelter, Zhang, and Smolarkiewicz</label><mixed-citation>Szmelter, J., Zhang, Z., and Smolarkiewicz, P. K.: An unstructured-mesh atmospheric model for nonhydrostatic dynamics: Towards optimal mesh resolution, J. Comput. Phys., 294, 363–381, <ext-link xlink:href="https://doi.org/10.1016/j.jcp.2015.03.054" ext-link-type="DOI">10.1016/j.jcp.2015.03.054</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx44"><label>Takle and Russell(1988)</label><mixed-citation>Takle, E. and Russell, R.: Applications of the finite element method to modeling the atmospheric boundary layer, Comput. Math. Appl., 16, 57–68, <ext-link xlink:href="https://doi.org/10.1016/0898-1221(88)90024-7" ext-link-type="DOI">10.1016/0898-1221(88)90024-7</ext-link>, 1988.</mixed-citation></ref>
      <ref id="bib1.bibx45"><label>Tang et al.(2023)Tang, Cui, Zhang, Zhou, Wu, Gong, and Zhang</label><mixed-citation>Tang, J., Cui, P., Zhang, J., Zhou, N., Wu, X., Gong, X., and Zhang, Y.: Review of mesh adaptation for fluid numerical simulation, Advances in Mechanics, 53, 661, <ext-link xlink:href="https://doi.org/10.6052/1000-0992-23-013" ext-link-type="DOI">10.6052/1000-0992-23-013</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx46"><label>Thibault and Senocak(2012)</label><mixed-citation>Thibault, J. and Senocak, I.: CUDA Implementation of a Navier-Stokes Solver on Multi-GPU Desktop Platforms for Incompressible Flows, in: 47th AIAA Aerospace Sciences Meeting including The New Horizons Forum and Aerospace Exposition, <ext-link xlink:href="https://doi.org/10.2514/6.2009-758" ext-link-type="DOI">10.2514/6.2009-758</ext-link>, 2012.</mixed-citation></ref>
      <ref id="bib1.bibx47"><label>Tissaoui et al.(2023)Tissaoui, Marras, Quaini, de Brangaca Alves, and Giraldo</label><mixed-citation>Tissaoui, Y., Marras, S., Quaini, A., de Brangaca Alves, F. A. V., and Giraldo, F. X.: A non-column based, fully unstructured implementation of Kessler's microphysics with warm rain using continuous and discontinuous spectral elements, J. Adv. Model. Earth Sy., 15, e2022MS003283, <ext-link xlink:href="https://doi.org/10.1029/2022MS003283" ext-link-type="DOI">10.1029/2022MS003283</ext-link>, 2023.</mixed-citation></ref>

  </ref-list></back>
    <!--<article-title-html>GPU-accelerated finite-element method for the three-dimensional unstructured mesh atmospheric dynamic framework</article-title-html>
<abstract-html/>
<ref-html id="bib1.bib1"><label>AMCG(2014)</label><mixed-citation>
      
AMCG: Fluidity manual, GitHub [code], <a href="https://fluidityproject.github.io/get-fluidity.html" target="_blank"/> (last access: 10 August 2026), 2014.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib2"><label>AMD(2025)</label><mixed-citation>
      
AMD: CUDA to HIP API Function Comparison, <a href="https://rocm.docs.amd.com/projects/HIP/en/latest/reference/api_syntax.html" target="_blank"/> (last access: 10 August 2026), 2025.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib3"><label>Bogenschutz et al.(2025)Bogenschutz, Clevenger, Bradley, Caldwell, Beydoun, Mahfouz, Keen, Guba, Bertagna, Foucar, Zhang, and Donahue</label><mixed-citation>
      
Bogenschutz, P. A., Clevenger, T. C., Bradley, A. M., Caldwell, P. M., Beydoun, H., Mahfouz, N., Keen, N. D., Guba, O., Bertagna, L., Foucar, J., Zhang, J., and Donahue, A. S.:
High Performance, High Fidelity: A GPU-Accelerated Doubly-Periodic Configuration of the Simple Cloud-Resolving E3SM Atmosphere Model Version 1 (DP-SCREAMv1), J. Adv. Model. Earth Sy., 17, e2025MS005127, <a href="https://doi.org/10.1029/2025MS005127" target="_blank">https://doi.org/10.1029/2025MS005127</a>, 2025.


    </mixed-citation></ref-html>
<ref-html id="bib1.bib4"><label>Böhm et al.(2025)Böhm, Bauer, Kohl, Alappat, Thönnes, Mohr, Köstler, and Rüde</label><mixed-citation>
      
Böhm, F., Bauer, D., Kohl, N., Alappat, C. L., Thönnes, D., Mohr, M., Köstler, H., and Rüde, U.:
Code Generation and Performance Engineering for Matrix-Free Finite Element Methods on Hybrid Tetrahedral Grids, SIAM J. Sci. Comput., 47, B131–B159, <a href="https://doi.org/10.1137/24M1653756" target="_blank">https://doi.org/10.1137/24M1653756</a>, 2025.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib5"><label>Conde et al.(2025)Conde, Ferreira, Canelas, Ricardo, and Mendes</label><mixed-citation>
      
Conde, D. A. S., Ferreira, R. M. L., Canelas, R., Ricardo, A. M., and Mendes, L.:
A Distributed-Heterogeneous Design for Explicit Hyperbolic Solvers. Application to Tsunami Urban Run-Up Modelling, J. Adv. Model. Earth Sy., 17, e2024MS004602, <a href="https://doi.org/10.1029/2024MS004602" target="_blank">https://doi.org/10.1029/2024MS004602</a>, 2025.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib6"><label>Czarnul et al.(2020)Czarnul, Proficz, and Drypczewski</label><mixed-citation>
      
Czarnul, P., Proficz, J., and Drypczewski, K.:
Survey of Methodologies, Approaches, and Challenges in Parallel Programming Using High-Performance Computing Systems, Sci. Programming-Neth., 2020, 4176794, <a href="https://doi.org/10.1155/2020/4176794" target="_blank">https://doi.org/10.1155/2020/4176794</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib7"><label>Farrell et al.(2009)Farrell, Piggott, Pain, Gorman, and Wilson</label><mixed-citation>
      
Farrell, P. E., Piggott, M. D., Pain, C. C., Gorman, G. J., and Wilson, C. R.:
Conservative interpolation between unstructured meshes via supermesh construction, Comput. Method. Appl. M., 198, 2632–2642, <a href="https://doi.org/10.1016/j.cma.2009.03.004" target="_blank">https://doi.org/10.1016/j.cma.2009.03.004</a>, 2009.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib8"><label>Fischer et al.(2020)Fischer, Min, Rathnayake, Dutta, Kolev, Dobrev, Camier, Kronbichler, Warburton, Swirydowicz, and Brown</label><mixed-citation>
      
Fischer, P., Min, M., Rathnayake, T., Dutta, S., Kolev, T., Dobrev, V., Camier, J.-S., Kronbichler, M., Warburton, T., Swirydowicz, K., and Brown, J.:
Scalability of High-Performance PDE Solvers, Int. J. High Perform. C., 34, 562–586, <a href="https://doi.org/10.1177/1094342020915762" target="_blank">https://doi.org/10.1177/1094342020915762</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib9"><label>Fu(2026)</label><mixed-citation>
      
Fu, X.: GPU-Accelerated Implementation of the Fluidity-Atmosphere Dynamical Core, Zenodo [code], <a href="https://doi.org/10.5281/zenodo.18799735" target="_blank">https://doi.org/10.5281/zenodo.18799735</a>, 2026.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib10"><label>Gan et al.(2026a)Gan, Li, Fang, Wu, Zhu, Wang, Zhu, and Zou</label><mixed-citation>
      
Gan, P., Li, J., Fang, F., Wu, X., Zhu, J., Wang, Z., Zhu, M., and Zou, X.:
HyMeshAI: Deep learning enabled three-dimensional adaptive mesh generator for high-resolution atmospheric simulations, J. Comput. Phys., 554, 114760, <a href="https://doi.org/10.1016/j.jcp.2026.114760" target="_blank">https://doi.org/10.1016/j.jcp.2026.114760</a>, 2026a.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib11"><label>Gan et al.(2026b)Gan, Li, Wu, Wu, Zou, Wang, Zhu, Yuan, Xie, Tang, Li, and Fang</label><mixed-citation>
      
Gan, P., Li, J., Wu, X., Wu, Q., Zou, X., Wang, Z., Zhu, J., Yuan, H., Xie, F., Tang, X., Li, L., and Fang, F.:
Tibetan Plateau Mountain Wave Simulation Using AI-Driven 3D Adaptive Mesh Refinement, J. Geophys. Res.-Atmos., 131, <a href="https://doi.org/10.1029/2025JD045585" target="_blank">https://doi.org/10.1029/2025JD045585</a>, 2026b.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib12"><label>Georgescu et al.(2013)Georgescu, Chow, and Okuda</label><mixed-citation>
      
Georgescu, S., Chow, P., and Okuda, H.:
GPU Acceleration for FEM-Based Structural Analysis, Arch. Comput. Method. E., 20, 111–121, <a href="https://doi.org/10.1007/s11831-013-9082-8" target="_blank">https://doi.org/10.1007/s11831-013-9082-8</a>, 2013.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib13"><label>Giraldo(2020)</label><mixed-citation>
      
Giraldo, F. X.:
An Introduction to Element-Based Galerkin Methods, Texts in Computational Science and Engineering, 1 edn., Springer, <a href="https://doi.org/10.1007/978-3-030-55069-1" target="_blank">https://doi.org/10.1007/978-3-030-55069-1</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib14"><label>Goddeke et al.(2009)Goddeke, Buijssen, Wobker, and Turek</label><mixed-citation>
      
Goddeke, D., Buijssen, S., Wobker, H., and Turek, S.:
GPU acceleration of an unmodified parallel finite element Navier-Stokes solver, in: Proceedings of the 2009 International Conference on High Performance Computing and Simulation, HPCS 2009, pp. 12 – 21, <a href="https://doi.org/10.1109/HPCSIM.2009.5191718" target="_blank">https://doi.org/10.1109/HPCSIM.2009.5191718</a>, 2009.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib15"><label>Huang et al.(2022)Huang, Chen, Li, Shen, and Xiao</label><mixed-citation>
      
Huang, P., Chen, C., Li, X., Shen, X., and Xiao, F.:
An Adaptive Nonhydrostatic Atmospheric Dynamical Core Using a Multi-Moment Constrained Finite Volume Method, Adv. Atmos. Sci., 39, 487–501, <a href="https://doi.org/10.1007/s00376-021-1185-9" target="_blank">https://doi.org/10.1007/s00376-021-1185-9</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib16"><label>Jendersie et al.(2025)Jendersie, Lessig, and Richter</label><mixed-citation>
      
Jendersie, R., Lessig, C., and Richter, T.:
A GPU parallelization of the neXtSIM-DG dynamical core (v0.3.1), Geosci. Model Dev., 18, 3017–3040, <a href="https://doi.org/10.5194/gmd-18-3017-2025" target="_blank">https://doi.org/10.5194/gmd-18-3017-2025</a>, 2025.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib17"><label>Kam et al.(2025)Kam, Tam, Sze, Cheung, Ng, and Lee</label><mixed-citation>
      
Kam, P.-H., Tam, C.-Y., Sze, W.-P., Cheung, C.-C., Ng, K.-K., and Lee, S.-H.:
Dynamically Adapting Mesh Refinement in an Unstructured Grid Global Model for Numerical Weather Prediction, Weather Forecast., 40, 1445 – 1462, <a href="https://doi.org/10.1175/WAF-D-24-0239.1" target="_blank">https://doi.org/10.1175/WAF-D-24-0239.1</a>, 2025.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib18"><label>Kiran et al.(2020)Kiran, Gautam, and Sharma</label><mixed-citation>
      
Kiran, U., Gautam, S. S., and Sharma, D.:
GPU-based matrix-free finite element solver exploiting symmetry of elemental matrices, Computing, 102, 1941–1965, <a href="https://doi.org/10.1007/s00607-020-00827-4" target="_blank">https://doi.org/10.1007/s00607-020-00827-4</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib19"><label>Kopera and Giraldo(2014)</label><mixed-citation>
      
Kopera, M. A. and Giraldo, F. X.:
Analysis of adaptive mesh refinement for IMEX discontinuous Galerkin solutions of the compressible Euler equations with application to atmospheric simulations, J. Comput. Phys., 275, 92–117, <a href="https://doi.org/10.1016/j.jcp.2014.06.026" target="_blank">https://doi.org/10.1016/j.jcp.2014.06.026</a>, 2014.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib20"><label>Kronbichler and Kormann(2012)</label><mixed-citation>
      
Kronbichler, M. and Kormann, K.:
A generic interface for parallel cell-based finite element operator application, Comput. Fluids, 63, 135–147, <a href="https://doi.org/10.1016/j.compfluid.2012.04.012" target="_blank">https://doi.org/10.1016/j.compfluid.2012.04.012</a>, 2012.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib21"><label>Kronbichler and Kormann(2019)</label><mixed-citation>
      
Kronbichler, M. and Kormann, K.:
Fast Matrix-Free Evaluation of Discontinuous Galerkin Finite Element Operators, ACM T. Math. Software, 45, 1–40, <a href="https://doi.org/10.1145/3325864" target="_blank">https://doi.org/10.1145/3325864</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib22"><label>Li et al.(2021)Li, Fang, Steppeler, Zhu, Cheng, and Wu</label><mixed-citation>
      
Li, J., Fang, F., Steppeler, J., Zhu, J., Cheng, Y., and Wu, X.:
Demonstration of a three-dimensional dynamically adaptive atmospheric dynamic framework for the simulation of mountain waves, Meteorol. Atmos. Phys., 133, 1627–1645, <a href="https://doi.org/10.1007/s00703-021-00828-8" target="_blank">https://doi.org/10.1007/s00703-021-00828-8</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib23"><label>Li et al.(2025)Li, Fu, Zheng, Li, and Li</label><mixed-citation>
      
Li, L., Fu, X., Zheng, X., Li, H., and Li, J.: Atmospheric Mountain Wave Simulation Dataset (Fluidity-Atmosphere), Zenodo [data set], <a href="https://doi.org/10.5281/zenodo.17824052" target="_blank">https://doi.org/10.5281/zenodo.17824052</a>, 2025.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib24"><label>Macioł et al.(2010)Macioł, Płaszewski, and Banaś</label><mixed-citation>
      
Macioł, P., Płaszewski, P., and Banaś, K.:
3D finite element numerical integration on GPUs, Procedia Comput. Sci., 1, 1093–1100, <a href="https://doi.org/10.1016/j.procs.2010.04.121" target="_blank">https://doi.org/10.1016/j.procs.2010.04.121</a>, 2010.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib25"><label>Marras et al.(2015)Marras, Kelly, Moragues Ginard, Müller, Kopera, Vázquez, Giraldo, Houzeaux, and Jorba</label><mixed-citation>
      
Marras, S., Kelly, J., Moragues Ginard, M., Müller, A., Kopera, M., Vázquez, M., Giraldo, F., Houzeaux, G., and Jorba, O.:
A Review of Element-Based Galerkin Methods for Numerical Weather Prediction: Finite Elements, Spectral Elements, and Discontinuous Galerkin, Arch. Comput. Method. E., 23, 673–722, <a href="https://doi.org/10.1007/s11831-015-9152-1" target="_blank">https://doi.org/10.1007/s11831-015-9152-1</a>, 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib26"><label>Michalakes(2020)</label><mixed-citation>
      
Michalakes, J.:
HPC for Weather Forecasting, in: Parallel Algorithms in Computational Science and Engineering, edited by: Grama, A. and Sameh, A. H., Springer International Publishing, Cham, pp. 297–323, <a href="https://doi.org/10.1007/978-3-030-43736-7_10" target="_blank">https://doi.org/10.1007/978-3-030-43736-7_10</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib27"><label>Mills et al.(2021)Mills, Adams, Balay, Brown, Dener, Knepley, Kruger, Morgan, Munson, Rupp, Smith, Zampini, Zhang, and Zhang</label><mixed-citation>
      
Mills, R. T., Adams, M. F., Balay, S., Brown, J., Dener, A., Knepley, M., Kruger, S. E., Morgan, H., Munson, T., Rupp, K., Smith, B. F., Zampini, S., Zhang, H., and Zhang, J.:
Toward performance-portable PETSc for GPU-based exascale systems, Parallel Comput., 108, 102831, <a href="https://doi.org/10.1016/j.parco.2021.102831" target="_blank">https://doi.org/10.1016/j.parco.2021.102831</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib28"><label>Müller et al.(2013)Müller, Behrens, Giraldo, and Wirth</label><mixed-citation>
      
Müller, A., Behrens, J., Giraldo, F. X., and Wirth, V.:
Comparison between adaptive and uniform discontinuous Galerkin simulations in dry 2D bubble experiments, J. Comput. Phys., 235, 371–393, <a href="https://doi.org/10.1016/j.jcp.2012.10.038" target="_blank">https://doi.org/10.1016/j.jcp.2012.10.038</a>, 2013.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib29"><label>Müller et al.(2015)Müller, Scheichl, and Vainikko</label><mixed-citation>
      
Müller, E. H., Scheichl, R., and Vainikko, E.:
Petascale solvers for anisotropic PDEs in atmospheric modelling on GPU clusters, Parallel Comput., 50, 53–69, <a href="https://doi.org/10.1016/j.parco.2015.10.007" target="_blank">https://doi.org/10.1016/j.parco.2015.10.007</a>, 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib30"><label>NVIDIA(2025a)</label><mixed-citation>
      
NVIDIA: CUDA C+ +  Programming Guide, <a href="https://docs.nvidia.com/cuda/cuda-c-programming-guide/index.html" target="_blank"/> (last access: 10 August 2026), 2025a.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib31"><label>NVIDIA(2025b)</label><mixed-citation>
      
NVIDIA: Sparse Matrix Formats, <a href="https://docs.nvidia.com/nvpl/latest/sparse/storage_format/sparse_matrix.html" target="_blank"/> (last access: 10 August 2026), 2025b.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib32"><label>Petersen et al.(2026)Petersen, Asay-Davis, Barthel, Begeman, Bishnu, Brus, Jones, Kang, Kim, Mametjanov, O'Neill, Ringel, Smith, Sreepathi, Van Roekel, and Waruszewski</label><mixed-citation>
      
Petersen, M. R., Asay-Davis, X. S., Barthel, A. M., Begeman, C. B., Bishnu, S., Brus, S. R., Jones, P. W., Kang, H.-G., Kim, Y., Mametjanov, A., O'Neill, B. J., Overfelt, J. R., Ringel, K. K., Smith, K. M., Sreepathi, S., Van Roekel, L. P., and Waruszewski, M.: The ocean model for E3SM global applications: Omega version 0.1.0 – a new high-performance computing code for exascale architectures, Geosci. Model Dev., 19, 3569–3594, <a href="https://doi.org/10.5194/gmd-19-3569-2026" target="_blank">https://doi.org/10.5194/gmd-19-3569-2026</a>, 2026.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib33"><label>Piggott et al.(2008)Piggott, Gorman, Pain, Allison, Candy, Martin, and Wells</label><mixed-citation>
      
Piggott, M. D., Gorman, G. J., Pain, C. C., Allison, P. A., Candy, A. S., Martin, B. T., and Wells, M. R.:
A new computational framework for multi-scale ocean modelling based on adapting unstructured meshes, Int. J. Numer. Meth. Fl., 56, 1003–1015, <a href="https://doi.org/10.1002/fld.1663" target="_blank">https://doi.org/10.1002/fld.1663</a>, 2008.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib34"><label>Ratnakar et al.(2021)Ratnakar, Sanfui, and Sharma</label><mixed-citation>
      
Ratnakar, S. K., Sanfui, S., and Sharma, D.:
Graphics Processing Unit-Based Element-by-Element Strategies for Accelerating Topology Optimization of Three-Dimensional Continuum Structures Using Unstructured All-Hexahedral Mesh, J. Comput. Inf. Sci. Eng., 22, 021013, <a href="https://doi.org/10.1115/1.4052892" target="_blank">https://doi.org/10.1115/1.4052892</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib35"><label>Rudi et al.(2015)Rudi, Malossi, Isaac, Stadler, Gurnis, Staar, Ineichen, Bekas, Curioni, and Ghattas</label><mixed-citation>
      
Rudi, J., Malossi, A. C. I., Isaac, T., Stadler, G., Gurnis, M., Staar, P. W. J., Ineichen, Y., Bekas, C., Curioni, A., and Ghattas, O.:
An extreme-scale implicit solver for complex PDEs: highly heterogeneous flow in earth's mantle, in: SC '15: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, pp. 1–12, <a href="https://doi.org/10.1145/2807591.2807675" target="_blank">https://doi.org/10.1145/2807591.2807675</a>, 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib36"><label>Sanfui and Sharma(2020)</label><mixed-citation>
      
Sanfui, S. and Sharma, D.:
A three-stage graphics processing unit-based finite element analyses matrix generation strategy for unstructured meshes, Int. J. Numer. Meth. Eng., 121, 3824–3848, <a href="https://doi.org/10.1002/nme.6383" target="_blank">https://doi.org/10.1002/nme.6383</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib37"><label>Sanfui and Sharma(2021)</label><mixed-citation>
      
Sanfui, S. and Sharma, D.:
Symbolic and Numeric Kernel Division for Graphics Processing Unit-Based Finite Element Analysis Assembly of Regular Meshes With Modified Sparse Storage Formats, J. Comput. Inf. Sci. Eng., 22, <a href="https://doi.org/10.1115/1.4051123" target="_blank">https://doi.org/10.1115/1.4051123</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib38"><label>Savre et al.(2016)Savre, Percival, Herzog, and Pain</label><mixed-citation>
      
Savre, J., Percival, J., Herzog, M., and Pain, C.:
Two-Dimensional Evaluation of ATHAM-Fluidity, a Nonhydrostatic Atmospheric Model Using Mixed Continuous/Discontinuous Finite Elements and Anisotropic Grid Optimization, Mon. Weather Rev., 144, 4349 – 4372, <a href="https://doi.org/10.1175/MWR-D-15-0398.1" target="_blank">https://doi.org/10.1175/MWR-D-15-0398.1</a>, 2016.


    </mixed-citation></ref-html>
<ref-html id="bib1.bib39"><label>Simek et al.(2009)Simek, Dvorak, Zboril, and Kunovsky</label><mixed-citation>
      
Simek, V., Dvorak, R., Zboril, F., and Kunovsky, J.:
Towards Accelerated Computation of Atmospheric Equations Using CUDA, in: 2009 11th International Conference on Computer Modelling and Simulation, pp. 449–454, <a href="https://doi.org/10.1109/UKSIM.2009.25" target="_blank">https://doi.org/10.1109/UKSIM.2009.25</a>, 2009.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib40"><label>Smolarkiewicz et al.(2013)Smolarkiewicz, Szmelter, and Wyszogrodzki</label><mixed-citation>
      
Smolarkiewicz, P. K., Szmelter, J., and Wyszogrodzki, A. A.:
An unstructured-mesh atmospheric model for nonhydrostatic dynamics, J. Comput. Phys., 254, 184–199, <a href="https://doi.org/10.1016/j.jcp.2013.07.027" target="_blank">https://doi.org/10.1016/j.jcp.2013.07.027</a>, 2013.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib41"><label>Stone et al.(2021)Stone, Walden, Zubair, and Nielsen</label><mixed-citation>
      
Stone, C. P., Walden, A., Zubair, M., and Nielsen, E. J.:
Accelerating unstructured-grid CFD algorithms on NVIDIA and AMD GPUs, in: 2021 IEEE/ACM 11th Workshop on Irregular Applications: Architectures and Algorithms (IA3), pp. 19–26, <a href="https://doi.org/10.1109/IA354616.2021.00010" target="_blank">https://doi.org/10.1109/IA354616.2021.00010</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib42"><label>Sulyok et al.(2019)Sulyok, Balogh, Reguly, and Mudalige</label><mixed-citation>
      
Sulyok, A. A., Balogh, G. D., Reguly, I. Z., and Mudalige, G. R.:
Locality optimized unstructured mesh algorithms on GPUs, J. Parallel Distr. Com., 134, 50–64, <a href="https://doi.org/10.1016/j.jpdc.2019.07.011" target="_blank">https://doi.org/10.1016/j.jpdc.2019.07.011</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib43"><label>Szmelter et al.(2015)Szmelter, Zhang, and Smolarkiewicz</label><mixed-citation>
      
Szmelter, J., Zhang, Z., and Smolarkiewicz, P. K.:
An unstructured-mesh atmospheric model for nonhydrostatic dynamics: Towards optimal mesh resolution, J. Comput. Phys., 294, 363–381, <a href="https://doi.org/10.1016/j.jcp.2015.03.054" target="_blank">https://doi.org/10.1016/j.jcp.2015.03.054</a>, 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib44"><label>Takle and Russell(1988)</label><mixed-citation>
      
Takle, E. and Russell, R.:
Applications of the finite element method to modeling the atmospheric boundary layer, Comput. Math. Appl., 16, 57–68, <a href="https://doi.org/10.1016/0898-1221(88)90024-7" target="_blank">https://doi.org/10.1016/0898-1221(88)90024-7</a>, 1988.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib45"><label>Tang et al.(2023)Tang, Cui, Zhang, Zhou, Wu, Gong, and Zhang</label><mixed-citation>
      
Tang, J., Cui, P., Zhang, J., Zhou, N., Wu, X., Gong, X., and Zhang, Y.:
Review of mesh adaptation for fluid numerical simulation, Advances in Mechanics, 53, 661, <a href="https://doi.org/10.6052/1000-0992-23-013" target="_blank">https://doi.org/10.6052/1000-0992-23-013</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib46"><label>Thibault and Senocak(2012)</label><mixed-citation>
      
Thibault, J. and Senocak, I.:
CUDA Implementation of a Navier-Stokes Solver on Multi-GPU Desktop Platforms for Incompressible Flows, in: 47th AIAA Aerospace Sciences Meeting including The New Horizons Forum and Aerospace Exposition, <a href="https://doi.org/10.2514/6.2009-758" target="_blank">https://doi.org/10.2514/6.2009-758</a>, 2012.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib47"><label>Tissaoui et al.(2023)Tissaoui, Marras, Quaini, de Brangaca Alves, and Giraldo</label><mixed-citation>
      
Tissaoui, Y., Marras, S., Quaini, A., de Brangaca Alves, F. A. V., and Giraldo, F. X.:
A non-column based, fully unstructured implementation of Kessler's microphysics with warm rain using continuous and discontinuous spectral elements, J. Adv. Model. Earth Sy., 15, e2022MS003283, <a href="https://doi.org/10.1029/2022MS003283" target="_blank">https://doi.org/10.1029/2022MS003283</a>, 2023.

    </mixed-citation></ref-html>--></article>
