<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing with OASIS Tables v3.0 20080202//EN" "journalpub-oasis3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:oasis="http://docs.oasis-open.org/ns/oasis-exchange/table" dtd-version="3.0">
  <front>
    <journal-meta>
<journal-id journal-id-type="publisher">GMD</journal-id>
<journal-title-group>
<journal-title>Geoscientific Model Development</journal-title>
<abbrev-journal-title abbrev-type="publisher">GMD</abbrev-journal-title>
<abbrev-journal-title abbrev-type="nlm-ta">Geosci. Model Dev.</abbrev-journal-title>
</journal-title-group>
<issn pub-type="epub">1991-9603</issn>
<publisher><publisher-name>Copernicus GmbH</publisher-name>
<publisher-loc>Göttingen, Germany</publisher-loc>
</publisher>
</journal-meta>

    <article-meta>
      <article-id pub-id-type="doi">10.5194/gmd-8-2067-2015</article-id><title-group><article-title>Experiences with distributed computing for meteorological applications: grid computing and cloud computing</article-title>
      </title-group><?xmltex \runningtitle{Grid computing and cloud computing}?><?xmltex \runningauthor{F.~Oesterle et al.}?>
      <contrib-group>
        <contrib contrib-type="author" corresp="yes" rid="aff1 aff3">
          <name><surname>Oesterle</surname><given-names>F.</given-names></name>
          <email>felix.oesterle@uibk.ac.at</email>
        <ext-link>https://orcid.org/0000-0002-7772-6884</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff2">
          <name><surname>Ostermann</surname><given-names>S.</given-names></name>
          
        </contrib>
        <contrib contrib-type="author" corresp="no" rid="aff2">
          <name><surname>Prodan</surname><given-names>R.</given-names></name>
          
        </contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1">
          <name><surname>Mayr</surname><given-names>G. J.</given-names></name>
          
        <ext-link>https://orcid.org/0000-0001-6661-9453</ext-link></contrib>
        <aff id="aff1"><label>1</label><institution>Institute of Atmospheric and Cryospheric Science,
University of Innsbruck,
Innrain 52,
6020 Innsbruck, Austria</institution>
        </aff>
        <aff id="aff2"><label>2</label><institution>Institute of Computer Science, University of Innsbruck,
Innsbruck, Austria</institution>
        </aff>
        <aff id="aff3"><label>*</label><institution>née Schüller</institution>
        </aff>
      </contrib-group>
      <author-notes><corresp id="corr1">F. Oesterle (felix.oesterle@uibk.ac.at)</corresp></author-notes><pub-date><day>13</day><month>July</month><year>2015</year></pub-date>
      
      <volume>8</volume>
      <issue>7</issue>
      <fpage>2067</fpage><lpage>2078</lpage>
      <history>
        <date date-type="received"><day>22</day><month>December</month><year>2014</year></date>
           <date date-type="rev-request"><day>10</day><month>February</month><year>2015</year></date>
           <date date-type="rev-recd"><day>17</day><month>June</month><year>2015</year></date>
           <date date-type="accepted"><day>29</day><month>June</month><year>2015</year></date>
      </history>
      <permissions>
<license license-type="open-access">
<license-p>This work is licensed under a Creative Commons Attribution 3.0 Unported License. To view a copy of this license, visit <ext-link ext-link-type="uri" xlink:href="http://creativecommons.org/licenses/by/3.0/">http://creativecommons.org/licenses/by/3.0/</ext-link></license-p>
</license>
</permissions><self-uri xlink:href="https://gmd.copernicus.org/articles/8/2067/2015/gmd-8-2067-2015.html">This article is available from https://gmd.copernicus.org/articles/8/2067/2015/gmd-8-2067-2015.html</self-uri>
<self-uri xlink:href="https://gmd.copernicus.org/articles/8/2067/2015/gmd-8-2067-2015.pdf">The full text article is available as a PDF file from https://gmd.copernicus.org/articles/8/2067/2015/gmd-8-2067-2015.pdf</self-uri>


      <abstract>
    <p>Experiences with three practical meteorological applications with different
characteristics are used to highlight the core computer science aspects and applicability
of distributed computing to meteorology.  Through presenting cloud and grid computing this paper
shows use case scenarios fitting a wide range of meteorological applications from
operational to research studies.  The paper concludes that distributed computing
complements and extends existing high performance computing concepts and allows for
simple, powerful and cost-effective access to computing capacity.</p>
  </abstract>
    </article-meta>
  </front>
<body>
      

      <?xmltex \hack{\allowdisplaybreaks}?>
<sec id="Ch1.S1" sec-type="intro">
  <title>Introduction</title>
      <p>Meteorology has an ever growing need for substantial amounts of computing power, be it for
sophisticated numerical models of the atmosphere itself, modelling systems and workflows,
coupled ocean and atmospheric models or the accompanying activities
visualisation or dissemination. In addition to the increased need for computing power,
more data are being produced, transferred and stored, which increases the problem.
Consequently, concepts and methods to supply the compute power and data handling capacity
also have to evolve.</p>
      <p>Until the beginning of this century, high performance clusters, local consortia and/or
buying cycles on commercial clusters were the main methods to acquire sufficient capacity.
Starting in the mid-1990s, the concept of grid computing, in which geographical and
institutional boundaries only play a minor role, became a powerful tool for scientists.
<xref ref-type="bibr" rid="bib1.bibx16" id="text.1"/> published the first and most cited definition of the grid:
“A computational grid is a hardware and software infrastructure that provides
dependable, consistent, pervasive, and inexpensive access to high-end computational
capabilities”. In the following years the definition changed to viewing the grid not as
a computing paradigm, but as an infrastructure that brings together different resources in
order to provide computing support for various applications, emphasising the social aspect
(<xref ref-type="bibr" rid="bib1.bibx18 bib1.bibx7" id="altparen.2"/>).  Grid initiatives can be classified as
<italic>compute grids</italic>, i.e. solely concentrated on raw computing power, or
<italic>data grids</italic> concentrating on storage/exchange of data.</p>
      <p>Many initiatives in the atmospheric sciences have utilised compute grids.  One of the
first climatological applications to use a compute grid is the Fast Ocean Atmospheric
Model (FOAM) <xref ref-type="bibr" rid="bib1.bibx26" id="paren.3"/>. They performed ensemble simulations of a coupled
climate model on the Teragrid, a US-based grid project sponsored by the National Science
Foundation.  More recently, <xref ref-type="bibr" rid="bib1.bibx15" id="text.4"/> provided an example with
the Community Atmospheric Model (CAM) for a climatological sensitivity study investigating
the connection of sea surface temperature and precipitation in the El Niño area.
<xref ref-type="bibr" rid="bib1.bibx38" id="text.5"/> presents three Bulgarian projects investigating air pollution and
climate change impacts. WRF4SG utilises grid computing with the Weather Research and
Forecast Model (WRF) <xref ref-type="bibr" rid="bib1.bibx5" id="paren.6"/> for various applications in weather forecasting and
extreme weather case studies.  TIGGE, the THORPEX Interactive Grand Global Ensemble,
partly uses grid computing to generate and share atmospheric data between various partner
<xref ref-type="bibr" rid="bib1.bibx8" id="paren.7"/>. The Earth system grid ESGF (Earth System Grid Federation) is a US–European data grid project
concentrating on storage and dissemination of climate simulation data
<xref ref-type="bibr" rid="bib1.bibx41" id="paren.8"/>.</p>
      <p>Cloud computing is slightly newer than grid computing.  Resources are also
pooled, but this time usually within one organisational unit, mostly within commercial
companies.  Similar to grids, applications range from services based on demand to simply
cutting ongoing costs or determining expected capacity needs.</p>
      <p>The most important characteristics of clouds are condensed into one of the most recent
definitions by <xref ref-type="bibr" rid="bib1.bibx23" id="text.9"/>: “Cloud computing is a model for enabling
ubiquitous, convenient, on-demand network access to a shared pool of configurable
computing resources, (e.g. networks, servers, storage, applications, and services) that
can be rapidly provisioned and released with minimal management effort or service provider
interaction”. Further definitions can be found in <xref ref-type="bibr" rid="bib1.bibx21" id="text.10"/>, <xref ref-type="bibr" rid="bib1.bibx39" id="text.11"/>
or <xref ref-type="bibr" rid="bib1.bibx39" id="text.12"/>.  One of the few papers to apply cloud technology to
meteorological research is <xref ref-type="bibr" rid="bib1.bibx13" id="text.13"/>, who conducted a feasibility study for cloud
computing with a coupled atmosphere–ocean model.</p>
      <p>In this paper, we discuss advantages and disadvantages of both infrastructures for atmospheric
research, show the supporting software ASKALON, and present three examples of
meteorological applications, which we have developed for different kinds of distributed
computing:  projects <italic>MeteoAG</italic> and <italic>MeteoAG2</italic> for a compute grid, and
<italic>RainCloud</italic> for cloud computing.  We look at issues and benefits mainly from our
perspective as users of distributed computing. Please note we describe our experiences but do not
show a direct, quantitative comparison, as we did not have the resources to run experiments on
both infrastructures with identical applications.</p>
</sec>
<sec id="Ch1.S2">
  <title>Aspects of distributed computing in meteorology</title>
<sec id="Ch1.S2.SS1">
  <title>Grid and cloud computing</title>
      <p>Our experiences in grid computing come from the projects MeteoAG and
MeteoAG2 within the national effort AustrianGrid (AGrid), including partners and
supercomputer centres from all over Austria <xref ref-type="bibr" rid="bib1.bibx40" id="paren.14"/>. AGrid phase 1
started in 2005 and concentrated on research of basic grid technology and application.
Phase 2, started in 2008, continued to build on research of phase 1 and additionally tried
to make AGrid self-sustaining. The research aim of this project was not to develop
conventional parallel applications that can be executed on individual grid machines, but
rather to unleash the power of the grid for single distributed program runs. To simplify
this task, all grid sites are required to run a similar Linux operating system. At the
height of the project AGrid consisted of nine clusters distributed over five locations in
Austria including various smaller sites with ad hoc desktop PC networks. The progress of
the project, its challenges and solutions were documented in several technical reports and
other publications <xref ref-type="bibr" rid="bib1.bibx6" id="paren.15"/>.</p>
      <p>For cloud computing, plenty of providers offer services, e.g. Rackspace or Google Compute
Engine.  Our cloud computing project <italic>RainCloud</italic> uses Amazon Web Services (AWS),
simply because it is the most well known and widely used.  AWS offers different services
for computing, different levels of data storage and data transfer, as well as tools for
monitoring and planning.  The services most interesting for meteorological computing
purposes are Amazon Elastic Compute cloud (EC2) for computing and Amazon Simple Storage
Service (S3) for data storage.  For computing, so-called <italic>instances</italic>, (i.e. virtual
computers) are defined according to their compute power relative to a reference CPU,
available memory, storage and network performance.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F1" specific-use="star"><caption><p>Schematic set-up of our computing environment for grid (left) and cloud
(right) computing. End
users interact with the ASKALON middleware via a Graphical User
Interface (GUI). The number of CPUs per cluster provided by the base grid varies,
whereas the instance types of cloud providers can be chosen. Execute engine, scheduler
and resource manager interact to effectively use the available resources and react
to changes in the provided computing infrastructure.</p></caption>
          <?xmltex \igopts{width=369.885827pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/8/2067/2015/gmd-8-2067-2015-f01.pdf"/>

        </fig>

      <p>Figure <xref ref-type="fig" rid="Ch1.F1"/> shows the basic structure of cloud computing on the right side
and AGrid as a grid example on the left side. In both cases an additional layer, so-called
middleware, is applied between the compute resources and the end user. The Middleware
layer handles all necessary scheduling, transfer of data and set up of cloud nodes.  Our
middleware is ASKALON <xref ref-type="bibr" rid="bib1.bibx28" id="paren.16"/>, which is described in more detail in
Sect. <xref ref-type="sec" rid="Ch1.S2.SS2"/>.</p>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T1" specific-use="star"><caption><p>Overview of advantages/disadvantages of grids and clouds affecting our
applications most. For a detailed discussion see Sect. 2.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="2">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="left"/>
     <oasis:thead>
       <oasis:row rowsep="1">  
         <oasis:entry colname="col1">grids</oasis:entry>  
         <oasis:entry colname="col2">clouds</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>  
         <oasis:entry colname="col1"><inline-formula><mml:math display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula> Massive amounts of data</oasis:entry>  
         <oasis:entry colname="col2"><inline-formula><mml:math display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula> Cost</oasis:entry>
       </oasis:row>
       <oasis:row>  
         <oasis:entry colname="col1"><inline-formula><mml:math display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula> Access to parallel computing enabled high performance  computing (HPC)</oasis:entry>  
         <oasis:entry colname="col2"><inline-formula><mml:math display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula> Full control of software set-up</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">  
         <oasis:entry colname="col1"/>  
         <oasis:entry colname="col2"><inline-formula><mml:math display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula> Simple on-demand access</oasis:entry>
       </oasis:row>
       <oasis:row>  
         <oasis:entry colname="col1"><inline-formula><mml:math display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula> Inhomogeneous hardware architectures</oasis:entry>  
         <oasis:entry colname="col2"><inline-formula><mml:math display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula> Slow data transfer</oasis:entry>
       </oasis:row>
       <oasis:row>  
         <oasis:entry colname="col1"><inline-formula><mml:math display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula> Complicated set-up and inflexible handling</oasis:entry>  
         <oasis:entry colname="col2"><inline-formula><mml:math display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula> Not suitable for MPI computing</oasis:entry>
       </oasis:row>
       <oasis:row>  
         <oasis:entry colname="col1"><inline-formula><mml:math display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula> Special compilation of source code needed</oasis:entry>  
         <oasis:entry colname="col2"/>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p>In the following sections, we list advantages and disadvantages of grid and cloud
concepts, which affected our research most (see Table <xref ref-type="table" rid="Ch1.T1"/> for a brief
overview). Criteria are extracted from literature, most notably <xref ref-type="bibr" rid="bib1.bibx17" id="normal.17"/>
containing a general comparison with all vital issues, <xref ref-type="bibr" rid="bib1.bibx23" id="text.18"/>,
<xref ref-type="bibr" rid="bib1.bibx21" id="normal.19"/> and <xref ref-type="bibr" rid="bib1.bibx16" id="normal.20"/>.  The discussed issue of security of
sensitive and valuable data did not apply to our research and operational setting.
However, for big and advanced operational weather forecasting this might be an issue due
to its monetary value. Because the hardware and network is completely out of the end
user's control, possible security breaches are harder or even impossible to detect.  If
security is a concern, detailed discussions can be found in <xref ref-type="bibr" rid="bib1.bibx10" id="normal.21"/> for grid
computing, and <xref ref-type="bibr" rid="bib1.bibx9" id="normal.22"/> and <xref ref-type="bibr" rid="bib1.bibx14" id="normal.23"/> for cloud computing.</p>
<sec id="Ch1.S2.SS1.SSS1">
  <title>Advantages and disadvantages grid</title>
      <p><def-list>
              <def-item><term>
                  <inline-formula><mml:math display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>
                </term><def>

                <p><italic>Handle massive amounts of data</italic>. The full atmospheric model in
MeteoAG generated large amounts of data. Through grid tools like
<italic>gridftp</italic> <xref ref-type="bibr" rid="bib1.bibx1" id="paren.24"/> we were able to  efficiently transfer and
store all simulation data.</p>
              </def></def-item>
              <def-item><term>
                  <inline-formula><mml:math display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>
                </term><def>

                <p><italic>Access to high performance  computing  (HPC)  which suits parallel applications, (e.g. Message Passing Interface, MPI).</italic> The model used in MeteoAG, as many other  meteorological models,
is  a massive parallel application parallelised with MPI. On single systems they
run efficiently; however, across different HPC clusters latencies become too
high.  A middleware can leverage the advantage of access to multiple machines and
run applications on suitable machines and appropriately distribute parts of
workflows in parallel.</p>
              </def></def-item>
              <def-item><term>
                  <inline-formula><mml:math display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>
                </term><def>

                <p><italic>Different hardware architectures.</italic> During tests in MeteoAG, we
discovered problems due to different hardware architectures <xref ref-type="bibr" rid="bib1.bibx35" id="paren.25"/>.
We tested different systems with
exactly the same set-up and software and got consistently different results. In our
case this affected our complex full model, but not our simple model. The exact
cause is unclear, but most likely a combination of programming, the libraries
and set-up down to the hardware level.</p>
              </def></def-item>
              <def-item><term>
                  <inline-formula><mml:math display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>
                </term><def>

                <p><italic>Difficult to set up and maintain as well as inflexible handling.</italic> For
us, the process of getting necessary updates, patches or special libraries needed
in meteorology onto all grid sites was complex and lengthy or sometimes even
impossible due to operating system limitations.</p>
              </def></def-item>
              <def-item><term>
                  <inline-formula><mml:math display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>
                </term><def>

                <p><italic>Special compilation of source code.</italic> To get the most out of the available
resources, the executables in MeteoAG needed to be compiled
for each architecture, with possible side effects. Even in a tightly managed
project like AGrid, we had to supply three different executables for the
meteorological model, with changes only during compilation, not in the
model code itself.</p>
              </def></def-item>
            </def-list></p>
      <p>Other typical characteristics are not as important for us.  The “limited amount of
resources” never influenced us as they were always vast enough to not hinder our models.
The “need to bring your own hardware/connections” is also a small hindrance, since
this is usually negotiable or the grid project might have different levels of
partnership.</p>
</sec>
<sec id="Ch1.S2.SS1.SSS2">
  <title>Advantages and disadvantages cloud computing</title>
      <p><def-list>
              <def-item><term>
                  <inline-formula><mml:math display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>
                </term><def>

                <p><italic>Cost</italic>. Costs can easily be determined and planned. More about costs
can be found in Sect. <xref ref-type="sec" rid="Ch1.S4"/>.</p>
              </def></def-item>
              <def-item><term>
                  <inline-formula><mml:math display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>
                </term><def>

                <p><italic>Full control of software environment, including operating system (OS) with root access</italic>. This proved to be one of the biggest advantages for our
workflows. It is easy to install software, special libraries or modify any
component of the system.  Cloud providers usually offer most standard operating
systems as images/Amazon Machine Image (AMI), but tuned images can also
be saved permanently and made publicly available (with additional storage costs).</p>
              </def></def-item>
              <def-item><term>
                  <inline-formula><mml:math display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>
                </term><def>

                <p><italic>Simple on-demand self-service.</italic> For applications with varying
requirements for compute resources or with repeated but short needs for compute
power, simple on-demand self-service is an important characteristic.  As long as funds are available the
required amount of compute power can be purchased. Our workflow was never forced
to wait for instances to be available.  Usually our standard on-demand Linux
instances were up and running within 5–10 <inline-formula><mml:math display="inline"><mml:mi mathvariant="normal">s</mml:mi></mml:math></inline-formula> (Amazon's documentation states
a maximum of 10 <inline-formula><mml:math display="inline"><mml:mi mathvariant="normal">min</mml:mi></mml:math></inline-formula>).</p>
              </def></def-item>
              <def-item><term>
                  <inline-formula><mml:math display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>
                </term><def>

                <p><italic>Slow data transfer and hardly any support for MPI computing.</italic> Data
transfer to and from cloud instances is slow as well as higher network latency
between the instances.  Only a subset of instance types are suitable for MPI
computing. This limitation makes cloud computing unsuitable for large-scale
complex atmospheric models.</p>
              </def></def-item>
            </def-list></p>
      <p>“Missing information about underlying hardware” has no impact on our workflow, as
we are not trying to optimise a single model execution. “No common standard
between clouds” and the possibility of “a cloud provider going out of business”
is also
unimportant for us.  Our software relies on common protocols like ssh and adaptation
to a new cloud provider could be done easily by adjusting the script requesting the
instances.</p>
</sec>
</sec>
<sec id="Ch1.S2.SS2">
  <title>Middleware ASKALON</title>
      <p>To make it as simple as possible for a (meteorological) end user to use distributed
computing resources, we make use of a so-called middleware system.  ASKALON, an existing
middleware from the Distributed and Parallel Systems group in Innsbruck, provides
integrated environments to support the development and execution of scientific workflows
on dynamic grid and cloud environments <xref ref-type="bibr" rid="bib1.bibx28" id="paren.26"/>.</p>
      <p>To account for the heterogeneity and the loosely coupled nature of resources from grid and
cloud providers, ASKALON has adopted a workflow paradigm <xref ref-type="bibr" rid="bib1.bibx37" id="paren.27"/> based
on loosely coupled coordination of atomic activities.  Distributed applications are split
in reasonably small execution parts, which can be executed in parallel on distributed
systems, allowing the runtime system to optimise resource usage, file transfers, load
balancing, reliability, scalability and handle failed parts. To overcome problems
resulting from unexpected job crashes and network interruptions, ASKALON is able to handle
most of the common failures. Jobs and file transfers are resubmitted on failure and jobs
might also be rescheduled to a different resource if transfers or jobs failed more than 5
times on a resource (<xref ref-type="bibr" rid="bib1.bibx30" id="altparen.28"/>). These features still exist in the
cloud version but play a less important role as resources showed to be more reliable in
the cloud case.</p>
      <p>Figure <xref ref-type="fig" rid="Ch1.F1"/> shows the design of the ASKALON system. Workflows can be
generated in a scientist-friendly Graphical User Interface (GUI) and submitted for
execution by a service. This allows for long lasting workflows without the need for the user
to be online throughout the whole execution period.</p>
      <p>Three main components handle the execution of the workflow:
<list list-type="custom"><list-item><label>1.</label><p><italic>Scheduler.</italic> Activities are mapped to physical (or virtualised) resources for their
execution with the end user deciding which pool of resources are used.  A wide set of
scheduling algorithms is available, e.g. Heterogeneous
Earliest
Finish Time
(HEFT) <xref ref-type="bibr" rid="bib1.bibx42" id="paren.29"/> or Dynamical
Critical Path - Clouds (DCP-C)
<xref ref-type="bibr" rid="bib1.bibx27" id="paren.30"/>. HEFT, for example, takes as input tasks, a set of resources, the
times to execute each task on each resource and the times to communicate results
between each job on each pair of resources.  Each task is assigned a priority and then
distributed onto the resources accordingly.  For the best possible scheduling,
a training phase is needed to get a function that relates the problem size to the
processing time.  Advanced techniques in prediction and machine learning are used to
achieve this goal (<xref ref-type="bibr" rid="bib1.bibx24 bib1.bibx25" id="altparen.31"/>).</p></list-item><list-item><label>2.</label><p><italic>Resource manager.</italic> cloud resources are known to “scale by credit card” and
theoretically an infinite amount of resources is available. The resource manager has
the task to provision the right amount of resources at the right moment to allow the
execute engine to run the workflow as the scheduler decided. Cost constraints must be
strictly adhered to as budgets are in practice limited. More on costs can be found in
Sect. <xref ref-type="sec" rid="Ch1.S4"/>.</p></list-item><list-item><label>3.</label><p><italic>Execute engine.</italic> Submission of jobs and transfer of data to the compute resources is
done with a suitable protocol, e.g. ssh or Globus resource allocation manager
(GRAM) in a Globus environment.</p></list-item><list-item><label>–</label><p><italic>System reliability.</italic> An important feature which is distributed over several
components of ASKALON is the capability to handle faults in distributed systems. Resources or network connections might
fail any time and mechanisms as described in <xref ref-type="bibr" rid="bib1.bibx30" id="text.32"/> are integrated in the execution engine <xref ref-type="bibr" rid="bib1.bibx32" id="text.33"/> allowing workflows to finish even when parts of the system fail.</p></list-item></list></p>
</sec>
</sec>
<sec id="Ch1.S3">
  <title>Applications in meteorology</title>
      <p>In the following subsections, we detail  the three applications we developed for usage
with distributed computing.  All projects investigate orographic precipitation over
complex terrain.  The most important distributed computing characteristics  of
the projects are shown in Table <xref ref-type="table" rid="Ch1.T3"/>.</p>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T2" specific-use="star"><caption><p>Prices and specifications for Amazon EC2 on-demand instances mentioned in this
paper,  running Linux OS in region <italic>EU-west</italic> as of November 2014. m1.xlarge,
m1.medium and m2.4xlarge are previous generations which were used in our experiments.
Storage is included in the instance, additional storage is available for purchase. One
elastic compute unit (ECU) provides the equivalent CPU capacity of a 1.0–1.2ĠHz 2007
Opteron or 2007 Xeon processor.  </p></caption><oasis:table frame="topbot"><oasis:tgroup cols="7">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="left"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:colspec colnum="5" colname="col5" align="right"/>
     <oasis:colspec colnum="6" colname="col6" align="right"/>
     <oasis:colspec colnum="7" colname="col7" align="right"/>
     <oasis:thead>
       <oasis:row rowsep="1">  
         <oasis:entry colname="col1">Instance family</oasis:entry>  
         <oasis:entry colname="col2">Instance type</oasis:entry>  
         <oasis:entry colname="col3">vCPU</oasis:entry>  
         <oasis:entry colname="col4">ECU</oasis:entry>  
         <oasis:entry colname="col5">Memory (GiB)</oasis:entry>  
         <oasis:entry colname="col6">Storage  (GB)</oasis:entry>  
         <oasis:entry colname="col7">Cost <inline-formula><mml:math display="inline"><mml:mrow><mml:mi mathvariant="normal">USD</mml:mi><mml:mspace width="0.125em" linebreak="nobreak"/><mml:msup><mml:mi mathvariant="normal">h</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>  
         <oasis:entry colname="col1">General purpose</oasis:entry>  
         <oasis:entry colname="col2">m3.medium</oasis:entry>  
         <oasis:entry colname="col3">1</oasis:entry>  
         <oasis:entry colname="col4">3</oasis:entry>  
         <oasis:entry colname="col5">3.75</oasis:entry>  
         <oasis:entry colname="col6">1 <inline-formula><mml:math display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 4 SSD</oasis:entry>  
         <oasis:entry colname="col7">0.077</oasis:entry>
       </oasis:row>
       <oasis:row>  
         <oasis:entry colname="col1"/>  
         <oasis:entry colname="col2">m3.2xlarge</oasis:entry>  
         <oasis:entry colname="col3">8</oasis:entry>  
         <oasis:entry colname="col4">26</oasis:entry>  
         <oasis:entry colname="col5">30</oasis:entry>  
         <oasis:entry colname="col6">2 <inline-formula><mml:math display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 80 SSD</oasis:entry>  
         <oasis:entry colname="col7">0.616</oasis:entry>
       </oasis:row>
       <oasis:row>  
         <oasis:entry colname="col1">Compute optim.</oasis:entry>  
         <oasis:entry colname="col2">c3.8xlarge</oasis:entry>  
         <oasis:entry colname="col3">32</oasis:entry>  
         <oasis:entry colname="col4">108</oasis:entry>  
         <oasis:entry colname="col5">60</oasis:entry>  
         <oasis:entry colname="col6">2 <inline-formula><mml:math display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 320 SSD</oasis:entry>  
         <oasis:entry colname="col7">1.912</oasis:entry>
       </oasis:row>
       <oasis:row>  
         <oasis:entry colname="col1">Storage optim.</oasis:entry>  
         <oasis:entry colname="col2">hs1.8xlarge</oasis:entry>  
         <oasis:entry colname="col3">16</oasis:entry>  
         <oasis:entry colname="col4">35</oasis:entry>  
         <oasis:entry colname="col5">117</oasis:entry>  
         <oasis:entry colname="col6">24 <inline-formula><mml:math display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 2048</oasis:entry>  
         <oasis:entry colname="col7">4.900</oasis:entry>
       </oasis:row>
       <oasis:row>  
         <oasis:entry colname="col1">Micro instances</oasis:entry>  
         <oasis:entry colname="col2">t1.micro</oasis:entry>  
         <oasis:entry colname="col3">1</oasis:entry>  
         <oasis:entry colname="col4">Var</oasis:entry>  
         <oasis:entry colname="col5">0.615</oasis:entry>  
         <oasis:entry colname="col6">EBS only</oasis:entry>  
         <oasis:entry colname="col7">0.014</oasis:entry>
       </oasis:row>
       <oasis:row>  
         <oasis:entry colname="col1">General purpose</oasis:entry>  
         <oasis:entry colname="col2">m1.xlarge</oasis:entry>  
         <oasis:entry colname="col3">4</oasis:entry>  
         <oasis:entry colname="col4">8</oasis:entry>  
         <oasis:entry colname="col5">15</oasis:entry>  
         <oasis:entry colname="col6"><inline-formula><mml:math display="inline"><mml:mrow><mml:mn mathvariant="normal">4</mml:mn><mml:mo>×</mml:mo><mml:mn>420</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>  
         <oasis:entry colname="col7">0.520</oasis:entry>
       </oasis:row>
       <oasis:row>  
         <oasis:entry colname="col1"/>  
         <oasis:entry colname="col2">m1.medium</oasis:entry>  
         <oasis:entry colname="col3">1</oasis:entry>  
         <oasis:entry colname="col4">2</oasis:entry>  
         <oasis:entry colname="col5">3.75</oasis:entry>  
         <oasis:entry colname="col6"><inline-formula><mml:math display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>×</mml:mo><mml:mn>410</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>  
         <oasis:entry colname="col7">0.130</oasis:entry>
       </oasis:row>
       <oasis:row>  
         <oasis:entry colname="col1">Memory optim.</oasis:entry>  
         <oasis:entry colname="col2">m2.4xlarge</oasis:entry>  
         <oasis:entry colname="col3">8</oasis:entry>  
         <oasis:entry colname="col4">26</oasis:entry>  
         <oasis:entry colname="col5">68.4</oasis:entry>  
         <oasis:entry colname="col6"><inline-formula><mml:math display="inline"><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mo>×</mml:mo><mml:mn>840</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>  
         <oasis:entry colname="col7">1.840</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

<sec id="Ch1.S3.SS1">
  <title>MeteoAG</title>
      <p>MeteoAG started as part of the AGrid computing initiative. Using ASKALON we
created a workflow to run a full numerical atmospheric model and visualisation on a grid
infrastructure <xref ref-type="bibr" rid="bib1.bibx33 bib1.bibx35 bib1.bibx34" id="paren.34"/>. The model is the non-hydrostatic Regional Atmospheric Modeling System (RAMS; version 6), a fully MPI
parallelised Fortran-based code <xref ref-type="bibr" rid="bib1.bibx11" id="paren.35"/>. The National Center for Atmospheric Research (NCAR) graphics library is used for
visualisation. Due to all AGrid sites running a similar Linux OS, no special code
adaptations to grid computing were needed.</p>
      <p>We simulated real cases as well as idealised test cases in the AGrid environment. Most
often these were parameter studies testing sensitivities to certain input parameters with
many slightly different runs. The investigated area in the realistic simulations covered
Europe and a target area over western Austria. Several nested domains are used with
a horizontal resolution of the innermost domain of 500 <inline-formula><mml:math display="inline"><mml:mi mathvariant="normal">m</mml:mi></mml:math></inline-formula> and 60 vertical levels (approx. 7.5 million grid points). Figure <xref ref-type="fig" rid="Ch1.F2"/> shows the workflow deployed to the AGrid.
Starting with many simulations with a shorter simulation time, it was then decided which
runs to extend further. Only runs where heavy precipitation occurs above a certain
threshold were chosen.  Post-processing done on the compute grid includes extraction of
variables and preliminary visualisation, but the main visualisation was done on a local
machine.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F2" specific-use="star"><caption><p>Workflow of MeteoAG using the Regional Atmospheric Modelling System (RAMS) and
supporting software REVU (extracts variables) and RAVER (analyses variables). Each case
represents a different weather event. <bold>(a)</bold> Meteorological
representation with indication which activities are parallelised using Message Passing Interface (MPI).
<bold>(b)</bold> Workflow representation of the activities as used by ASKALON middleware. In addition to the
different cases, selected variables are varied within each case. Same colours between the subfigures.
</p></caption>
          <?xmltex \igopts{width=369.885827pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/8/2067/2015/gmd-8-2067-2015-f02.pdf"/>

        </fig>

      <p>The workflow characteristics relevant for distributed computing are a few (20–50) model
instances but highly CPU intensive as well as lots of interprocess communications.
Results of this workflow require a substantial amount of data transfer between the
different grid sites and the end user (O(200 Gb)).</p>
      <p>Upon investigation of our first runs it was necessary to provide different executables for
specific architectures (32 bit, 64 bit, 64 bit Intel) to get optimum speed.  We ran into
a problem while executing the full model on different architectures. Using the exact same
static executable with the same input parameters and set-up led to consistently different
results across different clusters <xref ref-type="bibr" rid="bib1.bibx35" id="paren.36"/>. For real case simulations, these
errors are negligible compared to errors in the model itself. But for idealised
simulations, e.g. investigation of turbulence with an atmosphere initially at rest, where
tiny perturbations play a major role, this might lead to serious problems. We were not
able to determine the cause of these differences. It seems to be a problem of the complex
code of the full model and its interaction with the underlying libraries. While we can
only speculate on the exact cause, we strongly advise using a simple and quick test such
as simulating an atmosphere at rest or linear orographic precipitation to test for such
differences.</p>
</sec>
<sec id="Ch1.S3.SS2">
  <title>MeteoAG2</title>
      <p>MeteoAG2 is the continuation of MeteoAG and also part of AGrid
<xref ref-type="bibr" rid="bib1.bibx31" id="paren.37"/>.  Based on the experience from the MeteoAG experiments, we
hypothesise that it would be much more effective to deploy an application consisting of
serial CPU jobs. ASKALON is optimised for submission of single core parts of a workflow,
which avoids internal parallelism and communication of activities and allows for the best control
over the execution within ASKALON. Thus MeteoAG2 uses a simpler meteorological model, the linear model (LM) of
orographic precipitation <xref ref-type="bibr" rid="bib1.bibx36" id="paren.38"/>. The model computes only very simple
linear equations of orographic precipitation, is not parallelised, and has short runtime,
O(10 s),  even with high resolutions (500 <inline-formula><mml:math display="inline"><mml:mi mathvariant="normal">m</mml:mi></mml:math></inline-formula>) over large domains. LM is written in Fortran.
ASKALON is again used for workflow execution and Matlab routines for visualisation.</p>
      <p>With this workflow, rainfall over the Alps was investigated by taking input from the European
Centre for Medium-Range Weather Forecasts (ECMWF) model,  splitting the Alps into
subdomains (see Fig. <xref ref-type="fig" rid="Ch1.F3"/>a) and running the model within each subdomain
with variations in the input parameters. The last step combines the results from all
subdomains and visualises them. Using grid computing allowed us to run many O(50 000)
simulations in a relatively short amount of time O(h). This compares to about 50 typical,
albeit a lot more complex runs in current operational meteorological set-ups.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F3" specific-use="star"><caption><p>Set-up and workflow of MeteoAG2 using the linear model (LM) of orographic
precipitation. <bold>(a)</bold> Grid set-up of experiments in MeteoAG2 with dots representing grid
points of the European Center of Medium Range Weather Forecast (ECMWF) used to drive the
LM. Topography height in kilometres  a.m.s.l.
<bold>(b)</bold> Workflow representation of the activities as used by ASKALON. Activity MakeNML
prepares all input sequentially.
ProdNCfile is the main activity with the linear model run in parallel on the grid.
Panel
<bold>(a)</bold> courtesy of <xref ref-type="bibr" rid="bib1.bibx31" id="text.39"/>.</p></caption>
          <?xmltex \igopts{width=426.791339pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/8/2067/2015/gmd-8-2067-2015-f03.pdf"/>

        </fig>

      <p>The workflow deployed to the grid (Fig. <xref ref-type="fig" rid="Ch1.F3"/>b) is simple with only two
main activities: preparing all the input parameters for all subdomains and then the
parallel execution of all runs.  One of the drawbacks of MeteoAG2 is the very strict set-up
that was necessary due to the state of ASKALON at that time, e.g. no robust if-construct
yet, and the direct use of model executables without wrappers. The workflow could not
easily be changed to suit different research needs, e.g. change to different input
parameters for LM or to using a different model.</p>
</sec>
<sec id="Ch1.S3.SS3">
  <title>RainCloud</title>
      <p>In switching to cloud computing, RainCloud uses an extended version of the same simple model
of orographic precipitation as MeteoAG2. The main extension to LM is the ability to
simulate different layers, while still retaining its fast execution time
<xref ref-type="bibr" rid="bib1.bibx2" id="paren.40"/>.  The software stack again includes ASKALON, the
Fortran-based LM, python scripts and Matplotlib for visualisation.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F4"><caption><p>Workflow of RainClouds operational setting for the Avalanche Warning Service
Tyrol (LWD) using the double layer linear model (LM) of orographic precipitation.
Input data from the European Center for Medium Range Weather Forecast (ECMWF). <bold>(a)</bold>
Meteorological flow chart with parts not executed on cloud (in red).  <bold>(b)</bold> Workflow with
activities as used by ASKALON. Same colours between subfigures. </p></caption>
          <?xmltex \igopts{width=241.848425pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/8/2067/2015/gmd-8-2067-2015-f04.pdf"/>

        </fig>

      <p>The inclusion of if-constructs in ASKALON and a different approach to the scripting of
activities, (e.g. wrapping the model executables in python scripts and calling these)
allows RainCloud to be used in different set-ups. We are now able to run the workflow in
three
flavours without any changes: idealised, semi-idealised and realistic simulations as well as
different settings, operational and research. Figure <xref ref-type="fig" rid="Ch1.F4"/>b depicts the
workflow run on cloud computing. Only the first two activities, <italic>PrepareLM</italic> and
<italic>LinearModel</italic> have to be run, the others are optional. This workflow fits a lot of
meteorological applications as it has the following building blocks:
<list list-type="bullet"><list-item><p>preparation of the simulations (<italic>PrepareLM</italic>);</p></list-item><list-item><p>execution of a meteorological model (<italic>LinearModel</italic>);</p></list-item><list-item><p>post-processing of each individual run, e.g. for producing derived variables
(<italic>PostProcessSingle</italic>);</p></list-item><list-item><p>post-processing of all runs (<italic>PostprocessFinal</italic>).</p></list-item></list></p>
      <p>All activities are wrapped in Python scripts. As long as the input and output between
these activities are named the same, everything within the activity can be changed. We
use archives for transfer between the activities, again allowing different files to be
packed into these archives.</p>
      <p>The operational set-up produces spatially detailed, daily probabilistic precipitation
forecasts for the Avalanche Service Tyrol (Lawinenwarndienst Tirol) to help forecast
avalanche danger. Figure <xref ref-type="fig" rid="Ch1.F4"/>a shows the schematic of our operational
workflow. Starting with data from the ECMWF, we forecast and visualise precipitation
probabilities over Tyrol with a spatial resolution of 500 <inline-formula><mml:math display="inline"><mml:mi mathvariant="normal">m</mml:mi></mml:math></inline-formula>. Additionally, research type
experiments are used to test, explore and run experiments with new developments in LM
through parameter studies.</p>
      <p>Our workflow set-ups vary substantially in required computation power as well as data
size. The operational job is run daily during winter, whereas research types are run in
bursts. Data usage <italic>within</italic> the cloud can be substantial O(500 Gb) with all
flavours, but with big differences of data transfer from the cloud back to the local
machine. Operational results are small, of the order of O(100 Mb), while research results
can amount to O(100 Gb), influencing the overall runtime and costs due to the additional
data transfer time.</p>
</sec>
</sec>
<sec id="Ch1.S4">
  <title>Costs, performance and usage scenarios</title>
<sec id="Ch1.S4.SS1">
  <title>Costs</title>
      <p>To define the exact costs for a <italic>dedicated server</italic> system or the
participation in a <italic>grid</italic> initiative is not trivial, and often even unknown to the
provider; we contacted several of them, but due to complicated budgeting methodologies the
final costs are not obvious.  <xref ref-type="bibr" rid="bib1.bibx19" id="normal.41"/> discussed costs for operating a server
environment for data services from a provider perspective, including
servers, infrastructure, power requirements and networking. However, the authors did not
include the cost of human resources for, e.g., system administration. <xref ref-type="bibr" rid="bib1.bibx29" id="normal.42"/>
included human resources and establish a cost model for set-up and maintenance of a data
centre.  Grids may have different and negotiable levels of access and participation, with
varying associated costs to the user.  Some initiatives, e.g.  PRACE <xref ref-type="bibr" rid="bib1.bibx20" id="paren.43"/>,
offer free access to grid resources  after a proposal/review process.</p>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T3" specific-use="star"><caption><p>Overview of our projects and their workflow characteristics.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="4">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="justify" colwidth="113.811024pt"/>
     <oasis:colspec colnum="3" colname="col3" align="justify" colwidth="113.811024pt"/>
     <oasis:colspec colnum="4" colname="col4" align="justify" colwidth="99.584646pt"/>
     <oasis:thead>
       <oasis:row rowsep="1">  
         <oasis:entry colname="col1">Project</oasis:entry>  
         <oasis:entry colname="col2">MeteoAG</oasis:entry>  
         <oasis:entry colname="col3">MeteoAG2</oasis:entry>  
         <oasis:entry colname="col4">RainCloud</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row rowsep="1">  
         <oasis:entry colname="col1">Type</oasis:entry>  
         <oasis:entry colname="col2">grid</oasis:entry>  
         <oasis:entry colname="col3">grid</oasis:entry>  
         <oasis:entry colname="col4">cloud</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">  
         <oasis:entry colname="col1">Meteorological model</oasis:entry>  
         <oasis:entry colname="col2">RAMS (Regional Atmospheric Modeling System)</oasis:entry>  
         <oasis:entry colname="col3">single layer linear model of orographic precipitation</oasis:entry>  
         <oasis:entry colname="col4">double layer linear model of orographic precipitation</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">  
         <oasis:entry colname="col1">Model type</oasis:entry>  
         <oasis:entry colname="col2">complex full numerical model parallelised with Message Passing Interface (MPI)</oasis:entry>  
         <oasis:entry colname="col3">simplified model</oasis:entry>  
         <oasis:entry colname="col4">double layer simplified model</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">  
         <oasis:entry colname="col1">Parallel runs</oasis:entry>  
         <oasis:entry colname="col2">20–50</oasis:entry>  
         <oasis:entry colname="col3">approx. 50 000</oasis:entry>  
         <oasis:entry colname="col4"><inline-formula><mml:math display="inline"><mml:mrow><mml:mo>&gt;</mml:mo><mml:mn>5000</mml:mn></mml:mrow></mml:math></inline-formula> operational,<?xmltex \hack{\hfill\break}?> <inline-formula><mml:math display="inline"><mml:mrow><mml:mo>&gt;</mml:mo><mml:mn>10 000</mml:mn></mml:mrow></mml:math></inline-formula> research</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">  
         <oasis:entry colname="col1">Runtime</oasis:entry>  
         <oasis:entry colname="col2">several days</oasis:entry>  
         <oasis:entry colname="col3">several hours</oasis:entry>  
         <oasis:entry colname="col4">1–2 <inline-formula><mml:math display="inline"><mml:mi mathvariant="normal">h</mml:mi></mml:math></inline-formula> operational/<?xmltex \hack{\hfill\break}?> <inline-formula><mml:math display="inline"><mml:mo>&lt;</mml:mo></mml:math></inline-formula>1 <inline-formula><mml:math display="inline"><mml:mi mathvariant="normal">h</mml:mi></mml:math></inline-formula> research</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">  
         <oasis:entry colname="col1">Data transfer</oasis:entry>  
         <oasis:entry colname="col2">O(200GB)</oasis:entry>  
         <oasis:entry colname="col3">O(1GB)</oasis:entry>  
         <oasis:entry colname="col4">O(MB) – O(1GB)</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">  
         <oasis:entry colname="col1">Workflow flexibility</oasis:entry>  
         <oasis:entry colname="col2">strict</oasis:entry>  
         <oasis:entry colname="col3">strict</oasis:entry>  
         <oasis:entry colname="col4">flexible</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">  
         <oasis:entry colname="col1">Applications</oasis:entry>  
         <oasis:entry colname="col2">parameter studies,<?xmltex \hack{\hfill\break}?>case studies</oasis:entry>  
         <oasis:entry colname="col3">downscaling</oasis:entry>  
         <oasis:entry colname="col4">parameter studies,<?xmltex \hack{\hfill\break}?>downscaling,<?xmltex \hack{\hfill\break}?>probabilistic forecasts,<?xmltex \hack{\hfill\break}?>model testing</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">  
         <oasis:entry colname="col1">Intent</oasis:entry>  
         <oasis:entry colname="col2">research</oasis:entry>  
         <oasis:entry colname="col3">research</oasis:entry>  
         <oasis:entry colname="col4">operational, research</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">  
         <oasis:entry colname="col1">Frequency</oasis:entry>  
         <oasis:entry colname="col2">on demand</oasis:entry>  
         <oasis:entry colname="col3">on demand</oasis:entry>  
         <oasis:entry colname="col4">operational: daily; research: on demand</oasis:entry>
       </oasis:row>
       <oasis:row>  
         <oasis:entry colname="col1">Programming</oasis:entry>  
         <oasis:entry colname="col2">shell scripts, Fortran, NCAR Graphics, MPI</oasis:entry>  
         <oasis:entry colname="col3">shell scripts, Fortran, Matlab</oasis:entry>  
         <oasis:entry colname="col4">python, Fortran</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p><?xmltex \hack{\newpage}?>Cloud computing on the other hand offers simpler and transparent costs.  Pricing varies
depending on the provider, capability of a resource and also on the geographical region.
Prices (as of November 2014) of AWS on-demand compute instances for Linux OS can
be found in Table <xref ref-type="table" rid="Ch1.T2"/> and range from USD 0.014 up to <inline-formula><mml:math display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">5</mml:mn></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="normal">h</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>
(region Ireland).  Cheaper instance pricing is available through
spot instances where one bids on spare resources. These resources might get
cancelled if demand rises, but are a valid option for interruption-tolerant workflows or
for developing a workflow.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F5" specific-use="star"><caption><p>Bars show overall runtime of one operational run on various Amazon EC2 instance
types, each with a total of 32 cores (left <inline-formula><mml:math display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula> axis). Each bar represents one workflow
invocation with the corresponding instance type. Dots show costs for on-demand instances
(x) and spot instances (circle; right <inline-formula><mml:math display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula> axis). Only the execution part is shown, spin-up time, i.e. preparation and installation (2–5 <inline-formula><mml:math display="inline"><mml:mi mathvariant="normal">min</mml:mi></mml:math></inline-formula>) is not included.  See Table <xref ref-type="table" rid="Ch1.T3"/> for exact specifications.  All experiments were run during March 2014
with the exact same set-up.  </p></caption>
          <?xmltex \igopts{width=312.980315pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/8/2067/2015/gmd-8-2067-2015-f05.pdf"/>

        </fig>

      <p>Figure <xref ref-type="fig" rid="Ch1.F5"/> shows the difference between spot and on-demand pricing for 25
test runs of our operational workflow (circle and x; right <inline-formula><mml:math display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula> axis). All
runs use 32 cores but a different number of instances, i.e. only one c3.8xlarge (32 cores)
instance, but 32 m1.medium (1 core) instances. Runtime only includes the actual workflow,
not the spin-up needed to prepare the instances. It usually takes 5–10 <inline-formula><mml:math display="inline"><mml:mi mathvariant="normal">s</mml:mi></mml:math></inline-formula> for an
instance to become available and another 2–5 <inline-formula><mml:math display="inline"><mml:mi mathvariant="normal">min</mml:mi></mml:math></inline-formula> to set up the system and install
necessary libraries and software.  Spot and on demand only differ in the pricing scheme
not in the computational resources themselves. With spot pricing we achieved savings
between 65 and 90 %, however with an additional start-up latency of 2–3 <inline-formula><mml:math display="inline"><mml:mi mathvariant="normal">min</mml:mi></mml:math></inline-formula> (compared to
5–10 <inline-formula><mml:math display="inline"><mml:mi mathvariant="normal">s</mml:mi></mml:math></inline-formula>).</p>
      <p>To give an idea, a very simplified cost comparison can be done with the purchasing costs
of dedicated hardware, excluding costs for system administration, cooling or power. The
operational part of RainCloud runs on 32 cores for approximately 3 h per day for
6
months of the year, i.e. 550 h per year.
<list list-type="bullet"><list-item><p>A dedicated 32 core server with 64 GB RAM costs around USD 5500 (various
brands, excluding tax, Austria, November 2014).</p></list-item><list-item><p>A comparable on-demand AWS instance
(c3.x8large; 32 cores, 60 GB RAM) could run for USD <inline-formula><mml:math display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn>2800</mml:mn></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math display="inline"><mml:mi mathvariant="normal">h</mml:mi></mml:math></inline-formula> at
1.91 <inline-formula><mml:math display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="normal">h</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> pricing.</p></list-item></list>
Assuming no instance price variance, our operational workflow could be run on AWS for
approximately 5 years, the usual depreciation time for hardware. This suggests that AWS
is the cheaper alternative for RainCloud, since hardware is only one part of the total
cost of ownership of a dedicated system.</p>
</sec>
<sec id="Ch1.S4.SS2">
  <title>Performance</title>
      <p>For our operational RainCloud workflow, Fig. <xref ref-type="fig" rid="Ch1.F5"/> shows the effect of
different instance types on the runtime.  First, a clear difference between the instance
types is evident, with the longest running taking nearly twice as long as the shortest
one.  Second, even within one instance type, runtime varies by 10–20 percent.  Serial
execution on a 1 core desktop PC takes about 12 <inline-formula><mml:math display="inline"><mml:mi mathvariant="normal">h</mml:mi></mml:math></inline-formula>, i.e. a speedup of
<inline-formula><mml:math display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn>18</mml:mn></mml:mrow></mml:math></inline-formula> (a runtime of  <inline-formula><mml:math display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn>0.66</mml:mn></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math display="inline"><mml:mi mathvariant="normal">h</mml:mi></mml:math></inline-formula> as seen in Fig. <xref ref-type="fig" rid="Ch1.F5"/>).  Based
on these experiments our daily operational workflow uses four m3.2xlarge instances.</p>
      <p>To put this into relation, <xref ref-type="bibr" rid="bib1.bibx35" id="text.44"/> showed a speedup for MeteoAG of multiple
cores vs. 1 core for a short running test set-up of <inline-formula><mml:math display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">5</mml:mn></mml:mrow></mml:math></inline-formula>, with higher speedups
possible for a full complex workflow run. For MeteoAG2, <xref ref-type="bibr" rid="bib1.bibx31" id="text.45"/>
showed
a speedup of <inline-formula><mml:math display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn>120</mml:mn></mml:mrow></mml:math></inline-formula> when executing that workflow on several grid machines compared to
the execution on a single desktop PC.  However, as these are different workflows, no
comparison between the type of computing resources can be made from these performance
measures.</p>
</sec>
<sec id="Ch1.S4.SS3">
  <title>Usage scenarios</title>
      <p>Different usage scenarios are commonly found in meteorology. For choosing the right type
of computing system, several issues need to be taken into account. Only above a certain
workflow scale is it worth the effort to move  away from a local machine.
grids usually have a steep learning curve, clouds offer simple (web) interfaces and local
clusters are somewhere in the middle.  To make the most out of cloud computing (and to
some extent out of grid computing), it is best to have a workflow which can be split into
small, independent components.</p>
      <p>In an “operational scenario with frequent invocations”, either clouds and grids
might be suitable depending on the amount of data transferred and the complexity of the
model.  Time critical data dissemination of forecast products can be sped up with (data)
grids.  “Operational scenarios with infrequent invocations” might benefit from
using grid or even cloud computing, avoiding the need for a local cluster. Examples are
recalculation/reanalysis of seasonal/climate simulations or updating of model output
statistics (MOS) equations.  One important consideration for operational workflows is the
scheduling latency, i.e. the time between submitting a job and its actual execution.
<xref ref-type="bibr" rid="bib1.bibx3" id="normal.46"/> and <xref ref-type="bibr" rid="bib1.bibx22" id="normal.47"/> show  median latencies of 100 s for Enabling Grids for  E-Science in Europe (EGEE) grid,
but with frequent outliers upwards to 30 min and more (RainCloud 10–120 s).</p>
      <p>For a “research scenario with bursts of high activity with many small tasks”, cloud
computing fits perfectly. The costs are fully controllable and only little set-up is
required. Examples of such use cases include parameter studies with simple models or
computation of MOS. If a lot of data transfer is needed, grid
computing is the better alternative.  “Research applications with big, long running, data
intensive simulations” such as high-resolution complex models are best run on grids or
local clusters.</p>
</sec>
</sec>
<sec id="Ch1.S5" sec-type="conclusions">
  <title>Conclusions</title>
      <p>We successfully deployed meteorological applications on distributed computing
infrastructure of both grids and clouds. Our meteorological applications range from
a complex atmospheric limited-area model to a simplified model of orographic precipitation.
Adhering to some limitations/considerations, distributed computing can cater to both.</p>
      <p>A consideration to be taken into account for both concepts is security. With grids,
it is relatively easy to determine users and potential access to data as all resources and
locations are known. With clouds, this is nearly impossible/impractical to do this and potential
breaches are hard to detect.</p>
      <p>If the grid is seen as an agglomeration of individual supercomputers, complex parallelised
models are simple to deploy and efficient to use in a research setting.
The compute power is usually substantially larger than what a single institution could
afford. However, in an operational setting the immediate  availability of resources might
not be a given. This is an issue that needs to be addressed in advance. For data storage and
transfer, e.g.  dissemination of forecasts, grids are a powerful tool.</p>
      <p>Taking grid as a structure, workflows involving MPI are not simple to
exploit. As with clouds, it is much more effective to deploy an application consisting
of serial jobs with as little interprocess communication as possible.</p>
      <p>Heterogeneity of the underlying hardware cannot be ignored for grid computing as quality
tests showed <xref ref-type="bibr" rid="bib1.bibx35" id="paren.48"/>. Differences arising solely based on the used hardware
might influence very sensitive applications. However, this is application-specific and
needs to be tested for each set-up.</p>
      <p>The set-up and access to cloud infrastructure is a lot simpler and involves less effort than
participation in a grid project. Grids require hardware and more complex software to
access, whereas access to clouds is usually kept as simple as possible.</p>
      <p>Cloud (commercial) computing is very effective and cost saving tool for certain
meteorological applications. Individual projects with high-burst needs or an operational
setting with a simple model are two examples. Elasticity, i.e. access to a larger
scale of resources, is one of the biggest advantages of clouds. Undetermined or volatile
needs can be easily catered for. One option is to use clouds to baseline workflow requirements
and then build and move to a correctly sized in-house cluster/set-up based on this
prototyping.</p>
      <p>Disadvantages of clouds include above-mentioned security issues,
but one of the biggest problems for meteorological applications is data transfer.
Transfer to and from the cloud and within the cloud infrastructure is considerably slower
than for a dedicated cluster set-up or grids. Recently new instance types for massively
parallel computing have been emerging, (e.g. Amazon), but high computation applications
with only modest data needs are best suited for most clouds.</p>
      <p>Private clouds remove some of the disadvantages of public clouds, security and data
transfer are the most notable ones. However, using private clouds also removes the
advantage of not needing  hardware and system administration. We used a small private
cloud to develop our workflow before going full scale on Amazon AWS with our operational
set-up.</p>
      <p>In a meteorological research setting with specialised software, clouds offer
a flexible system with full control over operating system, installed software and
libraries. Grids on the other hand are managed on individual grid sites and are more
strict and less flexible. The same is true for customer service. clouds offer one
contact for all problems and offer (paid) premium support as opposed to having to contact
each system administration for every grid site.</p>
      <p>In conclusion, both concepts are an alternative or a supplement to self-hosted high-performance computing infrastructure.  We have laid out guidelines with which to decide
whether one's own application is suitable to either or both alternatives.</p>
</sec>

      
      </body>
    <back><ack><title>Acknowledgements</title><p>This research is supported by AustrianGrid, funded by the bm:bwk (Federal Ministry for
Education, Science and Culture) BMBWK GZ 4003/2-VI/4c/2004 (MeteoAG),
GZ BMWF  10.220/002-II/10/2007 (MeteoAG2), and Standortagentur Tirol:
RainCloud.<?xmltex \hack{\newline}?><?xmltex \hack{\newline}?>Edited by:  S. Unterstrasser</p></ack><ref-list>
    <title>References</title>

      <ref id="bib1.bibx1"><label>Allcock et al.(2002)Allcock, Bester, Bresnahan, Chervenak, Foster,
Kesselman, Meder, Nefedova, Quesnel, and Tuecke</label><mixed-citation>
Allcock, B., Bester, J., Bresnahan, J., Chervenak, A. L., Foster, I. T.,
Kesselman, C., Meder, S., Nefedova, V., Quesnel, D., and Tuecke, S.: Data
management and transfer in high-performance computational Grid
environments, Parallel Comput., 28, 749–771, 2002.</mixed-citation></ref>
      <ref id="bib1.bibx2"><label>Barstad and Schüller(2011)</label><mixed-citation>
Barstad, I. and Schüller, F.: An extension of Smith's linear theory of
orographic precipitation: introduction of vertical layers, J. Atmos. Sci.,
68, 2695–2709, 2011.</mixed-citation></ref>
      <ref id="bib1.bibx3"><label>Berger et al.(2009)Berger, Zangerl, and Fahringer</label><mixed-citation>
Berger, M., Zangerl, T., and Fahringer, T.: Analysis of overhead and waiting
time in the EGEE production Grid, in: Proceedings of the Cracow Grid
Workshop, 2008, 287–294, 2009.</mixed-citation></ref>
      <ref id="bib1.bibx4"><label>Berriman et al.(2013)Berriman, Deelman, Juve, Rynge, and Vöckler</label><mixed-citation>Berriman, G. B., Deelman, E., Juve, G., Rynge, M., and Vöckler, J.-S.:
The application of cloud computing to scientific workflows: a study of cost
and performance, Philos. T. R. Soc. A., 371, 20120066, <ext-link xlink:href="http://dx.doi.org/10.1098/rsta.2012.0066" ext-link-type="DOI">10.1098/rsta.2012.0066</ext-link>,
2013.</mixed-citation></ref>
      <ref id="bib1.bibx5"><label>Blanco et al.(2013)Blanco, Cofino, and Fernandez-Quiruelas</label><mixed-citation>
Blanco, C., Cofino, A. S., and Fernandez-Quiruelas, V.: WRF4SG: a scientific
gateway for climate experiment workflows, Geophys. Res. Abstr.,
EGU2013-11535, EGU General Assembly 2013, Vienna, Austria, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx6"><label>Bosa and Schreiner(2009)</label><mixed-citation>Bosa, K. and Schreiner, W.: A supercomputing API for the Grid, in:
Proceedings of 3rd Austrian Grid Symposium 2009, edited by: Volkert, J.,
Fahringer, T., Kranzlmuller, D., Kobler, R., and Schreiner, W., Austrian
Grid, Austrian Computer Society (OCG), 38–52, available at:
<uri>http://www.austriangrid.at/index.php?id=symposium</uri> (last access: 17 June 2015), 2009.</mixed-citation></ref>
      <ref id="bib1.bibx7"><label>Bote-Lorenzo et al.(2004)Bote-Lorenzo, Dimitriadis, and Sanchez</label><mixed-citation>
Bote-Lorenzo, M. L., Dimitriadis, Y. A., and Sanchez, E. G. A.: Grid
Characteristics and Uses: A Grid Definition, Vol. 2970, Springer, Berlin,
Heidelberg, 2004.</mixed-citation></ref>
      <ref id="bib1.bibx8"><label>Bougeault et al.(2010)Bougeault, Toth, Bishop, Brown, Burridge, Chen, Ebert, Fuentes, Hamill, Mylne, Nicolau, Paccagnella, Park, Parsons, Raoult, Schuster, Dias, Swinbank, Takeuchi, Tennant, Wilson, and Worley</label><mixed-citation>
Bougeault, P., Toth, Z., Bishop, C., Brown, B., Burridge, D., Chen, D. H.,
Ebert, B., Fuentes, M., Hamill, T. M., Mylne, K., Nicolau, J.,
Paccagnella, T., Park, Y.-Y., Parsons, D., Raoult, B., Schuster, D.,
Dias, P. S., Swinbank, R., Takeuchi, Y., Tennant, W., Wilson, L., and
Worley, S.: The THORPEX interactive grand global ensemble, B. Am. Meteorol.
Soc., 91, 1059–1072, 2010.</mixed-citation></ref>
      <ref id="bib1.bibx9"><label>Catteddu(2010)</label><mixed-citation>
Catteddu, D.: Cloud computing: benefits, risks and recommendations for
information security, in: Web Application Security SE – 9, edited by:
Serrão, C., Aguilera Díaz, V., and Cerullo, F., Vol. 72 of
Communications in Computer and Information Science, Springer, Berlin,
Heidelberg, p. 17, 2010.</mixed-citation></ref>
      <ref id="bib1.bibx10"><label>Cody et al.(2008)Cody, Sharman, Rao, and Upadhyaya</label><mixed-citation>
Cody, E., Sharman, R., Rao, R. H., and Upadhyaya, S.: Security in grid
computing: a review and synthesis, Decis. Support Syst., 44, 749–764, 2008.</mixed-citation></ref>
      <ref id="bib1.bibx11"><label>Cotton et al.(2003)Cotton, Sr., Walko, Liston, Tremback, Jiang, McAnelly, Harrington, Nicholls, Carrio, and McFadden</label><mixed-citation>
Cotton, W. R., Pielke Sr., R. A., Walko, R. L., Liston, G. E.,
Tremback, C. J., Jiang, H., McAnelly, R. L., Harrington, J. Y.,
Nicholls, M. E., Carrio, G. G., and McFadden, J. P.: RAMS 2001: Current
status and future directions, Meteorol. Atmos. Phys., 82, 5–29, 2003.</mixed-citation></ref>
      <ref id="bib1.bibx12"><label>Deelman et al.(2008)Deelman, Singh, Livny, Berriman, and Good</label><mixed-citation>
Deelman, E., Singh, G., Livny, M., Berriman, B., and Good, J.: The cost of
doing science on the cloud: the montage example, in: Proceedings of the 2008
ACM/IEEE conference on Supercomputing, p. 50, 2008.</mixed-citation></ref>
      <ref id="bib1.bibx13"><label>Evangelinos and Hill(2008)</label><mixed-citation>
Evangelinos, C. and Hill, C.: Cloud computing for parallel scientific HPC
applications: feasibility of running coupled atmosphere-ocean climate models
on Amazon's EC2, Ratio, 2, 2–34, 2008.</mixed-citation></ref>
      <ref id="bib1.bibx14"><label>Feng et al.(2011)Feng, Zhang, Zhang, and Xu</label><mixed-citation>
Feng, D.-G., Zhang, M., Zhang, Y., and Xu, Z.: Study on cloud computing
security, J. Softw., 22, 71–83, 2011.</mixed-citation></ref>
      <ref id="bib1.bibx15"><label>Fernández-Quiruelas et al.(2011)Fernández-Quiruelas, Fernández, Cofiño, Fita, and Gutiérrez</label><mixed-citation>
Fernández-Quiruelas, V., Fernández, J., Cofiño, A., Fita, L., and
Gutiérrez, J.: Benefits and requirements of grid computing for climate
applications. An example with the community atmospheric model, Environ.
Modell. Softw., 26, 1057–1069, 2011.</mixed-citation></ref>
      <ref id="bib1.bibx16"><label>Foster and Kesselman(2003)</label><mixed-citation>
Foster, I. and Kesselman, C.: The Grid 2: Blueprint for a New Computing
Infrastructure, Morgan Kaufmann Publishers Inc., San Francisco, CA, USA, 2003.</mixed-citation></ref>
      <ref id="bib1.bibx17"><label>Foster et al.(2008)Foster, Zhao, Raicu, and Lu</label><mixed-citation>
Foster, I., Zhao, Y., Raicu, I., and Lu, S.: Cloud Computing and Grid
Computing 360-Degree Compared, 2008 Grid Computing Environments Workshop,
1–10, 2008.</mixed-citation></ref>
      <ref id="bib1.bibx18"><label>Foster and Kesselman(2004)</label><mixed-citation>
Foster, I. T. and Kesselman, C.: The Grid: Blueprint for a New Computing
Infrastructure, 2nd Edn., Morgan Kaufmann, Amsterdam, 2004.</mixed-citation></ref>
      <ref id="bib1.bibx19"><label>Greenberg and Hamilton(2008)</label><mixed-citation>
Greenberg, A. and Hamilton, J.: The cost of a cloud: research problems in
data center networks, ACM SIGCOMM Computer Communication Review, 39, 68–73,
2008.</mixed-citation></ref>
      <ref id="bib1.bibx20"><label>Guest et al.(2012)Guest, Aloisio, and Kenway</label><mixed-citation>
Guest, M., Aloisio, G., and Kenway, R.: The scientific case for HPC in
Europe 2012–2020, tech. report, PRACE, 2012.</mixed-citation></ref>
      <ref id="bib1.bibx21"><label>Hamdaqa and Tahvildari(2012)</label><mixed-citation>
Hamdaqa, M. and Tahvildari, L.: Cloud computing uncovered: a research landscape, Adv. Comput., 86, 41–85, 2012.</mixed-citation></ref>
      <ref id="bib1.bibx22"><label>Lingrand and Montagnat(2009)</label><mixed-citation>
Lingrand, D. and Montagnat, J.: Analyzing the EGEE production grid workload:
application to jobs submission optimization, Lect. Notes Comput. Sc., 5798,
37–58, 2009.</mixed-citation></ref>
      <ref id="bib1.bibx23"><label>Mell and Grance(2011)</label><mixed-citation>Mell, P. and Grance, T.: The NIST definition of cloud computing
recommendations of the National Institute of Standards and Technology,
Special Publication 800–145, NIST, Gaithersburg, available at:
<uri>http://csrc.nist.gov/publications/nistpubs/800-145/SP800-145.pdf</uri> (last
access: 9 February 2015), 2011.</mixed-citation></ref>
      <ref id="bib1.bibx24"><label>Nadeem and Fahringer(2009)</label><mixed-citation>Nadeem, F. and Fahringer, T.: Predicting the execution time of grid workflow
applications through local learning, in: High Performance Computing
Networking, Proceedings of the Conference on Storage and Analysis, 1, 1–12,
<ext-link xlink:href="http://dx.doi.org/10.1145/1654059.1654093" ext-link-type="DOI">10.1145/1654059.1654093</ext-link>, 2009.</mixed-citation></ref>
      <ref id="bib1.bibx25"><label>Nadeem et al.(2007)Nadeem, Prodan, and Fahringer</label><mixed-citation>Nadeem, F., Prodan, R., and Fahringer, T.: Optimizing Performance of
Automatic Training Phase for Application Performance Prediction in the Grid,
in: High Performance Computing and Communications, Third International
Conference, HPCC 2007, Houston, USA, 26–28 September 2007, 309–321, 2007.
 </mixed-citation></ref><?xmltex \hack{\newpage}?>
      <ref id="bib1.bibx26"><label>Nefedova et al.(2006)Nefedova, Jacob, Foster, Liu, Liu, Deelman, Mehta, Su, and Vahi</label><mixed-citation>Nefedova, V., Jacob, R., Foster, I., Liu, Z., Liu, Y., Deelman, E.,
Mehta, G., Su, M.-H., and Vahi, K.: Automating climate science: large
ensemble simulations on the TeraGrid with the GriPhyN virtual data system,
e-science, 0, 32, <ext-link xlink:href="http://dx.doi.org/10.1109/E-SCIENCE.2006.261116" ext-link-type="DOI">10.1109/E-SCIENCE.2006.261116</ext-link>, 2006.</mixed-citation></ref>
      <ref id="bib1.bibx27"><label>Ostermann and Prodan(2012)</label><mixed-citation>
Ostermann, S. and Prodan, R.: Impact of variable priced cloud resources on
scientific workflow scheduling, in: Euro-Par 2012 Parallel Processing, edited
by: Kaklamanis, C., Papatheodorou, T., and Spirakis, P., Vol. 7484 of Lecture
Notes in Computer Science, Springer, Berlin, Heidelberg, 350–362, 2012.</mixed-citation></ref>
      <ref id="bib1.bibx28"><label>Ostermann et al.(2008)Ostermann, Plankensteiner, Prodan, Fahringer, and Iosup</label><mixed-citation>
Ostermann, S., Plankensteiner, K., Prodan, R., Fahringer, T., and Iosup, A.:
Workflow monitoring and analysis tool for ASKALON, in: Grid and Services
Evolution, Barcelona, Spain, 73–86, 2008.</mixed-citation></ref>
      <ref id="bib1.bibx29"><label>Patel and Shah(2005)</label><mixed-citation>
Patel, C. D. and Shah, A. J.: Cost Model for Planning, Development and
Operation of a Data Center, Technical Report HP Laboratories Palo Alto,
HPL-2005-107(R.1), 09 June 2005.</mixed-citation></ref>
      <ref id="bib1.bibx30"><label>Plankensteiner et al.(2009a)Plankensteiner, Prodan, and Fahringer</label><mixed-citation>
Plankensteiner, K., Prodan, R., and Fahringer, T.: A new fault tolerance
heuristic for scientific workflows in highly distributed environments based
on resubmission impact, in: eScience'09, 313–320, 2009a.</mixed-citation></ref>
      <ref id="bib1.bibx31"><label>Plankensteiner et al.(2009b)Plankensteiner, Vergeiner, Prodan, Mayr, and Fahringer</label><mixed-citation>
Plankensteiner, K., Vergeiner, J., Prodan, R., Mayr, G., and Fahringer, T.:
Porting LinMod to predict precipitation in the Alps using ASKALON on the
Austrian Grid, in: 3rd Austrian Grid Symposium, edited by: Volkert, J.,
Fahringer, T., Kranzlmüller, D., Kobler, R., and Schreiner, W., Vol. 269,
Austrian Computer Society, 103–114, 2009b.</mixed-citation></ref>
      <ref id="bib1.bibx32"><label>Qin et al.(2007)Qin, Wieczorek, Plankensteiner, and Fahringer</label><mixed-citation>
Qin, J., Wieczorek, M., Plankensteiner, K., and Fahringer, T.: Towards a
light-weight workflow engine in the ASKALON Grid environment, in:
Proceedings of the CoreGRID Symposium, Springer-Verlag, Rennes, France, 2007.</mixed-citation></ref>
      <ref id="bib1.bibx33"><label>Schüller(2008)</label><mixed-citation>
Schüller, F.: Grid Computing in Meteorology: Grid Computing with – and
Standard Test Cases for – A Meteorological Limited Area Model, VDM, Saarbrücken, Germany, 2008.</mixed-citation></ref>
      <ref id="bib1.bibx34"><label>Schüller and Qin(2006)</label><mixed-citation>
Schüller, F. and Qin, J.: Towards a workflow model for meteorologcial
simulations on the Austrian Grid, Austrian Computer Society, 210, 179–190,
2006.</mixed-citation></ref>
      <ref id="bib1.bibx35"><label>Schüller et al.(2007)Schüller, Qin, Nadeem, Prodan, Fahringer, and Mayr</label><mixed-citation>
Schüller, F., Qin, J., Nadeem, F., Prodan, R., Fahringer, T., and
Mayr, G.: Performance, Scalability and Quality of the Meteorological Grid
Workflow MeteoAG, Austrian Computer Society, 221, 155–165, 2007.</mixed-citation></ref>
      <ref id="bib1.bibx36"><label>Smith and Barstad(2004)</label><mixed-citation>
Smith, R. B. and Barstad, I.: A linear theory of orographic precipitation,
J. Atmos. Sci., 61, 1377–1391, 2004.</mixed-citation></ref>
      <ref id="bib1.bibx37"><label>Taylor et al.(2007)Taylor, Deelman, Gannon, Shields et al.</label><mixed-citation>
Taylor, I. J., Deelman, E., Gannon, D., and Shields, M.: Workflows for
e-Science, Springer-Verlag, London Limited, 2007.</mixed-citation></ref>
      <ref id="bib1.bibx38"><label>Todorova et al.(2010)Todorova, Syrakov, Gadjhev, Georgiev, Ganev, Prodanova, Miloshev, Spiridonov, Bogatchev, and Slavov</label><mixed-citation>
Todorova, A., Syrakov, D., Gadjhev, G., Georgiev, G., Ganev, K. G.,
Prodanova, M., Miloshev, N., Spiridonov, V., Bogatchev, A., and Slavov, K.:
Grid computing for atmospheric composition studies in Bulgaria, Earth
Sci. Inf., 3, 259–282, 2010.</mixed-citation></ref>
      <ref id="bib1.bibx39"><label>Vaquero and Rodero-Merino(2008)</label><mixed-citation>
Vaquero, L. and Rodero-Merino, L.: A break in the clouds: towards a cloud
definition, ACM SIGCOMM Computer Communication Review, 39, 50–55, 2008.</mixed-citation></ref>
      <ref id="bib1.bibx40"><label>Volkert(2004)</label><mixed-citation>
Volkert, J.: The Austrian Grid Initiative – high level extensions to Grid
middleware, in: PVM/MPI, edited by: Kranzlmüller, D., Kacsuk, P., and
Dongarra, J. J., Vol. 3241 of Lecture Notes in Computer Science, Springer,
p. 5, 2004.</mixed-citation></ref>
      <ref id="bib1.bibx41"><label>Williams et al.(2009)Williams, Drach, Ananthakrishnan, Foster, Fraser, Siebenlist, Bernholdt, Chen, Schwidder, Bharathi, Chervenak, Schuler, Su, Brown, Cinquini, Fox, Garcia, Middleton, Strand, Wilhelmi, Hankin, Schweitzer, Jones, Shoshani, and Sim</label><mixed-citation>Williams, D. N., Drach, R., Ananthakrishnan, R., Foster, I. T., Fraser, D.,
Siebenlist, F., Bernholdt, D. E., Chen, M., Schwidder, J., Bharathi, S.,
Chervenak, a. L., Schuler, R., Su, M., Brown, D., Cinquini, L., Fox, P.,
Garcia, J., Middleton, D. E., Strand, W. G., Wilhelmi, N., Hankin, S.,
Schweitzer, R., Jones, P., Shoshani, A., and Sim, A.: The Earth System Grid:
enabling access to multimodel climate simulation data, B. Am. Meteorol.
Soc., 90, 195–205, 2009.
 </mixed-citation></ref><?xmltex \hack{\newpage}?>
      <ref id="bib1.bibx42"><label>Zhao and Sakellariou(2003)</label><mixed-citation>
Zhao, H. and Sakellariou, R.: An experimental investigation into the rank
function of the heterogeneous earliest finish time scheduling algorithm, in:
Euro-Par Conference, 189–194, 2003.</mixed-citation></ref>

  </ref-list><app-group content-type="float"><app><title/>

    </app></app-group></back>
    </article>
