<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing with OASIS Tables v3.0 20080202//EN" "journalpub-oasis3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:oasis="http://docs.oasis-open.org/ns/oasis-exchange/table" dtd-version="3.0">
  <front>
    <journal-meta><journal-id journal-id-type="publisher">GMD</journal-id><journal-title-group>
    <journal-title>Geoscientific Model Development</journal-title>
    <abbrev-journal-title abbrev-type="publisher">GMD</abbrev-journal-title><abbrev-journal-title abbrev-type="nlm-ta">Geosci. Model Dev.</abbrev-journal-title>
  </journal-title-group><issn pub-type="epub">1991-9603</issn><publisher>
    <publisher-name>Copernicus Publications</publisher-name>
    <publisher-loc>Göttingen, Germany</publisher-loc>
  </publisher></journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.5194/gmd-10-4619-2017</article-id><title-group><article-title>A data model of the Climate and Forecast metadata
conventions (CF-1.6) with a software implementation (cf-python
v2.1)</article-title>
      </title-group><?xmltex \runningtitle{A CF-1.6 data model}?><?xmltex \runningauthor{D. Hassell et al.}?>
      <contrib-group>
        <contrib contrib-type="author" corresp="yes" rid="aff1">
          <name><surname>Hassell</surname><given-names>David</given-names></name>
          <email>david.hassell@ncas.ac.uk</email>
        <ext-link>https://orcid.org/0000-0001-5106-7502</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1 aff2">
          <name><surname>Gregory</surname><given-names>Jonathan</given-names></name>
          
        <ext-link>https://orcid.org/0000-0003-1296-8644</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff3">
          <name><surname>Blower</surname><given-names>Jon</given-names></name>
          
        <ext-link>https://orcid.org/0000-0002-5014-1026</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1">
          <name><surname>Lawrence</surname><given-names>Bryan N.</given-names></name>
          
        <ext-link>https://orcid.org/0000-0001-9262-7860</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff4">
          <name><surname>Taylor</surname><given-names>Karl E.</given-names></name>
          
        <ext-link>https://orcid.org/0000-0002-6491-2135</ext-link></contrib>
        <aff id="aff1"><label>1</label><institution>National Centre for Atmospheric Science, Department of
Meteorology, University of Reading, Reading, UK</institution>
        </aff>
        <aff id="aff2"><label>2</label><institution>Met Office Hadley Centre, Exeter, Exeter, UK</institution>
        </aff>
        <aff id="aff3"><label>3</label><institution>Institute for Environmental Analytics, University of Reading, Reading, UK</institution>
        </aff>
        <aff id="aff4"><label>4</label><institution>Program for Climate Model Diagnosis and Intercomparison,
Lawrence Livermore National Laboratory, Livermore, CA, USA</institution>
        </aff>
      </contrib-group>
      <author-notes><corresp id="corr1">David Hassell (david.hassell@ncas.ac.uk)</corresp></author-notes><pub-date><day>19</day><month>December</month><year>2017</year></pub-date>
      
      <volume>10</volume>
      <issue>12</issue>
      <fpage>4619</fpage><lpage>4646</lpage>
      <history>
        <date date-type="received"><day>29</day><month>June</month><year>2017</year></date>
           <date date-type="rev-request"><day>10</day><month>July</month><year>2017</year></date>
           <date date-type="rev-recd"><day>31</day><month>October</month><year>2017</year></date>
           <date date-type="accepted"><day>2</day><month>November</month><year>2017</year></date>
      </history>
      <permissions>
        
        
      <license license-type="open-access"><license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p></license></permissions><self-uri xlink:href="https://gmd.copernicus.org/articles/10/4619/2017/gmd-10-4619-2017.html">This article is available from https://gmd.copernicus.org/articles/10/4619/2017/gmd-10-4619-2017.html</self-uri><self-uri xlink:href="https://gmd.copernicus.org/articles/10/4619/2017/gmd-10-4619-2017.pdf">The full text article is available as a PDF file from https://gmd.copernicus.org/articles/10/4619/2017/gmd-10-4619-2017.pdf</self-uri>
      <abstract>
    <p id="d1e135">The CF (Climate and Forecast) metadata conventions are designed to promote
the creation, processing, and sharing of climate and forecasting data using
Network Common Data Form (netCDF) files and libraries. The CF conventions
provide a description of the physical meaning of data and of their spatial
and temporal properties, but they depend on the netCDF file encoding which
can currently only be fully understood and interpreted by someone familiar
with the rules and relationships specified in the conventions documentation.
To aid in development of CF-compliant software and to capture with a minimal
set of elements all of the information contained in the CF conventions, we
propose a formal data model for CF which is independent of netCDF and
describes all possible CF-compliant data. Because such data will often be
analysed and visualised using software based on other data models, we compare
our CF data model with the ISO 19123 coverage model, the Open Geospatial
Consortium CF netCDF standard, and the Unidata Common Data Model. To
demonstrate that this CF data model can in fact be implemented, we present
cf-python, a Python software library that conforms to the model and can
manipulate any CF-compliant dataset.</p>
  </abstract>
    </article-meta>
  </front>
<body>
      

<sec id="Ch1.S1" sec-type="intro">
  <title>Introduction</title>
      <p id="d1e145">Network Common Data Form (netCDF) supports a view of data as a
collection of self-describing, portable objects that can be accessed
through standardised software libraries. For climate scientists, as
well as others, it has become a popular way to create, access, and
share array-orientated scientific data <xref ref-type="bibr" rid="bib1.bibx10 bib1.bibx12" id="paren.1"/>. In this context, “self-describing” means that
a file contains, for each data array, an associated description of
what it represents scientifically, i.e. metadata. NetCDF was
developed and is maintained at Unidata, part of the US University
Corporation for Atmospheric Research (UCAR).</p>
      <p id="d1e151">The CF (Climate and Forecast) metadata conventions
<xref ref-type="bibr" rid="bib1.bibx3" id="paren.2"><named-content content-type="post"><uri>http://cfconventions.org</uri></named-content></xref> are a set of rules for storing
geoscientific data in netCDF files, with the aims of describing the data,
enabling users to identify comparable data held in different files, and
facilitating the development of software to extract, process, analyse, and
display the data. Initially CF was developed for gridded data from climate
and forecast models of the atmosphere and ocean, but its use has subsequently
extended to other geosciences, and to observations as well as numerical
models. The use of CF is recommended where applicable by Unidata.</p>
      <p id="d1e160">CF metadata are designed to be interpretable without reference to
external tables, readable by humans, easily parsable by programmes, and
minimally redundant, which reduces the potential for
inconsistencies. The development of CF began in 1999 and has proceeded
incrementally, with new features added only when called for by common
use cases, and with consideration of how the change might impact data
producers. Archival of data is a major purpose of netCDF, for which
reason backwards compatibility is an important consideration. So far,
no backwards-incompatible change has been made to the CF conventions,
meaning that a file written with a previous release of CF would still
be compliant with the most recent version.</p>
      <p id="d1e163">In general, CF tells data producers how they can provide information they
think is important for understanding their data, but it mandates very little
metadata. Projects which recommend or require the use of the CF conventions
may of course impose additional requirements on data producers, as is done,
for instance, by the Coupled Model Intercomparison Project
(<uri>https://pcmdi.llnl.gov/mips/cmip5/CMIP5_output_metadata_requirements.pdf</uri>),
which in its fifth phase (CMIP5) serves more than 4.2 petabytes of
CF-compliant netCDF (CF-netCDF) data.</p>
      <p id="d1e170">In this paper, we present a data model which is based on the version of the CF
conventions (1.6) which was the latest release at the time of writing. (Since
then, 1.7 has been released; see Sect. <xref ref-type="sec" rid="Ch1.S7"/>.) By a “data
model”, we mean an abstract interpretation of the data, that identifies the
elements of the dataset and their scientific intent, and describes how they
are related to one another and to the real or model world from which the data
were derived. A data model is necessary because it imposes the rules,
constraints, and relationships connecting metadata to the data that are
needed to imagine how the quantities included in the dataset should be
combined and processed scientifically.</p>
      <p id="d1e175">The netCDF interface that underlies CF has an explicit data model (the yellow
layer in Fig. <xref ref-type="fig" rid="Ch1.F1"/>). CF is defined by the CF-netCDF conventions
(the blue layer). The conventions have been widely adopted and there are many
software applications that work with CF datasets, but up to now a
comprehensive CF data model has not been explicitly proposed (the Open
Geospatial Consortium is discussed in Sect. <xref ref-type="sec" rid="Ch1.S5.SS2"/>). Those writing
software to process CF-compliant files have implicitly or explicitly adopted
data models to serve their own needs, not necessarily considering the whole
of the CF convention. Possible CF data models may differ regarding concepts
which are more abstract than the storage syntax of netCDF files, and which
are therefore not spelled out by the CF convention but become relevant when
the data are manipulated or visualised. Divergent interpretations of the data
can lead to misunderstandings, inconsistencies, and inefficiencies, and impair
the linking of independently developed software tools which might be needed
together for the analysis of CF data.</p>
      <p id="d1e182">Our aim is to create an explicit data model for CF (the green layer in
Fig. <xref ref-type="fig" rid="Ch1.F1"/>) to provide an interpretation of the conceptual
structure of CF which is consistent, comprehensive, and as far as possible
independent of the netCDF data model (the yellow layer). We believe that an
explicit comprehensive data model will lead to the CF conventions being
better understood, will provide guidance during the development of future
extensions to the CF conventions, and will help software developers to design
CF-compliant data-processing applications and to build interfaces to other
explicit data models.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F1"><caption><p id="d1e189">The benefits of having a CF data model. The CF-netCDF
conventions (CN) rely on the netCDF data model (NC) and, at
present, a software application is forced to make its own
interpretation of the CF-netCDF conventions – an interpretation
that is likely to be different from that of other
applications. The aim of this paper is to propose a comprehensive
CF data model that provides a consistent interpretation of the
conventions, thereby facilitating compatibility across
applications that adopt it. This by no means precludes a variety
of software implementations, as there is still considerable
flexibility in mapping data model elements onto the data objects
needed for a particular application.</p></caption>
        <?xmltex \igopts{width=210.550394pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/10/4619/2017/gmd-10-4619-2017-f01.pdf"/>

      </fig>

<sec id="Ch1.S1.SS1">
  <title>Design criteria for a CF data model</title>
      <p id="d1e203">The primary requirement of a data model is that it should be able to
describe all existing and conceivable CF-compliant datasets. If we
have been successful, then software libraries that adopt our CF data
model in constructing their internal data structures will be able to
represent and manipulate any CF-compliant dataset.</p>
      <p id="d1e206">For our data model, we define a minimal set of elements that are
sufficient for accommodating all aspects of the CF conventions. We
restrict the elements of a data model to those that are explicitly
mentioned in CF, but our data model elements do not have to be
irreducible in that a data model element could describe more than one
CF entity. For example, in CF, coordinates and coordinate bounds are
distinct entities, but coordinate bounds cannot exist without
coordinates. Therefore, it makes sense in our data model to group them
into a single element.</p>
      <p id="d1e209">Similarly, while it is possible to introduce additional elements not
presently needed or used by CF, we believe this would not be desirable
because it would increase the likelihood of a data model becoming
outdated or inconsistent with future versions of CF.</p>
      <p id="d1e212">The CF data model should also be independent of the encoding, meaning
that it should not be constrained by the parts of the CF conventions
which describe explicitly how to store (i.e. encode) metadata in a
netCDF file. The virtue of this is that should netCDF ever fail to
meet the community needs, we shall already have set the groundwork for
applying CF to other file formats.</p>
</sec>
<sec id="Ch1.S1.SS2">
  <title>Layout of the paper</title>
      <p id="d1e222">In Sect. <xref ref-type="sec" rid="Ch1.S3"/>, we introduce the key elements of the CF
conventions and describe how they are encoded in netCDF files. The
relationships between the elements of the CF conventions and our proposed CF
data model are described in Sect. <xref ref-type="sec" rid="Ch1.S4"/>. This data model is
compared with other data models in Sect. <xref ref-type="sec" rid="Ch1.S5"/>, and a software
implementation is presented in Sect. <xref ref-type="sec" rid="Ch1.S6"/>. How our CF data model
and its software implementation may evolve is discussed in
Sect. <xref ref-type="sec" rid="Ch1.S7"/>, and a summary and conclusions are given in
Sect. <xref ref-type="sec" rid="Ch1.S8"/>.</p>
</sec>
</sec>
<sec id="Ch1.S2">
  <title>The netCDF data model</title>
      <p id="d1e245">The existing CF conventions are for use with netCDF files following the
netCDF “classic” data model (the yellow layer in Fig. <xref ref-type="fig" rid="Ch1.F1"/>). A
brief summary of this explicit data model is useful since the CF conventions
cannot be described without reference to elements of netCDF.</p>
      <p id="d1e250">The netCDF classic data model is described using Unified Modeling Language
(UML) in Fig. <xref ref-type="fig" rid="Ch1.F2"/>. UML provides a standard way to visualise the
components of a system and how they relate to each other. In UML, different
styles of arrows denote different types of relationship
(Table <xref ref-type="table" rid="Ch1.T1"/>). Appendix <xref ref-type="sec" rid="App1.Ch1.S1"/> provides a primer on the
subset of UML used in this paper, and is recommended for readers new to this
style of diagram.</p>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T1" specific-use="star"><caption><p id="d1e262">The UML class associations used in this paper. See
Fig. <xref ref-type="fig" rid="App1.Ch1.F1"/> for a worked example.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="2">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="left"/>
     <oasis:thead>
       <oasis:row rowsep="1">  
         <oasis:entry colname="col1">UML association</oasis:entry>  
         <oasis:entry colname="col2">Description</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>  
         <oasis:entry colname="col1"><?xmltex \igopts{width=85.358268pt}?><inline-graphic xlink:href="https://gmd.copernicus.org/articles/10/4619/2017/gmd-10-4619-2017-g01.pdf"/></oasis:entry>  
         <oasis:entry colname="col2">Class-B is a special kind of class-A</oasis:entry>
       </oasis:row>
       <oasis:row>  
         <oasis:entry colname="col1"><?xmltex \igopts{width=85.358268pt}?><inline-graphic xlink:href="https://gmd.copernicus.org/articles/10/4619/2017/gmd-10-4619-2017-g02.pdf"/></oasis:entry>  
         <oasis:entry colname="col2">Class-D is a special kind of class-E (but class-E is not shown on the diagram)</oasis:entry>
       </oasis:row>
       <oasis:row>  
         <oasis:entry colname="col1"><?xmltex \igopts{width=85.358268pt}?><inline-graphic xlink:href="https://gmd.copernicus.org/articles/10/4619/2017/gmd-10-4619-2017-g03.pdf"/></oasis:entry>  
         <oasis:entry colname="col2">Instances of class-C can be included in an instance of class-B but cannot exist independently</oasis:entry>
       </oasis:row>
       <oasis:row>  
         <oasis:entry colname="col1"><?xmltex \igopts{width=85.358268pt}?><inline-graphic xlink:href="https://gmd.copernicus.org/articles/10/4619/2017/gmd-10-4619-2017-g04.pdf"/></oasis:entry>  
         <oasis:entry colname="col2">Instances of class-F can be included in an instance of class-B or can exist independently</oasis:entry>
       </oasis:row>
       <oasis:row>  
         <oasis:entry colname="col1"><?xmltex \igopts{width=85.358268pt}?><inline-graphic xlink:href="https://gmd.copernicus.org/articles/10/4619/2017/gmd-10-4619-2017-g05.pdf"/></oasis:entry>  
         <oasis:entry colname="col2">An instance of class-B is associated with an instance of class-D but class-D is independent of class-B</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d1e351">NetCDF classic files contain data in named variables, which can be
single numbers (with no dimensions), one-dimensional arrays (vectors),
or multi-dimensional arrays, and the dimensions are declared by name in
the file. Variables can be of integer, floating point, or character
data types. Variables may have attributes, of any data type,
attached. Attributes can have a single value or consist of a
one-dimensional array. NetCDF files also have “global” file
attributes which provide information about the dataset as a
whole. NetCDF library software has functions to define dimensions,
variables, and attributes, and write and read data.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F2"><caption><p id="d1e357">Key components of the netCDF classic data model
(corresponding to the yellow “NC” layer in
Fig. <xref ref-type="fig" rid="Ch1.F1"/>) described using UML
(Appendix <xref ref-type="sec" rid="App1.Ch1.S1"/>). Files consist of global attributes,
dimensions, and variables. Variables contain attributes and data,
and attributes also contain data. Variables, attributes, and
dimensions all contain properties, such as a “name” which
identifies them in the file. A data array has a data type for all
of its elements (e.g. “double” for 64-bit floating point
numbers).</p></caption>
        <?xmltex \igopts{width=221.931496pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/10/4619/2017/gmd-10-4619-2017-f02.pdf"/>

      </fig>

      <p id="d1e370">It is important to appreciate that netCDF itself has no other semantics; for
example, while coordinates can be stored in variables and described by
attributes, the meanings of these variables and attributes and relationships
between them and the variables containing data are not defined by netCDF.
NetCDF makes no prescriptions or restrictions regarding the type of metadata
which may be stored in the simple data structures that it offers. This
flexibility is intended to provide a scope for users and scientific disciplines
to develop their own conventions for encoding semantics so that datasets are
sufficiently described by those who create them and that they remain valid
for those who store and use them. CF is an example of this.</p>
      <p id="d1e373"><?xmltex \hack{\newpage}?>The original classic netCDF data model has been “enhanced” with the
addition of several new features, including the ability to organise variables
in hierarchical groups. Here, we adopt only one of the new features: we regard
the character string as a data type, whereas the classic model treats strings
as arrays of individual characters. Logically, these treatments are
equivalent, but because strings are easier to manipulate in software codes,
it is very likely that they will become a part of CF in the future.</p>
</sec>
<sec id="Ch1.S3">
  <title>The CF conventions</title>
      <p id="d1e383">In this section, we briefly describe how the most important of the CF
conventions are encoded in netCDF files. We do not consider any conventions
accepted after version 1.6 (see Sect. <xref ref-type="sec" rid="Ch1.S7"/> for a discussion on the
inclusion of newer conventions). The comprehensive definition of CF, which
includes many extra details, can be found on the CF website
(<uri>http://cfconventions.org</uri>). A netCDF file that encodes an example of
each aspect of the conventions that we describe is shown in Fig. <xref ref-type="fig" rid="Ch1.F3"/>.
This example file is presented in Common Data Language (CDL;
<xref ref-type="bibr" rid="bib1.bibx11" id="altparen.3"/>) – a human-readable notation for netCDF data that is
easily produced by netCDF library software.</p>
      <p id="d1e396"><?xmltex \hack{\newpage}?>In order to reduce storage occupied by netCDF files, the CF conventions
provide for lossy packing of data values and non-lossy compression
eliminating missing data values. Although practically valuable, these
mechanisms do not affect our conceptual data model and so we have chosen to
describe them in Appendix <xref ref-type="sec" rid="App1.Ch1.S2"/> rather than in this section.</p>
<sec id="Ch1.S3.SS1">
  <?xmltex \opttitle{Conventions from the netCDF user guide\hack{\break} and COARDS}?><title>Conventions from the netCDF user guide<?xmltex \hack{\break}?> and COARDS</title>
      <p id="d1e410">Unidata provides a netCDF user guide (NUG) <xref ref-type="bibr" rid="bib1.bibx11" id="paren.4"/>, in which they
propose some netCDF conventions. CF makes use of these conventions, so we
regard them as part of the CF conventions as well (the blue layer in
Fig. <xref ref-type="fig" rid="Ch1.F1"/>), and they are illustrated in Fig. <xref ref-type="fig" rid="Ch1.F3"/>. A
one-dimensional variable that has the same name as its dimension (<monospace>z</monospace>,
<monospace>x</monospace>, and <monospace>y</monospace> in Fig. <xref ref-type="fig" rid="Ch1.F3"/>) is regarded as a coordinate
variable. Since CF introduces other types of variables for coordinate data, we
sometimes refer to the kind defined in the netCDF user guide as a coordinate
variable “in the NUG sense”. (By the phrase “coordinate variable”, the CF
standard document consistently means “coordinate variable in the NUG
sense”.) We describe the various kinds of coordinate variables in more
detail later (Sect. <xref ref-type="sec" rid="Ch1.S3.SS3"/>). The netCDF user guide proposes a number of
conventional attributes; some of these are explicitly included in CF. The
guide also contains a statement that it is always allowable to make use of
attributes that are not standardised by CF. Some of the Unidata attributes
recognised by CF contain scientific metadata, e.g. <monospace>source</monospace>, for the
provenance of the data, and the <monospace>units</monospace> of data values (further
discussed in Sect. <xref ref-type="sec" rid="Ch1.S3.SS8"/>). Others concern the encoding in netCDF
files, e.g. <monospace>Conventions</monospace>, stating the netCDF conventions to which
the file adheres (line 73 of Fig. <xref ref-type="fig" rid="Ch1.F3"/>) and the specification of a
missing data value with <monospace>missing_value</monospace> (line 53 of Fig. <xref ref-type="fig" rid="Ch1.F3"/>,
indicating that values of <inline-formula><mml:math id="M1" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>10<inline-formula><mml:math id="M2" display="inline"><mml:msup><mml:mi/><mml:mn mathvariant="normal">30</mml:mn></mml:msup></mml:math></inline-formula> correspond to cells for which no data
are available, such as ocean points for a quantity measured only over land).</p>
      <p id="d1e469">When originally conceived, CF was an extension of the pre-existing COARDS
(Cooperative Ocean/Atmosphere Research Data Service) netCDF conventions
(<uri>http://www.ferret.noaa.gov/noaa_coop/coop_cdf_profile.html</uri>). For the
sake of backward compatibility of datasets, although CF is now much more
comprehensive and flexible than COARDS, CF explicitly upholds some COARDS
conventions, which we therefore regard also as part of CF.</p>

      <?xmltex \floatpos{p}?><fig id="Ch1.F3" specific-use="star"><caption><p id="d1e477">A CDL representation of the CF-netCDF file used for
examples in Sect. <xref ref-type="sec" rid="Ch1.S3"/> and to demonstrate the
software implementation in Sect. <xref ref-type="sec" rid="Ch1.S6"/>. Data values
have been omitted for brevity. Each line has a comment on the
right-hand side (beginning with <monospace>//</monospace>) that gives the line
number and notes the data model constructs
(Sect. <xref ref-type="sec" rid="Ch1.S4"/>) which correspond to netCDF
dimensions, variables, and attributes.</p></caption>
          <?xmltex \igopts{width=412.564961pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/10/4619/2017/gmd-10-4619-2017-f03.pdf"/>

        </fig>

</sec>
<sec id="Ch1.S3.SS2">
  <title>The data and the domain</title>
      <p id="d1e501">The overarching purpose of the conventions is to provide conforming datasets
with sufficient metadata that they are self-describing, in the sense that
each variable in the file has an associated description of what it
represents, and that each value can be located (usually in space and time).
To meet this objective, we define a data variable <inline-formula><mml:math id="M3" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula> (which might, for
example, represent air temperature), over a domain <inline-formula><mml:math id="M4" display="inline"><mml:mi>d</mml:mi></mml:math></inline-formula>,
            <disp-formula id="Ch1.E1" content-type="numbered"><mml:math id="M5" display="block"><mml:mrow><mml:mi>V</mml:mi><mml:mo>≡</mml:mo><mml:mi>V</mml:mi><mml:mo>(</mml:mo><mml:mi>d</mml:mi><mml:mo>)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
          where <inline-formula><mml:math id="M6" display="inline"><mml:mi>d</mml:mi></mml:math></inline-formula> represents a set of discrete “locations” in what generally would
be a multi-dimensional space, either in the real world or in a model's
simulated world. Thus, <inline-formula><mml:math id="M7" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula> is a function of all its independent dimensions.
For example, a variable that is a function of physical location alone would
have a three-dimensional discretised domain,
            <disp-formula id="Ch1.E2" content-type="numbered"><mml:math id="M8" display="block"><mml:mrow><mml:mi>d</mml:mi><mml:mo>≡</mml:mo><mml:mi>d</mml:mi><mml:mo>(</mml:mo><mml:mi>z</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
          comprising discretised axes of height (<inline-formula><mml:math id="M9" display="inline"><mml:mi>z</mml:mi></mml:math></inline-formula>), latitude (<inline-formula><mml:math id="M10" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula>), and longitude
(<inline-formula><mml:math id="M11" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula>) (Fig. <xref ref-type="fig" rid="Ch1.F4"/>). In CF, the domain may have fewer than
three spatial axes, and it may also have any number of non-spatial axes, as
in the common case of a variable that is a function time (<inline-formula><mml:math id="M12" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>). A CF-netCDF
file may contain <inline-formula><mml:math id="M13" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula> data variables and <inline-formula><mml:math id="M14" display="inline"><mml:mi>M</mml:mi></mml:math></inline-formula> domains, where <inline-formula><mml:math id="M15" display="inline"><mml:mrow><mml:mi>M</mml:mi><mml:mo>≤</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:math></inline-formula>.
Conversely, this means that a given domain may have one or more data
variables defined at each of its locations. For instance, there could be
values for both air temperature and relative humidity at each location in the
domain <inline-formula><mml:math id="M16" display="inline"><mml:mrow><mml:mi>d</mml:mi><mml:mo>(</mml:mo><mml:mi>z</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>.</p>
      <p id="d1e660">In CF-netCDF, the values and the description of <inline-formula><mml:math id="M17" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula> are stored in a netCDF
variable, called a “data variable”. The concept of a domain is not
mentioned in the CF conventions, because it does not correspond to any single
entity in the netCDF file. The domain is stored in a number of other
variables and attributes that are linked to the data variable in various ways
defined by the conventions. For instance, <monospace>temp</monospace> is a data variable
(line 52 in Fig. <xref ref-type="fig" rid="Ch1.F3"/>) with a four-dimensional domain. Its <inline-formula><mml:math id="M18" display="inline"><mml:mi>z</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M19" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula>,
<inline-formula><mml:math id="M20" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula>, and <inline-formula><mml:math id="M21" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> dimensions each have an associated coordinate variable specifying the
location at each point along the dimension (the <inline-formula><mml:math id="M22" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> dimension being implied
by the <monospace>t</monospace> scalar coordinate variable; see Sect. <xref ref-type="sec" rid="Ch1.S3.SS3"/> for
details). Note that it is not possible to store a domain in the absence of
any data variables (i.e. <inline-formula><mml:math id="M23" display="inline"><mml:mrow><mml:mi>N</mml:mi><mml:mo>≥</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula>), because in the absence of a data
variable, CF-netCDF lacks mechanisms to associate the variables from which a
domain could be defined. This does not mean that it is disallowed to create a
dataset that contains only elements of a domain, but rather that the CF
conventions only allow for them to be interpreted collectively as a domain
when they are associated with at least one data variable (see
Sect. “Interpreting CF-netCDF files” for an
example and further discussion).</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F4"><caption><p id="d1e730">An example domain defined by three dimensions, one of which
is single valued (height).</p></caption>
          <?xmltex \igopts{width=150.799606pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/10/4619/2017/gmd-10-4619-2017-f04.pdf"/>

        </fig>

      <p id="d1e739">Within a CF-netCDF file, dimensions and coordinate variables may be used in
the definition of multiple domains, thus reducing redundancy. In our example
file, <monospace>total_wv</monospace> is a data variable containing the vertical integral
of atmospheric water vapour (line 61 in Fig. <xref ref-type="fig" rid="Ch1.F3"/>) that has a different
domain than the <monospace>temp</monospace> data variable. Its <inline-formula><mml:math id="M24" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M25" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula>, and <inline-formula><mml:math id="M26" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> dimensions
and their coordinates (see Sect. <xref ref-type="sec" rid="Ch1.S3.SS3"/>) are identical to those of
<monospace>temp</monospace>, so they need not be replicated, but it does not require the <inline-formula><mml:math id="M27" display="inline"><mml:mi>z</mml:mi></mml:math></inline-formula>
dimension.</p>
</sec>
<sec id="Ch1.S3.SS3">
  <title>Dimensions and coordinates</title>
      <p id="d1e791">NetCDF dimensions establish the size of the index space of data variables,
e.g. lines 3–5 in Fig. <xref ref-type="fig" rid="Ch1.F3"/>, which specify sizes of 106, 110, and 20
for the <inline-formula><mml:math id="M28" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M29" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula>, and <inline-formula><mml:math id="M30" display="inline"><mml:mi>z</mml:mi></mml:math></inline-formula> dimensions, respectively. Each point of the domain of
<monospace>temp</monospace> is thus defined by a unique set of three indices <inline-formula><mml:math id="M31" display="inline"><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula>,
…, 105, <inline-formula><mml:math id="M32" display="inline"><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula>, …, 109, and <inline-formula><mml:math id="M33" display="inline"><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula>, …, 19. NetCDF
coordinate variables (in the NUG sense) supply the independent variables on
which the data depend. Coordinate variables must be numeric and strictly
monotonic, so that each element has a unique value. In our example, we have
three coordinate variables, with values <inline-formula><mml:math id="M34" display="inline"><mml:mrow><mml:mi>z</mml:mi><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M35" display="inline"><mml:mrow><mml:mi>y</mml:mi><mml:mo>(</mml:mo><mml:mi>j</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, and <inline-formula><mml:math id="M36" display="inline"><mml:mrow><mml:mi>x</mml:mi><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. Each
dimension, with its coordinate variable if it has one, constitutes an axis of
the multi-dimensional space of the domain. The CF conventions quite often use
the word “axis” to refer to the physical interpretation of the dimensions
of the data.<?xmltex \hack{\newpage}?></p>
      <p id="d1e900">In many cases, each dimension of a domain can be fully described by a single,
strictly monotonic coordinate variable (e.g. time, height, latitude,
longitude). For more complicated cases, however, such as parametric vertical
coordinates (e.g. dimensionless atmosphere sigma coordinates), CF provides a
way to record how to compute, from the original dimensional
coordinates, dimensional coordinates identifying the location
of the data in physical space (in the case of sigma, the air pressure). This
information is encoded with the <monospace>standard_name</monospace> and
<monospace>formula_terms</monospace> attributes of a parametric coordinate variable, e.g. lines 14 and 17 in Fig. <xref ref-type="fig" rid="Ch1.F3"/>. The <monospace>standard_name</monospace> attribute
defines the formula for calculating the dimensional coordinates, which needs
to be looked up in the CF conventions document, and the formula_terms
attribute specifies the values of the formula's terms. The
<monospace>atmosphere_sigma_coordinate</monospace> formula specified in Fig. <xref ref-type="fig" rid="Ch1.F3"/>
calculates air pressure from
<inline-formula><mml:math id="M37" display="inline"><mml:mrow><mml:mtext>ptop</mml:mtext><mml:mo>+</mml:mo><mml:mtext>sigma</mml:mtext><mml:mo>*</mml:mo><mml:mo>(</mml:mo><mml:mtext>ps</mml:mtext><mml:mo>-</mml:mo><mml:mtext>ptop</mml:mtext><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, where the values of
“ptop” (pressure at the top of the model), “sigma” (the dimensionless
coordinates), and “ps” (surface pressure) are taken from the netCDF
variables referenced by the <monospace>formula_terms</monospace> attribute.</p>
      <p id="d1e947">CF also defines “auxiliary coordinate variables” to provide mandatory or
optional coordinate information which is additional or alternative to that
contained in the coordinate variables in the NUG sense. Auxiliary coordinate
variables can be string valued, may contain missing values, and are not
necessarily monotonic. For example, we might like to associate the
coordinates of a vertical axis with model level number as well as sigma
coordinate or to provide location information and station names for the
points in a time series (as in Fig. <xref ref-type="fig" rid="Ch1.F5"/>). An auxiliary
coordinate variable is encoded as a netCDF variable that is referenced by the
<monospace>coordinates</monospace> attribute of a data variable and spans at least one of
that data variable's dimensions, e.g. lines 27, 30, and 57 in Fig. <xref ref-type="fig" rid="Ch1.F3"/>.
Coordinate variables (in the NUG sense) and auxiliary coordinate variables
(defined by CF) rely on different semantics, and the latter is not a special
type of the former, even though they share many characteristics.</p>
      <p id="d1e957">An important and mandatory use of auxiliary coordinates is to supply latitude
and longitude locations of each point when the horizontal axes of a grid are
themselves not latitude and longitude (e.g. if they refer to a rotated North
Pole or are based on a map projection, as is the case for <inline-formula><mml:math id="M38" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M39" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula> in the
example of Fig. <xref ref-type="fig" rid="Ch1.F3"/>, and sketched in Fig. <xref ref-type="fig" rid="Ch1.F6"/>). For this
case, the latitude and longitude auxiliary coordinate variables are
two-dimensional and can be used to indirectly locate a point horizontally, so
that <inline-formula><mml:math id="M40" display="inline"><mml:mrow><mml:mi>V</mml:mi><mml:mo>=</mml:mo><mml:mi>V</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> with <inline-formula><mml:math id="M41" display="inline"><mml:mrow><mml:mtext>longitude</mml:mtext><mml:mo>=</mml:mo><mml:mtext>longitude</mml:mtext><mml:mo>(</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>,
and <inline-formula><mml:math id="M42" display="inline"><mml:mrow><mml:mtext>latitude</mml:mtext><mml:mo>=</mml:mo><mml:mtext>latitude</mml:mtext><mml:mo>(</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, as in lines 27 and 30 of
Fig. <xref ref-type="fig" rid="Ch1.F3"/>. Given this information, the data can be located in space by
generic applications even if they are ignorant of the rules used to construct
the projection. However, CF also provides for information about the grid
construction to be included in a netCDF file by defining a “grid mapping”
variable, which is referenced by the <monospace>grid_mapping</monospace> attribute of a
data variable, e.g. lines 41 and 59 in Fig. <xref ref-type="fig" rid="Ch1.F3"/>.</p>
      <p id="d1e1071">Some axes have only a single coordinate value. Regrettably, single-valued
coordinates are often omitted from metadata, although they are very useful;
for example, the time information for a field sampled at a single time, for
instance, 12:15 Z on 14 July 2015, or the level of a single-level field, e.g. air
temperature at a height of 1.5 m (Fig. <xref ref-type="fig" rid="Ch1.F4"/>). For
convenience in storing single-valued coordinates, CF defines a third type of
variable containing coordinate data, namely a “scalar coordinate variable”,
which requires less netCDF machinery than a dimension of size unity. It is a
zero-dimensional netCDF variable that is referenced by the
<monospace>coordinates</monospace> attribute of a data variable, e.g. lines 8 and 57 in
Fig. <xref ref-type="fig" rid="Ch1.F3"/>.</p>
      <p id="d1e1081">Calendar time in CF (year, month, day, hour, minute, second) is encoded with
units “time unit since reference date–time” (e.g. line 9 in
Fig. <xref ref-type="fig" rid="Ch1.F3"/>). The encoded coordinates are the elapsed times since the
reference and as such are useful for computing time differences. CF does not
use strings for time because they cannot be used for such computations, are
inconvenient to standardise, take more storage space, and cannot always be
ordered monotonically. The encoding depends on the calendar (e.g. line 10 in
Fig. <xref ref-type="fig" rid="Ch1.F3"/>), which defines the permitted values of the reference date
(year, month, and day). The format of the units string conforms to udunits
syntax (Sect. <xref ref-type="sec" rid="Ch1.S3.SS8"/>), but the udunits software supports only the
real-world Julian/Gregorian calendar and hence is not sufficient for use
with CF, which recognises a wide selection of calendars, including those for
climate models and palaeoclimate. For instance, 31 August 2003 is a valid
date in the real-world Gregorian calendar but not in the “360-day”
calendar, which has 12 30-day months. In the Gregorian calendar, 12:00 Z on
29 February 2000 is 36 583.5 days since 00:00 Z on 1 January 1900, but it is
36 058.5 days in the 360-day calendar.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F5"><caption><p id="d1e1092">Auxiliary coordinate variables store “alternative”
coordinates for dimensions.</p></caption>
          <?xmltex \igopts{width=128.037402pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/10/4619/2017/gmd-10-4619-2017-f05.pdf"/>

        </fig>

<?xmltex \hack{\newpage}?>
</sec>
<sec id="Ch1.S3.SS4">
  <title>Discrete axes and sampling geometries</title>

      <?xmltex \floatpos{t}?><fig id="Ch1.F6"><caption><p id="d1e1111">For grid axes based on a map projection, two-dimensional
auxiliary coordinate variables must be used to store longitude and
latitude values for each location (latitude–longitude lines
dashed; grid lines solid).</p></caption>
          <?xmltex \igopts{width=150.799606pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/10/4619/2017/gmd-10-4619-2017-f06.pdf"/>

        </fig>

      <p id="d1e1120">A “discrete axis” is one which is not associated with any “continuous”
coordinate or auxiliary coordinate variables. A variable is continuous along
an axis if it makes physical sense to interpolate along that axis between its
values. If that is not the case, then either there are no coordinate values or
the coordinate values are discrete indices, whose order may or may not be
meaningful. Consider, for example, an ensemble of model experiments, each of
which produces a data variable <inline-formula><mml:math id="M43" display="inline"><mml:mrow><mml:mi>V</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>,</mml:mo><mml:mi>z</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> containing air pressure as a
function of time and spatial location. It may be convenient to combine the
data variables into a single variable <inline-formula><mml:math id="M44" display="inline"><mml:mrow><mml:mi>V</mml:mi><mml:mo>(</mml:mo><mml:mi>e</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>,</mml:mo><mml:mi>z</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, where <inline-formula><mml:math id="M45" display="inline"><mml:mi>e</mml:mi></mml:math></inline-formula> is the
ensemble dimension, defining a discrete axis of ensemble members. The members
could be identified by a numeric monotonic coordinate variable with the same
name as the dimension containing a member number. Alternatively, it is common
for them to be identified by one or more strings, which we could store in
auxiliary coordinate variables with the ensemble dimension containing, for
instance, model names or experiment names. It usually would not make sense to
interpolate between ensemble members, and their order may be immaterial.</p>
      <p id="d1e1187">An important use of discrete axes in CF is to store data from a collection of
“discrete sampling geometries” (DSGs) in a single data variable. In a DSG,
the data have a lower dimensionality than the space–time domain, because
they apply to a point or path within the domain. For example, a collection of
time series of surface air temperature at meteorological stations can be
stored in a two-dimensional data variable (Fig. <xref ref-type="fig" rid="Ch1.F5"/>) with a
dimension that is the discrete axis for the stations (the “point” axis) and
a dimension for the time series values (the “time” axis). Auxiliary
coordinate variables for the discrete axis can be used to provide location
information and station names (Fig. <xref ref-type="fig" rid="Ch1.F5"/>). Although the stations
are physically located in two or three spatial dimensions, these have been
combined into the one discrete axis. Other examples of DSGs are a vertical
profile (variation along a vertical axis at a fixed time and spatial
location) and a trajectory (variation along a path through space as a
function of time). A DSG may also be called a “feature” and the type of DSG
is called its “feature type”. The feature type (point, time series,
trajectory, profile, etc.) describes the DSG and specifies the dimensionality
of the data within the space–time domain. Prior to the introduction of DSGs
in version 1.6 of CF, the concept of DSG features was implicit in the sense
that they could be inferred from the existence of a discrete axis and
well-defined coordinate and auxiliary coordinate variables, but they were not
formally described. In general, the netCDF file attribute <monospace>featureType</monospace>
specifies the DSG feature type for every data variable in the file; but in
the cases where the featureType attribute is permitted to be missing, the
feature type may be inferred from the dimensions and space–time coordinates
alone.</p>
      <p id="d1e1197">Many DSGs many be stored in one file, but they might have different
coordinates, e.g. each time series might have its own set of sampling
times (different days or hours of observation), or each profile might
have its own set of vertical levels (e.g. air pressure reported by
radiosondes). If each feature in a large collection is stored as an
individual data variable with its own dimensions and coordinate, the
file will be cumbersome. If they are combined into a single data
variable, a one-dimensional coordinate variable would need to contain
the union of all the coordinates (times, levels, etc.) required, and
the dimension of the combined data variable might be much larger than
needed for the available data, containing a lot of missing data
elements.</p>
      <p id="d1e1201">As an alternative, CF provides three other methods for storing collections of
data on DSGs, all intended to allow data with different dimensionality to be
stored in a single data variable without wasting so much space. In the
“incomplete multi-dimensional array” representation, the dimension required
for the longest feature is used for all features, so that the shorter
features must be padded with missing values; this sacrifices storage space to
achieve simplicity for reading and writing. The “contiguous ragged array”
and “indexed ragged array” representations eliminate the need for padding
and thus reduce further the storage required, but they are more complex to pack
and unpack. In the former case, each feature in the collection occupies a
contiguous block, requiring the size of each feature to be known at the time
that it is created. In the latter case, the values of each feature in the
collection are interleaved. This representation can therefore be used for
real-time data streams that contain reports from many sources, with the data
being written as they arrive. The ragged array representations are described
in more detail in Appendix <xref ref-type="sec" rid="App1.Ch1.S2"/>.</p>
      <p id="d1e1206">Because these storage methods were introduced (in CF version 1.6) at
the same time as the recognition and definition of feature types, the
two are often thought of as belonging together, but this causes
confusion. The featureType is metadata, and it refers to the physical
construction and interpretation of a DSG data variable. The three new
storage mechanisms for DSGs do not involve any new or distinct
physical concepts.</p>
</sec>
<sec id="Ch1.S3.SS5">
  <title>Bounds and cells</title>
      <p id="d1e1215">It is often necessary to know the extent of a cell as well as the
grid point location, e.g. to calculate the area of a
latitude–longitude box or the thickness of a vertical layer. If cell
bounds are not provided, then there is no default assumption about cell
sizes (an application might reasonably assume that grid points are at
the centres of non-overlapping cells, but that is not required by CF).</p>
      <p id="d1e1218">CF provides a way to attach bounds variables to any variable containing
coordinate data. A bounds variable has an extra dimension to index the
vertices of the cells. The simplest case is shown for a one-dimensional
coordinate variable in Fig. <xref ref-type="fig" rid="Ch1.F7"/>. In this case, the values <inline-formula><mml:math id="M46" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>
may be a series of successive time instants, for example, midday on 6 and
7 November, bounded by <inline-formula><mml:math id="M47" display="inline"><mml:mrow><mml:mi>b</mml:mi><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>q</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> at midnight on 6, 7, and 8 November, with <inline-formula><mml:math id="M48" display="inline"><mml:mrow><mml:mi>q</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula>. While the bounds for a one-dimensional coordinate variable of dimension
<inline-formula><mml:math id="M49" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mi>n</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> could often be stored in a vector of dimension (<inline-formula><mml:math id="M50" display="inline"><mml:mrow><mml:mi>n</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula>), CF uses <inline-formula><mml:math id="M51" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mi>n</mml:mi><mml:mo>,</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>
instead as shown because it is convenient for use with the netCDF unlimited
dimension, and because it allows cells to be non-contiguous or overlapping.
The bounds can be used to test contiguity; in the figure, cell <inline-formula><mml:math id="M52" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula> and cell
<inline-formula><mml:math id="M53" display="inline"><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> are contiguous because <inline-formula><mml:math id="M54" display="inline"><mml:mrow><mml:mi>b</mml:mi><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mi>b</mml:mi><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. For multi-dimensional
auxiliary coordinate variables, such as the two-dimensional latitude and
longitude variables illustrated above, we have to supply the coordinates of
each vertex of the polygon and contiguity can similarly be tested by
coincidence of vertices. A bounds variable is encoded as a netCDF variable
that is referenced by the <monospace>bounds</monospace> attribute of a coordinate or
auxiliary coordinate variable and spans the same dimensions (as well as the
extra dimension defining the number of vertices), e.g. lines 6, 22, and 36 in
Fig. <xref ref-type="fig" rid="Ch1.F3"/>.</p>
      <p id="d1e1374">Some applications require information about the size, shape, or location of
the cells that cannot be deduced without specialist knowledge which is not
guaranteed to be available. For example, in computing the mean of several
cell values, it is often appropriate to “weight” the values by area, but
for some grids (such as some types of spherical geodesic grids) the cell
perimeter is not uniquely defined by its vertices and so the area cannot be
inferred from the available information. For this case, CF provides cell
measures variables which contain such information and are encoded as netCDF
variables which are referenced by the <monospace>cell_measures</monospace> attribute of a
data variable and span a subset of the data variable's dimensions, e.g. lines 38 and 58 in Fig. <xref ref-type="fig" rid="Ch1.F3"/>.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F7"><caption><p id="d1e1384">A one-dimensional coordinate variable with grid points
<inline-formula><mml:math id="M55" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and cell boundaries <inline-formula><mml:math id="M56" display="inline"><mml:mrow><mml:mi>b</mml:mi><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>q</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M57" display="inline"><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:math></inline-formula>; <inline-formula><mml:math id="M58" display="inline"><mml:mrow><mml:mi>q</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula>.</p></caption>
          <?xmltex \igopts{width=122.34685pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/10/4619/2017/gmd-10-4619-2017-f07.pdf"/>

        </fig>

<?xmltex \hack{\newpage}?>
</sec>
<sec id="Ch1.S3.SS6">
  <title>Variation within cells</title>
      <p id="d1e1469">CF describes variation within cells by use of “cell methods”. By default,
it is assumed that intensive quantities apply at grid points, e.g. temperature values apply at the spatial points and instants of time specified
by their coordinates, while extensive quantities apply to the entire
grid cell, e.g. a precipitation amount (kg m<inline-formula><mml:math id="M59" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>) is an accumulation in
time. The method may be different for each axis, e.g. precipitation amount is
intensive in space even though it is extensive in time. Because the default
is not always obvious, it is recommended that the method be stated explicitly
for every axis. Non-default methods include operations such as mean, maximum,
minimum, and standard deviation. A zonal-mean variable, for instance, has a
cell methods attribute that specifies it is a mean over longitude. A
time series of daily maximum values has cell methods indicating that the
values are maxima within their cells in time. The operations recorded by cell
methods might affect more than one axis at once, e.g. for the cell methods
necessary to describe maximum of the ocean meridional overturning
stream function within a depth–latitude cell.</p>
      <p id="d1e1484">By default, the method for horizontal cells is assumed to have been evaluated
over the entire area of the cell. It is, however, possible to limit
consideration to only a portion of a cell, e.g. to record that values apply
only to the fractions of cells which are land (as opposed to sea).</p>
      <p id="d1e1487">A further use of cell methods is to characterise climatological
statistics where a series of data points represent sets of
subintervals which are not contiguous. There are three kinds to
consider:
<list list-type="order"><list-item>
      <p id="d1e1492">corresponding portions of the annual cycle in a set of years,
e.g. decadal averages for January;</p></list-item><list-item>
      <p id="d1e1496">corresponding portions of a range of days, e.g. the average
diurnal cycle in April 1997; and</p></list-item><list-item>
      <p id="d1e1500">both at once, e.g. the average winter daily minimum
temperature from the years 1961 to 1990.</p></list-item></list>
In the latter example, the bounds are 00:00 Z
1 December 1961 (beginning of the first day of the first winter) and 00:00 Z
1 March 1991 (end of the last day of the last winter), and the cell methods
indicate the values are a minimum within days, a mean over a season, and a
mean over years.<?xmltex \hack{\newpage}?></p>
      <p id="d1e1505">Cell methods are encoded in the <monospace>cell_methods</monospace> attribute of a data
variable, e.g. line 56 in Fig. <xref ref-type="fig" rid="Ch1.F3"/>. In this example, the cell method
<monospace>t: mean (interval: 1 day)</monospace> specifies that data values are means over
the time dimension (i.e. temporal averages) calculated from daily samples.
Note that this netCDF file does not contain a dimension for time, but it is
sufficient that one is implied by the time scalar coordinate variable on
line 8.</p>
</sec>
<sec id="Ch1.S3.SS7">
  <title>Ancillary data</title>
      <p id="d1e1522">When metadata to describe the data depend on location within the domain,
they are stored in independent variables called ancillary data variables. For
example, each value of an array of instrument data may have associated
measures of uncertainty or of the status of the recording instrument. An
ancillary data variable is encoded as a netCDF variable that is referenced by
the <monospace>ancillary_variables</monospace> attribute of a data variable and spans a
subset of the data variable's dimensions, e.g. lines 60 and 68 in
Fig. <xref ref-type="fig" rid="Ch1.F3"/>.</p>
</sec>
<sec id="Ch1.S3.SS8">
  <title>Units and standard name</title>
      <p id="d1e1536">A range of attributes is available, introduced by the netCDF user
guide or CF, providing metadata for interpreting the values of
individual variables or about the dataset as a whole. In this section,
we discuss the two most important of these.</p>
      <p id="d1e1539">CF requires all variables with values (data variables, coordinate variables,
etc.) to have units unless they contain dimensionless numbers or cell
boundary values. The units are specified by a string attribute (e.g. lines 9
and 55 of Fig. <xref ref-type="fig" rid="Ch1.F3"/>), formatted according to the Unidata udunits
conventions <xref ref-type="bibr" rid="bib1.bibx4" id="paren.5"/>, which support many possible units and varieties
of syntax, e.g. metre, meter, meters, m, km, cm, second, s, kelvin, K, Pa, W m-2, W/m<inline-formula><mml:math id="M60" display="inline"><mml:msup><mml:mi/><mml:mo>∧</mml:mo></mml:msup></mml:math></inline-formula>2,
kg/m2/s, 1 (or any number). Many non-SI units are also supported by udunits,
e.g. degree, degree_north, degree_N, percent (equivalent to 0.01), ppm
(equivalent to 1e-6), mbar, mile, degC, degF, hours, and days.</p>
      <p id="d1e1556">For systematic identification of the physical quantity contained in
variables, CF defines a “standard name” string attribute (e.g. lines 28,
54 and 62 of Fig. <xref ref-type="fig" rid="Ch1.F3"/>), with permissible values listed in the standard
name table (<uri>http://cfconventions.org/standard-names.html</uri>), which
includes precise definitions. The standard name table is managed by a
community process and is continually expanding – version 44 of the table,
released in May 2017, contains 2847 standard names.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F8" specific-use="star"><caption><p id="d1e1566">The relationships between CF-netCDF elements (corresponding
to the blue CN layer in Fig. <xref ref-type="fig" rid="Ch1.F1"/>) and their
corresponding netCDF variables, dimensions, and attributes (the
yellow NC layer in Fig. <xref ref-type="fig" rid="Ch1.F1"/>) described using
UML (Appendix <xref ref-type="sec" rid="App1.Ch1.S1"/>). It is useful to define an abstract
generic coordinate variable that can be used to refer to
coordinates when the their type (coordinate, auxiliary, or scalar
coordinate variable) is not an issue.
The CF convention details the mechanisms which are used in the
netCDF file to express the relationships among the CF-netCDF elements,
but these are not shown in the UML.</p></caption>
          <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/10/4619/2017/gmd-10-4619-2017-f08.pdf"/>

        </fig>

      <p id="d1e1582">CF also upholds the use of the “long name” defined by the netCDF user
guide, but this is ad hoc. In contrast, the CF standard names are
consistently constructed and documented. As CF is applicable to many areas of
geoscience, the standard names have to be more self-explanatory and
informative than would suffice for any one area. For instance, there is no
name for plain “potential temperature”, since we have to distinguish air
potential temperature and sea water potential temperature. Standard names are
often longer than the terms familiarly used by the experts in particular
discipline, because they answer the question, “What does this mean?”,
rather than the question, “What do you call this?”. For example, the
quantity often called “precipitable water” by meteorologists has the
standard name of atmosphere_mass_content_of_water_vapor. Standard names
have a detailed description which further defines parts of the name; for
example, the description of the standard name land_ice_calving_rate notes
that “land ice” means glaciers, ice caps, and ice sheets resting on
bedrock, and the land ice calving rate is the rate at which ice is lost per unit area
through calving into the ocean. Each standard name also implies particular
physical dimensions (mass, length, time, and other dimensions corresponding to
SI base units, expressed as a “canonical unit”); for example, large-scale
rainfall amount (canonical unit kg m<inline-formula><mml:math id="M61" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>), large-scale rainfall flux
(kg m<inline-formula><mml:math id="M62" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> s<inline-formula><mml:math id="M63" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>), and large-scale rainfall rate (m s<inline-formula><mml:math id="M64" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>) are all
different in CF, although they might all be vaguely referred to as
“large-scale rain”.</p>
      <p id="d1e1633">Standard names have been defined for both more general and more specific
quantities, for different applications, e.g. ocean_mixed_layer_thickness
and ocean_mixed_layer_thickness_defined_by_temperature. Some standard
names require the existence of additional metadata and/or constraints on the
values of the variables with which they are associated. For example, the
standard name of downwelling_radiance_per_unit_wavelength_in_air
requires there to be a coordinate variable storing the radiation wavelength.</p>
      <p id="d1e1636">The CF conventions use size one or scalar coordinate variables
(Sect. <xref ref-type="sec" rid="Ch1.S3.SS3"/>) and the cell_methods attribute
(Sect. <xref ref-type="sec" rid="Ch1.S3.SS6"/>) to describe some aspects of a variable, and this
means standard names do not always correspond to identities of variables in
other file formats. For instance, to describe the time-mean air temperature
at 1.5 m above the ground, air_temperature alone is the standard name;
“time-mean” is described by cell_methods and the height as a coordinate.</p>
</sec>
</sec>
<sec id="Ch1.S4">
  <title>A CF data model</title>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T2" specific-use="star"><caption><p id="d1e1653">The elements of the CF-netCDF conventions, a brief
description of each, and the section in which it is described in
more detail. The relationships to netCDF entities are shown in
Fig. <xref ref-type="fig" rid="Ch1.F8"/>.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="3">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="left"/>
     <oasis:colspec colnum="3" colname="col3" align="left"/>
     <oasis:thead>
       <oasis:row rowsep="1">  
         <oasis:entry colname="col1">CF-netCDF element</oasis:entry>  
         <oasis:entry colname="col2">Description</oasis:entry>  
         <oasis:entry colname="col3">Section</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>  
         <oasis:entry colname="col1">Data variable</oasis:entry>  
         <oasis:entry colname="col2">Scientific data discretised within a domain</oasis:entry>  
         <oasis:entry colname="col3"><xref ref-type="sec" rid="Ch1.S3.SS2"/></oasis:entry>
       </oasis:row>
       <oasis:row>  
         <oasis:entry colname="col1">Dimension</oasis:entry>  
         <oasis:entry colname="col2">Independent axis of the domain</oasis:entry>  
         <oasis:entry colname="col3"><xref ref-type="sec" rid="Ch1.S3.SS3"/></oasis:entry>
       </oasis:row>
       <oasis:row>  
         <oasis:entry colname="col1">Coordinate variable</oasis:entry>  
         <oasis:entry colname="col2">Unique coordinates for a single axis</oasis:entry>  
         <oasis:entry colname="col3"><xref ref-type="sec" rid="Ch1.S3.SS3"/></oasis:entry>
       </oasis:row>
       <oasis:row>  
         <oasis:entry colname="col1">Auxiliary coordinate variable</oasis:entry>  
         <oasis:entry colname="col2">Additional or alternative coordinates for any axes</oasis:entry>  
         <oasis:entry colname="col3"><xref ref-type="sec" rid="Ch1.S3.SS3"/></oasis:entry>
       </oasis:row>
       <oasis:row>  
         <oasis:entry colname="col1">Scalar coordinate variable</oasis:entry>  
         <oasis:entry colname="col2">Coordinate for an implied size one axis</oasis:entry>  
         <oasis:entry colname="col3"><xref ref-type="sec" rid="Ch1.S3.SS3"/></oasis:entry>
       </oasis:row>
       <oasis:row>  
         <oasis:entry colname="col1">Grid mapping variable</oasis:entry>  
         <oasis:entry colname="col2">Horizontal coordinate system</oasis:entry>  
         <oasis:entry colname="col3"><xref ref-type="sec" rid="Ch1.S3.SS3"/></oasis:entry>
       </oasis:row>
       <oasis:row>  
         <oasis:entry colname="col1">Boundary variable</oasis:entry>  
         <oasis:entry colname="col2">Cell vertices</oasis:entry>  
         <oasis:entry colname="col3"><xref ref-type="sec" rid="Ch1.S3.SS5"/></oasis:entry>
       </oasis:row>
       <oasis:row>  
         <oasis:entry colname="col1">Cell measure variable</oasis:entry>  
         <oasis:entry colname="col2">Cell areas or volumes</oasis:entry>  
         <oasis:entry colname="col3"><xref ref-type="sec" rid="Ch1.S3.SS5"/></oasis:entry>
       </oasis:row>
       <oasis:row>  
         <oasis:entry colname="col1">Ancillary data variable</oasis:entry>  
         <oasis:entry colname="col2">Metadata that depend on the domain</oasis:entry>  
         <oasis:entry colname="col3"><xref ref-type="sec" rid="Ch1.S3.SS7"/></oasis:entry>
       </oasis:row>
       <oasis:row>  
         <oasis:entry colname="col1">Formula terms attribute</oasis:entry>  
         <oasis:entry colname="col2">Vertical coordinate system</oasis:entry>  
         <oasis:entry colname="col3"><xref ref-type="sec" rid="Ch1.S3.SS3"/></oasis:entry>
       </oasis:row>
       <oasis:row>  
         <oasis:entry colname="col1">Feature type attribute</oasis:entry>  
         <oasis:entry colname="col2">Characteristics of discrete sampling geometry</oasis:entry>  
         <oasis:entry colname="col3"><xref ref-type="sec" rid="Ch1.S3.SS4"/></oasis:entry>
       </oasis:row>
       <oasis:row>  
         <oasis:entry colname="col1">Cell methods attribute</oasis:entry>  
         <oasis:entry colname="col2">Description of variation within cells</oasis:entry>  
         <oasis:entry colname="col3"><xref ref-type="sec" rid="Ch1.S3.SS6"/></oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <?xmltex \floatpos{t}?><fig id="Ch1.F9"><caption><p id="d1e1835">The nine constructs of our CF data model (corresponding to
the green layer in Fig. <xref ref-type="fig" rid="Ch1.F1"/>) described using UML
(Appendix <xref ref-type="sec" rid="App1.Ch1.S1"/>). In this and all other CF data model
diagrams, the CF data model constructs are labelled with the
“construct” stereotype to distinguish them from other data model
elements which appear in other diagrams as green boxes with no
label. The field construct corresponds to a CF-netCDF data
variable. Relationships between other constructs and CF-netCDF are
given in Figs. <xref ref-type="fig" rid="Ch1.F10"/> and <xref ref-type="fig" rid="Ch1.F11"/>. The
domain provides the linkage between the field construct and the
constructs which describe measurement locations and cell
properties. It is not a construct of the data model (see
Sect. <xref ref-type="sec" rid="Ch1.S4.SS1"/>) but an abstract concept that is useful
for understanding it. Similarly, it is useful to define an
abstract generic coordinate construct that can be used to refer to
coordinates when the their type (dimension or auxiliary coordinate
construct) is not an issue.</p></caption>
        <?xmltex \igopts{width=221.931496pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/10/4619/2017/gmd-10-4619-2017-f09.pdf"/>

      </fig>

      <p id="d1e1854">The aspects of the CF conventions discussed in Sect. <xref ref-type="sec" rid="Ch1.S3"/> (i.e. CF-netCDF elements) are listed in Table <xref ref-type="table" rid="Ch1.T2"/> and shown (in
blue) with their interrelationships in UML (Appendix <xref ref-type="sec" rid="App1.Ch1.S1"/>) in
Fig. <xref ref-type="fig" rid="Ch1.F8"/>. The CF-netCDF elements and the relationships
among them could be regarded as defining a CF data model (rather like that
described in Sect. <xref ref-type="sec" rid="Ch1.S5.SS2"/>) to which any CF-compliant dataset may be
mapped. It is not the data model that we propose, however, because it does
not meet all of the design criteria described in Sect. <xref ref-type="sec" rid="Ch1.S1.SS1"/>.
Our CF data model (the green layer in Fig. <xref ref-type="fig" rid="Ch1.F1"/>) has been derived
from these CF-netCDF elements and relationships with the aims of removing
aspects specific to the netCDF encoding, and reducing the number of elements,
whilst retaining the ability to describe the CF conventions fully. The
elements of our CF data model are called “constructs”, a term chosen to
differentiate from the CF-netCDF elements previously defined and to be
programming language-neutral (i.e. as opposed to “object” or
“structure”). In this section, we relate the constructs to the CF-netCDF
elements of Sect. <xref ref-type="sec" rid="Ch1.S3"/>, which in turn relates the CF-netCDF
elements to the components of netCDF files. To clarify these connections, the
example netCDF file shown in Fig. <xref ref-type="fig" rid="Ch1.F3"/> indicates how its elements relate
to the constructs of our CF data model.</p>
<sec id="Ch1.S4.SS1">
  <title>The field construct</title>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T3" specific-use="star"><caption><p id="d1e1885">The constructs of our CF data model, a brief description of
each, and the section in which it is described in more detail. The
relationships between the constructs and CF-netCDF elements are
shown in Figs. <xref ref-type="fig" rid="Ch1.F9"/>–<xref ref-type="fig" rid="Ch1.F11"/>.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="3">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="left"/>
     <oasis:colspec colnum="3" colname="col3" align="left"/>
     <oasis:thead>
       <oasis:row rowsep="1">  
         <oasis:entry colname="col1">CF construct</oasis:entry>  
         <oasis:entry colname="col2">Description</oasis:entry>  
         <oasis:entry colname="col3">Section</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>  
         <oasis:entry colname="col1">Field</oasis:entry>  
         <oasis:entry colname="col2">Scientific data discretised within a domain</oasis:entry>  
         <oasis:entry colname="col3"><xref ref-type="sec" rid="Ch1.S4.SS1"/></oasis:entry>
       </oasis:row>
       <oasis:row>  
         <oasis:entry colname="col1">Domain axis</oasis:entry>  
         <oasis:entry colname="col2">Independent axes of the domain</oasis:entry>  
         <oasis:entry colname="col3"><xref ref-type="sec" rid="Ch1.S4.SS2"/></oasis:entry>
       </oasis:row>
       <oasis:row>  
         <oasis:entry colname="col1">Dimension coordinate</oasis:entry>  
         <oasis:entry colname="col2">Domain cell locations</oasis:entry>  
         <oasis:entry colname="col3"><xref ref-type="sec" rid="Ch1.S4.SS3"/></oasis:entry>
       </oasis:row>
       <oasis:row>  
         <oasis:entry colname="col1">Auxiliary coordinate</oasis:entry>  
         <oasis:entry colname="col2">Domain cell locations</oasis:entry>  
         <oasis:entry colname="col3"><xref ref-type="sec" rid="Ch1.S4.SS3"/></oasis:entry>
       </oasis:row>
       <oasis:row>  
         <oasis:entry colname="col1">Coordinate reference</oasis:entry>  
         <oasis:entry colname="col2">Domain coordinate systems</oasis:entry>  
         <oasis:entry colname="col3"><xref ref-type="sec" rid="Ch1.S4.SS4"/></oasis:entry>
       </oasis:row>
       <oasis:row>  
         <oasis:entry colname="col1">Domain ancillary</oasis:entry>  
         <oasis:entry colname="col2">Cell locations in alternative coordinate systems</oasis:entry>  
         <oasis:entry colname="col3"><xref ref-type="sec" rid="Ch1.S4.SS5"/></oasis:entry>
       </oasis:row>
       <oasis:row>  
         <oasis:entry colname="col1">Cell measure</oasis:entry>  
         <oasis:entry colname="col2">Domain cell size or shape</oasis:entry>  
         <oasis:entry colname="col3"><xref ref-type="sec" rid="Ch1.S4.SS6"/></oasis:entry>
       </oasis:row>
       <oasis:row>  
         <oasis:entry colname="col1">Field ancillary</oasis:entry>  
         <oasis:entry colname="col2">Ancillary metadata which vary within the domain</oasis:entry>  
         <oasis:entry colname="col3"><xref ref-type="sec" rid="Ch1.S4.SS7"/></oasis:entry>
       </oasis:row>
       <oasis:row>  
         <oasis:entry colname="col1">Cell method</oasis:entry>  
         <oasis:entry colname="col2">Describes how data represent variation within cells</oasis:entry>  
         <oasis:entry colname="col3"><xref ref-type="sec" rid="Ch1.S4.SS8"/></oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d1e2030">The field construct is central to our CF data model and includes all the
other constructs (Fig. <xref ref-type="fig" rid="Ch1.F9"/>). A field corresponds to a
CF-netCDF data variable with all of its metadata. All CF-netCDF elements are
mapped to some element of the CF field construct, as we describe in following
subsections, and the field constructs completely contain all the data and
metadata which can be extracted from the file using the CF conventions. Note
that the constructs contained by the field construct cannot exist
independently, as is indicated by the nature of the class associations shown
in Fig. <xref ref-type="fig" rid="Ch1.F9"/> (see Table <xref ref-type="table" rid="Ch1.T1"/> for details).</p>
      <p id="d1e2039">The field construct consists of a data array and the definition of its domain
(i.e. <inline-formula><mml:math id="M65" display="inline"><mml:mrow><mml:mi>V</mml:mi><mml:mo>(</mml:mo><mml:mi>d</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> in Eq. <xref ref-type="disp-formula" rid="Ch1.E1"/>), ancillary metadata fields defined over the
same domain, and cell method constructs to describe how the cell values
represent the variation of the physical quantity within the cells of the
domain (Fig. <xref ref-type="fig" rid="Ch1.F9"/>). Because the CF conventions do not mention
the concept of the domain, we do not regard it as a construct of the data
model. Instead, the domain is defined collectively by various other
constructs included in the field. All of the data model constructs are listed
in Table <xref ref-type="table" rid="Ch1.T3"/>, which refers to the sections and figures in
which they are fully described. All of the constructs contained by the field
construct are optional (as indicated by “0..*” in
Fig. <xref ref-type="fig" rid="Ch1.F9"/>). The only component of the field which is mandatory
is the data array.</p>
      <p id="d1e2064">The field construct also has optional properties to describe aspects of the
data that are independent of the domain. These correspond to some netCDF
attributes of variables (e.g. the units, long_name, and standard_name;
Sect. <xref ref-type="sec" rid="Ch1.S3.SS8"/>), and some netCDF global file attributes (e.g. “history”
and “institution”). We use the term “property”, rather than
“attribute”, because not all CF-netCDF attributes are properties in this
sense – some CF-netCDF attributes are used to point to (i.e. reference)
other netCDF variables and so only describe the data indirectly (e.g. the
“coordinates” attribute; Sect. <xref ref-type="sec" rid="Ch1.S3.SS3"/>), and others have structural
functions in the CF-netCDF file (e.g. the
“Conventions” attribute). In the
data model, we consider that netCDF global file attributes apply to every
data variable in the file, except where they are superseded by netCDF data
variable attributes with the same name. This interpretation of global file
attributes is not stated in the CF conventions, but for our data model it is
necessary because there is no notion of a file. Hence, metadata stored in
attributes of the file as a whole have to be transferred to the field
construct. If present, the global file attribute featureType applies to every
data variable in the file with a discrete sampling geometry
(Sect. <xref ref-type="sec" rid="Ch1.S3.SS4"/>). Hence, we regard feature type as a property of the field
construct.</p>
      <p id="d1e2074">The standard_name property (Sect. <xref ref-type="sec" rid="Ch1.S3.SS8"/>) constrains the units property
(i.e. only certain units are consistent with each standard name) and in some
cases also the dimensions that a data variable must have. These constraints,
however, do not supply any further information – they are just for
self-consistency. This is also the case for the feature type property, which
imposes some requirements on the axes the domain must have. Following our aim
of constructing a minimal data model, we do not regard the standard name nor
feature type as separate constructs within the field, because they do not
depend on any other construct for their interpretation. This is unlike a cell
method, for instance, which depends on the data variable's dimensions for its
interpretation.</p>
</sec>
<sec id="Ch1.S4.SS2">
  <title>Domain axis construct and the data array</title>
      <p id="d1e2085">A domain axis construct (Fig. <xref ref-type="fig" rid="Ch1.F10"/>) specifies the number of
points along an independent axis of the domain. It comprises a positive
integer representing the size of the axis. In CF-netCDF, it is usually defined
either by a netCDF dimension or by a scalar coordinate variable, which
implies a domain axis of size one (Sect. <xref ref-type="sec" rid="Ch1.S3.SS3"/>). The field construct's
data array spans the domain axis constructs of the domain, with the optional
exception of size one axes, because their presence makes no difference to the
order of the elements. Hence, the data array may be zero-dimensional (i.e. scalar) if there are no domain axis constructs of size greater than one.</p>
      <p id="d1e2092">When a collection of DSG features has been combined in a data variable using
the incomplete orthogonal or ragged representations to save space, the axis
size has to be inferred, but we regard this as an aspect of unpacking the
data, rather than its conceptual description. In practice, the unpacked data
array may be dominated by missing values (as could occur, for example, if all
features in a collection of time series had no common time coordinates), in
which case it may be preferable to view the collection as if each DSG feature
were a separate variable (Sect. <xref ref-type="sec" rid="Ch1.S3.SS4"/>), each one corresponding to a
different field construct.</p>
</sec>
<sec id="Ch1.S4.SS3">
  <?xmltex \opttitle{Coordinates: dimension coordinate and\hack{\break} auxiliary constructs}?><title>Coordinates: dimension coordinate and<?xmltex \hack{\break}?> auxiliary constructs</title>

      <?xmltex \floatpos{t}?><fig id="Ch1.F10" specific-use="star"><caption><p id="d1e2108">The relationship between domain axis, dimension coordinate,
and auxiliary coordinate constructs (Sect. <xref ref-type="sec" rid="Ch1.S4.SS2"/>
and <xref ref-type="sec" rid="Ch1.S4.SS3"/>) and CF-netCDF described using UML
(Appendix <xref ref-type="sec" rid="App1.Ch1.S1"/>). A dimension or auxiliary coordinate
construct is defined by a CF-netCDF coordinate, scalar coordinate,
or auxiliary coordinate variable, and the associated CF-netCDF
boundary variable if it exists. A generic coordinate construct
spans one or more domain axis constructs, but the mapping of which
ones is only held by the parent field construct.</p></caption>
          <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/10/4619/2017/gmd-10-4619-2017-f10.pdf"/>

        </fig>

      <p id="d1e2123">Coordinate constructs (Fig. <xref ref-type="fig" rid="Ch1.F10"/>) provide information which
locate the cells of the domain and which depend on a subset of the domain
axis constructs. As previously discussed, there are two distinct types of
coordinate construct: a dimension coordinate construct provides monotonic
numeric coordinates for a single domain axis, and an auxiliary coordinate
construct provides any type of coordinate information for one or more of the
domain axes.</p>
      <p id="d1e2128">In both cases, the coordinate construct consists of a data array of the
coordinate values which spans a subset of the domain axis constructs, an
optional array of cell bounds recording the extents of each cell, and
properties to describe the coordinates (in the same sense as for the field
construct). An array of cell bounds spans the same domain axes as its
coordinate array, with the addition of an extra dimension whose size is that
of the number of vertices of each cell. This extra dimension does not
correspond to a domain axis construct since it does not relate to an
independent axis of the domain (for example, the <monospace>bounds</monospace> dimension
defined at line 6 in Fig. <xref ref-type="fig" rid="Ch1.F3"/> does not correspond to a domain axis
construct). Note that, for climatological time axes, the bounds are
interpreted in a special way indicated by the cell method constructs.</p>
      <p id="d1e2137">The dimension coordinate construct is able to unambiguously describe cell
locations because a domain axis can be associated with at most one dimension
coordinate construct, whose data array values must all be non-missing and
strictly monotonically increasing or decreasing. They must also all be of the
same numeric data type. If cell bounds are provided, then each cell must have
exactly two vertices. CF-netCDF coordinate variables and numeric scalar
coordinate variables correspond to dimension coordinate constructs.</p>
      <p id="d1e2140">Auxiliary coordinate constructs have to be used, instead of dimension
coordinate constructs, when a single domain axis requires more then one set
of coordinate values, when coordinate values are not numeric, strictly
monotonic, or contain missing values, or when they vary along more than one
domain axis construct simultaneously. CF-netCDF auxiliary coordinate
variables and non-numeric scalar coordinate variables correspond to auxiliary
coordinate constructs.</p>
      <p id="d1e2143">If a domain axis construct does not correspond to a continuous physical
quantity, then it is not necessary for it to be associated with a dimension
coordinate construct. For example, this is the case for an axis that runs
over ocean basins or area types, or for a domain axis that indexes a
time series at scattered points. In such cases, one-dimensional auxiliary
coordinate constructs could be used to store coordinate values. These axes
are discrete axes in CF-netCDF.</p>
</sec>
<sec id="Ch1.S4.SS4">
  <title>Coordinate reference construct</title>
      <p id="d1e2152">The domain may contain various coordinate systems, each of which is
constructed from a subset of the dimension and auxiliary coordinate
constructs. For example, the domain of a four-dimensional field construct may
contain horizontal (<inline-formula><mml:math id="M66" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula>–<inline-formula><mml:math id="M67" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula>), vertical (<inline-formula><mml:math id="M68" display="inline"><mml:mi>z</mml:mi></mml:math></inline-formula>), and temporal (<inline-formula><mml:math id="M69" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>) coordinate
systems. There may be more than one of each of these, if there is more than
one coordinate construct applying to a particular spatiotemporal dimension
(for example, there could be both latitude–longitude and <inline-formula><mml:math id="M70" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula>–<inline-formula><mml:math id="M71" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> projection
coordinate systems). In general, a coordinate system may be constructed
implicitly from any subset of the coordinate constructs, yet a coordinate
construct does not need to be explicitly or exclusively associated with any
coordinate system.</p>
      <p id="d1e2198">A coordinate system of the field construct can be explicitly defined by a
coordinate reference construct (Fig. <xref ref-type="fig" rid="Ch1.F11"/>) which relates
the coordinate values of the coordinate system to locations in a planetary
reference frame and consists of the following:
<list list-type="bullet"><list-item>
      <p id="d1e2205">The dimension coordinate and auxiliary coordinate constructs
that define the coordinate system to which the coordinate reference
construct applies. Note that the coordinate values are not relevant
to the coordinate reference construct, only their properties.</p></list-item><list-item>
      <p id="d1e2209">A definition of a datum specifying the zeroes of the dimension
and auxiliary coordinate constructs which define the coordinate
system. The datum may be explicitly indicated via properties, or it
may be implied by the metadata of the contained dimension and
auxiliary coordinate constructs. Note that the datum may contain the
definition of a geophysical surface which corresponds to the zero of
a vertical coordinate construct, and this may be required for both
horizontal and vertical coordinate systems.</p></list-item><list-item>
      <p id="d1e2213">A coordinate conversion, which defines a formula for converting
coordinate values taken from the dimension or auxiliary coordinate
constructs to a different coordinate system. A term of the
conversion formula can be a scalar or vector parameter which does
not depend on any domain axis constructs, may have units (such as a
reference pressure value), or may be a descriptive string (such as
the projection name “mercator”), or it can be a domain ancillary
construct (such as one containing spatially varying orography data).</p></list-item></list></p>
      <p id="d1e2216">For <inline-formula><mml:math id="M72" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula>–<inline-formula><mml:math id="M73" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> coordinates, the coordinate conversion is either a map
projection, which converts between Cartesian coordinates and spherical
or ellipsoidal coordinates on the vertical datum, or a conversion
between different spherical coordinate systems (as in the case of
rotated pole coordinates). In the case of <inline-formula><mml:math id="M74" display="inline"><mml:mi>z</mml:mi></mml:math></inline-formula> coordinates, the
conversion is between a coordinate construct with parameterised values
(such as ocean sigma coordinates) and a coordinate construct with
dimensional values (such as depths), again with respect to the
vertical datum.</p>
      <p id="d1e2240">In some cases, the datum is not required as it is already described by the
dimension and auxiliary coordinate constructs. This is the case in CF for the
two-dimensional geographical latitude–longitude coordinate system based upon
a spherical Earth, which is assumed to have a datum at 0<inline-formula><mml:math id="M75" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> N, 0<inline-formula><mml:math id="M76" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> E. Similarly, the coordinate conversion
is not required if no other coordinate systems are described. Some parts of
the coordinate reference construct may not be relevant to a given coordinate
construct which it contains. The relevant parts are determined by an
application using the coordinate reference construct. For example, for a
coordinate reference construct which contained coordinate constructs for
<inline-formula><mml:math id="M77" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula>–<inline-formula><mml:math id="M78" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> projection and latitude and longitude coordinates, a datum
comprising a reference ellipsoid would apply to all of them, but projection
parameters would only apply to the projection coordinates.</p>
      <p id="d1e2276">In CF-netCDF, coordinate system information that is not found in coordinate
or auxiliary coordinate variables is stored in a grid mapping variable or the
formula_terms attribute of a coordinate variable, for horizontal or vertical
coordinate variables, respectively. Although these two cases are arranged
differently in CF-netCDF, each one contains, sometimes implicitly, a datum or
a coordinate conversion formula (or both) and so may be mapped to a
coordinate reference construct. A grid mapping name or the standard name of a
parametric vertical coordinate corresponds to a string-valued scalar
parameter of a coordinate conversion formula. A grid mapping parameter which
has more than one value (as is possible with the “standard parallel”
attribute) corresponds to a vector parameter of a coordinate conversion
formula. A data variable referenced by a formula_terms attribute corresponds
to the term of a coordinate conversion formula – either a domain ancillary
construct or, if it is zero-dimensional, a scalar parameter.</p>
</sec>
<sec id="Ch1.S4.SS5">
  <title>Domain ancillary construct</title>
      <p id="d1e2285">A domain ancillary construct (Fig. <xref ref-type="fig" rid="Ch1.F11"/>) provides
information which is needed for computing the location of cells in an
alternative coordinate system. It is the value of a term of a coordinate
conversion formula that contains a data array, which is zero-dimensional or
which depends on one or more of the domain axes.</p>
      <p id="d1e2290">It also contains an optional array of cell bounds recording the
extents of each cell (only applicable if the array contains coordinate
data) and properties to describe the data (in the same sense as for
the field construct). An array of cell bounds spans the same domain
axes as the data array, with the addition of an extra dimension whose
size is that of the number of vertices of each cell.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F11" specific-use="star"><caption><p id="d1e2295">The relationship between coordinate reference and domain
ancillary constructs (Sect. <xref ref-type="sec" rid="Ch1.S4.SS4"/> and
<xref ref-type="sec" rid="Ch1.S4.SS5"/>) and CF-netCDF described using
UML (Appendix <xref ref-type="sec" rid="App1.Ch1.S1"/>). A coordinate reference construct is
defined either by a grid mapping variable or a
<monospace>formula_terms</monospace> attribute of a CF-netCDF coordinate
variable. The coordinate reference construct is composed of
generic coordinate constructs, a datum, and a coordinate
conversion formula. The coordinate conversion formula is usually
defined by a named formula in the CF conventions. A domain
ancillary construct term of a coordinate conversion formula is
defined by a CF-netCDF data variable or a CF-netCDF generic
coordinate variable.</p></caption>
          <?xmltex \igopts{width=332.897244pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/10/4619/2017/gmd-10-4619-2017-f11.pdf"/>

        </fig>

      <p id="d1e2313">CF-netCDF variables named by the formula_terms attribute of a
CF-netCDF coordinate variable correspond to domain ancillary
constructs. These CF-netCDF variables may be coordinate, scalar
coordinate, or auxiliary coordinate variables, or they may be data variables. For
example, in a coordinate conversion for converting between ocean sigma
and height coordinate systems, the value of the “depth” term for
horizontally varying distance from ocean datum to sea floor would
correspond to a domain ancillary construct. In the case of a named
term being a type of coordinate variable, that variable will
correspond to an independent domain ancillary construct in addition to
the coordinate construct.</p>
</sec>
<sec id="Ch1.S4.SS6">
  <title>Cell measure construct</title>
      <p id="d1e2322">A cell measure (Fig. <xref ref-type="fig" rid="Ch1.F9"/>) construct provides information that
is needed about the size or shape of the cells and that depends on a subset
of the domain axis constructs. Cell measure constructs have to be used when
the size or shape of the cells cannot be deduced from the dimension or
auxiliary coordinate constructs without special knowledge that a generic
application cannot be expected to have.<?xmltex \hack{\newpage}?></p>
      <p id="d1e2328">The cell measure construct consists of a numeric array of the metric
data which span a subset of the domain axis constructs, and
properties to describe the data (in the same sense as for the field
construct). The properties must contain a “measure” property, which
indicates which metric of the space it supplies, e.g. cell horizontal
areas, and a units property consistent with the measure property,
e.g. m<inline-formula><mml:math id="M79" display="inline"><mml:msup><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:math></inline-formula>. It is assumed that the metric does not depend on axes
of the domain which are not spanned by the array, along which the
values are implicitly propagated. CF-netCDF cell measure variables
correspond to cell measure constructs.</p>
</sec>
<sec id="Ch1.S4.SS7">
  <title>Field ancillary constructs</title>
      <p id="d1e2347">The field ancillary construct (Fig. <xref ref-type="fig" rid="Ch1.F9"/>) provides metadata
which are distributed over the same sampling domain as the field itself. For
example, if a data variable holds a variable retrieved from a satellite
instrument, a related ancillary data variable might provide the uncertainty
estimates for those retrievals (varying over the same spatiotemporal domain).</p>
      <p id="d1e2352">The field ancillary construct consists of an array of the ancillary
data, which is zero-dimensional or which depends on one or more of the
domain axes, and properties to describe the data (in the same sense as
for the field construct). It is assumed that the data do not depend
on axes of the domain which are not spanned by the array, along which
the values are implicitly propagated. CF-netCDF ancillary data
variables correspond to field ancillary constructs. Note that a field
ancillary construct is constrained by the domain definition of the
parent field construct but does not contribute to the domain's
definition, unlike, for instance, an auxiliary coordinate construct or domain
ancillary construct.</p>
</sec>
<sec id="Ch1.S4.SS8">
  <title>Cell method construct</title>
      <p id="d1e2361">The cell method constructs (Fig. <xref ref-type="fig" rid="Ch1.F9"/>) describe how the cell
values represent the variation of the physical quantity within its cells –
the structure of the data at a higher resolution. A single cell method
construct consists of a set of axes (see below), a “method” property which
describes how a value of the field construct's data array describes the
variation of the quantity within a cell over those axes (e.g. a value might
represent the cell area average), and properties serving to indicate more
precisely how the method was applied (e.g. recording the spacing of the
original data, or the fact the method was applied only over El Niño years).</p>
      <p id="d1e2366">The field construct may contain an ordered sequence of cell method constructs
describing multiple processes which have been applied to the data, e.g. a
temporal maximum of the areal mean has two components – a mean and a
maximum – each acting over different sets of axes. It is an ordered sequence
because the methods specified are not necessarily commutative. There are
properties to indicate climatological time processing, e.g. multiannual
means of monthly maxima, in which case multiple cell method constructs need
to be considered together to define a special interpretation of boundary
coordinate array values. The cell_methods attribute of a CF-netCDF data
variable corresponds to one or more cell method constructs.</p>
      <p id="d1e2369"><?xmltex \hack{\newpage}?>The axes over which a cell method applies are either a subset of the domain
axis constructs or a collection of strings which identify axes that are not
part of the domain. The latter case is particularly useful when the
coordinate range for an axis cannot be precisely defined, making it
impossible to define a domain axis construct. For example, a climatological
time mean might be based on data which are not available over the same time
periods at every horizontal location – useful information can still be
conveyed by recording the fact the data have been temporally averaged without
specifying the range of times. The strings which identify such axes are well
defined in that they must be standard names (e.g. time, longitude) or the
special string “area”, indicating a combination of horizontal axes.</p>
</sec>
</sec>
<sec id="Ch1.S5">
  <title>Relationship to other data models</title>
      <p id="d1e2381">A data model does not exist on its own, and those exploiting it will need to
interpret it in the context of other data models with which they already
work, whether they are implicit or explicit (Sect. <xref ref-type="sec" rid="Ch1.S1.SS1"/>).
Often a clear, unambiguous, and universally agreed (or even agreeable) mapping
between data models is not possible. <xref ref-type="bibr" rid="bib1.bibx8" id="text.6"/>, who establish a mapping
between the Unidata Common Data Model and relevant international standards,
discuss many of the relevant issues. Here, we confine ourselves to drawing
some parallels between our explicit CF data model and selected other data
modelling activities. We begin with the ISO 19123 coverage model, which
provides an abstract view of the problem arena.</p>
      <p id="d1e2389">Readers who are not familiar with other data models may wish to omit this
section on a first reading, as it is not required to understand the CF
conventions, the CF data model presented here, nor the software
implementation of Sect. <xref ref-type="sec" rid="Ch1.S6"/>.</p>
<sec id="Ch1.S5.SS1">
  <title>The ISO 19123 coverage model</title>
      <p id="d1e2399">ISO 19123 <xref ref-type="bibr" rid="bib1.bibx6" id="paren.7"/> provides a language and conceptual schema for
describing “coverages”, that is, for datasets which assign data values to
specified data locations (in space and/or time). ISO 19123 builds on a range
of other ISO standards for geographical information, collectively known as
the ISO 191xx series.</p>

      <?xmltex \floatpos{p}?><fig id="Ch1.F12" specific-use="star"><caption><p id="d1e2407">Key concepts within the ISO 19123 view of coverages and the
associated coordinate systems: <bold>(a)</bold> as applied to a specific
coverage: the discrete coverage, which relates a domain of objects
to a set of attribute values; <bold>(b)</bold> an expanded view of the
relationship of discrete grid point coverages to grid cells, grid
footprints, and three different ways of thinking about the grids
(set of coordinates, rectified, and referenceable grids). Note the
comment box defining the two-letter labels of each element
(e.g. CV indicates that this element comes from the ISO 19123 data
model). See also Fig. 9 in <xref ref-type="bibr" rid="bib1.bibx2" id="text.8"/>. Further
relationships and details are discussed in the text.</p></caption>
          <?xmltex \igopts{width=355.659449pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/10/4619/2017/gmd-10-4619-2017-f12.pdf"/>

        </fig>

      <p id="d1e2425">An ISO 191xx coverage may be viewed as a function whose inputs are
spatiotemporal positions (the “domain”) which are related to outputs
comprising values of one or more geographical features (the “range” of the
coverage). Thus, a coverage is notionally a function over a domain which has a
range of values. Within ISO 19123, two types of coverage are defined: discrete
and continuous. The former is the most relevant here (see
Fig. <xref ref-type="fig" rid="Ch1.F12"/>a), having a domain which consists of a set of
“domain objects”, themselves described by spatial objects and temporal
primitives. In such a coverage, all the objects must share the same
coordinate reference system. Coordinate reference systems themselves can be
compound, but the simplest single coordinate reference system effectively
consists of a datum and a set of coordinates (which might include one or more
parametric coordinates) defining a coordinate system. Although the language
of much of the ISO 19123 specification is cast in terms of simple geospatial
coordinates, these abstract coordinate systems are in fact fully general –
although many ISO-compliant implementations restrict them to geospatial
coordinates – and cannot reflect the full generality of CF coordinates such as
wavelengths, ensembles, etc. UML diagrams (Appendix <xref ref-type="sec" rid="App1.Ch1.S1"/>) for the full
ISO 19123 model are given in Fig. <xref ref-type="fig" rid="Ch1.F12"/>.</p>
      <p id="d1e2434">For point data, a discrete coverage is nearly identical to a CF field
construct (<inline-formula><mml:math id="M80" display="inline"><mml:mrow><mml:mi>V</mml:mi><mml:mo>(</mml:mo><mml:mi>d</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, as described in Sect. <xref ref-type="sec" rid="Ch1.S3"/>). However, a CF
field construct explicitly describes a single variable on a domain, whilst a
discrete coverage allows for multiple variables on a shared domain. (In
ISO 19123, such variables are termed “features”, not to be confused with
“sampling features” discussed below.) It is also worth noting that many,
but not all, netCDF files containing CF data may well have CF data with a
shared implied spatiotemporal domain, and such files may well map rather
directly onto the coverage model, but as we have seen, this is not always the
case, not least because multiple domain descriptions may appear in a given
file.</p>
      <p id="d1e2454">Discrete coverages themselves can be further specialised into a set of
more specialised coverages: sampling the domain with sets of points, a
grid of points, or sets of curves, surfaces, or solids.</p>
      <p id="d1e2457">Of these, the most important (in terms of our CF data model) is the
DiscreteGridPointCoverage (Fig. <xref ref-type="fig" rid="Ch1.F12"/>b). This relates a domain
of grid points laid out in a grid to a range of values in GridValueMatrix.
The grid can be defined in three different ways: as a set of grid points
(with their own coordinates), as a “rectified” grid (a grid utilising a
datum and coordinate vectors for which there is an affine transformation
between the coordinates and those of an external coordinate reference
system), or a referenceable grid which provides parametric relations between a
position using a coordinate tuple and a position in a planetary reference
frame using a specific coordinate reference system. The values can appear in
a grid matrix, which defines how the sequence of values is laid out against
the positioning exposed via the domain grid (in any of the three forms).</p>
      <p id="d1e2462">There is a clear correspondence between the CF dimension coordinate
construct and an ISO coordinate reference system as used in a
rectified grid, and between the underlying concepts of ISO parametric
coordinate reference systems within a referenceable grid and a CF
domain described using CF auxiliary coordinate constructs. This
correspondence together supports the identification of an ISO grid
(which itself carries little information apart from a name and a list
of axes) with the abstract notion of a CF domain described by CF
coordinate reference constructs.</p>
      <p id="d1e2465">Even with an ISO rectified grid, which has the easiest correspondence
with CF, there are subtle but important differences in the treatment
of coordinates, probably the most important of which is that the CF
equivalent of the ISO datum is often held in the standard name of the
coordinate construct. For example, a CF coordinate construct with a
standard name of height means the coordinate is with reference to the
surface, i.e. the bottom of the atmosphere (distinct from other valid
vertical coordinates such as height_above_reference_ellipsoid and
height_above_sea_floor). These coordinates are all distinct
geophysical quantities, with vertical datums of the surface, the
reference ellipsoid, and the sea floor, respectively, though they all
have the same canonical unit of measure (metres) and direction (values
increase for locations further above the datum).</p>
      <p id="d1e2468">Where a more precise specification of the datum may be needed (for
example, the figure of the reference ellipsoid or the reference point
for a latitude–longitude coordinate system where it is not the
default of the intersection of the Equator and the Greenwich meridian),
it can be supplied by the coordinate reference construct, not the
standard name. This CF separation of grid mapping datum from
coordinates adds value because changing the datum does not alter the
geophysical nature of the coordinate and its interpretation. The
partitioning is suitable and convenient for many purposes of data
analysis, in which coordinate constructs are processed independently,
without the need for awareness of a full ISO coordinate reference
system (CRS). It arises from the generality of CF, in which
spatiotemporal coordinates are used having a wider variety than in
the geographic information system (GIS); non-spatiotemporal coordinates are also needed; and the data
are often from idealised worlds (such as in climate models).</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F13"><caption><p id="d1e2474">The relationship between ISO grid cells and footprints, and
CF cells described by cell bounds.</p></caption>
          <?xmltex \igopts{width=128.037402pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/10/4619/2017/gmd-10-4619-2017-f13.pdf"/>

        </fig>

      <p id="d1e2483">Whether grid cells are described directly or implied via referenceable or
rectified grids, it is important to note that in the ISO world, the cells lie
between the edges laid out by the coordinates, whereas by contrast the notion
of a cell in CF – defined by the cell measure construct
(Sect. <xref ref-type="sec" rid="Ch1.S4.SS6"/>) – is more directly analogous
(Fig. <xref ref-type="fig" rid="Ch1.F13"/>) to the FootPrint in an ISO coverage which
represents the sample space associated with a grid point
(Fig. <xref ref-type="fig" rid="Ch1.F12"/>b). There is also a richer set of semantics in CF
associated with the cell method construct (Sect. <xref ref-type="sec" rid="Ch1.S4.SS8"/>).<?xmltex \hack{\newpage}?></p>
      <p id="d1e2495">It might appear that some of the more complex geometries underlying CF fields
which are expressed on domains sampled using the CF DSG features best be
mapped onto other specialisations of DiscreteCoverages – this is the
approach taken by <xref ref-type="bibr" rid="bib1.bibx8" id="normal.9"/> who use the DiscreteCurveCoverage for
mapping ISO coverages onto Unidata Common Data Model (Sect. <xref ref-type="sec" rid="Ch1.S5.SS3"/>)
profile and trajectory data. However, we would assert that the underlying
semantics of CF as expressed by the concept of cells are generalisable to
multi-dimensional sampling inside the DiscreteGridPoint coverage. This is in
part why we see the CF sampling feature types as simply specialisations of
the CF field, possibly with a specific storage pattern. This simple mapping
between CF “domains” and ISO DiscreteGridPoint coverage is also the
approach taken by the Open Geospatial Consortium (OGC) CF-netCDF extension
standard (Sect. <xref ref-type="sec" rid="Ch1.S5.SS2"/> below) which describes another set of possible
relationships between CF-netCDF and ISO 19123, albeit without the
higher-level CF constructs we introduce here.</p>
      <p id="d1e2505">ISO 19156 Observations and Measurements <xref ref-type="bibr" rid="bib1.bibx7" id="paren.10"/> introduces sampling
features with a range of geometric spatial properties (e.g. sampling along a
curve, such as a trajectory). CF discrete sampling geometries can be mapped
onto these sampling features, but it is important to note that ISO 19156
explicitly expects actual observation (or simulation) values to be
subsampled. Thus, the observations are amenable to recording with more
general discrete coverages, just as we have done here by treating all CF
discrete sampling geometries as CF field constructs, optionally labelled with
a feature type property to help understand the intended use of the axes.</p>
</sec>
<sec id="Ch1.S5.SS2">
  <title>The OGC CF-netCDF standard</title>
      <p id="d1e2517">The Open Geospatial Consortium standard introduced above <xref ref-type="bibr" rid="bib1.bibx2" id="paren.11"/>
in the context of complex geometries presents their own CF data model: the
CF-netCDF extension model. Their model differs from ours in four major ways:
<list list-type="order"><list-item>
      <p id="d1e2525">It is not the complete CF version 1.6 (for example, it does not
appear to include ancillary data variables).</p></list-item><list-item>
      <p id="d1e2529">Their model makes some elements of CF mandatory, in order to
facilitate the ISO 19123 coverage interoperability, which is their
target.</p></list-item><list-item>
      <p id="d1e2533">It is tied to the netCDF format.</p></list-item><list-item>
      <p id="d1e2537">It is constructed in order to map as closely as possible onto
the ISO 19123 coverage model but without being faithful to CF; so,
for example, it introduces the notion of a CF coordinate system
including a notional HorizontalCRS, which is independent of
explicitly identified horizontal and vertical coordinates (their
Fig. 4). By contrast, we have only introduced new concepts as
abstractions where they help interpret and use CF itself (again, for
example, in our case the domain and abstract coordinate).</p></list-item></list>
The last of these points is of course subjective. We would argue that
our approach is the most consistent with a faithful model of CF, but
it is clear that CF itself currently admits a multitude of possible
interpretative models as well as a multitude of correct (if limited
and/or constrained) implementations such as this one.</p>
</sec>
<sec id="Ch1.S5.SS3">
  <title>The Unidata Common Data Model</title>
      <p id="d1e2548">The Unidata Common Data Model (CDM; <xref ref-type="bibr" rid="bib1.bibx13" id="altparen.12"/>) is an abstract data
model for scientific datasets that is a superset of the netCDF classic and
enhanced data models (Sect. <xref ref-type="sec" rid="Ch1.S2"/>). In addition to netCDF, and
similarly to the CF data model presented here, it consists of multiple
layers:
<list list-type="bullet"><list-item>
      <p id="d1e2558">a data access layer, which handles data reading and writing,
and merges the netCDF enhanced, OPeNDAP (Open-source Project for a
Network Data Access Protocol,
<uri>https://www.opendap.org</uri>) and HDF (Hierarchical
Data Format, <uri>https://www.hdfgroup.org</uri>) data models
to create a common application programming interface (API);</p></list-item><list-item>
      <p id="d1e2568">a coordinate system layer, which handles the coordinates of data
arrays;</p></list-item><list-item>
      <p id="d1e2572">a feature type layer, which handles similar notions to those
we express with CF field constructs and the CF sampling feature
types; and</p></list-item><list-item>
      <p id="d1e2576">a mature Java-based implementation which reads, manipulates,
and writes the CDM sampling features.</p></list-item></list></p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F14"><caption><p id="d1e2581">Key characteristics of the Unidata Common Data Model. There
is a wider variety of fundamental data types than is supported by
the netCDF classic data model; and the coordinate system includes
the option of coordinate axes of specific types for use in the
feature types, which limits the flexibility of the CDM data
model.</p></caption>
          <?xmltex \igopts{width=230.467323pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/10/4619/2017/gmd-10-4619-2017-f14.pdf"/>

        </fig>

      <p id="d1e2590">The CDM data access layer has a broader scope than ours (being about more than
just netCDF). If we consider that most of the CF standard as expressed in our
data model is about handling coordinates, cells, and domains, then our CF
data model corresponds to CDM coordinate system layer, with the various CF
feature types of our CF data model field constructs corresponding to the CDM
feature type layer. The cf-python software (Sect. <xref ref-type="sec" rid="Ch1.S6"/>)
corresponds to the CDM Java implementation, but our CF data model is intended
to be useful outside the context of our cf-python application.</p>
      <p id="d1e2595">The CDM data access layer handles more data types (Fig. <xref ref-type="fig" rid="Ch1.F14"/>), in
particular structures and sequences which allow more complex data types and
the iteration through data of (a priori) unknown length. These data types are
not yet required by CF-netCDF (Sect. <xref ref-type="sec" rid="Ch1.S2"/>), so we do not consider
them in our CF data model.</p>
      <p id="d1e2603">Within the coordinate system layer, there is much closer correspondence
between the CDM and our CF data model.
<list list-type="bullet"><list-item>
      <p id="d1e2608">A CF dimension or auxiliary coordinate construct maps to a
CDM CoordinateAxis.</p></list-item><list-item>
      <p id="d1e2612">The datum and coordinate conversion components of a CF
coordinate reference construct are components of a CDM
CoordinateTransform.</p></list-item><list-item>
      <p id="d1e2616">A CF coordinate reference construct maps to a CDM
CoordinateSystem.</p></list-item></list>
The last of these has one exception: a CDM CoordinateSystem must
contain at least one CDM CoordinateAxis, whereas CF dimension and
auxiliary coordinate constructs are optional in a CF coordinate
reference construct. In other words, our CF data model can record a
coordinate system datum in the absence of coordinate values. This is
useful when the CF field construct properties, rather than its
coordinates, define the extent of the data array.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F15"><caption><p id="d1e2622">The relationship between features, feature types, and
feature collections in the Unidata Common Data Model. Only the
simplest feature types are shown; more complicated features are
built from these basic elements <xref ref-type="bibr" rid="bib1.bibx14" id="paren.13"/>.</p></caption>
          <?xmltex \igopts{width=236.157874pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/10/4619/2017/gmd-10-4619-2017-f15.pdf"/>

        </fig>

      <?xmltex \floatpos{t}?><fig id="Ch1.F16" specific-use="star"><caption><p id="d1e2636">Reading a file using cf-python.</p></caption>
          <?xmltex \igopts{width=497.923228pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/10/4619/2017/gmd-10-4619-2017-f16.pdf"/>

        </fig>

      <p id="d1e2645">A CDM CoordinateAxis may be subtyped into axes which can be specifically
exploited by the sampling feature types in the CDM feature type layer, where
there are significant differences from our CF data model. The CDM feature
type implementation is discussed in <xref ref-type="bibr" rid="bib1.bibx14" id="text.14"/>; CDM implements
feature collections as collections of point features in such a way that a
profile, for example, is a collection of point features along a vertical line,
and so a profile collection is a nested point feature collection. Similarly,
a trajectory feature is a collection of points along a path through time and
space, with collections which are nested point feature collections
(Fig. <xref ref-type="fig" rid="Ch1.F15"/>). Other CDM feature types are built from this base. The CDM
data model exposes these concepts directly. By contrast, our CF data model
does not expose any of these ideas directly, with the interpretation left
entirely to software implementations: our CF data model simply exposes the
appropriate coordinates and their interpretation is either inferable from the
nature of those coordinates or is made explicit via the feature type
property, which is effectively a constraint with a label (Sect. <xref ref-type="sec" rid="Ch1.S4.SS1"/>).</p>

      <?xmltex \floatpos{p}?><fig id="Ch1.F17" specific-use="star"><caption><p id="d1e2658">A detailed inspection of a field object's metadata.</p></caption>
          <?xmltex \igopts{width=469.470472pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/10/4619/2017/gmd-10-4619-2017-f17.pdf"/>

        </fig>

</sec>
</sec>
<sec id="Ch1.S6">
  <title>cf-python: a data model implementation</title>
      <p id="d1e2674">A key use of our data model is to enable the creation of wholly CF-compliant
software, i.e. software that can represent and manipulate any CF-compliant
dataset. Such software corresponds to one of the application boxes in
Fig. <xref ref-type="fig" rid="Ch1.F1"/>. cf-python is a data analysis software library written
for the Python programming language that implements the CF data model
presented here for its internal data structures and so is able to process any
CF-compliant dataset. It is, however, not strict about CF compliance so that
partially conformant datasets may be ingested and their deficiencies
corrected with the API. cf-python is
open-source software and is free to download and install from the Python
package index at <uri>https://pypi.python.org/pypi/cf-python</uri>.</p>
      <p id="d1e2682">cf-python implements the data model constructs and their relationships
exactly as shown in Figs. <xref ref-type="fig" rid="Ch1.F9"/>–<xref ref-type="fig" rid="Ch1.F11"/>, but we
refer to the cf-python realisations of CF data model constructs as
“objects” in order to distinguish between abstract concepts and their
instantiated counterparts. cf-python (version 2.1) can read field objects
from netCDF files or create new field objects, manipulate field objects in
memory, and write field objects to netCDF files. Once a field object exists a
range of operations are possible, including to
<list list-type="bullet"><list-item>
      <p id="d1e2691">create, delete, and modify a field object's data and
metadata;</p></list-item><list-item>
      <p id="d1e2695">select and subspace field objects according to their
metadata;</p></list-item><list-item>
      <p id="d1e2699">perform arithmetic, comparison, and other mathematical
operations involving field objects;</p></list-item><list-item>
      <p id="d1e2703">collapse axes by statistical operations;</p></list-item><list-item>
      <p id="d1e2707">perform operations with date–time data;</p></list-item><list-item>
      <p id="d1e2711">regrid fields to new domains using the Earth System Modeling
Framework high-performance software infrastructure
<xref ref-type="bibr" rid="bib1.bibx9" id="paren.15"/>; and</p></list-item><list-item>
      <p id="d1e2718">visualise field objects by interfacing with the cf-plot
Python package, which is also open source and freely available
at <uri>https://pypi.python.org/pypi/cf-plot</uri>.</p></list-item></list></p>
      <p id="d1e2724">All of these operations are “metadata aware”, which means that
parameters needed for an operation need not be fully specified by the
user, provided that field objects have sufficient metadata to infer
the parameters unambiguously. This is greatly facilitated by having a
data model, because all standardised metadata are stored in a fully
defined manner and so the required parameters may be inferred
unambiguously. In practice, a field object's metadata may be
incomplete, in which case the user should use the cf-python API to
supplement the metadata. For example, the cf-python command
<monospace>h=f.regrids(g)</monospace> will create a new field object <monospace>h</monospace>
which has the data from field object <monospace>f</monospace> regridded to the
latitude–longitude plane of the domain of field object <monospace>g</monospace>. If
the domains of <monospace>f</monospace> and <monospace>g</monospace> are not sufficiently
described for this operation, then an error will be raised that states
which information is missing. The full API documentation is available
as part of the cf-python installation, as well as via the Python
package index.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F18" specific-use="star"><caption><p id="d1e2748">The cf-python class instances which correspond to CF data
model constructs.</p></caption>
        <?xmltex \igopts{width=497.923228pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/10/4619/2017/gmd-10-4619-2017-f18.pdf"/>

      </fig>

      <p id="d1e2758">How cf-python implements the data model may be seen by using the library to
read the CF-netCDF file described in Fig. <xref ref-type="fig" rid="Ch1.F3"/> and inspecting the field
objects that it creates. The code listed in Fig. <xref ref-type="fig" rid="Ch1.F16"/> demonstrates
importing the library and reading this netCDF file within the interactive
Python shell, in which a command is preceded by the <inline-formula><mml:math id="M81" display="inline"><mml:mrow><mml:mi mathvariant="italic">&gt;&gt;</mml:mi><mml:mo>&gt;</mml:mo></mml:mrow></mml:math></inline-formula> prompt and followed
by any printed output. After importing the cf-python library (<monospace>import cf</monospace>), the netCDF file <monospace>example_file.nc</monospace> described in Fig. <xref ref-type="fig" rid="Ch1.F3"/>
is read into the variable <monospace>f</monospace>, which is a list of the file's two field
objects – one containing air temperature data and the other containing the vertical
integral of atmospheric water vapour data. See Sect. “Interpreting CF-netCDF
files” for a discussion on why only two field
objects were created from the 17 netCDF variables in the file. The
one-line description of each field object shows the size and physical nature
of the data array dimensions and the units of the data array values. In
Fig. <xref ref-type="fig" rid="Ch1.F17"/>, the first of the two field objects is selected and
labelled <monospace>t</monospace>, and a more detailed look at its metadata is generated
with the <monospace>dump</monospace> function. This exposes the constructs of the data
model which have been instantiated for this field object and shows the field
object's properties (<monospace>standard_name</monospace> and <monospace>source</monospace> instantiated
from lines 54 and 74 of Fig. <xref ref-type="fig" rid="Ch1.F3"/>). The output describes</p>
      <p id="d1e2804"><list list-type="bullet">
          <list-item>

      <p id="d1e2809">one field object;</p>
          </list-item>
          <list-item>

      <p id="d1e2815">four domain axis objects and their sizes, including one which
is implied by the “time” CF-netCDF scalar coordinate variable;</p>
          </list-item>
          <list-item>

      <p id="d1e2821">one cell method object indicating that each data array value
is a time average constructed from daily samples;</p>
          </list-item>
          <list-item>

      <p id="d1e2827">one field ancillary object describing the uncertainty of the
data array values;</p>
          </list-item>
          <list-item>

      <p id="d1e2833">four dimension coordinate objects, each one spanning a unique
domain axis object;</p>
          </list-item>
          <list-item>

      <p id="d1e2840">two multi-dimensional auxiliary coordinate objects for true
latitude and true longitude coordinates (as required by the CF
conventions when the horizontal dimension coordinates are not
canonical geographical latitudes and longitudes);</p>
          </list-item>
          <list-item>

      <p id="d1e2846">three domain ancillary objects utilised by the coordinate
reference objects;</p>
          </list-item>
          <list-item>

      <p id="d1e2852">two coordinate reference objects: a vertical, atmosphere sigma
coordinate system which references the domain ancillary
objects and the vertical dimension coordinate object, and a
horizontal Lambert conformal conic coordinate system which
references the horizontal auxiliary and dimension coordinate
objects; and</p>
          </list-item>
          <list-item>

      <p id="d1e2858">one cell measure object containing horizontal cell areas.</p>
          </list-item>
        </list></p>
      <p id="d1e2863">The one-to-one correspondence between the data model and cf-python's
interpretation of CF may also be demonstrated by inspecting the objects from
which field object <monospace>t</monospace> is composed. In Fig. <xref ref-type="fig" rid="Ch1.F18"/>, the
<monospace>constructs</monospace> function is used to return all of these objects, each
having a similar name in camel case to its CF data model counterpart (e.g. a
domain axis construct is represented by a DomainAxis object). This output
demonstrates that it is only the field object which stores information on the
whole domain. For example, the latitude auxiliary coordinate object has a
two-dimensional data array shape of <inline-formula><mml:math id="M82" display="inline"><mml:mrow><mml:mn mathvariant="normal">110</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">106</mml:mn></mml:mrow></mml:math></inline-formula> but does not record what
these dimensions physically represent (Fig. <xref ref-type="fig" rid="Ch1.F10"/>). The cell
method object does, however, store references to the domain axes to which it
applies (Sect. <xref ref-type="sec" rid="Ch1.S4.SS8"/>) – the “dim3” in this example refers to
the size one domain axis object, which is identified by the field object
alone as being a time axis by virtue of the “time” dimension coordinate
object associated with it.</p>
      <p id="d1e2891">Whilst field object <monospace>t</monospace> contains at least one instance of every
type of data model construct, it is more common for field objects to
contain a subset of the possible data
constructs. Figure <xref ref-type="fig" rid="Ch1.F19"/> shows the detailed
description of two other cf-python field objects. The first of these
(<monospace>p</monospace>) is of medium complexity and contains only domain axis,
cell method, and dimension coordinate objects. In this case, the cell
method and time dimension coordinate objects collectively state that
the data are 30-year averages of monthly minima. The second of the
field objects (<monospace>q</monospace>) is minimally complex and contains no other
data model constructs, yet is still CF compliant. In this case, the
data array is scalar and there are no coordinates, so domain axes are
not necessary.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F19" specific-use="star"><caption><p id="d1e2907">Examples of cf-python field objects of medium and minimal
complexity.</p></caption>
        <?xmltex \igopts{width=497.923228pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/10/4619/2017/gmd-10-4619-2017-f19.pdf"/>

      </fig>

<sec id="Ch1.S6.SSx1" specific-use="unnumbered">
  <title>Interpreting CF-netCDF files</title>
      <p id="d1e2922">Any variable in a CF-netCDF file can always be viewed as a data variable in
addition to any metadata role it may have, simply by choosing to ignore any
other variables that may reference it. For example, a variable that is named
by the “coordinates” attribute of a data variable is always an auxiliary
coordinate variable (Sect. <xref ref-type="sec" rid="Ch1.S3.SS3"/>), but this metadata status is
conferred solely by the “coordinates” attribute, so by ignoring it this
variable also becomes a data variable.</p>
      <p id="d1e2927">When a CF-netCDF file is read, a decision must be taken as to which variables
are the data variables. By default, cf-python assumes that only unreferenced
variables are data variables that instantiate field objects (variables
<monospace>temp</monospace> and <monospace>total_wv</monospace> in Fig. <xref ref-type="fig" rid="Ch1.F3"/>). It is possible,
however, to override this default behaviour so that some or all referenced
variables instantiate field objects in addition to instantiating other,
metadata objects. For example, the variable <monospace>PS</monospace> in Fig. <xref ref-type="fig" rid="Ch1.F3"/>
will always create a domain ancillary object but may also, if requested,
create an independent field object for surface air pressure.</p>
      <p id="d1e2943">An interesting situation arises if a netCDF file contains only CF-netCDF
coordinate variables and their associated dimensions. It may be natural to
assume that these coordinate variables define a single domain, but without
the explicit links provided by a data variable the existence of a domain
cannot be assumed by the software. These coordinate variables are not referenced
by a data variable but they are explicitly defined as “coordinates”. When
reading such a file, cf-python by default creates no field objects (because
only coordinates have been defined), and also no dimension coordinate objects
(as dimension coordinate constructs can only exist within a field construct;
see Sect. <xref ref-type="sec" rid="Ch1.S4.SS1"/>). As is the case with referenced variables, it is
possible to override the default behaviour such that an independent field
object is instantiated from each coordinate variable. This example highlights
the issues arising from the absence of an explicit domain element in CF
(Sect. <xref ref-type="sec" rid="Ch1.S3.SS2"/>).</p>
</sec>
</sec>
<sec id="Ch1.S7">
  <title>Evolution of a CF data model and cf-python</title>
      <p id="d1e2957">A number of new features have recently been introduced in version 1.7 of the
CF conventions, published in September 2017, and there is no doubt that the
CF conventions will continue to evolve and meet the needs of the scientific
community for representing more types of data. Any CF data model will
therefore have to adapt to future enhancements. This is why our CF data model
was designed with a minimal number of simple constructs in mind
(Sect. <xref ref-type="sec" rid="Ch1.S1.SS1"/>). We suggest that such a design is more flexible
and therefore more likely (but not guaranteed) to meet future requirements.</p>
      <p id="d1e2962">A CF data model could guide the development of CF by providing a framework
for ensuring that proposed changes fit into CF in a logical, rather than just
a pragmatic, way. A proposed enhancement would be assessed to see if its new
features map onto the existing data model. If they do, then the enhancement
may be incorporated in the CF conventions with no change to the existing data
model.</p>
      <p id="d1e2965">As an example, it is instructive to consider one of the enhancements accepted
into CF at version 1.7, namely the ability to store a cell measure variable
(Sect. <xref ref-type="sec" rid="Ch1.S3.SS5"/>) in a different netCDF file to that of its data
variable (<uri>http://cf-trac.llnl.gov/trac/ticket/145</uri>). This is the first
occurrence in CF of an atomic dataset not being identical to a single file
and poses important questions on how software might be expected to link
multiple files and on the consequences of one of the files going missing,
thereby losing information from the dataset. From a CF data model
perspective, however, this is clearly just an artefact of the dataset
encoding, from which the data model is independent by design. Whether a
dataset is in one file, or spread across many files, the logical
relationships between a field construct and its various metadata constructs,
such as the cell measure construct, remain unchanged.</p>
      <p id="d1e2973">If a change does not map onto the existing data model and cannot be modified
to do so, the data model will need to be modified to accommodate the new
features. This modification will be either backwards compatible or backwards
incompatible. The former (preferable) case occurs if the data model may be
extended or generalised in some way that allows the new features but does not
affect its existing constructs and relationships. The latter case occurs if a
change is required that would affect the interpretation of existing datasets
and the design of software built around the data model. Much effort has been
already been put into avoiding backward incompatibilities in CF, so that
older datasets are still parsable by newer software, but doing so is not a
rule but a “best practice” that could be overridden if the community
consensus were that the benefits in doing so outweighed any inconvenience.</p>
      <p id="d1e2977">It is the authors' intention to ensure that the cf-python software library is
kept up to date with the latest version of CF conventions and the CF data
model presented here. To facilitate this, work is underway to create a
reference implementation of this data model, which in essence will be like
cf-python but without any of its higher-level functionality (such as
regridding methods). This reference implementation will then be imported back
into cf-python to provide not only its data model representation but also a
“read” method for mapping datasets onto field constructs and a “write”
method for mapping field constructs to new netCDF files. The reference
implementation will be easier to understand and maintain than cf-python,
could be used by other Python packages for manipulating CF datasets, and also
has the potential to be used as a test bed for proposed new features of CF.
This reference implementation will be open source (like cf-python), so any
interested parties may contribute to its maintenance and development.</p>
</sec>
<sec id="Ch1.S8" sec-type="conclusions">
  <title>Summary and conclusions</title>
      <p id="d1e2986">In this paper, we have presented a formal data model for the CF conventions,
identifying the fundamental elements of CF and showing how they relate to
each other. We have described the CF conventions in terms of their
relationship to the physical world (real or simulated) and in terms of their
netCDF encoding, and these steps led to our identifying the elements which
contribute to a CF data model. The CF conventions themselves have been
influenced by their netCDF encoding, and therefore our CF data model is
indirectly influenced by netCDF, although it aims to be independent of the
encoding. We have discussed the relationships of our CF data model to other
data models which address the problem of storing data and metadata, and we
have presented a software implementation of this CF data model capable of
manipulating any CF-compliant dataset. We have described possible ways in
which this CF data model and the cf-python software library may evolve over
time.</p>
      <p id="d1e2989">It is important to note that our CF data model is a description of what CF
is, rather than what it ought to be, either in our opinion or anyone else's.
We believe that there is little doubt that a CF data model is of considerable
value, and this has been recognised by the CF community, which highlights that a
CF data model will aid future developments in the CF conventions and make it
easier to create CF-compliant software
(<uri>http://cf-trac.llnl.gov/trac/ticket/88</uri>). In addition, the existence of
a data model can help resolve conflicts in the interpretation of the
conventions document. For example, a discussion on the CF mailing list
regarding the interpretation of CF-netCDF scalar coordinate variables was
resolved with some assistance from a developmental version of the data model
(<uri>http://cf-trac.llnl.gov/trac/ticket/104</uri>). There are other discussions
on unintended ambiguities that have not been resolved, however, and this is
unsatisfactory for those who need to manipulate datasets or create software
for that purpose.</p>
      <p id="d1e2998">Creating an explicit data model before the CF conventions were written would
arguably have been preferable. A data model created a priori increases the
likelihood that the problem space (i.e. storing and manipulating data and
metadata) is fully spanned and encourages coherent implementations, which
could be file storage syntaxes or software codes, the latter being a stated
goal of CF. For example, in CF-netCDF, horizontal and vertical coordinate
reference systems are described with very different structures – the grid
mapping variable and formula_terms attribute, respectively – a situation
that would likely not have occurred if a comprehensive CF data model already
existed. Writing a CF data model a posteriori clearly cannot bring about all
of these benefits, as the coverage of the problem space and file storage
syntax is a given, but it can still be of use to software implementations and
future developments in the conventions.</p>
      <p id="d1e3001">We believe that the data model proposed here is a complete and correct
description of CF, because we have yet to find a case for which our
implementation in the cf-python library fails to represent or misrepresents a
CF-compliant dataset. Moreover, the development of cf-python proves that is
possible to implement our CF data model. We consider that our CF data model
is simpler and more flexible than other such models, because it defines a
small number of general constructs rather than many specialised ones. While
the latter approach is closer to an object-orientated software
implementation, our aim is to describe CF in a way which is independent of
any software.</p>
      <p id="d1e3005">If this CF data model were to be accepted by the community as a formal part
of the CF conventions, then any future enhancements would have to be
incorporated into the data model as part of the public discussion that leads
to the acceptance of every enhancement. Version 1.7 of <?xmltex \hack{\vadjust{\newpage}}?>the CF conventions has
been recently published and it is the authors' intention to review all of the
new features for compatibility with this CF data model. As these enhancements
have already been finalised, any conflict will necessarily force a change in
the data model. Once up to date with version 1.7, the data model may then be
considered in parallel with the discussions on enhancements for subsequent
releases. Structural differences between different versions of the CF
conventions would be plain to see if each release contains a data model, thus
making it easier to write software that can cope with any backward
incompatibilities that may have been introduced.</p>
</sec>

      
      </body>
    <back><notes notes-type="codeavailability">

      <p id="d1e3014">The code of cf-python is open source and freely
downloadable at <uri>https://doi.org/10.5281/zenodo.832255</uri>
(<xref ref-type="bibr" rid="bib1.bibx5" id="altparen.16"/>). It is also available from its online repository at
<uri>https://bitbucket.org/cfpython/cf-python</uri> and from the Python package
index at <uri>https://pypi.python.org/pypi/cf-python</uri>.</p>
  </notes><?xmltex \hack{\clearpage}?><app-group>

<app id="App1.Ch1.S1">
  <title>A UML primer</title>
      <p id="d1e3038">Throughout this paper, we rely on UML to
construct diagrams that define the key relationships of the entities
described in CF-netCDF files and in our data model. These diagrams show
relationships between “classes” like those used in an object-orientated
programming language or like data types in Fortran. The relationship of an
“instance of a class” to its class is like that of a particular variable to
its data type. A class is like a species of animal, and an instance of a
class is like an individual animal. Classes can be included in other classes,
just as components are included in the definitions of derived data types in
Fortran, and organs comprise the body of an animal.</p>
      <p id="d1e3041">For reference in interpreting our UML diagrams, we describe the subset of UML
used here. As depicted in Table <xref ref-type="table" rid="Ch1.T1"/>, arrows and symbols
are used to show different types of relationship between classes. Some
relationships include a “cardinality” which indicates the number of
instances of one class that may be associated with an instance of another. If
there is a number (for instance, <inline-formula><mml:math id="M83" display="inline"><mml:mi>n</mml:mi></mml:math></inline-formula>) present at class <inline-formula><mml:math id="M84" display="inline"><mml:mi>Y</mml:mi></mml:math></inline-formula> where there is an
arrow from class <inline-formula><mml:math id="M85" display="inline"><mml:mi>X</mml:mi></mml:math></inline-formula> to class <inline-formula><mml:math id="M86" display="inline"><mml:mi>Y</mml:mi></mml:math></inline-formula>, it indicates that there must be exactly
<inline-formula><mml:math id="M87" display="inline"><mml:mi>n</mml:mi></mml:math></inline-formula> instances of class <inline-formula><mml:math id="M88" display="inline"><mml:mi>Y</mml:mi></mml:math></inline-formula> associated with class <inline-formula><mml:math id="M89" display="inline"><mml:mi>X</mml:mi></mml:math></inline-formula>. These cardinalities can
also be associated with ranges. For example, 0..1 means zero or one
instance(s) of class <inline-formula><mml:math id="M90" display="inline"><mml:mi>Y</mml:mi></mml:math></inline-formula> may be associated with class <inline-formula><mml:math id="M91" display="inline"><mml:mi>X</mml:mi></mml:math></inline-formula>, and <inline-formula><mml:math id="M92" display="inline"><mml:mrow><mml:mn mathvariant="normal">0</mml:mn><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>*</mml:mo></mml:mrow></mml:math></inline-formula> means
any number of instances of class <inline-formula><mml:math id="M93" display="inline"><mml:mi>Y</mml:mi></mml:math></inline-formula> may be associated with class <inline-formula><mml:math id="M94" display="inline"><mml:mi>X</mml:mi></mml:math></inline-formula>. All of
these relationships, along with techniques used to add further information to
classes and associations, are shown in the worked example of
Fig. <xref ref-type="fig" rid="App1.Ch1.F1"/>.</p>
      <p id="d1e3141">The UML diagram elements relating to netCDF (Sect. <xref ref-type="sec" rid="Ch1.S2"/>), the
CF-netCDF encoding and our CF data model (Sect. <xref ref-type="sec" rid="Ch1.S4"/>) are
coloured yellow, blue, and green, respectively. In addition, an element from a
data model that is not the main focus of the diagram has its name prefixed
with an identifier for its model – “NC” for netCDF and “CN” for
CF-netCDF. For example, in Fig. <xref ref-type="fig" rid="Ch1.F8"/>, which is focused on
the CF-netCDF conventions, the yellow “NC::Dimension” element is the same
as the “Dimension” element from Fig. <xref ref-type="fig" rid="Ch1.F2"/>, the main diagram for
the netCDF data model.</p><?xmltex \hack{\newpage}?><?xmltex \floatpos{h!}?><fig id="App1.Ch1.F1"><caption><p id="d1e3154">A worked example demonstrating the subset of UML used in
this paper (see Table <xref ref-type="table" rid="Ch1.T1"/> for definitions of
the class associations). Class-B is a subclass of class-A, and
class-D is subclass of class-E. An instance of class-B includes
one instance of class-C (that cannot exist independently) and
may include zero or one instance(s) of class-F (that can exist
independently). An instance of class-B is related to any number
of instances of class-D (with the relationship being described
by the label “Association”). An instance of class-C is
constrained to exhibit some behaviour. There is a general
comment concerning class-D and class-C.</p></caption>
        <?xmltex \igopts{width=193.47874pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/10/4619/2017/gmd-10-4619-2017-f20.pdf"/>

      </fig>

<?xmltex \hack{\clearpage}?>
</app>

<app id="App1.Ch1.S2">
  <?xmltex \opttitle{Optimising dataset storage and\hack{\break} creation in CF-netCDF}?><title>Optimising dataset storage and<?xmltex \hack{\break}?> creation in CF-netCDF</title>
      <p id="d1e3176">An important goal of the CF conventions is that datasets should be efficient
to create, store, and subsequently read, where efficiency is a measure of the
time taken for software to carry out a task or the amount of computer data
storage required for a file. The conventions describe various techniques for
optimising these requirements. A CF-netCDF file that uses any of the
optimisation techniques can always be recast without them and still contain
exactly the same scientific information; therefore, the optimisation
mechanisms do not affect the data model described in
Sect. <xref ref-type="sec" rid="Ch1.S4"/>. It is common for a CF-netCDF file to not require
optimisation, but in the cases where it may be applied, an unoptimised file
could suffer by being less readable by humans, consuming more storage or
being slower to create. The example file shown in Fig. <xref ref-type="fig" rid="Ch1.F3"/> does not
include any of these techniques, but many examples may be found in the CF
conventions document <xref ref-type="bibr" rid="bib1.bibx3" id="paren.17"/>.</p>
      <p id="d1e3186">These parts of the CF convention were devised because the netCDF
classic file format and API do not
offer any methods for compression. However, the netCDF-4 API supports
lossless compression of variables stored in files
<xref ref-type="bibr" rid="bib1.bibx12" id="paren.18"/>. This method does not affect the variables
as they appear to the user of the data and hence has no impact on our
CF data model.</p>
<sec id="App1.Ch1.S2.SS1">
  <title>Packing</title>
      <p id="d1e3197">Storage space in netCDF files may be reduced by a packing, i.e. by altering
the data in a way that reduces their precision. Lossy compression may be
essential for the archiving of the huge volumes of data produced by modern
high-resolution models <xref ref-type="bibr" rid="bib1.bibx1" id="paren.19"/>. This is achieved
through the simple use of the variable attributes <monospace>scale_factor</monospace> and
<monospace>add_offset</monospace>. After the data values of a variable have been read,
they are to be multiplied by the <monospace>scale_factor</monospace>, and have
<monospace>add_offset</monospace> added to them (if both attributes are present, the data
are scaled before the offset is added). Unpacked values are assumed to have
the same data type as the packing attributes, thus making it possible to
store 64-bit floating point data as 16-bit unsigned integers, for instance. In this
example, a loss of precision is likely to arise because an unpacked value can
only take one of <inline-formula><mml:math id="M95" display="inline"><mml:mrow><mml:msup><mml:mn mathvariant="normal">2</mml:mn><mml:mn mathvariant="normal">16</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> possible values.</p>
</sec>
<sec id="App1.Ch1.S2.SS2">
  <title>Compression</title>
      <p id="d1e3233">As well as external methods of compression applied to the file, CF has
support for space saving by identifying unwanted missing data. Such
compression techniques store the data more efficiently and result in no
precision loss.</p><?xmltex \hack{\newpage}?>
<sec id="App1.Ch1.S2.SS2.SSS1">
  <title>Gathering</title>
      <p id="d1e3242">Compression by gathering combines axes of a multi-dimensional array into a
new, discrete axis (the “list” dimension) whilst omitting the missing
values and thus reducing the number of values that need to be stored. The
information needed to uncompress the data is stored in a separate variable
(the “list” variable) that contains the indices needed to uncompress the
data. A list variable is encoded as a coordinate variable that has a
<monospace>compress</monospace> attribute which names the dimensions that have been
compressed. For example, a variable that spans <inline-formula><mml:math id="M96" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M97" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula>, and <inline-formula><mml:math id="M98" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> axes but has
missing data values at all points over the ocean could have its <inline-formula><mml:math id="M99" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M100" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula>
dimensions compressed to a new dimension (called “landpoint”, for instance)
whose size is the number of land points. The stored variable would then span
only the landpoint and <inline-formula><mml:math id="M101" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> dimensions.</p>
</sec>
<sec id="App1.Ch1.S2.SS2.SSS2">
  <title>Ragged array representations</title>
      <p id="d1e3297">A collection of DSG features may be stored using
the contiguous or indexed ragged array representation, which minimises the
amount of file storage required (Sect. <xref ref-type="sec" rid="Ch1.S3.SS4"/>). In both cases, the
“instance” dimension that distinguishes between different features is
combined with the number of elements of each feature to create a compressed
“sample” dimension. The entire collection may then be stored in an array
that spans the sample dimension, and the contiguous and ragged array
representations provide different techniques for populating this array and
uncompressing it to find the values for individual features.</p>
      <p id="d1e3302">In the contiguous case, each feature in the collection occupies a contiguous
block, and so can be used only if the size of each feature is known at the
time that it is created. It requires a “count” variable that gives the size
of each block and is encoded as a netCDF variable with a
<monospace>sample_dimension</monospace> attribute that names the sample dimension.</p>
      <p id="d1e3308">For indexed ragged arrays, the values of each feature in the collection are
interleaved along the sample dimension. The canonical use case for this
representation is the storage of real-time data streams that contain reports
from many sources; the data can be written as they arrive. It requires an
“index” variable that specifies the feature that each element of the sample
dimension belongs to and is encoded as a netCDF variable with an
<monospace>instance_dimension</monospace> attribute that names the instance dimension.</p>
      <p id="d1e3314">It is also possible to combine contiguous and indexed ragged array
representations, which is useful for cases such as writing real-time data
streams that contain vertical profiles from many trajectories, arriving
randomly, with the data for each entire profile written all at once.</p><?xmltex \hack{\clearpage}?>
</sec>
</sec>
</app>
  </app-group><notes notes-type="competinginterests">

      <p id="d1e3325">The authors declare that they have no conflict of
interest.</p>
  </notes><ack><title>Acknowledgements</title><p id="d1e3331">We would like to thank Mark Hedley, Antonio Cofiño, Martin Juckes,
Alison Pamment, and Paulo Ceppi for comments that greatly improved the
manuscript. We are also indebted to members of the CF community, whose
considerable efforts ensure the continuing success of the CF conventions – in
particular, those who took part in the data model discussions that took place
on the CF mailing list.</p><p id="d1e3333">The research leading to these results has received funding from the core
budget of the UK National Centre for Atmospheric Science, the European
Research Council, and the European Commission's Seventh Framework programme
(from ERC project “Seachange”, number 247220; and FW7 project “IS-ENES2”,
number 312979). Work by Karl E. Taylor was performed under the auspices of
the US Department of Energy (USDOE) by Lawrence Livermore National Laboratory
under contract DE-AC52-07NA27344 with support from the Regional and Global
Climate Modeling Program of the USDOE's Office of Science.<?xmltex \hack{\newline}?><?xmltex \hack{\newline}?> Edited by: Steve Easterbrook<?xmltex \hack{\newline}?> Reviewed by:
Venkatramani Balaji and Brian Eaton</p></ack><ref-list>
    <title>References</title>

      <ref id="bib1.bibx1"><label>Baker et al.(2016)Baker, Hammerling, Mickelson, Xu, Stolpe, Naveau,
Sanderson, Ebert-Uphoff, Samarasinghe, De Simone, Carbone, Gencarelli,
Dennis, Kay, and Lindstrom</label><mixed-citation>Baker, A. H., Hammerling, D. M., Mickelson, S. A., Xu, H., Stolpe, M. B.,
Naveau, P., Sanderson, B., Ebert-Uphoff, I., Samarasinghe, S., De Simone, F.,
Carbone, F., Gencarelli, C. N., Dennis, J. M., Kay, J. E., and Lindstrom, P.:
Evaluating lossy data compression on climate simulation data within a large
ensemble, Geosci. Model Dev., 9, 4381–4403,
<ext-link xlink:href="https://doi.org/10.5194/gmd-9-4381-2016" ext-link-type="DOI">10.5194/gmd-9-4381-2016</ext-link>, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx2"><label>Dominico and Nativi(2013)</label><mixed-citation>
Dominico, B. and Nativi, S. (Eds.): CF-netCDF3 Data Model Extension
Standard, no. OGC 11-165r2 in Open GIS Standard, Open Geospatial
Consortium, 3.1rd Edn., Wayland, MA, USA, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx3"><label>Eaton et al.(2011)Eaton, Gregory, Drach, Taylor, Hankin, Caron,
Signell, Bentley, Rappa, Höck, Pamment, and Juckes</label><mixed-citation>Eaton, B., Gregory, J., Drach, B., Taylor, K., Hankin, S., Caron, J.,
Signell,
R., Bentley, P., Rappa, G., Höck, H., Pamment, A., and Juckes, M.:
NetCDF Climate and Forecast (CF) Metadata Conventions V1.6,
available at:
<uri>http://cfconventions.org/cf-conventions/v1.6.0/cf-conventions.html</uri>
(last access: 11 December 2017),
2011.
</mixed-citation></ref><?xmltex \hack{\newpage}?>
      <ref id="bib1.bibx4"><label>Emmerson(2007)</label><mixed-citation>Emmerson, S.: UDUNITS-2 package, available at:
<uri>http://www.unidata.ucar.edu/software/udunits</uri> (last access: 11 December 2017), 2007.</mixed-citation></ref>
      <ref id="bib1.bibx5"><label>Hassell and Gregory(2017)</label><mixed-citation>Hassell, D. and Gregory, J.: cf-python, <ext-link xlink:href="https://doi.org/10.5281/zenodo.832255" ext-link-type="DOI">10.5281/zenodo.832255</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx6"><label>International Standards Organisation(2007)</label><mixed-citation>
International Standards Organisation: ISO19123: Geographic
Information; Schema for Coverage Geometry and Functions, ISO, Geneva,
Switzerland, 2007.</mixed-citation></ref>
      <ref id="bib1.bibx7"><label>International Standards Organisation(2011)</label><mixed-citation>
International Standards Organisation: ISO19156: Geographic
Information – Observations and Measurements, ISO, Geneva, 2011.</mixed-citation></ref>
      <ref id="bib1.bibx8"><label>Nativi et al.(2008)Nativi, Caron, Domenico, and Bigagli</label><mixed-citation>Nativi, S., Caron, J., Domenico, B., and Bigagli, L.: Unidata's Common Data
Model Mapping to the ISO 19123 Data Model, Earth Sci. Inform., 1, 59–78, <ext-link xlink:href="https://doi.org/10.1007/s12145-008-0011-6" ext-link-type="DOI">10.1007/s12145-008-0011-6</ext-link>, 2008.</mixed-citation></ref>
      <ref id="bib1.bibx9"><label>O'Kuinghttons et al.(2016)O'Kuinghttons, Koziol, Oehmke, DeLuca,
Theurich, Li, and Jacob</label><mixed-citation>
O'Kuinghttons, R., Koziol, B., Oehmke, R., DeLuca, C., Theurich, G., Li, P.,
and Jacob, J.: ESMPy and OpenClimateGIS: Python Interfaces for
High Performance Grid Remapping and Geospatial Dataset Manipulation,
Geophys. Res. Abstr.,
EGU2016-10050, EGU General Assembly 2016, Vienna, Austria, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx10"><label>Rew and Davis(1990)</label><mixed-citation>Rew, R. and Davis, G.: NetCDF: An Interface for Scientific Data Access,
IEEE Computer Graphics and Applications, 10, 76–82, <ext-link xlink:href="https://doi.org/10.1109/38.56302" ext-link-type="DOI">10.1109/38.56302</ext-link>,
1990.</mixed-citation></ref>
      <ref id="bib1.bibx11"><label>Rew et al.(1997)Rew, Davis, Emmerson, and Davies</label><mixed-citation>Rew, R., Davis, G., Emmerson, S., and Davies, H.: NetCDF User's Guide,
available at:
<uri>http://www.unidata.ucar.edu/software/netcdf/docs/user_guide.html</uri> (last access: 11 December 2017), 1997.</mixed-citation></ref>
      <ref id="bib1.bibx12"><label>Rew et al.(2006)Rew, Hartnett, and Caron</label><mixed-citation>
Rew, R., Hartnett, E., and Caron, J.: NetCDF-4: Software Implementing
an Enhanced Data Model for the Geosciences, AMS, Atlanta, 2006.</mixed-citation></ref>
      <ref id="bib1.bibx13"><label>Unidata(2014a)</label><mixed-citation>Unidata: Common Data Model, available at:
<uri>http://www.unidata.ucar.edu/software/thredds/current/netcdf-java/CDM/</uri> (last access: 11 December 2017),
2014a.</mixed-citation></ref>
      <ref id="bib1.bibx14"><label>Unidata(2014b)</label><mixed-citation>Unidata: Point Feature Datasets, available at:
<uri>http://www.unidata.ucar.edu/software/thredds/current/netcdf-java/reference/FeatureDatasets/PointFeatures.html</uri> (last access: 11 December 2017), 2014b.</mixed-citation></ref>

  </ref-list><app-group content-type="float"><app><title/>

    </app></app-group></back>
    <!--<article-title-html>A data model of the Climate and Forecast metadata conventions (CF-1.6) with a software implementation (cf-python v2.1)</article-title-html>
<abstract-html><p class="p">The CF (Climate and Forecast) metadata conventions are designed to promote
the creation, processing, and sharing of climate and forecasting data using
Network Common Data Form (netCDF) files and libraries. The CF conventions
provide a description of the physical meaning of data and of their spatial
and temporal properties, but they depend on the netCDF file encoding which
can currently only be fully understood and interpreted by someone familiar
with the rules and relationships specified in the conventions documentation.
To aid in development of CF-compliant software and to capture with a minimal
set of elements all of the information contained in the CF conventions, we
propose a formal data model for CF which is independent of netCDF and
describes all possible CF-compliant data. Because such data will often be
analysed and visualised using software based on other data models, we compare
our CF data model with the ISO 19123 coverage model, the Open Geospatial
Consortium CF netCDF standard, and the Unidata Common Data Model. To
demonstrate that this CF data model can in fact be implemented, we present
cf-python, a Python software library that conforms to the model and can
manipulate any CF-compliant dataset.</p></abstract-html>
<ref-html id="bib1.bib1"><label>Baker et al.(2016)Baker, Hammerling, Mickelson, Xu, Stolpe, Naveau,
Sanderson, Ebert-Uphoff, Samarasinghe, De Simone, Carbone, Gencarelli,
Dennis, Kay, and Lindstrom</label><mixed-citation>
Baker, A. H., Hammerling, D. M., Mickelson, S. A., Xu, H., Stolpe, M. B.,
Naveau, P., Sanderson, B., Ebert-Uphoff, I., Samarasinghe, S., De Simone, F.,
Carbone, F., Gencarelli, C. N., Dennis, J. M., Kay, J. E., and Lindstrom, P.:
Evaluating lossy data compression on climate simulation data within a large
ensemble, Geosci. Model Dev., 9, 4381–4403,
<a href="https://doi.org/10.5194/gmd-9-4381-2016" target="_blank">https://doi.org/10.5194/gmd-9-4381-2016</a>, 2016.
</mixed-citation></ref-html>
<ref-html id="bib1.bib2"><label>Dominico and Nativi(2013)</label><mixed-citation>
Dominico, B. and Nativi, S. (Eds.): CF-netCDF3 Data Model Extension
Standard, no. OGC 11-165r2 in Open GIS Standard, Open Geospatial
Consortium, 3.1rd Edn., Wayland, MA, USA, 2013.
</mixed-citation></ref-html>
<ref-html id="bib1.bib3"><label>Eaton et al.(2011)Eaton, Gregory, Drach, Taylor, Hankin, Caron,
Signell, Bentley, Rappa, Höck, Pamment, and Juckes</label><mixed-citation>
Eaton, B., Gregory, J., Drach, B., Taylor, K., Hankin, S., Caron, J.,
Signell,
R., Bentley, P., Rappa, G., Höck, H., Pamment, A., and Juckes, M.:
NetCDF Climate and Forecast (CF) Metadata Conventions V1.6,
available at:
<a href="http://cfconventions.org/cf-conventions/v1.6.0/cf-conventions.html" target="_blank">http://cfconventions.org/cf-conventions/v1.6.0/cf-conventions.html</a>
(last access: 11 December 2017),
2011.

</mixed-citation></ref-html>
<ref-html id="bib1.bib4"><label>Emmerson(2007)</label><mixed-citation>
Emmerson, S.: UDUNITS-2 package, available at:
<a href="http://www.unidata.ucar.edu/software/udunits" target="_blank">http://www.unidata.ucar.edu/software/udunits</a> (last access: 11 December 2017), 2007.
</mixed-citation></ref-html>
<ref-html id="bib1.bib5"><label>Hassell and Gregory(2017)</label><mixed-citation>
Hassell, D. and Gregory, J.: cf-python, <a href="https://doi.org/10.5281/zenodo.832255" target="_blank">https://doi.org/10.5281/zenodo.832255</a>, 2017.
</mixed-citation></ref-html>
<ref-html id="bib1.bib6"><label>International Standards Organisation(2007)</label><mixed-citation>
International Standards Organisation: ISO19123: Geographic
Information; Schema for Coverage Geometry and Functions, ISO, Geneva,
Switzerland, 2007.
</mixed-citation></ref-html>
<ref-html id="bib1.bib7"><label>International Standards Organisation(2011)</label><mixed-citation>
International Standards Organisation: ISO19156: Geographic
Information – Observations and Measurements, ISO, Geneva, 2011.
</mixed-citation></ref-html>
<ref-html id="bib1.bib8"><label>Nativi et al.(2008)Nativi, Caron, Domenico, and Bigagli</label><mixed-citation>
Nativi, S., Caron, J., Domenico, B., and Bigagli, L.: Unidata's Common Data
Model Mapping to the ISO 19123 Data Model, Earth Sci. Inform., 1, 59–78, <a href="https://doi.org/10.1007/s12145-008-0011-6" target="_blank">https://doi.org/10.1007/s12145-008-0011-6</a>, 2008.
</mixed-citation></ref-html>
<ref-html id="bib1.bib9"><label>O'Kuinghttons et al.(2016)O'Kuinghttons, Koziol, Oehmke, DeLuca,
Theurich, Li, and Jacob</label><mixed-citation>
O'Kuinghttons, R., Koziol, B., Oehmke, R., DeLuca, C., Theurich, G., Li, P.,
and Jacob, J.: ESMPy and OpenClimateGIS: Python Interfaces for
High Performance Grid Remapping and Geospatial Dataset Manipulation,
Geophys. Res. Abstr.,
EGU2016-10050, EGU General Assembly 2016, Vienna, Austria, 2016.
</mixed-citation></ref-html>
<ref-html id="bib1.bib10"><label>Rew and Davis(1990)</label><mixed-citation>
Rew, R. and Davis, G.: NetCDF: An Interface for Scientific Data Access,
IEEE Computer Graphics and Applications, 10, 76–82, <a href="https://doi.org/10.1109/38.56302" target="_blank">https://doi.org/10.1109/38.56302</a>,
1990.
</mixed-citation></ref-html>
<ref-html id="bib1.bib11"><label>Rew et al.(1997)Rew, Davis, Emmerson, and Davies</label><mixed-citation>
Rew, R., Davis, G., Emmerson, S., and Davies, H.: NetCDF User's Guide,
available at:
<a href="http://www.unidata.ucar.edu/software/netcdf/docs/user_guide.html" target="_blank">http://www.unidata.ucar.edu/software/netcdf/docs/user_guide.html</a> (last access: 11 December 2017), 1997.
</mixed-citation></ref-html>
<ref-html id="bib1.bib12"><label>Rew et al.(2006)Rew, Hartnett, and Caron</label><mixed-citation>
Rew, R., Hartnett, E., and Caron, J.: NetCDF-4: Software Implementing
an Enhanced Data Model for the Geosciences, AMS, Atlanta, 2006.
</mixed-citation></ref-html>
<ref-html id="bib1.bib13"><label>Unidata(2014a)</label><mixed-citation>
Unidata: Common Data Model, available at:
<a href="http://www.unidata.ucar.edu/software/thredds/current/netcdf-java/CDM/" target="_blank">http://www.unidata.ucar.edu/software/thredds/current/netcdf-java/CDM/</a> (last access: 11 December 2017),
2014a.
</mixed-citation></ref-html>
<ref-html id="bib1.bib14"><label>Unidata(2014b)</label><mixed-citation>
Unidata: Point Feature Datasets, available at:
<a href="http://www.unidata.ucar.edu/software/thredds/current/netcdf-java/reference/FeatureDatasets/PointFeatures.html" target="_blank">http://www.unidata.ucar.edu/software/thredds/current/netcdf-java/reference/FeatureDatasets/PointFeatures.html</a> (last access: 11 December 2017), 2014b.
</mixed-citation></ref-html>--></article>
