<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing with OASIS Tables v3.0 20080202//EN" "journalpub-oasis3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:oasis="http://docs.oasis-open.org/ns/oasis-exchange/table" xml:lang="en" dtd-version="3.0" article-type="research-article"><?xmltex \bartext{Model description paper}?>
  <front>
    <journal-meta><journal-id journal-id-type="publisher">GMD</journal-id><journal-title-group>
    <journal-title>Geoscientific Model Development</journal-title>
    <abbrev-journal-title abbrev-type="publisher">GMD</abbrev-journal-title><abbrev-journal-title abbrev-type="nlm-ta">Geosci. Model Dev.</abbrev-journal-title>
  </journal-title-group><issn pub-type="epub">1991-9603</issn><publisher>
    <publisher-name>Copernicus Publications</publisher-name>
    <publisher-loc>Göttingen, Germany</publisher-loc>
  </publisher></journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.5194/gmd-15-9015-2022</article-id><title-group><article-title>Predicting peak daily maximum 8 h ozone and linkages to emissions and
meteorology in Southern California using machine learning methods
(SoCAB-8HR V1.0)</article-title><alt-title>Predicting peak daily maximum 8 h ozone, and linkages to emissions and
meteorology</alt-title>
      </title-group><?xmltex \runningtitle{Predicting peak daily maximum 8\,h ozone, and linkages to emissions and
meteorology}?><?xmltex \runningauthor{Z.~Gao et al.}?>
      <contrib-group>
        <contrib contrib-type="author" corresp="yes" rid="aff1">
          <name><surname>Gao</surname><given-names>Ziqi</given-names></name>
          <email>zgao71@gatech.edu</email>
        </contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1">
          <name><surname>Wang</surname><given-names>Yifeng</given-names></name>
          
        </contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1">
          <name><surname>Vasilakos</surname><given-names>Petros</given-names></name>
          
        </contrib>
        <contrib contrib-type="author" corresp="no" rid="aff2 aff4">
          <name><surname>Ivey</surname><given-names>Cesunica E.</given-names></name>
          
        </contrib>
        <contrib contrib-type="author" corresp="no" rid="aff2 aff3">
          <name><surname>Do</surname><given-names>Khanh</given-names></name>
          
        </contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1">
          <name><surname>Russell</surname><given-names>Armistead G.</given-names></name>
          
        <ext-link>https://orcid.org/0000-0003-2027-8870</ext-link></contrib>
        <aff id="aff1"><label>1</label><institution>School of Civil and Environmental Engineering, Georgia Institute of
Technology, Atlanta, GA 30332, USA</institution>
        </aff>
        <aff id="aff2"><label>2</label><institution>Department of Chemical and Environmental Engineering, University of
California, Riverside, Riverside, CA 92521, USA</institution>
        </aff>
        <aff id="aff3"><label>3</label><institution>Center for Environmental Research and Technology, University of
California, Riverside, Riverside, CA 92521,
USA</institution>
        </aff>
        <aff id="aff4"><label>a</label><institution>now at: Department of Civil and Environmental Engineering,
University of California, Berkeley, Berkeley, CA 94720, USA</institution>
        </aff>
      </contrib-group>
      <author-notes><corresp id="corr1">Ziqi Gao (zgao71@gatech.edu)</corresp></author-notes><pub-date><day>16</day><month>December</month><year>2022</year></pub-date>
      
      <volume>15</volume>
      <issue>24</issue>
      <fpage>9015</fpage><lpage>9029</lpage>
      <history>
        <date date-type="received"><day>24</day><month>May</month><year>2022</year></date>
           <date date-type="rev-request"><day>28</day><month>July</month><year>2022</year></date>
           <date date-type="rev-recd"><day>24</day><month>November</month><year>2022</year></date>
           <date date-type="accepted"><day>2</day><month>December</month><year>2022</year></date>
      </history>
      <permissions>
        <copyright-statement>Copyright: © 2022 </copyright-statement>
        <copyright-year>2022</copyright-year>
      <license license-type="open-access"><license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p></license></permissions><self-uri xlink:href="https://gmd.copernicus.org/articles/.html">This article is available from https://gmd.copernicus.org/articles/.html</self-uri><self-uri xlink:href="https://gmd.copernicus.org/articles/.pdf">The full text article is available as a PDF file from https://gmd.copernicus.org/articles/.pdf</self-uri>
      <abstract><title>Abstract</title>

      <p id="d1e148">The growing abundance of data is conducive to using numerical
methods to relate air quality, meteorology and emissions to address which
factors impact pollutant concentrations. Often, it is the extreme values
that are of interest for health and regulatory purposes (e.g., the National
Ambient Air Quality Standard for ozone uses the annual maximum daily
fourth highest 8 h average (MDA8) ozone), though such values are the
most challenging to predict using empirical models. We developed four
different computational models, including the generalized additive model
(GAM), multivariate adaptive regression splines, random forest, and
support vector regression, to develop observation-based relationships
between the fourth highest MDA8 ozone in the South Coast Air Basin and
precursor emissions, meteorological factors and large-scale climate
patterns. All models had similar predictive performance, though the GAM
showed a relatively higher <inline-formula><mml:math id="M1" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> value (0.96) with a lower root mean
square error and mean bias.</p>
  </abstract>
    </article-meta>
  </front>
<body>
      

<sec id="Ch1.S1" sec-type="intro">
  <label>1</label><title>Introduction</title>
      <p id="d1e171">Tropospheric ozone has proven to be one of the most difficult air pollutants
to control, especially in the South Coast Air Basin (SoCAB) of California,
which includes the city of Los Angeles and parts of four counties with a
2020 population exceeding 18 million. Exposure to ozone can be harmful to
human health, leading to a variety of adverse outcomes, including premature
mortality (U.S. EPA, 2020), climate warming and decreased agricultural
production (Ainsworth et al., 2012; Hong et al., 2020). Ozone is formed
by chemical reactions between volatile organic compounds (VOCs) and nitrogen
oxides (<inline-formula><mml:math id="M2" display="inline"><mml:mrow class="chem"><mml:msub><mml:mi mathvariant="normal">NO</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) in the presence of sunlight (Seinfeld and Pandis,
2016). In addition to VOC and <inline-formula><mml:math id="M3" display="inline"><mml:mrow class="chem"><mml:msub><mml:mi mathvariant="normal">NO</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> emissions, meteorology and large-scale
climate patterns affect ozone (Aw and Kleeman, 2003; Blanchard et al.,
2014; Gorai et al., 2015; Kelley et al., 2020; Kleeman, 2008; Lu et al.,
2019; Mahmud et al., 2010; Mcglynn et al., 2018). As such, the resulting
relationships among ozone, emissions and meteorology are complex and
difficult to model accurately. However, the rise of machine learning
methods, along with an increasingly long observational record, suggests that
observation-based models can be used to understand those relationships. Since the 17th
century statistics have been used to record information about the wealth and
population in Europe (Porter, 1981). For example, William Petty, a British scientist and economist, estimated the census data of
Ireland through statistics (Banta, 1987). While the application of
statistics had been restricted to a few fields until the 19th century,
it gradually extended to other areas since then including physics,
astronomy and recently air quality (Porter, 1995). At their core,
statistical models aim at approximating a relationship between dependent and
independent variables, with regression being the most commonly used method,
a term that was coined by British statistician Francis Galton back in 1885
when he studied the trend of heights within families (Galton, 1889, 1888; Benirschke, 2004). The method however precedes the name, with
the use of regression starting years before the term was introduced, dating
back to the beginning of the 19th century with linear regression being
applied to questions in astronomy, such as determining orbits of comets,
while the least-squares method attributed to Adrien-Marie Legendre and Carl Friedrich Gauss was developed in the early 1800s (Stephen, 1981;
Agarwal and Sen, 2014). At the start of the 20th century, some
statisticians introduced the idea of nonlinear regression, trying to
explain more complex systems (Fisher, 1922). Since then,
as computational capacity increased dramatically in the past few decades,
regression analysis has been widely used in most scientific fields.</p>
      <p id="d1e196">The US Environmental Protection Agency's (EPA) National Ambient Air Quality
Standard (NAAQS) for ozone is based on the annual maximum daily fourth
highest 8 h average (MDA8) ozone observations, which is an extreme
statistic, and extreme statistics are often difficult to accurately predict
using empirical modeling, though different approaches have been used for
various purposes. For example, the US EPA adjusted the MDA8 ozone predictions
with meteorological observations using generalized linear modeling (GLM)
with natural spline smoothing functions in the R program to develop a
generalized additive model (GAM) (Camalier et al., 2007; Wells et al.,
2021) that meteorologically adjusts ozone trends to help isolate the impact
of emissions. The GAM is an extension of the GLM, which was introduced in
1986 (Hastie and Tibshirani, 1986, 1990). It is more flexible than
the GLM due to the smoothing functions on independent variables. Previous
studies suggested the GAM was useful to deal with the nonlinear relationship
between MDA8 ozone concentrations and meteorological indicators. About
40 % to 90 % of the variance of the MDA8 ozone concentrations could be
explained at different sites with meteorologically adjusted GAMs (Aldrin
and Haff, 2005; Blanchard et al., 2014, 2019; Camalier et
al., 2007; Flynn et al., 2021; Gong et al., 2018, 2017; Hu et
al., 2021; Huang et al., 2020; Jeong et al., 2020; Ma et al., 2020; McClure
and Jaffe, 2018; Pearce et al., 2011; Gao et al., 2022). GAMs can assess
each independent variable's contribution to the dependent variable. The
multivariate adaptive regression splines (MARS) model (Friedman,
1991) has been used to model the nonlinear relationship between ozone
concentrations and precursors' concentrations/meteorological factors,
including the interactions between the independent indicators
(García Nieto and Álvarez Antón, 2014; Roy et al., 2018).
Support vector regression (SVR) is an extension of the support vector
machine (SVM) (Drucker et al., 1996; Rodríguez-Pérez et al.,
2017; Smola and Schölkopf, 2004). Past studies have shown that the SVR
model with kernel functions can fit the nonlinear relationships between
ozone concentrations and meteorological factors and can obtain accurate
predictions (Liu et al., 2017; Luna et al., 2014; Rybarczyk and
Zalakeviciute, 2018; Sotomayor-Olmedo et al., 2013; Vong et al., 2012).
Random forest (RF) is a machine learning method (Tin Kam, 1995)
derived from the traditional decision tree method. Compared to the
traditional method, it is more accurate because it contains multiple
decision trees. The RF model can be used to fit nonlinear relationships and
deal with interaction effects. It can accurately predict ozone
concentrations using meteorological variables and emissions and capture
about 70 % to 95 % of the variability in ozone concentrations (Keller
and Evans, 2019; Pernak et al., 2019; Stafoggia et al., 2020; Zhan et al.,
2018). However, most prior empirical-model applications to simulate peak
MDA8 ozone levels were biased low, especially when considering capturing the
fourth highest annual MDA8 ozone concentrations.</p>
      <p id="d1e199">In this study, we develop observation-based models (SoCAB-8HR V1.0) using
four different methods (GAM, RF, SVR and MARS) with a broad range of
potential independent indicators that impact ozone formation (e.g.,
precursors emissions, meteorological conditions, large-scale climate events,
chemical reactions, seasonal variations and weekend effects) to predict the
annual fourth highest MDA8 ozone in the SoCAB from 1990 to 2019. We assess and
compare model performance and their applicability to help understand how
emissions and meteorology, independently and combined, impact high ozone
levels.</p>
</sec>
<sec id="Ch1.S2">
  <label>2</label><title>Methods and data</title>
<sec id="Ch1.S2.SS1">
  <label>2.1</label><title>Methods</title>
      <p id="d1e217">Brief descriptions of the four methods (GAM, MARS, RF and SVR) are provided
below, and they are described in greater detail in the referenced material.</p>
<sec id="Ch1.S2.SS1.SSS1">
  <label>2.1.1</label><title>Generalized additive model (GAM)</title>
      <p id="d1e227">A GAM uses flexible, nonlinear relationships defined between “knots” in
the explanatory variables using smoothing functions (Hastie and
Tibshirani, 1986, 1990). The knot is the point of the link of two
polynomial curves (Wood, 2017). Since the GAM is an additive model,
which means each indicator's function adds together to form the model
equation, the indicators can have a variety of relationships with the
response variable. The general form of the GAM is written as (Hastie and
Tibshirani, 1986, 1990; Wood, 2011, 2017)
              <disp-formula id="Ch1.Ex1"><mml:math id="M4" display="block"><mml:mrow><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:mi>a</mml:mi><mml:mo>+</mml:mo><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:munderover><mml:mi>f</mml:mi><mml:mfenced close=")" open="("><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:mfenced><mml:mo>+</mml:mo><mml:mi>e</mml:mi><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
            where <inline-formula><mml:math id="M5" display="inline"><mml:mi>a</mml:mi></mml:math></inline-formula> is the intercept, <inline-formula><mml:math id="M6" display="inline"><mml:mi>e</mml:mi></mml:math></inline-formula> is the error term, <inline-formula><mml:math id="M7" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> refers to each independent
indicator and <inline-formula><mml:math id="M8" display="inline"><mml:mi>f</mml:mi></mml:math></inline-formula> means the function applied to the predictors.</p>
      <p id="d1e299">There are multiple choices of the functions based on the relationship
between each independent and dependent variable, such as splines, linear
functions and polynomials. Splines (often cubic) are commonly applied to
capture nonlinear relationships. Cubic splines can provide a comparatively
more flexible curve than low-order splines. In addition, a cubic spline can
avoid overfitting with a smaller curviness and be more effective with less
computational time than high-order splines. The basis function of a cubic
spline is a third-order polynomial equation:
              <disp-formula id="Ch1.Ex2"><mml:math id="M9" display="block"><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>⋅</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mn mathvariant="normal">3</mml:mn></mml:msup><mml:mo>+</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>⋅</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>+</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>⋅</mml:mo><mml:mi>x</mml:mi><mml:mo>+</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
            where <inline-formula><mml:math id="M10" display="inline"><mml:mi>a</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M11" display="inline"><mml:mi>b</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M12" display="inline"><mml:mi>c</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M13" display="inline"><mml:mi>d</mml:mi></mml:math></inline-formula> are the estimated coefficients of each basis function and the
subscript <inline-formula><mml:math id="M14" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula> indicates the number of basis functions (equal to the number of
knots). Based on the number of knots, several basis functions are built with
different estimated coefficients. Each spline is given by the weighted sum
of the basis functions. Three to five knots typically are sufficient in
practice, and the knots are evenly distributed based on the percentiles of
each indicator (Harrell, 2015). The “mgcv” package in the R program
was used to build the GAM between the peak MDA8 ozone concentrations and
indicators (Hastie, 1991; Hastie and Tibshirani, 1990, 1986; Wood, 2011,
2017).</p>
</sec>
<sec id="Ch1.S2.SS1.SSS2">
  <label>2.1.2</label><title>Multivariate adaptive regression splines (MARS)</title>
      <p id="d1e406">The MARS model is a nonparametric, multivariate, piecewise regression model
that can be used to develop the nonlinear relationships between the
dependent variable and a set of indicators (Friedman, 1991). Similar
to the GAM, linear splines (referred to as “hinge functions”) are applied
to independent variables in the MARS model. The resulting model is formed by
a weighted sum of basis functions. The MARS model can deal with nonlinear
relationships and provide a more flexible curve than simple linear
regression models and polynomial regression models due to the linear splines
between each pair of knots. It is simpler, and the resulting associations
between the dependent and indicator variables are easier to interpret than
the complex machine learning methods (e.g., random forest and neural
network). The general equation of the MARS model is as follows (Friedman, 1991;
Leathwick et al., 2006; Oduro et al., 2015; Roy et al., 2018):
              <disp-formula id="Ch1.Ex3"><mml:math id="M15" display="block"><mml:mrow><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="italic">β</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:munder><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:munder><mml:msub><mml:mi mathvariant="italic">β</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msub><mml:mi>H</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
            where <inline-formula><mml:math id="M16" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">β</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> is the intercept, <inline-formula><mml:math id="M17" display="inline"><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> shows hinge functions and
<inline-formula><mml:math id="M18" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">β</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the coefficients of hinge functions. The hinge functions in
the MARS model are pairwise, and the form is

                  <disp-formula specific-use="gather"><mml:math id="M19" display="block"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>-</mml:mo><mml:mi>k</mml:mi><mml:msub><mml:mo>)</mml:mo><mml:mo>+</mml:mo></mml:msub><mml:mo>=</mml:mo><mml:mo movablelimits="false">max⁡</mml:mo><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>-</mml:mo><mml:mi>k</mml:mi><mml:mo>,</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>-</mml:mo><mml:mi>x</mml:mi><mml:msub><mml:mo>)</mml:mo><mml:mo>+</mml:mo></mml:msub><mml:mo>=</mml:mo><mml:mo movablelimits="false">max⁡</mml:mo><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>-</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>

              where <inline-formula><mml:math id="M20" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula> is the knot. When applying the MARS model, a two-stage approach is
used that includes forward and backward stages. The forward stage is similar
to the forward stepwise regression. At first, the model only includes the
intercept term. Then, the generated pairwise hinge functions are added into
the model continuously if they can reduce the residual error of the model.
This process will be terminated when the change of error is small (e.g.,
less than a threshold) or the model reaches the defined maximum number of
terms. A backward stage is applied to avoid overfitting and reduce the
number of terms, removing terms that do not significantly impact the error
(Wikipedia Contributors, 2022). Generalized cross validation (GCV) is used
to find the final MARS model after obtaining multiple models that have
different terms (Friedman, 1991; Friedman and Silverman, 1989; Hastie and
Tibshirani, 1996; Leathwick et al., 2006; Oduro et al., 2015; Roy et al.,
2018). The “earth” package in R was applied to build the relationship
between the top MDA8 ozone concentrations and independent indicators using
the MARS model (Friedman and Silverman, 1989; Milborrow, 2021; Hastie et
al., 2009), and this package chose the independent variables, the position
of the knots and the interaction of the terms automatically.</p>
</sec>
<sec id="Ch1.S2.SS1.SSS3">
  <label>2.1.3</label><title>Random forest model (RF)</title>
      <p id="d1e579">Random forest is a supervised machine learning method that can be used for
regression and classification. It is an ensemble of multiple decision trees.
The RF model resolves the limitation of the decision tree that the model can
be overfitting if the depth of the trees is deeper by applying the bagging
algorithm. The bagging algorithm effectively reduces the variance of the
model results and makes the RF model quite stable and robust. In regression,
the predicted result of the RF model is the average of the results of all
decision trees. The total error of RF is computed by the average of the
error of all the decision trees.</p>
      <p id="d1e582">Suppose we build a random forest model which contains <inline-formula><mml:math id="M21" display="inline"><mml:mi>m</mml:mi></mml:math></inline-formula> trees
(i.e., <inline-formula><mml:math id="M22" display="inline"><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mi>b</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M23" display="inline"><mml:mrow><mml:mi>b</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:mi>m</mml:mi></mml:mrow></mml:math></inline-formula>) and has a testing
point <inline-formula><mml:math id="M24" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula>. The predicted value of input <inline-formula><mml:math id="M25" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> would be
              <disp-formula id="Ch1.Ex6"><mml:math id="M26" display="block"><mml:mrow><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mi>m</mml:mi></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>m</mml:mi></mml:munderover><mml:msub><mml:mi>T</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>
            The following steps construct each decision tree in a random forest model.
First, randomly select a subset of the training dataset with replacement.
Then, at each decision node, randomly select a subset of variables. In order
to find the optimal variable and its corresponding value that can lead to
the best fit, we usually define a target function and compare all variables
in the subset to find the variable with the lowest or highest value. Once we
find the optimal variable and corresponding value, we next divide the
decision node based on the optimal variable and value. Repeat the previous
step until all decision nodes reach the minimal node size. Finally, for each
leaf node, suppose <inline-formula><mml:math id="M27" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula> data points,
<inline-formula><mml:math id="M28" display="inline"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:msub><mml:mi>x</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, belong to a
leaf node and the corresponding response variables are
<inline-formula><mml:math id="M29" display="inline"><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:msub><mml:mi>y</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> respectively.
The predicted result of a testing point <inline-formula><mml:math id="M30" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> which falls into this leaf node
should be
              <disp-formula id="Ch1.Ex7"><mml:math id="M31" display="block"><mml:mrow><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mi>k</mml:mi></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>k</mml:mi></mml:munderover><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>
            We used the “randomForest” package in the R software to build the RF model
(Liaw and Wiener, 2002). RF models can select interaction terms
between the independent variables automatically.</p>
</sec>
<sec id="Ch1.S2.SS1.SSS4">
  <label>2.1.4</label><title>Support vector regression (SVR)</title>
      <p id="d1e787">The support vector machine (SVM) method is a supervised machine learning
approach that is used for classification. The SVR model, which is an
extension of the SVM, can be used to describe the nonlinear relationships
between the response variable and independent indicators.</p>
      <p id="d1e790">Suppose we have a set of indicators
<inline-formula><mml:math id="M32" display="inline"><mml:mrow><mml:mi>X</mml:mi><mml:mo>=</mml:mo><mml:mo mathvariant="italic">{</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:msub><mml:mi>x</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:msub><mml:mi>x</mml:mi><mml:mi>m</mml:mi></mml:msub><mml:mo mathvariant="italic">}</mml:mo></mml:mrow></mml:math></inline-formula>
and a set of response variables <inline-formula><mml:math id="M33" display="inline"><mml:mrow><mml:mi>Y</mml:mi><mml:mo>=</mml:mo><mml:mo mathvariant="italic">{</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:msub><mml:mi>y</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:msub><mml:mi>y</mml:mi><mml:mi>m</mml:mi></mml:msub><mml:mo mathvariant="italic">}</mml:mo></mml:mrow></mml:math></inline-formula>.
We need to find a hyperplane to minimize error and achieve the best fit,
which can be written as <inline-formula><mml:math id="M34" display="inline"><mml:mrow><mml:msup><mml:mi>w</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:mi>x</mml:mi><mml:mo>+</mml:mo><mml:mi>b</mml:mi></mml:mrow></mml:math></inline-formula>. We can define the loss function as
              <disp-formula id="Ch1.Ex8"><mml:math id="M35" display="block"><mml:mrow><mml:mtext>Loss</mml:mtext><mml:mfenced open="(" close=")"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:mfenced><mml:mo>=</mml:mo><mml:mo movablelimits="false">max⁡</mml:mo><mml:mo>(</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">|</mml:mi><mml:msup><mml:mi>w</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:mi>b</mml:mi><mml:mo>-</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mi mathvariant="normal">|</mml:mi><mml:mo>-</mml:mo><mml:mi mathvariant="italic">ε</mml:mi><mml:mo>)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
            where <inline-formula><mml:math id="M36" display="inline"><mml:mi mathvariant="italic">ε</mml:mi></mml:math></inline-formula> (epsilon) is the margin of error, a user-defined
variable that can be manipulated to adjust the accuracy of the model. Then,
the problem can be written as
              <disp-formula id="Ch1.Ex9"><mml:math id="M37" display="block"><mml:mrow><mml:munder><mml:mo movablelimits="false">min⁡</mml:mo><mml:mrow><mml:mi>w</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi></mml:mrow></mml:munder><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:mfrac></mml:mstyle><mml:msup><mml:mfenced close="∥" open="∥"><mml:mi>w</mml:mi></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>+</mml:mo><mml:mi>C</mml:mi><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>m</mml:mi></mml:munderover><mml:mtext>Loss</mml:mtext><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
            where <inline-formula><mml:math id="M38" display="inline"><mml:mi>C</mml:mi></mml:math></inline-formula> is the cost, another user-defined variable that determines the
tolerance of the model to points outside the bounds set by <inline-formula><mml:math id="M39" display="inline"><mml:mi mathvariant="italic">ε</mml:mi></mml:math></inline-formula>,
and <inline-formula><mml:math id="M40" display="inline"><mml:mi>m</mml:mi></mml:math></inline-formula> is the number of dataset. To let the loss function result of each
training point be 0, we introduced the slack variables. Then, the
development of a nonlinear relationship between the response variable and
indicators can be converted a Lagrangian dual problem
(Schölkopf and Smola, 2001; Smola and Schölkopf, 2004).</p>
      <p id="d1e1042">We need to consider the interactions among features sometimes when we build
computational models, so we need to map the data into a nonlinear feature
space. The nonlinear feature space increases the dimension of the data
space, and consequently the computational complexity grows dramatically. We
introduced a kernel function to account for the interactions and reduce the
computational complexity. We used the package “e1071” in R software to build
the SVR model (Chang and Lin, 2011; Fan et al., 2005).</p>
</sec>
</sec>
<sec id="Ch1.S2.SS2">
  <label>2.2</label><title>Model evaluation</title>
      <p id="d1e1054">We used the coefficient of determination (<inline-formula><mml:math id="M41" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>), mean bias (MB) and root
mean squared error (RMSE) of the observed and predicted peak MDA8 ozone
concentrations from 1990 to 2019 to compare the performance of these four
models.

                <disp-formula specific-use="gather"><mml:math id="M42" display="block"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mtext>Mean bias</mml:mtext><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:munderover><mml:mo>(</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>x</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mi>i</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mi>n</mml:mi></mml:mfrac></mml:mstyle><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mtext>RMSE</mml:mtext><mml:mo>=</mml:mo><mml:msqrt><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:munderover><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>x</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mi>i</mml:mi></mml:msub><mml:msup><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow><mml:mi>n</mml:mi></mml:mfrac></mml:mstyle></mml:msqrt><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>

            where <inline-formula><mml:math id="M43" display="inline"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M44" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi>x</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> are the observed and predicted MDA8 ozone
concentrations and <inline-formula><mml:math id="M45" display="inline"><mml:mi>n</mml:mi></mml:math></inline-formula> is the total number of measurements. In addition, we
used 10-fold cross validation (CV) to evaluate the prediction accuracy and
stability of these four models. In the 10-fold CV, the dataset is randomly
divided into two subsets, in which 90 % is used to train the model and
10 % is the testing dataset. These two subsets are not overlapped, and
this separation process repeats 10 times. The averages of the <inline-formula><mml:math id="M46" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>, MB
and RMSE in these 10 runs are the final evaluation results of numerical
models.</p>
</sec>
<sec id="Ch1.S2.SS3">
  <label>2.3</label><title>Study domain</title>
      <p id="d1e1222">The SoCAB includes urban and suburban parts of Los Angeles County, Riverside
County and San Bernardino County and all of Orange County. This area
historically and still experiences some of the worst air quality in the US,
and various air pollutants at multiple sites in the SoCAB do not meet the
NAAQS, even with strict regulations leading to significant reductions in
pollutant emissions. The poor air quality is because SoCAB is one of the
most urbanized and populated regions in the US and is surrounded by
mountains on three sides, while the Pacific Ocean lies on the west side.
Temperature inversions are formed frequently along the coast due to the warm
subsiding air from North Pacific highs, suppressing vertical mixing. This
unique geographical and meteorological environment leads to reduced dilution
of air pollutants. In addition, most days in a year are sunny, leading to
warm and dry conditions with high solar radiation, exacerbating the
formation of photochemically derived pollutants, such as ozone.</p>
      <p id="d1e1225">We first focused on the Crestline site to develop the initial regression
models to predict the fourth highest MDA8 ozone concentrations in the
SoCAB. This site had the annual fourth highest MDA8 ozone concentrations
during about 77 % of this project's period. The other 23 % of the time,
the maximum site was close to the Crestline site, such as Glendora,
Redlands and Fontana.</p>
</sec>
<sec id="Ch1.S2.SS4">
  <label>2.4</label><title>Data</title>
      <p id="d1e1236">The daily MDA8 ozone concentrations from 1990 to 2019 in the South Coast Air
Basin was retrieved from California Air Resources Board (CARB) archives and
EPA Air Quality System (AQS) pre-generated data files (CARB, 2020). The total number of
days of daily MDA8 ozone levels is 10 957. We used the top 30 MDA8 ozone days
each year to develop the models for the fourth highest MDA8 ozone
concentrations to build robust computational models, since multiple factors
have impacts on the peak MDA8 ozone concentrations. Significant factors may
be missed if only the fourth highest MDA8 ozone concentrations are
considered, such as the day of the year, day of the week and meteorological-variable
impacts, as there would only be 30 observations for model training.
Furthermore, the size of the 30 years' fourth highest MDA8 ozone dataset
is too small to have sufficient statistical power. A small dataset may cause
a type II error (failing to identify a statistically significant effect) for
some significant features, which would then affect the accuracy of the
predictions.</p>
      <p id="d1e1239">We selected 25 independent indicators, including precursors' emissions,
meteorological factors suggested in previous studies (Blanchard et al.,
2014, 2019; Camalier et al., 2007), Niño 3.4 monthly
indices, the day of the week and the day of the year. A detailed description of all
the variables applied to test the final computational models is in Table S1 in the Supplement.</p>
      <p id="d1e1242">Estimated <inline-formula><mml:math id="M47" display="inline"><mml:mrow class="chem"><mml:msub><mml:mi mathvariant="normal">NO</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and VOC emissions in the SoCAB from 2000 to 2019 were acquired
from CARB archives using the emissions in 2012 (Cox et al., 2013). The
emissions between 1990 and 2000 were projected with the emissions in 2008
and 2012 (Cox et al., 2009, 2013). The detailed calculation is in the Supplement.</p>
      <p id="d1e1256">We included two kinds of meteorological data: surface meteorological data
and upper-air meteorological data. We obtained the surface meteorological
data, including temperature, wind speed and wind direction at Los Angeles
International Airport (LAX) and Barstow-Daggett Airport (Barstow Airport) from National Oceanic and
Atmospheric Administration (NOAA) archives and CARB archives (Menne et
al., 2012a, b). The upper-air meteorological data at the
Miramar site was provided by NOAA and contains geopotential height,
temperature, dew point temperature, wind speed and wind direction at 500
and 850 mb (millibar). Using temperature and dew point temperature, we
computed the relative humidity (RH) at 500 and 850 mb with the Clausius–Clapeyron equation
(Alduchov and Eskridge, 1996; Lawrence, 2005). The height of 500 mb is
around 5500 m (NOAA, 2020), and the 850 mb height is about 1500 m,
which is close to the boundary layer height. The upper-air meteorology is
related to the synoptic-scale weather and has an impact on the surface
meteorology (Blanchard et al., 2014; Camalier et al., 2007).</p>
      <p id="d1e1260">Past studies have shown there is a relationship between the El
Niño–Southern Oscillation (ENSO) events and the variability of MDA8
ozone concentrations by affecting the local meteorology (Lu et al., 2019;
Oman et al., 2013, 2011; Xu et al., 2017). Niño 3.4 monthly
indices were obtained from the Climate Prediction Center (CPC) to represent ENSO
events. To account for the daily variations and weekend effects of MDA8
ozone levels, we included the day of the week and the day of the year in the models
(Seinfeld and Pandis, 2016).</p>
</sec>
</sec>
<sec id="Ch1.S3">
  <label>3</label><title>Results</title>
<sec id="Ch1.S3.SS1">
  <label>3.1</label><title>Model application and performance</title>
<sec id="Ch1.S3.SS1.SSS1">
  <label>3.1.1</label><title>GAM model</title>
      <p id="d1e1286">We combined stepwise regression and <inline-formula><mml:math id="M48" display="inline"><mml:mi>F</mml:mi></mml:math></inline-formula> values to assess the statistical
significance of each independent indicator to refine the model equation to
provide the smallest Akaike information criterion (AIC) value after
excluding the highly correlated indicators (Fig. S1) (Pope and
Webster, 1972). However, the stepwise regression may exclude some factors
that are known to be tied to ozone formation from the final equation,
including VOC emissions, the day of the year and the day of the week. We used both
statistical indicators and knowledge of important relationships in the final
model to avoid losing significant factors that affect the peak ozone levels.
Furthermore, a limitation of the GAM is that it does not identify
interaction terms, so interaction terms were introduced with the spline
function manually in the style of <inline-formula><mml:math id="M49" display="inline"><mml:mrow><mml:mi>s</mml:mi><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mi>x</mml:mi><mml:mn mathvariant="normal">2</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>.</p>
      <p id="d1e1318">We applied cubic splines to the emissions and meteorological variables due
to the nonlinear relationship between peak MDA8 ozone concentrations and
meteorology/emissions. Also, the cubic spline was used for the day of the
year to add the daily and seasonal variation of the precursors' emissions.
In addition, we included the day of the week in a factor style to represent
the weekend effect and Niño 3.4 monthly indices to show the large-scale
climate pattern impacts on ozone formation with linear functions. We used
the annual top 30 MDA8 ozone concentrations from 1990 to 2019 on a log scale
as the dependent variable at the Crestline site because the ozone
concentrations follow a log-normal distribution (U.S. EPA, 2020; Henneman et
al., 2015; Hogrefe et al., 2000; Rao et al., 1997; Blanchard et al., 2014;
Camalier et al., 2007).</p>
      <p id="d1e1321">The final GAM (GAM-SoCAB-8HR V1.0) included emissions, meteorological
factors, large-scale climate indices and temporal variables at the Crestline
site from 1990 to 2019 (Eq. 1). The detailed description of each variable
is in Table 1 (e: error term):
              <disp-formula id="Ch1.E1" content-type="numbered"><label>1</label><mml:math id="M50" display="block"><mml:mtable rowspacing="0.2ex" class="split" displaystyle="true" columnalign="right left"><mml:mtr><mml:mtd><mml:mrow><mml:mi>log⁡</mml:mi><mml:mfenced open="(" close=")"><mml:mtext>MDA8</mml:mtext></mml:mfenced></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mo>=</mml:mo><mml:mi>a</mml:mi><mml:mo>+</mml:mo><mml:mtext>dayofweek (factor)</mml:mtext><mml:mo>+</mml:mo><mml:mtext>dayofyear</mml:mtext></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mo>+</mml:mo><mml:mi>s</mml:mi><mml:mfenced open="(" close=")"><mml:mtext>TMAXBarstow</mml:mtext></mml:mfenced><mml:mo>+</mml:mo><mml:mi>s</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="normal">Mir</mml:mi><mml:mn mathvariant="normal">850</mml:mn><mml:mi mathvariant="normal">RH</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mo>+</mml:mo><mml:mi>s</mml:mi><mml:mtext>(AWNDLAX)</mml:mtext><mml:mo>+</mml:mo><mml:mi>s</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="normal">e</mml:mi><mml:mrow class="chem"><mml:msub><mml:mi mathvariant="normal">NO</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow><mml:mo>)</mml:mo><mml:mo>+</mml:mo><mml:mi>s</mml:mi><mml:mtext>(eROG)</mml:mtext></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mo>+</mml:mo><mml:mi>s</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="normal">e</mml:mi><mml:mrow class="chem"><mml:msub><mml:mi mathvariant="normal">NO</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mtext>TMAXBarstow</mml:mtext><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mo>+</mml:mo><mml:mi>s</mml:mi><mml:mtext>(eROG,TMAXBarstow)</mml:mtext></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mo>+</mml:mo><mml:mi>s</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="normal">e</mml:mi><mml:mrow class="chem"><mml:msub><mml:mi mathvariant="normal">NO</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mtext>eROG</mml:mtext><mml:mo>)</mml:mo><mml:mo>+</mml:mo><mml:mtext>ENSOmonthly</mml:mtext><mml:mo>+</mml:mo><mml:mi mathvariant="normal">e</mml:mi><mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
            The correlation (<inline-formula><mml:math id="M51" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>) between independent variables in the final GAM was
tested (Fig. 1). The correlation between VOC and <inline-formula><mml:math id="M52" display="inline"><mml:mrow class="chem"><mml:msub><mml:mi mathvariant="normal">NO</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> emissions was high at
close to 1. However, both are the precursors of MDA8 ozone, so these factors
were not removed during the model development. Other than emissions, the
correlation among all the significant independent variables in the final
models is negligible (Fig. 1).</p>
      <p id="d1e1501">A total of 84 % of the variability of the peak MDA8 ozone concentrations can be
explained using this GAM (Fig. 2a). The 10-fold validation results show that the
<inline-formula><mml:math id="M53" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> value was 0.85 using the testing dataset, only 0.01 higher than the
training dataset. Also, the RMSE of the testing data was only slightly
different from that of the training data (Table S4), which indicated that
this GAM could predict peak MDA8 ozone concentrations stably. This model had
an <inline-formula><mml:math id="M54" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> value equal to 0.96, and RMSE is 11.1 ppbv for the fourth
highest MDA8 ozone predictions from 1990 to 2019 (Fig. 3a).</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F1"><?xmltex \currentcnt{1}?><?xmltex \def\figurename{Figure}?><label>Figure 1</label><caption><p id="d1e1529">Correlation value between the independent variables (only valid
for GAM).</p></caption>
            <?xmltex \igopts{width=236.157874pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/15/9015/2022/gmd-15-9015-2022-f01.png"/>

          </fig>

</sec>
<sec id="Ch1.S3.SS1.SSS2">
  <label>3.1.2</label><title>MARS model</title>
      <p id="d1e1546">We used the same dataset as the GAM (GAM-SoCAB-8HR V1.0) to be comparable
with the GAM's results. The final model contained six indicators, including
the <inline-formula><mml:math id="M55" display="inline"><mml:mrow class="chem"><mml:msub><mml:mi mathvariant="normal">NO</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and VOC emissions, the maximum temperature at Barstow-Daggett Airport, the average wind speed at LAX, Niño 3.4 monthly indices, the day of
the year and 11 interaction terms between indicators. Similar to GAM,
we applied the log function to the ozone concentrations (Epa, 2020; Henneman
et al., 2015; Hogrefe et al., 2000; Rao et al., 1997). The final equation of
the MARS model (MARS-SoCAB-8HR V1.0) is shown below (Eq. 2). A detailed
description of each variable is in Table 1:
              <disp-formula id="Ch1.E2" content-type="numbered"><label>2</label><mml:math id="M56" display="block"><mml:mtable rowspacing="0.2ex" class="split" displaystyle="true" columnalign="right left"><mml:mtr><mml:mtd><mml:mrow><mml:mi>log⁡</mml:mi><mml:mfenced open="(" close=")"><mml:mtext>MDA8</mml:mtext></mml:mfenced></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mo>=</mml:mo><mml:mn mathvariant="normal">4.59</mml:mn><mml:mo>+</mml:mo><mml:mo>(</mml:mo><mml:mn mathvariant="normal">1135.1</mml:mn><mml:mo>-</mml:mo><mml:mi mathvariant="normal">e</mml:mi><mml:mrow class="chem"><mml:msub><mml:mi mathvariant="normal">NO</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mo>⋅</mml:mo><mml:mo>(</mml:mo><mml:mo>-</mml:mo><mml:mn mathvariant="normal">9.57</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">4</mml:mn></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mo>+</mml:mo><mml:mo>(</mml:mo><mml:mi mathvariant="normal">e</mml:mi><mml:mrow class="chem"><mml:msub><mml:mi mathvariant="normal">NO</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1135.1</mml:mn><mml:mo>)</mml:mo><mml:mo>⋅</mml:mo><mml:mo>(</mml:mo><mml:mn mathvariant="normal">1.55</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mo>+</mml:mo><mml:mo>(</mml:mo><mml:mn mathvariant="normal">643.4</mml:mn><mml:mo>-</mml:mo><mml:mtext>eROG</mml:mtext><mml:mo>)</mml:mo><mml:mo>⋅</mml:mo><mml:mo>(</mml:mo><mml:mn mathvariant="normal">2.67</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mo>+</mml:mo><mml:mo>(</mml:mo><mml:mtext>TMAXBarstow</mml:mtext><mml:mo>-</mml:mo><mml:mn mathvariant="normal">36.7</mml:mn><mml:mo>)</mml:mo><mml:mo>⋅</mml:mo><mml:mn mathvariant="normal">0.022</mml:mn></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mo>+</mml:mo><mml:mo>(</mml:mo><mml:mn mathvariant="normal">1135.1</mml:mn><mml:mo>-</mml:mo><mml:mi mathvariant="normal">e</mml:mi><mml:mrow class="chem"><mml:msub><mml:mi mathvariant="normal">NO</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow><mml:mo>)</mml:mo><mml:mo>⋅</mml:mo><mml:mo>(</mml:mo><mml:mtext>eROG</mml:mtext><mml:mo>-</mml:mo><mml:mn mathvariant="normal">450.7</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mo>⋅</mml:mo><mml:mo>(</mml:mo><mml:mn mathvariant="normal">4.29</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">6</mml:mn></mml:mrow></mml:msup><mml:mo>)</mml:mo><mml:mo>+</mml:mo><mml:mo>(</mml:mo><mml:mi mathvariant="normal">e</mml:mi><mml:mrow class="chem"><mml:msub><mml:mi mathvariant="normal">NO</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1415.8</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mo>⋅</mml:mo><mml:mo>(</mml:mo><mml:mtext>eROG</mml:mtext><mml:mo>-</mml:mo><mml:mn mathvariant="normal">643.4</mml:mn><mml:mo>)</mml:mo><mml:mo>⋅</mml:mo><mml:mo>(</mml:mo><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1.08</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">6</mml:mn></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mo>+</mml:mo><mml:mo>(</mml:mo><mml:mi mathvariant="normal">e</mml:mi><mml:mrow class="chem"><mml:msub><mml:mi mathvariant="normal">NO</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1225</mml:mn><mml:mo>)</mml:mo><mml:mo>⋅</mml:mo><mml:mo>(</mml:mo><mml:mtext>TMAXBarstow</mml:mtext><mml:mo>-</mml:mo><mml:mn mathvariant="normal">36.7</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mo>⋅</mml:mo><mml:mo>(</mml:mo><mml:mn mathvariant="normal">7.21</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">4</mml:mn></mml:mrow></mml:msup><mml:mo>)</mml:mo><mml:mo>+</mml:mo><mml:mo>(</mml:mo><mml:mn mathvariant="normal">1081.5</mml:mn><mml:mo>-</mml:mo><mml:mtext>eROG</mml:mtext><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mo>⋅</mml:mo><mml:mo>(</mml:mo><mml:mtext>TMAXBarstow</mml:mtext><mml:mo>-</mml:mo><mml:mn mathvariant="normal">36.7</mml:mn><mml:mo>)</mml:mo><mml:mo>⋅</mml:mo><mml:mo>(</mml:mo><mml:mo>-</mml:mo><mml:mn mathvariant="normal">2.18</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">5</mml:mn></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mo>+</mml:mo><mml:mo>(</mml:mo><mml:mtext>eROG</mml:mtext><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1081.5</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mo>⋅</mml:mo><mml:mo>(</mml:mo><mml:mtext>TMAXBarstow</mml:mtext><mml:mo>-</mml:mo><mml:mn mathvariant="normal">36.7</mml:mn><mml:mo>)</mml:mo><mml:mo>⋅</mml:mo><mml:mo>(</mml:mo><mml:mo>-</mml:mo><mml:mn mathvariant="normal">3.95</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">4</mml:mn></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mo>+</mml:mo><mml:mo>(</mml:mo><mml:mtext>eROG</mml:mtext><mml:mo>-</mml:mo><mml:mn mathvariant="normal">643.4</mml:mn><mml:mo>)</mml:mo><mml:mo>⋅</mml:mo><mml:mo>(</mml:mo><mml:mn mathvariant="normal">2.8</mml:mn><mml:mo>-</mml:mo><mml:mtext>AWNDLAX</mml:mtext><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mo>⋅</mml:mo><mml:mo>(</mml:mo><mml:mo>-</mml:mo><mml:mn mathvariant="normal">3.34</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">4</mml:mn></mml:mrow></mml:msup><mml:mo>)</mml:mo><mml:mo>+</mml:mo><mml:mo>(</mml:mo><mml:mtext>eROG</mml:mtext><mml:mo>-</mml:mo><mml:mn mathvariant="normal">643.4</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mo>⋅</mml:mo><mml:mo>(</mml:mo><mml:mtext>AWNDLAX</mml:mtext><mml:mo>-</mml:mo><mml:mn mathvariant="normal">2.8</mml:mn><mml:mo>)</mml:mo><mml:mo>⋅</mml:mo><mml:mo>(</mml:mo><mml:mo>-</mml:mo><mml:mn mathvariant="normal">4.99</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">5</mml:mn></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mo>+</mml:mo><mml:mo>(</mml:mo><mml:mn mathvariant="normal">26.91</mml:mn><mml:mo>-</mml:mo><mml:mtext>ENSOmonthly</mml:mtext><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mo>⋅</mml:mo><mml:mo>(</mml:mo><mml:mtext>eROG</mml:mtext><mml:mo>-</mml:mo><mml:mn mathvariant="normal">643.4</mml:mn><mml:mo>)</mml:mo><mml:mo>⋅</mml:mo><mml:mo>(</mml:mo><mml:mo>-</mml:mo><mml:mn mathvariant="normal">3.03</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">4</mml:mn></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mo>+</mml:mo><mml:mo>(</mml:mo><mml:mn mathvariant="normal">28.3</mml:mn><mml:mo>-</mml:mo><mml:mtext>ENSOmonthly</mml:mtext><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mo>⋅</mml:mo><mml:mo>(</mml:mo><mml:mtext>eROG</mml:mtext><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1081.5</mml:mn><mml:mo>)</mml:mo><mml:mo>⋅</mml:mo><mml:mo>(</mml:mo><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1.61</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mo>+</mml:mo><mml:mo>(</mml:mo><mml:mn mathvariant="normal">28.3</mml:mn><mml:mo>-</mml:mo><mml:mtext>ENSOmonthly</mml:mtext><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mo>⋅</mml:mo><mml:mo>(</mml:mo><mml:mtext>eROG</mml:mtext><mml:mo>-</mml:mo><mml:mn mathvariant="normal">982.7</mml:mn><mml:mo>)</mml:mo><mml:mo>⋅</mml:mo><mml:mo>(</mml:mo><mml:mn mathvariant="normal">1.34</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mo>+</mml:mo><mml:mo>(</mml:mo><mml:mn mathvariant="normal">235</mml:mn><mml:mo>-</mml:mo><mml:mtext>dayofyear</mml:mtext><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mo>⋅</mml:mo><mml:mo>(</mml:mo><mml:mtext>TMAXBarstow</mml:mtext><mml:mo>-</mml:mo><mml:mn mathvariant="normal">36.7</mml:mn><mml:mo>)</mml:mo><mml:mo>⋅</mml:mo><mml:mo>(</mml:mo><mml:mn mathvariant="normal">1.22</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">4</mml:mn></mml:mrow></mml:msup><mml:mo>)</mml:mo><mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
            The <inline-formula><mml:math id="M57" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> when applied to predict the top 30 MDA8 ozone predictions was
0.83 and showed no overfitting (Table S4). The model also had a high <inline-formula><mml:math id="M58" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>
(0.95), and RMSE equaled 11.2 ppbv when predicting the fourth highest MDA8
ozone concentrations (Fig. 3b).</p>
      <p id="d1e2328">Multiple tests were performed using the different number of the remaining
terms in the output model with 10-fold CV to improve the MARS model
performance. The best model was obtained when there were 14 terms maintained
in the MARS model (Fig. S2). The performance of the MARS model with 14 terms
(<inline-formula><mml:math id="M59" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.83</mml:mn></mml:mrow></mml:math></inline-formula>, RMSE <inline-formula><mml:math id="M60" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 10.19) was similar to the MARS model with 16 terms
(<inline-formula><mml:math id="M61" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.83</mml:mn></mml:mrow></mml:math></inline-formula>, RMSE <inline-formula><mml:math id="M62" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 10.27).</p>
</sec>
<sec id="Ch1.S3.SS1.SSS3">
  <label>3.1.3</label><title>RF model</title>
      <p id="d1e2383">We first applied the same indicators and dataset as the GAM (GAM-SoCAB-8HR
V1.0) in order to compare the results of the above two regression methods.
In the base case run, we tried 0–500 trees to find the optimal number of
trees. Each tree chose two variables randomly that was equal to one-third of
the total number of variables by default. The optimal number of trees was
467 based on the RMSE value (Fig. S3). The majority of the top 30 and
fourth highest MDA8 ozone concentrations can be explained by the RF model
(RF-SoCAB-8HR V1.0) (<inline-formula><mml:math id="M63" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.81</mml:mn></mml:mrow></mml:math></inline-formula> and RMSE <inline-formula><mml:math id="M64" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 10.9 ppbv for the top 30 MDA8
ozone concentrations and <inline-formula><mml:math id="M65" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.97</mml:mn></mml:mrow></mml:math></inline-formula> and RMSE <inline-formula><mml:math id="M66" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 14.0 ppbv for the fourth highest MDA8
ozone concentration; Figs. 2c and 3c). The <inline-formula><mml:math id="M67" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> and RMSE values of the 10-fold CV
results were similar to those using the original RF model, with only a 0.01
difference in <inline-formula><mml:math id="M68" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> and about a 5 % reduction of the RMSE value that
indicated this RF model had a high prediction accuracy and no overfitting
(Table S4).</p>
      <p id="d1e2453">Two main hyperparameters affect the performance of RF models and can be
tuned: the number of trees used in the RF model and the number of random
variables in each tree. To improve the model performance further, we created
a grid with hyperparameters that the number of indicators considered at each
split from 2 to 8, and the number of trees was 1000 to tune the RF. The
optimal number of predictors in each tree was 2 due to the lowest out-of-bag
(OOB) error, the same as the default run. Also, the optimal number of trees
after model tuning was the same as the default run.</p>
      <p id="d1e2456">Next, we included all the available indicators in the RF model after
excluding the strongly correlated independent variables. Then we removed the
statistically insignificant indicators based on the <inline-formula><mml:math id="M69" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula> value and the variable
importance to find the optimal combination of independent variables in the
RF model. The final model contained two more variables than the one above:
maximum solar radiation and height at 850 mb. The importance of the
additional variables was minor and had negligible impacts on the model
performance (Fig. 4c). The optimal number of trees was equal to 495. The
<inline-formula><mml:math id="M70" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> and RMSE values for the top 30 MDA8 ozone predictions were similar
to those using the RF model with fewer variables, although the mean bias was
reduced (Table S3). In addition, the model performance for the fourth
highest MDA8 ozone predictions was worse than that using the RF model with
the same variables as GAM (Table 2). Therefore, the RF model with the same
GAM's variables fit the peak and the annual fourth highest MDA8 ozone
concentrations well.</p>
</sec>
<sec id="Ch1.S3.SS1.SSS4">
  <label>3.1.4</label><title>SVR model</title>
      <p id="d1e2485">We first built the SVR model (SVR-SoCAB-8HR V1.0) using the same variables
as the built GAM (GAM-SoCAB-8HR V1.0) above with the default setting (the
cost was 1, and epsilon was 0.1). We used kernel functions to consider the
interactions between the independent indicators. Several kernel functions
have been used in machine learning models, including the linear kernel,
polynomial kernel, radial kernel, etc. In practice, we used the linear
kernel for the linear relationship and the radial kernel for the nonlinear
relationship. Owing to the nonlinear relationship between the peak MDA8
ozone levels and emissions/meteorology, we applied the radial kernel to the
independent variables. The regression method we used was epsilon regression,
and the epsilon value is related to the margin tolerance.</p>
      <p id="d1e2488">The <inline-formula><mml:math id="M71" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> and RMSE values of the top 30 MDA8 ozone predictions were very
similar to the RF model's results, but the MB was larger than that of the RF
model (Table S3). Results for predicting the fourth highest MDA8 ozone
predictions found that the method did not capture the variability as well as
the other methods (<inline-formula><mml:math id="M72" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.89</mml:mn></mml:mrow></mml:math></inline-formula> and RMSE <inline-formula><mml:math id="M73" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 14.0 ppbv). The CV results
indicated that this SVR model is stable and has no overfitting (Table S4).</p>
      <p id="d1e2524">Two parameters significantly impact the improvement of predictions and can
be defined by users: the value of cost and epsilon. So we ran the SVR model
with a hyperparameter grid with the cost value from 1 to 512 and the
epsilon from 0 to 1 with an interval of 0.1. The model achieved the best
performance when the epsilon was 0.3, and the cost was 1. The predicted top
30 and fourth highest MDA8 ozone concentrations were similar to those
using the built SVR with default settings (Tables 2 and S3, Fig. S5).</p>
      <p id="d1e2527">We then built the SVR model with all the independent variables we had and
removed the insignificant variables using the <inline-formula><mml:math id="M74" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula> value and variable
importance. The optimal SVR model (SVRoptimal-SoCAB-8HR V1.0) contained the
variables in the above GAM and height at 850 mb and maximum solar
radiation. The ideal epsilon value was 0.1, and the cost value was 1, the
same as the default setting. Though the importance of these two additional
variables was close to 0, the model performance of simulations of the top 30 MDA8
ozone days improved slightly (<inline-formula><mml:math id="M75" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.83</mml:mn></mml:mrow></mml:math></inline-formula> and RMSE <inline-formula><mml:math id="M76" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 10.4 ppbv)
(Fig. 4d and Table S3). However, the fourth highest MDA8 ozone predictions
were less accurate compared to the <inline-formula><mml:math id="M77" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> and RMSE using this SVR model
(SVRoptimal-SoCAB-8HR V1.0) and the SVR model (SVR-SoCAB-8HR V1.0) with the
same GAM variables (<inline-formula><mml:math id="M78" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.87</mml:mn></mml:mrow></mml:math></inline-formula> and RMSE <inline-formula><mml:math id="M79" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 14.3 ppbv for the
SVRoptimal-SoCAB-8HR V1.0 and <inline-formula><mml:math id="M80" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.89</mml:mn></mml:mrow></mml:math></inline-formula> and RMSE <inline-formula><mml:math id="M81" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 14.0 ppbv for the
SVR-SoCAB-8HR V1.0) (Table 2).</p>

      <?xmltex \floatpos{p}?><fig id="Ch1.F2" specific-use="star"><?xmltex \currentcnt{2}?><?xmltex \def\figurename{Figure}?><label>Figure 2</label><caption><p id="d1e2618">Comparison between the top 30 observed and predicted MDA8 ozone
concentrations using the GAM-SoCAB-8HR V1.0 model <bold>(a)</bold>, MARS-SoCAB-8HR V1.0
model <bold>(b)</bold>, RF-SoCAB-8HR V1.0 model <bold>(c)</bold>, SVR-SoCAB-8HR V1.0 model <bold>(d)</bold> and the
SVRoptimal-SoCAB-8HR V1.0 model <bold>(e)</bold>.</p></caption>
            <?xmltex \igopts{width=398.338583pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/15/9015/2022/gmd-15-9015-2022-f02.png"/>

          </fig>

      <?xmltex \floatpos{t}?><fig id="Ch1.F3"><?xmltex \currentcnt{3}?><?xmltex \def\figurename{Figure}?><label>Figure 3</label><caption><p id="d1e2644">Comparison between the fourth highest observed and predicted
MDA8 ozone concentrations using the GAM model built for the top 30 MDA8
ozone days at the Crestline site using the GAM-SoCAB-8HR V1.0 model (blue),
MARS-SoCAB-8HR V1.0 model (orange), RF-SoCAB-8HR V1.0 model (green),
SVR-SoCAB-8HR V1.0 model (red) and SVRoptimal-SoCAB-8HR V1.0 model (purple).</p></caption>
            <?xmltex \igopts{width=236.157874pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/15/9015/2022/gmd-15-9015-2022-f03.png"/>

          </fig>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T1" specific-use="star"><?xmltex \currentcnt{1}?><label>Table 1</label><caption><p id="d1e2656">Predictors used in the GAM and MARS model equations.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="3">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="left"/>
     <oasis:colspec colnum="3" colname="col3" align="left"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Variable</oasis:entry>
         <oasis:entry colname="col2">Abbreviation</oasis:entry>
         <oasis:entry colname="col3">Unit</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">Day of the week (factor, from Monday to Sunday)</oasis:entry>
         <oasis:entry colname="col2">dayofweek</oasis:entry>
         <oasis:entry colname="col3">None</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Day of the year (from 1 to 365/366)</oasis:entry>
         <oasis:entry colname="col2">dayofyear</oasis:entry>
         <oasis:entry colname="col3">None</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Daily maximum surface temperature at the Barstow Airport site</oasis:entry>
         <oasis:entry colname="col2">TMAXBarstow</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M82" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula>C</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Daily average wind speed at the LAX site</oasis:entry>
         <oasis:entry colname="col2">AWNDLAX</oasis:entry>
         <oasis:entry colname="col3">m s<inline-formula><mml:math id="M83" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Daily RH at 850 mb</oasis:entry>
         <oasis:entry colname="col2">Mir850RH</oasis:entry>
         <oasis:entry colname="col3">%</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Monthly Niño 3.4 indices</oasis:entry>
         <oasis:entry colname="col2">ENSOmonthly</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M84" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula>C</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Annual averaged <inline-formula><mml:math id="M85" display="inline"><mml:mrow class="chem"><mml:msub><mml:mi mathvariant="normal">NO</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> emissions</oasis:entry>
         <oasis:entry colname="col2">e<inline-formula><mml:math id="M86" display="inline"><mml:mrow class="chem"><mml:msub><mml:mi mathvariant="normal">NO</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">t d<inline-formula><mml:math id="M87" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Annual averaged VOC emissions</oasis:entry>
         <oasis:entry colname="col2">eROG</oasis:entry>
         <oasis:entry colname="col3">t d<inline-formula><mml:math id="M88" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T2"><?xmltex \currentcnt{2}?><label>Table 2</label><caption><p id="d1e2858">Summary of statistical results of the fourth MDA8 ozone
predictions using four methods at the Crestline site.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="4">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:thead>
       <oasis:row>
         <oasis:entry colname="col1">Method</oasis:entry>
         <oasis:entry colname="col2">Mean bias</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M90" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">RMSE</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">(ppbv)</oasis:entry>
         <oasis:entry colname="col3"/>
         <oasis:entry colname="col4">(ppbv)</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">GAM</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M91" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">9.71</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">0.96</oasis:entry>
         <oasis:entry colname="col4">11.1</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">MARS model</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M92" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">9.28</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">0.95</oasis:entry>
         <oasis:entry colname="col4">11.2</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">RF model<inline-formula><mml:math id="M93" display="inline"><mml:msup><mml:mi/><mml:mi mathvariant="normal">a</mml:mi></mml:msup></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M94" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">12.5</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">0.97</oasis:entry>
         <oasis:entry colname="col4">14.0</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">RF model<inline-formula><mml:math id="M95" display="inline"><mml:msup><mml:mi/><mml:mi mathvariant="normal">b</mml:mi></mml:msup></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M96" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">12.7</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">0.95</oasis:entry>
         <oasis:entry colname="col4">14.5</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">SVR model<inline-formula><mml:math id="M97" display="inline"><mml:msup><mml:mi/><mml:mi mathvariant="normal">a</mml:mi></mml:msup></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M98" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">10.6</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">0.89</oasis:entry>
         <oasis:entry colname="col4">14.0</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">SVR model<inline-formula><mml:math id="M99" display="inline"><mml:mrow><mml:msup><mml:mi/><mml:mi mathvariant="normal">a</mml:mi></mml:msup><mml:mo>+</mml:mo></mml:mrow></mml:math></inline-formula> tune</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M100" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">10.3</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">0.89</oasis:entry>
         <oasis:entry colname="col4">13.7</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">SVR model<inline-formula><mml:math id="M101" display="inline"><mml:msup><mml:mi/><mml:mi mathvariant="normal">b</mml:mi></mml:msup></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M102" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">10.4</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">0.87</oasis:entry>
         <oasis:entry colname="col4">14.3</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table><table-wrap-foot><p id="d1e2861"><inline-formula><mml:math id="M89" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mi mathvariant="normal">a</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="normal">b</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> RF/SVR model with the same variables as the GAM (GAM-SoCAB-8HR
V1.0) and RF/SVR model with the optimal combination of the indicators.</p></table-wrap-foot></table-wrap>

</sec>
</sec>
<sec id="Ch1.S3.SS2">
  <label>3.2</label><title>Comparisons among the nonlinear methods</title>
<sec id="Ch1.S3.SS2.SSS1">
  <label>3.2.1</label><title>Statistical results and computational time (efficiency)</title>
      <p id="d1e3156">We compared the <inline-formula><mml:math id="M103" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>, MB and RMSE of the peak and fourth highest MDA8
ozone predictions using all these four models. The statistical results of
simulations of the top 30 MDA8 ozone days showed that all these four methods
explain most of the variability of observations, especially the GAM (Table S3). The GAM (GAM-SoCAB-8HR V1.0) had the lowest MB and RMSE and the highest
<inline-formula><mml:math id="M104" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> for the top 30 MDA8 ozone simulations among all these four
methods. Also, the GAM (GAM-SoCAB-8HR V1.0) showed the best stability of the
top 30 MDA8 ozone predictions based on CV results (Table S4).</p>
      <p id="d1e3181">In addition, these four numerical methods can capture the fourth highest
MDA8 ozone variations well. The RF model using the same variables as the
built GAM (RF-SoCAB-8HR V1.0) of the fourth highest MDA8 ozone predictions
with an <inline-formula><mml:math id="M105" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> of 0.97, MB of <inline-formula><mml:math id="M106" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">12.53</mml:mn></mml:mrow></mml:math></inline-formula> ppbv and RMSE of 14.02 ppbv showed
a lower model performance when compared to the GAM whose <inline-formula><mml:math id="M107" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> equaled
0.96, MB was <inline-formula><mml:math id="M108" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">9.71</mml:mn></mml:mrow></mml:math></inline-formula> ppbv and RMSE was 11.07 ppbv. The <inline-formula><mml:math id="M109" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> for the MARS
model (MARS-SoCAB-8HR V1.0) equaled 0.95 with an MB of <inline-formula><mml:math id="M110" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">9.28</mml:mn></mml:mrow></mml:math></inline-formula> ppbv and RMSE of
11.16 ppbv. In comparison to the performance of the GAM, the MARS model had
a better MB value but a worse <inline-formula><mml:math id="M111" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> and RMSE value. The SVR model
(SVR-SoCAB-8HR V1.0) showed the highest MB and RMSE value and lowest <inline-formula><mml:math id="M112" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>
among all these four methods, implying that the SVR model predictions gave
the highest variations and lowest prediction accuracy. In general, all these
four methods showed a similar performance to the fourth highest MDA8 ozone
predictions. The predicted fourth highest MDA8 ozone levels with RF and
SVR using the optimal variable combination had a lower <inline-formula><mml:math id="M113" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> and higher MB
and RMSE than those using the same variables as the GAM. Therefore, the
variables used in the GAM (GAM-SoCAB-8HR V1.0) were the best combination to
build the models for peak ozone levels.</p>
      <p id="d1e3281">The statistical results and computational time need to be considered
together to compare the model performance of all the models, especially for
a large size dataset. There were no significant differences among these four
methods in terms of the top 30 and the annual fourth highest ozone
predictions. The GAM was marginally better compared to the other three
models outside of cost effectiveness. The computational requirements for
each model in this work is small due to the small dataset size (Table 3). If
computational time is a key factor, the MARS model can be a good choice for
a larger dataset (Table 3).</p>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T3"><?xmltex \currentcnt{3}?><label>Table 3</label><caption><p id="d1e3288">Summary of the computational time of each model.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="2">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Method</oasis:entry>
         <oasis:entry colname="col2">Computational time (s)</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">GAM</oasis:entry>
         <oasis:entry colname="col2">14</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">MARS model</oasis:entry>
         <oasis:entry colname="col2">0.04</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">RF model</oasis:entry>
         <oasis:entry colname="col2">1.2</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">SVR model</oasis:entry>
         <oasis:entry colname="col2">4.9</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

</sec>
<sec id="Ch1.S3.SS2.SSS2">
  <label>3.2.2</label><title>Two-step method</title>
      <p id="d1e3359">The <inline-formula><mml:math id="M114" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> values of the fourth highest MDA8 ozone predictions using
these four regression methods were similar and agreed with the observations,
but the RMSE and MB values were larger than desired. In order to reduce the
bias, we applied a two-step method using the least-squares method to the
fourth highest MDA8 ozone predictions. The steps are shown below:
<list list-type="order"><list-item>
      <p id="d1e3375">Predict the top 30 MDA8 ozone concentrations from 1990 to 2019 using the
models built in Sect. 3.1.</p></list-item><list-item>
      <p id="d1e3379">Extract the annual predicted fourth highest MDA8 ozone concentrations
based on the date of the observations.</p></list-item><list-item>
      <p id="d1e3383">Apply the regression equation derived using the observations and the
predictions in step 2 to the fourth maximum value in each year's top 30
MDA8 ozone predictions (as the response variable) to get the updated
fourth highest MDA8 ozone predictions.</p></list-item><list-item>
      <p id="d1e3387">Use the regression equation from step 3 with the updated predictions to get
the improved fourth highest MDA8 ozone predictions.</p></list-item></list>
The mean bias of the improved predictions was removed; the <inline-formula><mml:math id="M115" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> value was
increased; and the RMSE was reduced. After applying the two-step method, the
GAM showed the best model performance among all the models with the highest
<inline-formula><mml:math id="M116" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> and lowest RMSE value. The performance of the MARS model and RF
model were almost the same. The improved SVR model results were still the
least accurate due to the lowest <inline-formula><mml:math id="M117" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> and highest MB and RMSE.</p>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T4"><?xmltex \currentcnt{4}?><label>Table 4</label><caption><p id="d1e3427">Summary of statistical results of the fourth MDA8 ozone
predictions after applying the two-step method using four methods at
the Crestline site.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="4">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:thead>
       <oasis:row>
         <oasis:entry colname="col1">Method</oasis:entry>
         <oasis:entry colname="col2">Mean bias</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M118" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">RMSE</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">(ppbv)</oasis:entry>
         <oasis:entry colname="col3"/>
         <oasis:entry colname="col4">(ppbv)</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">GAM</oasis:entry>
         <oasis:entry colname="col2">0</oasis:entry>
         <oasis:entry colname="col3">0.98</oasis:entry>
         <oasis:entry colname="col4">3.85</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">MARS model</oasis:entry>
         <oasis:entry colname="col2">0</oasis:entry>
         <oasis:entry colname="col3">0.97</oasis:entry>
         <oasis:entry colname="col4">4.54</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">RF model</oasis:entry>
         <oasis:entry colname="col2">0</oasis:entry>
         <oasis:entry colname="col3">0.97</oasis:entry>
         <oasis:entry colname="col4">4.55</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">SVR model</oasis:entry>
         <oasis:entry colname="col2">0</oasis:entry>
         <oasis:entry colname="col3">0.90</oasis:entry>
         <oasis:entry colname="col4">8.75</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

</sec>
<sec id="Ch1.S3.SS2.SSS3">
  <label>3.2.3</label><title>Relative importance of the independent variables</title>
      <p id="d1e3555">There are multiple methods to determine the importance of each independent
variable of computational models, but the differences among all the methods
are negligible. The algorithms used to calculate the variable importance of
the GAM, the RF model and the SVR model are similar, based on the
differences between the simulations using the original dataset and the
dataset with one indicator's value randomly permutated. If the change of the
simulations is significant, then that indicator is important or vice versa.
The variable importance shown is <inline-formula><mml:math id="M119" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:mi>r</mml:mi></mml:mrow></mml:math></inline-formula> (<inline-formula><mml:math id="M120" display="inline"><mml:mi>r</mml:mi></mml:math></inline-formula>: the Pearson correlation coefficient
between the simulations using the original and random-permutation datasets)
when the GAM is used. The RF and SVR model used the change of mean square
error between the simulations using the original and random-permutation
datasets. The MARS model computed the variable importance by adding the
indicator into the model and evaluating the error changes by GCV.</p>
      <p id="d1e3577">The precursors' emissions routinely are the most important indicators among
all the variables in these four models that indicate that the emissions have
more impact on the peak MDA8 ozone formation than the meteorology in the
SoCAB. The maximum temperature is quite significant among all the
meteorological factors. The GAM and the MARS model also included the
interaction terms between the emissions and maximum temperature, which
capture more variability of the peak MDA8 ozone concentrations. The maximum
temperature is related to solar radiation, which has an influence on the
rate of photolysis reactions. RH at 850 mb showed relatively high importance
in the RF and SVR models. It had a negative correlation with peak MDA8 ozone
concentrations because of its relationship with precipitation and cloud
cover and, in consequence, reduced solar radiation and affected photolysis
reactions.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F4" specific-use="star"><?xmltex \currentcnt{4}?><?xmltex \def\figurename{Figure}?><label>Figure 4</label><caption><p id="d1e3582">Variable importance for simulations of the top 30 MDA8 ozone days
using the GAM-SoCAB-8HR V1.0 model <bold>(a)</bold>, MARS-SoCAB-8HR V1.0 model <bold>(b)</bold>,
RF-SoCAB-8HR V1.0 model <bold>(c)</bold> and SVR-SoCAB-8HR V1.0 model <bold>(d)</bold>. The
variable importance for each model is calculated with different methods (see
text). SR: solar radiation.</p></caption>
            <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://gmd.copernicus.org/articles/15/9015/2022/gmd-15-9015-2022-f04.png"/>

          </fig>

</sec>
</sec>
<sec id="Ch1.S3.SS3">
  <label>3.3</label><title>Limitations</title>
      <p id="d1e3612">There are several limitations in the comparisons among these four models.
Some significant factors to the fourth highest MDA8 ozone concentrations
may be excluded from the models due to the relatively small dataset that
can consequently affect the prediction accuracy of the models; adding more
available meteorological factors (e.g., cloud coverage, planetary boundary
layer, surface wind direction and solar irradiance) and other large-scale
climate indices (e.g., Atlantic Multidecadal Oscillation (AMO) and tropical
Pacific sea surface temperature anomalies (TROP)), we can expect an
improvement in the model performance, albeit at the cost of performance.
Second, the running time of these four models with a small dataset does not
show any significant differences. The computational time, however, will be a
key criterion if the dataset is quite large and may affect the best model
choice. Third, since we included the local meteorological variables in the
model equations and the models were developed for the Crestline site, these
models may not offer the same performance with the peak ozone levels at
other sites in the SoCAB. Given that Crestline is downwind of Los Angeles,
which then is bordered by the Pacific Ocean, the models using SoCAB
emissions capture the upwind conditions. In other regions, such models could
be expanded to include both local emissions and upwind states' emissions.
Previous studies showed the emissions, maximum temperature, RH, wind speed,
wind direction and large-scale climate patterns have impacts on the daily
MDA8 ozone concentrations in different regions in the world (Blanchard et
al., 2014, 2019; Camalier et al., 2007; García Nieto
and Álvarez Antón, 2014; Gong et al., 2018, 2017; Jeong
et al., 2020; Jin et al., 2013; Ling et al., 2013; Liu et al., 2013; Lu and
Turco, 1996; Lu et al., 2019; Luna et al., 2014; Ma et al., 2020; McClure
and Jaffe, 2018; Sun et al., 2019). This is similar to the variable-importance results in this study. In addition, although the GAM and MARS
model outperform the RF and SVR model in this study, the machine learning
methods (e.g., RF, neural network and SVR) may potentially offer better
performance than the GAM and the MARS model with a significantly larger
dataset. Finally, the models were not developed to predict daily MDA8 ozone
concentrations because they are trained using the highest 30 MDA8 ozone
levels of each year. The relationships between inputs and predicted ozone
are very different at lower ozone levels.</p>
</sec>
</sec>
<sec id="Ch1.S4" sec-type="conclusions">
  <label>4</label><title>Conclusions</title>
      <p id="d1e3624">This study compared four observation-based approaches to predict the
fourth highest MDA8 ozone concentrations as a function of emissions,
meteorological factors and large-scale climate patterns. The statistical
results showed that these four models with estimated emissions and observed
meteorological factors can explain most of the variations of the top 30 and fourth highest MDA8
ozone concentrations (<inline-formula><mml:math id="M121" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.81</mml:mn></mml:mrow></mml:math></inline-formula>–0.84 for the top 30 MDA8 ozone concentrations
and <inline-formula><mml:math id="M122" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.89</mml:mn></mml:mrow></mml:math></inline-formula>–0.97 for the fourth highest MDA8 ozone concentrations). Among the top 30 MDA8
ozone models, the GAM (GAM-SoCAB-8HR V1.0) achieved the highest <inline-formula><mml:math id="M123" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>
(0.84) and lowest RMSE value (9.74 ppbv), and the SVR (SVR-SoCAB-8HR V1.0)
and RF (RF-SoCAB-8HR V1.0) achieved a lower <inline-formula><mml:math id="M124" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> value (0.81) and a
higher RMSE value (10.9 ppbv). So, in terms of the top 30 highest MDA8 ozone
predictions, there was little difference among these four models. These
models showed a better performance for predicting the fourth highest MDA8
ozone predictions than the peak ozone level. All models had a high <inline-formula><mml:math id="M125" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> value (close to or higher than 0.9), but after considering RMSE and MB
values, the GAM and the MARS model described the dataset better and provided
a significantly better prediction accuracy as compared to the RF and SVR
models. Although the computational time of each model was small for the dataset employed here, the MARS model required the least. The order of the
variable importance of the factors of each model was similar. The
precursors' emissions were the most significant factors for indicating the
importance of the emissions impact on peak ozone levels. Maximum temperature
presented relatively high importance among all the meteorological variables.</p>
</sec>

      
      </body>
    <back><notes notes-type="codedataavailability"><title>Code and data availability</title>

      <p id="d1e3695">Current versions of the models are available
from the project website at <ext-link xlink:href="https://doi.org/10.5281/zenodo.6892066" ext-link-type="DOI">10.5281/zenodo.6892066</ext-link> (Gao, 2022a)
under the Creative Commons Attribution 4.0 International license. The
dataset used to develop the model is available from the project website at
<ext-link xlink:href="https://doi.org/10.5281/zenodo.6892062" ext-link-type="DOI">10.5281/zenodo.6892062</ext-link> (Gao, 2022b) under the Creative Commons
Attribution 4.0 International license.</p>
  </notes><app-group>
        <supplementary-material position="anchor"><p id="d1e3704">The supplement related to this article is available online at: <inline-supplementary-material xlink:href="https://doi.org/10.5194/gmd-15-9015-2022-supplement" xlink:title="pdf">https://doi.org/10.5194/gmd-15-9015-2022-supplement</inline-supplementary-material>.</p></supplementary-material>
        </app-group><notes notes-type="authorcontribution"><title>Author contributions</title>

      <p id="d1e3713">PV, CEI and AGR conceived the research.
ZG and AGR designed and performed the research.
ZG and KD collected the observed data.
ZG analyzed the data, built the models and interpreted the results.
All authors contributed to editing the manuscript.</p>
  </notes><notes notes-type="competinginterests"><title>Competing interests</title>

      <p id="d1e3719">The contact author has declared that none of the authors has any competing interests.</p>
  </notes><notes notes-type="disclaimer"><title>Disclaimer</title>

      <p id="d1e3725">Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.</p>
  </notes><ack><title>Acknowledgements</title><p id="d1e3731">We would like to thank
Charles L. Blanchard for his comments that helped us significantly improve
the manuscript.</p></ack><notes notes-type="financialsupport"><title>Financial support</title>

      <p id="d1e3736">This research has been supported by the South Coast Air Quality Management District (grant no. 20058), a generous donation from Howard T. Tellepsen, NASA's Health and Air Quality Applied Sciences Team (HAQAST) program (grant
no. 80NSSC21K0506), the Phillips 66 Company, and the US Environmental Protection Agency (EPA; grant no. 84024601).</p>
  </notes><notes notes-type="reviewstatement"><title>Review statement</title>

      <p id="d1e3742">This paper was edited by Volker Grewe and reviewed by William Stockwell and one anonymous referee.</p>
  </notes><ref-list>
    <title>References</title>

      <ref id="bib1.bib1"><label>1</label><?label 1?><mixed-citation>Agarwal, R. and Sen, S.: Creators of Mathematical and Computational
Sciences, <ext-link xlink:href="https://doi.org/10.1007/978-3-319-10870-4" ext-link-type="DOI">10.1007/978-3-319-10870-4</ext-link>, 2014.</mixed-citation></ref>
      <ref id="bib1.bib2"><label>2</label><?label 1?><mixed-citation>Ainsworth, E. A., Yendrek, C. R., Sitch, S., Collins, W. J., and Emberson,
L. D.: The Effects of Tropospheric Ozone on Net Primary Productivity and
Implications for Climate Change, Annu. Rev. Plant Biol., 63,
637–661, <ext-link xlink:href="https://doi.org/10.1146/annurev-arplant-042110-103829" ext-link-type="DOI">10.1146/annurev-arplant-042110-103829</ext-link>, 2012.</mixed-citation></ref>
      <ref id="bib1.bib3"><label>3</label><?label 1?><mixed-citation>Aldrin, M. and Haff, I.: Generalised additive modelling of air pollution,
traffic volume and meteorology, Atmos. Environ., 39, 2145–2155,
<ext-link xlink:href="https://doi.org/10.1016/j.atmosenv.2004.12.020" ext-link-type="DOI">10.1016/j.atmosenv.2004.12.020</ext-link>, 2005.</mixed-citation></ref>
      <ref id="bib1.bib4"><label>4</label><?label 1?><mixed-citation>Alduchov, O. A. and Eskridge, R. E.: Improved Magnus Form Approximation of
Saturation Vapor Pressure, J. Appl. Meteorol., 35, 601–609,
<ext-link xlink:href="https://doi.org/10.1175/1520-0450(1996)035&lt;0601:imfaos&gt;2.0.co;2" ext-link-type="DOI">10.1175/1520-0450(1996)035&lt;0601:imfaos&gt;2.0.co;2</ext-link>, 1996.</mixed-citation></ref>
      <ref id="bib1.bib5"><label>5</label><?label 1?><mixed-citation>Aw, J. and Kleeman, M. J.: Evaluating the first-order effect of intraannual
temperature variability on urban air pollution, J. Geophys.
Res., 108, 4365, <ext-link xlink:href="https://doi.org/10.1029/2002jd002688" ext-link-type="DOI">10.1029/2002jd002688</ext-link>, 2003.</mixed-citation></ref>
      <ref id="bib1.bib6"><label>6</label><?label 1?><mixed-citation>Banta, J. E.: Sir William Petty: Modern epidemiologist (1623–1687), J. Commun. He., 12, 185–198, <ext-link xlink:href="https://doi.org/10.1007/bf01323480" ext-link-type="DOI">10.1007/bf01323480</ext-link>, 1987.</mixed-citation></ref>
      <ref id="bib1.bib7"><label>7</label><?label 1?><mixed-citation>Benirschke, K.: Francis Galton: Pioneer of Heredity and Biometry, J.
Heredity, 95, 273–273, <ext-link xlink:href="https://doi.org/10.1093/jhered/esh039" ext-link-type="DOI">10.1093/jhered/esh039</ext-link>, 2004.</mixed-citation></ref>
      <ref id="bib1.bib8"><label>8</label><?label 1?><mixed-citation>Blanchard, C. L., Hidy, G. M., and Tanenbaum, S.: Ozone in the southeastern
United States: An observation-based model using measurements from the SEARCH
network, Atmos. Environ., 88, 192–200, <ext-link xlink:href="https://doi.org/10.1016/j.atmosenv.2014.02.006" ext-link-type="DOI">10.1016/j.atmosenv.2014.02.006</ext-link>, 2014.</mixed-citation></ref>
      <ref id="bib1.bib9"><label>9</label><?label 1?><mixed-citation>Blanchard, C. L., Shaw, S. L., Edgerton, E. S., and Schwab, J. J.: Emission
influences on air pollutant concentrations in New York State: I. ozone,
Atmos. Environ. X, 3, 100033, <ext-link xlink:href="https://doi.org/10.1016/j.aeaoa.2019.100033" ext-link-type="DOI">10.1016/j.aeaoa.2019.100033</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bib10"><label>10</label><?label 1?><mixed-citation>Camalier, L., Cox, W., and Dolwick, P.: The effects of meteorology on ozone
in urban areas and their use in assessing ozone trends, Atmos.
Environ., 41, 7127–7137, <ext-link xlink:href="https://doi.org/10.1016/j.atmosenv.2007.04.061" ext-link-type="DOI">10.1016/j.atmosenv.2007.04.061</ext-link>, 2007.</mixed-citation></ref>
      <ref id="bib1.bib11"><label>11</label><?label 1?><mixed-citation>CARB: Air Quality and Meteorological Information System (AQMIS), <uri>https://www.arb.ca.gov/aqmis2/aqdselect.php</uri>,
last access: 27 May 2020.</mixed-citation></ref>
      <ref id="bib1.bib12"><label>12</label><?label 1?><mixed-citation>Chang, C.-C. and Lin, C.-J.: LIBSVM, ACM Transactions on Intelligent Systems
and Technology, 2, 1–27, <ext-link xlink:href="https://doi.org/10.1145/1961189.1961199" ext-link-type="DOI">10.1145/1961189.1961199</ext-link>, 2011.</mixed-citation></ref>
      <ref id="bib1.bib13"><label>13</label><?label 1?><mixed-citation>
Cox, P., Delao, A., Komorniczak, A., and Weller, R.: The California Almanac of Emissions and Air Quality – 2009 edition, Planning and Technical Support Division California Air Resources Board,
2009.</mixed-citation></ref>
      <ref id="bib1.bib14"><label>14</label><?label 1?><mixed-citation>
Cox, P., Delao, A., and Komorniczak, A.: The California Almanac of Emissions and Air Quality – 2013 edition, Air Quality Planning and Science Division California Air Resources Board,
2013.</mixed-citation></ref>
      <ref id="bib1.bib15"><label>15</label><?label 1?><mixed-citation>
Drucker, H., Burges, C. J. C., Kaufman, L., Smola, A., and Vapnik, V.:
Support vector regression machines, Proceedings of the 9th International
Conference on Neural Information Processing Systems, Denver, Colorado, 1996.</mixed-citation></ref>
      <ref id="bib1.bib16"><label>16</label><?label 1?><mixed-citation>
Fan, R.-E., Chen, P.-H., and Lin, C.-J.: Working Set Selection Using Second
Order Information for Training Support Vector Machines, J. Mach. Learn.
Res., 6, 1889–1918, 2005.</mixed-citation></ref>
      <ref id="bib1.bib17"><label>17</label><?label 1?><mixed-citation>Fisher, R. A.: On the mathematical foundations of theoretical statistics,
Philos. T. Roy. Soc. Lond. A, 222, 309–368,
<ext-link xlink:href="https://doi.org/10.1098/rsta.1922.0009" ext-link-type="DOI">10.1098/rsta.1922.0009</ext-link>, 1922.</mixed-citation></ref>
      <ref id="bib1.bib18"><label>18</label><?label 1?><mixed-citation>Flynn, M. T., Mattson, E. J., Jaffe, D. A., and Gratz, L. E.: Spatial
patterns in summertime surface ozone in the Southern Front Range of the U.S.
Rocky Mountains, Elementa: Science of the Anthropocene, 9, 00104,
<ext-link xlink:href="https://doi.org/10.1525/elementa.2020.00104" ext-link-type="DOI">10.1525/elementa.2020.00104</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bib19"><label>19</label><?label 1?><mixed-citation>Friedman, J. H.: Multivariate Adaptive Regression Splines, Ann.
Stat., 19, 1–67, <ext-link xlink:href="https://doi.org/10.1214/aos/1176347963" ext-link-type="DOI">10.1214/aos/1176347963</ext-link>, 1991.</mixed-citation></ref>
      <ref id="bib1.bib20"><label>20</label><?label 1?><mixed-citation>Friedman, J. H. and Silverman, B. W.: Flexible Parsimonious Smoothing and
Additive Modeling, Technometrics, 31, 3–21, <ext-link xlink:href="https://doi.org/10.2307/1270359" ext-link-type="DOI">10.2307/1270359</ext-link>, 1989.</mixed-citation></ref>
      <ref id="bib1.bib21"><label>21</label><?label 1?><mixed-citation>
Galton, F.: Co-Relations and Their Measurement, Chiefly from Anthropometric
Data, P. Roy. Soc. Lond., 45, 135–145, 1888.</mixed-citation></ref>
      <ref id="bib1.bib22"><label>22</label><?label 1?><mixed-citation>Galton, F. S.: Natural inheritance, Macmillan, London, <ext-link xlink:href="https://doi.org/10.5962/bhl.title.32181" ext-link-type="DOI">10.5962/bhl.title.32181</ext-link>, 1889.</mixed-citation></ref>
      <ref id="bib1.bib23"><label>23</label><?label 1?><mixed-citation>Gao, Z.: Predicting peak daily maximum 8-hour ozone, and linkages to emissions and meteorology, in Southern California using machine learning methods, Zenodo [code], <ext-link xlink:href="https://doi.org/10.5281/zenodo.6892066" ext-link-type="DOI">10.5281/zenodo.6892066</ext-link>, 2022a.</mixed-citation></ref>
      <ref id="bib1.bib24"><label>24</label><?label 1?><mixed-citation>Gao, Z.: Predicting peak daily maximum 8-hour ozone, and linkages to emissions and meteorology, in Southern California using machine learning methods, Zenodo [data set], <ext-link xlink:href="https://doi.org/10.5281/zenodo.6892062" ext-link-type="DOI">10.5281/zenodo.6892062</ext-link>, 2022b.</mixed-citation></ref>
      <ref id="bib1.bib25"><label>25</label><?label 1?><mixed-citation>Gao, Z., Ivey, C. E., Blanchard, C. L., Do, K., Lee, S.-M., and Russell, A.
G.: Separating emissions and meteorological impacts on peak ozone
concentrations in Southern California using generalized additive modeling,
Environ. Pollut., 307, 119503, <ext-link xlink:href="https://doi.org/10.1016/j.envpol.2022.119503" ext-link-type="DOI">10.1016/j.envpol.2022.119503</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bib26"><label>26</label><?label 1?><mixed-citation>García Nieto, P. J. and Álvarez Antón, J. C.: Nonlinear air
quality modeling using multivariate adaptive regression splines in Gijón
urban area (Northern Spain) at local scale, Appl. Math.
Comput., 235, 50–65, <ext-link xlink:href="https://doi.org/10.1016/j.amc.2014.02.096" ext-link-type="DOI">10.1016/j.amc.2014.02.096</ext-link>, 2014.</mixed-citation></ref>
      <ref id="bib1.bib27"><label>27</label><?label 1?><mixed-citation>Gong, X., Kaulfus, A., Nair, U., and Jaffe, D. A.: Quantifying O3 Impacts in
Urban Areas Due to Wildfires Using a Generalized Additive Model,
Environ. Sci. Technol., 51, 13216–13223,
<ext-link xlink:href="https://doi.org/10.1021/acs.est.7b03130" ext-link-type="DOI">10.1021/acs.est.7b03130</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bib28"><label>28</label><?label 1?><mixed-citation>Gong, X., Hong, S., and Jaffe, D. A.: Ozone in China: Spatial Distribution
and Leading Meteorological Factors Controlling O<inline-formula><mml:math id="M126" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">3</mml:mn></mml:msub></mml:math></inline-formula> in 16 Chinese Cities,
Aerosol Air Qual. Res., 18, 2287–2300, <ext-link xlink:href="https://doi.org/10.4209/aaqr.2017.10.0368" ext-link-type="DOI">10.4209/aaqr.2017.10.0368</ext-link>,
2018.</mixed-citation></ref>
      <ref id="bib1.bib29"><label>29</label><?label 1?><mixed-citation>Gorai, A. K., Tuluri, F., Tchounwou, P. B., and Ambinakudige, S.: Influence
of local meteorology and NO<inline-formula><mml:math id="M127" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> conditions on ground-level ozone concentrations
in the eastern part of Texas, USA, Air Quality, Atmos. He., 8,
81–96, <ext-link xlink:href="https://doi.org/10.1007/s11869-014-0276-5" ext-link-type="DOI">10.1007/s11869-014-0276-5</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bib30"><label>30</label><?label 1?><mixed-citation>Harrell, F. E.: General Aspects of Fitting Regression Models, Springer
International Publishing, 13–44, <ext-link xlink:href="https://doi.org/10.1007/978-3-319-19425-7_2" ext-link-type="DOI">10.1007/978-3-319-19425-7_2</ext-link>,
2015.</mixed-citation></ref>
      <ref id="bib1.bib31"><label>31</label><?label 1?><mixed-citation>
Hastie, T.: Generalized additive models, Chapter 7, in: Statistical Models in S, edited by: Chambers, J. M. and Hastie,
T. J., Wadsworth &amp; Brooks/Cole, 1991.</mixed-citation></ref>
      <ref id="bib1.bib32"><label>32</label><?label 1?><mixed-citation>Hastie, T. and Tibshirani, R.: Generalized Additive Models, Stat.
Sci., 1, 297–310, <ext-link xlink:href="https://doi.org/10.1214/ss/1177013604" ext-link-type="DOI">10.1214/ss/1177013604</ext-link>, 1986.</mixed-citation></ref>
      <ref id="bib1.bib33"><label>33</label><?label 1?><mixed-citation>
Hastie, T. and Tibshirani, R.: Generalized additive models, Chapman &amp;
Hall/CRC, London, 1990.</mixed-citation></ref>
      <ref id="bib1.bib34"><label>34</label><?label 1?><mixed-citation>
Hastie, T. and Tibshirani, R.: Discriminant Analysis by Gaussian Mixtures,
J. Roy. Stat. Soc. B, 58,
155–176, 1996.</mixed-citation></ref>
      <ref id="bib1.bib35"><label>35</label><?label 1?><mixed-citation>
Hastie, T., Tibshirani, R., and Friedman, J. H.: The elements of statistical
learning: data mining, inference, and prediction, 2nd edn.,
Springer, New York, 2009.</mixed-citation></ref>
      <ref id="bib1.bib36"><label>36</label><?label 1?><mixed-citation>Henneman, L. R. F., Holmes, H. A., Mulholland, J. A., and Russell, A. G.:
Meteorological detrending of primary and secondary pollutant concentrations:
Method application and evaluation using long-term (2000–2012) data in
Atlanta, Atmos. Environ., 119, 201–210, <ext-link xlink:href="https://doi.org/10.1016/j.atmosenv.2015.08.007" ext-link-type="DOI">10.1016/j.atmosenv.2015.08.007</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bib37"><label>37</label><?label 1?><mixed-citation>Hogrefe, C., Rao, S. T., Zurbenko, I. G., and Porter, P. S.: Interpreting
the Information in Ozone Observations and Model Predictions Relevant to
Regulatory Policies in the Eastern United States, B. Am.
Meteorol. Soc., 81, 2083–2106, <ext-link xlink:href="https://doi.org/10.1175/1520-0477(2000)081&lt;2083:itiioo&gt;2.3.co;2" ext-link-type="DOI">10.1175/1520-0477(2000)081&lt;2083:itiioo&gt;2.3.co;2</ext-link>, 2000.</mixed-citation></ref>
      <ref id="bib1.bib38"><label>38</label><?label 1?><mixed-citation>Hong, C., Mueller, N. D., Burney, J. A., Zhang, Y., AghaKouchak, A., Moore,
F. C., Qin, Y., Tong, D., and Davis, S. J.: Impacts of ozone and climate
change on yields of perennial crops in California, Nature Food, 1, 166–172,
<ext-link xlink:href="https://doi.org/10.1038/s43016-020-0043-8" ext-link-type="DOI">10.1038/s43016-020-0043-8</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bib39"><label>39</label><?label 1?><mixed-citation>Hu, C., Kang, P., Jaffe, D. A., Li, C., Zhang, X., Wu, K., and Zhou, M.:
Understanding the impact of meteorology on ozone in 334 cities of China,
Atmos. Environ., 248, 118221, <ext-link xlink:href="https://doi.org/10.1016/j.atmosenv.2021.118221" ext-link-type="DOI">10.1016/j.atmosenv.2021.118221</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bib40"><label>40</label><?label 1?><mixed-citation>Huang, X. G., Shao, T. J., Zhao, J. B., Cao, J. J., and Lü, X. H.:
Influencing Factors of Ozone Concentration in Xi'an Based on Generalized
Additive Models, Huan Jing Ke Xue, 41, 1535–1543,
<ext-link xlink:href="https://doi.org/10.13227/j.hjkx.201906067" ext-link-type="DOI">10.13227/j.hjkx.201906067</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bib41"><label>41</label><?label 1?><mixed-citation>Jeong, Y., Lee, H. W., and Jeon, W.: Regional Differences of Primary
Meteorological Factors Impacting O<inline-formula><mml:math id="M128" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">3</mml:mn></mml:msub></mml:math></inline-formula> Variability in South Korea, Atmosphere,
11, 74, <ext-link xlink:href="https://doi.org/10.3390/atmos11010074" ext-link-type="DOI">10.3390/atmos11010074</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bib42"><label>42</label><?label 1?><mixed-citation>Jin, L., Loisy, A., and Brown, N. J.: Role of meteorological processes in
ozone responses to emission controls in California's San Joaquin Valley,
J. Geophys. Res.-Atmos., 118, 8010–8022,
<ext-link xlink:href="https://doi.org/10.1002/jgrd.50559" ext-link-type="DOI">10.1002/jgrd.50559</ext-link>, 2013.</mixed-citation></ref>
      <ref id="bib1.bib43"><label>43</label><?label 1?><mixed-citation>Keller, C. A. and Evans, M. J.: Application of random forest regression to the calculation of gas-phase chemistry within the GEOS-Chem chemistry model v10, Geosci. Model Dev., 12, 1209–1225, <ext-link xlink:href="https://doi.org/10.5194/gmd-12-1209-2019" ext-link-type="DOI">10.5194/gmd-12-1209-2019</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bib44"><label>44</label><?label 1?><mixed-citation>Kelley, M. C., Brown, M. M., Fedler, C. B., and Ardon-Dryer, K.: Long Term
Measurements of PM<inline-formula><mml:math id="M129" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2.5</mml:mn></mml:msub></mml:math></inline-formula> Concentrations in Lubbock, Texas, Aerosol Air
Qual. Res., 20, 1306–1318, <ext-link xlink:href="https://doi.org/10.4209/aaqr.2019.09.0469" ext-link-type="DOI">10.4209/aaqr.2019.09.0469</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bib45"><label>45</label><?label 1?><mixed-citation>Kleeman, M. J.: A preliminary assessment of the sensitivity of air quality
in California to global change, Clim. Change, 87, 273–292,
<ext-link xlink:href="https://doi.org/10.1007/s10584-007-9351-3" ext-link-type="DOI">10.1007/s10584-007-9351-3</ext-link>, 2008.</mixed-citation></ref>
      <ref id="bib1.bib46"><label>46</label><?label 1?><mixed-citation>Lawrence, M. G.: The Relationship between Relative Humidity and the Dewpoint
Temperature in Moist Air: A Simple Conversion and Applications, B.
Am. Meteorol. Soc., 86, 225–234, <ext-link xlink:href="https://doi.org/10.1175/BAMS-86-2-225" ext-link-type="DOI">10.1175/BAMS-86-2-225</ext-link>,
2005.</mixed-citation></ref>
      <ref id="bib1.bib47"><label>47</label><?label 1?><mixed-citation>Leathwick, J. R., Elith, J., and Hastie, T.: Comparative performance of
generalized additive models and multivariate adaptive regression splines for
statistical modelling of species distributions, Ecol. Model., 199,
188–196, <ext-link xlink:href="https://doi.org/10.1016/j.ecolmodel.2006.05.022" ext-link-type="DOI">10.1016/j.ecolmodel.2006.05.022</ext-link>, 2006.</mixed-citation></ref>
      <ref id="bib1.bib48"><label>48</label><?label 1?><mixed-citation>
Liaw, A. and Wiener, M.: Classification and Regression by randomForest, R
News, 2, 18–22, 2002.</mixed-citation></ref>
      <ref id="bib1.bib49"><label>49</label><?label 1?><mixed-citation>Ling, Z. H., Guo, H., Zheng, J. Y., Louie, P. K. K., Cheng, H. R., Jiang,
F., Cheung, K., Wong, L. C., and Feng, X. Q.: Establishing a conceptual
model for photochemical ozone pollution in subtropical Hong Kong,
Atmos. Environ., 76, 208–220, <ext-link xlink:href="https://doi.org/10.1016/j.atmosenv.2012.09.051" ext-link-type="DOI">10.1016/j.atmosenv.2012.09.051</ext-link>, 2013.</mixed-citation></ref>
      <ref id="bib1.bib50"><label>50</label><?label 1?><mixed-citation>Liu, B.-C., Binaykia, A., Chang, P.-C., Tiwari, M. K., and Tsao, C.-C.:
Urban air quality forecasting based on multi-dimensional collaborative
Support Vector Regression (SVR): A case study of
Beijing-Tianjin-Shijiazhuang, PLOS ONE, 12, e0179763,
<ext-link xlink:href="https://doi.org/10.1371/journal.pone.0179763" ext-link-type="DOI">10.1371/journal.pone.0179763</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bib51"><label>51</label><?label 1?><mixed-citation>Liu, T., Li, T. T., Zhang, Y. H., Xu, Y. J., Lao, X. Q., Rutherford, S.,
Chu, C., Luo, Y., Zhu, Q., Xu, X. J., Xie, H. Y., Liu, Z. R., and Ma, W. J.:
The short-term effect of ambient ozone on mortality is modified by
temperature in Guangzhou, China, Atmos. Environ., 76, 59–67,
<ext-link xlink:href="https://doi.org/10.1016/j.atmosenv.2012.07.011" ext-link-type="DOI">10.1016/j.atmosenv.2012.07.011</ext-link>, 2013.</mixed-citation></ref>
      <ref id="bib1.bib52"><label>52</label><?label 1?><mixed-citation>Lu, R. and Turco, R. P.: Ozone distributions over the los angeles basin:
Three-dimensional simulations with the smog model, Atmos. Environ.,
30, 4155–4176, <ext-link xlink:href="https://doi.org/10.1016/1352-2310(96)00153-7" ext-link-type="DOI">10.1016/1352-2310(96)00153-7</ext-link>, 1996.</mixed-citation></ref>
      <ref id="bib1.bib53"><label>53</label><?label 1?><mixed-citation>Lu, X., Zhang, L., and Shen, L.: Meteorology and Climate Influences on
Tropospheric Ozone: a Review of Natural Sources, Chemistry, and Transport
Patterns, Current Pollut. Rep., 5, 238–260, <ext-link xlink:href="https://doi.org/10.1007/s40726-019-00118-3" ext-link-type="DOI">10.1007/s40726-019-00118-3</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bib54"><label>54</label><?label 1?><mixed-citation>Luna, A. S., Paredes, M. L. L., De Oliveira, G. C. G., and Corrêa, S.
M.: Prediction of ozone concentration in tropospheric levels using
artificial neural networks and support vector machine at Rio de Janeiro,
Brazil, Atmos. Environ., 98, 98–104, <ext-link xlink:href="https://doi.org/10.1016/j.atmosenv.2014.08.060" ext-link-type="DOI">10.1016/j.atmosenv.2014.08.060</ext-link>,
2014.</mixed-citation></ref>
      <ref id="bib1.bib55"><label>55</label><?label 1?><mixed-citation>Ma, Y., Ma, B., Jiao, H., Zhang, Y., Xin, J., and Yu, Z.: An analysis of the
effects of weather and air pollution on tropospheric ozone using a
generalized additive model in Western China: Lanzhou, Gansu, Atmos.
Environ., 224, 117342, <ext-link xlink:href="https://doi.org/10.1016/j.atmosenv.2020.117342" ext-link-type="DOI">10.1016/j.atmosenv.2020.117342</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bib56"><label>56</label><?label 1?><mixed-citation>Mahmud, A., Hixson, M., Hu, J., Zhao, Z., Chen, S.-H., and Kleeman, M. J.: Climate impact on airborne particulate matter concentrations in California using seven year analysis periods, Atmos. Chem. Phys., 10, 11097–11114, <ext-link xlink:href="https://doi.org/10.5194/acp-10-11097-2010" ext-link-type="DOI">10.5194/acp-10-11097-2010</ext-link>, 2010.</mixed-citation></ref>
      <ref id="bib1.bib57"><label>57</label><?label 1?><mixed-citation>McClure, C. D. and Jaffe, D. A.: Investigation of high ozone events due to
wildfire smoke in an urban area, Atmos. Environ., 194, 146–157,
<ext-link xlink:href="https://doi.org/10.1016/j.atmosenv.2018.09.021" ext-link-type="DOI">10.1016/j.atmosenv.2018.09.021</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bib58"><label>58</label><?label 1?><mixed-citation>McGlynn, D., Mao, H., Sive, B., and Sharac, T.: Understanding Long-Term
Variations in Surface Ozone in United States (U.S.) National Parks,
Atmosphere, 9, 125, <ext-link xlink:href="https://doi.org/10.3390/atmos9040125" ext-link-type="DOI">10.3390/atmos9040125</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bib59"><label>59</label><?label 1?><mixed-citation>Menne, M. J., Durre, I., Korzeniewski, B., McNeal, S., Thomas, K., Yin, X.,
Anthony, S., Ray, R., Vose, R. S., Gleason, B. E., and Houston, T. G.:
Global Historical Climatology Network - Daily (GHCN-Daily), Version 3,
<ext-link xlink:href="https://doi.org/10.7289/V5D21VHZ" ext-link-type="DOI">10.7289/V5D21VHZ</ext-link>, 2012a.</mixed-citation></ref>
      <ref id="bib1.bib60"><label>60</label><?label 1?><mixed-citation>Menne, M. J., Durre, I., Vose, R. S., Gleason, B. E., and Houston, T. G.: An
Overview of the Global Historical Climatology Network-Daily Database,
J. Atmos. Ocean. Tech., 29, 897–910,
<ext-link xlink:href="https://doi.org/10.1175/jtech-d-11-00103.1" ext-link-type="DOI">10.1175/jtech-d-11-00103.1</ext-link>, 2012b.</mixed-citation></ref>
      <ref id="bib1.bib61"><label>61</label><?label 1?><mixed-citation>Milborrow, S.: Regression Splines,
FIM Forecast Model: 500 mb Wind Speed and 500 mb Height Contours – Real-time:
<uri>https://CRAN.R-project.org/package=earth</uri> (last access: 13 November 2021),
R program [code], 2021.</mixed-citation></ref>
      <ref id="bib1.bib62"><label>62</label><?label 1?><mixed-citation>NOAA: FIM Forecast Model: 500 mb Wind Speed and 500 mb Height Contours – Real-time, <ext-link xlink:href="https://sos.noaa.gov/datasets/fim-forecast-model-500mb-wind-speed-and-500mb-height-contours-real-time/">https://sos.noaa.gov/datasets/fim-forecast-model-500mb-wind-speed-and-500mb-height-contours-real-time/</ext-link> (last access: 23 May 2021), 2020.</mixed-citation></ref>
      <ref id="bib1.bib63"><label>63</label><?label 1?><mixed-citation>Oduro, S. D., Metia, S., Duc, H., Hong, G., and Ha, Q. P.: Multivariate
adaptive regression splines models for vehicular emission prediction,
Visualization in Engineering, 3, 13, <ext-link xlink:href="https://doi.org/10.1186/s40327-015-0024-4" ext-link-type="DOI">10.1186/s40327-015-0024-4</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bib64"><label>64</label><?label 1?><mixed-citation>Oman, L. D., Ziemke, J. R., Douglass, A. R., Waugh, D. W., Lang, C.,
Rodriguez, J. M., and Nielsen, J. E.: The response of tropical tropospheric
ozone to ENSO, Geophys. Res. Lett., 38, 13706, <ext-link xlink:href="https://doi.org/10.1029/2011GL047865" ext-link-type="DOI">10.1029/2011GL047865</ext-link>, 2011.</mixed-citation></ref>
      <ref id="bib1.bib65"><label>65</label><?label 1?><mixed-citation>Oman, L. D., Douglass, A. R., Ziemke, J. R., Rodriguez, J. M., Waugh, D. W.,
and Nielsen, J. E.: The ozone response to ENSO in Aura satellite
measurements and a chemistry-climate simulation, J. Geophys. Res.-Atmos., 118, 965–976,
<ext-link xlink:href="https://doi.org/10.1029/2012jd018546" ext-link-type="DOI">10.1029/2012jd018546</ext-link>, 2013.</mixed-citation></ref>
      <ref id="bib1.bib66"><label>66</label><?label 1?><mixed-citation>Pearce, J. L., Beringer, J., Nicholls, N., Hyndman, R. J., and Tapper, N.
J.: Quantifying the influence of local meteorology on air quality using
generalized additive models, Atmos. Environ., 45, 1328–1336,
<ext-link xlink:href="https://doi.org/10.1016/j.atmosenv.2010.11.051" ext-link-type="DOI">10.1016/j.atmosenv.2010.11.051</ext-link>, 2011.</mixed-citation></ref>
      <ref id="bib1.bib67"><label>67</label><?label 1?><mixed-citation>Pernak, R., Alvarado, M., Lonsdale, C., Mountain, M., Hegarty, J., and
Nehrkorn, T.: Forecasting Surface O<inline-formula><mml:math id="M130" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">3</mml:mn></mml:msub></mml:math></inline-formula> in Texas Urban Areas Using Random
Forest and Generalized Additive Models, Aerosol Air Qual. Res., 9,
2815–2826, <ext-link xlink:href="https://doi.org/10.4209/aaqr.2018.12.0464" ext-link-type="DOI">10.4209/aaqr.2018.12.0464</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bib68"><label>68</label><?label 1?><mixed-citation>
Pope, P. T. and Webster, J. T.: The Use of an F-Statistic in Stepwise
Regression Procedures, Technometrics, 14, 327–340, 1972.</mixed-citation></ref>
      <ref id="bib1.bib69"><label>69</label><?label 1?><mixed-citation>Porter, T. M.: Social Interests and Statistical Theory: Statistics in
Britain, 1865–1930, Science, 214, 784–784,
<ext-link xlink:href="https://doi.org/10.1126/science.214.4522.784.a" ext-link-type="DOI">10.1126/science.214.4522.784.a</ext-link>, 1981.</mixed-citation></ref>
      <ref id="bib1.bib70"><label>70</label><?label 1?><mixed-citation>
Porter, T. M.: Trust in Numbers, Princeton University Press, ISBN 9780691208411, 1995.</mixed-citation></ref>
      <ref id="bib1.bib71"><label>71</label><?label 1?><mixed-citation>Rao, S. T., Zurbenko, I. G., Neagu, R., Porter, P. S., Ku, J. Y., and Henry,
R. F.: Space and Time Scales in Ambient Ozone Data, B. Am.
Meteorol. Soc., 78, 2153–2166, <ext-link xlink:href="https://doi.org/10.1175/1520-0477(1997)078&lt;2153:satsia&gt;2.0.co;2" ext-link-type="DOI">10.1175/1520-0477(1997)078&lt;2153:satsia&gt;2.0.co;2</ext-link>, 1997.</mixed-citation></ref>
      <ref id="bib1.bib72"><label>72</label><?label 1?><mixed-citation>Rodríguez-Pérez, R., Vogt, M., and Bajorath, J.: Support Vector
Machine Classification and Regression Prioritize Different Structural
Features for Binary Compound Activity and Potency Value Prediction, ACS
Omega, 2, 6371–6379, <ext-link xlink:href="https://doi.org/10.1021/acsomega.7b01079" ext-link-type="DOI">10.1021/acsomega.7b01079</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bib73"><label>73</label><?label 1?><mixed-citation>Roy, S. S., Pratyush, C., and Barna, C.: Predicting Ozone Layer
Concentration Using Multivariate Adaptive Regression Splines, Random Forest
and Classification and Regression Tree, Springer International
Publishing, 140–152, <ext-link xlink:href="https://doi.org/10.1007/978-3-319-62524-9_11" ext-link-type="DOI">10.1007/978-3-319-62524-9_11</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bib74"><label>74</label><?label 1?><mixed-citation>Rybarczyk, Y. and Zalakeviciute, R.: Machine Learning Approaches for Outdoor
Air Quality Modelling: A Systematic Review, Appl. Sci., 8, 2570,
<ext-link xlink:href="https://doi.org/10.3390/app8122570" ext-link-type="DOI">10.3390/app8122570</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bib75"><label>75</label><?label 1?><mixed-citation>
Schölkopf, B. and Smola, A. J.: Learning with Kernels: Support Vector
Machines, Regularization, Optimization, and Beyond, MIT Press, ISBN 9780262536578, 2001.</mixed-citation></ref>
      <ref id="bib1.bib76"><label>76</label><?label 1?><mixed-citation>
Seinfeld, J. H. and Pandis, S. N.: Atmospheric Chemistry and Physics: From
Air Pollution to Climate Change, Wiley, ISBN 978-1-118-94740-1, 2016.</mixed-citation></ref>
      <ref id="bib1.bib77"><label>77</label><?label 1?><mixed-citation>Smola, A. J. and Schölkopf, B.: A tutorial on support vector regression,
Stat. Comput., 14, 199–222, <ext-link xlink:href="https://doi.org/10.1023/b:stco.0000035301.49549.88" ext-link-type="DOI">10.1023/b:stco.0000035301.49549.88</ext-link>,
2004.</mixed-citation></ref>
      <ref id="bib1.bib78"><label>78</label><?label 1?><mixed-citation>Sotomayor-Olmedo, A., Aceves-Fernández, M. A., Gorrostieta-Hurtado, E.,
Pedraza-Ortega, C., Ramos-Arreguín, J. M., and Vargas-Soto, J. E.:
Forecast Urban Air Pollution in Mexico City by Using Support Vector
Machines: A Kernel Performance Approach, International Journal of
Intelligence Science, 03, 126–135, <ext-link xlink:href="https://doi.org/10.4236/ijis.2013.33014" ext-link-type="DOI">10.4236/ijis.2013.33014</ext-link>, 2013.</mixed-citation></ref>
      <ref id="bib1.bib79"><label>79</label><?label 1?><mixed-citation>Stafoggia, M., Johansson, C., Glantz, P., Renzi, M., Shtein, A., De Hoogh,
K., Kloog, I., Davoli, M., Michelozzi, P., and Bellander, T.: A Random
Forest Approach to Estimate Daily Particulate Matter, Nitrogen Dioxide, and
Ozone at Fine Spatial Resolution in Sweden, Atmosphere, 11, 239,
<ext-link xlink:href="https://doi.org/10.3390/atmos11030239" ext-link-type="DOI">10.3390/atmos11030239</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bib80"><label>80</label><?label 1?><mixed-citation>Stephen, M. S.: Gauss and the Invention of Least Squares, Ann.
Stat., 9, 465–474, <ext-link xlink:href="https://doi.org/10.1214/aos/1176345451" ext-link-type="DOI">10.1214/aos/1176345451</ext-link>, 1981.</mixed-citation></ref>
      <ref id="bib1.bib81"><label>81</label><?label 1?><mixed-citation>Sun, L., Xue, L., Wang, Y., Li, L., Lin, J., Ni, R., Yan, Y., Chen, L., Li, J., Zhang, Q., and Wang, W.: Impacts of meteorology and emissions on summertime surface ozone increases over central eastern China between 2003 and 2015, Atmos. Chem. Phys., 19, 1455–1469, <ext-link xlink:href="https://doi.org/10.5194/acp-19-1455-2019" ext-link-type="DOI">10.5194/acp-19-1455-2019</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bib82"><label>82</label><?label 1?><mixed-citation>Tin Kam, H.: Random decision forests, Proceedings of 3rd International
Conference on Document Analysis and Recognition, 14–16 August 1995,
271, 278–282, <ext-link xlink:href="https://doi.org/10.1109/ICDAR.1995.598994" ext-link-type="DOI">10.1109/ICDAR.1995.598994</ext-link>, 1995.</mixed-citation></ref>
      <ref id="bib1.bib83"><label>83</label><?label 1?><mixed-citation>U.S. EPA:
Trends in Ozone Adjusted for Weather Conditions, <ext-link xlink:href="https://www.epa.gov/air-trends/trends-ozone-adjusted-weather-conditions">https://www.epa.gov/air-trends/trends-ozone-adjusted-weather-conditions</ext-link> (last access: 13 November 2021),
U.S. EPA, 2016.
</mixed-citation></ref><?xmltex \hack{\newpage}?>
      <ref id="bib1.bib84"><label>84</label><?label 1?><mixed-citation>
U.S. EPA: Integrated Science Assessment (ISA) for Ozone and Related Photochemical Oxidants (Final Report, Apr 2020), U.S. Environmental Protection Agency, Washington, D.C., EPA/600/R-20/012, 2020.</mixed-citation></ref>
      <ref id="bib1.bib85"><label>85</label><?label 1?><mixed-citation>Vong, C.-M., Ip, W.-F., Wong, P.-K., and Yang, J.-Y.: Short-Term Prediction
of Air Pollution in Macau Using Support Vector Machines, J. Control
Sci. Eng., 2012, 1–11, <ext-link xlink:href="https://doi.org/10.1155/2012/518032" ext-link-type="DOI">10.1155/2012/518032</ext-link>, 2012.</mixed-citation></ref>
      <ref id="bib1.bib86"><label>86</label><?label 1?><mixed-citation>Wells, B., Dolwick, P., Eder, B., Evangelista, M., Foley, K., Mannshardt,
E., Misenis, C., and Weishampel, A.: Improved estimation of trends in U.S.
ozone concentrations adjusted for interannual variability in meteorological
conditions, Atmos. Environ., 248, 118234,
<ext-link xlink:href="https://doi.org/10.1016/j.atmosenv.2021.118234" ext-link-type="DOI">10.1016/j.atmosenv.2021.118234</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bib87"><label>87</label><?label 1?><mixed-citation>Wikipedia Contributors: Multivariate adaptive regression spline, <uri>https://en.wikipedia.org/w/index.php?title=Multivariate_adaptive_regression_spline&amp;oldid=1083057440</uri>, last access: 19 April 2022.</mixed-citation></ref>
      <ref id="bib1.bib88"><label>88</label><?label 1?><mixed-citation>Wood, S. N.: Fast stable restricted maximum likelihood and marginal
likelihood estimation of semiparametric generalized linear models, J. Roy. Statist. Soc. B, 73,
3–36, <ext-link xlink:href="https://doi.org/10.1111/j.1467-9868.2010.00749.x" ext-link-type="DOI">10.1111/j.1467-9868.2010.00749.x</ext-link>, 2011.</mixed-citation></ref>
      <ref id="bib1.bib89"><label>89</label><?label 1?><mixed-citation>Wood, S. N.: Generalized Additive Models: An Introduction with R, 2nd edn.,
Chapman and Hall/CRC,  ISBN 9781315370279, <ext-link xlink:href="https://doi.org/10.1201/9781315370279" ext-link-type="DOI">10.1201/9781315370279</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bib90"><label>90</label><?label 1?><mixed-citation>Xu, L., Yu, J.-Y., Schnell, J. L., and Prather, M. J.: The Seasonality and
Geographic Dependence of ENSO Impacts on US Surface Ozone Variability, Geophys. Res. Lett.,
44, 3420–3428,
<ext-link xlink:href="https://doi.org/10.1002/2017gl073044" ext-link-type="DOI">10.1002/2017gl073044</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bib91"><label>91</label><?label 1?><mixed-citation>Zhan, Y., Luo, Y., Deng, X., Grieneisen, M. L., Zhang, M., and Di, B.:
Spatiotemporal prediction of daily ambient ozone levels across China using
random forest for human exposure assessment, Environ. Pollut., 233,
464–473, <ext-link xlink:href="https://doi.org/10.1016/j.envpol.2017.10.029" ext-link-type="DOI">10.1016/j.envpol.2017.10.029</ext-link>, 2018.</mixed-citation></ref>

  </ref-list></back>
    <!--<article-title-html>Predicting peak daily maximum 8&thinsp;h ozone and linkages to emissions and meteorology in Southern California using machine learning methods (SoCAB-8HR V1.0)</article-title-html>
<abstract-html/>
<ref-html id="bib1.bib1"><label>1</label><mixed-citation>
Agarwal, R. and Sen, S.: Creators of Mathematical and Computational
Sciences, <a href="https://doi.org/10.1007/978-3-319-10870-4" target="_blank">https://doi.org/10.1007/978-3-319-10870-4</a>, 2014.
</mixed-citation></ref-html>
<ref-html id="bib1.bib2"><label>2</label><mixed-citation>
Ainsworth, E. A., Yendrek, C. R., Sitch, S., Collins, W. J., and Emberson,
L. D.: The Effects of Tropospheric Ozone on Net Primary Productivity and
Implications for Climate Change, Annu. Rev. Plant Biol., 63,
637–661, <a href="https://doi.org/10.1146/annurev-arplant-042110-103829" target="_blank">https://doi.org/10.1146/annurev-arplant-042110-103829</a>, 2012.
</mixed-citation></ref-html>
<ref-html id="bib1.bib3"><label>3</label><mixed-citation>
Aldrin, M. and Haff, I.: Generalised additive modelling of air pollution,
traffic volume and meteorology, Atmos. Environ., 39, 2145–2155,
<a href="https://doi.org/10.1016/j.atmosenv.2004.12.020" target="_blank">https://doi.org/10.1016/j.atmosenv.2004.12.020</a>, 2005.
</mixed-citation></ref-html>
<ref-html id="bib1.bib4"><label>4</label><mixed-citation>
Alduchov, O. A. and Eskridge, R. E.: Improved Magnus Form Approximation of
Saturation Vapor Pressure, J. Appl. Meteorol., 35, 601–609,
<a href="https://doi.org/10.1175/1520-0450(1996)035&lt;0601:imfaos&gt;2.0.co;2" target="_blank">https://doi.org/10.1175/1520-0450(1996)035&lt;0601:imfaos&gt;2.0.co;2</a>, 1996.
</mixed-citation></ref-html>
<ref-html id="bib1.bib5"><label>5</label><mixed-citation>
Aw, J. and Kleeman, M. J.: Evaluating the first-order effect of intraannual
temperature variability on urban air pollution, J. Geophys.
Res., 108, 4365, <a href="https://doi.org/10.1029/2002jd002688" target="_blank">https://doi.org/10.1029/2002jd002688</a>, 2003.
</mixed-citation></ref-html>
<ref-html id="bib1.bib6"><label>6</label><mixed-citation>
Banta, J. E.: Sir William Petty: Modern epidemiologist (1623–1687), J. Commun. He., 12, 185–198, <a href="https://doi.org/10.1007/bf01323480" target="_blank">https://doi.org/10.1007/bf01323480</a>, 1987.
</mixed-citation></ref-html>
<ref-html id="bib1.bib7"><label>7</label><mixed-citation>
Benirschke, K.: Francis Galton: Pioneer of Heredity and Biometry, J.
Heredity, 95, 273–273, <a href="https://doi.org/10.1093/jhered/esh039" target="_blank">https://doi.org/10.1093/jhered/esh039</a>, 2004.
</mixed-citation></ref-html>
<ref-html id="bib1.bib8"><label>8</label><mixed-citation>
Blanchard, C. L., Hidy, G. M., and Tanenbaum, S.: Ozone in the southeastern
United States: An observation-based model using measurements from the SEARCH
network, Atmos. Environ., 88, 192–200, <a href="https://doi.org/10.1016/j.atmosenv.2014.02.006" target="_blank">https://doi.org/10.1016/j.atmosenv.2014.02.006</a>, 2014.
</mixed-citation></ref-html>
<ref-html id="bib1.bib9"><label>9</label><mixed-citation>
Blanchard, C. L., Shaw, S. L., Edgerton, E. S., and Schwab, J. J.: Emission
influences on air pollutant concentrations in New York State: I. ozone,
Atmos. Environ. X, 3, 100033, <a href="https://doi.org/10.1016/j.aeaoa.2019.100033" target="_blank">https://doi.org/10.1016/j.aeaoa.2019.100033</a>, 2019.
</mixed-citation></ref-html>
<ref-html id="bib1.bib10"><label>10</label><mixed-citation>
Camalier, L., Cox, W., and Dolwick, P.: The effects of meteorology on ozone
in urban areas and their use in assessing ozone trends, Atmos.
Environ., 41, 7127–7137, <a href="https://doi.org/10.1016/j.atmosenv.2007.04.061" target="_blank">https://doi.org/10.1016/j.atmosenv.2007.04.061</a>, 2007.
</mixed-citation></ref-html>
<ref-html id="bib1.bib11"><label>11</label><mixed-citation>
CARB: Air Quality and Meteorological Information System (AQMIS), <a href="https://www.arb.ca.gov/aqmis2/aqdselect.php" target="_blank"/>,
last access: 27 May 2020.
</mixed-citation></ref-html>
<ref-html id="bib1.bib12"><label>12</label><mixed-citation>
Chang, C.-C. and Lin, C.-J.: LIBSVM, ACM Transactions on Intelligent Systems
and Technology, 2, 1–27, <a href="https://doi.org/10.1145/1961189.1961199" target="_blank">https://doi.org/10.1145/1961189.1961199</a>, 2011.
</mixed-citation></ref-html>
<ref-html id="bib1.bib13"><label>13</label><mixed-citation>
Cox, P., Delao, A., Komorniczak, A., and Weller, R.: The California Almanac of Emissions and Air Quality – 2009 edition, Planning and Technical Support Division California Air Resources Board,
2009.
</mixed-citation></ref-html>
<ref-html id="bib1.bib14"><label>14</label><mixed-citation>
Cox, P., Delao, A., and Komorniczak, A.: The California Almanac of Emissions and Air Quality – 2013 edition, Air Quality Planning and Science Division California Air Resources Board,
2013.
</mixed-citation></ref-html>
<ref-html id="bib1.bib15"><label>15</label><mixed-citation>
Drucker, H., Burges, C. J. C., Kaufman, L., Smola, A., and Vapnik, V.:
Support vector regression machines, Proceedings of the 9th International
Conference on Neural Information Processing Systems, Denver, Colorado, 1996.
</mixed-citation></ref-html>
<ref-html id="bib1.bib16"><label>16</label><mixed-citation>
Fan, R.-E., Chen, P.-H., and Lin, C.-J.: Working Set Selection Using Second
Order Information for Training Support Vector Machines, J. Mach. Learn.
Res., 6, 1889–1918, 2005.
</mixed-citation></ref-html>
<ref-html id="bib1.bib17"><label>17</label><mixed-citation>
Fisher, R. A.: On the mathematical foundations of theoretical statistics,
Philos. T. Roy. Soc. Lond. A, 222, 309–368,
<a href="https://doi.org/10.1098/rsta.1922.0009" target="_blank">https://doi.org/10.1098/rsta.1922.0009</a>, 1922.
</mixed-citation></ref-html>
<ref-html id="bib1.bib18"><label>18</label><mixed-citation>
Flynn, M. T., Mattson, E. J., Jaffe, D. A., and Gratz, L. E.: Spatial
patterns in summertime surface ozone in the Southern Front Range of the U.S.
Rocky Mountains, Elementa: Science of the Anthropocene, 9, 00104,
<a href="https://doi.org/10.1525/elementa.2020.00104" target="_blank">https://doi.org/10.1525/elementa.2020.00104</a>, 2021.
</mixed-citation></ref-html>
<ref-html id="bib1.bib19"><label>19</label><mixed-citation>
Friedman, J. H.: Multivariate Adaptive Regression Splines, Ann.
Stat., 19, 1–67, <a href="https://doi.org/10.1214/aos/1176347963" target="_blank">https://doi.org/10.1214/aos/1176347963</a>, 1991.
</mixed-citation></ref-html>
<ref-html id="bib1.bib20"><label>20</label><mixed-citation>
Friedman, J. H. and Silverman, B. W.: Flexible Parsimonious Smoothing and
Additive Modeling, Technometrics, 31, 3–21, <a href="https://doi.org/10.2307/1270359" target="_blank">https://doi.org/10.2307/1270359</a>, 1989.
</mixed-citation></ref-html>
<ref-html id="bib1.bib21"><label>21</label><mixed-citation>
Galton, F.: Co-Relations and Their Measurement, Chiefly from Anthropometric
Data, P. Roy. Soc. Lond., 45, 135–145, 1888.
</mixed-citation></ref-html>
<ref-html id="bib1.bib22"><label>22</label><mixed-citation>
Galton, F. S.: Natural inheritance, Macmillan, London, <a href="https://doi.org/10.5962/bhl.title.32181" target="_blank">https://doi.org/10.5962/bhl.title.32181</a>, 1889.
</mixed-citation></ref-html>
<ref-html id="bib1.bib23"><label>23</label><mixed-citation>
Gao, Z.: Predicting peak daily maximum 8-hour ozone, and linkages to emissions and meteorology, in Southern California using machine learning methods, Zenodo [code], <a href="https://doi.org/10.5281/zenodo.6892066" target="_blank">https://doi.org/10.5281/zenodo.6892066</a>, 2022a.
</mixed-citation></ref-html>
<ref-html id="bib1.bib24"><label>24</label><mixed-citation>
Gao, Z.: Predicting peak daily maximum 8-hour ozone, and linkages to emissions and meteorology, in Southern California using machine learning methods, Zenodo [data set], <a href="https://doi.org/10.5281/zenodo.6892062" target="_blank">https://doi.org/10.5281/zenodo.6892062</a>, 2022b.
</mixed-citation></ref-html>
<ref-html id="bib1.bib25"><label>25</label><mixed-citation>
Gao, Z., Ivey, C. E., Blanchard, C. L., Do, K., Lee, S.-M., and Russell, A.
G.: Separating emissions and meteorological impacts on peak ozone
concentrations in Southern California using generalized additive modeling,
Environ. Pollut., 307, 119503, <a href="https://doi.org/10.1016/j.envpol.2022.119503" target="_blank">https://doi.org/10.1016/j.envpol.2022.119503</a>, 2022.
</mixed-citation></ref-html>
<ref-html id="bib1.bib26"><label>26</label><mixed-citation>
García Nieto, P. J. and Álvarez Antón, J. C.: Nonlinear air
quality modeling using multivariate adaptive regression splines in Gijón
urban area (Northern Spain) at local scale, Appl. Math.
Comput., 235, 50–65, <a href="https://doi.org/10.1016/j.amc.2014.02.096" target="_blank">https://doi.org/10.1016/j.amc.2014.02.096</a>, 2014.
</mixed-citation></ref-html>
<ref-html id="bib1.bib27"><label>27</label><mixed-citation>
Gong, X., Kaulfus, A., Nair, U., and Jaffe, D. A.: Quantifying O3 Impacts in
Urban Areas Due to Wildfires Using a Generalized Additive Model,
Environ. Sci. Technol., 51, 13216–13223,
<a href="https://doi.org/10.1021/acs.est.7b03130" target="_blank">https://doi.org/10.1021/acs.est.7b03130</a>, 2017.
</mixed-citation></ref-html>
<ref-html id="bib1.bib28"><label>28</label><mixed-citation>
Gong, X., Hong, S., and Jaffe, D. A.: Ozone in China: Spatial Distribution
and Leading Meteorological Factors Controlling O<sub>3</sub> in 16 Chinese Cities,
Aerosol Air Qual. Res., 18, 2287–2300, <a href="https://doi.org/10.4209/aaqr.2017.10.0368" target="_blank">https://doi.org/10.4209/aaqr.2017.10.0368</a>,
2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib29"><label>29</label><mixed-citation>
Gorai, A. K., Tuluri, F., Tchounwou, P. B., and Ambinakudige, S.: Influence
of local meteorology and NO<sub>2</sub> conditions on ground-level ozone concentrations
in the eastern part of Texas, USA, Air Quality, Atmos. He., 8,
81–96, <a href="https://doi.org/10.1007/s11869-014-0276-5" target="_blank">https://doi.org/10.1007/s11869-014-0276-5</a>, 2015.
</mixed-citation></ref-html>
<ref-html id="bib1.bib30"><label>30</label><mixed-citation>
Harrell, F. E.: General Aspects of Fitting Regression Models, Springer
International Publishing, 13–44, <a href="https://doi.org/10.1007/978-3-319-19425-7_2" target="_blank">https://doi.org/10.1007/978-3-319-19425-7_2</a>,
2015.
</mixed-citation></ref-html>
<ref-html id="bib1.bib31"><label>31</label><mixed-citation>
Hastie, T.: Generalized additive models, Chapter 7, in: Statistical Models in S, edited by: Chambers, J. M. and Hastie,
T. J., Wadsworth &amp; Brooks/Cole, 1991.
</mixed-citation></ref-html>
<ref-html id="bib1.bib32"><label>32</label><mixed-citation>
Hastie, T. and Tibshirani, R.: Generalized Additive Models, Stat.
Sci., 1, 297–310, <a href="https://doi.org/10.1214/ss/1177013604" target="_blank">https://doi.org/10.1214/ss/1177013604</a>, 1986.
</mixed-citation></ref-html>
<ref-html id="bib1.bib33"><label>33</label><mixed-citation>
Hastie, T. and Tibshirani, R.: Generalized additive models, Chapman &amp;
Hall/CRC, London, 1990.
</mixed-citation></ref-html>
<ref-html id="bib1.bib34"><label>34</label><mixed-citation>
Hastie, T. and Tibshirani, R.: Discriminant Analysis by Gaussian Mixtures,
J. Roy. Stat. Soc. B, 58,
155–176, 1996.
</mixed-citation></ref-html>
<ref-html id="bib1.bib35"><label>35</label><mixed-citation>
Hastie, T., Tibshirani, R., and Friedman, J. H.: The elements of statistical
learning: data mining, inference, and prediction, 2nd edn.,
Springer, New York, 2009.
</mixed-citation></ref-html>
<ref-html id="bib1.bib36"><label>36</label><mixed-citation>
Henneman, L. R. F., Holmes, H. A., Mulholland, J. A., and Russell, A. G.:
Meteorological detrending of primary and secondary pollutant concentrations:
Method application and evaluation using long-term (2000–2012) data in
Atlanta, Atmos. Environ., 119, 201–210, <a href="https://doi.org/10.1016/j.atmosenv.2015.08.007" target="_blank">https://doi.org/10.1016/j.atmosenv.2015.08.007</a>, 2015.
</mixed-citation></ref-html>
<ref-html id="bib1.bib37"><label>37</label><mixed-citation>
Hogrefe, C., Rao, S. T., Zurbenko, I. G., and Porter, P. S.: Interpreting
the Information in Ozone Observations and Model Predictions Relevant to
Regulatory Policies in the Eastern United States, B. Am.
Meteorol. Soc., 81, 2083–2106, <a href="https://doi.org/10.1175/1520-0477(2000)081&lt;2083:itiioo&gt;2.3.co;2" target="_blank">https://doi.org/10.1175/1520-0477(2000)081&lt;2083:itiioo&gt;2.3.co;2</a>, 2000.
</mixed-citation></ref-html>
<ref-html id="bib1.bib38"><label>38</label><mixed-citation>
Hong, C., Mueller, N. D., Burney, J. A., Zhang, Y., AghaKouchak, A., Moore,
F. C., Qin, Y., Tong, D., and Davis, S. J.: Impacts of ozone and climate
change on yields of perennial crops in California, Nature Food, 1, 166–172,
<a href="https://doi.org/10.1038/s43016-020-0043-8" target="_blank">https://doi.org/10.1038/s43016-020-0043-8</a>, 2020.
</mixed-citation></ref-html>
<ref-html id="bib1.bib39"><label>39</label><mixed-citation>
Hu, C., Kang, P., Jaffe, D. A., Li, C., Zhang, X., Wu, K., and Zhou, M.:
Understanding the impact of meteorology on ozone in 334 cities of China,
Atmos. Environ., 248, 118221, <a href="https://doi.org/10.1016/j.atmosenv.2021.118221" target="_blank">https://doi.org/10.1016/j.atmosenv.2021.118221</a>, 2021.
</mixed-citation></ref-html>
<ref-html id="bib1.bib40"><label>40</label><mixed-citation>
Huang, X. G., Shao, T. J., Zhao, J. B., Cao, J. J., and Lü, X. H.:
Influencing Factors of Ozone Concentration in Xi'an Based on Generalized
Additive Models, Huan Jing Ke Xue, 41, 1535–1543,
<a href="https://doi.org/10.13227/j.hjkx.201906067" target="_blank">https://doi.org/10.13227/j.hjkx.201906067</a>, 2020.
</mixed-citation></ref-html>
<ref-html id="bib1.bib41"><label>41</label><mixed-citation>
Jeong, Y., Lee, H. W., and Jeon, W.: Regional Differences of Primary
Meteorological Factors Impacting O<sub>3</sub> Variability in South Korea, Atmosphere,
11, 74, <a href="https://doi.org/10.3390/atmos11010074" target="_blank">https://doi.org/10.3390/atmos11010074</a>, 2020.
</mixed-citation></ref-html>
<ref-html id="bib1.bib42"><label>42</label><mixed-citation>
Jin, L., Loisy, A., and Brown, N. J.: Role of meteorological processes in
ozone responses to emission controls in California's San Joaquin Valley,
J. Geophys. Res.-Atmos., 118, 8010–8022,
<a href="https://doi.org/10.1002/jgrd.50559" target="_blank">https://doi.org/10.1002/jgrd.50559</a>, 2013.
</mixed-citation></ref-html>
<ref-html id="bib1.bib43"><label>43</label><mixed-citation>
Keller, C. A. and Evans, M. J.: Application of random forest regression to the calculation of gas-phase chemistry within the GEOS-Chem chemistry model v10, Geosci. Model Dev., 12, 1209–1225, <a href="https://doi.org/10.5194/gmd-12-1209-2019" target="_blank">https://doi.org/10.5194/gmd-12-1209-2019</a>, 2019.
</mixed-citation></ref-html>
<ref-html id="bib1.bib44"><label>44</label><mixed-citation>
Kelley, M. C., Brown, M. M., Fedler, C. B., and Ardon-Dryer, K.: Long Term
Measurements of PM<sub>2.5</sub> Concentrations in Lubbock, Texas, Aerosol Air
Qual. Res., 20, 1306–1318, <a href="https://doi.org/10.4209/aaqr.2019.09.0469" target="_blank">https://doi.org/10.4209/aaqr.2019.09.0469</a>, 2020.
</mixed-citation></ref-html>
<ref-html id="bib1.bib45"><label>45</label><mixed-citation>
Kleeman, M. J.: A preliminary assessment of the sensitivity of air quality
in California to global change, Clim. Change, 87, 273–292,
<a href="https://doi.org/10.1007/s10584-007-9351-3" target="_blank">https://doi.org/10.1007/s10584-007-9351-3</a>, 2008.
</mixed-citation></ref-html>
<ref-html id="bib1.bib46"><label>46</label><mixed-citation>
Lawrence, M. G.: The Relationship between Relative Humidity and the Dewpoint
Temperature in Moist Air: A Simple Conversion and Applications, B.
Am. Meteorol. Soc., 86, 225–234, <a href="https://doi.org/10.1175/BAMS-86-2-225" target="_blank">https://doi.org/10.1175/BAMS-86-2-225</a>,
2005.
</mixed-citation></ref-html>
<ref-html id="bib1.bib47"><label>47</label><mixed-citation>
Leathwick, J. R., Elith, J., and Hastie, T.: Comparative performance of
generalized additive models and multivariate adaptive regression splines for
statistical modelling of species distributions, Ecol. Model., 199,
188–196, <a href="https://doi.org/10.1016/j.ecolmodel.2006.05.022" target="_blank">https://doi.org/10.1016/j.ecolmodel.2006.05.022</a>, 2006.
</mixed-citation></ref-html>
<ref-html id="bib1.bib48"><label>48</label><mixed-citation>
Liaw, A. and Wiener, M.: Classification and Regression by randomForest, R
News, 2, 18–22, 2002.
</mixed-citation></ref-html>
<ref-html id="bib1.bib49"><label>49</label><mixed-citation>
Ling, Z. H., Guo, H., Zheng, J. Y., Louie, P. K. K., Cheng, H. R., Jiang,
F., Cheung, K., Wong, L. C., and Feng, X. Q.: Establishing a conceptual
model for photochemical ozone pollution in subtropical Hong Kong,
Atmos. Environ., 76, 208–220, <a href="https://doi.org/10.1016/j.atmosenv.2012.09.051" target="_blank">https://doi.org/10.1016/j.atmosenv.2012.09.051</a>, 2013.
</mixed-citation></ref-html>
<ref-html id="bib1.bib50"><label>50</label><mixed-citation>
Liu, B.-C., Binaykia, A., Chang, P.-C., Tiwari, M. K., and Tsao, C.-C.:
Urban air quality forecasting based on multi-dimensional collaborative
Support Vector Regression (SVR): A case study of
Beijing-Tianjin-Shijiazhuang, PLOS ONE, 12, e0179763,
<a href="https://doi.org/10.1371/journal.pone.0179763" target="_blank">https://doi.org/10.1371/journal.pone.0179763</a>, 2017.
</mixed-citation></ref-html>
<ref-html id="bib1.bib51"><label>51</label><mixed-citation>
Liu, T., Li, T. T., Zhang, Y. H., Xu, Y. J., Lao, X. Q., Rutherford, S.,
Chu, C., Luo, Y., Zhu, Q., Xu, X. J., Xie, H. Y., Liu, Z. R., and Ma, W. J.:
The short-term effect of ambient ozone on mortality is modified by
temperature in Guangzhou, China, Atmos. Environ., 76, 59–67,
<a href="https://doi.org/10.1016/j.atmosenv.2012.07.011" target="_blank">https://doi.org/10.1016/j.atmosenv.2012.07.011</a>, 2013.
</mixed-citation></ref-html>
<ref-html id="bib1.bib52"><label>52</label><mixed-citation>
Lu, R. and Turco, R. P.: Ozone distributions over the los angeles basin:
Three-dimensional simulations with the smog model, Atmos. Environ.,
30, 4155–4176, <a href="https://doi.org/10.1016/1352-2310(96)00153-7" target="_blank">https://doi.org/10.1016/1352-2310(96)00153-7</a>, 1996.
</mixed-citation></ref-html>
<ref-html id="bib1.bib53"><label>53</label><mixed-citation>
Lu, X., Zhang, L., and Shen, L.: Meteorology and Climate Influences on
Tropospheric Ozone: a Review of Natural Sources, Chemistry, and Transport
Patterns, Current Pollut. Rep., 5, 238–260, <a href="https://doi.org/10.1007/s40726-019-00118-3" target="_blank">https://doi.org/10.1007/s40726-019-00118-3</a>, 2019.
</mixed-citation></ref-html>
<ref-html id="bib1.bib54"><label>54</label><mixed-citation>
Luna, A. S., Paredes, M. L. L., De Oliveira, G. C. G., and Corrêa, S.
M.: Prediction of ozone concentration in tropospheric levels using
artificial neural networks and support vector machine at Rio de Janeiro,
Brazil, Atmos. Environ., 98, 98–104, <a href="https://doi.org/10.1016/j.atmosenv.2014.08.060" target="_blank">https://doi.org/10.1016/j.atmosenv.2014.08.060</a>,
2014.
</mixed-citation></ref-html>
<ref-html id="bib1.bib55"><label>55</label><mixed-citation>
Ma, Y., Ma, B., Jiao, H., Zhang, Y., Xin, J., and Yu, Z.: An analysis of the
effects of weather and air pollution on tropospheric ozone using a
generalized additive model in Western China: Lanzhou, Gansu, Atmos.
Environ., 224, 117342, <a href="https://doi.org/10.1016/j.atmosenv.2020.117342" target="_blank">https://doi.org/10.1016/j.atmosenv.2020.117342</a>, 2020.
</mixed-citation></ref-html>
<ref-html id="bib1.bib56"><label>56</label><mixed-citation>
Mahmud, A., Hixson, M., Hu, J., Zhao, Z., Chen, S.-H., and Kleeman, M. J.: Climate impact on airborne particulate matter concentrations in California using seven year analysis periods, Atmos. Chem. Phys., 10, 11097–11114, <a href="https://doi.org/10.5194/acp-10-11097-2010" target="_blank">https://doi.org/10.5194/acp-10-11097-2010</a>, 2010.
</mixed-citation></ref-html>
<ref-html id="bib1.bib57"><label>57</label><mixed-citation>
McClure, C. D. and Jaffe, D. A.: Investigation of high ozone events due to
wildfire smoke in an urban area, Atmos. Environ., 194, 146–157,
<a href="https://doi.org/10.1016/j.atmosenv.2018.09.021" target="_blank">https://doi.org/10.1016/j.atmosenv.2018.09.021</a>, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib58"><label>58</label><mixed-citation>
McGlynn, D., Mao, H., Sive, B., and Sharac, T.: Understanding Long-Term
Variations in Surface Ozone in United States (U.S.) National Parks,
Atmosphere, 9, 125, <a href="https://doi.org/10.3390/atmos9040125" target="_blank">https://doi.org/10.3390/atmos9040125</a>, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib59"><label>59</label><mixed-citation>
Menne, M. J., Durre, I., Korzeniewski, B., McNeal, S., Thomas, K., Yin, X.,
Anthony, S., Ray, R., Vose, R. S., Gleason, B. E., and Houston, T. G.:
Global Historical Climatology Network - Daily (GHCN-Daily), Version 3,
<a href="https://doi.org/10.7289/V5D21VHZ" target="_blank">https://doi.org/10.7289/V5D21VHZ</a>, 2012a.
</mixed-citation></ref-html>
<ref-html id="bib1.bib60"><label>60</label><mixed-citation>
Menne, M. J., Durre, I., Vose, R. S., Gleason, B. E., and Houston, T. G.: An
Overview of the Global Historical Climatology Network-Daily Database,
J. Atmos. Ocean. Tech., 29, 897–910,
<a href="https://doi.org/10.1175/jtech-d-11-00103.1" target="_blank">https://doi.org/10.1175/jtech-d-11-00103.1</a>, 2012b.
</mixed-citation></ref-html>
<ref-html id="bib1.bib61"><label>61</label><mixed-citation>
Milborrow, S.: Regression Splines,
FIM Forecast Model: 500&thinsp;mb Wind Speed and 500&thinsp;mb Height Contours – Real-time:
<a href="https://CRAN.R-project.org/package=earth" target="_blank"/> (last access: 13 November 2021),
R program [code], 2021.
</mixed-citation></ref-html>
<ref-html id="bib1.bib62"><label>62</label><mixed-citation>
NOAA: FIM Forecast Model: 500&thinsp;mb Wind Speed and 500&thinsp;mb Height Contours – Real-time, <a href="https://sos.noaa.gov/datasets/fim-forecast-model-500mb-wind-speed-and-500mb-height-contours-real-time/" target="_blank">https://sos.noaa.gov/datasets/fim-forecast-model-500mb-wind-speed-and-500mb-height-contours-real-time/</a> (last access: 23 May 2021), 2020.
</mixed-citation></ref-html>
<ref-html id="bib1.bib63"><label>63</label><mixed-citation>
Oduro, S. D., Metia, S., Duc, H., Hong, G., and Ha, Q. P.: Multivariate
adaptive regression splines models for vehicular emission prediction,
Visualization in Engineering, 3, 13, <a href="https://doi.org/10.1186/s40327-015-0024-4" target="_blank">https://doi.org/10.1186/s40327-015-0024-4</a>, 2015.
</mixed-citation></ref-html>
<ref-html id="bib1.bib64"><label>64</label><mixed-citation>
Oman, L. D., Ziemke, J. R., Douglass, A. R., Waugh, D. W., Lang, C.,
Rodriguez, J. M., and Nielsen, J. E.: The response of tropical tropospheric
ozone to ENSO, Geophys. Res. Lett., 38, 13706, <a href="https://doi.org/10.1029/2011GL047865" target="_blank">https://doi.org/10.1029/2011GL047865</a>, 2011.
</mixed-citation></ref-html>
<ref-html id="bib1.bib65"><label>65</label><mixed-citation>
Oman, L. D., Douglass, A. R., Ziemke, J. R., Rodriguez, J. M., Waugh, D. W.,
and Nielsen, J. E.: The ozone response to ENSO in Aura satellite
measurements and a chemistry-climate simulation, J. Geophys. Res.-Atmos., 118, 965–976,
<a href="https://doi.org/10.1029/2012jd018546" target="_blank">https://doi.org/10.1029/2012jd018546</a>, 2013.
</mixed-citation></ref-html>
<ref-html id="bib1.bib66"><label>66</label><mixed-citation>
Pearce, J. L., Beringer, J., Nicholls, N., Hyndman, R. J., and Tapper, N.
J.: Quantifying the influence of local meteorology on air quality using
generalized additive models, Atmos. Environ., 45, 1328–1336,
<a href="https://doi.org/10.1016/j.atmosenv.2010.11.051" target="_blank">https://doi.org/10.1016/j.atmosenv.2010.11.051</a>, 2011.
</mixed-citation></ref-html>
<ref-html id="bib1.bib67"><label>67</label><mixed-citation>
Pernak, R., Alvarado, M., Lonsdale, C., Mountain, M., Hegarty, J., and
Nehrkorn, T.: Forecasting Surface O<sub>3</sub> in Texas Urban Areas Using Random
Forest and Generalized Additive Models, Aerosol Air Qual. Res., 9,
2815–2826, <a href="https://doi.org/10.4209/aaqr.2018.12.0464" target="_blank">https://doi.org/10.4209/aaqr.2018.12.0464</a>, 2019.
</mixed-citation></ref-html>
<ref-html id="bib1.bib68"><label>68</label><mixed-citation>
Pope, P. T. and Webster, J. T.: The Use of an F-Statistic in Stepwise
Regression Procedures, Technometrics, 14, 327–340, 1972.
</mixed-citation></ref-html>
<ref-html id="bib1.bib69"><label>69</label><mixed-citation>
Porter, T. M.: Social Interests and Statistical Theory: Statistics in
Britain, 1865–1930, Science, 214, 784–784,
<a href="https://doi.org/10.1126/science.214.4522.784.a" target="_blank">https://doi.org/10.1126/science.214.4522.784.a</a>, 1981.
</mixed-citation></ref-html>
<ref-html id="bib1.bib70"><label>70</label><mixed-citation>
Porter, T. M.: Trust in Numbers, Princeton University Press, ISBN 9780691208411, 1995.
</mixed-citation></ref-html>
<ref-html id="bib1.bib71"><label>71</label><mixed-citation>
Rao, S. T., Zurbenko, I. G., Neagu, R., Porter, P. S., Ku, J. Y., and Henry,
R. F.: Space and Time Scales in Ambient Ozone Data, B. Am.
Meteorol. Soc., 78, 2153–2166, <a href="https://doi.org/10.1175/1520-0477(1997)078&lt;2153:satsia&gt;2.0.co;2" target="_blank">https://doi.org/10.1175/1520-0477(1997)078&lt;2153:satsia&gt;2.0.co;2</a>, 1997.
</mixed-citation></ref-html>
<ref-html id="bib1.bib72"><label>72</label><mixed-citation>
Rodríguez-Pérez, R., Vogt, M., and Bajorath, J.: Support Vector
Machine Classification and Regression Prioritize Different Structural
Features for Binary Compound Activity and Potency Value Prediction, ACS
Omega, 2, 6371–6379, <a href="https://doi.org/10.1021/acsomega.7b01079" target="_blank">https://doi.org/10.1021/acsomega.7b01079</a>, 2017.
</mixed-citation></ref-html>
<ref-html id="bib1.bib73"><label>73</label><mixed-citation>
Roy, S. S., Pratyush, C., and Barna, C.: Predicting Ozone Layer
Concentration Using Multivariate Adaptive Regression Splines, Random Forest
and Classification and Regression Tree, Springer International
Publishing, 140–152, <a href="https://doi.org/10.1007/978-3-319-62524-9_11" target="_blank">https://doi.org/10.1007/978-3-319-62524-9_11</a>, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib74"><label>74</label><mixed-citation>
Rybarczyk, Y. and Zalakeviciute, R.: Machine Learning Approaches for Outdoor
Air Quality Modelling: A Systematic Review, Appl. Sci., 8, 2570,
<a href="https://doi.org/10.3390/app8122570" target="_blank">https://doi.org/10.3390/app8122570</a>, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib75"><label>75</label><mixed-citation>
Schölkopf, B. and Smola, A. J.: Learning with Kernels: Support Vector
Machines, Regularization, Optimization, and Beyond, MIT Press, ISBN 9780262536578, 2001.
</mixed-citation></ref-html>
<ref-html id="bib1.bib76"><label>76</label><mixed-citation>
Seinfeld, J. H. and Pandis, S. N.: Atmospheric Chemistry and Physics: From
Air Pollution to Climate Change, Wiley, ISBN 978-1-118-94740-1, 2016.
</mixed-citation></ref-html>
<ref-html id="bib1.bib77"><label>77</label><mixed-citation>
Smola, A. J. and Schölkopf, B.: A tutorial on support vector regression,
Stat. Comput., 14, 199–222, <a href="https://doi.org/10.1023/b:stco.0000035301.49549.88" target="_blank">https://doi.org/10.1023/b:stco.0000035301.49549.88</a>,
2004.
</mixed-citation></ref-html>
<ref-html id="bib1.bib78"><label>78</label><mixed-citation>
Sotomayor-Olmedo, A., Aceves-Fernández, M. A., Gorrostieta-Hurtado, E.,
Pedraza-Ortega, C., Ramos-Arreguín, J. M., and Vargas-Soto, J. E.:
Forecast Urban Air Pollution in Mexico City by Using Support Vector
Machines: A Kernel Performance Approach, International Journal of
Intelligence Science, 03, 126–135, <a href="https://doi.org/10.4236/ijis.2013.33014" target="_blank">https://doi.org/10.4236/ijis.2013.33014</a>, 2013.
</mixed-citation></ref-html>
<ref-html id="bib1.bib79"><label>79</label><mixed-citation>
Stafoggia, M., Johansson, C., Glantz, P., Renzi, M., Shtein, A., De Hoogh,
K., Kloog, I., Davoli, M., Michelozzi, P., and Bellander, T.: A Random
Forest Approach to Estimate Daily Particulate Matter, Nitrogen Dioxide, and
Ozone at Fine Spatial Resolution in Sweden, Atmosphere, 11, 239,
<a href="https://doi.org/10.3390/atmos11030239" target="_blank">https://doi.org/10.3390/atmos11030239</a>, 2020.
</mixed-citation></ref-html>
<ref-html id="bib1.bib80"><label>80</label><mixed-citation>
Stephen, M. S.: Gauss and the Invention of Least Squares, Ann.
Stat., 9, 465–474, <a href="https://doi.org/10.1214/aos/1176345451" target="_blank">https://doi.org/10.1214/aos/1176345451</a>, 1981.
</mixed-citation></ref-html>
<ref-html id="bib1.bib81"><label>81</label><mixed-citation>
Sun, L., Xue, L., Wang, Y., Li, L., Lin, J., Ni, R., Yan, Y., Chen, L., Li, J., Zhang, Q., and Wang, W.: Impacts of meteorology and emissions on summertime surface ozone increases over central eastern China between 2003 and 2015, Atmos. Chem. Phys., 19, 1455–1469, <a href="https://doi.org/10.5194/acp-19-1455-2019" target="_blank">https://doi.org/10.5194/acp-19-1455-2019</a>, 2019.
</mixed-citation></ref-html>
<ref-html id="bib1.bib82"><label>82</label><mixed-citation>
Tin Kam, H.: Random decision forests, Proceedings of 3rd International
Conference on Document Analysis and Recognition, 14–16 August 1995,
271, 278–282, <a href="https://doi.org/10.1109/ICDAR.1995.598994" target="_blank">https://doi.org/10.1109/ICDAR.1995.598994</a>, 1995.
</mixed-citation></ref-html>
<ref-html id="bib1.bib83"><label>83</label><mixed-citation>
U.S. EPA:
Trends in Ozone Adjusted for Weather Conditions, <a href="https://www.epa.gov/air-trends/trends-ozone-adjusted-weather-conditions" target="_blank">https://www.epa.gov/air-trends/trends-ozone-adjusted-weather-conditions</a> (last access: 13 November 2021),
U.S. EPA, 2016.

</mixed-citation></ref-html>
<ref-html id="bib1.bib84"><label>84</label><mixed-citation>
U.S. EPA: Integrated Science Assessment (ISA) for Ozone and Related Photochemical Oxidants (Final Report, Apr 2020), U.S. Environmental Protection Agency, Washington, D.C., EPA/600/R-20/012, 2020.
</mixed-citation></ref-html>
<ref-html id="bib1.bib85"><label>85</label><mixed-citation>
Vong, C.-M., Ip, W.-F., Wong, P.-K., and Yang, J.-Y.: Short-Term Prediction
of Air Pollution in Macau Using Support Vector Machines, J. Control
Sci. Eng., 2012, 1–11, <a href="https://doi.org/10.1155/2012/518032" target="_blank">https://doi.org/10.1155/2012/518032</a>, 2012.
</mixed-citation></ref-html>
<ref-html id="bib1.bib86"><label>86</label><mixed-citation>
Wells, B., Dolwick, P., Eder, B., Evangelista, M., Foley, K., Mannshardt,
E., Misenis, C., and Weishampel, A.: Improved estimation of trends in U.S.
ozone concentrations adjusted for interannual variability in meteorological
conditions, Atmos. Environ., 248, 118234,
<a href="https://doi.org/10.1016/j.atmosenv.2021.118234" target="_blank">https://doi.org/10.1016/j.atmosenv.2021.118234</a>, 2021.
</mixed-citation></ref-html>
<ref-html id="bib1.bib87"><label>87</label><mixed-citation>
Wikipedia Contributors: Multivariate adaptive regression spline, <a href="https://en.wikipedia.org/w/index.php?title=Multivariate_adaptive_regression_spline&amp;oldid=1083057440" target="_blank"/>, last access: 19 April 2022.
</mixed-citation></ref-html>
<ref-html id="bib1.bib88"><label>88</label><mixed-citation>
Wood, S. N.: Fast stable restricted maximum likelihood and marginal
likelihood estimation of semiparametric generalized linear models, J. Roy. Statist. Soc. B, 73,
3–36, <a href="https://doi.org/10.1111/j.1467-9868.2010.00749.x" target="_blank">https://doi.org/10.1111/j.1467-9868.2010.00749.x</a>, 2011.
</mixed-citation></ref-html>
<ref-html id="bib1.bib89"><label>89</label><mixed-citation>
Wood, S. N.: Generalized Additive Models: An Introduction with R, 2nd edn.,
Chapman and Hall/CRC,  ISBN 9781315370279, <a href="https://doi.org/10.1201/9781315370279" target="_blank">https://doi.org/10.1201/9781315370279</a>, 2017.
</mixed-citation></ref-html>
<ref-html id="bib1.bib90"><label>90</label><mixed-citation>
Xu, L., Yu, J.-Y., Schnell, J. L., and Prather, M. J.: The Seasonality and
Geographic Dependence of ENSO Impacts on US Surface Ozone Variability, Geophys. Res. Lett.,
44, 3420–3428,
<a href="https://doi.org/10.1002/2017gl073044" target="_blank">https://doi.org/10.1002/2017gl073044</a>, 2017.
</mixed-citation></ref-html>
<ref-html id="bib1.bib91"><label>91</label><mixed-citation>
Zhan, Y., Luo, Y., Deng, X., Grieneisen, M. L., Zhang, M., and Di, B.:
Spatiotemporal prediction of daily ambient ozone levels across China using
random forest for human exposure assessment, Environ. Pollut., 233,
464–473, <a href="https://doi.org/10.1016/j.envpol.2017.10.029" target="_blank">https://doi.org/10.1016/j.envpol.2017.10.029</a>, 2018.
</mixed-citation></ref-html>--></article>
