<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">IJG</journal-id><journal-title-group><journal-title>International Journal of Geosciences</journal-title></journal-title-group><issn pub-type="epub">2156-8359</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/ijg.2015.610090</article-id><article-id pub-id-type="publisher-id">IJG-60721</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Earth&amp;Environmental Sciences</subject></subj-group></article-categories><title-group><article-title>
 
 
  Correspondence Analysis on a Space-Time Data Set for Multiple Environmental Variables
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>alma</surname><given-names>Monica</given-names></name><xref ref-type="aff" rid="aff1"><sub>1</sub></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib></contrib-group><aff id="aff1"><label>1</label><addr-line>Universitá del Salento, Lecce, Italy</addr-line></aff><author-notes><corresp id="cor1">* E-mail:<email>monica.palma@unisalento.it</email></corresp></author-notes><pub-date pub-type="epub"><day>20</day><month>10</month><year>2015</year></pub-date><volume>06</volume><issue>10</issue><fpage>1154</fpage><lpage>1165</lpage><history><date date-type="received"><day>3</day>	<month>August</month>	<year>2015</year></date><date date-type="rev-recd"><day>accepted</day>	<month>26</month>	<year>October</year>	</date><date date-type="accepted"><day>29</day>	<month>October</month>	<year>2015</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
  Applications of the multivariate technique called correspondence analysis for environmental studies are relatively new and are limited to spatial multivariate data set. In this paper, a procedure of applying correspondence analysis to a large space-time data set for multiple environmental variables is shown. In particular, nitrogen dioxide and carbon monoxide hourly concentrations measured during January 1999 at several monitored stations in a district of Northern Italy are analyzed. The procedure consists in transforming the continuous variables into categorical ones by the means of appropriate indicator variables, generating special contingency tables and applying correspondence analysis. The use of this classical multivariate technique allows the identification of important relationships among pollution levels and monitoring stations and/or relationships among pollution levels and observation times.
 
</p></abstract><kwd-group><kwd>Space-Time Data</kwd><kwd> Indicator Transform</kwd><kwd> Correspondence Analysis</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>Usually environmental monitoring networks collect a huge amount of data such as pollutant concentrations, atmospheric variates, weather conditions, and so on, which are of particular interest for public policies oriented to environmental and human health protection.</p><p>Such data sets may have the following features:</p><p>• they are multivariate, as several variables are simultaneously measured;</p><p>• they present a spatio-temporal structure, since the measurements are taken in several point of the study area and for a certain period of time.</p><p>Classical multivariate techniques represent useful tools for analyzing multiple va-riables. Their main goal is to obtain a summary description of the data: Principal Component Analysis (PCA) finds a smaller number of variates representing all those collected, without loss of essential information; Correspondence Analysis (CA) studies the association between two or more categorical variables by representing the categories of the variables as points in a low-dimensional space; Canonical Correlation Analysis (CCA) describes the relationships between two groups of several variables. Classical multivariate techniques can be also applied to space-time data sets in order to summarize the spatial and temporal profiles which characterize the information, finding relationships among the data. In De Iaco et al. [<xref ref-type="bibr" rid="scirp.60721-ref1">1</xref>] , the use of PCA allowed summarizing a very large data set of space-time observations for three contaminants. The authors identified a single measure of total air pollution which synthesized the original data without loss of information. Moreover, lately a space-time data set for air pollution and atmospheric variables has been analyzed through CCA in De Iaco [<xref ref-type="bibr" rid="scirp.60721-ref2">2</xref>] . The author emphasized the features of that multiva-riate technique which allowed describing very important relationships between three contaminants (nitrix oxide, nitrogen dioxide and ozone) and atmospheric indicators (humidity, temperature and wind speed).</p><p>Hence, when multiple variables are measured at several locations of the area under study and for a period of time, in other words, when a space-time multivariate data set is available, and the aim is studying the simul- taneous behaviour of the va-riables in order to understand the relationships among the space-time observations, a multivariate technique is the most useful tool. CA is one of the multivariate techniques with a wide range of applications in several fields such as social and political sciences, marketing research, economy, ecology and biology. This technique is usually applied as an exploratory method, with the aim to describe the structure of the data under study with minimal constraints on the form of the same structure [<xref ref-type="bibr" rid="scirp.60721-ref3">3</xref>] .</p><p>In this paper, it will be shown that even CA can be applied to a space-time multivariate data set, finding very important results which other techniques may not highlight. In particular, in this paper CA will be applied to an air pollution data set involving two contaminants measured at monitoring stations in northern Italy during January 1999. The analysis will identify relationships in space among pollution levels and monitoring stations and relationships in time among pollution levels and observation times.</p><p>After a presentation of CA (Section 2) and a review of its theory (Section 2.1), the description of compu- tational aspects follows (Section 2.2). Then, the data set (Section 3) and the most important results from the applied CA and their interpretation are given (Section 4).</p></sec><sec id="s2"><title>2. Correspondence Analysis</title><p>CA is an algebraic technique analogous to PCA, but, while PCA is used for tables of continuous measurements, CA is more appropriate for categorical variates. Hence, CA is suitable for analyzing qualitative information represented by a contingency table. Lebart et al. [<xref ref-type="bibr" rid="scirp.60721-ref4">4</xref>] suggest that CA is useful for the analysis of large data matrices, particularly when there is little auxiliary information concerning the data. The original development of the method was driven by the need to analyze occurrence frequencies in a contingency table [<xref ref-type="bibr" rid="scirp.60721-ref5">5</xref>] . This technique can be viewed as finding the best simultaneous representation of two data sets that comprise the rows and columns of a data matrix with non-negative entries [<xref ref-type="bibr" rid="scirp.60721-ref4">4</xref>] .</p><p>For a long time CA has been applied by European statistical community for psycho-metric and economic studies. This technique has been very popular in France, mainly owing to the efforts of Jean-Paul Benz&#233;cri [<xref ref-type="bibr" rid="scirp.60721-ref5">5</xref>] ; it came to occupy such a strong position in the analysis methodology that it almost became synonymous with data analysis. In the 80’s this technique started to be used in English-speaking countries since some books and papers presented the method in relatively simple form [<xref ref-type="bibr" rid="scirp.60721-ref4">4</xref>] [<xref ref-type="bibr" rid="scirp.60721-ref6">6</xref>] and now several American statistical package, such as SAS [<xref ref-type="bibr" rid="scirp.60721-ref7">7</xref>] and SPSS [<xref ref-type="bibr" rid="scirp.60721-ref8">8</xref>] , include procedures to perform correspondence analysis. In the geostatistical context, applications of CA are relatively new. Avila et al. in [<xref ref-type="bibr" rid="scirp.60721-ref9">9</xref>] and [<xref ref-type="bibr" rid="scirp.60721-ref10">10</xref>] analyzed a data set consisting of the concen- trations of chemical elements measured in a lake; Dutot et al. [<xref ref-type="bibr" rid="scirp.60721-ref11">11</xref>] applied this method to an aerosol collected in a simple atmospheric environment; Jim&#233;nez-Espinosa et al. [<xref ref-type="bibr" rid="scirp.60721-ref12">12</xref>] used CA on 602 soil samples taken in a region of NW Spain to identify geochemical patterns and anomalies.</p><p>All CA applications for environmental studies are limited to spatial multivariate data sets, where observations for several variables are spatially located [<xref ref-type="bibr" rid="scirp.60721-ref13">13</xref>] . Actually, most, if not all, environmental data are collected in space and time and exhaustive time series are often available for several monitored stations inside the area of interest. One of the major goal for an environmental quality control system is to obtain summary information about pollution conditions [<xref ref-type="bibr" rid="scirp.60721-ref14">14</xref>] [<xref ref-type="bibr" rid="scirp.60721-ref15">15</xref>] . Knowing the area inside the monitored region and/or interval of time within the observed period which need of closer controls because of frequent exceeding fixed pollution levels, is definitely a very important issue. CA allows achieving this goal simultaneously for several contaminants. Therefore, it is useful to develop a procedure of applying CA to space-time multivariate data sets.</p><sec id="s2_1"><title>2.1. The Method</title><p>The theory of CA is discussed in several books, [<xref ref-type="bibr" rid="scirp.60721-ref4">4</xref>] - [<xref ref-type="bibr" rid="scirp.60721-ref6">6</xref>] , so only the main features of the method are reviewed here.</p><p>From an initial data matrix <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/5-2801061x5.png" xlink:type="simple"/></inline-formula> with non-negative entries<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/5-2801061x6.png" xlink:type="simple"/></inline-formula>, CA determines the best simultaneous geometrical representation of rows and columns in a low-dimensional space (usually in a two- dimensional space).</p><p>Let <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/5-2801061x7.png" xlink:type="simple"/></inline-formula> be the relative frequency matrix, whose entries are:</p><disp-formula id="scirp.60721-formula1477"><label>(1)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/5-2801061x8.png"  xlink:type="simple"/></disp-formula><p>Two different matrices are used to re-scale<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/5-2801061x9.png" xlink:type="simple"/></inline-formula>, these are:</p><disp-formula id="scirp.60721-formula1478"><label>(2)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/5-2801061x10.png"  xlink:type="simple"/></disp-formula><p>and</p><disp-formula id="scirp.60721-formula1479"><label>(3)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/5-2801061x11.png"  xlink:type="simple"/></disp-formula><p>where:</p><disp-formula id="scirp.60721-formula1480"><label>(4)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/5-2801061x12.png"  xlink:type="simple"/></disp-formula><p>and</p><disp-formula id="scirp.60721-formula1481"><label>(5)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/5-2801061x13.png"  xlink:type="simple"/></disp-formula><p>CA consists in finding a vector u, in a p-dimensional space, which maximizes</p><disp-formula id="scirp.60721-formula1482"><label>(6)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/5-2801061x14.png"  xlink:type="simple"/></disp-formula><p>subject to the constraint:</p><disp-formula id="scirp.60721-formula1483"><label>(7)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/5-2801061x15.png"  xlink:type="simple"/></disp-formula><p>It is known that this is equivalent to finding the vector v, in an l-dimensional space, which maximizes</p><disp-formula id="scirp.60721-formula1484"><label>(8)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/5-2801061x16.png"  xlink:type="simple"/></disp-formula><p>subject to the constraint:</p><disp-formula id="scirp.60721-formula1485"><label>(9)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/5-2801061x17.png"  xlink:type="simple"/></disp-formula><p>The eigenvectors <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/5-2801061x18.png" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/5-2801061x19.png" xlink:type="simple"/></inline-formula> are related by:</p><disp-formula id="scirp.60721-formula1486"><label>(10)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/5-2801061x20.png"  xlink:type="simple"/></disp-formula><p>where <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/5-2801061x21.png" xlink:type="simple"/></inline-formula> is the same eigenvalue for either maximization problems in (6) and (8).</p><p>This duality formula permits displaying the row and column projections in the same graph (called biplots) and this CA feature has been considered as its advantage with respect to others multivariate techniques.</p><p>Sequentially, the method searches for new solutions orthogonal to the previous ones; in particular, orthogo- nality is considered with respect to the inner product defined by the weighting matrices (2) and (3). There will be <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/5-2801061x22.png" xlink:type="simple"/></inline-formula> non-trivial solutions.</p><p>The factors</p><disp-formula id="scirp.60721-formula1487"><label>(11)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/5-2801061x23.png"  xlink:type="simple"/></disp-formula><p>and</p><disp-formula id="scirp.60721-formula1488"><label>(12)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/5-2801061x24.png"  xlink:type="simple"/></disp-formula><p>define the plane where rows and columns of the data matrix are projected.</p><p>Results from CA consist of graphical representations of the projections of rows and columns of the data matrix onto factorial planes, in order to find and understand underlying relationships [<xref ref-type="bibr" rid="scirp.60721-ref4">4</xref>] . There are also con- venient diagnostics that help in the interpretation of the results; in particular:</p><p>• the percentage of explained variation, which is a measure of fit when a particular factor is retained, so that the cumulative percentage of explained variation</p><disp-formula id="scirp.60721-formula1489"><label>(13)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/5-2801061x25.png"  xlink:type="simple"/></disp-formula><p>represents a global measure of fit when K factors, <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/5-2801061x26.png" xlink:type="simple"/></inline-formula>are retained, each <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/5-2801061x26.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/5-2801061x27.png" xlink:type="simple"/></inline-formula> giving the</p><p>contribution of a particular factor. Note that the terminology is similar to that one used in PCA, but in CA the term variation does not refer to variance in the statistical sense; it is an increasing function of K and it is used to choose the number of factors to be kept;</p><p>• the absolute contributions of the h-th row <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/5-2801061x28.png" xlink:type="simple"/></inline-formula> and the i-th column <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/5-2801061x28.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/5-2801061x29.png" xlink:type="simple"/></inline-formula> to the k-th factor, <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/5-2801061x28.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/5-2801061x29.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/5-2801061x30.png" xlink:type="simple"/></inline-formula>, explain the composition of the retained factor. They are respectively:</p><disp-formula id="scirp.60721-formula1490"><label>(14)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/5-2801061x31.png"  xlink:type="simple"/></disp-formula><p>and</p><disp-formula id="scirp.60721-formula1491"><label>(15)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/5-2801061x32.png"  xlink:type="simple"/></disp-formula><p>• the relative contributions of a retained factor with the h-th row <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/5-2801061x33.png" xlink:type="simple"/></inline-formula> or the i-th column <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/5-2801061x33.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/5-2801061x34.png" xlink:type="simple"/></inline-formula> provide a measure of the row or column variation explained by the factor. They are respectively:</p><disp-formula id="scirp.60721-formula1492"><label>(16)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/5-2801061x35.png"  xlink:type="simple"/></disp-formula><p>and</p><disp-formula id="scirp.60721-formula1493"><label>(17)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/5-2801061x36.png"  xlink:type="simple"/></disp-formula><p>Note that the ACs serve primarily as guides to the interpretation of the dimension defined by the retained factors; whereas the RCs indicate how well a point is described by the retained factors. Usually, a large AC implies a large RC, but not conversely [<xref ref-type="bibr" rid="scirp.60721-ref6">6</xref>] .</p></sec><sec id="s2_2"><title>2.2. Computational Aspects</title><p>The application of CA to a space-time data set for multiple environmental variables is based on special con- tingency matrices generated as follows.</p><p>Let<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/5-2801061x37.png" xlink:type="simple"/></inline-formula>, be the space-time data for R variables measured at a-th location, <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/5-2801061x37.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/5-2801061x38.png" xlink:type="simple"/></inline-formula>and w-th observation time,<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/5-2801061x37.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/5-2801061x38.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/5-2801061x39.png" xlink:type="simple"/></inline-formula>. For semplicity, consider</p><p><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/5-2801061x40.png" xlink:type="simple"/></inline-formula>although the procedure can be used to analyze variables measured at different sets of spatial locations.</p><p>Let <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/5-2801061x41.png" xlink:type="simple"/></inline-formula> be J non-overlapping classes of values defined for each of the R variables under study.</p><p>Through the indicator transform, the belonging of <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/5-2801061x42.png" xlink:type="simple"/></inline-formula> to a certain class of values is described:</p><disp-formula id="scirp.60721-formula1494"><label>(18)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/5-2801061x43.png"  xlink:type="simple"/></disp-formula><p>From the four dimensional matrix (variable, station, time, class of values) obtained after the indicator trans- formation (18), the following two dimensional matrices are generated.</p><p>• Matrix<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/5-2801061x44.png" xlink:type="simple"/></inline-formula>, where</p><disp-formula id="scirp.60721-formula1495"><label>(19)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/5-2801061x45.png"  xlink:type="simple"/></disp-formula><disp-formula id="scirp.60721-formula1496"><graphic  xlink:href="http://html.scirp.org/file/5-2801061x46.png"  xlink:type="simple"/></disp-formula><p>In A, the <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/5-2801061x47.png" xlink:type="simple"/></inline-formula> rows represent all survey stations for each variable and the columns represent the J classes of values, so that the entries (19) indicate the number of times, values belonging to the j-th class, are recorded at the a-th station.</p><p>• Matrix<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/5-2801061x48.png" xlink:type="simple"/></inline-formula>, where</p><disp-formula id="scirp.60721-formula1497"><label>(20)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/5-2801061x49.png"  xlink:type="simple"/></disp-formula><disp-formula id="scirp.60721-formula1498"><graphic  xlink:href="http://html.scirp.org/file/5-2801061x50.png"  xlink:type="simple"/></disp-formula><p>In B, the <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/5-2801061x51.png" xlink:type="simple"/></inline-formula> rows represent the observation times for each variable and the columns represent the J classes of values, so that the entries (20) indicate, for each variable, how many stations in the w-th observation time have values belonging to the j-th class.</p><p>The indicator transform allows the user to categorize continuous variables, synthesizing a large multivariate space-time data set. The above two dimensional matrices relate different classes of values (in the case study pollution levels) to locations (matrix A) or to observation times (matrix B), jointly for the variables (pollutants) under study. Thus, CA applied to each matrix, A and B, will allow describing relationships</p><p>• in space, among pollution levels and monitored stations,</p><p>• in time, among pollution levels and observation times,</p><p>simultaneously for the variables under study.</p><p>CA results will also identify clusters of survey stations and intervals of time which need of closer controls when the contaminants frequently exceed fixed thresholds.</p></sec></sec><sec id="s3"><title>3. The Data Set</title><p>The data set consists of concentration values of two pollutants over a particular period of time and at stations of the monitoring network in Milan district, Lombardy (this is one of the northern Italy regions which suffers a serious air pollution pro-blem). The air quality monitoring network covers a wide area with about 190 stations where the main atmospheric contaminants, such as sulphur dioxide (SO<sub>2</sub>), ozone (O<sub>3</sub>), nitric oxide (NO), nitrogen dioxide (NO<sub>2</sub>), carbon monoxide (CO), and meteorological variates, such as humidity, wind velocity, temperature, solar radiation, are continuously measured.</p><p>In the Milan district, air pollution is mainly caused by traffic and industrial activities. Two pollutants, which are primarily generated by the human activities, considered among the most dangerous ones for the atmosphere and human health and have been analyzed in this paper: NO<sub>2</sub> and CO. Nitrogen dioxide is a secondary pollutant generated by the thermic and photochemical reactions among the primary pollutants; it is caused, mainly in winter, by civil and industrial heating systems and by traffic. Therefore its concentration values are very high in urban areas characte-rized by high population density. Carbon monoxide is a primary pollutant caused by the motor vehicles emissions and its values are very high in areas with heavy traffic and poor ventilation. These characteristics are considered to choose the period of the year to be analyzed: January 1999. Indeed, most of the highest values for both pollutants under study were observed during the first month of the year. The box plot of the hourly averages for each pollutant, measured during January 1999 (<xref ref-type="fig" rid="fig1">Figure 1</xref>), highlights exceeding the so called level of attention for several times during the month.</p><p>The national laws, particularly the Premier’s Decree of the 12th of November, 1992, according to the European settlements, lay down, for each pollutant, a specific threshold called level of attention. When the pollution concentrations exceed this level for a long time and at several monitoring stations, air quality is poor and the situation is considered dangerous for the public health.</p><p>The analysis is limited to stations in the Milan district where data for both contaminants are available at all the desidered time points. In <xref ref-type="fig" rid="fig2">Figure 2</xref>, the 27 selected survey stations are shown. They have been classified, according to the Premier’s Decree of the 20th of May, 1991, in two types:</p><p>• stations C, which are located in areas with heavy traffic and poor ventilation; in these areas the CO plume is more evident;</p><p>• stations B, which are located in areas with high density population, therefore these areas are subject to both NO<sub>2</sub> and CO pollution.</p><fig id="fig1"  position="float"><label><xref ref-type="fig" rid="fig1">Figure 1</xref></label><caption><title> Pollution concentration values for CO and NO<sub>2</sub> during January 1999</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/5-2801061x52.png"/></fig><fig id="fig2"  position="float"><label><xref ref-type="fig" rid="fig2">Figure 2</xref></label><caption><title> Posting map of the selected survey stations in Milan district</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/5-2801061x53.png"/></fig><p>In order to split each spatial-temporal distribution into non-overlapping classes of values, the following thresholds:</p><p>a) 1.6 2.3 3 3.9 5.4 mg/m<sup>3</sup></p><p>b) 52 64 75 90 115 mg/m<sup>3</sup></p><p>corresponding to the 0.17, 0.33, 0.50, 0.67, 0.83 quantiles of the distributions of CO a) and NO<sub>2</sub> b) hourly averages, are considered. Hence, six classes of CO and NO<sub>2</sub> concentrations are defined as follows:</p><disp-formula id="scirp.60721-formula1499"><graphic  xlink:href="http://html.scirp.org/file/5-2801061x54.png"  xlink:type="simple"/></disp-formula><disp-formula id="scirp.60721-formula1500"><graphic  xlink:href="http://html.scirp.org/file/5-2801061x55.png"  xlink:type="simple"/></disp-formula><disp-formula id="scirp.60721-formula1501"><graphic  xlink:href="http://html.scirp.org/file/5-2801061x56.png"  xlink:type="simple"/></disp-formula><disp-formula id="scirp.60721-formula1502"><graphic  xlink:href="http://html.scirp.org/file/5-2801061x57.png"  xlink:type="simple"/></disp-formula><p>Then, through the indicator transform, two dimensional matrices are generated as described in (2.2); so that:</p><p>• A is a <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/5-2801061x58.png" xlink:type="simple"/></inline-formula> matrix, whose entries are:</p><disp-formula id="scirp.60721-formula1503"><label>(21)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/5-2801061x59.png"  xlink:type="simple"/></disp-formula><p><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/5-2801061x60.png" xlink:type="simple"/></inline-formula>;</p><p>• B is a <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/5-2801061x61.png" xlink:type="simple"/></inline-formula> matrix, whose entries are<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/5-2801061x61.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/5-2801061x62.png" xlink:type="simple"/></inline-formula>, as defined in (20), cumulated every 24 hours, that is:</p><disp-formula id="scirp.60721-formula1504"><label>(22)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/5-2801061x63.png"  xlink:type="simple"/></disp-formula><disp-formula id="scirp.60721-formula1505"><graphic  xlink:href="http://html.scirp.org/file/5-2801061x64.png"  xlink:type="simple"/></disp-formula><p>CA is applied to these matrices.</p></sec><sec id="s4"><title>4. Results</title><p>A French package software, SPAD [<xref ref-type="bibr" rid="scirp.60721-ref16">16</xref>] , is used for the data analysis since it performs most of multivariate techniques, giving graphical results and diagnostics, in a very simple and fast manner.</p><p>Even if it is a commercial software, it is a very powerful software for data mining, indeed it can perform many statistical data analysis, as Factorial Analysis, Classification, Segmentation, as well as Textual analysis. Moreover, SPAD has a good graphical tools and is easy to use (user-friendly) [<xref ref-type="bibr" rid="scirp.60721-ref17">17</xref>] .</p><p>CA is applied to matrix A and matrix B, since information from both analysis are useful for the aim of the paper, as it will be shown.</p><p>The results from CA are displayed in a series of tables and graphs. In particular, <xref ref-type="table" rid="table1">Table 1</xref> and <xref ref-type="table" rid="table2">Table 2</xref> show the eigenvalues and the percentages of variation explained by the five non-trivial factors from CA applied to the matrix A and B, respectively.</p><p>On the other hand, <xref ref-type="table" rid="table3">Table 3</xref> and <xref ref-type="table" rid="table4">Table 4</xref> list, for the first two factors from CA applied to the matrix A, stations/pollutants with the highest absolute (<xref ref-type="table" rid="table3">Table 3</xref>) and relative (<xref ref-type="table" rid="table4">Table 4</xref>) contributions. Similarly <xref ref-type="table" rid="table5">Table 5</xref> and <xref ref-type="table" rid="table6">Table 6</xref>, which refer to the diagnostics from CA applied to the matrix B.</p><p><xref ref-type="fig" rid="fig3">Figure 3</xref> and <xref ref-type="fig" rid="fig4">Figure 4</xref> show the projections of rows and columns of each matrix on the respective first factorial plane. Note that in <xref ref-type="fig" rid="fig3">Figure 3</xref>, which refers to matrix A, columns (points labeled<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/5-2801061x65.png" xlink:type="simple"/></inline-formula>) and rows (points labeled with the station code and a symbol related to the pollutants) are displayed together on the same plane. Similarly in <xref ref-type="fig" rid="fig4">Figure 4</xref>, where columns (points<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/5-2801061x65.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/5-2801061x66.png" xlink:type="simple"/></inline-formula>) and rows (observation hours, <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/5-2801061x65.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/5-2801061x66.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/5-2801061x67.png" xlink:type="simple"/></inline-formula>, labeled with different symbols for the pollutants under study) of the matrix B are projected together on the same plane.</p><fig id="fig3"  position="float"><label><xref ref-type="fig" rid="fig3">Figure 3</xref></label><caption><title> Plot of the first two factors from CA applied to the matrix A</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/5-2801061x68.png"/></fig><fig id="fig4"  position="float"><label><xref ref-type="fig" rid="fig4">Figure 4</xref></label><caption><title> Plot of the first two factors from CA applied to the matrix B</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/5-2801061x69.png"/></fig><table-wrap id="table1" ><label><xref ref-type="table" rid="table1">Table 1</xref></label><caption><title> Eigenvalues and percentages of variation explained by the factors from CA applied to the matrix A</title></caption><table><tbody><thead><tr><th align="center" valign="middle" ></th><th align="center" valign="middle" >Eigenvalues</th><th align="center" valign="middle" >Variation Explained (%)</th></tr></thead><tr><td align="center" valign="middle" >Factor 1</td><td align="center" valign="middle" >0.0975</td><td align="center" valign="middle" >59.90</td></tr><tr><td align="center" valign="middle" >Factor 2</td><td align="center" valign="middle" >0.0408</td><td align="center" valign="middle" >25.05</td></tr><tr><td align="center" valign="middle" >Factor 3</td><td align="center" valign="middle" >0.0161</td><td align="center" valign="middle" >9.89</td></tr><tr><td align="center" valign="middle" >Factor 4</td><td align="center" valign="middle" >0.0048</td><td align="center" valign="middle" >2.94</td></tr><tr><td align="center" valign="middle" >Factor 5</td><td align="center" valign="middle" >0.0036</td><td align="center" valign="middle" >2.22</td></tr></tbody></table></table-wrap><table-wrap id="table2" ><label><xref ref-type="table" rid="table2">Table 2</xref></label><caption><title> Eigenvalues and percentages of variation explained by the factors from CA applied to the matrix B</title></caption><table><tbody><thead><tr><th align="center" valign="middle" ></th><th align="center" valign="middle" >Eigenvalues</th><th align="center" valign="middle" >Variation Explained (%)</th></tr></thead><tr><td align="center" valign="middle" >Factor 1</td><td align="center" valign="middle" >0.0951</td><td align="center" valign="middle" >85.26</td></tr><tr><td align="center" valign="middle" >Factor 2</td><td align="center" valign="middle" >0.0115</td><td align="center" valign="middle" >10.29</td></tr><tr><td align="center" valign="middle" >Factor 3</td><td align="center" valign="middle" >0.0030</td><td align="center" valign="middle" >2.69</td></tr><tr><td align="center" valign="middle" >Factor 4</td><td align="center" valign="middle" >0.0013</td><td align="center" valign="middle" >1.21</td></tr><tr><td align="center" valign="middle" >Factor 5</td><td align="center" valign="middle" >0.0006</td><td align="center" valign="middle" >0.56</td></tr></tbody></table></table-wrap><table-wrap id="table3" ><label><xref ref-type="table" rid="table3">Table 3</xref></label><caption><title> Highest absolute contributions to the first two factors from CA applied to the matrix A (percentages in parentheses)</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Factors</th><th align="center" valign="middle" >Stations/Pollutants</th><th align="center" valign="middle" >Classes of Values</th></tr></thead><tr><td align="center" valign="middle" >Factor 1</td><td align="center" valign="middle" >86/CO (11) 15/CO (8) 41/NO<sub>2</sub> (8) 93/CO (5) 45/CO (4) 125/NO<sub>2</sub> (4)</td><td align="center" valign="middle" >c6 (39) c1 (29)</td></tr><tr><td align="center" valign="middle" >Factor 2</td><td align="center" valign="middle" >97/NO<sub>2</sub> (27) 62/CO (9) 97/CO (7) 111/CO (6)</td><td align="center" valign="middle" >c1 (50) c3 (21)</td></tr></tbody></table></table-wrap><table-wrap id="table4" ><label><xref ref-type="table" rid="table4">Table 4</xref></label><caption><title> Highest relative contributions to the first two factors from CA applied to the matrix A (percentages in parentheses)</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Factors</th><th align="center" valign="middle" >Stations/Pollutants</th><th align="center" valign="middle" >Classes of Values</th></tr></thead><tr><td align="center" valign="middle" >Factor 1</td><td align="center" valign="middle" >41/NO<sub>2</sub> (95) 101/NO<sub>2</sub> (95) 86/CO (93) 103/NO<sub>2</sub> (91) 125/NO<sub>2</sub> (89) 42/NO<sub>2</sub> (87)</td><td align="center" valign="middle" >c6 (84) c5 (76)</td></tr><tr><td align="center" valign="middle" >Factor 2</td><td align="center" valign="middle" >85/NO<sub>2</sub> (92) 107/CO (89) 97/CO (86)</td><td align="center" valign="middle" >c3 (65) c1 (41)</td></tr></tbody></table></table-wrap><p>The position of the points and the absolute and relative contributions suggest the following comments.</p><p>CA applied to the matrix A.</p><p>As previously described, matrix A relates six non-overlapping classes of values to CO and NO<sub>2</sub> survey stations, so that, by analyzing this matrix, it is possible to finding underlying relationships in space among</p><table-wrap id="table5" ><label><xref ref-type="table" rid="table5">Table 5</xref></label><caption><title> Highest absolute contributions to the first two factors from CA applied to the matrix B (percentages in parentheses)</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Factors</th><th align="center" valign="middle" >Hours/Pollutants</th><th align="center" valign="middle" >Classes of Values</th></tr></thead><tr><td align="center" valign="middle" >Factor 1</td><td align="center" valign="middle" >6/NO<sub>2</sub> (7) 5/NO<sub>2</sub> (6) 6/CO (6) 5/CO (6) 4/NO<sub>2</sub> (5) 19/CO (5) 20/CO (5) 4/CO (4)</td><td align="center" valign="middle" >c6 (42) c1 (32)</td></tr><tr><td align="center" valign="middle" >Factor 2</td><td align="center" valign="middle" >12/CO (9) 13/CO (7) 13/NO<sub>2</sub> (7) 12/NO<sub>2</sub> (4)</td><td align="center" valign="middle" >c6 (34) c1 (25)</td></tr></tbody></table></table-wrap><table-wrap id="table6" ><label><xref ref-type="table" rid="table6">Table 6</xref></label><caption><title> Highest relative contributions to the first two factors from CA applied to the matrix B (percentages in parentheses)</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Factors</th><th align="center" valign="middle" >Hours/Pollutants</th><th align="center" valign="middle" >Classes of Values</th></tr></thead><tr><td align="center" valign="middle" >Factor 1</td><td align="center" valign="middle" >6/NO<sub>2</sub> (97) 6/CO (97) 4/CO (96) 5/CO (95) 5/NO<sub>2</sub> (92) 19/CO (89) 20/CO (88) 4/NO<sub>2</sub> (86)</td><td align="center" valign="middle" >c2 (92) c6 (91)</td></tr><tr><td align="center" valign="middle" >Factor 2</td><td align="center" valign="middle" >13/CO (83) 12/CO (82)</td><td align="center" valign="middle" >c3 (44) c4 (43)</td></tr></tbody></table></table-wrap><p>different pollution levels and monitored locations. The first two factors are retained since they explain together about 85% of the total variation (<xref ref-type="table" rid="table3">Table 3</xref>).</p><p>The last class of values (c6) and the first one (c1) have the highest absolute contributions to the first factor, respectively, 39% and 29% (<xref ref-type="table" rid="table2">Table 2</xref>); whereas the first class (c1) and the third one (c3) have the highest absolute contribution to the second factor (respectively, 50% and 21%). Hence, the first factor better explains the variation of high pollution levels, i.e. levels which are greater than the last quantile values (5.4 mg/m<sup>3</sup> for CO and 115 mg/m<sup>3</sup> for NO<sub>2</sub>), while the second factor better explains the variation of low pollution levels, i.e. levels which are smaller than 1,6 mg/m<sup>3</sup> for CO and 52 mg/m<sup>3</sup> for NO<sub>2</sub>. <xref ref-type="table" rid="table2">Table 2</xref> list also the stations/pollutants which have the highest absolute contributions to the retained factors. The first factor better represents stations located at the central area of the Milan district (86, 125, 41, 93, 15), while the second factor better represents stations located at the peripheral areas (62, 97, 111).</p><p>The cumulative relative contributions to the first two factors (<xref ref-type="table" rid="table4">Table 4</xref> shows those stations/pollutants and classes with the highest relative contributions) are always greater than 80%, highlighting the good quality of representation of rows and columns in the space determining by the first two factors.</p><p>The projection of the classes to the first factorial plane (<xref ref-type="fig" rid="fig3">Figure 3</xref>) shows a horseshoe effect [<xref ref-type="bibr" rid="scirp.60721-ref6">6</xref>] which corresponds to a non-linear relationships between the two axes, even if they are linearly orthogonal.</p><p>In <xref ref-type="fig" rid="fig3">Figure 3</xref>, classes and stations/pollutants are displayed together so that and it is possible to identify two clusters of stations/pollutants:</p><p>1) points 111, 11, 45, 81, referred to CO, and point 41, referred to NO<sub>2</sub>, with positive first and second co-ordinate;</p><p>2) points 86, 15, 93, 113, 101, referred to CO, and points 102, 15, referred to NO<sub>2</sub>, with negative first co-ordinate.</p><p>The position of the second cluster on the factorial plane, being the points closer to point c6 with respect to the other points, highlights that most of the highest pollutant concentrations was read during January 1999 at those locations.</p><p>CA applied to the matrix B.</p><p>Matrix B summarizes the spatial aspect for each hour, since in this matrix each entry indicates how many monitoring stations, at a fixed hour, have recorded pollution levels belonging to a given class of values. Hence, by analyzing this matrix, underling relationships among observation times (hours) and different pollution levels can be identified.</p><p><xref ref-type="table" rid="table2">Table 2</xref> shows the eigenvalues and the percentages of variation explained by each of the 5 non-trivial factors. In this case, the first factor explains a greater part of the total variation (85.26%) than in the previous analysis. The greater the percentage of explained variation, the greater the association between rows and columns of the data matrix, then the high percentage of variation explained by this factor is due to a strong association between observation times and classes of pollution levels. Once more, the first two factors are retained since they explain together more than 95% of the total variation.</p><p><xref ref-type="fig" rid="fig4">Figure 4</xref> shows the projections of rows (hours/pollutants) and columns (classes of values) to the first factorial plane. Now, a horseshoe effect is evident not only in the projections of the classes, but also in the projections of the hours/pollutants on the factorial plane: this means that distant pairs of hours can be considered as equidistant, while neighbouring hours are progressively dissimilar.</p><p><xref ref-type="table" rid="table5">Table 5</xref> and <xref ref-type="table" rid="table6">Table 6</xref> list the classes of values and the hours/pollutants with the highest absolute (<xref ref-type="table" rid="table5">Table 5</xref>) and relative (<xref ref-type="table" rid="table6">Table 6</xref>) contributions. Hours 4, 5, 6 referred to both contaminants have the highest absolute contribu- tions to the first factor and, by looking to the factorial plane (<xref ref-type="fig" rid="fig4">Figure 4</xref>), the position of these observation times closer to point c1 with respect to the other points highlights that most of CO and NO<sub>2</sub> low readings was measured from the 4-th to the 6-th hour. Instead, most of the high pollution concentrations was observed during the evening, particularly during the 19-th to the 22-nd hour, for CO and from the 12-th to 14-th hour, for NO<sub>2</sub>.</p></sec><sec id="s5"><title>5. Conclusion</title><p>In this work, an application of CA to an air pollution space-time data set for CO and NO<sub>2</sub> hourly concentrations, recorded at some monitoring stations in Milan district, is given. The transformation of the original continuous variables into new categorical ones has been formally presented in this paper by the means of the indicator approach. By counting the indicator data over both spatial locations and observation times, two contingency matrices are generated. Each of them accounts information of both pollutants examined in this paper. CA is applied to these matrices providing a summary description of spatial and temporal profiles, simultaneously for the contaminants under study. The data analysis allows identifying relationships in space among CO and NO<sub>2</sub> pollution levels and monitored stations and relationships in time among CO and NO<sub>2</sub> pollution levels and observation times. The aim of each air quality control system is to obtain information about the atmospherical conditions and evaluate the opportunity of major restrictions and closer controls. The application of CA carried out in this paper makes it possible, since its graphical results and diagnostics help in identifying stations inside the area under study and intervals of time during the day for which the contaminants of interest need closer controls because of joint exceeding of fixed pollution levels.</p></sec><sec id="s6"><title>Acknowledgements</title><p>The author would like to thank Prof. Donato Posa of University of Salento, Apulian region (Italy), whose suggestions have been helpful and improved this paper.</p></sec><sec id="s7"><title>Cite this paper</title><p>PalmaMonica, (2015) Correspondence Analysis on a Space-Time Data Set for Multiple Environmental Variables. International Journal of Geosciences,06,1154-1165. doi: 10.4236/ijg.2015.610090</p></sec></body><back><ref-list><title>References</title><ref id="scirp.60721-ref1"><label>1</label><mixed-citation publication-type="book" xlink:type="simple">De Iaco, S., Myers, D.E. and Posa, D. (2000) Total Air Pollution and Space-Time Modeling. In: Monestiez P., Allard D. and Froidevaux R., Eds., GeoEnv III, Geostatistics for Environmental Applications, Kluwer Academic Publishers, Norwell, 45-56.</mixed-citation></ref><ref id="scirp.60721-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">De Iaco, S. (2011) A New Space-Time Multivariate Approach for Environmental Data Analysis. Journal of Applied Statistics, 38, 2471-2483. http://dx.doi.org/10.1080/02664763.2011.559206</mixed-citation></ref><ref id="scirp.60721-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">Blasius, J., Greenacre, M., Groenen, P.J.F. and van de Velden, M. (2009) Special Issue on Correspondence Analysis and Related Methods. Computational Statistics and Data Analysis, 53, 3103-3106.  
http://dx.doi.org/10.1016/j.csda.2008.11.010</mixed-citation></ref><ref id="scirp.60721-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">Lebart, L., Morineau, A. and Warwick, K.M. (1984) Multivariate Descriptive Statistical Analysis. John Wiley &amp; Sons, New York.</mixed-citation></ref><ref id="scirp.60721-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">Benzécri, J.P. (1983) Histoire et préhistoire de l’analyse des données. Dunod, Paris.</mixed-citation></ref><ref id="scirp.60721-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">Greenacre, M.J. (1989) Theory and Applications of Correspondence Analysis. Academic Press, London.</mixed-citation></ref><ref id="scirp.60721-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">SAS/Stat (1990) SAS Institute Inc., Cary.</mixed-citation></ref><ref id="scirp.60721-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">(1999) SPSS 8.0, SPSS Inc., Chicago.</mixed-citation></ref><ref id="scirp.60721-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">Avila, F. and Myers, D.E. (1991) Correspondence Analysis Applied to Environmental Data Sets: A Study of Chautauqua Lake sediments. Chemometrics and Intelligent Laboratory Systems, 11, 229-249.  
http://dx.doi.org/10.1016/0169-7439(91)85002-7</mixed-citation></ref><ref id="scirp.60721-ref10"><label>10</label><mixed-citation publication-type="other" xlink:type="simple">Avila, F., Myers, D.E. and Palmer, C. (1991) Correspondence Analysis and Adsorbate Selection for Chemical Sensor Arrays. Journal of Chemometrics, 5, 455-465. http://dx.doi.org/10.1002/cem.1180050505</mixed-citation></ref><ref id="scirp.60721-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">Dutot, A.L., Bergametti, G. and Buat-Menard, P. (1988) Application of Correspondence Analysis to Apportion Sources of Ambient Particles. Athmospheric Environment, 22, 1737-1743. http://dx.doi.org/10.1016/0004-6981(88)90403-9</mixed-citation></ref><ref id="scirp.60721-ref12"><label>12</label><mixed-citation publication-type="book" xlink:type="simple">Jiménez-Espinosa, R., Sousa, A.J. and Chica-Olmo, M. (1992) Application of Correspondence Analysis and Factorial Kriging Analysis: A Case Study on Geochemical Exploration in Geostatistics. 2, Soares, A. Ed., Troia, 853-864.</mixed-citation></ref><ref id="scirp.60721-ref13"><label>13</label><mixed-citation publication-type="other" xlink:type="simple">Greenacre, M.J. and Primicerio, R. (2013) Multivariate Analysis for Ecological Data. Fundación BBVA, Bilbao.</mixed-citation></ref><ref id="scirp.60721-ref14"><label>14</label><mixed-citation publication-type="other" xlink:type="simple">De Iaco, S., Palma, M. and Posa, D. (2013) Prediction of Particle Pollution through Spatio-Temporal Multivariate Geostatistical Analysis: Spatial Special Issue. AStA Advanced Statistical Analysis, 97, 133-150.  
http://dx.doi.org/10.1007/s10182-012-0199-0</mixed-citation></ref><ref id="scirp.60721-ref15"><label>15</label><mixed-citation publication-type="book" xlink:type="simple">De Iaco, S., Maggio, S., Palma, M. and Posa, D. (2012) Advances in Spatio-Temporal Modeling and Prediction for Environmental Risk Assessment. In: Haryanto, B., Ed., Air Pollution—A Comprehensive Perspective, InTech, Croazia, 365-390. http://dx.doi.org/10.5772/51227</mixed-citation></ref><ref id="scirp.60721-ref16"><label>16</label><mixed-citation publication-type="other" xlink:type="simple">(1999) SPAD 4.0, Cisia, Montreuil Cedex, France.</mixed-citation></ref><ref id="scirp.60721-ref17"><label>17</label><mixed-citation publication-type="book" xlink:type="simple">Morineau, A. and Lebart, L. (1986) Specific Clustering Algorithms for Large Data Sets and Implementation in SPAD Software, in Classification as a Tool of Research. In: Gaul, W. and Schader, M., Eds., Classification as a Tool of Research: Proceedings of the 9th Annual Meeting of the Classification Society (F.R.G.), University of Karlsrube, F.R.G., North Holland, 321-329.</mixed-citation></ref></ref-list></back></article>