<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">JCC</journal-id><journal-title-group><journal-title>Journal of Computer and Communications</journal-title></journal-title-group><issn pub-type="epub">2327-5219</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/jcc.2020.810008</article-id><article-id pub-id-type="publisher-id">JCC-103821</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Computer Science&amp;Communications</subject></subj-group></article-categories><title-group><article-title>
 
 
  Hybrid Warehouse Model and Solutions for Climate Data Analysis
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Hasan</surname><given-names>Hashim</given-names></name><xref ref-type="aff" rid="aff1"><sub>1</sub></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib></contrib-group><aff id="aff1"><label>1</label><addr-line>College of Computer Science and Engineering, Yanbu Taibah University, Taibah, KSA</addr-line></aff><pub-date pub-type="epub"><day>22</day><month>10</month><year>2020</year></pub-date><volume>08</volume><issue>10</issue><fpage>75</fpage><lpage>98</lpage><history><date date-type="received"><day>29,</day>	<month>September</month>	<year>2020</year></date><date date-type="rev-recd"><day>27,</day>	<month>October</month>	<year>2020</year>	</date><date date-type="accepted"><day>30,</day>	<month>October</month>	<year>2020</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
  Recently, due to the rapid growth increment of data sensors, a massive volume of data is generated from different sources. The way of administering such data in a sense storing, managing, analyzing, and extracting insightful information from the massive volume of data is a challenging task. Big data analytics is becoming a vital research area in domains such as climate data analysis which demands fast access to data. Nowadays, an open-source platform namely MapReduce which is a distributed computing framework is widely used in many domains of big data analysis. In our work, we have developed a conceptual framework of data modeling essentially useful for the implementation of a hybrid data warehouse model to store the features of National Climatic Data Center (NCDC) climate data. The hybrid data warehouse model for climate big data enables for the identification of weather patterns that would be applicable in agricultural and other similar climate change-related studies that will play a major role in recommending actions to be taken by domain experts and make contingency plans over extreme cases of weather variability.
 
</p></abstract><kwd-group><kwd>Data Warehouse</kwd><kwd> Hadoop</kwd><kwd> NCDC Data Set</kwd><kwd> Weather</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>Nowadays, the volume of data generated from different sources is rapidly increasing. Hence, administering and processing such a massive volume of data is challenging. It is time-consuming, costly, and has many obstacles to researchers involving it. The way of administering such data in a sense storing, managing, analyzing and extracting insightful information from the huge size of data is a challenging task. The volume of data produced on a daily basis is a massive amount whereby this accelerating growth of data is due to the growth of the Internet of Things (IoT), Artificial Intelligence and Data Science. To make use of the huge volume of data, researchers in the domain often use cutting-edge techniques such as analytical tools and techniques with the help of Artificial Intelligence and machine learning methods. The main concern with big data is privacy, security, and discrimination. Researchers are then to start building support systems such as data warehouse.</p><p>The huge amount of weather data and climatic variables are recorded manually or digitally using many resources such as weather stations and satellites around the country. Thus, the storage and manipulation task of the information has to be effective, integrated and more flexible in Meteorological and climatology studies to achieve an effective and intensive analysis. The access to recorded Meteorological data varies according to the certain process, for instance, in weather forecasting tasks the raw data needed to be accessed rapidly and in Climatology, to increase the accuracy it is important to have high-quality historical weather data [<xref ref-type="bibr" rid="scirp.103821-ref1">1</xref>].</p><p>With technological advances in the area of tools and the different types of equipment to collect weather data, researchers can access and share huge data. Generally, the NCDC and Daily Global Weather Measurements 1901-2020 (GSOD, NCDC) is the largest active archive of climate data available online so far. In addition, the history of the NCDC dataset goes back to more than 150 years of data and the size of new data collected daily is close to 224 gigabytes. Moreover, it is easily accessible and downloads NCDC and GSOD datasets from NCDC website [<xref ref-type="bibr" rid="scirp.103821-ref2">2</xref>].</p><p>In the literature, it is suggested that a clear definition of a data warehouse which is a single, consistent, complete storage of data acquired from various sources in order to analyze it using a business intelligence tool [<xref ref-type="bibr" rid="scirp.103821-ref3">3</xref>] [<xref ref-type="bibr" rid="scirp.103821-ref4">4</xref>]. Data warehouse technology is applicable in a lot of domains in industry that use historical data for prediction, statistical analysis, and decision making, for instance, banking, consumer goods, finance industry, weather, healthcare, and Internet of Things (IoT) [<xref ref-type="bibr" rid="scirp.103821-ref1">1</xref>] [<xref ref-type="bibr" rid="scirp.103821-ref5">5</xref>] [<xref ref-type="bibr" rid="scirp.103821-ref6">6</xref>] [<xref ref-type="bibr" rid="scirp.103821-ref7">7</xref>] [<xref ref-type="bibr" rid="scirp.103821-ref8">8</xref>] [<xref ref-type="bibr" rid="scirp.103821-ref9">9</xref>]. Moreover, the huge amount of sensor data collected from the Meteorological domain has many problems to be addressed in the preprocessing stages. It demands careful manipulation so as to carry out weather analysis in an accurate way by extracting relevant patterns [<xref ref-type="bibr" rid="scirp.103821-ref10">10</xref>].</p><p>The traditional way of storing Meteorological data was file-based. Whereas, recently saving weather data in relational database management systems (RDMS) is attracting the attention of researchers in the area of climatology studies. As a result, many institutions operating in this domain are shifting towards storing their data in terms of RDMS [<xref ref-type="bibr" rid="scirp.103821-ref11">11</xref>]. Chen [<xref ref-type="bibr" rid="scirp.103821-ref12">12</xref>] describes a data warehouse framework namely Cheetah which is designed using the MapReduce platform. Dimri et al. [<xref ref-type="bibr" rid="scirp.103821-ref13">13</xref>] proposed a data warehouse model for storing weather data using On-Line Analysis Processing (OLAP) method to generate proper data for weather analysis and provide a multi-dimensional report. Jos&#233; et al. [<xref ref-type="bibr" rid="scirp.103821-ref14">14</xref>] developed a data warehouse to save climatic variables in the weather stations of Mexico. The authors have used the SQL Server for storing the final version of data in the proposed model. Sameer and Madhu [<xref ref-type="bibr" rid="scirp.103821-ref15">15</xref>] discussed how to design a data warehousing model using Hadoop.</p><p>The main contributions of this work are:</p><p>• The main contribution of this work is the definition and implementation of hybrid data warehouse infrastructure to support the distribution of weather data storage, computing, and parallel programming.</p><p>• The new data warehouse can be implemented in different ways to store huge data sets and workloads for distribution in hybrid Hadoop platform.</p><p>• The main aim is to preserve and improve a traditional data warehouse for reporting, OLAP, and performance management while new development in data platforms for advanced analytics.</p><p>• The hybrid data warehouse model for climate big data enables for the identification of weather patterns that would be applicable in agricultural and other similar climate change-related studies that will play a major role in recommending actions to be taken by domain experts and make contingency plans over extreme cases of weather variability.</p><p>The rest of the paper is organized as follows. Section 2 presents a review of the related works in the data warehouse and big data context. Section 3 deals with the dataset description. Section 4 describes the concept and architecture of the data warehouse. Section 5 presents the concept of Big Data and the required tools such as Hadoop, Pig, Hive, and Sqoop. Section 6 shows the proposed data warehouse model for the weather data under consideration. Finally, concluding remarks and summaries are presented in Section 7.</p></sec><sec id="s2"><title>2. Related Work</title><p>There are many research works that have been done to design a data warehouse based on Big Data and Hadoop framework.</p><p>Kalra and Steiner [<xref ref-type="bibr" rid="scirp.103821-ref7">7</xref>] explained the quality and content of data vary over time based on the types of information for instance, weather, and health data. Thus, this data needs to be gathered, processed and stored in different formats. The authors developed a weather data warehouse model that enables dynamic and smooth integration of new information sources and data formats. The proposed architecture depicts an active and flexible weather data warehouse that provides a broad variety of weather data from various sources to different weather-based applications.</p><p>N&#233;stor et al. [<xref ref-type="bibr" rid="scirp.103821-ref16">16</xref>] discussed the problems of data recorded by meteorological which require strategies for capturing, delivering, storing and processing to increase the quality and stability of the data. They proposed a star schema model for a data warehouse that allows storage and analysis of historical multidimensional hydro-climatological data. Moreover, the proposed data warehouse provides efficient data storage in which data collected from two networks of hydro-meteorological stations goes back to more than 50 years in the city of Manizales, Colombia.</p><p>Doreswamy et al., [<xref ref-type="bibr" rid="scirp.103821-ref1">1</xref>] proposed a scalable architecture for a hybrid data warehouse approach for climate data using Hadoop and various big data tools. The proposed schema enables the identification of weather patterns and it is better to derive knowledge from the data in comparison to the traditional database.</p><p>Vuong et al. [<xref ref-type="bibr" rid="scirp.103821-ref17">17</xref>] show how designing and developing a data warehouse for the agriculture field has the main role in establishing a crop intelligence platform. They describe the requirements for efficient agricultural data-warehouses such as privacy, security, and real-time access among its stakeholders. Thus, the proposed system architecture and a database schema for designing and implementing an efficient agricultural data warehouse in Big Data and data mining.</p></sec><sec id="s3"><title>3. NCDC Dataset Details</title><p>This section presents specific descriptions about the data produced by the NCDC dataset for Saudia country. The NCDC is a large dataset that has more than 9000 stations around the globe and is available online from NCDC meteorological site [<xref ref-type="bibr" rid="scirp.103821-ref2">2</xref>]. <xref ref-type="fig" rid="fig1">Figure 1</xref> shows the selected Saudi Arabia weather stations from the NCDC dataset and each station has 16 attributes. <xref ref-type="table" rid="table1">Table 1</xref> presents the column names and their corresponding description.</p><table-wrap id="table1" ><label><xref ref-type="table" rid="table1">Table 1</xref></label><caption><title> Description of the attributes for each station</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >No.</th><th align="center" valign="middle" >Attribute</th><th align="center" valign="middle" >Description</th></tr></thead><tr><td align="center" valign="middle" >1</td><td align="center" valign="middle" >STN</td><td align="center" valign="middle" >Station number</td></tr><tr><td align="center" valign="middle" >2</td><td align="center" valign="middle" >WBAN</td><td align="center" valign="middle" >Weather Bureau Air force Navy</td></tr><tr><td align="center" valign="middle" >3</td><td align="center" valign="middle" >YEARMODA</td><td align="center" valign="middle" >The year, month and day</td></tr><tr><td align="center" valign="middle" >4</td><td align="center" valign="middle" >TEMP</td><td align="center" valign="middle" >Mean temperature</td></tr><tr><td align="center" valign="middle" >5</td><td align="center" valign="middle" >DEWP</td><td align="center" valign="middle" >Mean dew point</td></tr><tr><td align="center" valign="middle" >6</td><td align="center" valign="middle" >SLP</td><td align="center" valign="middle" >Mean sea level pressure in millibars</td></tr><tr><td align="center" valign="middle" >7</td><td align="center" valign="middle" >STP</td><td align="center" valign="middle" >Mean station pressure</td></tr><tr><td align="center" valign="middle" >8</td><td align="center" valign="middle" >VISIB</td><td align="center" valign="middle" >Visibility</td></tr><tr><td align="center" valign="middle" >9</td><td align="center" valign="middle" >WDSP</td><td align="center" valign="middle" >Mean wind speed in knots</td></tr><tr><td align="center" valign="middle" >10</td><td align="center" valign="middle" >MXSPD</td><td align="center" valign="middle" >Maximum sustained wind speed</td></tr><tr><td align="center" valign="middle" >11</td><td align="center" valign="middle" >GUST</td><td align="center" valign="middle" >Maximum wind gust</td></tr><tr><td align="center" valign="middle" >12</td><td align="center" valign="middle" >MAX</td><td align="center" valign="middle" >Maximum temperature</td></tr><tr><td align="center" valign="middle" >13</td><td align="center" valign="middle" >MIN</td><td align="center" valign="middle" >Minimum temperature</td></tr><tr><td align="center" valign="middle" >14</td><td align="center" valign="middle" >PRCP</td><td align="center" valign="middle" >Precipitation</td></tr><tr><td align="center" valign="middle" >15</td><td align="center" valign="middle" >SNDP</td><td align="center" valign="middle" >Snow depth in inches</td></tr><tr><td align="center" valign="middle" >16</td><td align="center" valign="middle" >FRSHTT</td><td align="center" valign="middle" >Fog rain snow hail thunder Tornado</td></tr></tbody></table></table-wrap><p><xref ref-type="table" rid="table2">Table 2</xref> presents information such as station number, station name, latitude, longitude, begin and end date for an illustrative case of the selected Saudi Arabia weather stations from the NCDC dataset. Moreover, <xref ref-type="table" rid="table3">Table 3</xref> demonstrates a sample data of one station of the NCDC dataset and the corresponding values for each feature. The problem associated with the NCDC dataset is that it has several missing values. The missing data for the selected attributes are taking the following values: 9999.9, 999.9 or 99.99. For example, a 9999.9 shows a missing value for the variables TEMP, DEWP, SLP, STP, MAX and MIN, whereas 999.9 indicates a missing value for the column names VISIB, WDSP, MXSPD, GUST and SNDP. Moreover, 99.99 shows a missing value for the column PRCP.</p></sec><sec id="s4"><title>4. Data Warehouse (DW)</title><p>In the 1990s, Bill Inmon suggested the first architecture of data warehouse (DW) model. While Gartner, in 2005, provided a clear concept of the DW. The main task of DW schema is to accumulate and keep data from various sources for future analysis and decision making. Generally, the classic relational database schema is used to store, manage and query structured data. The present DW model is briefed as follows [<xref ref-type="bibr" rid="scirp.103821-ref4">4</xref>]:</p><p>• Subject oriented: In this case, the entire data is manipulated to classify into various domain areas, for instance, each domain will have complete data related to each subject.</p><p>• Integrated: This means that the following conditions should be satisfied: the logical model should be integrated and consistent. For instance, the values of the data should be standardized such as female/male representations should be consistent.</p><table-wrap id="table2" ><label><xref ref-type="table" rid="table2">Table 2</xref></label><caption><title> The selected Saudi Arabia stations from the NCDC dataset</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >STATION STN</th><th align="center" valign="middle" >STATION WBN</th><th align="center" valign="middle" >STATION Name</th><th align="center" valign="middle" >C</th><th align="center" valign="middle" >LATITUDE</th><th align="center" valign="middle" >LONGITUDE</th><th align="center" valign="middle" >ELEV</th><th align="center" valign="middle" >Begin DATE</th><th align="center" valign="middle" >End DATE</th></tr></thead><tr><td align="center" valign="middle" >410060</td><td align="center" valign="middle" >99999</td><td align="center" valign="middle" >MUWAIH</td><td align="center" valign="middle" >SA</td><td align="center" valign="middle" >+22.433</td><td align="center" valign="middle" >+041.750</td><td align="center" valign="middle" >+0971.0</td><td align="center" valign="middle" >19851001</td><td align="center" valign="middle" >20110614</td></tr><tr><td align="center" valign="middle" >410080</td><td align="center" valign="middle" >99999</td><td align="center" valign="middle" >ZULM</td><td align="center" valign="middle" >SA</td><td align="center" valign="middle" >+22.717</td><td align="center" valign="middle" >+042.167</td><td align="center" valign="middle" >+0870.0</td><td align="center" valign="middle" >20041015</td><td align="center" valign="middle" >20041015</td></tr><tr><td align="center" valign="middle" >410200</td><td align="center" valign="middle" >99999</td><td align="center" valign="middle" >JEDDAH I.E.</td><td align="center" valign="middle" >SA</td><td align="center" valign="middle" >+21.417</td><td align="center" valign="middle" >+039.217</td><td align="center" valign="middle" >+0010.0</td><td align="center" valign="middle" >19910605</td><td align="center" valign="middle" >20050729</td></tr><tr><td align="center" valign="middle" >410240</td><td align="center" valign="middle" >99999</td><td align="center" valign="middle" >KING ABDULAZIZ INTL</td><td align="center" valign="middle" >SA</td><td align="center" valign="middle" >+21.680</td><td align="center" valign="middle" >+039.157</td><td align="center" valign="middle" >+0014.6</td><td align="center" valign="middle" >19830101</td><td align="center" valign="middle" >20161211</td></tr><tr><td align="center" valign="middle" >410260</td><td align="center" valign="middle" >99999</td><td align="center" valign="middle" >JEDDAH</td><td align="center" valign="middle" >SA</td><td align="center" valign="middle" >+21.500</td><td align="center" valign="middle" >+039.200</td><td align="center" valign="middle" >+0015.0</td><td align="center" valign="middle" >19560101</td><td align="center" valign="middle" >20061026</td></tr><tr><td align="center" valign="middle" >410300</td><td align="center" valign="middle" >99999</td><td align="center" valign="middle" >MAKKAH</td><td align="center" valign="middle" >SA</td><td align="center" valign="middle" >+21.433</td><td align="center" valign="middle" >+039.767</td><td align="center" valign="middle" >+0240.0</td><td align="center" valign="middle" >19830701</td><td align="center" valign="middle" >20161211</td></tr><tr><td align="center" valign="middle" >410360</td><td align="center" valign="middle" >99999</td><td align="center" valign="middle" >TAIF</td><td align="center" valign="middle" >SA</td><td align="center" valign="middle" >+21.483</td><td align="center" valign="middle" >+040.544</td><td align="center" valign="middle" >+1477.7</td><td align="center" valign="middle" >19830101</td><td align="center" valign="middle" >20161211</td></tr><tr><td align="center" valign="middle" >411140</td><td align="center" valign="middle" >99999</td><td align="center" valign="middle" >KING KHALED AB</td><td align="center" valign="middle" >SA</td><td align="center" valign="middle" >+18.297</td><td align="center" valign="middle" >+042.804</td><td align="center" valign="middle" >+2065.9</td><td align="center" valign="middle" >19830101</td><td align="center" valign="middle" >20161211</td></tr><tr><td align="center" valign="middle" >411280</td><td align="center" valign="middle" >99999</td><td align="center" valign="middle" >NEJRAN</td><td align="center" valign="middle" >SA</td><td align="center" valign="middle" >+17.611</td><td align="center" valign="middle" >+044.419</td><td align="center" valign="middle" >+1213.7</td><td align="center" valign="middle" >19830101</td><td align="center" valign="middle" >20161211</td></tr><tr><td align="center" valign="middle" >411400</td><td align="center" valign="middle" >99999</td><td align="center" valign="middle" >KING ABDULLAH BIN ABDULAZIZ</td><td align="center" valign="middle" >SA</td><td align="center" valign="middle" >+16.901</td><td align="center" valign="middle" >+042.586</td><td align="center" valign="middle" >+0006.1</td><td align="center" valign="middle" >19830101</td><td align="center" valign="middle" >20161211</td></tr></tbody></table></table-wrap><table-wrap id="table3" ><label><xref ref-type="table" rid="table3">Table 3</xref></label><caption><title> A sample of records and the attributes of one station</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >STN-WBAN</th><th align="center" valign="middle" >YEARMODA</th><th align="center" valign="middle" >TEMP</th><th align="center" valign="middle" >DEWP</th><th align="center" valign="middle" >SLP</th><th align="center" valign="middle" >STP</th><th align="center" valign="middle" >VISIB</th><th align="center" valign="middle" >WDSP</th><th align="center" valign="middle" >MXSPD</th><th align="center" valign="middle" >GUST</th><th align="center" valign="middle" >MAX</th><th align="center" valign="middle" >MIN</th><th align="center" valign="middle" >PRCP</th><th align="center" valign="middle" >SNDP</th></tr></thead><tr><td align="center" valign="middle" >410240 99999</td><td align="center" valign="middle" >20180101</td><td align="center" valign="middle" >75.1</td><td align="center" valign="middle" >59.1</td><td align="center" valign="middle" >1013.3</td><td align="center" valign="middle" >1011.4</td><td align="center" valign="middle" >6.2</td><td align="center" valign="middle" >12.0</td><td align="center" valign="middle" >20.0</td><td align="center" valign="middle" >999.9</td><td align="center" valign="middle" >78.8</td><td align="center" valign="middle" >69.8</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >999.9</td></tr><tr><td align="center" valign="middle" >410240 99999</td><td align="center" valign="middle" >20180102</td><td align="center" valign="middle" >73.1</td><td align="center" valign="middle" >47.5</td><td align="center" valign="middle" >1013.1</td><td align="center" valign="middle" >1011.1</td><td align="center" valign="middle" >6.2</td><td align="center" valign="middle" >8.4</td><td align="center" valign="middle" >12.0</td><td align="center" valign="middle" >999.9</td><td align="center" valign="middle" >78.8</td><td align="center" valign="middle" >66.2</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >999.9</td></tr><tr><td align="center" valign="middle" >410240 99999</td><td align="center" valign="middle" >20180103</td><td align="center" valign="middle" >77.0</td><td align="center" valign="middle" >49.4</td><td align="center" valign="middle" >1011.6</td><td align="center" valign="middle" >1009.7</td><td align="center" valign="middle" >6.0</td><td align="center" valign="middle" >7.2</td><td align="center" valign="middle" >15.0</td><td align="center" valign="middle" >999.9</td><td align="center" valign="middle" >87.8</td><td align="center" valign="middle" >67.6</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >999.9</td></tr><tr><td align="center" valign="middle" >410240 99999</td><td align="center" valign="middle" >20180104</td><td align="center" valign="middle" >74.7</td><td align="center" valign="middle" >58.2</td><td align="center" valign="middle" >1013.5</td><td align="center" valign="middle" >1011.5</td><td align="center" valign="middle" >6.0</td><td align="center" valign="middle" >6.1</td><td align="center" valign="middle" >15.0</td><td align="center" valign="middle" >999.9</td><td align="center" valign="middle" >80.6</td><td align="center" valign="middle" >69.8</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >999.9</td></tr><tr><td align="center" valign="middle" >410240 99999</td><td align="center" valign="middle" >20180105</td><td align="center" valign="middle" >71.4</td><td align="center" valign="middle" >53.3</td><td align="center" valign="middle" >1014.4</td><td align="center" valign="middle" >1012.5</td><td align="center" valign="middle" >6.2</td><td align="center" valign="middle" >7.6</td><td align="center" valign="middle" >14.0</td><td align="center" valign="middle" >999.9</td><td align="center" valign="middle" >78.8</td><td align="center" valign="middle" >66.2</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >999.9</td></tr><tr><td align="center" valign="middle" >410240 99999</td><td align="center" valign="middle" >20180106</td><td align="center" valign="middle" >71.5</td><td align="center" valign="middle" >52.8</td><td align="center" valign="middle" >1014.4</td><td align="center" valign="middle" >1012.5</td><td align="center" valign="middle" >6.2</td><td align="center" valign="middle" >9.9</td><td align="center" valign="middle" >15.0</td><td align="center" valign="middle" >999.9</td><td align="center" valign="middle" >80.6</td><td align="center" valign="middle" >65.5</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >999.9</td></tr></tbody></table></table-wrap><p>• Nonvolatile: This means that the unmodified data shall be saved for long period of time in the DW.</p><p>• Time variant: In this case, the DW has the ability to store the newly modified versions of records.</p><p>• Not virtual: In this regard, the DW saves the data in the physical storage area for a longer period of time persistently.</p><p>The ETL stands for Extract, Transform and Load that stands for three database operations that are combined into one tool to extract data from different sources and place it into another data warehouse. The extract is the operation of reading data from multiple sources and different natures data such as structured, semistructured and unstructured data. The transform operation is used to fit the data into the analytical model. Moreover, the load operation deals with physically storing the data into the data warehouse. Finally, analysis, visualization and report generating are carried out as a final core activity in constructing a DW model (See <xref ref-type="fig" rid="fig2">Figure 2</xref>) [<xref ref-type="bibr" rid="scirp.103821-ref3">3</xref>].</p><sec id="s4_1"><title>4.1. The OLTP Data Processing</title><p>The Online Transactional Processing (OLTP) is a type of data processing technique</p><p>that deals with transaction-related tasks and sub-tasks [<xref ref-type="bibr" rid="scirp.103821-ref18">18</xref>]. The common tasks in OLTP are inserting, updating, and deleting of data in a database by a large number of users concurrently. The OLTP handles recent operational data in small to medium-sized data with the goal to perform daily operations. It uses simple queries with read/write operations essential for faster transaction speeds. OLTP architecture is used to manage the day-to-day transaction of a business entity. The main aim of OLTP is for data processing and not for data analysis. Some example of OLTP related-system, such as the ATM center where it makes sure that withdrawal of a certain amount from ATM machine can not be more than the amount deposited in the saving account. Therefore, OLTP systems are highly optimized for transactional related tasks such as online banking, online flight ticket booking, sending a text message, order entry and adding a textbook to shopping carts [<xref ref-type="bibr" rid="scirp.103821-ref19">19</xref>].</p><p>The OLTP is known for the following two features namely concurrency and atomicity whereby the atomicity guarantees if a single step is incomplete during the transaction, then the entire process automatically stops. On the other hand, the concurrency property of the system prevents the altering of data by multiple users simultaneously. The main benefit of the OLTP system is it allows the administration of daily transactions of organizations’ data and widens customer satisfaction by simplifying routine processes. <xref ref-type="table" rid="table4">Table 4</xref> illustrates the differences between OLTP and OLAP.</p></sec><sec id="s4_2"><title>4.2. Data Warehouse: Terminologies</title><p>The famous approaches used for data modeling are: 1) The Dimensional Model or Star Schema is created using two types of tables namely fact and dimension. 2) The Normalized Model is designed similarly to the way OLTP is designed. Moreover, a star schema is easier and faster in terms of executing queries, while the normalized model is easier when the process of updating information is done [<xref ref-type="bibr" rid="scirp.103821-ref1">1</xref>] [<xref ref-type="bibr" rid="scirp.103821-ref4">4</xref>]. <xref ref-type="table" rid="table5">Table 5</xref> presents a short summary between the dimensional model and the normalized approach in Data Warehouse.</p><table-wrap id="table4" ><label><xref ref-type="table" rid="table4">Table 4</xref></label><caption><title> The comparison between OLTP and OLAP [<xref ref-type="bibr" rid="scirp.103821-ref20">20</xref>] [<xref ref-type="bibr" rid="scirp.103821-ref21">21</xref>] [<xref ref-type="bibr" rid="scirp.103821-ref22">22</xref>]</title></caption><table><tbody><thead><tr><th align="center" valign="middle" ></th><th align="center" valign="middle" >OLTP</th><th align="center" valign="middle" >OLAP</th></tr></thead><tr><td align="center" valign="middle" >Organization</td><td align="center" valign="middle" >It uses workflow per application</td><td align="center" valign="middle" >It uses dimension and business subject</td></tr><tr><td align="center" valign="middle" >Data Retention</td><td align="center" valign="middle" >Short term (2 - 6 months)</td><td align="center" valign="middle" >Long term (2 - 5 years)</td></tr><tr><td align="center" valign="middle" >Database Design and Storage</td><td align="center" valign="middle" >- The design in OLTP is highly normalized. - Gigabytes</td><td align="center" valign="middle" >- The design is typically de-normalized and contains fewer tables - Terabytes</td></tr><tr><td align="center" valign="middle" >Data Integration</td><td align="center" valign="middle" >Minimal or none</td><td align="center" valign="middle" >High, as part of ETL process</td></tr><tr><td align="center" valign="middle" >Application</td><td align="center" valign="middle" >- Real time (short and fast inserts and updates) - write &amp; update - Transactional data - controlling and running fundamental business tasks</td><td align="center" valign="middle" >- Batch load (includes periodic long-running batch jobs that refresh the data) - Reporting, read-only - Spiked usage - Planning, problem solving and decision support</td></tr><tr><td align="center" valign="middle" >Adavnatges</td><td align="center" valign="middle" >- Involves standardized and simple queries that return few records hence, it is faster - A large number of short on-line transaction</td><td align="center" valign="middle" >- Involves complex queries along with aggregations that return a huge amount of data. - OLAP applications are widely used by Data Mining techniques.</td></tr></tbody></table></table-wrap><table-wrap id="table5" ><label><xref ref-type="table" rid="table5">Table 5</xref></label><caption><title> The Dimensional model vs the normalized approach in data warehouse</title></caption><table><tbody><thead><tr><th align="center" valign="middle" ></th><th align="center" valign="middle" >The Dimensional Model (star schema or snowflake)</th><th align="center" valign="middle" >The normalized approach (3NF model)</th></tr></thead><tr><td align="center" valign="middle" >Data warehouse</td><td align="center" valign="middle" >States that the data warehouse should be modeled using a Dimensional Model (star schema or snowflake)</td><td align="center" valign="middle" >States that the data warehouse should be modeled using an E-R model/normalized model</td></tr><tr><td align="center" valign="middle" >Data</td><td align="center" valign="middle" >The data is partitioned into either: facts: which are generally numeric transaction data; dimensions: which are the reference information that gives context to the facts.</td><td align="center" valign="middle" >The data in the data warehouse are stored following database normalization rules. Tables are grouped together by subject areas that reflect general data categories (e.g., data on customers, products, finance, etc.). The normalized structure divides data into entities, which creates several tables in a relational database. Each of the created entities is converted into separate physical tables when the database is implemented.</td></tr><tr><td align="center" valign="middle" >Advantages</td><td align="center" valign="middle" >A key advantage of a dimensional approach is that the data warehouse is easier for the user to understand and to use. The retrieval of data from the data warehouse tends to operate very quickly.</td><td align="center" valign="middle" >The main advantage of this approach is that it is straightforward to add information into the database.</td></tr><tr><td align="center" valign="middle" >Disadvantages</td><td align="center" valign="middle" >The main disadvantage of the dimensional approach is that In order to maintain the integrity of facts and dimensions, loading the data warehouse with data from different operational systems is complicated.</td><td align="center" valign="middle" >A disadvantage of this approach is that, because of the number of tables involved, it can be difficult for users both to join data from different sources into meaningful information and then access the information without a precise understanding of the sources of data and of the data structure of the data warehouse.</td></tr></tbody></table></table-wrap></sec><sec id="s4_3"><title>4.3. Fundamentals of Data Warehouse</title><p>To design the data warehouse model, there are a set of basic fundamentals such as grain, additivity, facts, dimension, and calendar tables. The description of these fundamentals are presented as follow:</p><p>• Grain: It presents to the scale or level of granularity of fact table such that all facts should have the same grain.</p><p>• Additivity: This is the property of numeric facts that can be used in competitions such as averaging, min, max, while the other types of facts are also taken into counted.</p><p>• Facts tables: These tables are connected with dimension tables using one or more foreign keys.</p><p>• Dimension tables: these types of tables are used to execute the required query and generates the reports.</p><p>• Calendar dimension: This table is the basic table to make simple dates and is linked with both fact and dimension tables.</p><p>The quality of the data warehouse has a significant role in the accuracy of data analysis [<xref ref-type="bibr" rid="scirp.103821-ref23">23</xref>]. The main criteria to measure the quality of a data warehouse is described as follows: 1) Access to information should be easy. 2) Recorded data should be consistent and integrating correctly. 3) Data warehouses should be adapted to any change. Finally, 4) data warehouse should get acceptance by end-users [<xref ref-type="bibr" rid="scirp.103821-ref17">17</xref>].</p></sec></sec><sec id="s5"><title>5. Big Data</title><p>In recent years, big data indicates the size, velocity, variety, and diversity of the data sets that are alarmingly growing to create a challenge in storing and analyzing using the conventional database systems. [<xref ref-type="bibr" rid="scirp.103821-ref23">23</xref>]. Recently, many popular data technologies trying to address the challenges of the new Big Data and Internet of Things (IoT) applications that generate data of various sizes are currently ranging from terabytes to petabytes and also considered as a synthetic data generator for structured, semi-structured, and unstructured data [<xref ref-type="bibr" rid="scirp.103821-ref4">4</xref>]. Thus, there are different types of data sources in many domains which create a huge volume of data (Big Data), for instance, video archives, sensor data, Internet text and documents, social networks, tweets, blogs, log files, biochemical, medical records, and transactional records [<xref ref-type="bibr" rid="scirp.103821-ref3">3</xref>]. Recently, the datasets in big data are characterized by n Vs whereby the n refers to the 9 characteristics of the datasets such as Veracity, Variety, Velocity, Volume, Validity, Variability, Volatility, Visualization and Value [<xref ref-type="bibr" rid="scirp.103821-ref24">24</xref>].</p><sec id="s5_1"><title>5.1. Hadoop</title><p>Apache Hadoop is an open-source distributed framework that is widely used for parallel storage, and efficacy in processing the big data on the cluster of machines using high-level programming languages. Hadoop modules have many features to help developers and researchers such as graphical user interfaces, simple administration tools and provide high-level languages such as Java and Python. Moreover, the advantages of this framework are high availability, fault tolerance and scalability to process petabytes of data by the Hadoop cluster. The Hadoop cluster is a group of thousands of computers connected with each other to store Big Data and to run the MapReduce programs in parallel. Generally, users can execute remotely jobs using the Hadoop cluster [<xref ref-type="bibr" rid="scirp.103821-ref25">25</xref>] [<xref ref-type="bibr" rid="scirp.103821-ref26">26</xref>].</p><p><xref ref-type="fig" rid="fig3">Figure 3</xref> shows the principal components of the Hadoop framework in a</p><p>master/slave architecture that includes MapReduce, Hadoop Distributed File System (HDFS), YARN, and some other additional modules [<xref ref-type="bibr" rid="scirp.103821-ref27">27</xref>] [<xref ref-type="bibr" rid="scirp.103821-ref28">28</xref>]. Basically, HDFS is considered the primary storage of large datasets in Hadoop and distributed file system management. While, MapReduce is a general programming framework designed specifically for parallel processing the big size of unstructured data as shown in <xref ref-type="fig" rid="fig4">Figure 4</xref>.</p></sec><sec id="s5_2"><title>5.2. HDFS Architecture</title><p>The general structure of the HDFS system is simply based on the main communication master/slave framework. The HDFS cluster has two essential elements namely NameNode, and DataNode [<xref ref-type="bibr" rid="scirp.103821-ref29">29</xref>]. Moreover, the cluster has specific NameNode which is responsible for storing metadata, control access to files stored in DataNode and manages the file system. The cluster can have a limited number of DataNodes which are used to store the data as shown in <xref ref-type="fig" rid="fig5">Figure 5</xref>. Generally, the big files in HDFS are split into small blocks and stored in a specific number of Datanodes on the cluster. The default size of blocks is a little bit large (64 MB or 128 MB), to decrease the running time. In general, the replication factor determines how many copies of one block are saved in different Datanodes. If the replication factor is 2 the block should be on two Datanodes. <xref ref-type="fig" rid="fig5">Figure 5</xref> demonstrates a simple example of large files that are stored in the Namenode and four DataNodes. The main advantages of HDFS are 1) it is developed to be more scalable and has the capability of fault tolerance, 2) it saves big files for future use, these files split into blocks which replicated on two or more DataNodes, 3) the performance of HDFS is scaled by increasing the number of Data Nodes [<xref ref-type="bibr" rid="scirp.103821-ref30">30</xref>].</p></sec><sec id="s5_3"><title>5.3. MapReduce</title><p>Recently, Dean and Ghemawat developed the most important programming model MapReduce which is widely used for the Cloud computing framework</p><p>and is especially supported by Google under the Apache Hadoop project [<xref ref-type="bibr" rid="scirp.103821-ref31">31</xref>]. In general, the developers address big issues by writing the MapReduce programs where the input files split into small chunks to adapt to the HDFS system and the parallel computations [<xref ref-type="bibr" rid="scirp.103821-ref4">4</xref>].</p><p>The “Map” and “Reduce” classes are the fundamental components of MapReduce. <xref ref-type="fig" rid="fig6">Figure 6</xref> shows the main steps involved in the process flow of MapReduce. In the Map step, the input data divides into a number of small chunks such that all sub-files are allocated parallelly to different mappers. Moreover, the basic idea of the mapper method is to read the corresponding chunk as a bench of keys and their value pairs and similarly, the results of this method are a set of (key, value) pairs. The shuffles and sorts phase comes immediately after the map phase and the input for this stage is the result of the mapper functions, and it produces the (key, value) pairs that assigned to the reducers. In the final process, the output (key, value) pairs of previous tasks are grouped based on the key and assigned to the reducers to produce the final results which are stored in HDFS [<xref ref-type="bibr" rid="scirp.103821-ref32">32</xref>]. The Hadoop framework has a JobTracker service that assigns the tasks of MapReduce to clearly defined nodes within the Hadoop cluster, specifically the nodes that have the data. On the other hand, the TaskTracker is a node that accepts</p><p>operations from a JobTracker service such as Map, Reduce and Shuffle. During the execution of MapReduce job, JobTracker and TaskTracker tools are applicable to schedule, monitor, and restart processes in case of failing as shown in <xref ref-type="fig" rid="fig6">Figure 6</xref>.</p></sec><sec id="s5_4"><title>5.4. Hadoop Ecosystem</title><p>This section describes the Hadoop Ecosystem that has many modules, MapReduce and HDFS that are used in the ETL operations [<xref ref-type="bibr" rid="scirp.103821-ref33">33</xref>]. Moreover, Hadoop has different software for particular tasks such as Apache Spark, Apache Hive are used for data processing and Apache Oozie, Apache Drill for jobs orchestration, while Apache Spark and Apache Mahout are applied to build high-performance machine learning models [<xref ref-type="bibr" rid="scirp.103821-ref34">34</xref>].</p><p>1) Apache Sqoop</p><p>It is a command-line-based tool and an open-source under the Apache software, that allows for efficient importing of records from relational database tables to HDFS directories as database tables in the Hadoop framework. Therefore, the essential tasks of Sqoop are transferring stored data from Oracle or MySQL database into the HDFS database, then the MapReduce program executed on the imported data, and finally exports data in database tables or Data warehouse [<xref ref-type="bibr" rid="scirp.103821-ref3">3</xref>] [<xref ref-type="bibr" rid="scirp.103821-ref35">35</xref>] [<xref ref-type="bibr" rid="scirp.103821-ref36">36</xref>]. There are many versions of Hadoop connectors provided by famous EDW vendors such as IBM Db2 [<xref ref-type="bibr" rid="scirp.103821-ref37">37</xref>], HP Vertica [<xref ref-type="bibr" rid="scirp.103821-ref38">38</xref>] and Oracle [<xref ref-type="bibr" rid="scirp.103821-ref39">39</xref>] bulk-loads data from Hadoop to Oracle Database.</p><p>2) Apache Hive</p><p>This is an open-source project and an efficient query language that simplifies the development of applications using the MapReduce framework. Apache Hive has two main components namely Hive Query Language (HiveQL) and Hive metastore [<xref ref-type="bibr" rid="scirp.103821-ref3">3</xref>] [<xref ref-type="bibr" rid="scirp.103821-ref40">40</xref>] [<xref ref-type="bibr" rid="scirp.103821-ref41">41</xref>]. HiveQL is a query language for Hive software and used to facilitate writing queries data stored in Apache HBase and HDFS. While, the Hive metastore is the master repository of Hive metadata which is used to store metadata about data files, blocks in the HDFS NameNode [<xref ref-type="bibr" rid="scirp.103821-ref42">42</xref>]. Moreover, Hive is used in data warehousing implementation such that it facilitates the operations such as reading, writing, and managing large files stored in HDFS [<xref ref-type="bibr" rid="scirp.103821-ref43">43</xref>].</p><p>3) ODBC/JDBC Connectors</p><p>Basically, any version of Apache Hadoop software provides a JDBC/ODBC connectors for HBase and Hive and interface to make a connection with Business Intelligence and different visualization tools. Apache Hive ODBC and Hive JDBC drivers are software components used for translating from traditional SQL queries into HiveQL commands so that it can execute upon the data in the Apache Hadoop/Hive distributions [<xref ref-type="bibr" rid="scirp.103821-ref3">3</xref>] [<xref ref-type="bibr" rid="scirp.103821-ref44">44</xref>].</p></sec></sec><sec id="s6"><title>6. Implementation</title><p>The commonly used methods for building a meteorological data warehouse are the classical DW, big data tools such as Hadoop, Sqoop, Hive, Pig, and Spark. Finally, the hybrid DW model merges traditional DW with big data. In this work, the proposed data warehouse model constructed based on the third approach such that the following big data tools are applied to implement the model as shown in <xref ref-type="fig" rid="fig7">Figure 7</xref>. The list of tools of big data are described as follows:</p><p>• RDBMS: It is used to store the records of the collected data.</p><p>• SQOOP: It is used to import records from RDBMS into HDFS, and to transfer the final results of aggregation operations to the data warehouse.</p><p>• HDFS: It is used to save big files.</p><p>• Hive: It is used to create tables and databases on HDFS. Moreover, it is has the ability to perform join, partition, merge, or aggregation operations.</p><sec id="s6_1"><title>6.1. High Dimensional Schema for NCDC Dataset</title><p>The data-warehouse schema is a logical description of the whole database [<xref ref-type="bibr" rid="scirp.103821-ref8">8</xref>]. The main components of data-warehouse include fact tables, dimension tables, and their dependencies. A dimension is simply a row and column of the high-dimensional table that contains the number of samples and their corresponding</p><p>attributes. There are some main operations that can be applied to the high-dimensional table such as grouping, filtering and labeling. Moreover, the samples in dimension tables are identified by a unique Key and each column represents a range of values that are found by measuring using given units.</p><p>The available data set is collected for each year as compressed files. <xref ref-type="fig" rid="fig8">Figure 8</xref> shows the design of the relational schema for the collected data. For instance, the weather_station table contains information about each station that has columns namely station_STN, latitude, longitude, ..., etc. Similarly, the countries table has information about each country and it is linked to the states table. Finally, each parameter key in the Parameter table is used to connect with the corresponding table.</p><p><xref ref-type="fig" rid="fig9">Figure 9</xref> demonstrates the proposed star schema on NCDC weather datasets that contains two types of tables and it presents one fact table and 7-dimensional tables. The primary task of OLTP schema is used for performing the preprocessing stage such as reduce redundancy, normalize the data and check for its integrity. In addition, it also has another advantage such as it is possible to create, update or delete any column in a particular table. The weather fact table has seven dimension tables, namely STATION, TIME, TEMPERATURE, PARAMETER, SLP, DEWP, and PRCP. In general, the fact table consists of a number of primary keys and many keys that refer to their corresponding dimension table. On the other hand, the primary key in the dimension table has the same name as the table that contains observations collected from a particular station. <xref ref-type="table" rid="table6">Table 6</xref> describes the dimension tables including all the relevant attributes.</p><table-wrap id="table6" ><label><xref ref-type="table" rid="table6">Table 6</xref></label><caption><title> The dimension tables and attributes in OLTP Schema</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >TABLE</th><th align="center" valign="middle" >Attributes</th></tr></thead><tr><td align="center" valign="middle" >Countries</td><td align="center" valign="middle" >Country_key, Country_Name, Short_Name</td></tr><tr><td align="center" valign="middle" >States</td><td align="center" valign="middle" >State_key, Long_Name, Short_Name, Country_key</td></tr><tr><td align="center" valign="middle" >Cities</td><td align="center" valign="middle" >City_key, City_Name, Short_Name, State_key</td></tr><tr><td align="center" valign="middle" >Stations</td><td align="center" valign="middle" >Station_STN, Station_wban, Station_name, Country, call_st, latitude, longitude, Elevation, BEGIN_date, END_date</td></tr><tr><td align="center" valign="middle" >Parameters</td><td align="center" valign="middle" >Parameter_key, Station_STN, Station_wban, Yearmody, Temperature, DEWP, SLP, STP, VISIB, WDSP, MXSPD, GUST, MAX, MIN, PRCP, SNDP, I_FOG, I_RAIN_DZL, I_SNW_ICE, I_HAIL, I_THUNDER, I_TDO_FNL</td></tr><tr><td align="center" valign="middle" >Daily_Measures_Temperature</td><td align="center" valign="middle" >Temperature_key, Temperature_CNT, Month, Year, Day1, Day2, Day3, ..., Day31</td></tr><tr><td align="center" valign="middle" >Daily_Measures_DEWP</td><td align="center" valign="middle" >DEWP_key, DEWP_CNT, Month, Year, Day1, Day2, Day3, ..., Day31</td></tr><tr><td align="center" valign="middle" >Daily_Measures_SLP</td><td align="center" valign="middle" >SLP_key, SLP_CNT, Month, Year, Day1, Day2, Day3, ..., Day31</td></tr><tr><td align="center" valign="middle" >Daily_Measures_STP</td><td align="center" valign="middle" >STP_key, STP_CNT, Month, Year, Day1, Day2, Day3, ..., Day31</td></tr><tr><td align="center" valign="middle" >Daily_Measures_VISIB</td><td align="center" valign="middle" >VISIB_key, VISIB_CNT, Month, Year, Day1, Day2, Day3, ..., Day31</td></tr><tr><td align="center" valign="middle" >Daily_Measures_WDSP</td><td align="center" valign="middle" >WDSP_key, WDSP_CNT, Month, Year, Day1, Day2, Day3, ..., Day31</td></tr><tr><td align="center" valign="middle" >Daily_Measures_MXSPD</td><td align="center" valign="middle" >MXSPD_key, Month, Year, Day1, Day2, Day3, ..., Day31</td></tr><tr><td align="center" valign="middle" >Daily_Measures_GUST</td><td align="center" valign="middle" >GUST_key, Month, Year, Day1, Day2, Day3, ..., Day31</td></tr><tr><td align="center" valign="middle" >Daily_Measures_PRCP</td><td align="center" valign="middle" >PRCP_key, PRCP_FLAG, Month, Year, Day1, Day2, Day3, ..., Day31</td></tr></tbody></table></table-wrap></sec><sec id="s6_2"><title>6.2. The Proposed Hadoop Data Warehouse Model</title><p><xref ref-type="fig" rid="fig1">Figure 1</xref>0 presents the high level conceptual model for Hadoop data warehouse. The first step in the building of such a model is to import data from OLTP tables to HDFS. For instance, the weather_fact, parameter_fact and the rest of the tables are imported using the Sqoop tool. The Sqoop job for any table from OLTP imports all data and overwrite the existing old table. To solve this problem,</p><p>only the import jobs for each table transfer new records and consequently updates only the current records. Finally, the aggregation operations such as max, min, average, and count are done for the new version of the updated tables.</p><p><xref ref-type="fig" rid="fig1">Figure 1</xref>1 illustrates the simplified design of the proposed Hadoop data warehouse model. This design composes of three main tables namely Countries, Parameters, and Weather_Stations. However, the other tables from OLTP schema have been denormalized into one of these three main tables, for instance, the daily_measure_VISIB, daily_measure_PRCP, daily_measure_DEWP, and daily_measure_SLP tables have been denormalized into Parameters.</p><p>There are many reasons to use the Hadoop framework in the data warehouse model, it has the capability to store the same records such as OLTP, save large files that can be distributed over HDFS and both DW and Hadoop are allowed the data to be moved between them efficiently. However, the way to store data in Hadoop is totally different from OLTP. In OLTP, it is allowed to update data by one record at a time and only one operation is executed such as insert/update/ delete. So, the OLTP schema allows the following operations for the values of the records: updates/deletes/inserts to change these values. On the other hand, for the HDFS system in Hadoop, it is not allowed to update the values of records, it allows only deleting old data and replaces it by the new ones. To solve this problem in Hadoop, it is essential to create two versions of tables such that the second table is used to append only new records. For instance, the weather_stations table has another version Station_history which is an append-only table and is used to save the history of all stations. Sequentially, once the import job executed, the new data imported from OLTP appended to the end of the Station_history table and the final data are stored in weather_stations.</p><p>To make this clear, the following example illustrates the concept of insert/ update/delete in Hadoop, <xref ref-type="table" rid="table7">Table 7</xref> presents a sample of Stations table in the OLTP.</p><p>The same data of Stations table are stored in Hadoop as shown in <xref ref-type="table" rid="table8">Table 8</xref>.</p><p>In case of any update operation, it is done in the Weather_stations table such as an add a new station, so the data of this table needs to be updated in Hadoop and have only full records. Thus, the HDFS system has two versions of the same</p><table-wrap id="table7" ><label><xref ref-type="table" rid="table7">Table 7</xref></label><caption><title> A sample of rows of station table in OLTP database</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >USAF</th><th align="center" valign="middle" >WBAN</th><th align="center" valign="middle" >STATION NAME</th><th align="center" valign="middle" >CTRY</th><th align="center" valign="middle" >LAT</th><th align="center" valign="middle" >LON</th><th align="center" valign="middle" >ELEV(M)</th><th align="center" valign="middle" >BEGIN</th><th align="center" valign="middle" >END</th></tr></thead><tr><td align="center" valign="middle" >410090</td><td align="center" valign="middle" >99999</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" >19851001</td><td align="center" valign="middle" >20110614</td></tr><tr><td align="center" valign="middle" >410060</td><td align="center" valign="middle" >99999</td><td align="center" valign="middle" >MUWAIH</td><td align="center" valign="middle" >SA</td><td align="center" valign="middle" >+22.433</td><td align="center" valign="middle" >+041.750</td><td align="center" valign="middle" >+0971.0</td><td align="center" valign="middle" >19851001</td><td align="center" valign="middle" >20110614</td></tr><tr><td align="center" valign="middle" >410080</td><td align="center" valign="middle" >99999</td><td align="center" valign="middle" >ZULM</td><td align="center" valign="middle" >SA</td><td align="center" valign="middle" >+22.717</td><td align="center" valign="middle" >+042.167</td><td align="center" valign="middle" >+0870.0</td><td align="center" valign="middle" >20041015</td><td align="center" valign="middle" >20041015</td></tr><tr><td align="center" valign="middle" >410100</td><td align="center" valign="middle" >99999</td><td align="center" valign="middle" >LAYLA</td><td align="center" valign="middle" >SA</td><td align="center" valign="middle" >+22.333</td><td align="center" valign="middle" >+046.733</td><td align="center" valign="middle" >+0543.0</td><td align="center" valign="middle" >19910408</td><td align="center" valign="middle" >20030329</td></tr><tr><td align="center" valign="middle" >410140</td><td align="center" valign="middle" >99999</td><td align="center" valign="middle" >OBAYLAH (AUT)</td><td align="center" valign="middle" >SA</td><td align="center" valign="middle" >+22.217</td><td align="center" valign="middle" >+050.883</td><td align="center" valign="middle" >+0588.0</td><td align="center" valign="middle" >19830201</td><td align="center" valign="middle" >20030315</td></tr><tr><td align="center" valign="middle" >410160</td><td align="center" valign="middle" >99999</td><td align="center" valign="middle" >SHAWALAH</td><td align="center" valign="middle" >SA</td><td align="center" valign="middle" >+22.517</td><td align="center" valign="middle" >+054.050</td><td align="center" valign="middle" >+0468.0</td><td align="center" valign="middle" >19830402</td><td align="center" valign="middle" >20050526</td></tr></tbody></table></table-wrap><table-wrap id="table8" ><label><xref ref-type="table" rid="table8">Table 8</xref></label><caption><title> Stations table in Hadoop before update</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >USAF</th><th align="center" valign="middle" >WBAN</th><th align="center" valign="middle" >STATION NAME</th><th align="center" valign="middle" >CTRY</th><th align="center" valign="middle" >LAT</th><th align="center" valign="middle" >LON</th><th align="center" valign="middle" >ELEV(M)</th><th align="center" valign="middle" >BEGIN</th><th align="center" valign="middle" >END</th></tr></thead><tr><td align="center" valign="middle" >410090</td><td align="center" valign="middle" >99999</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" >19851001</td><td align="center" valign="middle" >20110614</td></tr><tr><td align="center" valign="middle" >410060</td><td align="center" valign="middle" >99999</td><td align="center" valign="middle" >MUWAIH</td><td align="center" valign="middle" >SA</td><td align="center" valign="middle" >+22.433</td><td align="center" valign="middle" >+041.750</td><td align="center" valign="middle" >+0971.0</td><td align="center" valign="middle" >19851001</td><td align="center" valign="middle" >20110614</td></tr><tr><td align="center" valign="middle" >410080</td><td align="center" valign="middle" >99999</td><td align="center" valign="middle" >ZULM</td><td align="center" valign="middle" >SA</td><td align="center" valign="middle" >+22.717</td><td align="center" valign="middle" >+042.167</td><td align="center" valign="middle" >+0870.0</td><td align="center" valign="middle" >20041015</td><td align="center" valign="middle" >20041015</td></tr><tr><td align="center" valign="middle" >410100</td><td align="center" valign="middle" >99999</td><td align="center" valign="middle" >LAYLA</td><td align="center" valign="middle" >SA</td><td align="center" valign="middle" >+22.333</td><td align="center" valign="middle" >+046.733</td><td align="center" valign="middle" >+0543.0</td><td align="center" valign="middle" >19910408</td><td align="center" valign="middle" >20030329</td></tr><tr><td align="center" valign="middle" >410140</td><td align="center" valign="middle" >99999</td><td align="center" valign="middle" >OBAYLAH (AUT)</td><td align="center" valign="middle" >SA</td><td align="center" valign="middle" >+22.217</td><td align="center" valign="middle" >+050.883</td><td align="center" valign="middle" >+0588.0</td><td align="center" valign="middle" >19830201</td><td align="center" valign="middle" >20030315</td></tr><tr><td align="center" valign="middle" >410160</td><td align="center" valign="middle" >99999</td><td align="center" valign="middle" >SHAWALAH</td><td align="center" valign="middle" >SA</td><td align="center" valign="middle" >+22.517</td><td align="center" valign="middle" >+054.050</td><td align="center" valign="middle" >+0468.0</td><td align="center" valign="middle" >19830402</td><td align="center" valign="middle" >20050526</td></tr></tbody></table></table-wrap><p>table in order to address the issue of update data in HDFS. On the other hand, Station_history includes all the imported records from the same version of this table in the OLTP schema as shown in <xref ref-type="table" rid="table8">Table 8</xref>. Similarly, <xref ref-type="table" rid="table9">Table 9</xref> demonstrates the final version of Weather_stations table that has only the complete rows.</p></sec><sec id="s6_3"><title>6.3. The Aggregates Design</title><p>The primary task of a data warehouse is to make great flexibility and efficiency for the query process. So, aggregations are used to decrease the query time using pre-computed summary data. In this section, the following tasks namely ingestion, aggregation, and data export are briefed.</p><p>1) Ingestion</p><p>The first stage of the conventional ETL approach is to transfer the stored data from the OLTP schema to a data warehouse. Generally, the Sqoop tool is used for ingesting data from OLTP into HDFS in Hadoop as shown in <xref ref-type="fig" rid="fig1">Figure 1</xref>2. Moreover, to ingest a small data from OLTP into Hadoop is require one task using Sqoop. On the other hand, if the data size is large, then it will need many tasks or repeat the Sqoop job many times as shown in <xref ref-type="fig" rid="fig1">Figure 1</xref>3. Moreover, the list of files transferred after Sqoop task done is shown in <xref ref-type="fig" rid="fig1">Figure 1</xref>4.</p><p>2) Aggregation</p><p>The aggregation operation is one of the most time-consuming operations to perform in the classical database. Thus, to reduce the running time of ETL and aggregation operations are implemented in Hadoop. As the Hadoop framework</p><table-wrap id="table9" ><label><xref ref-type="table" rid="table9">Table 9</xref></label><caption><title> Stations table in Hadoop after update</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >USAF</th><th align="center" valign="middle" >WBAN</th><th align="center" valign="middle" >STATION NAME</th><th align="center" valign="middle" >CTRY</th><th align="center" valign="middle" >LAT</th><th align="center" valign="middle" >LON</th><th align="center" valign="middle" >ELEV(M)</th><th align="center" valign="middle" >BEGIN</th><th align="center" valign="middle" >END</th></tr></thead><tr><td align="center" valign="middle" >410060</td><td align="center" valign="middle" >99999</td><td align="center" valign="middle" >MUWAIH</td><td align="center" valign="middle" >SA</td><td align="center" valign="middle" >+22.433</td><td align="center" valign="middle" >+041.750</td><td align="center" valign="middle" >+0971.0</td><td align="center" valign="middle" >19851001</td><td align="center" valign="middle" >20110614</td></tr><tr><td align="center" valign="middle" >410080</td><td align="center" valign="middle" >99999</td><td align="center" valign="middle" >ZULM</td><td align="center" valign="middle" >SA</td><td align="center" valign="middle" >+22.717</td><td align="center" valign="middle" >+042.167</td><td align="center" valign="middle" >+0870.0</td><td align="center" valign="middle" >20041015</td><td align="center" valign="middle" >20041015</td></tr><tr><td align="center" valign="middle" >410100</td><td align="center" valign="middle" >99999</td><td align="center" valign="middle" >LAYLA</td><td align="center" valign="middle" >SA</td><td align="center" valign="middle" >+22.333</td><td align="center" valign="middle" >+046.733</td><td align="center" valign="middle" >+0543.0</td><td align="center" valign="middle" >19910408</td><td align="center" valign="middle" >20030329</td></tr><tr><td align="center" valign="middle" >410140</td><td align="center" valign="middle" >99999</td><td align="center" valign="middle" >OBAYLAH (AUT)</td><td align="center" valign="middle" >SA</td><td align="center" valign="middle" >+22.217</td><td align="center" valign="middle" >+050.883</td><td align="center" valign="middle" >+0588.0</td><td align="center" valign="middle" >19830201</td><td align="center" valign="middle" >20030315</td></tr><tr><td align="center" valign="middle" >410160</td><td align="center" valign="middle" >99999</td><td align="center" valign="middle" >SHAWALAH</td><td align="center" valign="middle" >SA</td><td align="center" valign="middle" >+22.517</td><td align="center" valign="middle" >+054.050</td><td align="center" valign="middle" >+0468.0</td><td align="center" valign="middle" >19830402</td><td align="center" valign="middle" >20050526</td></tr></tbody></table></table-wrap><p>is distributed and can be run tasks in parallel, the aggregation tasks are executed faster than the OLTP database. Hive and Impala are the most popular tools known for aggregation over Hadoop. Finally, aggregation operations such as average, max, count, and summation are implemented cheaper and scalable using Hadoop that store a huge number of records. For example, the following Hive code calculates the records count and the average temperature for each station, respectively.</p><p>3) Data Export</p><p>This is the next step after ingestion and aggregation to transfer the cleaned data from HDFS to the real data warehouse. In this stage, the preferred tool is Sqoop that can be applied to shift the data from the Hadoop system to the traditional data warehouse with insert and update operations. For instance, the following Sqoop code exports data from avg_temp table in weather_dwh database stored in Hadoop to a database table.</p></sec></sec><sec id="s7"><title>7. Conclusion</title><p>In this work, we explored the different notions regarding big data, data warehouse, and presented in detail the proposed data warehouse for weather data that typically constructed on top of the strong Hadoop system. Moreover, the flexible meteorological data warehouse successfully produced using the suggested star schema model and different Big Data software. The presented schema includes all necessary fact and dimension tables in order to deal with scale and efficient analytical models. Furthermore, the suggested data warehouse model optimized for NCDC that was available. The advantages of this model are the possibility to add new variables; different queries are easily done in a flexible way, and similarly, it easy to extract data. Finally, the proposed model is flexible, adaptable, and quite qualified for the rapid increase of data from different weather variables without any major change.</p></sec><sec id="s8"><title>Conflicts of Interest</title><p>The author declares no conflicts of interest regarding the publication of this paper.</p></sec><sec id="s9"><title>Cite this paper</title><p>Hashim, H. (2020) Hybrid Warehouse Model and Solutions for Climate Data Analysis. Journal of Computer and Communications, 8, 75-98. https://doi.org/10.4236/jcc.2020.810008</p></sec></body><back><ref-list><title>References</title><ref id="scirp.103821-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">Doreswamy, I. and Manjunatha, B.R. (2017) Hybrid Data Warehouse Model for Climate Big Data Analysis. 2017 International Conference on Circuit, Power and Computing Technologies (ICCPCT) IEEE, Kollam, 20-21 April 2017. https://doi.org/10.1109/ICCPCT.2017.8074229</mixed-citation></ref><ref id="scirp.103821-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">N. Climatic Data Center (NCDC) (2016) Noaa’s National Centers for Environmental Information (NCEI). https://www.ncdc.noaa.gov/data-access/land-based-station-data/land-based-datasets</mixed-citation></ref><ref id="scirp.103821-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">I. Corporation (2013) Extract, Transform, and Load Big Data with Apache Hadoop. Intel. Corporation, Vol. White Paper Big Data Analytics, 1-5. https://software.intel.com/content/www/us/en/develop/download/white-paper-extract-transform-and-load-big-data-with-apache-hadoop.html</mixed-citation></ref><ref id="scirp.103821-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">Amr Awadallah, D.G. (2013) Hadoop and the Data Warehouse: When to Use Which. Cloudera Corporation and Teradata Corporation, Vol. White Paper, 1-19. https://kannandreams.files.wordpress.com/2013/10/hadoop-use-case-1.pdf</mixed-citation></ref><ref id="scirp.103821-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">Alexander, I., Rassetiadi, R., Garcia, S., Girsang, A.S. and Isa, S.M. (2018) Business Solution for Choosing Products Using Data Warehouse in Payment Solution. 2018 Indonesian Association for Pattern Recognition International Conference, Jakarta, 7-8 September 2018. https://doi.org/10.1109/INAPR.2018.8627028</mixed-citation></ref><ref id="scirp.103821-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">Chen, D.-Q., Wang, W.-Y. and Yang, H.-K. (2010) Application Research on Data Warehouse of Hydrological Data Comprehensive Analysis. 2010 3rd International Conference on Computer Science and Information Technology, Vol. 9, 140-143. https://doi.org/10.1109/ICCSIT.2010.5565123</mixed-citation></ref><ref id="scirp.103821-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">Kalra, G. and Steiner, D. (2005) Weather Data Warehouse: An Agent-Based Data Warehousing System. Proceedings of the 38th Annual Hawaii International Conference on System Sciences, Big Island, 6 January 2005.</mixed-citation></ref><ref id="scirp.103821-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">Ma, N., Yuan, M., Bao, Y., Jin, Z. and Zhou, H. (2010) The Design of Meteorological Data Warehouse and Multidimensional Data Report. 2010 Second International Conference on Information Technology and Computer Science, Kiev, 24-25 July 2010, 280-283. https://doi.org/10.1109/ITCS.2010.75</mixed-citation></ref><ref id="scirp.103821-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">Tian, Y. (2019) Hybrid Data Warehouse. In: Encyclopedia of Big Data Technologies, Springer International Publishing, Berlin, 979. https://doi.org/10.1007/978-3-319-77525-8_100167</mixed-citation></ref><ref id="scirp.103821-ref10"><label>10</label><mixed-citation publication-type="other" xlink:type="simple">Zhang, Z. and Li, J. (2020) Big Climate Data. In: Big Data Mining for Climate Change, Elsevier, Amsterdam, 1-18. https://doi.org/10.1016/B978-0-12-818703-6.00006-4</mixed-citation></ref><ref id="scirp.103821-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">Doreswamy and Harishkumar, K. (2018) Multidimensional Data Model for Air Pollution Data Analysis. 2018 International Conference on Advances in Computing, Communications and Informatics (ICACCI) IEEE, Bangalore, 19-22 September 2018. https://doi.org/10.1109/ICACCI.2018.8554621</mixed-citation></ref><ref id="scirp.103821-ref12"><label>12</label><mixed-citation publication-type="other" xlink:type="simple">Chen, S. (2010) Cheetah: A High Performance, Custom Data Warehouse on Top of Mapreduce. Proceedings of the VLDB Endowment, 3, 1459-1468. https://doi.org/10.14778/1920841.1921020</mixed-citation></ref><ref id="scirp.103821-ref13"><label>13</label><mixed-citation publication-type="other" xlink:type="simple">Dimri, P. and Gunwant, H. (2012) Conceptual Model for Developing Meteorological Data Warehouse in Uttarakhand—A Review. Journal of Information and Operations Management, 3, 107-110.</mixed-citation></ref><ref id="scirp.103821-ref14"><label>14</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>José Torres-Jiménez</surname><given-names> J.F.G. </given-names></name>,<etal>et al</etal>. (<year>2004</year>)<article-title>A Data Warehouse for Weather Information: A Pattern Recognition Solution for Climatic Conditions in México</article-title><source> Proceedings of the Sixth International Conference on Enterprise Information Systems</source><volume> 1</volume>,<fpage> 562</fpage>-<lpage>565</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.103821-ref15"><label>15</label><mixed-citation publication-type="other" xlink:type="simple">Wadkar, S. and Siddalingaiah, M. (2014) Data Warehousing Using Hadoop. In: Pro Apache Hadoop, Apress, New York, 217-239. https://doi.org/10.1007/978-1-4302-4864-4_10</mixed-citation></ref><ref id="scirp.103821-ref16"><label>16</label><mixed-citation publication-type="other" xlink:type="simple">Duque-Méndez, N.D., Orozco-Alzate, M. and Vélez, J.J. (2014) Hydro-Meteorological Data Analysis Using OLAP Techniques. DYNA, 81, 160-167. https://doi.org/10.15446/dyna.v81n185.37700</mixed-citation></ref><ref id="scirp.103821-ref17"><label>17</label><mixed-citation publication-type="other" xlink:type="simple">Vuong, N.-A.L.-K., Ngo, M. and Kechadi, M.-T. (2018) An Efficient Data Warehouse for Crop Yield Prediction. Proceedings of the 14th International Conference on Precision Agriculture, Montreal, 24-27 June 2018. https://researchrepository.ucd.ie/bitstream/10197/10118/2/1807.00035v1.pdf</mixed-citation></ref><ref id="scirp.103821-ref18"><label>18</label><mixed-citation publication-type="other" xlink:type="simple">Narayan, R. and Mehta, G. (2020) Design of Customer Information Management System. In: Data Communication and Networks, Springer, Berlin, 195-216. https://doi.org/10.1007/978-981-15-0132-6_13</mixed-citation></ref><ref id="scirp.103821-ref19"><label>19</label><mixed-citation publication-type="other" xlink:type="simple">Jin, Z.-H., Shi, H., Hu, Y.-X., Zha, L. and Lu, X. (2020) Cirrodata: Yet Another Sqlon-Hadoop Data Analytics Engine with High Performance. Journal of Computer Science and Technology, 35, 194-208. https://doi.org/10.1007/s11390-020-9536-z</mixed-citation></ref><ref id="scirp.103821-ref20"><label>20</label><mixed-citation publication-type="other" xlink:type="simple">Conn, S. (2005) OLTP and OLAP Data Integration: A Review of Feasible Implementation Methods and Architectures for Real Time Data Analysis. Proceedings. IEEE Southeast Con, Ft. Lauderdale, 8-10 April 2005, 515-520.</mixed-citation></ref><ref id="scirp.103821-ref21"><label>21</label><mixed-citation publication-type="other" xlink:type="simple">Chohan, M.A. and Javed, M.Y. (2010) OLAP and OLTP Data Integration for Operational Level Decision Making. 2010 International Conference on Networking and Information Technology, Manila, 11-12 June 2010, 493-496. https://doi.org/10.1055/s-0030-1258076</mixed-citation></ref><ref id="scirp.103821-ref22"><label>22</label><mixed-citation publication-type="other" xlink:type="simple">Giceva, J. and Sadoghi, M. (2019) Hybrid OLTP and OLAP. In: Encyclopedia of Big Data Technologies, Springer International Publishing, Berlin, 979-986. https://doi.org/10.1007/978-3-319-77525-8_179</mixed-citation></ref><ref id="scirp.103821-ref23"><label>23</label><mixed-citation publication-type="other" xlink:type="simple">Aftab, U. and Siddiqui, G.F. (2018) Big Data Augmentation with Data Warehouse: A Survey. 2018 IEEE International Conference on Big Data (Big Data), Seattle, 10-13 December 2018, 2785-2794. https://doi.org/10.1109/BigData.2018.8622206</mixed-citation></ref><ref id="scirp.103821-ref24"><label>24</label><mixed-citation publication-type="other" xlink:type="simple">Owais, S.S. and Hussein, N.S. (2016) Extract Five Categories CPIVW from the 9v’s Characteristics of the Big Data. International Journal of Advanced Computer Science &amp; Applications, 1, 254-258. https://doi.org/10.14569/IJACSA.2016.070337</mixed-citation></ref><ref id="scirp.103821-ref25"><label>25</label><mixed-citation publication-type="other" xlink:type="simple">Dean, J. and Ghemawat, S. (2008) Mapreduce: Simplified Data Processing on Large Clusters. Communications of the ACM, 51, 107-113. https://doi.org/10.1145/1327452.1327492</mixed-citation></ref><ref id="scirp.103821-ref26"><label>26</label><mixed-citation publication-type="other" xlink:type="simple">Nazari, E., Shahriari, M.H. and Tabesh, H. (2019) Big Data Analysis in Healthcare: Apache Hadoop, Apache Spark and Apache Flink. Frontiers in Health Informatics, 8, 14. https://doi.org/10.30699/fhi.v8i1.180</mixed-citation></ref><ref id="scirp.103821-ref27"><label>27</label><mixed-citation publication-type="other" xlink:type="simple">Anthony, B., et al. (2016) Ecosystem at Large: Hadoop with Apache Bigtop. In: Professional Hadoop, John Wiley &amp; Sons Inc., Hoboken, 141-160. https://doi.org/10.1002/9781119281320.ch7</mixed-citation></ref><ref id="scirp.103821-ref28"><label>28</label><mixed-citation publication-type="other" xlink:type="simple">Vavilapalli, V.K., Seth, S., Saha, B., Curino, C., O’Malley, O., Radia, S., Reed, B., Baldeschwieler, E., Murthy, A.C., Douglas, C., Agarwal, S., Konar, M., Evans, R., Graves, T., Lowe, J. and Shah, H. (2013) Apache Hadoop YARN. In: Proceedings of the 4th annual Symposium on Cloud Computing, ACM Press, New York. https://doi.org/10.1145/2523616.2523633</mixed-citation></ref><ref id="scirp.103821-ref29"><label>29</label><mixed-citation publication-type="other" xlink:type="simple">Tian, Y., &amp;Ouml;zcan, F., Zou, T., Goncalves, R. and Pirahesh, H. (2016) Building a Hybrid Warehouse: Efficient Joins between Data Stored in HDFS and Enterprise Warehouse. ACM Transactions on Database Systems, 41, 1-38. https://doi.org/10.1145/2972950</mixed-citation></ref><ref id="scirp.103821-ref30"><label>30</label><mixed-citation publication-type="other" xlink:type="simple">Pang, G. and Li, H. (2018) Caching for SQL-on-Hadoop. In: Encyclopedia of Big Data Technologies, Springer International Publishing, Berlin, 1-5. https://doi.org/10.1007/978-3-319-63962-8_249-1</mixed-citation></ref><ref id="scirp.103821-ref31"><label>31</label><mixed-citation publication-type="other" xlink:type="simple">Dean, J. and Ghemawat, S. (2008) MapReduce. Communications of the ACM, 51, 107. https://doi.org/10.1145/1327452.1327492</mixed-citation></ref><ref id="scirp.103821-ref32"><label>32</label><mixed-citation publication-type="other" xlink:type="simple">Sakr, S. and Zomaya, A.Y. (2019) Apache Hadoop. In: Encyclopedia of Big Data Technologies, Springer International Publishing, Berlin, 58. https://doi.org/10.1007/978-3-319-77525-8_100009</mixed-citation></ref><ref id="scirp.103821-ref33"><label>33</label><mixed-citation publication-type="other" xlink:type="simple">Koitzsch, K. (2017) Pro Hadoop Data Analytics. Apress, New York. https://doi.org/10.1007/978-1-4842-1910-2</mixed-citation></ref><ref id="scirp.103821-ref34"><label>34</label><mixed-citation publication-type="other" xlink:type="simple">Elahi, I. (2019) Hello Apache Spark. In: Scala Programming for Big Data Analytics, Apress, New York, 261-299. https://doi.org/10.1007/978-1-4842-4810-2_14</mixed-citation></ref><ref id="scirp.103821-ref35"><label>35</label><mixed-citation publication-type="other" xlink:type="simple">Ting, K. and Cecho, J.J. (2013) Apache Sqoop Cookbook. O’Reilly Media, Inc., Sebastopol.</mixed-citation></ref><ref id="scirp.103821-ref36"><label>36</label><mixed-citation publication-type="other" xlink:type="simple">Vohra, D. (2016) Apache Sqoop. In: Practical Hadoop Ecosystem, Apress, New York, 261-286. https://doi.org/10.1007/978-1-4842-2199-0_5</mixed-citation></ref><ref id="scirp.103821-ref37"><label>37</label><mixed-citation publication-type="other" xlink:type="simple">&amp;Ouml;zcan, F., Hoa, D., Beyer, K.S., Balmin, A., Liu, C.J. and Li, Y. (2011) Emerging Trends in the Enterprise Data Analytics. In: Proceedings of the 2011 International Conference on Management of Data, ACM Press, New York. https://doi.org/10.1145/1989323.1989446</mixed-citation></ref><ref id="scirp.103821-ref38"><label>38</label><mixed-citation publication-type="other" xlink:type="simple">Vertica (2016) Hadoop Integration Guide—Hp Vertica Analytic Database.https://softwaresupport.softwaregrp.com/doc/KM00681126?fileName=hp_man_Vertica_7.0.x_Hadoop_integration_pdf.pdf</mixed-citation></ref><ref id="scirp.103821-ref39"><label>39</label><mixed-citation publication-type="other" xlink:type="simple">Oracle (2012) High Performance Connectors for Load and Access of Data from Hadoop to Oracle Database.https://www.oracle.com/technetwork/bdc/hadoop-loader/connectors-hdfs-wp-1674035.pdf</mixed-citation></ref><ref id="scirp.103821-ref40"><label>40</label><mixed-citation publication-type="other" xlink:type="simple">Huai, Y., Zhang, X., Chauhan, A., Gates, A., Hagleitner, G., Hanson, E.N., Malley, O.O., Pandey, J., Yuan, Y. and Lee, R. (2014) Major Technical Advancements in Apache Hive. In: Proceedings of the 2014 ACM SIGMOD International Conference on Management of Data, ACM Press, New York. https://doi.org/10.1145/2588555.2595630</mixed-citation></ref><ref id="scirp.103821-ref41"><label>41</label><mixed-citation publication-type="other" xlink:type="simple">Thusoo, A., Sarma, J.S., Jain, N., Shao, Z., Chakka, P., Anthony, S., Liu, H., Wyckoff, P. and Murthy, R. (2009) Hive: A Warehousing Solution over a Map-Reduce Framework. Proceedings of the VLDB Endowment, 2, 1626-1629. https://doi.org/10.14778/1687553.1687609</mixed-citation></ref><ref id="scirp.103821-ref42"><label>42</label><mixed-citation publication-type="other" xlink:type="simple">Vohra, D. (2016) Apache Hive. In: Practical Hadoop Ecosystem, Apress, New York, 209-231. https://doi.org/10.1007/978-1-4842-2199-0_3</mixed-citation></ref><ref id="scirp.103821-ref43"><label>43</label><mixed-citation publication-type="other" xlink:type="simple">Oracle (2017) Apache Hive SQL Conformance (2017). https://cwiki.apache.org/confluence/display/Hive/Apache+Hive+SQL+Conformance</mixed-citation></ref><ref id="scirp.103821-ref44"><label>44</label><mixed-citation publication-type="other" xlink:type="simple">Khalifa, S. (2018) Tools and Libraries for Big Data Analysis. In: Encyclopedia of Big Data Technologies, Springer International Publishing, Berlin, 1-7. https://doi.org/10.1007/978-3-319-63962-8_282-1</mixed-citation></ref></ref-list></back></article>