<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">JTTs</journal-id><journal-title-group><journal-title>Journal of Transportation Technologies</journal-title></journal-title-group><issn pub-type="epub">2160-0473</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/jtts.2021.112016</article-id><article-id pub-id-type="publisher-id">JTTs-108625</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Engineering</subject></subj-group></article-categories><title-group><article-title>
 
 
  Demand Prediction of Ride-Hailing Pick-Up Location Using Ensemble Learning Methods
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Divine</surname><given-names>Carson-Bell</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Mawutor</surname><given-names>Adadevoh-Beckley</given-names></name><xref ref-type="aff" rid="aff2"><sup>2</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Kendra</surname><given-names>Kaitoo</given-names></name><xref ref-type="aff" rid="aff2"><sup>2</sup></xref></contrib></contrib-group><aff id="aff2"><addr-line>Department of Transport, Regional Maritime University, Accra, Ghana</addr-line></aff><aff id="aff1"><addr-line>College of Transport and Communication, Shanghai Maritime University, Shanghai, China</addr-line></aff><pub-date pub-type="epub"><day>25</day><month>02</month><year>2021</year></pub-date><volume>11</volume><issue>02</issue><fpage>250</fpage><lpage>264</lpage><history><date date-type="received"><day>16,</day>	<month>March</month>	<year>2021</year></date><date date-type="rev-recd"><day>22,</day>	<month>April</month>	<year>2021</year>	</date><date date-type="accepted"><day>25,</day>	<month>April</month>	<year>2021</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
  Ride-hailing and carpooling platforms have become a popular way to move around in urban cities. Based on the principle of matching riders with drivers, with Uber, Lyft and Didi having the largest market share. The challenge re
  mains being able to optimally match rider demand with driver supply, reducing congestion and emissions associated with Vehicle clustering, dead
  heading, ultimately leading to surge pricing where providers raise the price of the trip in order to attract drivers into such zones. This sudden spike in rates is seen by many riders as disincentive on the service provided. In this paper, data mining techniques are applied to ultimately develop an ensemble learning model based on historical data from City of Chicago Transport provider’s dataset. The objective is to develop a dynamic model capable of predicting rider drop-off location using pick-up location data then subsequently using 
  drop-off location data to predict pick-up points for effective driver
   deployment 
  under multiple scenarios of privacy and information. Results show neural
   network algorithms perform best in generalizing pick-up and drop-off points 
  when given only starting point information. Ensemble learning methods,
   Adaboost and Random forest algorithm are able to predict both drop-off and pick-up points with a MAE of one (1) community area knowing rider pick-up 
  point and Census Tract information only and in reverse predict potential 
  pick-up points using the Drop-off point as the new starting point.
 
</p></abstract><kwd-group><kwd>Ride-Hailing</kwd><kwd> Braess Paradox</kwd><kwd> Vehicle Clustering</kwd><kwd> Deadheading</kwd><kwd> Congestion</kwd><kwd> Predictive Modelling</kwd><kwd> Vehicle Deployment</kwd><kwd> Ensemble Learning</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>In recent years, ride-hailing and carpooling platforms have become increasingly popular and convenient way of moving around in most modern cities, matching riders with drivers, with Uber, Lyft and Didi being the biggest providers within the industry. In light of increased environmental awareness as well as concerns on minimizing carbon footprint, ridesharing and carpooling has become increasingly important.</p><p>Carpooling has numerous societal and individual benefits, including but not limited to reduction of Greenhouse-Gas emissions, cost savings in terms of shared travel costs for public agencies and employers [<xref ref-type="bibr" rid="scirp.108625-ref1">1</xref>].</p><p>In their paper, [<xref ref-type="bibr" rid="scirp.108625-ref2">2</xref>] present salient points in the understanding of the key aspects of the existing ridesharing system, going on to design a framework to identify challenges in the use of ridesharing thus fostering the development of mechanisms to overcome and promote widespread use.</p><p>Emerging studies [<xref ref-type="bibr" rid="scirp.108625-ref3">3</xref>] demonstrate psychological factors such as monetary and time benefits becoming more dominant factors in decisions to use ride-hailing and carpooling services. In relation to rider satisfaction, [<xref ref-type="bibr" rid="scirp.108625-ref4">4</xref>] found surge pricing not to bias Uber towards riders of higher income threshold, but rather, homophilous matching that is, matching riders to drivers of a similar age resulted in higher ratings and further went on to use these insights to predict driver and/or rider retention. Examining ridesharing platforms, [<xref ref-type="bibr" rid="scirp.108625-ref5">5</xref>] concluded moving forward, these platforms will do more good than harm, also, it was found that relatively little is known about their efficiency and equity but is likely to change with growing research interest. Using online reviews of drivers of popular ride-hailing companies, Uber and Lyft, [<xref ref-type="bibr" rid="scirp.108625-ref6">6</xref>] was able to demonstrate preference of Uber to Lyft. In addition, analysis show increased competition to attract more drivers, for which drivers counted job flexibility, and meeting new people as main advantages. In contrast, insufficient compensation, poor job security, poor rider behavior and poor customer service as impeding factors.</p></sec><sec id="s2"><title>2. Problem Statement</title><p>The Braess Paradox [<xref ref-type="bibr" rid="scirp.108625-ref7">7</xref>] [<xref ref-type="bibr" rid="scirp.108625-ref8">8</xref>] is a network phenomenon in which it is observed that the addition of extra capacity reduces overall network performance over time with lack of cooperation of users being the ultimate culprit for network breakdown</p><p>Congestion &amp; Vehicle-Clustering: Over the years, the number of vehicles engaged in ride-hailing has increased astronomically, surpassing taxis in many urban cities, [<xref ref-type="bibr" rid="scirp.108625-ref9">9</xref>]. A report from the (Union of Concerned Scientists, 2020) shows that ride hailing trips are responsible for 69 percent more emissions than the trips the service displaces with a significant amount of trips being Deadheading (Dead-mileage). This constitutes the period between drop-off and pick-up and is associated with increased costs [<xref ref-type="bibr" rid="scirp.108625-ref10">10</xref>]. Surge pricing i.e. where prices are adjusted upwards to meet acute driver shortage is viewed a disincentive to many riders, leading to lost revenue.</p><p>Solving the problem</p><p>In order to combat the problems above, it is necessary to develop a sound driver deployment strategy. Collective Intelligence (COIN) [<xref ref-type="bibr" rid="scirp.108625-ref11">11</xref>] was first suggested as a way of solving Braess paradox. This involves all networks users acting centrally for the benefit of all. [<xref ref-type="bibr" rid="scirp.108625-ref12">12</xref>], Observed that strategic repositioning is key to maximizing driver earnings as against surge chasing which increases Deadmileage. First and foremost will be to be able to predict and deploy vehicles accordingly. [<xref ref-type="bibr" rid="scirp.108625-ref13">13</xref>], conclude that centralized fleet coordination offers substantial benefits towards sustainable growth and market share.</p><p>Research Purpose and Objective</p><p>The objective is to develop a city-wide prediction algorithm capable of predicting trip pick-up and drop-off points, as well as potential pick-up locations after each drop-off based on historical data using Data mining Techniques.</p><p>Case study: City of Chicago, Illinois.</p></sec><sec id="s3"><title>3. Related Works</title><p>The growth of demand for ride-hailing services has disrupted urban transportation and is changing the way in which people travel. Modern ride-hailing services require the development of efficient recommendation systems in order to improve both riders and driver experience. In response, many researchers have conducted various experiments to help predict ride hailing demand in order to improve effective ride-hailing vehicle deployment.</p><p>In attempting to optimize the number of pick-ups whilst minimizing waiting time for taxi services, [<xref ref-type="bibr" rid="scirp.108625-ref14">14</xref>] developed a ride-hailing recommendation system. This is completed in 3 phases. The model starts by first effectively estimating future customer demand in different clusters within the area of interest. This is followed up with a taxi-to-region matching according to preset rules and conditions including driver preference and finally concluded with the design of an optimized geo-routing algorithm to help drivers minimize dead-mileage. The problem with this mainly lies with the instability of driver preference which changes frequently, making the approach difficult to deploy in real world situations.</p><p>Dead-mileage comprises a significant share of total travel covered by drivers within the ride-hailing industry in terms of miles travelled and number of trips overall. Accurate demand prediction within the ride-hailing industry can greatly improve vehicle utilization whilst reducing waiting time. Customers mainly desire minimization of waiting time whilst drivers on the other hand aim to minimize deadheading and idle time after trips. This subsection of the industry comprises another area of strong research interest.</p><p>[<xref ref-type="bibr" rid="scirp.108625-ref15">15</xref>], is one of the first to study this emerging field. He develops a model which predicts the gap between rider demand and driver supply within a given time period and specific geographic area using Point of Interest (POI), Traffic, Weather data as well as data from Car sharing orders. A data sampling techniques is used to determine patterns and generalizations which can be applied in real case scenarios forming the basis for future work. This concept of finding the supply and demand gap is important as it allows for the deployment of drivers to improve the level of service</p><p>Time based demand prediction is another research area fast gaining ground. This is based on the premise of predicting ride-hailing vehicle demand in the next hour.</p><sec id="s3_1"><title>3.1. Operational Research Mobility Optimization</title><p>The vast majority of human interaction takes place in one of two areas; home or work. In order to further understand mobility patterns of users of ridesharing services across home and work locations, as well as social ties between users, [<xref ref-type="bibr" rid="scirp.108625-ref16">16</xref>] developed an algorithm for matching users with similar mobility patterns under constraints and concluded, a decrease in social distance of as much as 31% when users shared rides with others. These findings indicate the importance of the study of mobility patterns and the benefits which can be derived from optimizing ride-hailing services at an operational level. Using a more flexible yet extendible mobility model representing ride-sharing users movement and habits, [<xref ref-type="bibr" rid="scirp.108625-ref17">17</xref>] deploy a Variable-Order Markov Model (VOMM) underplayed with a Partial Matching (PPM) algorithm for next location prediction, with a prediction accuracy ranging from 60% - 81%. A major limitation of the usage of the PPM algorithm hovers around the compression process which tends to limit performance over time. In comparing the use of privately owned vehicles and two Autonomous Mobility on-Demand (AMoD) simulated on a real transport network based on current situation, under different scenarios, [<xref ref-type="bibr" rid="scirp.108625-ref18">18</xref>] found the deployment of AMoD system resulted in a major decrease in both number of vehicles required in order to meet transport needs (that is, 43% in AMoD1 and 88% in AMoD2) and street parking space required (58% in AMoD1 and 83% in AMoD2). [<xref ref-type="bibr" rid="scirp.108625-ref19">19</xref>], also cite effective road utilization as another advantage of designing the matching algorithm. Comparing the use of privately owned vehicles and two autonomous mobility on-demand (AMoD) simulated on a real transport network based on current situation, under different scenarios. Autonomous Mobility on-Demand vehicles are viewed by many as the future of transport, however their effectiveness hinders largely on the ability to coordinate their movement and predict demand as accurately as possible using the vast quantity of data we have available at our disposal, for which this paper seeks to pursue further.</p><p>In an attempt to resolve the surge of homeward-bound persons during the holiday seasons, [<xref ref-type="bibr" rid="scirp.108625-ref20">20</xref>] proposed a large-scale ridesharing system called CountryRoads<sup>&#174;</sup> using an online greedy matching algorithm to match drivers and passengers, recording a success rate of 23.2%. Online Greedy matching algorithms have a comparatively low performance threshold when applied in complex systems such as ride-hailing services as experienced by the authors this is largely due to the level of rigidity of process making it not ideal for location prediction. Based on the concept of space-time windows, [<xref ref-type="bibr" rid="scirp.108625-ref21">21</xref>], develop a unique approach based on Lagrangian relaxation, and conclude that the adoption of flexible pickup and delivery will evidently reduce system-wide cost whilst improving service quality. This hypothesis although found to be true, defeats the purpose of ride hailing services. Flexible pickup and delivery have not been widely accepted even within the carpooling sphere as centralized pick-up location is yet to gather wide acceptance.</p></sec><sec id="s3_2"><title>3.2. Linear Programming &amp; Statistical Methods</title><p>In implementing optimization solutions based on linear programming, [<xref ref-type="bibr" rid="scirp.108625-ref22">22</xref>] deploy a Tabubased meta-heuristic algorithm with the aim of solving the mixed integer linear program (MILP) under differing scenarios. The algorithm is observed to have a higher computational accuracy than control, the introduction of meet points to the ridesharing system reduces total travel time by 2.7% - 3.8% for scaled tests. With meet-points not having been widely accepted within the ride-hailing and carpooling industry, the benefits of reduced travel time, and reduced travel costs associated with it cannot be fully quantified. Especially given Covid-19 social distancing protocols. This demonstrates the need to improve location prediction as a lasting solution.</p><p>From the domain of probability and statistics, [<xref ref-type="bibr" rid="scirp.108625-ref23">23</xref>] having collected data of taxi trips in New York, Singapore, San Francisco and Vienna compute shareability curves for each city, then through natural rescaling collapse them into a universal curve which is used to predict the potential of ridesharing in any given city based on a few qualities and parameters. The statistical methods employed here demonstrate the general overview of the potential of the growth of ride-hailing services in any given city. This is to help with city planning purposes and fails to examine rider-driver interaction.</p><p>Examining the relationship between the frequency and probability of ridesharing usage, and frequency of public transit usage, [<xref ref-type="bibr" rid="scirp.108625-ref24">24</xref>], develop a Zero-inflated negative binomial regression model.</p><p>Results show a positive relationship between ridesharing and public transit use particularly for people living in areas of high population density and comparatively fewer vehicles. The significance of this is to allow the measurement of ride-hailing service utilization across population densities across any given city taking into consideration anticipated demand and in the selection of the research Case study.</p></sec></sec><sec id="s4"><title>4. Research Framework and Design</title><p>To reduce the number of vehicles, alleviate traffic jams and curb pollution in transporting people in office hubs in Poland, [<xref ref-type="bibr" rid="scirp.108625-ref25">25</xref>] collected a representative sample of the population and used spatial data mining techniques to develop a set of parameters for the multi-agent system. Using the distributed model-free, system DeepPool<sup>&#174;</sup> based on deep Q-network (DQN) techniques, [<xref ref-type="bibr" rid="scirp.108625-ref26">26</xref>] develop an algorithm able to learn the optimal dispatch policy through interaction with the environment, incorporating travel demand statistics and a dataset of taxi trips in New York to dispatch vehicles and anticipate future demand. Deploying a convolutional neural network (CNN) based on deep learning for multi-step ride-hailing demand prediction using trip request data in Chengdu, [<xref ref-type="bibr" rid="scirp.108625-ref27">27</xref>] showcase faster training and prediction of CNN models compared to the use of Long Short Term Memory (LSTM) models.</p><sec id="s4_1"><title>4.1. Data</title><p>In conducting this research, a large scale dataset of rideshare and taxi trips spanning 2018/2019 in Chicago is collected, as shown in <xref ref-type="table" rid="table1">Table 1</xref>, with each observation consisting of the following elements:</p><p>The data is processed and cleaned. As a first step, a comprehensive understanding of the individual features within the dataset is required, as well as knowledge of trip distribution across the city, from origin (O) to Destination (D). Numerous studies have demonstrated the importance of regional partitioning in location prediction. Research and experiments by [<xref ref-type="bibr" rid="scirp.108625-ref28">28</xref>] demonstrated that regional partitioning led to better forecast and demand prediction of geospatial data.</p><p>This is followed up with followed by scenario development. <xref ref-type="fig" rid="fig1">Figure 1</xref> shows a color-coded layout of the City of Chicago, detailing its community areas as well as census tracts.</p><table-wrap id="table1" ><label><xref ref-type="table" rid="table1">Table 1</xref></label><caption><title> Data points used for data mining and the development of the predictive algorithms</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Trip ID</th><th align="center" valign="middle"  colspan="3"  >DATA FEATURES AND ATTRIBUTES PER OBSERVATION</th></tr></thead><tr><td align="center" valign="middle" >Trip Start Timestamp</td><td align="center" valign="middle" >Drop-off Census Tract</td><td align="center" valign="middle" >Pickup Census Tract</td><td align="center" valign="middle" >Fare</td></tr><tr><td align="center" valign="middle" >Trip End Timestamp</td><td align="center" valign="middle" >Drop-off Community Area</td><td align="center" valign="middle" >Pickup Community Area</td><td align="center" valign="middle" >Shared Trip Authorized</td></tr><tr><td align="center" valign="middle" >Trip Seconds</td><td align="center" valign="middle" >Drop-off Centroid Longitude</td><td align="center" valign="middle" >Pickup Centroid Longitude</td><td align="center" valign="middle" >Additional Charges</td></tr><tr><td align="center" valign="middle" >Trip Miles</td><td align="center" valign="middle" >Drop-off Centroid Latitude</td><td align="center" valign="middle" >Pickup Centroid Latitude</td><td align="center" valign="middle" >Trips Pooled</td></tr><tr><td align="center" valign="middle" >Trip Total</td><td align="center" valign="middle" >Drop-off Centroid Location</td><td align="center" valign="middle" >Pickup Centroid Location</td><td align="center" valign="middle" >Tip</td></tr></tbody></table></table-wrap></sec><sec id="s4_2"><title>4.2. Research Framework and Scope</title><p>Multidimensional Scenario Formulation</p><p>Scenario performance analysis allows for measuring performance under varied rider privacy limitations.</p><p>Scenario 1</p><p>Location prediction with no information i.e. drop-off community area (destination) prediction with only pick-up (origin) data, and vice versa. This is in order to allow for riders with strict privacy concerns in information release, measuring ability to predict trip start and end points given rider privacy restrictions.</p><p>Scenario 2</p><p>Location prediction with partial information. That is, drop-off community area (destination) prediction with pick-up data and Census Tract (destination zone) information, vice versa. It is based on the idea of being able to predict trip start and end points under rider uncertainty.</p><p>Steps and Methodological process</p><p><xref ref-type="fig" rid="fig2">Figure 2</xref> shows the steps taken in the design, evaluation and interpretation of the research framework employed in carrying out this work.</p><p>1) Perform Principal Component Analysis (PCA) on trip dataset. Record and analyze results against degree of variance covered by each principal component.</p><p>2) Perform feature scoring and ranking using Relief metrics. Record and analyze results.</p><p>3) Reevaluate steps 1 and 2. Determine features and variables with largest weight in designing and building the model.</p><p>4) Evaluation and scoring of prediction accuracy and error tolerance (MAE, MSE, and R2) under both scenario 1 and 2.</p><p>5) In-depth scenario analysis of both scenario 1 and 2, firstly on drop-off community area prediction and pick-up community area prediction.</p><p>6) Analyzing implications on surge pricing policy and ridesharing efficiency.</p><p>Principal Component Analysis (PCA)</p><p>Principal component analysis (PCA) is based on the use of an orthogonal transformation to convert a set of observations with possibly correlated variables in a set of linearly uncorrelated principal components using eigenvalues to measure the total degree of variance explained by each factor.</p><p>FEATURE RANK USING RRELIEFF</p><p>The RReliefF algorithm estimates the quality of an attribute according to the degree with which it discriminates between instances near each other. Here, an instance R is randomly selected, then the K-nearest instances with respect to class value are selected. The difference between the value of A of R as well as the value of the same attribute for one of the K-instances is then compared with respect to the difference of their class values. This process is repeated and ultimately yields a weight for each attribute ranging between −1 and 1.</p><p>Cross Validation Model Evaluation and Scoring</p><p>The Leave-P-Out Cross Validation (CV) approach leaves “p” data points out of the training data, with a sample size of n-p being used as the validation set. This process is repeated for all possible combinations, with error being averaged for all trials in order to determine overall effectiveness.</p><p>To measure the degree of error of the developed models, error metrics will then be used to judge model quality and compare the different regression models. The Mean Average Error (MAE), Mean Squared Error (MSE), and R-Squared (R<sup>2</sup>) will be used for evaluation.</p><p>MAE = 1 n ∑ i = 1 n | Y i − Y ⌢ i |</p><p>MSE = 1 n ∑ i = 1 n ( Y i − Y ⌢ i ) 2</p><p>where,</p><p>Y ⌢ —Predicted value of Y</p><p>Y &#175; —mean value of Y</p><p>R 2 = ( SSEM − SSER ) ( SSEM ) = 1 − SSER SSEM</p><p>where,</p><p>SSEM is the sum of Squared Errors by Mean line and</p><p>SSER is the sum of Squared Errors by Regression Line</p><p>Predictive Modelling using Ensemble Learning</p><p>Generally, ensemble learning is the term used to describe meta-algorithms that makes predictions based on inputs from different models, thus, by combining multiple individual models, the ensemble model tends to have less bias, variance, and avoids overfitting culminating in improved predictions.</p><p>Adaboost and Random Forest are the most commonly used.</p></sec></sec><sec id="s5"><title>5. Framework and Results</title><sec id="s5_1"><title>5.1. Principal Component Analysis (PCA)</title><p>In analyzing the weights of the individual features within the data sample collected, PCA analysis is performed, measuring the degree of variance covered by each principal component within the data set.</p><p>Analysis of PCA results reveals an increase in the degree of variance explained by each of the data attributes within the dataset.</p><p><xref ref-type="fig" rid="fig3">Figure 3</xref> describes the results obtained from PCA analysis. Results show that certain attributes within the dataset are able to explain 55.9% of the recorded variance, with 5 attributes able to explain 73.7% of the variance and so on. This aids in selecting the most important data attributes which will effectively improve the models prediction accuracy. Analysis reveals that 9 attributes to be the optimal number of features to incorporate in building the models.</p><p>Feature Scoring and Rank</p><p>After PCA analysis, the features within the dataset are then ranked in order according to feature influence on prediction output. <xref ref-type="fig" rid="fig4">Figure 4</xref> details the weight associated with each attribute used in designing the model, with some attributes being more critical to predictive performance than others.</p><p>RreliefF is used to rank and measure individual features by level of importance as shown above.</p></sec><sec id="s5_2"><title>5.2. Re-Evaluation and Model Calibration</title><p>Scenario 1</p><p>Predicting drop-off community area (destination) with only pick-up (origin) data. Model results show an ability of linear regression models to predict potential Drop-off areas within a radius of 13 blocks (community area). This is in the absence of any information other than pick-up point (origin).</p><p>The results are shown in <xref ref-type="fig" rid="fig5">Figure 5</xref>:</p><p>This figure is divided into 2 parts, with the first part (Top) displaying results from model evaluation whilst the 2nd displays location predictive results against actual. The dark column above displays actual drop-off community areas as against predicted values on its left.</p><p>Scenario 2</p><p>Predicting drop-off community area (destination) with partial information, that is, (destination zone) information</p><p>Model results show an ability of ensemble learning models such as Adaboost and Random Forest to predict potential Drop-off areas precisely with error under 1 block (community area).</p><p>This is in the absence of any information other than pick-up point (origin). The results are shown below.</p><p><xref ref-type="fig" rid="fig6">Figure 6</xref> is divided into 2 parts, with the first part (Top) displaying results from model evaluation whilst the 2nd displays location predictive results against actual. The dark column above displays actual drop-off community areas as against predicted values on its left.</p><p>Pick-up point Prediction after Drop-off</p><p>This part focuses on predicting demand centers within the city after the any given drop-off. The aim is to predict rideshare demand centers, anticipating demand and price surge before they happen.</p><p>In an effort to optimize rideshare vehicle distribution, it is imperative to be able to predict where demand will occur ahead of time, taking advantage of imbalance of supply and demand as well as revenue per trip, with the results displayed in <xref ref-type="fig" rid="fig7">Figure 7</xref> below.</p><p>This figure is divided into 2 parts, with the first part (Top) displaying results from model evaluation whilst the 2<sup>nd</sup> part displays location predictive results against actual. The dark column above displays actual drop-off community areas as against predicted values on its left.</p></sec><sec id="s5_3"><title>5.3. Discussion</title><p>Research into the field of mobility remains a hot topic amongst many researchers. Mobility-As-A-Service (MAAS) where vehicle trips are used to render services has come to stay in the era where we’ve experienced a boom in ride-hailing</p><p>services. The need to optimize the operations of these services remains of utmost importance. The results show neural network algorithms perform best in generalizing pick-up and drop-off points when provided with only starting point information. The significance of this is to allow for trip generalization in pooled trips, where riders are most likely to have a common drop-off point, e.g. coworker’s trip to work or trips to work or shared trip to a sporting event. Ensemble learning methods, Adaboost and Random forest algorithm are able to predict both drop-off and pick-up points with a MAE of 1 community area knowing rider pick-up point and Census Tract information only and in reverse predict potential pick-up points using the Drop-off point as the new starting point. This allows the algorithm to confidently predict the most likely pick-up point of potential riders following a drop-off in in so doing increasing supply of drivers into potential surge zones and thus being less reactive, more proactive in trip deployment. Here, it can be seen that the introduction of more data and ensemble learning techniques greatly increases the precision accuracy of the model. This demonstrates the influence of data management within the ride-hailing industry, especially in a time when privacy concerns and right to privacy have become a matter of safety and security, of which varies from rider to rider. Direct impacts on the ride-hailing industry and operations include:</p><p>Implications on ride-hailing Industry includes:</p><p>1) Improved vehicle utilization, and time efficiency.</p><p>2) Reduced dead-mileage and idle time after trips.</p><p>3) Improvement riders and driver experience.</p></sec></sec><sec id="s6"><title>6. Conclusions</title><p>In conclusion, results from the research indicate the ability to use predictive modelling and analytics to adequately maximize driver positioning and deployment by predicting surge zones before they occur irrespective of rider privacy settings.</p><p>The implications of these results on the transport industry includes:</p><p>• Reduced incidence of the surge and increasing rider satisfaction.</p><p>• Reduced transport costs.</p><p>• Increase in the ease of parking particularly in high-demand (downtown) areas.</p><p>• From a social and environmental point of view for fewer wasted miles would translate into less emissions overall.</p></sec><sec id="s7"><title>Conflicts of Interest</title><p>The authors declare no conflicts of interest regarding the publication of this paper.</p></sec><sec id="s8"><title>Cite this paper</title><p>Carson-Bell, D., Adadevoh-Beckley, M. and Kaitoo, K. (2021) Demand Prediction of Ride-Hailing Pick-Up Location Using Ensemble Learning Methods. Journal of Transportation Technologies, 11, 250-264. https://doi.org/10.4236/jtts.2021.112016</p></sec></body><back><ref-list><title>References</title><ref id="scirp.108625-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">Shaheen, S., Cohen, A. and Bayen, A. (2018) The Benefits of Carpooling. UC Berkeley. https://escholarship.org/uc/item/7jx6z631</mixed-citation></ref><ref id="scirp.108625-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">Furuhata, M., Dessouky, M., Ordó&amp;#241;ez, F., Brunet, M.-E., Wang, X. and Koenig, S. (2013) Ridesharing: The State-of-the-Art and Future Directions. Transportation Research Part B: Methodological, 57, 28-46. https://doi.org/10.1016/j.trb.2013.08.012</mixed-citation></ref><ref id="scirp.108625-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">Olsson, L.E., Maier, R. and Friman, M. (2019) Why Do They Ride with Others? Meta-Analysis of Factors Influencing Travelers to Carpool. Sustainability, 11, 2414. https://doi.org/10.3390/su11082414</mixed-citation></ref><ref id="scirp.108625-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">Kooti, F., Grbovic, M., Aiello, L.M., Djuric, N., Radosavljevic, V. and Lerman, K. (2017) Analyzing Uber’s Ride-Sharing Economy. International World Wide Web Conference Committee (IW3C2), Perth, 3-7 April 2017. https://doi.org/10.1145/3041021.3054194</mixed-citation></ref><ref id="scirp.108625-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">Hahn, R. and Metcalfe, R. (2017) The Ridesharing Revolution: Economic Survey and Synthesis. Volume IV: More Equal by Design: Economic Design Responses to Inequality. Oxford University Press, Oxford.</mixed-citation></ref><ref id="scirp.108625-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">Shokoohyar, S. (2018) Ride-Sharing Platforms from Drivers’ Perspective: Evidence from Uber and Lyft Drivers. International Journal of Data and Network Science, 2, 89-98. https://doi.org/10.5267/j.ijdns.2018.10.001</mixed-citation></ref><ref id="scirp.108625-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">Braess, D. (1968) über ein Paradoxon aus der Verkehrsplanung. Unternehmensforschung, 12, 258-268. https://doi.org/10.1007/BF01918335</mixed-citation></ref><ref id="scirp.108625-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">Braess, D., Nagurney, A. and Wakolbinger, T. (2005) On a Paradox of Traffic Planning. Transportation Science, 39, 446-450. https://doi.org/10.1287/trsc.1050.0127</mixed-citation></ref><ref id="scirp.108625-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">Schneider, T. (2019) Taxi and Ride-Hailing Usage in Chicago. http://toddwschneider.com/dashboards/chicago-taxi-ridehailing-data</mixed-citation></ref><ref id="scirp.108625-ref10"><label>10</label><mixed-citation publication-type="other" xlink:type="simple">Nair, G.S., Bhat, C.R., Batur, I., Pendyala, R.M. and Lam, W.H.K. (2020) A Model of Deadheading Trips and Pick-Up Locations for Ride-Hailing Service Vehicles. Transportation Research Part A: Policy and Practice, 135, 289-308. https://doi.org/10.1016/j.tra.2020.03.015</mixed-citation></ref><ref id="scirp.108625-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">Tumer, K. and Wolpert, D. (2002) Collective Intelligence and Braess’ Paradox. Journal of Artificial Intelligence Research, 16, 359-387. https://doi.org/10.1613/jair.995</mixed-citation></ref><ref id="scirp.108625-ref12"><label>12</label><mixed-citation publication-type="other" xlink:type="simple">Chaudhari, H.A., Byers, J.W. and Terzi, E. (2018) Putting Data in the Driver’s Seat: Optimizing Earnings for On-Demand Ride-Hailing. 11th Eleventh ACM International Conference on Web Search and Data Mining, New York, 5-9 February 2018, 9 p. https://doi.org/10.1145/3159652.3159721</mixed-citation></ref><ref id="scirp.108625-ref13"><label>13</label><mixed-citation publication-type="other" xlink:type="simple">Merlin, L.A. (2019) Transportation Sustainability Follows from More People in Fewer Vehicles, Not Necessarily Automation. Journal of the American Planning Association, 85, 501-510. https://doi.org/10.1080/01944363.2019.1637770</mixed-citation></ref><ref id="scirp.108625-ref14"><label>14</label><mixed-citation publication-type="other" xlink:type="simple">Wan, X., Ghazzai, H. and Massoud, Y. (2020) A Generic Data-Driven Recommendation System for Large-Scale Regular and Ride-Hailing Taxi Services. Electronics, 9, 648. https://doi.org/10.3390/electronics9040648</mixed-citation></ref><ref id="scirp.108625-ref15"><label>15</label><mixed-citation publication-type="other" xlink:type="simple">Wang, R. (2017) Supply-Demand Forecasting for a Ride-Hailing System. http://escholarship.org/uc/item/7hr5t5vv</mixed-citation></ref><ref id="scirp.108625-ref16"><label>16</label><mixed-citation publication-type="other" xlink:type="simple">Cici, B., Markopoulou, A., Laoutaris, F.-M. and Nikolaos, E.A. (2017) Assessing the Potential of Ride-Sharing Using Mobile and Social Data: A Tale of Four Cities. https://doi.org/10.1145/2632048.2632055</mixed-citation></ref><ref id="scirp.108625-ref17"><label>17</label><mixed-citation publication-type="other" xlink:type="simple">Roor, R., Karg, M., Liao, A., Lei, W. and Kirsch, A. (2017) Predictive Ridesharing Based on Personal Mobility Patterns. Intelligent Vehicles Symposium (IV), Los Angeles, 11-14 June 2017. https://doi.org/10.1109/IVS.2017.7995895</mixed-citation></ref><ref id="scirp.108625-ref18"><label>18</label><mixed-citation publication-type="other" xlink:type="simple">Dia, H. and Javanshour, F. (2017) Autonomous Shared Mobility-on-Demand: Melbourne Pilot Simulation Study. Transportation Research Procedia, 22, 285-292. https://doi.org/10.1016/j.trpro.2017.03.035</mixed-citation></ref><ref id="scirp.108625-ref19"><label>19</label><mixed-citation publication-type="other" xlink:type="simple">Sonet, K.M.H., Rahman, M.M., Mehedy, S.R. and Rahman, R.M. (2019) A Dynamic Ridesharing and Carpooling Solution Using Advanced Optimised Algorithm. International Journal of Knowledge Engineering and Data Mining, 6, 1-31. https://doi.org/10.1504/IJKEDM.2019.097355</mixed-citation></ref><ref id="scirp.108625-ref20"><label>20</label><mixed-citation publication-type="other" xlink:type="simple">Jiang, W., Dominguez, C.R., Zhang, P., Zhang, S., et al. (2018) Large-Scale Nationwide Ridesharing System: A Case Study of Chunyun. International Journal of Transportation Science and Technology, 7, 45-59. https://doi.org/10.1016/j.ijtst.2017.10.002</mixed-citation></ref><ref id="scirp.108625-ref21"><label>21</label><mixed-citation publication-type="other" xlink:type="simple">Zhao, M., Yin, J., An, S., Wang, J. and Feng, D. (2018) Ridesharing Problem with Flexible Pickup and Delivery Locations for App-Based Transportation Service: Mathematical Modeling and Decomposition Methods. Journal of Advanced Transportation, 2018, Article ID: 6430950. https://doi.org/10.1155/2018/6430950</mixed-citation></ref><ref id="scirp.108625-ref22"><label>22</label><mixed-citation publication-type="other" xlink:type="simple">Li, X., Hu, S., Fan, W. and Deng, K. (2018) Modeling an Enhanced Ridesharing System with Meet Points and Time Windows. PLoS ONE, 13, e0195927. https://doi.org/10.1371/journal.pone.0195927</mixed-citation></ref><ref id="scirp.108625-ref23"><label>23</label><mixed-citation publication-type="other" xlink:type="simple">Tachet, R., Sagarra, O., Santi, P., Resta, G., Szell, M., Strogatz, S.H. and Ratti, C. (2017) Scaling Law of Urban Ride Sharing. Scientific Reports, 7, Article No. 42868. https://doi.org/10.1038/srep42868</mixed-citation></ref><ref id="scirp.108625-ref24"><label>24</label><mixed-citation publication-type="other" xlink:type="simple">Zhang, Y. and Zhang, Y. (2018) Exploring the Relationship between Ridesharing and Public Transit Use in the United States. International Journal of Environmental Research and Public Health, 15, 1763. https://doi.org/10.3390/ijerph15081763</mixed-citation></ref><ref id="scirp.108625-ref25"><label>25</label><mixed-citation publication-type="other" xlink:type="simple">Olszewski, R., Palka, P. and Turek, A. (2018) Solving “Smart City” Transport Problems by Designing Carpooling Gamification Schemes with Multi-Agent Systems: The Case of the So-Called “Mordor of Warsaw”. MDPI Sensors, 18, 141. https://doi.org/10.3390/s18010141</mixed-citation></ref><ref id="scirp.108625-ref26"><label>26</label><mixed-citation publication-type="other" xlink:type="simple">Alabbasi, A., Ghosh, A. and Aggarwal, V. (2019) DeepPool: Distributed Model-Free Algorithm for Ride-Sharing Using Deep Reinforcement Learning. IEEE Transactions on Intelligent Transportation Systems, 20, 4714-4727. https://doi.org/10.1109/TITS.2019.2931830</mixed-citation></ref><ref id="scirp.108625-ref27"><label>27</label><mixed-citation publication-type="other" xlink:type="simple">Wang, C., Hou, Y. and Barth, M. (2019) Data-Driven Multi-Step Demand Prediction for Ride-Hailing Services Using Convolutional Neural Network. Computer Vision Conference (CVC), Las Vegas, 25-26 April 2019. http://www.researchgate.net/publication/329402005 https://doi.org/10.1007/978-3-030-17798-0_2</mixed-citation></ref><ref id="scirp.108625-ref28"><label>28</label><mixed-citation publication-type="other" xlink:type="simple">Niu, K., Wang, C., Zhou, X. and Zhou, T. (2019) Predicting Ride-Hailing Service Demand via RPA-LSTM. IEEE Transactions on Vehicular Technology, 68, 4213-4222. https://doi.org/10.1109/TVT.2019.2901284</mixed-citation></ref></ref-list></back></article>