<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">JSEA</journal-id><journal-title-group><journal-title>Journal of Software Engineering and Applications</journal-title></journal-title-group><issn pub-type="epub">1945-3116</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/jsea.2019.126014</article-id><article-id pub-id-type="publisher-id">JSEA-93452</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Computer Science&amp;Communications</subject></subj-group></article-categories><title-group><article-title>
 
 
  A Survey on Software Cost Estimation Techniques
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Sai</surname><given-names>Mohan Reddy Chirra</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Hassan</surname><given-names>Reza</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib></contrib-group><aff id="aff1"><addr-line>School of Electrical Engineering &amp;amp; Computer Science, University of North Dakota, Grand Forks, USA</addr-line></aff><pub-date pub-type="epub"><day>24</day><month>06</month><year>2019</year></pub-date><volume>12</volume><issue>06</issue><fpage>226</fpage><lpage>248</lpage><history><date date-type="received"><day>26,</day>	<month>January</month>	<year>2019</year></date><date date-type="rev-recd"><day>27,</day>	<month>June</month>	<year>2019</year>	</date><date date-type="accepted"><day>30,</day>	<month>June</month>	<year>2019</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
  The ability to accurately estimate the cost needed to complete a specific project has been a challenge over the past decades. For a successful software project, accurate prediction of the cost, time and effort is a very much essential task. This paper presents a systematic review of different models used for software cost estimation which includes algorithmic methods, non-algorithmic methods and learning-oriented methods. The models considered in this review include both the traditional and the recent approaches for software cost estimation. The main objective of this paper is to provide an overview of software cost estimation models and summarize their strengths, weakness, accuracy, amount of data needed, and validation techniques used. Our findings show, in general, neural network based models outperforms other cost estimation techniques. However, no one technique fits every problem and we recommend practitioners to search for the model that best fit their needs.
 
</p></abstract><kwd-group><kwd>Software Cost Estimation</kwd><kwd> Classical SCE Models</kwd><kwd> Algorithmic Models</kwd><kwd> Non-Algorithmic Models</kwd><kwd> Learning-Oriented Cost Estimation Techniques</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>Software cost estimation is one of the crucial activities of the software development which involves predicting the effort, size and cost required to develop a software system or a software project [<xref ref-type="bibr" rid="scirp.93452-ref1">1</xref>] [<xref ref-type="bibr" rid="scirp.93452-ref2">2</xref>] . Several cost estimation models have been developed to better estimate the cost of a project. To have better accuracy in software cost estimation it is important to consider appropriate approaches to apply. Inaccurate effort estimations have been found to be very risky in the field of industrial economics. Over the past few decades, conducting cost estimation for different projects was cumbersome, but with the implementation of the software cost estimation process, things have positively changed. As it will be highlighted in this paper, there are different types of software cost estimation models which can be categorized into Algorithmic models, non-algorithmic models and learning oriented models [<xref ref-type="bibr" rid="scirp.93452-ref3">3</xref>] . The Algorithmic models include COCOMO, Function point analysis, Putnam model, etc. The non-algorithmic models which are also called as the non-parametric models include Expert judgment, analogy based, price to win, top-down and bottom up. The learning-oriented models or the machine learning methods include Artificial Neural Networks (ANN’s), Fuzzy Logic (FL), analogy based, Bayesian Network, Regression tree, Support Vector Machines, Genetic Algorithm (GA) and case-based reasoning [<xref ref-type="bibr" rid="scirp.93452-ref1">1</xref>] . The different methods have their pros and cons which imply that there are not the same in the sense of performance. The discussion will give more credential to modern methods of machine learning and their impacts on the software engineering sector.</p><p>Software Cost Estimation as a topic entails several issues as it requires keen observation and frequent trials before stipulating that a certain technique is fit for the estimation purposes. The utilization of software cost estimation techniques makes it possible to predict the amount of effort and cost that will be incurred in a certain software project. Thus, the approximated amount of manpower needed, and the period required to complete the project in the required time are all made possible by the software cost estimation process (<xref ref-type="fig" rid="fig1">Figure 1</xref>).</p><p>However, it still remains to be one of the most difficult areas of software engineering as each achievement made in the area attracts more questions and research. The algorithmic models primarily rely on the mathematical formulas and expressions to give a prediction on a project. The formulas exploited in the sector are also dependent on different factors such as product factor to calculate the estimations [<xref ref-type="bibr" rid="scirp.93452-ref3">3</xref>] . The non-algorithmic methods are considered to be most advanced as they have incorporated artificial intelligence in achieving their results. The field of artificial intelligence (AI) is a different scientific spectrum that integrates automation with non-algorithmic methods to come up with more accurate results [<xref ref-type="bibr" rid="scirp.93452-ref1">1</xref>] . The current methods have been improvised through artificial intelligence such that they can easily estimate the cost of a project even where limited data is available. More automation and improvements have been performed on the process which has led to the emergence of more strong methods.</p><p>Software cost estimation has the potential to completely change the industry by providing an accurate prediction of the amount of resources a project might need to be completed. However, at the moment, these estimation techniques could lead to important negative effects. The capability to estimate the cost and effort a project will take is of great importance because overestimation can easily lead to incurring of financial losses in any organization [<xref ref-type="bibr" rid="scirp.93452-ref1">1</xref>] . Under-estimation on the other hand, can significantly contribute to poor quality service delivery leading to failure of the entire project [<xref ref-type="bibr" rid="scirp.93452-ref1">1</xref>] . In a study, it has been reported that there is an overestimation of up to 40% in the estimating of software projects [<xref ref-type="bibr" rid="scirp.93452-ref4">4</xref>] . There is a need for efficient and accurate estimations to reduce the risks and timely delivery of a software project within the budget constraint [<xref ref-type="bibr" rid="scirp.93452-ref1">1</xref>] . This paper will address the strengths and weakness of different cost estimation techniques used in software cost estimation. Through the discussion, it is plausible for the researchers and the practitioners to understand the correct area where each model could be deployed and the factors that make it best suited for the specific area. All these issues will be highlighted and discussed in this paper (<xref ref-type="fig" rid="fig2">Figure 2</xref>).</p></sec><sec id="s2"><title>2. Background and Related Work</title><p>Several studies and systematic reviews related to the software cost estimation to techniques have been conducted to date [<xref ref-type="bibr" rid="scirp.93452-ref1">1</xref>] - [<xref ref-type="bibr" rid="scirp.93452-ref10">10</xref>] . The SLR papers gave the timeline of the cost estimation methods while the studies conducted gave a discussion on the existing cost estimation methods. There is still a research gap in listing out the popular cost estimation techniques and discussing the strength and weakness of those cost estimation techniques which aids the researchers and practitioners to choose which estimation technique to opt for depending on their needs. According to Jianfeng et al., software development effort estimation (SDEE) is a process that is used by the project managers or the software developers in predicting the effort required to develop a software system [<xref ref-type="bibr" rid="scirp.93452-ref2">2</xref>] . Over the</p><p>years since the 1990’s, researchers have suggested the implementation of Machine Learning (ML) models and SDEE in improving on estimation accuracy. Although there have been numerous researches on the process still there lacks empirical evidence of comparisons of different software cost estimation techniques. There are ongoing studies and researches that are underway to give detailed literature on the issue of software cost estimations as depicted by Sharma et al. [<xref ref-type="bibr" rid="scirp.93452-ref1">1</xref>] . The current researchers are aimed at improving o the functionality of the different cost estimation approaches that are used in software cost estimation. According to Jorgensen et al., it is of great relevance for the users of these approaches to be well conversant with each approach so as to aid in selecting the appropriate technique to deploy in software cost estimation [<xref ref-type="bibr" rid="scirp.93452-ref5">5</xref>] .</p><p>Current studies show that in order for us to be able to understand the achievements that have been met since the software cost estimation process began it would be of importance to review the methods back to decades when they were exploited [<xref ref-type="bibr" rid="scirp.93452-ref3">3</xref>] . Since then to the current year, companies have been exploiting different taxonomies and classification criteria in identifying the suitable method to support their estimations [<xref ref-type="bibr" rid="scirp.93452-ref7">7</xref>] . The first journals and reports were published between 50’s and 70’s which proves that studies on software cost estimation existed in the past despite them being conducted manually [<xref ref-type="bibr" rid="scirp.93452-ref8">8</xref>] . With numerous studies being perfumed on the process we are aware that the entire process entails software plans, resources, coding, testing, development and design. Numerous organizations are dependent upon software development for their stability and sustainability in relation to cost estimations. According to [<xref ref-type="bibr" rid="scirp.93452-ref8">8</xref>] , it is important for the cost of software to be estimated in such a way that it will not compromise quality, efficiency and timeline. The former methods that were exploited in the sector were dependent on source line of code (SLOC), cost drivers and function points [<xref ref-type="bibr" rid="scirp.93452-ref3">3</xref>] .</p><p>However current studies prove that there is a lot of automation in the software cost estimation process with each technique being implemented to fit the desired functions [<xref ref-type="bibr" rid="scirp.93452-ref1">1</xref>] . There are basic reasons as to why there is a tremendous research going on in software cost estimation process; proper budgeting, accurate estimations, software improvement investment analysis, project planning and control, trade-off and risk analysis.</p></sec><sec id="s3"><title>3. Approaches</title><sec id="s3_1"><title>3.1. Algorithmic Methods</title><p>1) COCOMO (Cost Constructive Model)</p><p>COCOMO is one of the classical techniques used in cost estimation [<xref ref-type="bibr" rid="scirp.93452-ref11">11</xref>] . There are different approaches that can be undertaken when it comes to estimating the cost of software projects. Among these methods that are available, COCOMO (Constructive Cost Model) is the most common model that is majorly used [<xref ref-type="bibr" rid="scirp.93452-ref11">11</xref>] . It was developed by Barry Bohem in 1981 [<xref ref-type="bibr" rid="scirp.93452-ref11">11</xref>] . The COCOMO falls under the algorithmic technique. The model has been in existence for many years and since it is being up to date serves to indicate that it is highly reliable as a cost estimation technique.</p><p>COCOMO is a software cost estimation approach that uses mathematical formulas and calculations to estimate the cost of a project. It gives the estimates regarding the amount of the effort that is required as well as the schedule for the software project [<xref ref-type="bibr" rid="scirp.93452-ref11">11</xref>] . Thus, it can achieve the two most vital objectives of cost estimation which ascertain the cost of the essential resources and schedule for the software project. In addition to the above, the Constructive Cost Model uses a set of metrics that guide its operations. Function Points (FP) and Object Points (OP) are the metrics that guide the calculations, and at the same time, they endure that the calculations are in line with codes relating to KLOC and KDSI [<xref ref-type="bibr" rid="scirp.93452-ref11">11</xref>] . COCOMO II, as well as COCOMO 81, are the two versions of Constructive Cost Model. The model parameters in the Cost Constructive model are derived from fitting a regression formula using historical projects data [<xref ref-type="bibr" rid="scirp.93452-ref11">11</xref>] . A total of 61 projects are used for COCOMO 81 and 163 projects for COCOMO II [<xref ref-type="bibr" rid="scirp.93452-ref11">11</xref>] .</p><p>2) COCOMO (Cost Constructive Model) 81</p><p>To being with, COCOMO 81 was the first version of this technique. Under this technique, estimates that are produced, according to [<xref ref-type="bibr" rid="scirp.93452-ref11">11</xref>] , have a 20% accuracy margin for the actual value of the software while 68% is the accuracy margin for the time estimate. In the same COCOMO model, there are three sub-models which apply throughout the lifecycle of the project. They include; the basic model, the intermediate model, and advanced model. The basic model is applicable in early life of a project. It is essential in providing a rough estimate of what is to be expected in the later stages of the project. It offers a glimpse into what is to be expected and how activities need to be undertaken to meet a set of requirements [<xref ref-type="bibr" rid="scirp.93452-ref11">11</xref>] . The intermediate model, on the other hand, is different in its way in that it deals with the estimation of value and time after more details about the project have been acquired. With the detailed requirements, this model can effectively initiate the cost estimation process. The advanced model comes as the last applicable model with COCOMO 81. It only applies upon completion of the project. It offers a more refined estimate that is useful and reliable.</p><p>There are different equations that are used with COCOMO 81. The two equations that are used with COCOMO help in calculating the effort and scheduled time. The estimated schedule time is measured in months. The two equations [<xref ref-type="bibr" rid="scirp.93452-ref11">11</xref>] include;</p><p>PM = a ( KDSI ) * EAF (1)</p><p>TDEV = c ( PM ) d (2)</p><p>Each abbreviation in the equations stands for a factor that affects the cost of software. Thus, they present dynamics that help in cost estimation.</p><p>PM =&gt; Person-Months</p><p>EAF =&gt; Effort Adjustment Factor</p><p>TDEV =&gt; Scheduled time</p><p>KDSI =&gt; Number of lines of code (that is denoted in thousands)</p><p>Initials a, b, c and d are constants which result from the mode that is used in estimating cost.</p><p>The modes include organic, semi-embedded, as well as embedded. On the same note, there are cost drivers that related to COCOMO 81 [<xref ref-type="bibr" rid="scirp.93452-ref11">11</xref>] . For instance, the EAF serves the purpose of tailoring the estimate so that conditions that affect the development of the environment are considered. When it comes to the intermediate model, a total of 15 drivers are present that affect cost which can be manipulated into helping to calculate for the EAF [<xref ref-type="bibr" rid="scirp.93452-ref11">11</xref>] . The 15 cost drivers are segmented into four sections which are product attributes, personal attributes, computer attributes and project attributes [<xref ref-type="bibr" rid="scirp.93452-ref11">11</xref>] . Depending on the impact that the cost drivers have on the project development, they are categorized ranging from very low to extra high [<xref ref-type="bibr" rid="scirp.93452-ref11">11</xref>] .</p><p>On a different note, various advantages and disadvantages are associated with the use of COCOMO 81. Its major advantage is that it is simple to estimate cost with this technique [<xref ref-type="bibr" rid="scirp.93452-ref11">11</xref>] . It is useful to estimate in large projects and takes less time to estimate the cost. On the other hand, there are various disadvantages that are associated with the use of this model. With the Constructive Cost Model, there is a disadvantage of estimation failures. Estimation failures come about in that estimation is carried out at the early stages of the project. This leaves room for many errors, and most of these errors can result in failures in estimation. At the early stages of the project many variables have not evolved, and thus they cannot be considered in the estimation process. All the same, the early stages do not offer sufficient ground and room for cost estimation that is reliable. In doing so, this form of cost estimation is not reliable and cannot meet the needs of the current software dynamics.</p><p>3) COCOMO (Cost Constructive Model) II</p><p>COCOMO II is different from COCOMO 81 considering that is addresses most of the problems and challenges that were associated with COCOMO 81 is its inability to offer reliable estimates due to many failures [<xref ref-type="bibr" rid="scirp.93452-ref12">12</xref>] . Additionally, the inability of COCOMO 81 to gather all input parameters relating to size and time was another problem. However, COCOMO II addresses these issues considering that it does not base its estimates at the earlier stage of the project [<xref ref-type="bibr" rid="scirp.93452-ref12">12</xref>] . In addition to the above, COCOMO II also aims to develop database for software cost estimates that are characterized with tool capabilities meaning that it can enhance and ensure there is the model improvement. Thirdly, and lastly, the model also aims to allow the provision of the quantitative and analytical framework that enhances the evaluation of effects that result from software technology improvements. This distinct model is not very traditional considering that it allows for the incorporation of technological changes when it comes to the evaluation process. Thus, it is effective in that makes it possible to carry out an evaluation based on additions made to a software project.</p><p>All the same, [<xref ref-type="bibr" rid="scirp.93452-ref12">12</xref>] explains that this approach is not very different from COCOMO 81. In other terms, the minor changes that are notable in this model are the use of a higher number of cost drivers. The cost drivers that are used are different as compared to those that were used in COCOMO 81. When it comes to calculating to come up with the estimates, the different approach that is undertaken is that variables are used instead of constants. In addition to the above, lines of code are used as the main metric as opposed to function points that are used as the main metric in COCOMO 81 [<xref ref-type="bibr" rid="scirp.93452-ref12">12</xref>] . However, function points can also be used in place of lines of code for making estimates. In this case, the line of code metric tools is customized to act as the LOC. The three models that are characteristic of COCOMO II are Application Composition Model, Early Design Model, and Post-Architecture Model. The Application Composition Model is used for projects that have been built for rapid application [<xref ref-type="bibr" rid="scirp.93452-ref12">12</xref>] . Object points are useful when it comes to ascertaining size estimates. Prototyping efforts are employed to help in resolving high-risk issues. Thus, it easily meets the expectations and the needs of many users making it one of the most used traditional cots estimation technique favorable for cost estimation of software being currently developed.</p><p>Various advantages and disadvantages are associated with COCOMO II. One of its advantages is that COCOMO II has a calibration process that is clear and effective. It is clear and effective in that it has systems in place which clearly define the main metrics and the variables that are used are more detailed. In addition to the above, there is the advantage of allowing industries to function more effectively due to their ability to adopt a model that is flexible to changes. Thus, COCOMO II is an industry model.</p><sec id="s3_1_1"><title>3.1.1. Function Point Analysis</title><p>Due to diverse functional aspects in software systems, proper metric systems remain to be a major concern in software engineering. Therefore, software engineers focus on measuring the functionality size in software development projects. In reducing the complex task of software metrics in terms of functional size, functional point analysis method was developed [<xref ref-type="bibr" rid="scirp.93452-ref13">13</xref>] . Functional point analysis refers to the standardized methods of determining software sizes by using functional constraints which determine the key features to be designed [<xref ref-type="bibr" rid="scirp.93452-ref14">14</xref>] . Significantly, this method is universal because its application is not limited to programming languages and technologies. In function point analysis, two major components are measured. The most crucial aspects of software application measure in function point analysis comprise the data functionality and transaction functionality attributes [<xref ref-type="bibr" rid="scirp.93452-ref13">13</xref>] . Precisely, the metrics of the two basic features are determined by evaluating the scope of the system product, quality indicators, productivity, and the system performance.</p><p>Function point analysis (FPA) evaluates the system’s metrics from a functional perspective, thereby resolving issues associated with technology dependency in the development lifecycle [<xref ref-type="bibr" rid="scirp.93452-ref13">13</xref>] . The efficiency of FPA in software engineering is achieved through a comprehensive analysis of applications in three stages [<xref ref-type="bibr" rid="scirp.93452-ref13">13</xref>] . The first stage of function point analysis concerns identifying the forms of transactions to be made in the software applications. Secondly, the engineers evaluate and appraise the components of the software system. Lastly, the process involves evaluations of the general system characteristics. Fundamentally, general system characteristics are categorized using into 14 main features [<xref ref-type="bibr" rid="scirp.93452-ref13">13</xref>] . These are data processing, system performance, hardware configurations, transaction rates, data entry, end-user efficiency, online updates, reusability, ease of use, support of multiples sites and change facilitations [<xref ref-type="bibr" rid="scirp.93452-ref13">13</xref>] . There has also been tremendous research in enhancing function point analysis with non-functional requirements.</p></sec><sec id="s3_1_2"><title>3.1.2. Putnam’s Model</title><p>The Putnam’s model is a dynamic multivariate model which was developed by Larry Putnam in the late 1970s for effort estimation [<xref ref-type="bibr" rid="scirp.93452-ref15">15</xref>] . The model functions by examining the many software projects and analyzing the distribution of manpower. The relationship between the size and effort is non-linear and is highly sensitive to delivery time. It provides a simple and computationally plausible way of predicting software costs. It is used to calculate both effort and time required to complete a software project based upon the specified size of the project. The Putnam’s model equation is given as [<xref ref-type="bibr" rid="scirp.93452-ref15">15</xref>] :</p><p>B 1 / 3 ∗ Size Productivity = Effort 1 / 3 ∗ Time 4 / 3 (3)</p><p>With this model, SLIM is the tool that is useful when it comes to cost estimation and allows workforce scheduling. Thus, it achieves the objectives of cost estimation as it provides the schedule and cost of the essential resources. However, it’s capacity in estimating the total manpower requirements and development time at an early stage is still not satisfactory [<xref ref-type="bibr" rid="scirp.93452-ref16">16</xref>] . It has the setback of not accounting for other aspects of the software project [<xref ref-type="bibr" rid="scirp.93452-ref16">16</xref>] . There are a series of aspects relating to software development especially in its life cycle that needs to be considered. The uncertainty in the size of the software may lead to the inaccurate cost estimation. According to [<xref ref-type="bibr" rid="scirp.93452-ref17">17</xref>] SLIM’s error percentage is said to be 772.87%. On the contrary, its advantage is that the model is based on two variables that are keys to cost estimation which are time and size [<xref ref-type="bibr" rid="scirp.93452-ref16">16</xref>] and it needs fewer parameters compared to COCOMO 81 and COCOMO II.</p></sec></sec><sec id="s3_2"><title>3.2. Non-Algorithmic Methods</title><sec id="s3_2_1"><title>3.2.1. Expert Judgement</title><p>Expert judgment is one of the traditional techniques that are used in the software cost estimation in the early phases of the software development [<xref ref-type="bibr" rid="scirp.93452-ref3">3</xref>] . The technique is such that it relies heavily on the expertise and the experience of an expert at cost estimating. It depends on the domain knowledge of the expert rather than the historical data [<xref ref-type="bibr" rid="scirp.93452-ref3">3</xref>] . The experienced estimator is tasked with the responsibility of estimating the cost of software based on the fact that they have sufficient knowledge that ensures cost estimation is as accurate as possible [<xref ref-type="bibr" rid="scirp.93452-ref3">3</xref>] . On the same note, the expert judgement used to estimate cost on a given project is often limited to a specific filed of expertise of the estimator. An expert who has knowledge and experience on a specific field is more likely to have majored in the given project putting them in a better position to estimate the cost. Expert judgement comes in handy especially when there are limitations that limit effective and efficient data collection [<xref ref-type="bibr" rid="scirp.93452-ref18">18</xref>] . In doing so, expert judgement is required to make decisions based on the limitations and the stringent factors. Delphi technique is one of the examples that follow the expert judgement approach. An expert can be able to offer an honest and experienced opinion on the best course of action when it comes to an understanding of the impacts of a system being incorporated. There is also a setback of tedious processes that are undertaken in documenting the factors that an expert requires to make the judgement [<xref ref-type="bibr" rid="scirp.93452-ref3">3</xref>] . Most of the factors in many software projects are many and documenting them is tedious as well as difficult. There is no specific validation for this approach as the estimation depends solely on the domain knowledge and previous experiences of the expert.</p><p>Additionally, there is the disadvantage of obtaining cost estimates that are biased and optimistic. Expert judgement is given by experts who have human emotions and is very likely that their emotions influence the judgement process [<xref ref-type="bibr" rid="scirp.93452-ref3">3</xref>] [<xref ref-type="bibr" rid="scirp.93452-ref18">18</xref>] . Therefore, there is the possibility that pessimism, optimism, as well as bias might influence the judgement process.</p></sec><sec id="s3_2_2"><title>3.2.2. Top-Down Estimation</title><p>Top Down cost estimation approach focuses on estimating the cost of a project from the global properties of the overall project and using either algorithmic such as Putnam model or non-algorithmic methods [<xref ref-type="bibr" rid="scirp.93452-ref19">19</xref>] . The estimation is then split into various components in proportion [<xref ref-type="bibr" rid="scirp.93452-ref19">19</xref>] . This method can be followed when there is limited historical data available about the similar project. This technique is more beneficial while the project is still in its early stages. This is because, at this stage, there is no need for detailed information about the project [<xref ref-type="bibr" rid="scirp.93452-ref19">19</xref>] .</p><p>Top Down estimation is used for high-level decisions when the planning horizon is quite long. This technique is used when there is very little specific project information where we can get a ballpark estimate. Unlike other cost estimation techniques, this approach focuses on activities like management and integration which are overlooked in other techniques. The major disadvantage is that the ballpark estimate is very inaccurate [<xref ref-type="bibr" rid="scirp.93452-ref19">19</xref>] .</p></sec><sec id="s3_2_3"><title>3.2.3. Bottom-Up Estimation</title><p>This is the exact opposite of the top-down estimation methodology. In this particular technique, the cost of every component of the software is derived and then the final result is obtained by combining these elements to get the total estimated cost of the project [<xref ref-type="bibr" rid="scirp.93452-ref19">19</xref>] . The aim of this technique is to obtain a proper estimate that will be an accumulation of the estimates of the smaller components of the software. Depending upon the variety of the projects both the methodologies are useful [<xref ref-type="bibr" rid="scirp.93452-ref19">19</xref>] . The estimation methodology is best suited for small projects to estimate the cost.</p><p>In the bottom-up estimation, we take each work package or each activity in a schedule and assign a dollar cost to that and then add them up to get the overall cost of the project [<xref ref-type="bibr" rid="scirp.93452-ref19">19</xref>] . This can create a very detailed estimate that will be very accurate. However, there are some disadvantages of the bottom-up estimate.</p><p>1) If we add up each activity, we may be lacking any coordination between the activities such as resource overlap or if some of the dependencies that one activity has to be finished before another can start.</p><p>2) The accuracy is high, but it can also take a lot of time to create this level of estimate, which means that it can be expensive to create.</p><p>So, to overcome these disadvantages it is preferable to do a high-level estimate i.e. top-down estimate for the whole project and then do some detailed estimates for the activities that are coming up in the short term.</p></sec><sec id="s3_2_4"><title>3.2.4. Price-to-Win Estimation</title><p>In this approach, the estimation of the software project is directly proportional to the budget of the customer. The customer’s budget is more focused rather than the functionality of the software and the project costs whatever the customer has to spend on it [<xref ref-type="bibr" rid="scirp.93452-ref19">19</xref>] . This approach is not so recommended as instead of focusing on software functionality it focuses more on the client’s budget and capacity. Accuracy varies drastically based on the client’s budget and therefore it is rated as low accuracy. This approach does require extensive data or no previous data at all as the client will convey the requirements and does not depend on the historical data [<xref ref-type="bibr" rid="scirp.93452-ref19">19</xref>] . It is not a good practice as it may cause a delay in the development and delivery and may also force the development team to work overtime. The validation of this approach is based on the customer’s budget and person-month factor.</p><p>On the contrary, the major advantage of the price-to-win approach is that it the estimation doesn’t exceed the customer’s budget. And as the functionality of the software is restricted with the customer’s budget the quality of the software is compromised [<xref ref-type="bibr" rid="scirp.93452-ref19">19</xref>] .</p></sec></sec><sec id="s3_3"><title>3.3. Learning Oriented Models</title><p>Machine Learning can be said to be a method of data analysis that is used to automate analytical model. It is also a branch of Artificial Intelligence (AI) which can be said to be based on the notion that machines such as robots and other computerized devices can learn from the data [<xref ref-type="bibr" rid="scirp.93452-ref20">20</xref>] . These approaches have the ability to learn from previous data and predict the future outcome based on the previous data. Some of the common machines learning algorithms that have been exploited in the software cost estimation are neural networks, fuzzy logic, genetic algorithms, Bayesian networks, support vector regression and analogy based. Researchers have proposed several machine learning software cost estimation models in order to improve the estimation accuracy [<xref ref-type="bibr" rid="scirp.93452-ref2">2</xref>] . And several studies have also reported that under the same model when the model is constructed with different historical project datasets or different experimental designs there is a variation in the predicted accuracy [<xref ref-type="bibr" rid="scirp.93452-ref2">2</xref>] [<xref ref-type="bibr" rid="scirp.93452-ref21">21</xref>] [<xref ref-type="bibr" rid="scirp.93452-ref22">22</xref>] [<xref ref-type="bibr" rid="scirp.93452-ref23">23</xref>] [<xref ref-type="bibr" rid="scirp.93452-ref24">24</xref>] [<xref ref-type="bibr" rid="scirp.93452-ref25">25</xref>] .</p><p>A number of studies have concluded that machine learning models in software cost estimation outperform non-machine learning models [<xref ref-type="bibr" rid="scirp.93452-ref1">1</xref>] [<xref ref-type="bibr" rid="scirp.93452-ref2">2</xref>] [<xref ref-type="bibr" rid="scirp.93452-ref9">9</xref>] [<xref ref-type="bibr" rid="scirp.93452-ref10">10</xref>] . The use of these machine learning models in the software cost estimation process has gained significant popularity due to the wide margin of error in the classical estimation models [<xref ref-type="bibr" rid="scirp.93452-ref26">26</xref>] . There are remendous and continuous improvements in the machine learning algorithms which assists in achieving more accurate predictions when these are applied [<xref ref-type="bibr" rid="scirp.93452-ref2">2</xref>] [<xref ref-type="bibr" rid="scirp.93452-ref26">26</xref>] . These learning models consistently predict accurate results because of its learning nature from the previously completed projects. According to Monika et al. [<xref ref-type="bibr" rid="scirp.93452-ref10">10</xref>] investigation, it had been concluded that ANN was the prominent methods for developing estimating models. A brief overview of these models along with their strengths and weakness in the context of software cost estimation has been discussed. This will aid the researchers to understand which approach suits best for their methodology. Different machine learning models have different strengths and limitations and thus favor different estimation contexts [<xref ref-type="bibr" rid="scirp.93452-ref2">2</xref>] .</p><sec id="s3_3_1"><title>3.3.1. Artificial Neural Networks</title><p>ANN is one of the major approaches that are exploited in the sector of machine learning models. As the name suggests, it is usually inspired by the neural part of the brain system with an intention of imitating an intelligent living organism [<xref ref-type="bibr" rid="scirp.93452-ref27">27</xref>] . It is composed of two layers, that is the input and output layers, within the layers there is a hidden layer which constitutes of units whose main purpose is to assign weights to the data that is fed from the input. These weights are assigned randomly to the data [<xref ref-type="bibr" rid="scirp.93452-ref28">28</xref>] . They are among the exploited tools in cost estimation due to their excellent performance [<xref ref-type="bibr" rid="scirp.93452-ref1">1</xref>] . Neural networks have existed for long as they can be traced back to the year 1940 although the best way to exploit them was the missing link which has been currently accomplished. There are different types of neural networks that exist with each of them having a defined level of complexity coupled with use. The most common and general type of neural network (NN) is the feed forward neural network as the name suggests data travels in only one direction which is from input to output [<xref ref-type="bibr" rid="scirp.93452-ref27">27</xref>] . Other types include Recurrent Neural Network (RN), Convolution Neural Networks (CNN), LTSM Recurrent Neural Networks, etc. The ANN approach is used in cost estimation where the pattern of certain data requires to be understood prior to estimating the project cost. In the cost estimating field, it is used to classify data into predefined classes, clustering, and prediction [<xref ref-type="bibr" rid="scirp.93452-ref28">28</xref>] . For instance, in projects such as the stock exchange market, it is used to make predictions in the market through the exploitation of the previous data.</p><p>Artificial Neural Networks are usually excellent in capturing non-linear relationships which makes it an iconic advantage of the model. The model has a deep neural net which requires fewer features [<xref ref-type="bibr" rid="scirp.93452-ref27">27</xref>] . The neural net has the ability to develop its own features which are also an advantage when limited features are available in the dataset for training the model. It is also flexible in the sense that there are a plethora of features one can choose from which include CNN’s, RNN’s, LSTM RNN’s, etc. (<xref ref-type="fig" rid="fig3">Figure 3</xref>).</p><p>It has also some limitations whereby sometimes they tend to over-fit the data in the software cost estimation process [<xref ref-type="bibr" rid="scirp.93452-ref2">2</xref>] . Additionally, the model also requires an enormous amount of power for computation which may not be available at all times. It was better if the model is in a position to efficiently operate even in places where there is less consumption power as putting up the high computation power is costly to the involving firm.</p><p>Several Artificial Neural Networks models have been used in the software cost estimation processes which are mostly common according to the investigation [<xref ref-type="bibr" rid="scirp.93452-ref9">9</xref>] :</p><p>a) Feed-forward neural network</p><p>b) Recurrent Neural Network</p><p>c) Radial basis function (RBF) network</p><p>d) Neuro-fuzzy networks</p><p>Hamza et al. [<xref ref-type="bibr" rid="scirp.93452-ref9">9</xref>] concluded that choosing the right artificial neural network model is essential to get the accurate estimations. In their study [<xref ref-type="bibr" rid="scirp.93452-ref9">9</xref>] they have also concluded that feed-forward neural network works better than other models in ANN’s but, the technique needs data filtering to prevent noise. In the case of noisy data, the radial basis function is more suitable.</p></sec><sec id="s3_3_2"><title>3.3.2. Genetic Algorithms</title><p>They are defined as adaptive and heuristic search algorithms which are a subject to the theory of natural selection by Darwin. In a current, study they are described as one of the most active areas of research which have been designed through nature-inspired metaheuristics [<xref ref-type="bibr" rid="scirp.93452-ref29">29</xref>] . Genetic Algorithm (GA) forms one of the soft computing techniques in software cost estimation process whereby its main role is to change certain parameters of classical methods such as COCOMO approach to predict software cost in a more accurate manner [<xref ref-type="bibr" rid="scirp.93452-ref29">29</xref>] . GA has widely been exploited in different fields of cost estimation such as correcting the identification system, path-searching problems within a project [<xref ref-type="bibr" rid="scirp.93452-ref29">29</xref>] . Additionally, GA has been used to solve a variety of NP-hard computational problems [<xref ref-type="bibr" rid="scirp.93452-ref29">29</xref>] . This model takes inspiration from nature such as bullet train design based on fish. Used for optimization problems (Np-problem) few of them are firefly algorithm particle swarm optimization and cuckoo search just to mention a few.</p><p>GA model usually exploits the optimization problem by using an evolutionary process. The first benefit of the model is the fact that it is easy to set up than neural networks but is less flexible. Once the algorithm begins it is on its own. It learns its own features; hence we don’t have to supervise the process. On the other hand, it has a disadvantage which includes less flexibility, many hyper parameters which includes preference of functions, reproduction rates, the percentage of elitism and cross over, dealing with out of bound conditions, creating a strategy and setting the required tree sizes and depths within the model [<xref ref-type="bibr" rid="scirp.93452-ref30">30</xref>] . The model has no guarantee of finding an optimal solution, infinite time because it has asymptotic convergence, containing a number of parameters, sometimes the result is highly dependent on the parameters set. It has also self-adaptive parameters. It is computationally expensive and Meta models of functions are used in this process too.</p><p>The original use of the model was to establish manpower required to complete a certain project [<xref ref-type="bibr" rid="scirp.93452-ref29">29</xref>] . It is also used in scheduling different tasks within the field of cost estimation which implies that it has the ability to determine and evaluate a variety of tasks within a single project. It has also been used in mining data within the same scope which is considered as a complex exercise in case traditional methods are exploited in place. On the extreme end GA has been used in optimizing distributed tasks within the software cost estimation process. It has been used to solve distributed queries which imply that all relevant queries about a certain project can be accessed during the initial steps of its development. It is also easy to make assumption and predictions to the software being developed while using this method. It is somehow time conservative especially when operated by experienced personnel.</p></sec><sec id="s3_3_3"><title>3.3.3. Fuzzy Logic</title><p>The model is a computing approach that is based on the degrees of truth rather than the unusual true or false normally referred to as Boolean logic which most of the modern computers are based on [<xref ref-type="bibr" rid="scirp.93452-ref31">31</xref>] . Additionally, it can be said to be an approach to computing based on many-valued function for instance, instead of a task completed or not, one can say 50% is completed. Fuzzy Logic (FL) approach gives an acceptable but definite output which is usually in response to inaccurate (fuzzy), distorted, incomplete and ambiguous input. FL was developed in 1965 by Loft Zadeh from a concept of fuzzy set theory. The first system is the one that is exploited in the estimation sector as most of the systems produce crisps data as input and expect the same type of data as output. There are three steps that have to be followed while using FL; the first is the Fuzzification which converts a crisp into a fuzzy set [<xref ref-type="bibr" rid="scirp.93452-ref31">31</xref>] . The second is the Fuzzy Rule-based System, at the step after all the crisp input has been fuzzified into their respective linguistic values the inference engine then derives their linguistic values [<xref ref-type="bibr" rid="scirp.93452-ref32">32</xref>] . Defuzzification is the final step which involves the conversion of fuzzy output into crisp output.</p><p>Fuzzy logic is deployed for decision making whereby it can be implemented with various sizes and abilities ranging from small microcontrollers to large workstation-station based software development. In the sector of software cost estimation, FL has been used to give acceptable reasoning although it does not guarantee accurate reasoning [<xref ref-type="bibr" rid="scirp.93452-ref32">32</xref>] . This approach is very easy to use despite the fact that it has numerous functions that it can execute when it comes to cost estimation of the software development [<xref ref-type="bibr" rid="scirp.93452-ref32">32</xref>] . Also, with the approach, it is advantageous as estimators of software are in a position to estimate and give an anticipation of the different aspects of the project even before the initiation level. The method has also a fast learning ability compared to other methods under the same scope [<xref ref-type="bibr" rid="scirp.93452-ref31">31</xref>] . The major disadvantage that the approach lacks guarantee of accuracy which is a vital part of the software cost estimation as to avoid incurring financial losses. Through the defuzzification process, the model is able to interact the output into a numerical value as per the desire.</p><p>There are a number of three applied approaches under this method which include; 2FA-kmodes it is used in clustering where numerical datasets are shown through fuzzy sets however the fuzz k-modes algorithm is used to maintain categorical attributes. The second approach is 2FA-kprototype it is used in clustering also where project data are clustered into homogeneous sets [<xref ref-type="bibr" rid="scirp.93452-ref31">31</xref>] . The final approach is the classical analogy it is used to predict the effort of the target project using information from former similar projects. The first two models 2FA-knodes, 2FA-kprototypes perform same and are both better than classical analogy. The similarity in the functionality of the two models is technically for the purposes of increasing efficiency within the model [<xref ref-type="bibr" rid="scirp.93452-ref32">32</xref>] . The major disadvantage associated with the two models is that they are not very good in COCOMO dataset and classical analogy doesn’t recognize categories as “less risky”, “medium”, and “high”. Thus, there is a need for implementations to be conducted for the model so as to enable future proper operation in case they are deployed in COCOMO and other datasets.</p></sec><sec id="s3_3_4"><title>3.3.4. Bayesian Networks</title><p>The model is also referred to as a probabilistic directed acyclic graphical model. It exploits graphical models to represent sets of variables coupled with their conditional dependencies through a directed acyclic graph (DAG) [<xref ref-type="bibr" rid="scirp.93452-ref33">33</xref>] . In other terms it uses Bayesian inference to perform probability computations they aim at modeling conditional dependence while developing software cost estimation which is usually represented by the edges within a directed graph. In addition to that, the model uses three main inference tasks: inferring unobserved variables, parameter learning and structure learning [<xref ref-type="bibr" rid="scirp.93452-ref33">33</xref>] . Through developers of software understanding the different relationships that exist in the model, they can efficiently conduct inference on random variables within software [<xref ref-type="bibr" rid="scirp.93452-ref33">33</xref>] . Each edge in the model corresponds to a conditional dependency while at the same time each node corresponds to a unique random variable. Thus, a specific pattern is usually followed in the model so as to attain its functionality within the software cost estimation process.</p><p>In a recent study, results have indicated that the inclusion of Bayesian Network, a machine-learning model in software cost estimation models it could improve on the levels of accuracy in the project [<xref ref-type="bibr" rid="scirp.93452-ref4">4</xref>] . This approach is able to break down project data about the cost into a form that is easy to analyze while at the same time it offers cost items coupled with the probability which can capture the uncertainty of different items within the software to be developed [<xref ref-type="bibr" rid="scirp.93452-ref33">33</xref>] . In software cost estimation sector, it is exploited in initial planning also as it has the ability to evaluate different parameters of planning such as time, cost coupled with resources to be exploited in the development of software. This model only allows exactly tested cost estimation information to be integrated with natural master feelings.</p><p>The best advantage with exploiting this model is that it gives a better assurance on accuracy and it is also able to improve on the overall quality of software that is being developed. Most of the software cost estimation models have some difficulties in offering assurance which makes the model outstanding compared to the rest. It is also a model that saves time which is of great importance when it comes to software project development [<xref ref-type="bibr" rid="scirp.93452-ref4">4</xref>] . Since the model encodes all variables, then it is also able to handle different types of missing data. In case it is deployed learning casual relationships, they help better understand a problem domain as well as forecast the consequences should there be an intervention [<xref ref-type="bibr" rid="scirp.93452-ref4">4</xref>] . Through probabilistic and casual semantics, it is ideal to use the model for representing prior data and knowledge.</p><p>Despite it being a complex approach in software cost estimation it is easy to use even with staffs who have less experience in the field [<xref ref-type="bibr" rid="scirp.93452-ref33">33</xref>] . On the other hand, there are some disadvantages while using this model such as Information theoretically infeasible it turns out that specifying a prior is extremely difficult challenge. Thus, being in a position to efficiently exploit this model, then users have to be well familiarized with the model’s mode of learning. Being in a position to understand different languages, the model is able to execute more functions which are advantageous to the project running. Acquiring this knowledge takes some significant effort, this implies a lot of time will have to be consumed during the training period. Despite a lot of time being spent on training, also on the other hand it has to be the required training hence improper training would ruin the entire model. The models also do not always tell the truth as, they don’t specify their actual prior, but they give credentials to the convenient one.</p></sec><sec id="s3_3_5"><title>3.3.5. Support Vector Regression</title><p>It is also referred to as support vector networks or machines which are supervised models which are capable of learning algorithms and finally analyzing the same data that will be exploited for regression and classification analysis [<xref ref-type="bibr" rid="scirp.93452-ref34">34</xref>] . In other description, it is said to be a concept a set of related supervised learning methods usually used to analyze data coupled with recognized patterns. This model when applied in software cost estimation it takes a set of input data then it gives a prediction of each input [<xref ref-type="bibr" rid="scirp.93452-ref4">4</xref>] . The input must be a member of the support vector machine (SVM) which makes the model to be a non-probabilistic binary linear classifier [<xref ref-type="bibr" rid="scirp.93452-ref34">34</xref>] . In software cost estimation, it is used to solve problems related to pattern classification within the software. So as to apply it effectively in software cost estimation developers have to be in a position to design questions based on a problem and the design involved within the software.</p><p>The model has over the year’s utilized regression and classification as the main mode of analyzing data within software. In China, the model was famously used in predicting software development efforts [<xref ref-type="bibr" rid="scirp.93452-ref34">34</xref>] . Support vector reasoning has been found to be helpful in hypertext and text categorization within the software. This model is usually utilized in software cost estimation so as to improve on the levels of accuracy compared to the classical query refinement schemes. It is also exploited in image segmentation in different software depending on the role that the software is designed to perform [<xref ref-type="bibr" rid="scirp.93452-ref34">34</xref>] . With the model in operation, the software development process automatically avoids over fitting of data. Thus, we can technically say that the model is conservative when it comes to financial and time while planning for a project.</p><p>A major advantage of the model is that it has the ability to work best with text data known as string kernel [<xref ref-type="bibr" rid="scirp.93452-ref4">4</xref>] . On the other hand, there it has been found to be occupying a huge memory and it does not operate well when the data set is huge. Occupying a huge memory concerning storage of data sets makes it inefficient a factor that they need to reconsider. The fact that it is occupying a huge data set for storage makes is not being able to meet the cost-effective criterion which should be a model for each model. In this model, there are implementations and improvements which have been done on it such as dimensionality reduction such as PCA and ICA. The implementation has been done on the model so as to improve on the functionality of the model.</p></sec><sec id="s3_3_6"><title>3.3.6. Regression Tree</title><p>It is also referred to as a decision tree, is a logical model most exploited for decision analysis in software cost estimation. It is a common approach exploited in software cost estimation based on a number of factors and their consequences [<xref ref-type="bibr" rid="scirp.93452-ref35">35</xref>] . In other descriptions, the model is perceived as a procedure that is used for classification and regression in the sector of software cost estimation process [<xref ref-type="bibr" rid="scirp.93452-ref35">35</xref>] . The model is usually in the form of a tree structure with different internal nodes that stands for a test on an attribute. Individual branch on the tree usually represents an outcome of the test and the class label is held on the leaf node. Looking at the tree from a technical perspective, it is evident that all the branches stand for a specific function within the software development process. The model has usually trained through top-down training method which is specific to a certain direction [<xref ref-type="bibr" rid="scirp.93452-ref35">35</xref>] . However, current research shows there are efforts that are aimed at training the model so as to operate in a multidirectional purpose. This will be a breakthrough within the sector of software cost estimation as it means the model will be able to execute more functions than it does as for now [<xref ref-type="bibr" rid="scirp.93452-ref35">35</xref>] . Over the years it has been used to predict the amounts of effort required to develop a software system which has proved to be a good model as it saves on time and budgets that have been planned on. Thus, we can technically say that the model is cost efficient as it saves on the factors that are likely to jeopardize a developing project.</p><p>The regression tree has been widely used in inductive learning, in turn, it has proved to be good in terms of predictive accuracy especially in software cost estimation process despite its complexity [<xref ref-type="bibr" rid="scirp.93452-ref35">35</xref>] . One of the main advantages is the fact that it is fast compared to neural networks and support vector reasoning. The efficiency that is instilled in the model is because it has a simple structure which executes diverse roles is simple steps. Also, while exploiting the model they tend to be very interpretable thus one can almost make an anticipation of what is to be formed in the final project. This implies that there are possibilities of making assumptions on the outcome of the project that is being worked on.</p></sec><sec id="s3_3_7"><title>3.3.7. Analogy Based</title><p>The analogy was used in the past as a method of supporting or supporting explanations of a particular natural phenomenon which was later incorporated in the software development sector of software development of for the purposes of software cost estimations [<xref ref-type="bibr" rid="scirp.93452-ref36">36</xref>] . Analogy Based is one of the most efficient approaches in software cost estimation due to its outstanding performance and capability of handling complex datasets [<xref ref-type="bibr" rid="scirp.93452-ref37">37</xref>] . The model exploits comparison as the main form of a subject to compare software project under considerations with past historical projects which have prior known characteristics, schedule and efforts. Thus, this model has to rely on previous projects similar to the ones that are to be developed thus the levels of accuracy are usually very high [<xref ref-type="bibr" rid="scirp.93452-ref37">37</xref>] .</p><p>Conventional analogy-based models usually deploy the same number of analogies in all the projects that it is involved with similar data sets which improves on the estimation approximations. There are some researchers that claim using the same number of analogies could only improve on the data sets only and not the entire project. Thus, when exploiting this model, it is important to understand the</p><table-wrap id="table1" ><label><xref ref-type="table" rid="table1">Table 1</xref></label><caption><title> Software cost estimation techniques comparison</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >N0</th><th align="center" valign="middle" >Method</th><th align="center" valign="middle" >Type</th><th align="center" valign="middle" >Strengths</th><th align="center" valign="middle" >Weakness</th><th align="center" valign="middle" >Accuracy</th><th align="center" valign="middle" >Data</th><th align="center" valign="middle" >Validation</th></tr></thead><tr><td align="center" valign="middle" >1</td><td align="center" valign="middle" >COCOMO- II</td><td align="center" valign="middle" >AM</td><td align="center" valign="middle" >Simple to carryout estimations; takes less time and effort to estimate; Useful in large projects</td><td align="center" valign="middle" >Details of the past projects needed to estimate; May leave out hidden costs hence lead to more expensive estimation; Calibration is required; Cannot meet current software standards</td><td align="center" valign="middle" >MA</td><td align="center" valign="middle" >ED</td><td align="center" valign="middle" >EV</td></tr><tr><td align="center" valign="middle" >2</td><td align="center" valign="middle" >FPA</td><td align="center" valign="middle" >AM</td><td align="center" valign="middle" >Easy to estimate development costs at the requirements gathering stage/initial stage; Independent of languages and tools used</td><td align="center" valign="middle" >Quality attributes, development time and man power are not considered</td><td align="center" valign="middle" >MA</td><td align="center" valign="middle" >LD</td><td align="center" valign="middle" >EV</td></tr><tr><td align="center" valign="middle" >3</td><td align="center" valign="middle" >PM</td><td align="center" valign="middle" >AM</td><td align="center" valign="middle" >Estimation depends on only time and size which are key to cost estimation; Needs fewer parameters compared to COCOMO 81 and COCOMO II</td><td align="center" valign="middle" >Does not consider other important aspects of the software development</td><td align="center" valign="middle" >MA</td><td align="center" valign="middle" >LD</td><td align="center" valign="middle" >EV</td></tr><tr><td align="center" valign="middle" >4</td><td align="center" valign="middle" >EJ</td><td align="center" valign="middle" >NA</td><td align="center" valign="middle" >Best technique where limited data is available; Experience of experts makes the estimation more accurate and realistic</td><td align="center" valign="middle" >Biased opinions of the experts; Difficult to document the parameters used by the experts to estimate; Experts require experience of similar projects</td><td align="center" valign="middle" >LA</td><td align="center" valign="middle" >LD</td><td align="center" valign="middle" >EV</td></tr><tr><td align="center" valign="middle" >5</td><td align="center" valign="middle" >TDE</td><td align="center" valign="middle" >NA</td><td align="center" valign="middle" >Only few details are required to estimate; Simple and less time consuming</td><td align="center" valign="middle" >Difficult to identify low level problems which causes under estimation; Less details may overlook important attributes of the project</td><td align="center" valign="middle" >LA</td><td align="center" valign="middle" >LD</td><td align="center" valign="middle" >n/a</td></tr><tr><td align="center" valign="middle" >6</td><td align="center" valign="middle" >BUE</td><td align="center" valign="middle" >NA</td><td align="center" valign="middle" >More stable compared to top down approach; Errors are estimated and very stable</td><td align="center" valign="middle" >Development time and system-level activities are not considered</td><td align="center" valign="middle" >LA</td><td align="center" valign="middle" >LD</td><td align="center" valign="middle" >n/a</td></tr><tr><td align="center" valign="middle" >7</td><td align="center" valign="middle" >PTWE</td><td align="center" valign="middle" >NA</td><td align="center" valign="middle" >It depends only on the customer budget</td><td align="center" valign="middle" >Costs do not accurately reflect the work required; Low quality system is developed due to client budget constraint</td><td align="center" valign="middle" >LA</td><td align="center" valign="middle" >ND</td><td align="center" valign="middle" >n/a</td></tr><tr><td align="center" valign="middle" >8</td><td align="center" valign="middle" >NN</td><td align="center" valign="middle" >LM</td><td align="center" valign="middle" >Very good in capturing non-linear relationships; Deep neural Nets do not require a lot of features, they come up on their own; There is a lot of flexibility, you could choose different architectures; RNN’s, LSTM, etc.</td><td align="center" valign="middle" >They tend to over fit; They require enormous amount of computation power</td><td align="center" valign="middle" >VHA</td><td align="center" valign="middle" >ED</td><td align="center" valign="middle" >NI, DCS, CVM</td></tr><tr><td align="center" valign="middle" >9</td><td align="center" valign="middle" >GA</td><td align="center" valign="middle" >LM</td><td align="center" valign="middle" >Need not be supervised</td><td align="center" valign="middle" >Too many hyper parameters</td><td align="center" valign="middle" >HA</td><td align="center" valign="middle" >ED</td><td align="center" valign="middle" >CVM, VAF, RWS</td></tr><tr><td align="center" valign="middle" >10</td><td align="center" valign="middle" >FL</td><td align="center" valign="middle" >LM</td><td align="center" valign="middle" >Considers real valued states instead of binary</td><td align="center" valign="middle" >Does not perform well on complex datasets like COCOMO</td><td align="center" valign="middle" >VHA</td><td align="center" valign="middle" >ED</td><td align="center" valign="middle" >JM, DCS, CVM</td></tr><tr><td align="center" valign="middle" >11</td><td align="center" valign="middle" >BN</td><td align="center" valign="middle" >LM</td><td align="center" valign="middle" >Bayesian network encodes all variables; missing data entries can be handled</td><td align="center" valign="middle" >Difficult to model</td><td align="center" valign="middle" >HA</td><td align="center" valign="middle" >ED</td><td align="center" valign="middle" >CVM</td></tr><tr><td align="center" valign="middle" >12</td><td align="center" valign="middle" >SVR</td><td align="center" valign="middle" >LM</td><td align="center" valign="middle" >It works best with text data: string kernel</td><td align="center" valign="middle" >It takes a lot of memory (RAM); Doesn’t scale well when dataset is huge</td><td align="center" valign="middle" >HA</td><td align="center" valign="middle" >ED</td><td align="center" valign="middle" >LOM</td></tr><tr><td align="center" valign="middle" >13</td><td align="center" valign="middle" >RT</td><td align="center" valign="middle" >LM</td><td align="center" valign="middle" >more capable of handling noisy datasets</td><td align="center" valign="middle" >Models are unstable at times, suffer with high variance, low bias. (Keyword: bias-variance tradeoff)</td><td align="center" valign="middle" >MA</td><td align="center" valign="middle" >LD</td><td align="center" valign="middle" >CVM, DCS, LOM</td></tr><tr><td align="center" valign="middle" >14</td><td align="center" valign="middle" >ABE</td><td align="center" valign="middle" >NA</td><td align="center" valign="middle" >Simple and easy to use; No bootstrap cost</td><td align="center" valign="middle" >They do not work with categorical data</td><td align="center" valign="middle" >MA</td><td align="center" valign="middle" >LD</td><td align="center" valign="middle" >EV, ED</td></tr></tbody></table></table-wrap><p>characteristics of each data sets so as to be able to discover the optimum set of analogies for each project [<xref ref-type="bibr" rid="scirp.93452-ref38">38</xref>] . Accuracy levels can only be increased when users of the model are aware of the different data characteristics instilled by each data set. Through the comparison of various datasets, the model is used to approximate the time that a project would take for it to be completed. In cost estimation, the model is relied on to give empirical evidence on the probability of certain software features being accurate.</p><p>It is also flexible and intuitive in nature as it can be applied in a variety of circumstances where other algorithmic modeling and estimating techniques don’t operate [<xref ref-type="bibr" rid="scirp.93452-ref36">36</xref>] . As much as it is said to be dependent upon past history so as to execute its functions, it can still operate even where past data is not available. However, there are still some setbacks while using the model as it lacks appropriate analogs [<xref ref-type="bibr" rid="scirp.93452-ref38">38</xref>] . For instance, despite it being widely used in the industries it has not stipulated standards on how it should be exploited for expert opinion-based opinion. The common advantage that most users of the method are familiar with is the ability of the model to avoid bootstrapping of cost [<xref ref-type="bibr" rid="scirp.93452-ref38">38</xref>] . On the extreme end, it is accompanied by a major disadvantage which is its inability to operate with categorical data. Hence more implementations need to be done on the model so as to enable to operate with all types of data (<xref ref-type="table" rid="table1">Table 1</xref>).</p></sec></sec></sec><sec id="s4"><title>4. Conclusion &amp; Future Work</title><p>In this paper, we presented an overview of the software cost, effort and size estimation techniques based on algorithmic, non-algorithmic and learning-oriented approaches. We also tabulated all the techniques based on their type, strengths, weaknesses, amount of data and validation methods used by them. Our major findings from the previous studies show that the Neural Network based models outperform other cost estimation techniques in terms of accuracy followed by fuzzy logic [<xref ref-type="bibr" rid="scirp.93452-ref39">39</xref>] . Our survey also shows that no one technique is perfect and all of them have their own advantages and disadvantages. We recommend the researchers and practitioners to choose the best technique which fits their best needs. This study may also give an insight into the software cost estimation techniques for the researchers who are new to this area. We found that there has not been much research on estimating the development time or schedule for a project. Our future work may involve researching the application of the LSTM Recurrent Neural Networks model in predicting the time series for the project.</p></sec><sec id="s5"><title>Conflicts of Interest</title><p>The authors declare no conflicts of interest regarding the publication of this paper.</p></sec><sec id="s6"><title>Cite this paper</title><p>Chirra, S.M.R. and Reza, H. (2019) A Survey on Software Cost Estimation Techniques. Journal of Software Engineering and Applications, 12, 226-248. https://doi.org/10.4236/jsea.2019.126014</p></sec><sec id="s7"><title>Abbreviations</title><p>AM Algorithmic Method</p><p>NA Non-Algorithmic Method</p><p>LM Learning-Oriented Method</p><p>LA Low Accuracy</p><p>MA Moderate Accuracy</p><p>HA High Accuracy</p><p>VHA Very High Accuracy</p><p>ND No Data</p><p>LD Limited Data</p><p>ED Extensive Data</p><p>CVM Cross Validation Method</p><p>NI Number of Iterations</p><p>VAF Variance Accounted For</p><p>RWS Roulette Wheel Selection</p><p>JM Jackknife Method</p><p>LOM Leave-one-out Method</p><p>DCS Degree of Confidence and Significance</p><p>EV Empirical Validation</p><p>ED Euclidean Distance</p><p>WBS Work Breakdown Structure</p><p>n/a Not Applicable</p><p>COCOMO Cost Constructive Model</p><p>FPA Function Point Analysis</p><p>PM Putnams’s Model</p><p>EJ Expert Judgement</p><p>TDE Top Down Estimation</p><p>BUE Bottom Up Estimation</p><p>PTWE Price-To-Win Estimation</p><p>NN Neural Networks</p><p>GA Genetic Algorithm</p><p>FL Fuzzy Logic</p><p>BN Bayesian Networks</p><p>SVR Support Vector Regression</p><p>RT Regression Tree</p><p>ABE Analogy Based Estimation</p></sec></body><back><ref-list><title>References</title><ref id="scirp.93452-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">Pinkashia, S. and Singh, J. (2017) Systematic Literature Review on Software Effort Estimation Using Machine Learning Approaches. 2017 International Conference on Next Generation Computing and Information Systems, Jammu, India, 11-12 December 2017, 43-47. https://doi.org/10.1109/ICNGCIS.2017.33</mixed-citation></ref><ref id="scirp.93452-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">Wen, J., Li, S., Lin, Z., Hu, Y. and Huang. C. (2012) Systematic Literature Review of Machine Learning Based Software Development Effort Estimation Models. Information and Software Technology, 54, 41-59.  
https://doi.org/10.1016/j.infsof.2011.09.002</mixed-citation></ref><ref id="scirp.93452-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">Barry, B., Abts, C. and Chulani, S. (2000) Software Development Cost Estimation Approaches—A Survey. Annals of Software Engineering, 10, 177-205. 
https://doi.org/10.1023/A:1018991717352</mixed-citation></ref><ref id="scirp.93452-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">Moloekken-&amp;#216;Estvold, K., J&amp;#248;rgensen, M., Tanilkan, S.S. Gallis, H., Lien, A.C. and Hove, S.W. (2004) A Survey on Software Estimation in the Norwegian Industry. 10th International Symposium on Software Metrics, Chicago, IL, 11-17 September 2004, 208-219.</mixed-citation></ref><ref id="scirp.93452-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">Li, M.-S., He, M., Yang, D., Shu, F.-D. and Wang, Q. (2007) Software Cost Estimation Method and Application. Journal of Software, 18, 775-795.</mixed-citation></ref><ref id="scirp.93452-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">Magne, J. and Shepperd, M. (2007) A Systematic Review of Software Development Cost Estimation Studies. IEEE Transactions on Software Engineering, 33, 33-53.  
https://doi.org/10.1109/TSE.2007.256943</mixed-citation></ref><ref id="scirp.93452-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">Tomás, V., Ochoa, S.F. and Perovich, D. (2017) Survey of Software Development Effort Estimation Taxonomies. Technical Report. Pending ID. Computer Science Department, University of Chile, Chile.</mixed-citation></ref><ref id="scirp.93452-ref8"><label>8</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Rajeswari</surname><given-names> K. </given-names></name>,<etal>et al</etal>. (<year>2018</year>)<article-title>A Critique on Software Cost Estimation</article-title><source> International Journal of Pure and Applied Mathematics</source><volume> 118</volume>,<fpage> 3851</fpage>-<lpage>3862</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.93452-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">Haitham, H., Kamel, A. and Shams, K. (2013) Software Effort Estimation Using Artificial Neural Networks: A Survey of the Current Practices. 2013 10th Information Technology: New Generations, Las Vegas, NV, 15-17 April 2013, 731-733. 
https://doi.org/10.1109/ITNG.2013.111</mixed-citation></ref><ref id="scirp.93452-ref10"><label>10</label><mixed-citation publication-type="other" xlink:type="simple">Sangwan, O.P. (2017) Software Effort Estimation Using Machine Learning Techniques. 2017 7th International Conference on Cloud Computing, Data Science &amp; Engineering-Confluence, Noida, India, 12-13 January 2017, 92-98.</mixed-citation></ref><ref id="scirp.93452-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">Boehm, B.W. (1981) Software Engineering Economics. Prentice-Hall, Englewood Cliffs, NJ.</mixed-citation></ref><ref id="scirp.93452-ref12"><label>12</label><mixed-citation publication-type="other" xlink:type="simple">Barry, B., Clark, B., Horowitz, E., Westland, C., Madachy, R. and Selby, R. (1995) Cost Models for Future Software Life Cycle Processes: COCOMO 2.0. Annals of Software Engineering, 1, 57-94. https://doi.org/10.1007/BF02249046</mixed-citation></ref><ref id="scirp.93452-ref13"><label>13</label><mixed-citation publication-type="other" xlink:type="simple">IFPUG, FPCPM (2000) International Function Point Users Group (IFPUG) Function Point Counting Practices Manual.</mixed-citation></ref><ref id="scirp.93452-ref14"><label>14</label><mixed-citation publication-type="other" xlink:type="simple">Komal, G., Kaur, P., Kapoor, S. and Narula, S. (2014) Enhancement in COCOMO Model Using Function Point Analysis to Increase Effort Estimation. International Journal of Computer Science and Mobile Computing, 3, 265-572.</mixed-citation></ref><ref id="scirp.93452-ref15"><label>15</label><mixed-citation publication-type="other" xlink:type="simple">Putnam, L.H. (1978) A General Empirical Solution to the Macro Software Sizing and Estimating Problem. IEEE Transactions on Software Engineering, 4, 345-361. 
https://doi.org/10.1109/TSE.1978.231521</mixed-citation></ref><ref id="scirp.93452-ref16"><label>16</label><mixed-citation publication-type="other" xlink:type="simple">Warburton, R.D.H. (1983) Managing and Predicting the Costs of Real-Time Software. IEEE Transactions on Software Engineering, 5, 562-569. 
https://doi.org/10.1109/TSE.1983.235115</mixed-citation></ref><ref id="scirp.93452-ref17"><label>17</label><mixed-citation publication-type="other" xlink:type="simple">Kemerer, C.F. (1987) An Empirical Validation of Software Cost Estimation Models. Communications of the ACM, 30, 416-429. 
https://doi.org/10.1145/22899.22906</mixed-citation></ref><ref id="scirp.93452-ref18"><label>18</label><mixed-citation publication-type="other" xlink:type="simple">Christopher, R. and Roy, R. (2001) Expert Judgment in Cost Estimating: Modelling the Reasoning Process. Concurrent Engineering, 9, 271-284. 
https://doi.org/10.1177/1063293X0100900404</mixed-citation></ref><ref id="scirp.93452-ref19"><label>19</label><mixed-citation publication-type="other" xlink:type="simple">Hareton, L. and Zhang, F. (2002) Software Cost Estimation. In: Handbook of Software Engineering and Knowledge Engineering: Volume 2: Emerging Technologies, World Scientific Publishing Co Pte Ltd., Singapore, 307-324.  
https://doi.org/10.1142/9789812389701_0014</mixed-citation></ref><ref id="scirp.93452-ref20"><label>20</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Sharma</surname><given-names> S. </given-names></name>,<etal>et al</etal>. (<year>2017</year>)<article-title>Applications of Genetic Algorithm in Software Engineering, Distributed Computing and Machine Learning</article-title><source> International Journal of Computer Applications &amp; Information Technology</source><volume> 9</volume>,<fpage> 208</fpage>-<lpage>212</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.93452-ref21"><label>21</label><mixed-citation publication-type="other" xlink:type="simple">Gray, A.R. and Macdonell, S.G. (1999) Software Metrics Data Analysis—Exploring the Relative Performance of Some Commonly Used Modeling Techniques. Empirical Software Engineering, 4, 297-316. https://doi.org/10.1023/A:1009849100780</mixed-citation></ref><ref id="scirp.93452-ref22"><label>22</label><mixed-citation publication-type="other" xlink:type="simple">Ross, J., Ruhe, M and Wieczorek, I. (2000) A Comparative Study of Two Software Development Cost Modeling Techniques Using Multi-Organizational and Company-Specific Data. Information and Software Technology, 42, 1009-1016. 
https://doi.org/10.1016/S0950-5849(00)00153-1</mixed-citation></ref><ref id="scirp.93452-ref23"><label>23</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Abbas</surname><given-names> H. </given-names></name>,<etal>et al</etal>. (<year>2002</year>)<article-title>Comparison of Artificial Neural Network and Regression Models for Estimating Software Development Effort</article-title><source> Information and Software Technology</source><volume> 44</volume>,<fpage> 911</fpage>-<lpage>922</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.93452-ref24"><label>24</label><mixed-citation publication-type="other" xlink:type="simple">Dolado, J.J. (2001) On the Problem of the Software Cost Function. Information and Software Technology, 43, 61-72.  
https://doi.org/10.1016/S0950-5849(00)00137-3</mixed-citation></ref><ref id="scirp.93452-ref25"><label>25</label><mixed-citation publication-type="other" xlink:type="simple">Ingunn, M. and Stensrud, E. (1999) A Controlled Experiment to Assess the Benefits of Estimating with Analogy and Regression Models. IEEE Transactions on Software Engineering, 25, 510-525. https://doi.org/10.1109/32.799947</mixed-citation></ref><ref id="scirp.93452-ref26"><label>26</label><mixed-citation publication-type="other" xlink:type="simple">Ahmed, B.M. (2018) Predicting Software Effort Estimation Using Machine Learning Techniques. 2018 8th International Conference on Computer Science and Information Technology, Amman, 11-12 July 2018, 249-256.  
https://doi.org/10.1109/CSIT.2018.8486222</mixed-citation></ref><ref id="scirp.93452-ref27"><label>27</label><mixed-citation publication-type="other" xlink:type="simple">Poonam, R. and Jain, S. (2016) Enhanced Software Effort Estimation Using Multi Layered Feed Forward Artificial Neural Network Technique. Procedia Computer Science, 89, 307-312. https://doi.org/10.1016/j.procs.2016.06.073</mixed-citation></ref><ref id="scirp.93452-ref28"><label>28</label><mixed-citation publication-type="other" xlink:type="simple">Idri, A., Khoshgoftaar, T.M. and Abran, A. (2002) Can Neural Networks Be Easily Interpreted in Software Cost Estimation? 2002 IEEE World Congress on Computational Intelligence. 2002 IEEE International Conference on Fuzzy Systems, Honolulu, HI, 12-17 May 2002, 1162-1167.</mixed-citation></ref><ref id="scirp.93452-ref29"><label>29</label><mixed-citation publication-type="other" xlink:type="simple">Singh, B.K. and Misra, A.K. (2012) Software Effort Estimation by Genetic Algorithm Tuned Parameters of Modified Constructive Cost Model for NASA Software Projects. International Journal of Computer Applications, 59, 22-26.</mixed-citation></ref><ref id="scirp.93452-ref30"><label>30</label><mixed-citation publication-type="other" xlink:type="simple">Burgess, C.J. and Lefley, M. (2001) Can Genetic Programming Improve Software Effort Estimation? A Comparative Evaluation. Information and Software Technology, 43, 863-873. https://doi.org/10.1016/S0950-5849(01)00192-6</mixed-citation></ref><ref id="scirp.93452-ref31"><label>31</label><mixed-citation publication-type="other" xlink:type="simple">Anupama, K., Soni, A.K. and Soni, R. (2013) Radial Basis Function Network Using Intuitionistic Fuzzy C Means for Software Cost Estimation. International Journal of Computer Applications in Technology, 47, 86-95. 
https://doi.org/10.1504/IJCAT.2013.054305</mixed-citation></ref><ref id="scirp.93452-ref32"><label>32</label><mixed-citation publication-type="other" xlink:type="simple">Anish, M., Parkash, K. and Mittal, H. (2010) Software Cost Estimation Using Fuzzy Logic. ACM SIGSOFT Software Engineering Notes, 35, 1-7. 
https://doi.org/10.1145/1668862.1668866</mixed-citation></ref><ref id="scirp.93452-ref33"><label>33</label><mixed-citation publication-type="other" xlink:type="simple">Hrvoje, K. and Gotovac, S. (2015) Estimating Software Development Effort Using Bayesian Networks. 2015 23rd International Conference on Software, Telecommunications and Computer Networks, Split, Croatia, 16-18 September 2015, 229-233.  
https://doi.org/10.1109/SOFTCOM.2015.7314091</mixed-citation></ref><ref id="scirp.93452-ref34"><label>34</label><mixed-citation publication-type="other" xlink:type="simple">Bhavendra Kumar, S., Sinhal, A. and Verma, B. (2013) A Software Measurement Using Artificial Neural Network and Support Vector Machine. International Journal of Software Engineering &amp; Applications, 4, 41-52. 
https://doi.org/10.5121/ijsea.2013.4404</mixed-citation></ref><ref id="scirp.93452-ref35"><label>35</label><mixed-citation publication-type="other" xlink:type="simple">Magne, J. (2004) Regression Models of Software Development Effort Estimation Accuracy and Bias. Empirical Software Engineering, 9, 297-314. 
https://doi.org/10.1023/B:EMSE.0000039881.57613.cb</mixed-citation></ref><ref id="scirp.93452-ref36"><label>36</label><mixed-citation publication-type="other" xlink:type="simple">Ali, I., Amazal, F.A. and Abran, A. (2016) Accuracy Comparison of Analogy-Based Software Development Effort Estimation Techniques. International Journal of Intelligent Systems, 31, 128-152. https://doi.org/10.1002/int.21748</mixed-citation></ref><ref id="scirp.93452-ref37"><label>37</label><mixed-citation publication-type="other" xlink:type="simple">Hathaichanok, S. and Prompoon, N. (2012) Framework for Developing a Software Cost Estimation Model for Software Modification Based on a Relational Matrix of Project Profile and Software Cost Using an Analogy Estimation Method. International Journal of Computer and Communication Engineering, 1, 129-134. 
https://doi.org/10.7763/IJCCE.2012.V1.36</mixed-citation></ref><ref id="scirp.93452-ref38"><label>38</label><mixed-citation publication-type="other" xlink:type="simple">Manikavelan, D. and Ponnusamy, R. (2015) Improvised Analogy Based Software Cost Estimation with Ant Colony Optimization. Research Journal of Applied Sciences, Engineering and Technology, 10, 293-297. 
https://doi.org/10.19026/rjaset.10.2490</mixed-citation></ref><ref id="scirp.93452-ref39"><label>39</label><mixed-citation publication-type="other" xlink:type="simple">Papatheocharous, E. and Andreou, A.S. (2012) Software Cost Modelling and Estimation Using Artificial Neural Networks Enhanced by Input Sensitivity Analysis. Journal of Universal Computer Science, 18, 2041-2070.</mixed-citation></ref></ref-list></back></article>