<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">JSEA</journal-id><journal-title-group><journal-title>Journal of Software Engineering and Applications</journal-title></journal-title-group><issn pub-type="epub">1945-3116</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/jsea.2017.104020</article-id><article-id pub-id-type="publisher-id">JSEA-75710</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Computer Science&amp;Communications</subject></subj-group></article-categories><title-group><article-title>
 
 
  Estimation Models for Software Functional Test Effort
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Kamala</surname><given-names>Ramasubramani Jayakumar</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Alain</surname><given-names>Abran</given-names></name><xref ref-type="aff" rid="aff2"><sup>2</sup></xref></contrib></contrib-group><aff id="aff1"><addr-line>Amitysoft Technologies, Chennai, India</addr-line></aff><aff id="aff2"><addr-line>&amp;amp;Eacute;cole de technologie supérieure, University of Quebec, Montreal, Canada</addr-line></aff><pub-date pub-type="epub"><day>06</day><month>04</month><year>2017</year></pub-date><volume>10</volume><issue>04</issue><fpage>338</fpage><lpage>353</lpage><history><date date-type="received"><day>January</day>	<month>25,</month>	<year>2017</year></date><date date-type="rev-recd"><day>Accepted:</day>	<month>April</month>	<year>24,</year>	</date><date date-type="accepted"><day>April</day>	<month>27,</month>	<year>2017</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
  The International Software Benchmarking and Standards Group (ISBSG) data-base was used to build estimation models for estimating software functional test effort. The analysis of the data revealed three test productivity patterns representing economies or diseconomies of scale and these patterns served as a basis for investigating the characteristics of the corresponding projects. Three groups of projects related to the three different productivity patterns, characterized by domain, team size, elapsed time and rigor of verification and validation carried out during development, were found to be statistically significant. Within each project group, the variations in test effort can be explained, in addition to functional size, by 1) the processes executed during development, and 2) the processes adopted for testing. Portfolios of estimation models were built using combinations of the three independent variables. Performance of the estimation models built using the function point method innovated by the Common Software Measurement International Consortium (COSMIC) known as COSMIC Function Points, and the one advocated by the International Function Point Users Group (IFPUG) known as IFPUG Function Points, were compared to evaluate the impact of these respective sizing methods on test effort estimation.
 
</p></abstract><kwd-group><kwd>COSMIC Function Points</kwd><kwd> Estimation</kwd><kwd> Functional Sizing</kwd><kwd> Performance Measurement</kwd><kwd> Software Testing</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>This paper reports on a set of estimation models designed with data chosen from the ISBSG repository consisting of functional sizes reported both in IFPUG function points [<xref ref-type="bibr" rid="scirp.75710-ref1">1</xref>] and COSMIC function points. These estimation models were evaluated using criteria for measuring outputs from estimation models. The models were compared to understand their performance based on the measure of their predictability.</p><p>The motivation for this research work arises from the fact that existing techniques for estimating test effort (such as judgment-based, work breakdown, factors &amp; weights, and functional size based techniques) suffer from several limitations [<xref ref-type="bibr" rid="scirp.75710-ref2">2</xref>] [<xref ref-type="bibr" rid="scirp.75710-ref3">3</xref>] , while other innovative approaches for estimating testing effort (such as fuzzy inference, artificial neural networks, and case-based reasoning as proposed in the literature) are yet to be adopted in the industry. There is a growing body of work on the use of the COSMIC function points [<xref ref-type="bibr" rid="scirp.75710-ref4">4</xref>] [<xref ref-type="bibr" rid="scirp.75710-ref5">5</xref>] for estimation and performance measurement of software development projects which can be adapted for estimating software test effort too.</p><p>The remainder of this paper is structured as follows. Section 2 presents the data preparation; Section 3 is data analysis; Section 4 is the estimation models and Section 5 is the conclusions.</p></sec><sec id="s2"><title>2. Data Preparation</title><sec id="s2_1"><title>2.1. ISBSG Data</title><p>Release 12 of ISBSG data published in 2013 [<xref ref-type="bibr" rid="scirp.75710-ref6">6</xref>] consists of data related to parameters of software projects re-ported over the last two and half decades, providing industry and researchers with standardized data for benchmarking and estimation. The ISBSG dataset has been extensively reviewed for its applicability to building effort estimation models, including effects of outliers and missing values [<xref ref-type="bibr" rid="scirp.75710-ref7">7</xref>] [<xref ref-type="bibr" rid="scirp.75710-ref8">8</xref>] .</p><p>The attributes of interest for test effort estimation models are:</p><p>a. Functional size data based on international measurement standards such as IFPUG and COSMIC function points.</p><p>b. Schedule, team size, work effort information, project elapsed time and breakdown of work effort by project phase (planning, specifications, design, build, test and install).</p><p>c. Project process related data based on software life cycle activities (e.g. planning, specifications, design, build, test) and adoption of practices from standards or models such as ISO 9001, CMMI, SPICE, PSP etc. used in developing the software.</p><p>d. Grouping attributes: industry sector, application group (e.g., business, real time etc.), and development type (new development, enhancement or re-deve- lopment).</p><p>e. Development platform: PC, mid-range, main frame or multi-platform.</p><p>f. Architecture: whether the application is standalone, multi-tier, client/server or Web-based.</p><p>g. Language type: 3GL, 4GL, or application generators used in development.</p><p>h. Overall data quality rating assigned by the ISBSG: A, B, C or D indicating very good to unreliable.</p><p>i. Function points data quality rating assigned by the ISBSG: A, B, C or D ranging from very good to unreliable.</p></sec><sec id="s2_2"><title>2.2. Data Preprocessing</title><p>A set of criteria was defined to ensure data quality, relevance to current industry needs, suitability to the testing context and adequacy for statistical analysis, as follows:</p><p>1) Data Quality</p><p>a. ISBSG quality rating:</p><p>Data quality ratings of A and B were selected to reduce risk and improve confidence in the results.</p><p>b. Function point size quality:</p><p>When IFPUG function points were used for the measurement of size, only the un-adjusted function point value was considered. Function point data quality ratings of C and D were excluded from the data.</p><p>2) Data Relevance</p><p>ISBSG data consist of projects reported since the early 90s. Data prior to 2000 and projects with an architecture type of “standalone” were removed while client/server or Web-based projects were considered for modelling.</p><p>3) Data Suitability</p><p>To exclude trivial projects, the following filters were applied:</p><p>a. Total normalized work effort (full life cycle effort for project) equal to or greater than 80 hours.</p><p>b. Efforts reported for testing greater than or equal to 16 hours.</p><p>c. Types of testing other than functional testing were excluded.</p><p>4) Data Adequacy</p><p>a. Application group chosen: business.</p><p>b. Development type chosen: new development and re-development.</p></sec><sec id="s2_3"><title>2.3. Generation of Datasets</title><p>Applying the filters related to the criteria for data selection and removal of outliers resulted in 142 data points, which were then grouped to form four datasets:</p><p>Dataset A: This dataset consists of all 142 data points including project functional size measures reported in IFPUG 4.1 or COSMIC FP. For this study, they were not differentiated within dataset A as they correlate well even though the relationship is not the same across all size ranges [<xref ref-type="bibr" rid="scirp.75710-ref9">9</xref>] .</p><p>Dataset B: In the case of dataset A, projects with an architecture field value of “standalone” were eliminated from the original ISBSG data set, while “blanks” were retained. To be very specific about the architecture type, “blanks” were also eliminated from dataset A to arrive at dataset B, with 72 data points.</p><p>Dataset C: Data set C is made up of projects where functional size was reported in COSMIC function points. It is a subset of data set A and has 82 data points.</p><p>Dataset D: Dataset D includes only projects where functional size was reported in IFPUG function points. It is another subset of dataset A and contains 60 data points.</p></sec></sec><sec id="s3"><title>3. Data Analysis</title><sec id="s3_1"><title>3.1. Strategy</title><p>The following strategy was adopted for data analysis:</p><p>a) Identify data point subsets exhibiting different levels of testing productivity.</p><p>b) Analyze these subsets to identify the possible causes for the differences in productivity.</p></sec><sec id="s3_2"><title>3.2. Identification of Test Productivity Levels</title><p>The scatter diagram in <xref ref-type="fig" rid="fig1">Figure 1</xref> depicts a large dispersion between functional size and test effort, the independent and dependent variables, respectively. The pattern is closer to wedge-shaped and is typical of data from large repositories [<xref ref-type="bibr" rid="scirp.75710-ref10">10</xref>] .</p><p>Within the dataset of <xref ref-type="fig" rid="fig1">Figure 1</xref>, there are candidate groups exhibiting both large economies of scale and large diseconomies of scale. The rate of increase of test effort is not the same for all similar functional sizes. Analyzing various slices of data brought out different testing productivity levels (<xref ref-type="fig" rid="fig2">Figure 2</xref>).</p><p>As economies and diseconomies of scale correspond to different productivity levels, a new term “test delivery rate” (TDR) was defined to describe project testing productivity. TDR is the rate at which software functionality is tested as a factor of the effort required, and is expressed as hours per functional size unit</p><fig id="fig1"  position="float"><label><xref ref-type="fig" rid="fig1">Figure 1</xref></label><caption><title> Scatter diagram: size versus test effort</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/3-9302379x2.png"/></fig><fig id="fig2"  position="float"><label><xref ref-type="fig" rid="fig2">Figure 2</xref></label><caption><title> Multiple data groups representing different economies of scale</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/3-9302379x3.png"/></fig><p>(hr./FSU). Functional size unit (FSU) refers to either IFPUG or COSMIC function points, depending upon the sizing method used for measurement. The four varying levels of productivity are referred as “TDR levels”. TDR being the effect, the characteristics of the projects falling into each level were then investigated to identify the underlying causes. Due to the highly dispersed nature of TDR level 4, only TDR levels 1 to 3 were taken up for further analysis and development of the estimation models.</p></sec><sec id="s3_3"><title>3.3. Identification of Candidate Characteristics of Projects</title><p>Previous research work [<xref ref-type="bibr" rid="scirp.75710-ref11">11</xref>] [<xref ref-type="bibr" rid="scirp.75710-ref12">12</xref>] based on data from hundreds of software projects has indicated that team size and schedule (duration of the project) within a particular domain affect the productivity of development projects. As software testing is one of the phases of development, project attributes such as domain, team size and elapsed time are likely causes for test productivity, too. Testability of the software components, i.e., quality of the software delivered for testing, is critical for reducing the testing cost [<xref ref-type="bibr" rid="scirp.75710-ref13">13</xref>] and hence the effort for testing. The quality of the software delivered for testing can be determined by the extent of verification and validation activities carried out during the development process.</p><p>From these previous studies, therefore, in choosing candidates of interest for our investigations we selected the following project characteristics:</p><p>Team size: We classified team size into three categories typically present in the industry: (a) 1 - 4 persons, (b) 5 - 8 persons and (c) more than 8 persons resenting small, medium and large team sizes, respectively.</p><p>Elapsed Time: Elapsed time in calendar months was derived from the “project elapsed time”. Based on this attribute, projects were classified into three groups: (a) 1 - 3 months, (b) 4 - 6 months and (c) greater than 6 months, referred to as small, medium, and large, respectively.</p><p>V &amp; V rigor: This attribute was derived from the data fields related to the “Documents &amp; Techniques” category in ISBSG, which indicates the degree of rigor applied during verification and validation. Two ratings are proposed for V&amp;V rigor (<xref ref-type="table" rid="table1">Table 1</xref>).</p><p>Application domain: This attribute was derived from the ISBSG data field “Industry Sector”. Considering the number of data points available for different industry sectors, the application domain was classified into three categories, namely: (a) banking, financial services and insurance (BFSI) (b) education and (c) government (govt).</p><table-wrap id="table1" ><label><xref ref-type="table" rid="table1">Table 1</xref></label><caption><title> V&amp;V rigor rating scheme</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >V&amp;V Rigor Rating</th><th align="center" valign="middle" >Description</th></tr></thead><tr><td align="center" valign="middle" >Low</td><td align="center" valign="middle" >Little or no evidence of reviews/inspection</td></tr><tr><td align="center" valign="middle" >High</td><td align="center" valign="middle" >Reviews/inspection reported for at least one of the specification, design and build phases</td></tr></tbody></table></table-wrap><p>The group of projects contributing to each TDR level was termed the project group. Accordingly, project group 1 (PG1), project group 2 (PG2) and project group 3 (PG 3) refer to TDR levels 1, 2 and 3 discussed in Section 3.2. Based on the percentage of projects falling into each of the project groups for the four attributes of interest (<xref ref-type="table" rid="table2">Table 2</xref>), we were able to characterize the project groups.</p><p>Close to half of the BFSI projects (46%) fell into PG3 followed by a third in PG2. All education projects fell into PG1 while slightly more than half of the government projects fell in PG2. Close to two thirds of projects with a small team size fell into PG1, while 82% (46% + 36%) those with a medium team size were distributed between PG1 and PG2. Slightly less than 50% of the large team size projects fell into PG3. Close to two thirds of projects with a small elapsed time went into PG1, while 71% (38% + 33%) of those with a medium elapsed time were spread between PG2 and PG3. Similarly, PG2 and PG3 shared 72% (40% + 32%) of the projects with a large elapsed time. Projects with higher V&amp;V rigor had a two thirds presence in PG1 while 77% (38% + 39%) of the lower V&amp;V rigor projects were divided more or less evenly between PG2 and PG3.</p><p>The above analysis reveals that the three project groups have certain distinctions with respect to the domain, team size, elapsed time and V&amp;V rigor besides test productivity. To establish statistical significance, a test of hypothesis was performed and the p value computed. The significance test conducted on the three project groups with respect to these attributes resulted in a p value of less than 0.001 for V&amp;V rigor and domain and less than 0.1 for team size and elapsed time. This further establishes that the variations across the three project groups are reasonably significant and the attributes identified are potential contributors to test productivity.</p><table-wrap id="table2" ><label><xref ref-type="table" rid="table2">Table 2</xref></label><caption><title> Analysis of project characteristics―Dataset A</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Domain</th><th align="center" valign="middle" >No. %</th><th align="center" valign="middle" >PG1</th><th align="center" valign="middle" >PG2</th><th align="center" valign="middle" >PG3</th><th align="center" valign="middle" >Team Size</th><th align="center" valign="middle" >No. %</th><th align="center" valign="middle" >PG1</th><th align="center" valign="middle" >PG2</th><th align="center" valign="middle" >PG3</th></tr></thead><tr><td align="center" valign="middle"  rowspan="2"  >BFSI</td><td align="center" valign="middle" >No.</td><td align="center" valign="middle" >14</td><td align="center" valign="middle" >23</td><td align="center" valign="middle" >32</td><td align="center" valign="middle"  rowspan="2"  >Small</td><td align="center" valign="middle" >No.</td><td align="center" valign="middle" >10</td><td align="center" valign="middle" >4</td><td align="center" valign="middle" >2</td></tr><tr><td align="center" valign="middle" >%</td><td align="center" valign="middle" >20</td><td align="center" valign="middle" >33</td><td align="center" valign="middle" >46</td><td align="center" valign="middle" >%</td><td align="center" valign="middle" >63</td><td align="center" valign="middle" >25</td><td align="center" valign="middle" >13</td></tr><tr><td align="center" valign="middle"  rowspan="2"  >Education</td><td align="center" valign="middle" >No.</td><td align="center" valign="middle" >11</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle"  rowspan="2"  >Medium</td><td align="center" valign="middle" >No.</td><td align="center" valign="middle" >18</td><td align="center" valign="middle" >14</td><td align="center" valign="middle" >7</td></tr><tr><td align="center" valign="middle" >%</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >%</td><td align="center" valign="middle" >46</td><td align="center" valign="middle" >36</td><td align="center" valign="middle" >18</td></tr><tr><td align="center" valign="middle"  rowspan="2"  >Govt.</td><td align="center" valign="middle" >No.</td><td align="center" valign="middle" >6</td><td align="center" valign="middle" >10</td><td align="center" valign="middle" >2</td><td align="center" valign="middle"  rowspan="2"  >Large</td><td align="center" valign="middle" >No.</td><td align="center" valign="middle" >5</td><td align="center" valign="middle" >4</td><td align="center" valign="middle" >8</td></tr><tr><td align="center" valign="middle" >%</td><td align="center" valign="middle" >33</td><td align="center" valign="middle" >56</td><td align="center" valign="middle" >11</td><td align="center" valign="middle" >%</td><td align="center" valign="middle" >29</td><td align="center" valign="middle" >24</td><td align="center" valign="middle" >47</td></tr><tr><td align="center" valign="middle" >Elapsed Time</td><td align="center" valign="middle" >No. %</td><td align="center" valign="middle" >PG1</td><td align="center" valign="middle" >PG2</td><td align="center" valign="middle" >PG3</td><td align="center" valign="middle" >V &amp; V Rigor</td><td align="center" valign="middle" >No. %</td><td align="center" valign="middle" >PG1</td><td align="center" valign="middle" >PG2</td><td align="center" valign="middle" >PG3</td></tr><tr><td align="center" valign="middle"  rowspan="2"  >Small</td><td align="center" valign="middle" >No.</td><td align="center" valign="middle" >18</td><td align="center" valign="middle" >7</td><td align="center" valign="middle" >5</td><td align="center" valign="middle"  rowspan="2"  >Low</td><td align="center" valign="middle" >No.</td><td align="center" valign="middle" >25</td><td align="center" valign="middle" >42</td><td align="center" valign="middle" >43</td></tr><tr><td align="center" valign="middle" >%</td><td align="center" valign="middle" >60</td><td align="center" valign="middle" >23</td><td align="center" valign="middle" >17</td><td align="center" valign="middle" >%</td><td align="center" valign="middle" >23</td><td align="center" valign="middle" >38</td><td align="center" valign="middle" >39</td></tr><tr><td align="center" valign="middle"  rowspan="2"  >Medium</td><td align="center" valign="middle" >No.</td><td align="center" valign="middle" >6</td><td align="center" valign="middle" >8</td><td align="center" valign="middle" >7</td><td align="center" valign="middle"  rowspan="2"  >High</td><td align="center" valign="middle" >No.</td><td align="center" valign="middle" >21</td><td align="center" valign="middle" >7</td><td align="center" valign="middle" >4</td></tr><tr><td align="center" valign="middle" >%</td><td align="center" valign="middle" >29</td><td align="center" valign="middle" >38</td><td align="center" valign="middle" >33</td><td align="center" valign="middle" >%</td><td align="center" valign="middle" >66</td><td align="center" valign="middle" >22</td><td align="center" valign="middle" >13</td></tr><tr><td align="center" valign="middle"  rowspan="2"  >Large</td><td align="center" valign="middle" >No.</td><td align="center" valign="middle" >14</td><td align="center" valign="middle" >20</td><td align="center" valign="middle" >16</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >%</td><td align="center" valign="middle" >28</td><td align="center" valign="middle" >40</td><td align="center" valign="middle" >32</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr></tbody></table></table-wrap><p>The results of the analysis of the project characteristics in the three datasets A to C, excluding dataset D, demonstrate similar behavior. Characteristics of project groups PG1 to PG3 based on these attributes are summarized in <xref ref-type="table" rid="table3">Table 3</xref>.</p></sec><sec id="s3_4"><title>3.4. Identification of the Independent Variables</title><p>1) Size</p><p>It has been observed that functional size is the most accepted approach for measuring size, as sensitivity to changes in functional size has a greater impact on project effort [<xref ref-type="bibr" rid="scirp.75710-ref14">14</xref>] [<xref ref-type="bibr" rid="scirp.75710-ref15">15</xref>] . Here, correlation coefficients computed using dataset A, between size and test effort values of 0.9035 for PG1, 0.8572 for PG2 and 0.8572 for PG3, indicate good correlation of functional size with effort. Size was therefore chosen as the primary independent variable.</p><p>2) Non-Size Variables</p><p>Size being the main independent variable, other independent variables were next examined for significance of incorporating them into estimation models. It has been observed [<xref ref-type="bibr" rid="scirp.75710-ref13">13</xref>] that “testability of software components”, meaning the quality of the software delivered for testing and testing processes followed while testing, are critical factors for reducing testing effort and improving software quality. To accommodate these process factors two new variables representing development process quality and testing process quality were defined and investigated as follows.</p><p>a) Development Process Quality Rating (DevQ)</p><p>The process followed during development was rated by considering the nature of the development life cycle followed and the artefacts produced, based on the following project attributes:</p><p> Standards followed.</p><p> Distinct development life cycle phases followed.</p><p> Verification activities carried out during development.</p><p>The ISBSG data field “software process” has one of the values―CMMI, ISO, SPICE, PSP or any such standard followed during development. A set of fields representing “Documents and Techniques” exists in the ISBSG data providing information on the life cycle phases adopted and verification activities carried out during development. Based on these, a rating for DevQ was developed, as shown in <xref ref-type="table" rid="table4">Table 4</xref>.</p><p>b) Test Process Quality Rating (TestQ)</p><p>While reviewing, the data related to the testing process followed, it was found</p><table-wrap id="table3" ><label><xref ref-type="table" rid="table3">Table 3</xref></label><caption><title> Characteristics of project groups</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Attribute</th><th align="center" valign="middle" >PG1</th><th align="center" valign="middle" >PG2</th><th align="center" valign="middle" >PG3</th></tr></thead><tr><td align="center" valign="middle" >Domain</td><td align="center" valign="middle" >Educational</td><td align="center" valign="middle" >Government</td><td align="center" valign="middle" >BFSI</td></tr><tr><td align="center" valign="middle" >Team Size</td><td align="center" valign="middle" >Small/Medium</td><td align="center" valign="middle" >Small/ Medium</td><td align="center" valign="middle" >Large</td></tr><tr><td align="center" valign="middle" >Elapsed Time</td><td align="center" valign="middle" >Small</td><td align="center" valign="middle" >Medium/Large</td><td align="center" valign="middle" >Medium/Large</td></tr><tr><td align="center" valign="middle" >V&amp;V Rigour</td><td align="center" valign="middle" >High</td><td align="center" valign="middle" >Low</td><td align="center" valign="middle" >Low</td></tr></tbody></table></table-wrap><table-wrap id="table4" ><label><xref ref-type="table" rid="table4">Table 4</xref></label><caption><title> Rating for development process (DevQ)</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Software Process</th><th align="center" valign="middle" >Documents &amp; Techniques</th><th align="center" valign="middle" >DevQ Rating</th></tr></thead><tr><td align="center" valign="middle" >Not reported</td><td align="center" valign="middle" >Very little reporting to infer</td><td align="center" valign="middle" >0</td></tr><tr><td align="center" valign="middle" >Reported</td><td align="center" valign="middle" >Very little reporting to infer</td><td align="center" valign="middle" >1</td></tr><tr><td align="center" valign="middle" >Not reported</td><td align="center" valign="middle" >One or more phases has values</td><td align="center" valign="middle" >1</td></tr><tr><td align="center" valign="middle" >Reported</td><td align="center" valign="middle" >One or more phases has values</td><td align="center" valign="middle" >2</td></tr></tbody></table></table-wrap><p>that there were not enough fields in the ISBSG data to capture the details of the testing process, such as testing techniques adopted, levels of testing executed, test artefacts produced, reviews of test cases etc., to gauge the extent of testing. This notwithstanding, it was possible to classify the test process rating broadly into two categories (<xref ref-type="table" rid="table5">Table 5</xref>).</p></sec><sec id="s3_5"><title>3.5. Analysis of DevQ and TestQ</title><p>Projects in data set A were analyzed in terms of DevQ and TestQ:</p><p> 36%, 49%, and 15% of the projects were found to be in DevQ with ratings 0, 1 and 2, respectively.</p><p> 80% and 20% of the projects had TestQ ratings 0 and 1 respectively.</p><p>To further justify the inclusion of these variables, two statistical tests were carried out to quantify their significance (<xref ref-type="table" rid="table6">Table 6</xref>):</p><p> the Kruskal-Wallis Test for DevQ as it involved three categories, and</p><p> the Mann Whitney Test was applied for TestQ.</p><p>The p value indicated that size, DevQ and TestQ were statistically significant.</p></sec></sec><sec id="s4"><title>4. Estimation Models</title><sec id="s4_1"><title>4.1. Portfolio of Models</title><p>The linear form of relationship between input and output variables was chosen to build models for effort estimation. Linear regression analysis, a well-known and well understood algorithm in statistics and machine learning, does not require much training data, and is easily interpreted by project managers. Parametric models are objective, repeatable, fast and easy to use, and can be used early in the life cycle if they are properly calibrated and validated [<xref ref-type="bibr" rid="scirp.75710-ref16">16</xref>] . A set of 24 models under four portfolios were generated (<xref ref-type="table" rid="table7">Table 7</xref>) using datasets A to D.</p><p>Portfolio A models based on dataset A:</p><p>Models 1, 2 and 3 are for each project group using size as the independent variable.</p><p>Models 4, 5 and 6 use both size and DevQ as independent variables and relate to project groups 1, 2 and 3 respectively.</p><p>Models 7, 8 and 9 use size, DevQ and TestQ as independent variables and represent project groups 1, 2 and 3 respectively.</p><p>Portfolio B models based on dataset B:</p><p>Models 10, 11 and 12 relate to project groups 1, 2 and 3 respectively using size as independent variable.</p><table-wrap id="table5" ><label><xref ref-type="table" rid="table5">Table 5</xref></label><caption><title> Rating for test process (TestQ)</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Test Process Criteria</th><th align="center" valign="middle" >Test Process Rating (TestQ)</th></tr></thead><tr><td align="center" valign="middle" >No evidence of Test Artefacts</td><td align="center" valign="middle" >0</td></tr><tr><td align="center" valign="middle" >Evidence of Test Artefacts</td><td align="center" valign="middle" >1</td></tr></tbody></table></table-wrap><table-wrap id="table6" ><label><xref ref-type="table" rid="table6">Table 6</xref></label><caption><title> Test of significance for independent variables</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Statistical Test</th><th align="center" valign="middle" >Variable</th><th align="center" valign="middle" >p Value</th></tr></thead><tr><td align="center" valign="middle" >Chi Square</td><td align="center" valign="middle" >Size</td><td align="center" valign="middle" >&lt; 0.001</td></tr><tr><td align="center" valign="middle" >Kruskal-Wallis Test</td><td align="center" valign="middle" >DevQ</td><td align="center" valign="middle" >0.005</td></tr><tr><td align="center" valign="middle" >Mann-Whitney Test</td><td align="center" valign="middle" >TestQ</td><td align="center" valign="middle" >0.003</td></tr></tbody></table></table-wrap><table-wrap id="table7" ><label><xref ref-type="table" rid="table7">Table 7</xref></label><caption><title> Portfolio of estimation models</title></caption><table><tbody><thead><tr><th align="center" valign="middle"  rowspan="3"  >Portfolio</th><th align="center" valign="middle"  rowspan="3"  >ID</th><th align="center" valign="middle"  rowspan="3"  >PG</th><th align="center" valign="middle"  colspan="8"  >Model Coefficients</th></tr></thead><tr><td align="center" valign="middle"  rowspan="2"  >A</td><td align="center" valign="middle"  rowspan="2"  >B</td><td align="center" valign="middle"  colspan="2"  >D1</td><td align="center" valign="middle"  colspan="2"  >D2</td><td align="center" valign="middle" >T1</td><td align="center" valign="middle" >T2</td></tr><tr><td align="center" valign="middle" >DevQ = 0</td><td align="center" valign="middle" >DevQ = 1</td><td align="center" valign="middle" >DevQ = 0</td><td align="center" valign="middle" >DevQ = 1</td><td align="center" valign="middle" >TestQ = 0</td><td align="center" valign="middle" >TestQ = 0</td></tr><tr><td align="center" valign="middle"  rowspan="8"  >A</td><td align="center" valign="middle" >1</td><td align="center" valign="middle" >1</td><td align="center" valign="middle" >1.617</td><td align="center" valign="middle" >0.604</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >2</td><td align="center" valign="middle" >2</td><td align="center" valign="middle" >20.69</td><td align="center" valign="middle" >1.705</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >3</td><td align="center" valign="middle" >3</td><td align="center" valign="middle" >98.13</td><td align="center" valign="middle" >4.801</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >4</td><td align="center" valign="middle" >1</td><td align="center" valign="middle" >16.12</td><td align="center" valign="middle" >0.485</td><td align="center" valign="middle" >19.347</td><td align="center" valign="middle" >−39.375</td><td align="center" valign="middle" >−0.23</td><td align="center" valign="middle" >0.214</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >5</td><td align="center" valign="middle" >2</td><td align="center" valign="middle" >20.57</td><td align="center" valign="middle" >1.56</td><td align="center" valign="middle" >−94.1</td><td align="center" valign="middle" >34.077</td><td align="center" valign="middle" >0.562</td><td align="center" valign="middle" >−0.009</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >6</td><td align="center" valign="middle" >3</td><td align="center" valign="middle" >38.85</td><td align="center" valign="middle" >3.734</td><td align="center" valign="middle" >−55.913</td><td align="center" valign="middle" >92.609</td><td align="center" valign="middle" >2.14</td><td align="center" valign="middle" >0.852</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >7</td><td align="center" valign="middle" >1</td><td align="center" valign="middle" >−9.62</td><td align="center" valign="middle" >0.65</td><td align="center" valign="middle" >6.967</td><td align="center" valign="middle" >−41.78</td><td align="center" valign="middle" >0.003</td><td align="center" valign="middle" >0.193</td><td align="center" valign="middle" >38.124</td><td align="center" valign="middle" >−0.191</td></tr><tr><td align="center" valign="middle" >8</td><td align="center" valign="middle" >2</td><td align="center" valign="middle" >30.74</td><td align="center" valign="middle" >1.541</td><td align="center" valign="middle" >−19.755</td><td align="center" valign="middle" >62.481</td><td align="center" valign="middle" >−0.039</td><td align="center" valign="middle" >−0.338</td><td align="center" valign="middle" >−84.511</td><td align="center" valign="middle" >0.62</td></tr><tr><td align="center" valign="middle" ></td><td align="center" valign="middle" >9</td><td align="center" valign="middle" >3</td><td align="center" valign="middle" >38.85</td><td align="center" valign="middle" >3.734</td><td align="center" valign="middle" >−55.913</td><td align="center" valign="middle" >92.609</td><td align="center" valign="middle" >2.14</td><td align="center" valign="middle" >0.852</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td></tr><tr><td align="center" valign="middle"  rowspan="9"  >B</td><td align="center" valign="middle" >10</td><td align="center" valign="middle" >1</td><td align="center" valign="middle" >−8.3448</td><td align="center" valign="middle" >0.61</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >11</td><td align="center" valign="middle" >2</td><td align="center" valign="middle" >−30.569</td><td align="center" valign="middle" >1.929</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >12</td><td align="center" valign="middle" >3</td><td align="center" valign="middle" >−157.62</td><td align="center" valign="middle" >6.126</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >13</td><td align="center" valign="middle" >1</td><td align="center" valign="middle" >16.124</td><td align="center" valign="middle" >0.485</td><td align="center" valign="middle" >46.572</td><td align="center" valign="middle" >−52.672</td><td align="center" valign="middle" >−0.201</td><td align="center" valign="middle" >0.222</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >14</td><td align="center" valign="middle" >2</td><td align="center" valign="middle" >20.57</td><td align="center" valign="middle" >1.56</td><td align="center" valign="middle" >−180.84</td><td align="center" valign="middle" >−60.58</td><td align="center" valign="middle" >0.973</td><td align="center" valign="middle" >0.313</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >15</td><td align="center" valign="middle" >3</td><td align="center" valign="middle" >38.847</td><td align="center" valign="middle" >3.734</td><td align="center" valign="middle" >−375.38</td><td align="center" valign="middle" >5.027</td><td align="center" valign="middle" >3.881</td><td align="center" valign="middle" >1.449</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >16</td><td align="center" valign="middle" >1</td><td align="center" valign="middle" >−12.583</td><td align="center" valign="middle" >0.68</td><td align="center" valign="middle" >58.608</td><td align="center" valign="middle" >−43.272</td><td align="center" valign="middle" >−0.208</td><td align="center" valign="middle" >0.171</td><td align="center" valign="middle" >16.67</td><td align="center" valign="middle" >−0.188</td></tr><tr><td align="center" valign="middle" >17</td><td align="center" valign="middle" >2</td><td align="center" valign="middle" >56.462</td><td align="center" valign="middle" >1.492</td><td align="center" valign="middle" >2.443</td><td align="center" valign="middle" >129.634</td><td align="center" valign="middle" >−0.354</td><td align="center" valign="middle" >−1.025</td><td align="center" valign="middle" >−219.18</td><td align="center" valign="middle" >1.395</td></tr><tr><td align="center" valign="middle" >18</td><td align="center" valign="middle" >3</td><td align="center" valign="middle" >38.847</td><td align="center" valign="middle" >3.734</td><td align="center" valign="middle" >−375.38</td><td align="center" valign="middle" >5.027</td><td align="center" valign="middle" >3.881</td><td align="center" valign="middle" >1.449</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td></tr><tr><td align="center" valign="middle"  rowspan="3"  >C</td><td align="center" valign="middle" >19</td><td align="center" valign="middle" >1</td><td align="center" valign="middle" >−20.142</td><td align="center" valign="middle" >0.693</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >20</td><td align="center" valign="middle" >2</td><td align="center" valign="middle" >47.999</td><td align="center" valign="middle" >1.59</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >21</td><td align="center" valign="middle" >3</td><td align="center" valign="middle" >136.267</td><td align="center" valign="middle" >4.481</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle"  rowspan="3"  >D</td><td align="center" valign="middle" >22</td><td align="center" valign="middle" >1</td><td align="center" valign="middle" >37.588</td><td align="center" valign="middle" >0.455</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >23</td><td align="center" valign="middle" >2</td><td align="center" valign="middle" >−29.939</td><td align="center" valign="middle" >1.917</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" >24</td><td align="center" valign="middle" >3</td><td align="center" valign="middle" >77.585</td><td align="center" valign="middle" >6.087</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr></tbody></table></table-wrap><p>Models 13, 14 and 15 belong to project groups 1, 2 and 3 respectively using size and DevQ as independent variables.</p><p>Models 16, 17 and 18 refer to project groups 1, 2 and 3 using size, DevQ and TestQ as independent variables.</p><p>Portfolio C models based on dataset C (COSMIC FP Projects):</p><p>Model 19, 20 and 21 relate to project groups 1, 2 and 3 respectively using size as independent variable.</p><p>Portfolio D models based on dataset D (IFPUG FP Projects):</p><p>Model 22, 23 and 24 relate to project groups 1, 2 and 3 respectively using size as independent variable. Depending upon the number of independent variables, model equations have coefficients A, B, D1, D2, T1 and T2 (<xref ref-type="table" rid="table7">Table 7</xref>), which were used to estimate the value for test effort for specific values of size, DevQ and TestQ as explained next.</p><p>Using estimation models based on size:</p><p>Test effort for a particular functional size can be estimated from models using the following equation representing size based estimation models:</p><disp-formula id="scirp.75710-formula210"><label>(1)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/3-9302379x4.png"  xlink:type="simple"/></disp-formula><p>Test effort for a particular functional size can be computed by using the values of A and B from <xref ref-type="table" rid="table7">Table 7</xref> and substituting functional size for “size” in Equation (1).</p><p>Using estimation models based on size and DevQ:</p><p>Test effort for a particular value of functional size and DevQ can be estimated from models using the following equation:</p><disp-formula id="scirp.75710-formula211"><label>(2)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/3-9302379x5.png"  xlink:type="simple"/></disp-formula><p>Test effort for a particular functional size where rating for DevQ is available can be computed using Equation (2). D1 and D2 have different values based on the value of DevQ. Appropriate values from <xref ref-type="table" rid="table7">Table 7</xref> are to be chosen depending on whether DevQ = 0 or Dev Q = 1. For DevQ = 2, the value is 0, the base value considered while modelling.</p><p>Using estimation models based on size, DevQ and TestQ:</p><p>The equation for estimating Test Effort for particular values of size, DevQ and TestQ from the model has the form:</p><disp-formula id="scirp.75710-formula212"><label>(3)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/3-9302379x6.png"  xlink:type="simple"/></disp-formula><p>Equation (3) can be used for computing test effort estimate for a particular functional size when ratings for both DevQ and TestQ are available. Values for D1 and D2 are to be chosen from <xref ref-type="table" rid="table7">Table 7</xref> depending upon the input value of DevQ, is either 0 or 1. Values for T1 and T2 are provided for TestQ = 0. Values for DevQ = 2 and Test Q = 1 are zero, as they were the baseline for the modelling.</p><p>An estimator chooses the project group by mapping the characteristics of the project to be estimated to the attributes of project group and selects the related data set in order to choose the closest model for estimation.</p></sec><sec id="s4_2"><title>4.2. Evaluation of Estimation Models</title><p>The quality of estimation models was evaluated using criteria such as coefficient of determination (R<sup>2</sup>), Adj R<sup>2</sup>, magnitude of relative error (MRE), median magnitude of relative error (MedMRE) [<xref ref-type="bibr" rid="scirp.75710-ref10">10</xref>] (<xref ref-type="table" rid="table8">Table 8</xref>).</p><p>The value of R<sup>2</sup> for portfolio A ranged between 0.74 and 0.86, and that of Adj R<sup>2</sup> ranged between 0.73 and 0.83 indicating a strong relationship between the independent variables-size, DevQ and TestQ with the dependent variable test effort in all models.</p><p>The value of MedMRE ranging between 0.22 and 0.28 shows that the error levels between the estimate and actual are within the range of 22% to 28% for 50% or less of the samples, which is practical considering the multi-organiza- tional data used for building the models.</p><p>Similar observations can be made for rest of the models.</p><table-wrap id="table8" ><label><xref ref-type="table" rid="table8">Table 8</xref></label><caption><title> Quality of estimation models</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Portfolio</th><th align="center" valign="middle" >Model id</th><th align="center" valign="middle" >No. of projects</th><th align="center" valign="middle" >R<sup>2</sup></th><th align="center" valign="middle" >Adj R<sup>2</sup></th><th align="center" valign="middle" >MedMRE</th></tr></thead><tr><td align="center" valign="middle"  rowspan="9"  >A (N = 142)</td><td align="center" valign="middle" >1</td><td align="center" valign="middle" >46</td><td align="center" valign="middle" >0.82</td><td align="center" valign="middle" >0.81</td><td align="center" valign="middle" >0.24</td></tr><tr><td align="center" valign="middle" >2</td><td align="center" valign="middle" >49</td><td align="center" valign="middle" >0.74</td><td align="center" valign="middle" >0.73</td><td align="center" valign="middle" >0.27</td></tr><tr><td align="center" valign="middle" >3</td><td align="center" valign="middle" >47</td><td align="center" valign="middle" >0.77</td><td align="center" valign="middle" >0.79</td><td align="center" valign="middle" >0.25</td></tr><tr><td align="center" valign="middle" >4</td><td align="center" valign="middle" >46</td><td align="center" valign="middle" >0.85</td><td align="center" valign="middle" >0.83</td><td align="center" valign="middle" >0.24</td></tr><tr><td align="center" valign="middle" >5</td><td align="center" valign="middle" >49</td><td align="center" valign="middle" >0.75</td><td align="center" valign="middle" >0.73</td><td align="center" valign="middle" >0.28</td></tr><tr><td align="center" valign="middle" >6</td><td align="center" valign="middle" >47</td><td align="center" valign="middle" >0.79</td><td align="center" valign="middle" >0.77</td><td align="center" valign="middle" >0.22</td></tr><tr><td align="center" valign="middle" >7</td><td align="center" valign="middle" >46</td><td align="center" valign="middle" >0.86</td><td align="center" valign="middle" >0.83</td><td align="center" valign="middle" >0.23</td></tr><tr><td align="center" valign="middle" >8</td><td align="center" valign="middle" >49</td><td align="center" valign="middle" >0.78</td><td align="center" valign="middle" >0.74</td><td align="center" valign="middle" >0.24</td></tr><tr><td align="center" valign="middle" >9</td><td align="center" valign="middle" >47</td><td align="center" valign="middle" >0.79</td><td align="center" valign="middle" >0.77</td><td align="center" valign="middle" >0.22</td></tr><tr><td align="center" valign="middle"  rowspan="9"  >B (N = 72)</td><td align="center" valign="middle" >10</td><td align="center" valign="middle" >32</td><td align="center" valign="middle" >0.80</td><td align="center" valign="middle" >0.8</td><td align="center" valign="middle" >0.24</td></tr><tr><td align="center" valign="middle" >11</td><td align="center" valign="middle" >24</td><td align="center" valign="middle" >0.67</td><td align="center" valign="middle" >0.66</td><td align="center" valign="middle" >0.26</td></tr><tr><td align="center" valign="middle" >12</td><td align="center" valign="middle" >16</td><td align="center" valign="middle" >0.83</td><td align="center" valign="middle" >0.82</td><td align="center" valign="middle" >0.25</td></tr><tr><td align="center" valign="middle" >13</td><td align="center" valign="middle" >32</td><td align="center" valign="middle" >0.84</td><td align="center" valign="middle" >0.81</td><td align="center" valign="middle" >0.22</td></tr><tr><td align="center" valign="middle" >14</td><td align="center" valign="middle" >24</td><td align="center" valign="middle" >0.70</td><td align="center" valign="middle" >0.62</td><td align="center" valign="middle" >0.25</td></tr><tr><td align="center" valign="middle" >15</td><td align="center" valign="middle" >16</td><td align="center" valign="middle" >0.91</td><td align="center" valign="middle" >0.86</td><td align="center" valign="middle" >0.10</td></tr><tr><td align="center" valign="middle" >16</td><td align="center" valign="middle" >32</td><td align="center" valign="middle" >0.87</td><td align="center" valign="middle" >0.83</td><td align="center" valign="middle" >0.20</td></tr><tr><td align="center" valign="middle" >17</td><td align="center" valign="middle" >24</td><td align="center" valign="middle" >0.70</td><td align="center" valign="middle" >0.57</td><td align="center" valign="middle" >0.25</td></tr><tr><td align="center" valign="middle" >18</td><td align="center" valign="middle" >16</td><td align="center" valign="middle" >0.91</td><td align="center" valign="middle" >0.86</td><td align="center" valign="middle" >0.10</td></tr><tr><td align="center" valign="middle"  rowspan="3"  >C (N = 82)</td><td align="center" valign="middle" >19</td><td align="center" valign="middle" >27</td><td align="center" valign="middle" >0.87</td><td align="center" valign="middle" >0.86</td><td align="center" valign="middle" >0.19</td></tr><tr><td align="center" valign="middle" >20</td><td align="center" valign="middle" >26</td><td align="center" valign="middle" >0.73</td><td align="center" valign="middle" >0.71</td><td align="center" valign="middle" >0.30</td></tr><tr><td align="center" valign="middle" >21</td><td align="center" valign="middle" >29</td><td align="center" valign="middle" >0.82</td><td align="center" valign="middle" >0.82</td><td align="center" valign="middle" >0.23</td></tr><tr><td align="center" valign="middle"  rowspan="3"  >D (N = 60)</td><td align="center" valign="middle" >22</td><td align="center" valign="middle" >19</td><td align="center" valign="middle" >0.78</td><td align="center" valign="middle" >0.77</td><td align="center" valign="middle" >0.25</td></tr><tr><td align="center" valign="middle" >23</td><td align="center" valign="middle" >23</td><td align="center" valign="middle" >0.76</td><td align="center" valign="middle" >0.75</td><td align="center" valign="middle" >0.26</td></tr><tr><td align="center" valign="middle" >24</td><td align="center" valign="middle" >18</td><td align="center" valign="middle" >0.70</td><td align="center" valign="middle" >0.68</td><td align="center" valign="middle" >0.33</td></tr></tbody></table></table-wrap></sec><sec id="s4_3"><title>4.3. Comparison of Model Performance</title><sec id="s4_3_1"><title>4.3.1. Predictive Performance of Models</title><p>The criterion used to evaluate the predictive quality of an estimation model was PRED (l) = k/n, where k is the number of projects in a specific sample of size n for which MRE &lt;= l. In the software engineering literature, an estimation model is considered good when PRED (0.25) = 0.75 [<xref ref-type="bibr" rid="scirp.75710-ref17">17</xref>] or PRED (0.30) = 0.70 and PRED (0.20) = 0.80 [<xref ref-type="bibr" rid="scirp.75710-ref10">10</xref>] . PRED (0.25) = 0.75 means 75% of the samples should have MRE values less than or equal to 0.25. While an MRE error level in 75% of the population less than 0.25 is the expectation of this criterion, multi-organiza- tional data such as in ISBSG data exhibit large MRE for 75% of the population.</p><p>To compare the performance of models, MRE values for 50% and 25% of the population in addition to 75% were taken into consideration. Each vertical bar in the charts (Figures 3-7) depicts the MRE value for 50% in the middle with either extremes showing MRE values for 25% and 75% of the population for each of the identified model. The middle points of each bar (MRE for 50% of population) are connected by a line to visualize the difference between successive models. This point is referred simply as MRE in the following discussions and was used to compare the performance of the models.</p><fig id="fig3"  position="float"><label><xref ref-type="fig" rid="fig3">Figure 3</xref></label><caption><title> Performance of data set A models</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/3-9302379x7.png"/></fig><fig id="fig4"  position="float"><label><xref ref-type="fig" rid="fig4">Figure 4</xref></label><caption><title> Performance of size based models</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/3-9302379x8.png"/></fig><fig id="fig5"  position="float"><label><xref ref-type="fig" rid="fig5">Figure 5</xref></label><caption><title> Size &amp; DevQ models in data set A and B</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/3-9302379x9.png"/></fig><fig id="fig6"  position="float"><label><xref ref-type="fig" rid="fig6">Figure 6</xref></label><caption><title> Size, DevQ &amp; TestQ models in data set A and B</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/3-9302379x10.png"/></fig><fig id="fig7"  position="float"><label><xref ref-type="fig" rid="fig7">Figure 7</xref></label><caption><title> Performance of COSMIC vs. IFPUG models</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/3-9302379x11.png"/></fig></sec><sec id="s4_3_2"><title>4.3.2. Dataset A Models in Portfolio A</title><p>There are 9 models in portfolio A. A comparison of these nine models reveals how predictability varies between project groups and while using different independent variables. <xref ref-type="fig" rid="fig3">Figure 3</xref> depicts the MRE levels of models corresponding to PG1 (the leftmost three bars), PG2 (the next three bars) and PG3 (the rightmost three bars).</p><p>Within PG1, the model with size, DevQ and TestQ as independent variables (model 7) demonstrate lower MRE for 50% of the population compared to the model with size, Dev Q (model 4) which is lower than the model with size alone (model 1).</p><p>A similar pattern is observed in PG3 for models 3, 6 and 9. In the case of PG2, the model using size, DevQ and TestQ (model 5) exhibits higher MRE compared to other models in PG2 (models 2 &amp; 8) as well as for all the models in dataset A.</p></sec><sec id="s4_3_3"><title>4.3.3. Size-Based Models</title><p>A comparison of all models using only size as an independent variable across all portfolios shows under which context size-based models provide better predictability. <xref ref-type="fig" rid="fig4">Figure 4</xref> illustrates size-based models from each portfolio, the first three bars corresponding to each project group in portfolio A, the next three corresponding to each project group in portfolio B and so on.</p><p>In summary:</p><p> Size-based models for project groups PG1 and PG3 are better than PG2 with the exception of portfolio D.</p><p> PG3 size models, in general perform better than PG1 and PG2 except for the last model (Model ID 24).</p></sec><sec id="s4_3_4"><title>4.3.4. Models in Portfolios A and B</title><p>Portfolio B models were developed using a subset of data used for portfolio A. Portfolio B models were more specific to web or client/server architecture, unlike portfolio A models where there was an approximation due to differences in architecture. A comparison between the models across portfolios A and B using the independent variables DevQ and TestQ along with size helped to make certain observations. <xref ref-type="fig" rid="fig5">Figure 5</xref> depicts size &amp; DevQ models for PG1, PG2 and PG3 for dataset A and B, while <xref ref-type="fig" rid="fig6">Figure 6</xref> illustrates size, DevQ and TestQ models for PG1, PG2 and PG3 for data set A and B.</p><p>Examination of <xref ref-type="fig" rid="fig5">Figure 5</xref> reveals that models in portfolio B (models 13, 14,15,16 &amp; 18) performed much better than models in portfolio A (models 4 to 9) with model 17 being an exception.</p></sec><sec id="s4_3_5"><title>4.3.5. COSMIC and IFPUG Models</title><p>The performance of COSMIC (dataset C) and IFPUG (dataset D) models was compared next using size-based models from portfolio A as the reference. Both COSMIC and IFPUG data are subsets of dataset A consisting of projects measured using the corresponding sizing method. This comparison can help to evaluate prediction accuracy of COSMIC-based models versus IFPUG-based models. <xref ref-type="fig" rid="fig7">Figure 7</xref> depicts PG1 models for dataset A (model 1), COSMIC (model 19), IFPUG (model 22), PG2 models for data set A (model 2), COSMIC (model 20), IFPUG (model 23) and PG3 (model 3), COSMIC (model 21) and IFPUG (model 24).</p><p>COSMIC-based estimation models using dataset C had better performance than IFPUG-based estimation models using data set D, with the exception of PG2 (model 20). COSMIC-based PG3 model demonstrated the best predictability. Furthermore, the R<sup>2</sup> values for COSMIC-based models ranged from 0.73 to 0.87 while that of IFPUG-based models ranged from 0.70 to 0.78 (<xref ref-type="table" rid="table8">Table 8</xref>). The MedMRE value for COSMIC-based models ranged between 0.19 and 0.30 compared to IFPUG based models ranging between 0.25 and 0.33 (<xref ref-type="table" rid="table8">Table 8</xref>) demonstrating better accuracy of COSMIC based models.</p></sec></sec></sec><sec id="s5"><title>5. Conclusions</title><p>This research work explored software testing from the perspective of estimation of efforts for functional testing. The ISBSG database, with its wealth of project data from around the globe, was used for the first time in building effort models for functional testing. The analysis of the data revealed three test productivity patterns representing economies and diseconomies of scale, based on which characteristics of the corresponding projects were investigated. Three project groups, characterized by domain, team size, elapsed time and rigor of verification and validation, and related to three productivity patterns were found to be statistically significant. Within each project group, the variations in test effort could be explained, apart from the functional size, by 1) the processes executed during the development, and 2) the processes adopted for testing.</p><p>Two new independent variables, DevQ and TestQ were identified as influential in the estimation of effort. A total of 24 models were built, using combinations of the three independent variables. The quality of each model was evaluated using established criteria such as R<sup>2</sup>, Adj R<sup>2</sup>, MRE and MedMRE. As these models were built from ISBSG data, they could serve as an industry benchmark for functional test efforts. Test estimation models using projects measured in COSMIC function point exhibited better quality and resulted in more accurate estimates compared to projects measured in IFPUG function points.</p><p>The models are applicable only for the ranges of size in the data set and for testing of business applications. The models generated are not applicable for enhancement projects. These limitations can be overcome by generating specific models for enhancements or real-time projects, using an approach like the one followed in this work. This may require identification of additional project characteristics, as well as other variables influencing testing effort. PG4―the fourth group of project data points remains to be analyzed.</p><p>The process factors used for rating DevQ and TestQ can be further refined within organizational context. There could be other variables that influence test efforts in specific contexts, which would require further study and analysis. The estimation models designed can be further refined by considering testing techniques adopted as a parameter to evaluate their impact and then used to build estimation models.</p></sec><sec id="s6"><title>Acknowledgements</title><p>The authors would like to acknowledge the support provided by COSMIC consultant Srikanth Arvamudhan and statistician Sriram Ramachandran during this work.</p></sec><sec id="s7"><title>Cite this paper</title><p>Jayakumar, K.R. and Abran, A. (2017) Estimation Models for Software Functional Test Effort. Journal of Software Engineering and Applications, 10, 338-353. https://doi.org/10.4236/jsea.2017.104020</p></sec></body><back><ref-list><title>References</title><ref id="scirp.75710-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">Conte, S.D., Dunsmore, D.E. and Shen, V.Y. (1986) Software Engineering Metrics and Models. The Benjamin/Cummings Publishing Company, Inc., Menlo Park.</mixed-citation></ref><ref id="scirp.75710-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">Galorath, D. (2015) Why Can’t People Estimate: Estimation Bias and Mitigation. Conference presentation at IT Confidence Conference, Rio de Janerio, October 2015.</mixed-citation></ref><ref id="scirp.75710-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">Hill, P., ISBSG (2010) Practical Software Project Estimation: A Toolkit for Estimating Software Development Effort and Duration. McGraw-Hill, New York.</mixed-citation></ref><ref id="scirp.75710-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">Bharadwaj, M. and Rana, A. (2015) Estimation of Testing and Rework Efforts for software Development Projects. Asian Journal of Computer Science and Information Technology, 5.</mixed-citation></ref><ref id="scirp.75710-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">Taipale, O. (2007) Observations on Software Testing Practice. PhD Thesis, Lappeenranta University of Technology, Finland.</mixed-citation></ref><ref id="scirp.75710-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">Armel, K. (2012) Top Performing Projects Use Small Teams Deliver Lower Cost, Higher Quality, Blog Posting. Quantitative Software Management, Inc., USA.</mixed-citation></ref><ref id="scirp.75710-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">Putnam, D. (2005) Team Size Can Be the Key to a Successful Software Project. Quantitative Software Management Inc., USA. www.qsm.com</mixed-citation></ref><ref id="scirp.75710-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">Abran, A. (2015) Software Project Estimation: The Fundamentals for Providing High Quality Information to Decision Makers. John Wiley &amp; Sons &amp; IEEE Computer Society, New Jersey. https://doi.org/10.1002/9781118959312</mixed-citation></ref><ref id="scirp.75710-ref9"><label>9</label><mixed-citation publication-type="book" xlink:type="simple">Dumke, R. and Abran, A., Eds., (2011) COSMIC Function Points, Theory and Advanced Practices, Chapter 3.5: Measurement Convertibility—From Function Points to COSMIC FFP. CRC Press, New York.</mixed-citation></ref><ref id="scirp.75710-ref10"><label>10</label><mixed-citation publication-type="other" xlink:type="simple">Bala, A. and Abran, A. (2016) Use of the Multiple Imputation Strategy to Deal with Missing Data in the ISBSG Repository. Journal of Information Technology &amp; Software Engineering, 6, 171.</mixed-citation></ref><ref id="scirp.75710-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">Bala, A. (2013) Impact Analysis of a Multiple Imputation Technique for Handling Missing Value in the ISBSG Repository of Software Project. PhD Thesis, Ecole de technologie superieure, University of Quebec, Montreal (Canada), 17 October.</mixed-citation></ref><ref id="scirp.75710-ref12"><label>12</label><mixed-citation publication-type="other" xlink:type="simple">ISBSG (2013) Repository Data Release 12—Field Descriptions, “e.Field Descriptions —Data Release 12. Pdf” Document Provided as a Part of Data Set. International Software Benchmarking and Standards Group.</mixed-citation></ref><ref id="scirp.75710-ref13"><label>13</label><mixed-citation publication-type="other" xlink:type="simple">COSMIC (2015) The COSMIC Functional Size Measurement Method, Version 4.0.1, Measurement Manual, The COSMIC Implementation Guide for ISO/IEC 19761:2011. Common Software Measurements International Consortium, Canada, April 2015. www.cosmic-sizing.org</mixed-citation></ref><ref id="scirp.75710-ref14"><label>14</label><mixed-citation publication-type="other" xlink:type="simple">ISO/IEC 19761 (2011) Software Engineering—COSMIC—A Functional Size Measurement Method. International Organization for Standardization (ISO), Geneva.</mixed-citation></ref><ref id="scirp.75710-ref15"><label>15</label><mixed-citation publication-type="other" xlink:type="simple">Abran, A. (2010) Software Metrics and Software Metrology. Wiley &amp; IEEE Computer Society Press, New Jersey. https://doi.org/10.1002/9780470606834</mixed-citation></ref><ref id="scirp.75710-ref16"><label>16</label><mixed-citation publication-type="other" xlink:type="simple">Jayakumar, K.R. and Abran, A. (2013) A Survey of Software Test Estimation Techniques. Journal of Software Engineering and Applications, 6, 47-52.  
https://doi.org/10.4236/jsea.2013.610A006</mixed-citation></ref><ref id="scirp.75710-ref17"><label>17</label><mixed-citation publication-type="other" xlink:type="simple">ISO/IEC 20926:2009 (2009) Software and Systems Engineering—Software Measurement—IFPUG Functional Size Measurement Method. International Organization for Standardization (ISO), Geneva.</mixed-citation></ref></ref-list></back></article>