<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">OJS</journal-id><journal-title-group><journal-title>Open Journal of Statistics</journal-title></journal-title-group><issn pub-type="epub">2161-718X</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/ojs.2015.57077</article-id><article-id pub-id-type="publisher-id">OJS-62191</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Physics&amp;Mathematics</subject></subj-group></article-categories><title-group><article-title>
 
 
  Nonnegative Matrix Factorization with Zellner Penalty
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>atthew</surname><given-names>A. Corsetti</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Ernest</surname><given-names>Fokoué</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib></contrib-group><aff id="aff1"><addr-line>School of Mathematical Sciences, Rochester Institute of Technology, Rochester, NY, USA</addr-line></aff><author-notes><corresp id="cor1">* E-mail:<email>mac4017@rit.edu(AAC)</email>;<email>epfeqa@rit.edu(EF)</email>;</corresp></author-notes><pub-date pub-type="epub"><day>11</day><month>12</month><year>2015</year></pub-date><volume>05</volume><issue>07</issue><fpage>777</fpage><lpage>786</lpage><history><date date-type="received"><day>12</day>	<month>November</month>	<year>2015</year></date><date date-type="rev-recd"><day>accepted</day>	<month>21</month>	<year>December</year>	</date><date date-type="accepted"><day>24</day>	<month>December</month>	<year>2015</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
   Nonnegative matrix factorization (NMF) is a relatively new unsupervised learning algorithm that decomposes a nonnegative data matrix into a parts-based, lower dimensional, linear representation of the data. NMF has applications in image processing, text mining, recommendation systems and a variety of other fields. Since its inception, the NMF algorithm has been modified and explored by numerous authors. One such modification involves the addition of auxiliary constraints to the objective function of the factorization. The purpose of these auxiliary constraints is to impose task-specific penalties or restrictions on the objective function. Though many auxiliary constraints have been studied, none have made use of data-dependent penalties. In this paper, we propose Zellner nonnegative matrix factorization (ZNMF), which uses data-dependent auxiliary constraints. We assess the facial recognition performance of the ZNMF algorithm and several other well-known constrained NMF algorithms using the Cambridge ORL database. 
 
</p></abstract><kwd-group><kwd>Nonnegative Matrix Factorization</kwd><kwd> Zellner g-Prior</kwd><kwd> Auxiliary Constraints</kwd><kwd> Regularization</kwd><kwd> Penalty</kwd><kwd> Classification</kwd><kwd> Image Processing</kwd><kwd> Feature Extraction</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>Visual recognition tasks have become increasingly popular and complex in the last several decades as they often involve massively large datasets. Facial detection and recognition tasks are particularly of interest and can be severely complicated due to variation in illumination, emotional expression as well as physical location and orientation of the face within an image. Due to the often massive size of facial image datasets, subspace methods are frequently used to identify latent variables and reduce data dimensionality, so as to produce apposite representations of facial image databases.</p><p>Nonnegative matrix factorization (NMF) is a relatively new unsupervised learning subspace method that was first introduced in 1999 by Lee and Seung [<xref ref-type="bibr" rid="scirp.62191-ref1">1</xref>] . NMF factorizes a data matrix <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1240604x6.png" xlink:type="simple"/></inline-formula> while imposing a nonnegativity constraint on the matrix X. The subsequent nonnegative basis matrix <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1240604x7.png" xlink:type="simple"/></inline-formula> and nonnegative coefficient matrix <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1240604x8.png" xlink:type="simple"/></inline-formula> approximate X when multiplied together (i.e. X &#187; WH). NMF produces a sparse, part-based representation of the database as the nonnegativity constraint allows for additive, but not subtractive combinations of components. Because of this property, NMF is frequently used as a dimensionality reduction technique for tasks in which it is intuitive to combine parts to form a complete object such as in image processing, facial recognition [<xref ref-type="bibr" rid="scirp.62191-ref1">1</xref>] -[<xref ref-type="bibr" rid="scirp.62191-ref4">4</xref>] or community network visualizations [<xref ref-type="bibr" rid="scirp.62191-ref5">5</xref>] .</p><p>Suppose <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1240604x9.png" xlink:type="simple"/></inline-formula> is a database of faces, for which n represents the total number of images in the database and p represents the number of pixels within each image (assumed to be constant across all images in the data matrix X). NMF factorizes the nonnegative data matrix X into W and H by minimizing a cost function―most commonly a generalization of the square of the Euclidean Distance to matrix space</p><disp-formula id="scirp.62191-formula220"><label>(1)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/13-1240604x10.png"  xlink:type="simple"/></disp-formula><p>or the Kullback-Leibler Divergence</p><disp-formula id="scirp.62191-formula221"><label>(2)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/13-1240604x11.png"  xlink:type="simple"/></disp-formula><p>Many authors have adapted the NMF algorithm by altering either the cost function formulation [<xref ref-type="bibr" rid="scirp.62191-ref3">3</xref>] [<xref ref-type="bibr" rid="scirp.62191-ref4">4</xref>] [<xref ref-type="bibr" rid="scirp.62191-ref6">6</xref>] -[<xref ref-type="bibr" rid="scirp.62191-ref9">9</xref>] , the minimization method for solving (1) or (2) [<xref ref-type="bibr" rid="scirp.62191-ref10">10</xref>] -[<xref ref-type="bibr" rid="scirp.62191-ref12">12</xref>] , or the initialization strategy for W and/or H [<xref ref-type="bibr" rid="scirp.62191-ref13">13</xref>] -[<xref ref-type="bibr" rid="scirp.62191-ref15">15</xref>] . Relatively new adaptations of the NMF algorithm involve applying secondary constraints to the W and/or H matrix. These often take the form of smoothness constraints [<xref ref-type="bibr" rid="scirp.62191-ref8">8</xref>] [<xref ref-type="bibr" rid="scirp.62191-ref16">16</xref>] or sparsity constraints [<xref ref-type="bibr" rid="scirp.62191-ref17">17</xref>] -[<xref ref-type="bibr" rid="scirp.62191-ref19">19</xref>] . These constraints are added so as to encode prior information regarding the nature of the application under examination or to ensure preferred characteristics in the solution for W and H. For constrained NMF (CNMF), penalty terms are used to apply the secondary constraints on W and H. This results in an extension to the optimization task provided in (1):</p><disp-formula id="scirp.62191-formula222"><label>(3)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/13-1240604x12.png"  xlink:type="simple"/></disp-formula><p>Here J<sub>1</sub>(W) and J<sub>2</sub>(H) represent the penalty terms and 0 ≤ α ≤ 1 and 0 ≤ β ≤ 1 are the regularization parameters that specify the relationship between the constraints. Often a sparsity constraint and an approximation error constraint are used.</p><p>Though there are many adaptations of the NMF algorithm in which auxiliary constraints are imposed on W and H, none of these methodologies make use of data-dependent penalties. Inspired by the so-called Zellner’s g-Prior [<xref ref-type="bibr" rid="scirp.62191-ref20">20</xref>] , used in Bayesian Regression Analysis, we explore the use of two penalty terms that are data dependent. We use the ORL database to test the facial classification capability of the NMF algorithm when constrained by Zellner g-Prior penalties, henceforth referred to as Zellner nonnegative matrix factorization (ZNMF). We compare the facial classification capability of ZNMF with Constrained nonnegative matrix factorization (CNMF) [<xref ref-type="bibr" rid="scirp.62191-ref8">8</xref>] and show that it is superior across all selected factorization ranks. We also compare the ZNMF recognition performance with the algorithms described in [<xref ref-type="bibr" rid="scirp.62191-ref4">4</xref>] and determine that it outperforms many of them across many of the selected factorization ranks, most notably, the smaller of the factorization ranks.</p></sec><sec id="s2"><title>2. Nonnegative Matrix Factorization Algorithms</title><sec id="s2_1"><title>2.1. Traditional Nonnegative Matrix Factorization with Multiplicative Update</title><p>NMF factorizes a matrix <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1240604x13.png" xlink:type="simple"/></inline-formula> into a basis matrix <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1240604x14.png" xlink:type="simple"/></inline-formula> and a coefficient matrix <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1240604x15.png" xlink:type="simple"/></inline-formula> while imposing a nonnegativity constraint. Because of the nonnegativity constraint, the basis images (when considering X to be a database of faces) can be combined in an additive fashion to form a complete face. In traditional NMF [<xref ref-type="bibr" rid="scirp.62191-ref1">1</xref>] the two most commonly considered cost functions for determining the cost of factorizing X into W and H are the square of the Euclidean (1) and the Kullback-Leibler Divergence (2). Traditional NMF produces the W and H matrices by calculating minimizations of (1) or (2) using the following multiplicative update equations:</p><disp-formula id="scirp.62191-formula223"><label>(4)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/13-1240604x16.png"  xlink:type="simple"/></disp-formula><p>where</p><disp-formula id="scirp.62191-formula224"><label>(5)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/13-1240604x17.png"  xlink:type="simple"/></disp-formula><p>and</p><disp-formula id="scirp.62191-formula225"><label>(6)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/13-1240604x18.png"  xlink:type="simple"/></disp-formula><p>where</p><disp-formula id="scirp.62191-formula226"><label>(7)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/13-1240604x19.png"  xlink:type="simple"/></disp-formula><p>Traditional NMF using a multiplicative update is known to be slow to converge as it requires a large number of iterations. Gradient descent and alternating least squares algorithms are commonly used in place of traditional NMF as they require far fewer iterations resulting in a faster convergence; however, we will not explore them in this paper.</p><p>The standard NMF multiplicative updating algorithms have a continuous descent property. The descent will lead to a stationary point within the region under examination; however, it is uncertain as to whether or not this stationary point is a local minimum as it could certainly be a saddle point. This is due to the iterative optimization nature of the algorithm which optimizes W and H iteratively, though never simultaneously.</p></sec><sec id="s2_2"><title>2.2. Constrained Nonnegative Matrix Factorization</title><p>CNMF [<xref ref-type="bibr" rid="scirp.62191-ref8">8</xref>] expands the optimization task shown in (1) to include penalty terms J<sub>1</sub>(W) and J<sub>2</sub>(H) that serve to apply task-specific, auxiliary constraints on the solutions of (3). 0 ≤ α ≤ 1, 0 ≤ β ≤ 1 are regularization parameters. For our purposes we define J<sub>1</sub>(W) and J<sub>2</sub>(H) from (3) as follows:</p><disp-formula id="scirp.62191-formula227"><label>(8)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/13-1240604x20.png"  xlink:type="simple"/></disp-formula><p>and</p><disp-formula id="scirp.62191-formula228"><label>(9)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/13-1240604x21.png"  xlink:type="simple"/></disp-formula><p>while α and β are used as constraints on the sparsity and approximation error respectively. When the optimization task is that of (3), the multiplicative updates of (4) and (6) are modified as follows:</p><disp-formula id="scirp.62191-formula229"><label>(10)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/13-1240604x22.png"  xlink:type="simple"/></disp-formula><p>and</p><disp-formula id="scirp.62191-formula230"><label>(11)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/13-1240604x23.png"  xlink:type="simple"/></disp-formula></sec><sec id="s2_3"><title>2.3. Zellner Nonnegative Matrix Factorization</title><p>In regression analysis, for a Gaussian Distribution with <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1240604x24.png" xlink:type="simple"/></inline-formula> it is known that</p><disp-formula id="scirp.62191-formula231"><label>(12)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/13-1240604x25.png"  xlink:type="simple"/></disp-formula><p>has variance</p><disp-formula id="scirp.62191-formula232"><label>(13)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/13-1240604x26.png"  xlink:type="simple"/></disp-formula><p>The Zellner g-Prior exploits this fact, in the Bayesian setting to use</p><disp-formula id="scirp.62191-formula233"><label>(14)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/13-1240604x27.png"  xlink:type="simple"/></disp-formula><p>which corresponds to using the penalty’s empirical risk shown below:</p><disp-formula id="scirp.62191-formula234"><graphic  xlink:href="http://html.scirp.org/file/13-1240604x28.png"  xlink:type="simple"/></disp-formula><disp-formula id="scirp.62191-formula235"><graphic  xlink:href="http://html.scirp.org/file/13-1240604x29.png"  xlink:type="simple"/></disp-formula><p>We extend and adapt Zellner’s ideas as follows:</p><disp-formula id="scirp.62191-formula236"><label>(15)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/13-1240604x30.png"  xlink:type="simple"/></disp-formula><p>where S = (X<sup>T</sup>W) is n &#215; q, S<sup>T</sup> is q &#215; n and S<sup>T</sup>S is q &#215; q.</p><disp-formula id="scirp.62191-formula237"><label>(16)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/13-1240604x31.png"  xlink:type="simple"/></disp-formula><p>where R = XH<sup>T</sup> is p &#215; q and essentially represents the projection weighting. R<sup>T</sup>R is q &#215; q and its diagonal essentially represents the idiosyncratic variance of the projections onto the lower dimensional space.</p><disp-formula id="scirp.62191-formula238"><label>(17)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/13-1240604x32.png"  xlink:type="simple"/></disp-formula><disp-formula id="scirp.62191-formula239"><label>(18)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/13-1240604x33.png"  xlink:type="simple"/></disp-formula><p>As can be seen in (17) and (18), the updates of both W and H are simply post or pre-weighted by the input space variances or the data spaces variances.</p><p>Our objective function is</p><disp-formula id="scirp.62191-formula240"><label>(19)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/13-1240604x34.png"  xlink:type="simple"/></disp-formula><p>where</p><disp-formula id="scirp.62191-formula241"><label>(20)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/13-1240604x35.png"  xlink:type="simple"/></disp-formula><p>and</p><disp-formula id="scirp.62191-formula242"><label>(21)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/13-1240604x36.png"  xlink:type="simple"/></disp-formula><p>where</p><disp-formula id="scirp.62191-formula243"><label>(22)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/13-1240604x37.png"  xlink:type="simple"/></disp-formula><p>For (22) the Risk Inflation Criterion (RIC) of Foster and George [<xref ref-type="bibr" rid="scirp.62191-ref21">21</xref>] , which sets g = p<sup>2</sup> is combined with the Bayesian Information Criterion (BIC) to produce the so-called benchmark prior as g = n leads to the unit information prior found in BIC. And so, (22) will be found to be appropriate.</p><p>When using ZNMF, the updating equations of CNMF shown in (10) and (11) are modified as follows:</p><disp-formula id="scirp.62191-formula244"><label>(23)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/13-1240604x38.png"  xlink:type="simple"/></disp-formula><p>and</p><disp-formula id="scirp.62191-formula245"><label>(24)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/13-1240604x39.png"  xlink:type="simple"/></disp-formula></sec></sec><sec id="s3"><title>3. Experimental Results</title><p>In this section, we conducted a series of simulations to evaluate the classification performance of the ZNMF and CNMF algorithms. We replicated the ORL classification experiment conducted in Wang et al. [<xref ref-type="bibr" rid="scirp.62191-ref4">4</xref>] , which evaluated the classification performance of traditional NMF, Local NMF (LNMF) [<xref ref-type="bibr" rid="scirp.62191-ref6">6</xref>] , Fisher NMF (FNMF) [<xref ref-type="bibr" rid="scirp.62191-ref4">4</xref>] , Principle Component Analysis, and Principle Component Analysis NMF (PNMF) [<xref ref-type="bibr" rid="scirp.62191-ref4">4</xref>] using the Cambridge ORL database. By replicating the experiment in [<xref ref-type="bibr" rid="scirp.62191-ref4">4</xref>] , many hundreds of times, we created an avenue through which direct comparisons of the performances of the aforementioned algorithms could be carried out.</p><p>The Cambridge ORL database consists of 10 gray-scale facial images each of 36 male and 4 female subjects. The images vary in illumination, facial expression and position. The faces are forward-facing with slight rotations to the left and right. For each simulation, the training dataset X &#206; ℝ<sup>644&#215;200</sup> was produced by randomly selecting 5 images from each of the 40 subjects resulting in a training dataset of 200 images of 644 pixels each. The test datasets were comprised of the remaining 200 unselected images and were used to evaluate the facial recognition capabilities of CNMF and ZNMF using the first Nearest Neighbor classifier. In order to optimize the computational efficiency the resolution of the images was reduced from 112 &#215; 92 to 28 &#215; 23 in accordance with [<xref ref-type="bibr" rid="scirp.62191-ref4">4</xref>] , which found that reducing the resolution of the ORL faces to 25% of the original resolution had little effect on the accuracy of the facial recognition. The reduction in resolution is demonstrated for 9 images shown in <xref ref-type="fig" rid="fig1">Figure 1</xref>.</p><fig-group id="fig1"><label><xref ref-type="fig" rid="fig1">Figure 1</xref></label><caption><title> (Left) 9 ORL faces at full 112 &#215; 92 resolution; (Right) 9 ORL faces at reduced 28 &#215; 23 resolution.</title></caption><fig id ="fig1_1"><label></label><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/13-1240604x40.png"/></fig><fig id ="fig1_2"><label></label><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/13-1240604x41.png"/></fig></fig-group><p>The effects of the α, β, and g-Prior parameter settings on the average recognition rate were explored to great lengths through extensive computer simulations. We restricted the α, β relationship to two possible scenarios:</p><disp-formula id="scirp.62191-formula246"><label>(25)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/13-1240604x42.png"  xlink:type="simple"/></disp-formula><p>and</p><disp-formula id="scirp.62191-formula247"><label>(26)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/13-1240604x43.png"  xlink:type="simple"/></disp-formula><p>such that 0 ≤ α ≤ 1, 0 ≤ β ≤ 1. Optimal α and β settings were determined across all considered factorization ranks q &#206; {16, 25, 36, 49, 64, 81, 100} for the CNMF algorithm simulations using (25) and (26) (see <xref ref-type="table" rid="table1">Table 1</xref> and <xref ref-type="table" rid="table2">Table 2</xref>). This was not the case for the ZNMF simulations because of the addition of the g-Prior parameter which dramatically increased the number of possible settings for the regularization parameters. Because there were many</p><table-wrap id="table1" ><label><xref ref-type="table" rid="table1">Table 1</xref></label><caption><title> CNMF optimized regularization parameter settings and recognition performances using the α β relationships of (25) and (26)</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Factorization Rank (q)</th><th align="center" valign="middle" >α and β Relationship</th><th align="center" valign="middle" >Optimal α</th><th align="center" valign="middle" >Optimal β</th><th align="center" valign="middle" >Average Recognition Rate</th></tr></thead><tr><td align="center" valign="middle"  rowspan="2"  >16</td><td align="center" valign="middle" >α = β</td><td align="center" valign="middle" >0.81</td><td align="center" valign="middle" >0.81</td><td align="center" valign="middle" >0.88578</td></tr><tr><td align="center" valign="middle" >α = 1 − β</td><td align="center" valign="middle" >0.15</td><td align="center" valign="middle" >0.85</td><td align="center" valign="middle" >0.88602</td></tr><tr><td align="center" valign="middle"  rowspan="2"  >25</td><td align="center" valign="middle" >α = β</td><td align="center" valign="middle" >0.64</td><td align="center" valign="middle" >0.64</td><td align="center" valign="middle" >0.88631</td></tr><tr><td align="center" valign="middle" >α = 1 − β</td><td align="center" valign="middle" >0.39</td><td align="center" valign="middle" >0.61</td><td align="center" valign="middle" >0.88698</td></tr><tr><td align="center" valign="middle"  rowspan="2"  >36</td><td align="center" valign="middle" >α = β</td><td align="center" valign="middle" >0.56</td><td align="center" valign="middle" >0.56</td><td align="center" valign="middle" >0.88630</td></tr><tr><td align="center" valign="middle" >α = 1 − β</td><td align="center" valign="middle" >0.82</td><td align="center" valign="middle" >0.18</td><td align="center" valign="middle" >0.88347</td></tr><tr><td align="center" valign="middle"  rowspan="2"  >49</td><td align="center" valign="middle" >α = β</td><td align="center" valign="middle" >0.99</td><td align="center" valign="middle" >0.99</td><td align="center" valign="middle" >0.88685</td></tr><tr><td align="center" valign="middle" >α = 1 − β</td><td align="center" valign="middle" >0.82</td><td align="center" valign="middle" >0.18</td><td align="center" valign="middle" >0.88556</td></tr><tr><td align="center" valign="middle"  rowspan="2"  >64</td><td align="center" valign="middle" >α = β</td><td align="center" valign="middle" >0.42</td><td align="center" valign="middle" >0.42</td><td align="center" valign="middle" >0.88838</td></tr><tr><td align="center" valign="middle" >α = 1 − β</td><td align="center" valign="middle" >0.46</td><td align="center" valign="middle" >0.54</td><td align="center" valign="middle" >0.88561</td></tr><tr><td align="center" valign="middle"  rowspan="2"  >81</td><td align="center" valign="middle" >α = β</td><td align="center" valign="middle" >1.00</td><td align="center" valign="middle" >1.00</td><td align="center" valign="middle" >0.88868</td></tr><tr><td align="center" valign="middle" >α = 1 − β</td><td align="center" valign="middle" >0.14</td><td align="center" valign="middle" >0.86</td><td align="center" valign="middle" >0.88460</td></tr><tr><td align="center" valign="middle"  rowspan="2"  >100</td><td align="center" valign="middle" >α = β</td><td align="center" valign="middle" >0.94</td><td align="center" valign="middle" >0.94</td><td align="center" valign="middle" >0.88901</td></tr><tr><td align="center" valign="middle" >α = 1 − β</td><td align="center" valign="middle" >0.12</td><td align="center" valign="middle" >0.88</td><td align="center" valign="middle" >0.88496</td></tr></tbody></table></table-wrap><table-wrap id="table2" ><label><xref ref-type="table" rid="table2">Table 2</xref></label><caption><title> ZNMF optimized regularization parameter settings and recognition performances using the α β relationships of (25) and (26)</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Factorization Rank (q)</th><th align="center" valign="middle" >α and β Relationship</th><th align="center" valign="middle" >Optimal α</th><th align="center" valign="middle" >Optimal β</th><th align="center" valign="middle" >Optimal g-Prior</th><th align="center" valign="middle" >Average Recognition Rate</th></tr></thead><tr><td align="center" valign="middle"  rowspan="2"  >16</td><td align="center" valign="middle" >α = β</td><td align="center" valign="middle" >0.45</td><td align="center" valign="middle" >0.45</td><td align="center" valign="middle" >75</td><td align="center" valign="middle" >0.90283</td></tr><tr><td align="center" valign="middle" >α = 1 − β</td><td align="center" valign="middle" >0.60</td><td align="center" valign="middle" >0.40</td><td align="center" valign="middle" >80</td><td align="center" valign="middle" >0.90080</td></tr><tr><td align="center" valign="middle"  rowspan="2"  >25</td><td align="center" valign="middle" >α = β</td><td align="center" valign="middle" >0.45</td><td align="center" valign="middle" >0.45</td><td align="center" valign="middle" >75</td><td align="center" valign="middle" >0.89975</td></tr><tr><td align="center" valign="middle" >α = 1 − β</td><td align="center" valign="middle" >0.60</td><td align="center" valign="middle" >0.40</td><td align="center" valign="middle" >80</td><td align="center" valign="middle" >0.90200</td></tr><tr><td align="center" valign="middle"  rowspan="2"  >36</td><td align="center" valign="middle" >α = β</td><td align="center" valign="middle" >0.45</td><td align="center" valign="middle" >0.45</td><td align="center" valign="middle" >75</td><td align="center" valign="middle" >0.90039</td></tr><tr><td align="center" valign="middle" >α = 1 − β</td><td align="center" valign="middle" >0.60</td><td align="center" valign="middle" >0.40</td><td align="center" valign="middle" >80</td><td align="center" valign="middle" >0.89933</td></tr><tr><td align="center" valign="middle"  rowspan="2"  >49</td><td align="center" valign="middle" >α = β</td><td align="center" valign="middle" >0.45</td><td align="center" valign="middle" >0.45</td><td align="center" valign="middle" >75</td><td align="center" valign="middle" >0.90093</td></tr><tr><td align="center" valign="middle" >α = 1 − β</td><td align="center" valign="middle" >0.60</td><td align="center" valign="middle" >0.40</td><td align="center" valign="middle" >80</td><td align="center" valign="middle" >0.89952</td></tr><tr><td align="center" valign="middle"  rowspan="2"  >64</td><td align="center" valign="middle" >α = β</td><td align="center" valign="middle" >0.45</td><td align="center" valign="middle" >0.45</td><td align="center" valign="middle" >75</td><td align="center" valign="middle" >0.90055</td></tr><tr><td align="center" valign="middle" >α = 1 − β</td><td align="center" valign="middle" >0.60</td><td align="center" valign="middle" >0.40</td><td align="center" valign="middle" >80</td><td align="center" valign="middle" >0.90219</td></tr><tr><td align="center" valign="middle"  rowspan="2"  >81</td><td align="center" valign="middle" >α = β</td><td align="center" valign="middle" >0.45</td><td align="center" valign="middle" >0.45</td><td align="center" valign="middle" >75</td><td align="center" valign="middle" >0.90024</td></tr><tr><td align="center" valign="middle" >α = 1 − β</td><td align="center" valign="middle" >0.60</td><td align="center" valign="middle" >0.40</td><td align="center" valign="middle" >80</td><td align="center" valign="middle" >0.89918</td></tr><tr><td align="center" valign="middle"  rowspan="2"  >100</td><td align="center" valign="middle" >α = β</td><td align="center" valign="middle" >0.45</td><td align="center" valign="middle" >0.45</td><td align="center" valign="middle" >75</td><td align="center" valign="middle" >0.90111</td></tr><tr><td align="center" valign="middle" >α = 1 − β</td><td align="center" valign="middle" >0.60</td><td align="center" valign="middle" >0.40</td><td align="center" valign="middle" >80</td><td align="center" valign="middle" >0.90012</td></tr></tbody></table></table-wrap><p>more regularization parameter settings to consider for the ZNMF algorithm than the CNMF algorithm, α and β were optimized exclusively at a factorization rank q = 16 for (25) and (26) in the ZNMF simulations. The optimal tuning values of α and β for the ZNMF simulations, were then held constant across the remaining factorization ranks (25, 36, 49, 64, 81, 100), only differing depending upon the relationship of α with β specified by (25) and (26). There were 20 replications used at each unique setting of the regularization parameters for the CNMF algorithm; while only 5 replications were used at each unique setting of the regularization parameters in the ZNMF simulations. The noticeable difference between the number of replications for the CNMF and ZNMF algorithms was again due to the fact that there were far more parameter settings to explore using ZNMF than CNMF.</p><p>The recognition performances of the ZNMF simulations across various settings of α, β, and the g-Prior are on display in <xref ref-type="fig" rid="fig2">Figure 2</xref> and <xref ref-type="fig" rid="fig3">Figure 3</xref>. We were able to determine the optimal settings for α, β, and the g-Prior for the ZNMF simulations using these surfaces. Initially we explored two broad regions. The first, shown to the left in <xref ref-type="fig" rid="fig2">Figure 2</xref>, took into consideration the regularization parameter relationship specified by (25); while the second,</p><fig id="fig2"  position="float"><label><xref ref-type="fig" rid="fig2">Figure 2</xref></label><caption><title> (Left) Recognition performances of ZNMF with α = β; (Right) Recognition performances of ZNMF with α = 1 − β</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/13-1240604x44.png"/></fig><fig id="fig3"  position="float"><label><xref ref-type="fig" rid="fig3">Figure 3</xref></label><caption><title> (Left) Recognition performances of ZNMF with α = β in optimal condensed territory; (Right) Recognition performances of ZNMF with α = 1 − β in optimal condensed territory</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/13-1240604x45.png"/></fig><p>shown in the right of <xref ref-type="fig" rid="fig2">Figure 2</xref> considered the relationship specified by (26). Both surfaces in <xref ref-type="fig" rid="fig2">Figure 2</xref> depict a maximal region defined by relatively low g-Prior values and 0.40 ≤ α ≤ 0.60. The natures of these optimal regions were explored using the surfaces of <xref ref-type="fig" rid="fig3">Figure 3</xref>. There were 25 replications conducted at each of the unique regularization parameter settings in these condensed territories. The optimal parameter settings for the ZNMF algorithm under both condition (25) and condition (26), using a factorization rank q = 16, were discovered atop ridgelines in the optimal territories of <xref ref-type="fig" rid="fig3">Figure 3</xref> and are provided in <xref ref-type="table" rid="table2">Table 2</xref>.</p><p>After identifying optimal settings for the regularization parameters, 500 replications were conducted for both the CNMF and ZNMF algorithms, across the factorization ranks q &#206; {16, 25, 36, 49, 64, 81, 100}, using the optimal parameter settings. The results, provided in <xref ref-type="fig" rid="fig4">Figure 4</xref> were quite telling. The ZNMF algorithm had a better average recognition rate across all factorization ranks for both (25) and (26) than the CNMF algorithm. Furthermore, the ZNMF algorithm produced better average recognition rates than the NMF, LNMF, FNMF, PCA and PNMF algorithms used in [<xref ref-type="bibr" rid="scirp.62191-ref4">4</xref>] across the majority of the factorization ranks. The first exception to this occurred at factorization rank of q = 49, in which ZNMF performed better than NMF, LNMF, and PCA, and approximately equal to FNMF and PNMF. The second and third exceptions occurred at factorization ranks q = 64 and q = 81 where ZNMF outperformed NMF, LNMF, PCA and PNMF and performed approximately equal to FNMF. It should be noted that the ZNMF algorithm was able to maintain relatively higher recognition rates (about 90%) consistently across all factorization ranks, including smaller factorization ranks, such as q = 16 and q = 25 where other algorithms produced lower average recognition rates. This is quite exciting as it implies that ZNMF requires less information (lowered factorization ranks) to produce equally as impressive recognition rates on the ORL database as other algorithms [<xref ref-type="bibr" rid="scirp.62191-ref4">4</xref>] produce when provided with relatively more information (higher factorization ranks).</p></sec><sec id="s4"><title>4. Conclusion and Discussion</title><p>In this paper, we proposed the ZNMF algorithm for the assessment of facial recognition and assessed its capability in this regard using the Cambridge ORL Faces Database. We compared its facial recognition capabilities with traditional NMF and several constrained version of NMF across seven different factorization ranks. We found that ZNMF algorithm outperformed the other algorithms across the majority of the factorization ranks, most notably at the lower factorization ranks where the margin of improvement was the most significant. The FNMF algorithm approximately tied the facial recognition rate of the ZNMF algorithm at three factorization ranks (49, 64 and 81) and the PNMF algorithm approximately tied the ZNMF algorithm at just one factorization rank (49). Quite possibly the most important finding was that the ZNMF algorithm produced facial recognition</p><fig id="fig4"  position="float"><label><xref ref-type="fig" rid="fig4">Figure 4</xref></label><caption><title> Average correct recognition rates of the CNMF and ZNMF algorithms using the ORL database with 500 simulations at each factorization rank q &#206; {16, 25, 36, 49, 64, 81, 100}. q was determined in accordance with [<xref ref-type="bibr" rid="scirp.62191-ref4">4</xref>] </title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/13-1240604x46.png"/></fig><p>rates, using less information (lower factorization ranks), that either out-performed or were comparable to the results of other algorithms at higher factorization ranks. This finding implied that, for the ORL Dataset, the data-dependent ZNMF algorithm could classify facial images better than the other algorithms under examination and it could do so with less information, making it computationally less taxing.</p><p>This paper demonstrates the advantages of including data-dependent auxiliary constraints in the NMF algorithm through the introduction of ZNMF. In the future, we hope to explore other data-dependent auxiliary</p><p>constraints. A possibility would be to use <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1240604x47.png" xlink:type="simple"/></inline-formula> where<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1240604x47.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1240604x48.png" xlink:type="simple"/></inline-formula>, and <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1240604x47.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1240604x48.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1240604x49.png" xlink:type="simple"/></inline-formula> where<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1240604x47.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1240604x48.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1240604x49.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1240604x50.png" xlink:type="simple"/></inline-formula>. Notice that SS<sup>T</sup> is n &#215; n and is some-</p><p>what the linear Gram matrix of the projected X. RR<sup>T</sup> is p &#215; p and somewhat mirrors the covariance matrix in input space. We hope to explore these auxiliary constraints in the near future, again using the Cambridge ORL database and perhaps the Facial Recognition Technology (FERET) database as well.</p></sec><sec id="s5"><title>Acknowledgements</title><p>Ernest Fokou&#233; wishes to express his heartfelt gratitude and infinite thanks to our lady of perpetual help for her ever-present support and guidance, especially for the uninterrupted flow of inspiration received through her most powerful intercession.</p></sec><sec id="s6"><title>Cite this paper</title><p>Matthew A.Corsetti,ErnestFokou&#233;, (2015) Nonnegative Matrix Factorization with Zellner Penalty. Open Journal of Statistics,05,777-786. doi: 10.4236/ojs.2015.57077</p></sec></body><back><ref-list><title>References</title><ref id="scirp.62191-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">Lee, D.D. and Seung, H.S. (1999) Learning the Parts of Objects by Nonnegative matrix Factorization. Nature, 401, 788-791. &lt;/br&gt;http://dx.doi.org/10.1038/44565</mixed-citation></ref><ref id="scirp.62191-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">Li, S.Z., Hou, X., Zhang, H. and Cheng, Q. (2001) Learning Spatially Localized Parts-Based Representations. IEEE Conference on Computer Vision and Pattern Recognition, Kauai, 8-14 December 2001, 207-212. &lt;/br&gt;http://dx.doi.org/10.1109/cvpr.2001.990477</mixed-citation></ref><ref id="scirp.62191-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">Wang, Y., Jia, Y., Hu, C. and Turk, M. (2004) Fisher Nonnegative Matrix Factorization for Learning Local Features. Asian Conference of Computer Vision, Jeju, January 2004, 27-30.</mixed-citation></ref><ref id="scirp.62191-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">Wang, Y., Jia, Y., Hu, C. and Turk, M. (2005) Nonnegative Matrix Factorization Framework for Face Recognition. International Journal of Pattern Recognition and Artificial Intelligence, 19, 495-511.&lt;/br&gt;http://dx.doi.org/10.1142/S0218001405004198</mixed-citation></ref><ref id="scirp.62191-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">Huang, X., Zhao, J., Ash, J. and Lai, W. (2013) Clustering Student Discussion Messages on Online Forum by Visualization and Nonnegative matrix Factorization. Journal of Software Engineering and Applications, 6, 7-12. &lt;/br&gt;http://dx.doi.org/10.4236/jsea.2013.67B002</mixed-citation></ref><ref id="scirp.62191-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">Guillamet, D., Bressan, M. and Vitria, J. (2001) A Weighted Nonnegative Matrix Factorization for Local Representations. IEEE Conference of Computer Vision and Pattern Recognition, Kauai, 8-14 December 2001, 942-947.</mixed-citation></ref><ref id="scirp.62191-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">Cichocki, A., Zdunek, R. and Amari, S. (2006) Csiszár’s Divergence for Nonnegative Matrix Factorization: Family of New Algorithms. Lecture Notes in Computer Science, 3889, 32-39. &lt;/br&gt;http://dx.doi.org/10.1007/11679363_5</mixed-citation></ref><ref id="scirp.62191-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">Pauca, V.P., Piper, J. and Plemmons, R.J. (2006) Nonnegative Matrix Factorization for Spectral Data Analysis. Linear Algebra and its Applications, 416, 29-47. &lt;/br&gt;http://dx.doi.org/10.1016/j.laa.2005.06.025</mixed-citation></ref><ref id="scirp.62191-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">Hamza, A.B. and Brady, D.J. (2006) Reconstruction of Reflectance Spectra Using Robust Nonnegative Matrix Factorization. IEEE Transactions on Signal Processing, 54, 3637-3642. &lt;/br&gt;http://dx.doi.org/10.1109/TSP.2006.879282</mixed-citation></ref><ref id="scirp.62191-ref10"><label>10</label><mixed-citation publication-type="other" xlink:type="simple">Lin, C.J. (2007) Projected Gradient Methods for Nonnegative Matrix Factorization. Neural Computation, 19, 2756-2779. http://dx.doi.org/10.1162/neco.2007.19.10.2756</mixed-citation></ref><ref id="scirp.62191-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">Gonzales, E.F. and Zhang, Y. (2005) Accelerating the Lee-Seung Algorithm for Nonnegative Matrix Factorization. Technical Report, Department of Computational and Applied Mathematics, Rice University, Houston.</mixed-citation></ref><ref id="scirp.62191-ref12"><label>12</label><mixed-citation publication-type="other" xlink:type="simple">Zdunek, R. and Cichocki, A. (2006) Nonnegative Matrix Factorization with Quasi-Newton Optimization. Lecture Notes in Computer Science, 4029, 870-879. &lt;/br&gt;http://dx.doi.org/10.1007/11785231_91</mixed-citation></ref><ref id="scirp.62191-ref13"><label>13</label><mixed-citation publication-type="other" xlink:type="simple">Wild, S., Curry, J. and Dougherty, A. (2004) Improving Nonnegative Matrix Factorization through Structured Initialization. Pattern Recognition, 37, 2217-2232. &lt;/br&gt;http://dx.doi.org/10.1016/j.patcog.2004.02.013</mixed-citation></ref><ref id="scirp.62191-ref14"><label>14</label><mixed-citation publication-type="other" xlink:type="simple">Wild, S. (2002) Seeding Nonnegative Matrix Factorization with the Spherical K-Means Clustering. Master’s Thesis, University of Colorado, Colorado.</mixed-citation></ref><ref id="scirp.62191-ref15"><label>15</label><mixed-citation publication-type="other" xlink:type="simple">Wild, S., Curry, J. and Dougherty, A. (2003) Motivating Nonnegative Matrix Factorizations. Proceedings of the 8th SIAM Conference on Applied Linear Algebra, Williamsburg, 15-19 July 2003.&lt;/br&gt;http://www.siam.org/meetings/la03/proceedings/</mixed-citation></ref><ref id="scirp.62191-ref16"><label>16</label><mixed-citation publication-type="other" xlink:type="simple">Piper, J., Pauca, J.P., Plemmons, R.J. and Giffin, M. (2004) Object Characterization from Spectral Data Using Nonnegative Factorization and Information Theory. Proceedings of the 2004 AMOS Technical Conference, Maui, 9-12 September 2004.</mixed-citation></ref><ref id="scirp.62191-ref17"><label>17</label><mixed-citation publication-type="other" xlink:type="simple">Hoyer, P.O. (2002) Non-Negative Sparse Coding. Proceedings of the 12th IEEE Workshop on Neural Networks for Signal Processing, Martigny, Switzerland, 4-6 September 2002, 557-565. &lt;/br&gt;http://dx.doi.org/10.1109/nnsp.2002.1030067</mixed-citation></ref><ref id="scirp.62191-ref18"><label>18</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Hoyer</surname><given-names> P.O. </given-names></name>,<etal>et al</etal>. (<year>2004</year>)<article-title>Nonnegative Matrix Factorization with Sparseness Constraints</article-title><source> Journal of Machine Learning Research</source><volume> 5</volume>,<fpage> 1457</fpage>-<lpage>1469</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.62191-ref19"><label>19</label><mixed-citation publication-type="other" xlink:type="simple">Liu, W., Zheng, N. and Lu, X. (2003) Nonnegative Matrix Factorization for Visual Coding. 2003 IEEE International Conference on Acoustics, Speech and Signal Processing, 6, 293-296.</mixed-citation></ref><ref id="scirp.62191-ref20"><label>20</label><mixed-citation publication-type="book" xlink:type="simple">Zellner, A. (1986) On Assessing Prior Distributions and Bayesian Regression Analysis with g-Prior Distributions. In: Goel, P. and Zellner, A., Eds., Bayesian Inference and Decision Techniques: Essays in Honor of Bruno de Finetti, Elsevier Science Publishers, Inc., New York, 233-243.</mixed-citation></ref><ref id="scirp.62191-ref21"><label>21</label><mixed-citation publication-type="other" xlink:type="simple">Foster, D.P. and George, E.I. (1994) The Risk Inflation Criterion for Multiple Regression. Annals of Statistics, 22, 1947-1975.&lt;/br&gt; http://dx.doi.org/10.1214/aos/1176325766</mixed-citation></ref></ref-list></back></article>