<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">JCC</journal-id><journal-title-group><journal-title>Journal of Computer and Communications</journal-title></journal-title-group><issn pub-type="epub">2327-5219</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/jcc.2015.311031</article-id><article-id pub-id-type="publisher-id">JCC-61543</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Computer Science&amp;Communications</subject></subj-group></article-categories><title-group><article-title>
 
 
  A Bayesian Approach to Identify Photos Likely to Be More Popular in Social Media
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Arunabha</surname><given-names>Choudhury</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Sriram</surname><given-names>Nagaswamy</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib></contrib-group><aff id="aff1"><addr-line>Samsung R and D, Bangalore, India</addr-line></aff><pub-date pub-type="epub"><day>19</day><month>11</month><year>2015</year></pub-date><volume>03</volume><issue>11</issue><fpage>198</fpage><lpage>204</lpage><history><date date-type="received"><day>November</day>	<month>2015</month>	</date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
   With cameras becoming ubiquitous in Smartphones, it has become a very common trend to capture and share moments with friends and family in social media. Arguably, the 2 most relevant factors that contribute to the popularity are: the user’s social aspect and the content of the image (image quality, objects in the image etc.). In recent years, due to various security concerns, it has been increasingly difficult to derive social attributes from social media. Due to this limitation, in this paper we study what make images popular in social media based on the image content alone. We use Bayesian learning approach with variable likelihood function in order to predict image popularity. Our finding shows that a mapping between image content to image popularity can be achieved with a significant recall and precision. We then use our model to predict images that are likely to be more popular from a set of user images which eventually facilitate easy share. 
 
</p></abstract><kwd-group><kwd>Bayesian</kwd><kwd> Supervised Learning</kwd><kwd> Image Popularity</kwd><kwd> Classification</kwd><kwd> Data Mining</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>Image popularity for a user in social media can be attributed to variety of factors of which both image content as well as the social aspect plays significant role. With the explosion in number of photos clicked using Smartphones, it is important that the user can quickly search and find photos he may be interested in. According to report from KPCB analyst Mary Meeker in 2014, we now upload and share over 1.8 billion photos each day through various channels and social media. Since the numbers of people who have access to these images are also very large, getting one’s image viewed or liked by many people gives a feeling of being popular in one’s social circle and in turn an instant gratification. This makes image popularity an important driving factor for choosing and sharing the right image from a large collection of images. In this research we investigate what images are likely to be more popular so that a recommendation channel can be generated to allow the user to quickly access these photos.</p><p>Even though popularity of text (such as tweets) [<xref ref-type="bibr" rid="scirp.61543-ref1">1</xref>] [<xref ref-type="bibr" rid="scirp.61543-ref2">2</xref>] and videos (such as YouTube videos) [<xref ref-type="bibr" rid="scirp.61543-ref3">3</xref>]-[<xref ref-type="bibr" rid="scirp.61543-ref5">5</xref>] have been studied in recent years, image popularity prediction still remains a difficult problem. In our study, we have found that the most relevant research work in image popularity prediction [<xref ref-type="bibr" rid="scirp.61543-ref6">6</xref>] has been done using the Flickr dataset. This research is however a combination of both social aspects and image content. Also, the definition of popularity is number of views in Flickr. As opposed to these ideas, in our work, we address two issues in particular: First, we investigate whether there is a more relevant definition of image popularity and answer this question affirmatively. Second, we show that it is possible to predict image popularity with a significant accuracy based on image content alone. Although our definition of popularity is based on number of likes in Facebook, we have only used this information to label the training set. To the best of our knowledge, this work is the first to predict image popularity based on image content alone.</p></sec><sec id="s2"><title>2. Related Work</title><p>As already discussed, the most significant work on image popularity prediction has been done by Khosla et al. [<xref ref-type="bibr" rid="scirp.61543-ref6">6</xref>]. For other related work on popularity, there has been some work by Figueiredo et al. in video popularity prediction [<xref ref-type="bibr" rid="scirp.61543-ref7">7</xref>] [<xref ref-type="bibr" rid="scirp.61543-ref8">8</xref>]. However, their work is mostly dependent on social aspect like comments and tags associated with videos. Popularity prediction has been studied in other areas like popularity of online articles [<xref ref-type="bibr" rid="scirp.61543-ref9">9</xref>] and popularity of social marketing messages [<xref ref-type="bibr" rid="scirp.61543-ref10">10</xref>]. However, the study done in these papers comes under the category of predicting text popularity such as tweets or Facebook posts and predicting popularity of web pages. None of these in any way relate to the image popularity we discuss in this work. Xinran et al. in [<xref ref-type="bibr" rid="scirp.61543-ref11">11</xref>] have studied the prediction of clicks on Ads in Facebook as a classification problem. Although number of clicks associated with an ad can be considered as a form of popularity, click prediction of an Ad is not directly related to popularity.</p><p>In another interesting work by Justin Cheng et al. in [<xref ref-type="bibr" rid="scirp.61543-ref12">12</xref>] they study the prediction of cascading in social media such as Twitter and Facebook. Cascading refers to sharing and re-sharing posts in Facebook and re-tweeting tweets in Twitter. These sharing may include anything from text to images that is sharable through these channels. The growth of this cascade can also in some ways be considered as popularity of the sharable items. This however does not directly relate to our consideration of image popularity for individual users in social media particularly for two reasons. First, most of these sharable items consist of items that are strictly public and not private images from users. Second, along with content, there is a strict dependency on other social features that cannot easily be avoided. Although a separate study of popularity prediction for these sharable items is possible, in our context, due to their varied properties we have decided not to mix them up.</p></sec><sec id="s3"><title>3. Regression vs. Classification</title><p>The definition of popularity by Khosla et al. in [<xref ref-type="bibr" rid="scirp.61543-ref6">6</xref>] is number of views in Flickr. In their work they have considered the normalized view count as the dependent variable for regression and the performance measure is in terms of rank correlation. <xref ref-type="table" rid="table1">Table 1</xref> summarizes their results.</p><p>From <xref ref-type="table" rid="table1">Table 1</xref>, one can see that the contribution to prediction by image content alone is very low and most of the contribution is from social aspect. Another issue is the number of features used in [<xref ref-type="bibr" rid="scirp.61543-ref6">6</xref>] is very large for image content alone and yet contribution to prediction from all these features is very little. <xref ref-type="table" rid="table2">Table 2</xref> summarizes this issue.</p><table-wrap id="table1" ><label><xref ref-type="table" rid="table1">Table 1</xref></label><caption><title> Rank correlation for image content and social attributes in [<xref ref-type="bibr" rid="scirp.61543-ref6">6</xref>]</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Dataset Type</th><th align="center" valign="middle" >Image Content Only</th><th align="center" valign="middle" >Social Attributes Only</th><th align="center" valign="middle" >Content + Social</th></tr></thead><tr><td align="center" valign="middle" >One-per-user</td><td align="center" valign="middle" >0.31</td><td align="center" valign="middle" >0.77</td><td align="center" valign="middle" >0.81</td></tr><tr><td align="center" valign="middle" >User-mix</td><td align="center" valign="middle" >0.36</td><td align="center" valign="middle" >0.66</td><td align="center" valign="middle" >0.72</td></tr><tr><td align="center" valign="middle" >User-specific</td><td align="center" valign="middle" >0.40</td><td align="center" valign="middle" >0.21</td><td align="center" valign="middle" >0.48</td></tr></tbody></table></table-wrap><table-wrap id="table2" ><label><xref ref-type="table" rid="table2">Table 2</xref></label><caption><title> Rank correlation for individual feature type and combined [<xref ref-type="bibr" rid="scirp.61543-ref6">6</xref>]</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Dataset Type</th><th align="center" valign="middle" >Gist</th><th align="center" valign="middle" >Color Histogram</th><th align="center" valign="middle" >Texture</th><th align="center" valign="middle" >Color Patches</th><th align="center" valign="middle" >Gradient</th><th align="center" valign="middle" >Deep Learning</th><th align="center" valign="middle" >Objects</th><th align="center" valign="middle" >Combined</th></tr></thead><tr><td align="center" valign="middle" >One-per-user</td><td align="center" valign="middle" >0.07</td><td align="center" valign="middle" >0.12</td><td align="center" valign="middle" >0.20</td><td align="center" valign="middle" >0.23</td><td align="center" valign="middle" >0.26</td><td align="center" valign="middle" >0.28</td><td align="center" valign="middle" >0.23</td><td align="center" valign="middle" >0.31</td></tr><tr><td align="center" valign="middle" >User-mix</td><td align="center" valign="middle" >0.13</td><td align="center" valign="middle" >0.15</td><td align="center" valign="middle" >0.22</td><td align="center" valign="middle" >0.29</td><td align="center" valign="middle" >0.32</td><td align="center" valign="middle" >0.33</td><td align="center" valign="middle" >0.30</td><td align="center" valign="middle" >0.36</td></tr><tr><td align="center" valign="middle" >User-specific</td><td align="center" valign="middle" >0.16</td><td align="center" valign="middle" >0.23</td><td align="center" valign="middle" >0.32</td><td align="center" valign="middle" >0.36</td><td align="center" valign="middle" >0.34</td><td align="center" valign="middle" >0.26</td><td align="center" valign="middle" >0.33</td><td align="center" valign="middle" >0.40</td></tr></tbody></table></table-wrap><p>Each of these categories in image content (like Gist, Texture etc.) consists of a number of features ranging from 512 to 10,752. Computing all these features from a collection of images may not be a feasible scenario in devices such as Smartphones where one is limited in terms of memory and processing capabilities. Also in [<xref ref-type="bibr" rid="scirp.61543-ref6">6</xref>] the final prediction is number of views an image is likely to receive as a result of prediction but the conclusion whether or not the image is popular is a matter of subjective judgment (one may decide popularity based on number of predicted view count). In our work we look at this problem in slightly different way.</p><sec id="s3_1"><title>3.1. Definition of Image Popularity</title><p>As already mentioned in [<xref ref-type="bibr" rid="scirp.61543-ref6">6</xref>] there are various ways to define popularity. In our work we use the number of likes in Facebook as a measure of popularity. Since “Like” in Facebook refers to users consciously making a decision of whether or not they “Like” an image, in our consideration this definition is slightly better than number of views, which may be noisier in certain situation. Nonetheless, the definition of true popularity may still be argued and remains a difficult question to answer.</p><p>For our purpose, the raw like count cannot directly be used as measure of popularity. The reason being, in Facebook only people friends with the users are allowed to like photos. Thus, people with low friend count may have images which are popular in their circle but may not match up to the like count for users with many friends. For this reason we have normalized the like count per user basis. <xref ref-type="fig" rid="fig1">Figure 1</xref> shows the density plot for number of likes and number of likes normalized.</p><p>As expected, we see from <xref ref-type="fig" rid="fig1">Figure 1</xref> that the median normalized like count is about 0.1 and about a quarter of images have normalized like count more than 0.4. One can at this point consider 0.4 as the threshold to define popularity of an image. This definition of popularity however may still be improved. The third quartile threshold based on the entire collection is still subjected to the fact that popularity of an image is across multiple users. As already discussed, each user in Facebook may have a local sense of popularity for which the above definition may still affect images when the variance in like count is high between users with high and low friend count. In order to overcome this issue and to consider local popularity as the measure of popularity, we first find the mean of normalized like count for each user. Then, for each user, if an image has normalized like count more than the mean, we consider that image to be popular. <xref ref-type="fig" rid="fig2">Figure 2</xref> gives the boxplot for normalized like count for random users chosen from the dataset.</p><p>From <xref ref-type="fig" rid="fig2">Figure 2</xref> we can see that in many cases the distribution of normalized like count for users are skewed. Here one may argue median to be a better estimate. But in our observation we have seen that for most cases the distribution is right skewed. Hence mean gives a tighter bound than median. This is shown in <xref ref-type="fig" rid="fig3">Figure 3</xref>. As we can see in most cases mean is either comparable to median or more than median. Hence, for our purpose we consider mean normalized like count for each user to be our threshold for popularity.</p><fig-group id="fig1"><label><xref ref-type="fig" rid="fig1">Figure 1</xref></label><caption><title>Density plot.</title></caption><fig id ="fig1_1"><label></label><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/61543x4.png"/></fig><fig id ="fig1_2"><label></label><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/61543x5.png"/></fig></fig-group><fig id="fig2"  position="float"><label><xref ref-type="fig" rid="fig2">Figure 2</xref></label><caption><title> Boxplot for normalized like count (22 random users)</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/61543x6.png"/></fig><fig id="fig3"  position="float"><label><xref ref-type="fig" rid="fig3">Figure 3</xref></label><caption><title> Mean vs median for normalized like count (22 random users)</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/61543x7.png"/></fig></sec><sec id="s3_2"><title>3.2. Dataset</title><p>In this section we introduce the dataset and the features we have used for our experiment. Firstly, our data set comprises of user images from Facebook with number of likes for each image as the measure of popularity. For our analysis, we have collected about 1000 images from 35 different Facebook users. We normalized the like count for each user. We then convert the like count to a binary label 0 and 1 with 0 being the like count less than the mean normalized like count and 1 being like count more than the mean normalized like count for each user. Secondly, our final feature set after feature selection consists of eight low level computer vision features called blur, motion blur, brightness, overexposure, underexposure, contrast, feature number, out of focus background and one high level image feature whose values have been derived from face size and number of faces in the image. The first eight features are more relevant to the quality of the image whereas the last one correspond to objects in the image; in this case faces. Our assumption is that quality of image is a very important factor for images being popular but not the only factor. Presence of faces in the image in social media makes a big difference and hence, we have also considered face size and number of faces in the image.</p></sec><sec id="s3_3"><title>3.3. Regression</title><p>We first use the normalized like count as it is to predict the number of likes for each image. <xref ref-type="table" rid="table3">Table 3</xref> summarizes the performance in terms of rank correlation for SVR, Linear regression and ANN.</p><p>As we can see, the rank correlation using only nine features are comparable to the rank correlation in <xref ref-type="table" rid="table2">Table 2</xref>. Therefore, using image content alone may not be sufficient for this regression problem. Thus, for our purpose, we convert this problem to a classification problem by converting the normalized like count to popularity. The conversion is according to the definition of popularity introduced in Section 3.2.</p></sec><sec id="s3_4"><title>3.4. Classification</title><p>We then try to formalize this problem as a binary classification problem with,</p><p><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/61543x8.png" xlink:type="simple"/></inline-formula>the image data and <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/61543x9.png" xlink:type="simple"/></inline-formula> the <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/61543x10.png" xlink:type="simple"/></inline-formula> feature vector</p><p><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/61543x11.png" xlink:type="simple"/></inline-formula>the set of binary labels</p><p><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/61543x12.png" xlink:type="simple"/></inline-formula>a given distribution over <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/61543x13.png" xlink:type="simple"/></inline-formula></p><p>The goal is to find a function f such that minimizes</p><disp-formula id="scirp.61543-formula186"><label>(1)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/61543x14.png"  xlink:type="simple"/></disp-formula><p>where, <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/61543x15.png" xlink:type="simple"/></inline-formula>is an error function [<xref ref-type="bibr" rid="scirp.61543-ref13">13</xref>] that calculates error between original class label y and predicted labels<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/61543x16.png" xlink:type="simple"/></inline-formula>. We then train six different supervised learning algorithms. Test results are given in terms of precision, recall and accuracy and are averaged over twenty trials. These results are plotted in <xref ref-type="fig" rid="fig4">Figure 4</xref>. Here we can see that the classification accuracy is more or less uniform across all classifiers. Hence, in order to differentiate their performance we also plot precision and recall for all classifiers in <xref ref-type="fig" rid="fig4">Figure 4</xref>.</p><p>Since in our case it is more appropriate to weight both precision and recall equally, we consider f-measure as our performance matrix. In <xref ref-type="fig" rid="fig5">Figure 5</xref> we plot the f-measure for all the classifiers for eight random trials. From the plots we observe that performance of Na&#239;ve Bayes is relatively more consistent. Even though ANN does better in some cases the performance is not very consistent. For the purpose of consistency and overall performance measure we have considered Na&#239;ve Bayes as our base model.</p><fig id="fig4"  position="float"><label><xref ref-type="fig" rid="fig4">Figure 4</xref></label><caption><title> Performance plot for different classifiers</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/61543x17.png"/></fig><fig id="fig5"  position="float"><label><xref ref-type="fig" rid="fig5">Figure 5</xref></label><caption><title> Plot of f-measure</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/61543x18.png"/></fig><table-wrap id="table3" ><label><xref ref-type="table" rid="table3">Table 3</xref></label><caption><title> Rank correlation for different regression algorithms</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Regression model</th><th align="center" valign="middle" >Correlation</th></tr></thead><tr><td align="center" valign="middle" >SVR</td><td align="center" valign="middle" >0.22</td></tr><tr><td align="center" valign="middle" >Linear regression</td><td align="center" valign="middle" >0.20</td></tr><tr><td align="center" valign="middle" >ANN</td><td align="center" valign="middle" >0.23</td></tr></tbody></table></table-wrap></sec></sec><sec id="s4"><title>4. Experiment</title><p>For a given set of feature vector X and binary label set B Bayes theorem can be stated as:</p><disp-formula id="scirp.61543-formula187"><label>(2)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/61543x19.png"  xlink:type="simple"/></disp-formula><p>where, <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/61543x20.png" xlink:type="simple"/></inline-formula>is the conditional density function. In our Na&#239;ve Bayes implementation in the previous section, the assumption is that</p><disp-formula id="scirp.61543-formula188"><label>(3)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/61543x21.png"  xlink:type="simple"/></disp-formula><p>In order to verify this we plot the density function of one of the features Blur in <xref ref-type="fig" rid="fig6">Figure 6</xref>.</p><p>For a lot of features in our case the Normal distribution turns out to be a not so good fit. In our experiment we also observe that for certain features a Gamma or a Weibull distribution fits better than Normal distribution. <xref ref-type="table" rid="table4">Table 4</xref> summarizes this observation.</p><p>We measure goodness of fit in terms of p-value greater than 0.05. From the summary in <xref ref-type="table" rid="table4">Table 4</xref> we can see that some features even though fit well with Normal distribution, in other cases Weibull or Gamma distribution scores higher in terms of goodness of fit (<xref ref-type="fig" rid="fig6">Figure 6</xref> and <xref ref-type="fig" rid="fig7">Figure 7</xref>).</p><p>We finally use this variable density function based Na&#239;ve Bayes to predict popularity. <xref ref-type="fig" rid="fig8">Figure 8</xref> compares the precision and recall of Na&#239;ve Bayes with normal distribution (NBND) vs. Na&#239;ve Bayes with variable density function (NBVDF). With this modification we have been able to achieve a 52% precision and 80% recall.</p><fig-group id="fig6"><label><xref ref-type="fig" rid="fig6">Figure 6</xref></label><caption><title> Blur density plot with Normal and Weibull distribution respectively (red).</title></caption><fig id ="fig6_1"><label></label><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/61543x22.png"/></fig><fig id ="fig6_2"><label></label><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/61543x23.png"/></fig></fig-group><fig id="fig7"  position="float"><label><xref ref-type="fig" rid="fig7">Figure 7</xref></label><caption><title> Contrast density plot with Gamma distribution (red)</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/61543x24.png"/></fig><fig id="fig8"  position="float"><label><xref ref-type="fig" rid="fig8">Figure 8</xref></label><caption><title> Na&#239;ve Bayes (normal vs. variable distribution)</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/61543x25.png"/></fig><table-wrap id="table4" ><label><xref ref-type="table" rid="table4">Table 4</xref></label><caption><title> Goodness of fit (p-values)</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Features</th><th align="center" valign="middle" >Normal</th><th align="center" valign="middle" >Gamma</th><th align="center" valign="middle" >Weibull</th></tr></thead><tr><td align="center" valign="middle" >Brightness</td><td align="center" valign="middle" >0.641</td><td align="center" valign="middle" >0.003</td><td align="center" valign="middle" >0.60</td></tr><tr><td align="center" valign="middle" >Blur</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >0.175</td><td align="center" valign="middle" >0.455</td></tr><tr><td align="center" valign="middle" >Motion Blur</td><td align="center" valign="middle" >0.004</td><td align="center" valign="middle" >0.036</td><td align="center" valign="middle" >0.703</td></tr><tr><td align="center" valign="middle" >Over Exposure</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >0.228</td></tr><tr><td align="center" valign="middle" >Contrast</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >0.047</td><td align="center" valign="middle" >0.034</td></tr></tbody></table></table-wrap></sec><sec id="s5"><title>5. Conclusion</title><p>A regression based popularity prediction has constrains such as: low accuracy when only image content is used, system overhead due to large number of features and highly dependent on social data which may not always be available. In most cases one is concerned with the final decision on whether or not an image is going to be popular and not be concerned about the number of views or likes. Thus solving the classification problem is much more intuitive and a simple Na&#239;ve Bayes variant works well without much system overhead. In future we plan to add contextual information to improve the overall precision and recall. In our work we have been able to show that one may choose to ignore any social aspect associated with the image and still achieve significant precision and recall to predict popularity of an image purely based on image features.</p></sec><sec id="s6"><title>Cite this paper</title><p>Arunabha Choudhury,Sriram Nagaswamy, (2015) A Bayesian Approach to Identify Photos Likely to Be More Popular in Social Media. Journal of Computer and Communications,03,198-204. doi: 10.4236/jcc.2015.311031</p></sec></body><back><ref-list><title>References</title><ref id="scirp.61543-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">Petrovic, S., Osborne, M. and Lavrenko, V. (2011) Rt to Win! Predicting Message Propagation in Twitter. ICWSM.</mixed-citation></ref><ref id="scirp.61543-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">Hong, L., Dan, O. and Davison, B.D. (2011) Predicting Popular Messages in Twitter. WWW (Companion Volume), 57-58. http://dx.doi.org/10.1145/1963192.1963222</mixed-citation></ref><ref id="scirp.61543-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">Pinto, H., Almeida, J.M. and Goncalves, M.A. (2013) Using Early View Patterns to Predict the Popularity of Youtube Videos. WSDM, 365-374.</mixed-citation></ref><ref id="scirp.61543-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">Shamma, D.A., Yew, J., Kennedy, L. and Churchill, E.F. (2011) Viral Actions: Predicting Video View Counts Using Synchronous Sharing Behaviors. ICWSM.</mixed-citation></ref><ref id="scirp.61543-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">Nwana, A.O., Avestimehr, S. and Chen, T. (2013) A Latent Social Approach to Youtube Popularity Prediction. CoRR.</mixed-citation></ref><ref id="scirp.61543-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">Khosla, A., Sarma, A.D. and Hamid, R. (2014) What Makes an Image Popular? IW3C2.</mixed-citation></ref><ref id="scirp.61543-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">Figueiredo, F. (2013) On the Prediction of Popularity of Trends and Hits for User Generated Videos. Proceedings of the Sixth ACM International Conference on Web Search and Data Mining, 741-746.  
http://dx.doi.org/10.1145/2433396.2433489</mixed-citation></ref><ref id="scirp.61543-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">Figueiredo, F., Benevenuto, F. and Almeida, J.M. (2011) The Tube over Time: Characterizing Popularity Growth of YouTube Videos. Proceedings of the Fourth ACM International Con-ference on Web Search and Data Mining, 745- 754. http://dx.doi.org/10.1145/1935826.1935925</mixed-citation></ref><ref id="scirp.61543-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">Vanwinckelen, G. and Meert, W. (2014) Predicting the Popularity of Online Articles with Random Forests. ECML/ PKDD Discovery Challenge on Predictive Web Analytics, Nancy, September 2014.</mixed-citation></ref><ref id="scirp.61543-ref10"><label>10</label><mixed-citation publication-type="book" xlink:type="simple">Yu, B., Chen, M. and Kwok, L. (2011) Toward Predicting Popularity of Social Marketing Messages. Salerno, J., et al., Eds., SBP 2011, LNCS 6589, 317-324. http://dx.doi.org/10.1007/978-3-642-19656-0_44</mixed-citation></ref><ref id="scirp.61543-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">He, X., et al. (2014) Practical Lessons from Predicting Clicks on Ads at Facebook. ADKDD’14, 24-27 August 2014. 
http://dx.doi.org/10.1145/2648584.2648589</mixed-citation></ref><ref id="scirp.61543-ref12"><label>12</label><mixed-citation publication-type="other" xlink:type="simple">Cheng, J., Adamic, L.A., Dow, P.A., Kleinberg, J. and Leskovec, J. Can Cascades Be Predicted? WWW’14, Seoul, Republic of Korea.</mixed-citation></ref><ref id="scirp.61543-ref13"><label>13</label><mixed-citation publication-type="other" xlink:type="simple">Daume III, H. A Course in Machine Learning. Chapter 5.1, 69.</mixed-citation></ref></ref-list></back></article>