<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">JSEA</journal-id><journal-title-group><journal-title>Journal of Software Engineering and Applications</journal-title></journal-title-group><issn pub-type="epub">1945-3116</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/jsea.2016.94009</article-id><article-id pub-id-type="publisher-id">JSEA-65570</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Computer Science&amp;Communications</subject></subj-group></article-categories><title-group><article-title>
 
 
  Hand Gesture Recognition Using Appearance Features Based on 3D Point Cloud
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>anwen</surname><given-names>Chong</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Jianfeng</surname><given-names>Huang</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Shaoming</surname><given-names>Pan</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib></contrib-group><aff id="aff1"><addr-line>State Key Laboratory for Information Engineering in Surveying, Mapping and Remote Sensing, Wuhan 
University, Wuhan, China</addr-line></aff><pub-date pub-type="epub"><day>18</day><month>04</month><year>2016</year></pub-date><volume>09</volume><issue>04</issue><fpage>103</fpage><lpage>111</lpage><history><date date-type="received"><day>6</day>	<month>March</month>	<year>2016</year></date><date date-type="rev-recd"><day>accepted</day>	<month>15</month>	<year>April</year>	</date><date date-type="accepted"><day>18</day>	<month>April</month>	<year>2016</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
  This paper presents a method for hand gesture recognition based on 3D point cloud. Digital image processing technology is used in this research. Based on the 3D point from depth camera, the system firstly extracts some raw data of the hand. After the data segmentation and preprocessing, three kinds of appearance features are extracted, including the number of stretched fingers, the angles between fingers and the gesture region’s area distribution feature. Based on these features, the system implements the identification of the gestures by using decision tree method. The results of experiment demonstrate that the proposed method is pretty efficient to recognize common gestures with a high accuracy.
 
</p></abstract><kwd-group><kwd>Human-Computer-Interaction</kwd><kwd> Gesture Recognition</kwd><kwd> 3D Point Cloud</kwd><kwd> Depth Image</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>There has been a great emphasis lately on Human-Computer-Interaction (HCI) research to create easy-to-use interfaces by directly employing natural communication and manipulation skills of humans [<xref ref-type="bibr" rid="scirp.65570-ref1">1</xref>] . As an important part of body, naturally, the hand is given more and more attention. Gesture recognition is a key aspect of Human-Computer-Interaction. There are countless researches focus on this advanced topic to create natural user interface and to improve user experiences by using simple and intuitive hand gestures for free-hand controller [<xref ref-type="bibr" rid="scirp.65570-ref2">2</xref>] . How to detect the hands, segment them from the background and recognize the gestures become great challenges. And various methods are proposed to solve those issues.</p><p>A hierarchical method of static hand gesture recognition that combines finger detection and histogram of oriented gradient (HOG) features is proposed in [<xref ref-type="bibr" rid="scirp.65570-ref3">3</xref>] . An algorithm based on the spatial pyramid bag of features is proposed to describe the hand image in [<xref ref-type="bibr" rid="scirp.65570-ref4">4</xref>] . But both of them are based on RGB image, which means it’s difficult to distinguish the hand from complex background. What’s more, the intensity of light seriously affects the recognition results. In order to avoid these drawbacks, many scholars select depth image in the research of gesture recognition. Depth information has long been regarded as an essential part of successful gesture recognition [<xref ref-type="bibr" rid="scirp.65570-ref5">5</xref>] . Many researches [<xref ref-type="bibr" rid="scirp.65570-ref6">6</xref>] - [<xref ref-type="bibr" rid="scirp.65570-ref8">8</xref>] extract different features from the depth data, then various classifiers are employed for gesture recognition. These methods all get good effect, but they have to collect a large number of training samples. Based on 3D point cloud data, this paper uses geometric method for extracting the appearance features of gestures and classifying the given gestures. This method can obtain high gesture recognition accuracy without training sample. Compared with the previous methods, it’s more succinct and efficient.</p></sec><sec id="s2"><title>2. Proposed Gesture Recognition System</title><p>The proposed gesture recognition system (shown in <xref ref-type="fig" rid="fig1">Figure 1</xref>) composed of three parts. In the first part, the 3D point cloud data of the hand region is gotten from depth camera (SwissRanger 4000 depth camera), then after threshold segmentation and gray transformation the 3D point cloud becomes a binary image. Meanwhile, some preprocessing on the grayscale image is necessary. In the second part, some apparent features are extracted. Finally, on the basis of the features extracted in last step the gesture can be recognized.</p><fig id="fig1"  position="float"><label><xref ref-type="fig" rid="fig1">Figure 1</xref></label><caption><title> Architecture of the proposed gesture recognition system</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/1-9302195x6.png"/></fig></sec><sec id="s3"><title>3. Hand Segmentation</title><sec id="s3_1"><title>3.1. Image Collection</title><p>In this paper, the 3D point cloud data of gesture are collected from the SwissRanger 4000 (SR4000) depth camera. The SR4000 cameras are optical imaging systems which provide real time distance data at video frame rates. Based on the Time-of-Flight (ToF) principle, the cameras employ an integrated light source. The emitted light is reflected by objects in the scene and travels back to the camera, where the precise time of arrival is measured independently by each pixel of the image sensor, producing a per-pixel distance measurement. Finally, we can get the three-dimensional coordinates of each point from the camera. A typical 3D point cloud of gesture scene just like shown in <xref ref-type="fig" rid="fig2">Figure 2</xref>. In this image, the origin of the coordinate system (0, 0, 0) is at the intersection of the optical axis with the front face of the camera, and <xref ref-type="fig" rid="fig3">Figure 3</xref> shows the camera’s output coordinate system.</p><p>From <xref ref-type="fig" rid="fig2">Figure 2</xref> we can notice that the 3D point cloud contains not only hand region but also other region. Obviously, we need to extract the hand region G. G is given by (1):</p><disp-formula id="scirp.65570-formula69"><label>(1)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/1-9302195x7.png"  xlink:type="simple"/></disp-formula><p>In (1), Min<sub>x</sub>, Max<sub>x</sub>, Min<sub>v</sub>, Max<sub>v</sub>, Min<sub>z</sub> and Max<sub>z</sub> are the thresholds of three coordinate directions. We set Min<sub>x</sub> = −150 mm, Max<sub>x</sub> = 150 mm, Min<sub>v</sub> = −150 mm, Max<sub>v</sub> = 150 mm, Min<sub>z</sub> = 200 mm, Max<sub>z</sub> = 500 mm to ensure that the whole hand region is extracted. Next, G is transformed into a binary image.</p></sec><sec id="s3_2"><title>3.2. Image Preprocessing</title><p>Due to the nature of the depth sensor, the hand region on the depth map may be have holes and cracks [<xref ref-type="bibr" rid="scirp.65570-ref9">9</xref>] , which will seriously affect the accuracy of hand gesture. Usually the binary image always has some noisy. So image preprocessing is necessary, which contains filling the holes and image denoising. In other papers [<xref ref-type="bibr" rid="scirp.65570-ref10">10</xref>] - [<xref ref-type="bibr" rid="scirp.65570-ref12">12</xref>] , some inpainting and filtering methods reach a good result. However, the methods are always so complex. We just employ some simple morphological operations (erosion and dilation) in our preprocessing.</p><fig id="fig2"  position="float"><label><xref ref-type="fig" rid="fig2">Figure 2</xref></label><caption><title> 3D point cloud of gesture scene</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/1-9302195x8.png"/></fig><fig id="fig3"  position="float"><label><xref ref-type="fig" rid="fig3">Figure 3</xref></label><caption><title> (x, y, z) as delivered by the camera is given in this coordinate syste</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/1-9302195x9.png"/></fig></sec></sec><sec id="s4"><title>4. Extraction of Features</title><p>Appearance features are important in gesture recognition. Compared to other methods of feature extraction, appearance features are more intuitional and efficient. As appearance features, the number of stretched fingers and the angles between fingers are used for gesture recognition in [<xref ref-type="bibr" rid="scirp.65570-ref13">13</xref>] . In this paper, we also choose the number of stretched fingers and the angles between fingers as features. Besides, the gesture region’s area distribution feature is chosen as an appearance feature, too.</p><sec id="s4_1"><title>4.1. Extraction of the Central Point</title><p>Through the erosion operations of mathematical morphology we can locate the central point C. As we all know, the palm is the primary part of gesture. Through continuous erosion operations, the boundary of gesture region is removed over and over again. And the gesture region get smaller and smaller. Eventually, only a point is left, which is just the central point C of the gesture region.</p></sec><sec id="s4_2"><title>4.2. Extraction of Appearance Features</title><p>The appearance features used in this paper contain the number of stretched fingers, the angles between fingers and the gesture region’s area distribution feature. The following is the main steps of extraction of appearance features.</p><p>1) Firstly, the maximum distance value D between the central point and the edge of the gesture region is cal-</p><p>culated. Then we <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/1-9302195x10.png" xlink:type="simple"/></inline-formula> which represent ten different lengths of radius. Next, choosing</p><p>C as the centers of rhombuses and r<sub>n</sub> as the radiuses, we can draw 10 rhombuses (the innermost one is recorded as the first rhombus and the outermost one is recorded as the tenth rhombus), as shown in <xref ref-type="fig" rid="fig4">Figure 4</xref> (In order to highlight the effect, the color has been transformed).</p><p>2) From <xref ref-type="fig" rid="fig4">Figure 4</xref> we can notice that every rhombus has a different number of intersections with gesture region. In order to get the number of the stretched fingers N, as a rule thumb, we choose the sixth rhombus to calculate. First of all, in a clockwise direction, we record the points on the sixth rhombus whose color from blue change into red or from red change into blue. We define K<sub>i</sub> as the i-th (i = 1, 2, 3・・・) point whose color from blue change into red, and T<sub>i</sub> as the i-th point whose color from red change into blue. Obviously, the number of K or T is just the number of the stretched fingers N.</p><p>3) We define M<sub>i</sub> as the midpoint of K<sub>i</sub> and T<sub>i</sub> (i = 1, 2, 3・・・), then each midpoint M<sub>i</sub> and the central point C can be connected into a line. And we can calculate each angle between adjacent lines. We use A<sub>j</sub> (j = 1, 2, 3・・・i-1) to represent these angles.</p><p>4) As a rule thumb, the fifth rhombus is chosen as boundary line, so the gesture region is divided into two parts. We define P<sub>1</sub> to denote the first part which is inside of the fifth rhombus and define P<sub>2</sub> to denote another part which is outside of the fifth rhombus. Then we calculate the ratio of P<sub>1</sub> to P<sub>2</sub>, and we use R to represent this ratio. Naturally, R can be used to describe the gesture region’s area distribution feature.</p><fig id="fig4"  position="float"><label><xref ref-type="fig" rid="fig4">Figure 4</xref></label><caption><title> Extraction of appearance features</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/1-9302195x11.png"/></fig></sec></sec><sec id="s5"><title>5. Gesture Recognition</title><p>Basing on the above appearance features, we construct the decision tree for gesture recognition. In this paper, 9 common gestures (shown in <xref ref-type="fig" rid="fig5">Figure 5</xref>) are employed for recognition and classification.</p><p>Decision tree is a kind of mathematical method to classify the new data by using the decision rules, and the decision rules are get from training samples. The key to construct a good decision tree is to choose the proper logical judgment and attributes.</p><p>In this paper, the number of stretched fingers N, the angles between fingers A<sub>j</sub> and the ratio of different gesture regions’ area R are chosen as branch node of decision tree. First of all, notice that Gesture-9 is the most special gesture, because it doesn’t have a stretched finger. So we can choose R as root note of the decision tree to distinguish Gesture-9 and other gestures in the first step. Through some training of samples, we can easily find that Gesture-9’s R is always greater than 0.8, and other gestures’ R always less than 0.8. So 0.8 is set as a threshold of R. Then, in the rest kinds of gestures, the number of stretched fingers N is an important feature. According to the value of N, Gesture-1, Gesture-4 and Gesture-5 can be uniquely identified. However, if N = 2, the gesture may be Gesture-2 or Gesture-6, if N = 3, the gesture may be Gesture-3 or Gesture-7 or Gesture-8, which is why we need to choose A<sub>j</sub> as another appearance feature to distinguish them. Finally, the decision tree we construct is shown as <xref ref-type="fig" rid="fig6">Figure 6</xref>.</p><fig id="fig5"  position="float"><label><xref ref-type="fig" rid="fig5">Figure 5</xref></label><caption><title> 9 common gestures</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/1-9302195x12.png"/></fig><fig id="fig6"  position="float"><label><xref ref-type="fig" rid="fig6">Figure 6</xref></label><caption><title> The decision tree</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/1-9302195x13.png"/></fig></sec><sec id="s6"><title>6. Experimental Results</title><p>In order to validate the method proposed in this paper, we connected the SR4000 depth camera with computer to do a lot of experiments. The experiments were conducted to identify the 9 common gestures. A total of 2700 test samples from 5 people were tested under three different conditions, including under the sunlight (strong light), indoors with the light off (weak light) and indoors with the light on (ordinary light). Each gesture have 300 test samples including different light conditions. The recognition results are shown in <xref ref-type="table" rid="table1">Table 1</xref>, <xref ref-type="table" rid="table2">Table 2</xref> and <xref ref-type="table" rid="table3">Table 3</xref>. In the following tables, the recognition accuracy of each gesture refer to the ratio of the number of correct recognition to the total number of recognition. And the mean accuracy refer to the mean of the recognition accuracy of every gesture.</p><table-wrap id="table1" ><label><xref ref-type="table" rid="table1">Table 1</xref></label><caption><title> The recognition result (strong light)</title></caption><table><tbody><thead><tr><th align="center" valign="middle" ></th><th align="center" valign="middle" >G-1</th><th align="center" valign="middle" >G-2</th><th align="center" valign="middle" >G-3</th><th align="center" valign="middle" >G-4</th><th align="center" valign="middle" >G-5</th><th align="center" valign="middle" >G-6</th><th align="center" valign="middle" >G-7</th><th align="center" valign="middle" >G-8</th><th align="center" valign="middle" >G-9</th></tr></thead><tr><td align="center" valign="middle" >G-1</td><td align="center" valign="middle" >93</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td></tr><tr><td align="center" valign="middle" >G-2</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >98</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td></tr><tr><td align="center" valign="middle" >G-3</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >91</td><td align="center" valign="middle" >1</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >5</td><td align="center" valign="middle" >6</td><td align="center" valign="middle" >0</td></tr><tr><td align="center" valign="middle" >G-4</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >92</td><td align="center" valign="middle" >2</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td></tr><tr><td align="center" valign="middle" >G-5</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >5</td><td align="center" valign="middle" >97</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td></tr><tr><td align="center" valign="middle" >G-6</td><td align="center" valign="middle" >1</td><td align="center" valign="middle" >2</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td></tr><tr><td align="center" valign="middle" >G-7</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >4</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >1</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >89</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td></tr><tr><td align="center" valign="middle" >G-8</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >5</td><td align="center" valign="middle" >2</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >6</td><td align="center" valign="middle" >90</td><td align="center" valign="middle" >0</td></tr><tr><td align="center" valign="middle" >G-9</td><td align="center" valign="middle" >6</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >4</td><td align="center" valign="middle" >100</td></tr><tr><td align="center" valign="middle" >Total</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >100</td></tr><tr><td align="center" valign="middle" >Recognition Accuracy (%)</td><td align="center" valign="middle" >93.0</td><td align="center" valign="middle" >98.0</td><td align="center" valign="middle" >91.0</td><td align="center" valign="middle" >92.0</td><td align="center" valign="middle" >97.0</td><td align="center" valign="middle" >100.0</td><td align="center" valign="middle" >89.0</td><td align="center" valign="middle" >90.0</td><td align="center" valign="middle" >100.0</td></tr><tr><td align="center" valign="middle" >Mean Accuracy (%)</td><td align="center" valign="middle" >94.4</td><td align="center" valign="middle" >-</td><td align="center" valign="middle" >-</td><td align="center" valign="middle" >-</td><td align="center" valign="middle" >-</td><td align="center" valign="middle" >-</td><td align="center" valign="middle" >-</td><td align="center" valign="middle" >-</td><td align="center" valign="middle" >-</td></tr></tbody></table></table-wrap><table-wrap id="table2" ><label><xref ref-type="table" rid="table2">Table 2</xref></label><caption><title> The recognition result (weak light)</title></caption><table><tbody><thead><tr><th align="center" valign="middle" ></th><th align="center" valign="middle" >G-1</th><th align="center" valign="middle" >G-2</th><th align="center" valign="middle" >G-3</th><th align="center" valign="middle" >G-4</th><th align="center" valign="middle" >G-5</th><th align="center" valign="middle" >G-6</th><th align="center" valign="middle" >G-7</th><th align="center" valign="middle" >G-8</th><th align="center" valign="middle" >G-9</th></tr></thead><tr><td align="center" valign="middle" >G-1</td><td align="center" valign="middle" >94</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td></tr><tr><td align="center" valign="middle" >G-2</td><td align="center" valign="middle" >1</td><td align="center" valign="middle" >97</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >1</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td></tr><tr><td align="center" valign="middle" >G-3</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >1</td><td align="center" valign="middle" >93</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >3</td><td align="center" valign="middle" >3</td><td align="center" valign="middle" >0</td></tr><tr><td align="center" valign="middle" >G-4</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >93</td><td align="center" valign="middle" >1</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td></tr><tr><td align="center" valign="middle" >G-5</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >4</td><td align="center" valign="middle" >96</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td></tr><tr><td align="center" valign="middle" >G-6</td><td align="center" valign="middle" >1</td><td align="center" valign="middle" >2</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >99</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td></tr><tr><td align="center" valign="middle" >G-7</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >3</td><td align="center" valign="middle" >2</td><td align="center" valign="middle" >3</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >90</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td></tr><tr><td align="center" valign="middle" >G-8</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >4</td><td align="center" valign="middle" >1</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >7</td><td align="center" valign="middle" >94</td><td align="center" valign="middle" >0</td></tr><tr><td align="center" valign="middle" >G-9</td><td align="center" valign="middle" >4</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >3</td><td align="center" valign="middle" >100</td></tr><tr><td align="center" valign="middle" >Total</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >100</td></tr><tr><td align="center" valign="middle" >Recognition Accuracy (%)</td><td align="center" valign="middle" >94.0</td><td align="center" valign="middle" >97.0</td><td align="center" valign="middle" >93.0</td><td align="center" valign="middle" >93.0</td><td align="center" valign="middle" >96.0</td><td align="center" valign="middle" >99.0</td><td align="center" valign="middle" >90.0</td><td align="center" valign="middle" >94.0</td><td align="center" valign="middle" >100.0</td></tr><tr><td align="center" valign="middle" >Mean Accuracy (%)</td><td align="center" valign="middle" >95.1</td><td align="center" valign="middle" >-</td><td align="center" valign="middle" >-</td><td align="center" valign="middle" >-</td><td align="center" valign="middle" >-</td><td align="center" valign="middle" >-</td><td align="center" valign="middle" >-</td><td align="center" valign="middle" >-</td><td align="center" valign="middle" >-</td></tr></tbody></table></table-wrap><table-wrap id="table3" ><label><xref ref-type="table" rid="table3">Table 3</xref></label><caption><title> The recognition result (ordinary light)</title></caption><table><tbody><thead><tr><th align="center" valign="middle" ></th><th align="center" valign="middle" >G-1</th><th align="center" valign="middle" >G-2</th><th align="center" valign="middle" >G-3</th><th align="center" valign="middle" >G-4</th><th align="center" valign="middle" >G-5</th><th align="center" valign="middle" >G-6</th><th align="center" valign="middle" >G-7</th><th align="center" valign="middle" >G-8</th><th align="center" valign="middle" >G-9</th></tr></thead><tr><td align="center" valign="middle" >G-1</td><td align="center" valign="middle" >95</td><td align="center" valign="middle" >1</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >1</td></tr><tr><td align="center" valign="middle" >G-2</td><td align="center" valign="middle" >3</td><td align="center" valign="middle" >96</td><td align="center" valign="middle" >2</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >2</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td></tr><tr><td align="center" valign="middle" >G-3</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >2</td><td align="center" valign="middle" >90</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >1</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >4</td><td align="center" valign="middle" >4</td><td align="center" valign="middle" >0</td></tr><tr><td align="center" valign="middle" >G-4</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >1</td><td align="center" valign="middle" >94</td><td align="center" valign="middle" >1</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td></tr><tr><td align="center" valign="middle" >G-5</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >3</td><td align="center" valign="middle" >95</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td></tr><tr><td align="center" valign="middle" >G-6</td><td align="center" valign="middle" >1</td><td align="center" valign="middle" >1</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >98</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td></tr><tr><td align="center" valign="middle" >G-7</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >5</td><td align="center" valign="middle" >1</td><td align="center" valign="middle" >3</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >91</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td></tr><tr><td align="center" valign="middle" >G-8</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >2</td><td align="center" valign="middle" >2</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >5</td><td align="center" valign="middle" >93</td><td align="center" valign="middle" >0</td></tr><tr><td align="center" valign="middle" >G-9</td><td align="center" valign="middle" >1</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >0</td><td align="center" valign="middle" >3</td><td align="center" valign="middle" >99</td></tr><tr><td align="center" valign="middle" >Total</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >100</td></tr><tr><td align="center" valign="middle" >Recognition Accuracy (%)</td><td align="center" valign="middle" >95.0</td><td align="center" valign="middle" >96.0</td><td align="center" valign="middle" >90.0</td><td align="center" valign="middle" >94.0</td><td align="center" valign="middle" >95.0</td><td align="center" valign="middle" >98.0</td><td align="center" valign="middle" >91.0</td><td align="center" valign="middle" >93.0</td><td align="center" valign="middle" >99.0</td></tr><tr><td align="center" valign="middle" >Mean Accuracy (%)</td><td align="center" valign="middle" >94.6</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr></tbody></table></table-wrap><table-wrap id="table4" ><label><xref ref-type="table" rid="table4">Table 4</xref></label><caption><title> The running time of gesture recognition</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Gesture</th><th align="center" valign="middle" >G-1</th><th align="center" valign="middle" >G-2</th><th align="center" valign="middle" >G-3</th><th align="center" valign="middle" >G-4</th><th align="center" valign="middle" >G-5</th><th align="center" valign="middle" >G-6</th><th align="center" valign="middle" >G-7</th><th align="center" valign="middle" >G-8</th><th align="center" valign="middle" >G-9</th></tr></thead><tr><td align="center" valign="middle" >Mean Running Time (ms)</td><td align="center" valign="middle" >15.5</td><td align="center" valign="middle" >15.8</td><td align="center" valign="middle" >15.7</td><td align="center" valign="middle" >15.6</td><td align="center" valign="middle" >15.5</td><td align="center" valign="middle" >15.6</td><td align="center" valign="middle" >15.7</td><td align="center" valign="middle" >15.8</td><td align="center" valign="middle" >15.2</td></tr></tbody></table></table-wrap><table-wrap id="table5" ><label><xref ref-type="table" rid="table5">Table 5</xref></label><caption><title> Comparative results of the methods in [<xref ref-type="bibr" rid="scirp.65570-ref9">9</xref>] , [<xref ref-type="bibr" rid="scirp.65570-ref14">14</xref>] and the proposed method</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Method</th><th align="center" valign="middle" >Mean Accuracy</th><th align="center" valign="middle" >Running Time (s)</th></tr></thead><tr><td align="center" valign="middle" >Convex Shape Decomposition Method in [<xref ref-type="bibr" rid="scirp.65570-ref9">9</xref>]</td><td align="center" valign="middle" >91.9%</td><td align="center" valign="middle" >0.026</td></tr><tr><td align="center" valign="middle" >Thresholding Decomposition + FEMD in [<xref ref-type="bibr" rid="scirp.65570-ref14">14</xref>]</td><td align="center" valign="middle" >90.6%</td><td align="center" valign="middle" >0.5004</td></tr><tr><td align="center" valign="middle" >Near-convex Decomposition + FEMD in [<xref ref-type="bibr" rid="scirp.65570-ref14">14</xref>]</td><td align="center" valign="middle" >93.9%</td><td align="center" valign="middle" >4.0012</td></tr><tr><td align="center" valign="middle" >Proposed Method</td><td align="center" valign="middle" >94.7%</td><td align="center" valign="middle" >0.0156</td></tr></tbody></table></table-wrap><p>From <xref ref-type="table" rid="table1">Table 1</xref>, <xref ref-type="table" rid="table2">Table 2</xref> and <xref ref-type="table" rid="table3">Table 3</xref> we can notice that the method proposed in this paper has a high recognition accuracy, especially in Gesture-2, Gesture-6 and Gesture-9. The recognition accuracy of Gesture-3, Gesture-7 and Gesture-8 is slightly lower than other gestures, which result from the great difference of the expression of gestures. As a whole, the mean recognition accuracy reach 94.7%, which prove the effectiveness of this method. Meanwhile, another advantage of this method is that the recognition accuracy isn’t impacted by light intensity. It can work well in strong light or weak light environment, even in the darkness. In addition, from <xref ref-type="table" rid="table4">Table 4</xref> we can find that the proposed method has extremely short running time. Almost every recognition can be completed within 16ms, in other word, the gesture recognition speed can reach about 60 frames per second, which can completely meet the needs of real-time application. The efficient recognition process ensures that it can be used in real-time situation.</p><p>We also compare the proposed system with the previous work. The recognition methods proposed in [<xref ref-type="bibr" rid="scirp.65570-ref14">14</xref>] are geometry-based. Two methods in that paper can’t balance accuracy and running time well. And the convex shape decomposition method is employ in [<xref ref-type="bibr" rid="scirp.65570-ref9">9</xref>] . Although the running time get great improvement, the accuracy isn’t so satisfying. However, compared with these methods, both the mean accuracy and the running time in our method can reach a pretty good effect. Experiments demonstrate our method is much more efficient for real-time applications, which is shown in <xref ref-type="table" rid="table5">Table 5</xref>.</p></sec><sec id="s7"><title>7. Conclusion</title><p>Gesture recognition has a wide range of applications in Human-Computer-Interaction. This paper proposes an efficient and succinct method for gesture recognition, which takes advantages of 3D point cloud data. The 3D point cloud data is collected from depth camera, then it is transformed into binary image. Basing on binary image, three different appearance features are extracted, including the number of stretched fingers, the angles between fingers as features and the gesture region’s area distribution feature. Finally, the decision tree is constructed for gesture recognition. Extensive experimental results demonstrate accuracy and robustness of the method proposed in this paper. As a result, this method can play a role in the application of real-time gesture recognition.</p></sec><sec id="s8"><title>Acknowledgements</title><p>This paper is funded by National Natural Science Foundation of China (61572372, 41271398), the Foundation Research Funds for the Central Universities (204201kf0242, 204201kf0263), Shanghai Aerospace Science and Technology Innovation Fund Projects (SAST201425).</p></sec><sec id="s9"><title>Cite this paper</title><p>Yanwen Chong,Jianfeng Huang,Shaoming Pan, (2016) Hand Gesture Recognition Using Appearance Features Based on 3D Point Cloud. Journal of Software Engineering and Applications,09,103-111. doi: 10.4236/jsea.2016.94009</p></sec></body><back><ref-list><title>References</title><ref id="scirp.65570-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">Ren, Z., Meng, J. and Yuan, J. (2011) Depth Camera Based Hand Gesture Recognition and Its Applications in Human-Computer-Interaction. Information, Communications and Signal Processing (ICICS), Singapore, 13-16 December 2011, 1-5.</mixed-citation></ref><ref id="scirp.65570-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">Suarez, J. and Murphy, R.R. (2012) Hand Gesture Recognition with Depth Images: A Review. RO-MAN, IEEE, Paris, 9-13 September 2012, 411-417.</mixed-citation></ref><ref id="scirp.65570-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">Liu, S., Liu, Y., Yu, J. and Wang, Z. (2015) Hierarchical Static Hand Gesture Recognition by Combining Finger Detection and HOG Features. Journal of Image and Graphics, 20, 0781-0788.</mixed-citation></ref><ref id="scirp.65570-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">Yu, S., Cao, J., Li, P., et al. (2015) Hand Gesture Recognition Based on The Spatial Pyramid Bag of Features. CAAI Transactions on Intelligent Systems, 10, 429-435.</mixed-citation></ref><ref id="scirp.65570-ref5"><label>5</label><mixed-citation publication-type="book" xlink:type="simple">Janoch, A., Karayev, S., Jia, Y., Barron, J., Fritz, M., Saenko, K. and Darrell, T. (2013) A Category-Level 3D Object Dataset: Putting the Kinect to Work. In: Fossati, A., Gall, J., Grabner, H., Ren, X.F. and Konolige, K., Eds., Consumer Depth Cameras for Computer Vision, Springer London, London, 141-165. http://dx.doi.org/10.1007/978-1-4471-4640-7_8</mixed-citation></ref><ref id="scirp.65570-ref6"><label>6</label><mixed-citation publication-type="book" xlink:type="simple">Dominio, F., Donadeo, M. and Zanuttigh, P. (2014) Combining Multiple Depth-Based Descriptors ForHand Gesture Recognition. In: Borgefors, G., Sanniti di Baja, G. and Sarkar, S., Eds., Pattern Recognition Letters, Elsevier, Amsterdam, 101-111. http://dx.doi.org/10.1016/j.patrec.2013.10.010</mixed-citation></ref><ref id="scirp.65570-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">Nguyen, L.T., Thanh, C.D., Ba, T.N., Viet, C.T. and Thanh, H.L. (2013) Contour Based Hand Gesture Recognition Using Depth Data. Advanced Science and Technology Letters, 29, 60-65.</mixed-citation></ref><ref id="scirp.65570-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">Konda, K.R., Konigs, A., Schulz, H. and Schulz, D. (2012) Real Time Interaction with Mobile Robots Using Hand Gestures. Proceedings of the seventh annual ACM/IEEE international conference on Human-Robot Interaction, Boston, 5-8 March 2012, 177-178.</mixed-citation></ref><ref id="scirp.65570-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">Qin, S., Zhu, X., Yang, Y. and Jiang, Y. (2014) Real-Time Hand Gesture Recognition from Depth Images Using Convex Shape Decomposition Method. Journal of Signal Processing Systems, 74, 47-58. http://dx.doi.org/10.1007/s11265-013-0778-7</mixed-citation></ref><ref id="scirp.65570-ref10"><label>10</label><mixed-citation publication-type="other" xlink:type="simple">Daribo, I. and Saito, H. (2011) A Novel Inpainting-based Layered Depth Video for 3dtv. IEEE Transactions on Broadcasting, 57, 533-541. http://dx.doi.org/10.1109/TBC.2011.2125110</mixed-citation></ref><ref id="scirp.65570-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">Telea, A. (2004) AnImage Inpainting Technique Based on the Fast Marching Method. Journal of Graphics Tools, 9, 23-34. http://dx.doi.org/10.1080/10867651.2004.10487596</mixed-citation></ref><ref id="scirp.65570-ref12"><label>12</label><mixed-citation publication-type="other" xlink:type="simple">Kopf, J., Cohen, M.F., Lischinski, D. and Uyttendaele, M. (2007) Joint Bilateral Upsampling. ACM Transactions on Graphics, 26, 1-8. http://dx.doi.org/10.1145/1275808.1276497</mixed-citation></ref><ref id="scirp.65570-ref13"><label>13</label><mixed-citation publication-type="other" xlink:type="simple">Cao, C., Li, R. and Zhao, L. (2012) Hand Posture Recognition Method Based on Depth Image Technology. Computer Engineering, 38, 16-18.</mixed-citation></ref><ref id="scirp.65570-ref14"><label>14</label><mixed-citation publication-type="other" xlink:type="simple">Ren, Z., Yuan, J. and Zhang, Z. (2011) Robust Hand Gesture Recognition Based on Finger-Earth Mover’s Distance with a Commodity Depth Camera. ACM International Conference on Multimedia, Scottsdale, 28 November-1 December 2011, 1093-1096. http://dx.doi.org/10.1145/2072298.2071946</mixed-citation></ref></ref-list></back></article>