<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">OALibJ</journal-id><journal-title-group><journal-title>Open Access Library Journal</journal-title></journal-title-group><issn pub-type="epub">2333-9705</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/oalib.1104747</article-id><article-id pub-id-type="publisher-id">OALibJ-86539</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Biomedical&amp;Life Sciences</subject><subject> Business&amp;Economics</subject><subject> Chemistry&amp;Materials Science</subject><subject> Computer Science&amp;Communications</subject><subject> Earth&amp;Environmental Sciences</subject><subject> Engineering</subject><subject> Medicine&amp;Healthcare</subject><subject> Physics&amp;Mathematics</subject><subject> Social Sciences&amp;Humanities</subject></subj-group></article-categories><title-group><article-title>
 
 
  Testing the Menzerath-Altmann Law in the Sentence Level of Written Chinese
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Heng</surname><given-names>Chen</given-names></name><xref ref-type="aff" rid="aff1"><sub>1</sub></xref></contrib></contrib-group><aff id="aff1"><label>1</label><addr-line>Center for Linguistics and Applied Linguistics, Guangdong University of Foreign Studies, Guangzhou, China</addr-line></aff><author-notes><corresp id="cor1">* E-mail:</corresp></author-notes><pub-date pub-type="epub"><day>03</day><month>08</month><year>2018</year></pub-date><volume>05</volume><issue>08</issue><fpage>1</fpage><lpage>5</lpage><history><date date-type="received"><day>1,</day>	<month>July</month>	<year>2018</year></date><date date-type="rev-recd"><day>5,</day>	<month>August</month>	<year>2018</year>	</date><date date-type="accepted"><day>8,</day>	<month>August</month>	<year>2018</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
  Language unit is a fundamental conception in modern linguistics, but the boundaries are not clear between language levels. As language is a multi-level system, quantification rather than microscopic grammatical analysis shoul
  d be used to investigate into this question. In this paper, Menzerath-Altmann law is used to make out the basic language units in written Chinese in the sentence level.
 
</p></abstract><kwd-group><kwd>Menzerath-Altmann Law</kwd><kwd> Sentence</kwd><kwd> Language Level</kwd><kwd> Written Chinese</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>Language levels and language units are critical conceptions in a language system, and they are highly related with the entities in a language, as well as the methods used [<xref ref-type="bibr" rid="scirp.86539-ref1">1</xref>] [<xref ref-type="bibr" rid="scirp.86539-ref2">2</xref>] . Generally, five language units are recognized by grammarians: morpheme, word, phrase, clause and sentence [<xref ref-type="bibr" rid="scirp.86539-ref3">3</xref>] . However, different linguistic schools have different opinions upon the systematicness of language. Therefore, the methods and standards they use to divide language levels and units are different. These language units include sound, word, phrase, sentence, phone, phoneme, morph, morpheme, syllable, affix, word-group, etc.</p><p>The most characteristic feature of modern linguistics is structuralism. Briefly, this means that language is not a haphazard conglomerate of words and sounds but a well knit and coherent whole. However, linguistics is traditionally preoccupied with the fine detail of language structure, or in other words, the language phenomena at the microscopic scale rather than at the system level [<xref ref-type="bibr" rid="scirp.86539-ref4">4</xref>] . Therefore, it is not ordinarily feasible to analyze each language level separately, and the work must be carried on simultaneously on all levels. Moreover, the results should be stated in terms of an orderly hierarchy of levels [<xref ref-type="bibr" rid="scirp.86539-ref3">3</xref>] .</p><p>Menzerath-Altmann law is a general statement about the natural language constructions which says: the longer is a construction, the shorter are its constituents. Language is a whole complex system, and it is a set of relations. The language units correlate with each other in different levels and in complex ways. The whole is composed of its parts, which interact with each other. Language units in the same levels are relatively homogeneous. Therefore, the relation between two adjacent language levels is a “whole-part” relationship.</p><p>Actually, in quantitative linguistics, the relationship between “whole-part” has been extensively investigated [<xref ref-type="bibr" rid="scirp.86539-ref5">5</xref>] . This relation was investigated and tested on many linguistic levels and in many languages and even on some non-linguistic data [<xref ref-type="bibr" rid="scirp.86539-ref6">6</xref>] . [<xref ref-type="bibr" rid="scirp.86539-ref7">7</xref>] conducted the first empirical test of the Menzerath-Altmann law on “sentence &gt; clause &gt; word”, analyzing German and English short stories and philosophical texts. The tests on the data confirmed the validity of the law with high significance. The law has also been used to study phenomena on the supra-sentencial level and fractal structures of text [<xref ref-type="bibr" rid="scirp.86539-ref8">8</xref>] [<xref ref-type="bibr" rid="scirp.86539-ref9">9</xref>] . This is why this law is considered one of the most frequently corroborated laws in linguistics. The law is a good example of the importance of the quantitative linguistic methodology, since it clearly shows that the “independent language subsystems” are in fact interconnected by relationships which are hard to detect by a qualitative research.</p><p>In this paper, we will test the construction units in written Chinese in the sentence level, i.e., “sentence-clause-word”. The rest of this paper is organized as follows. Section 2 introduces the materials and methods of the present study. Section 3 presents the results of the tests for “sentence-clause-word” levels. Section 4 concludes the study and makes suggestions for further research.</p></sec><sec id="s2"><title>2. Materials and Methods</title><p>We use the Lancaster Chinese corpus (LCMC) as the testing material. The corpus is segmented and part of speech (POS) tagged, and its basic information is in <xref ref-type="table" rid="table1">Table 1</xref>.</p><p>The language units we will test in this paper are word, clause and sentence. The reason why we do not include phrase here is that, a complete sentence or</p><table-wrap id="table1" ><label><xref ref-type="table" rid="table1">Table 1</xref></label><caption><title> Basic information of LCMC</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Language units</th><th align="center" valign="middle" >Scale</th></tr></thead><tr><td align="center" valign="middle" >Character (tokens)</td><td align="center" valign="middle" >1,314,058</td></tr><tr><td align="center" valign="middle" >Character (types)</td><td align="center" valign="middle" >4705</td></tr><tr><td align="center" valign="middle" >Clauses (types)</td><td align="center" valign="middle" >126,455</td></tr><tr><td align="center" valign="middle" >Sentence (types)</td><td align="center" valign="middle" >45,969</td></tr><tr><td align="center" valign="middle" >Word (types)</td><td align="center" valign="middle" >847,521</td></tr></tbody></table></table-wrap><p>clause cannot be divided into several sequential phrases, both theoretically and practically.</p><p>All the language units are easy to get in LCMC by using some tools except clause. Therefore, in the following we will firstly define the other language units, and then give our methods of defining phrase.</p><p>In written Chinese, sentences are separated from one another by using special marks of punctuation (full-stop, question-mark, exclamation-mark). As for our case, the sentences are tagged in LCMC, so here there is no difficulty distinguishing sentence.</p><p>Clause is not tagged in LCMC, nor in any other corpus available. Generally speaking, clause is the smallest independent grammatical unit of expression. But this definition can hardly be used to obtain the clauses in LCMC. [<xref ref-type="bibr" rid="scirp.86539-ref10">10</xref>] analyzes a long sentence from a literary book, and claims that the constituents just between two punctuations (comma and period) can be defined as clauses roughly. We believe that although this method is not so exact in grammatical analyses, it can be used in large-scale-corpus studies. But we need to state that, since in LCMC sentences are tagged, we choose comma and semicolon as our marks of clause boundaries.</p><p>After obtaining all the statics with respect to language units in LCMC, we use Menzerath-Altmann law to fit the hierarchical data.</p><p>Menzerath-Altmann law (short for Menzerathian function) describes the mathematical relation between two adjacent language units, and its model function is</p><p>y = a x b e − c x (1)</p><p>In this function, y represents the length of the upper language unit, and x represent the mean length of the lower language unit; a, b, c are parameters which seem to depend mainly on the level of the language units under investigation- much more than on language, the kind of text, or author as previously expected., and e is natural constant, which equals 2.71828 approximately. The goodness of fit can be seen from determination coefficient R<sup>2</sup>. We say the result is accepted for R<sup>2</sup> &gt; 0.75, good for R<sup>2</sup> &gt; 0.80, and very good for R<sup>2</sup> &gt; 0.90.</p></sec><sec id="s3"><title>3. Results</title><p>The language units we examine in this paper are “word &gt; clause &gt; sentence” (here we use “&gt;” to direct to a higher-rank unit in written Chinese). <xref ref-type="table" rid="table2">Table 2</xref> shows the Menzerathian data of this group.</p><p>As can be seen in <xref ref-type="table" rid="table2">Table 2</xref>, the sentence length is measured in clause, and the clause length is measured in word.</p><p>The Menzerathian function is used to fit the data, and the fitting results are displayed in <xref ref-type="fig" rid="fig1">Figure 1</xref>.</p><p>As can be seen from the fitting results in <xref ref-type="fig" rid="fig1">Figure 1</xref>, the goodness of fit indicator R<sup>2</sup> (0.8498) indicates the result is good. This means that the group “word &gt; clause &gt; sentence” lines with Menzerath-Altmann law.</p><table-wrap id="table2" ><label><xref ref-type="table" rid="table2">Table 2</xref></label><caption><title> Hierarchical data of “word &gt; clause &gt; sentence”</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Sentence length (in clause)</th><th align="center" valign="middle" >Mean clause length (in word)</th><th align="center" valign="middle" >Sentence length (in clause)</th><th align="center" valign="middle" >Mean clause length (in word)</th></tr></thead><tr><td align="center" valign="middle" >1</td><td align="center" valign="middle" >7.7407</td><td align="center" valign="middle" >9</td><td align="center" valign="middle" >6.2194</td></tr><tr><td align="center" valign="middle" >2</td><td align="center" valign="middle" >7.0465</td><td align="center" valign="middle" >10</td><td align="center" valign="middle" >6.3932</td></tr><tr><td align="center" valign="middle" >3</td><td align="center" valign="middle" >6.7162</td><td align="center" valign="middle" >11</td><td align="center" valign="middle" >5.8068</td></tr><tr><td align="center" valign="middle" >4</td><td align="center" valign="middle" >6.4866</td><td align="center" valign="middle" >12</td><td align="center" valign="middle" >5.7661</td></tr><tr><td align="center" valign="middle" >5</td><td align="center" valign="middle" >6.3357</td><td align="center" valign="middle" >13</td><td align="center" valign="middle" >6.1723</td></tr><tr><td align="center" valign="middle" >6</td><td align="center" valign="middle" >6.2485</td><td align="center" valign="middle" >14</td><td align="center" valign="middle" >6.5510</td></tr><tr><td align="center" valign="middle" >7</td><td align="center" valign="middle" >6.1646</td><td align="center" valign="middle" >15</td><td align="center" valign="middle" >6.4500</td></tr><tr><td align="center" valign="middle" >8</td><td align="center" valign="middle" >6.2296</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td></tr></tbody></table></table-wrap></sec><sec id="s4"><title>4. Conclusions</title><p>In Section 3, we tested the Menzerath-Altmann law in the “word &gt; clause &gt; sentence” levels. The results show that they are in line with the Menzrathian law.</p><p>Language is a system. This view has been put forward for about 100 years, however, it has never been realized until quantification is introduced into linguistics. In this paper, we show that Menzerath-Altmann law can be an efficient way of finding the basic language units in a language in the sentence level. Since language is a complex adaptive system, in the future, we will investigate into this question from a diachronic perspective to see if these Menzerathian levels will change over time.</p></sec><sec id="s5"><title>Acknowledgements</title><p>This work is supported by the Education Department of Guangdong Province “Innovative Strong School Project” Youth Innovation Talents Project (Humanities and Social Sciences) (Project Number: 2017WQNCX046).</p></sec><sec id="s6"><title>Conflicts of Interest</title><p>The author declares no conflicts of interest regarding the publication of this paper.</p></sec><sec id="s7"><title>Cite this paper</title><p>Chen, H. (2018) Testing the Menzerath-Altmann Law in the Sentence Level of Written Chinese. Open Access Library Journal, 5: e4747. https://doi.org/10.4236/oalib.1104747</p></sec></body><back><ref-list><title>References</title><ref id="scirp.86539-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">Altmann, G. (1996) The Nature of Linguistic Units. Journal of Quantitative Linguistics, 3, 1-7. https://doi.org/10.1080/09296179608590059</mixed-citation></ref><ref id="scirp.86539-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">Chen, H. and Liu, H. (2016) How to Measure Word Length in Spoken and Written Chinese. Journal of Quantitative Linguistics, 23, 5-29.https://doi.org/10.1080/09296174.2015.1071147</mixed-citation></ref><ref id="scirp.86539-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">Lyons, J. (1968) Introduction to Theoretical Linguistics. Cambridge University Press, London.</mixed-citation></ref><ref id="scirp.86539-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">Liu, H. and Cong, J. (2014) Empirical Characterization of Modern Chinese as a Multi-Level System from the Complex Network Approach. Journal of Chinese Linguistics, 42, 1-38.</mixed-citation></ref><ref id="scirp.86539-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">Menzerath, P. (1954) Die Architektonik des deutschen Wortschatzes. Dümmler, Bonn.</mixed-citation></ref><ref id="scirp.86539-ref6"><label>6</label><mixed-citation publication-type="book" xlink:type="simple">Altmann, G. and Schwibbe, H. (Eds.) (1989) Das Menzerathsche Gesetz in informationsverarbeitenden Systemen. Georg OlmsVerlag, Hildesheim.</mixed-citation></ref><ref id="scirp.86539-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">Kohler, R. (2012) Quantitative Syntax Analysis. Walter de Gruyter, Ber-lin/Boston.https://doi.org/10.1515/9783110272925</mixed-citation></ref><ref id="scirp.86539-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">Hrebícek, L. (1995) Text Levels: Language Constructs, Constituents and the Menzerath-Altmann Law. WVT, Wiss. Verlag Trier.</mixed-citation></ref><ref id="scirp.86539-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">Andres, J. (2010) On a Conjecture about the Fractal Structure of Language. Journal of Quantitative Linguistics, 17, 101-122. https://doi.org/10.1080/09296171003643189</mixed-citation></ref><ref id="scirp.86539-ref10"><label>10</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Luke</surname><given-names> K. </given-names></name>,<etal>et al</etal>. (<year>2006</year>)<article-title>On the Status of the Clause in Chinese Grammar</article-title><source> Chinese Linguistics</source><volume> 15</volume>,<fpage> 2</fpage>-<lpage>14</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref></ref-list></back></article>