<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">JSIP</journal-id><journal-title-group><journal-title>Journal of Signal and Information Processing</journal-title></journal-title-group><issn pub-type="epub">2159-4465</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/jsip.2014.53010</article-id><article-id pub-id-type="publisher-id">JSIP-48303</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>COMPUTER SCIENCE &amp; COMMUNICATIONS</subject></subj-group></article-categories><title-group><article-title>An Intonation Speech Synthesis Model for Indonesian Using Pitch Pattern and Phrase Identification</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Yohanes</surname><given-names>Suyanto</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Subanar</surname><given-names> </given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Agus</surname><given-names>Harjoko</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Sri</surname><given-names>Hartati</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib></contrib-group><aff id="aff1"><addr-line>Department of Computer Science and Electronics, Gadjah Mada University, Yogyakarta, Indonesia</addr-line></aff><author-notes><corresp id="cor1">* E-mail:<email>yanto@ugm.ac.id(YS)</email>;<email>subanar@ugm.ac.id(S)</email>;<email>aharjoko@ugm.ac.id(AH)</email>;<email>shartati@ugm.ac.id(SH)</email>;</corresp></author-notes><pub-date pub-type="epub"><day>30</day><month>07</month><year>2014</year></pub-date><volume>05</volume><issue>03</issue><fpage>80</fpage><lpage>88</lpage><history><date date-type="received"><day>25</day>	<month>May</month>	<year>2014</year></date><date date-type="rev-recd"><day>20</day>	<month>June</month>	<year>2014</year>	</date><date date-type="accepted"><day>16</day>	<month>July</month>	<year>2014</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
	Prosody in speech
synthesis systems (text-to-speech) is a determinant of tone, duration, and
loudness of speech sound. Intonation is a part of prosody which determines the
speech tone. In Indonesian, intonation is determined by the structure of
sentences, types of sentences, and also the position of the word in a sentence.
In this study, a model of speech synthesis that focuses on its intonation is
proposed. The speech intonation is determined by sentence structure, intonation
patterns of the example sentences, and general rules of Indonesian
pronunciation. The model receives texts and intonation patterns as inputs.
Based on the general principle of Indonesian pronunciation, a prosody file was
made. Based on input text, sentence structure is determined and then interval
among parts of a sentence (phrase) can be determined. These intervals are used
to correct the duration of the initial prosody file. Furthermore, the
frequencies in prosody file were corrected using intonation patterns. The final
result is prosody file that can be pronounced by speech engine application.
Experiment results of studies using the original voice of radio news announcer
and the speech synthesis show that the peaks of F<sub>0</sub> are determined by general rules or
intonation patterns which are dominant. Similarity test with the PESQ method
shows that the result of the synthesis is 1.18 at MOS-LQO scale.
</p></abstract><kwd-group><kwd>Speech Synthesis</kwd><kwd> PESQ</kwd><kwd> Intonation</kwd><kwd> Indonesian</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>Modeling emotions in speech synthesis are made based on a number of parameters such as place, the level of the fundamental frequency<inline-formula><inline-graphic xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\b0570fc9-f513-445f-8901-ce9a85873f75.png" xlink:type="simple"/></inline-formula>, voice quality, articulation or accuracy. A different speech synthesis technique aimed for parameters in the different levels Schroder 2001. Formant synthesis, also known as the rule-based synthesis, creates the sound of speech based on the rules alone. No people’s voice recording is involved in it. The result is a speech sound like a robot voice.</p><p>In the concatenation synthesis, voice recording from speaker strung together to produce synthetic speech. In this process, diphone which cuts the signal in mid phoneme until next mid phoneme is often used. Diphone is recorded with a monotonous tone. At the synthesis, required contour <inline-formula><inline-graphic xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\6bda763a-b89e-42d0-b0d3-9ba74e3e5cd5.png" xlink:type="simple"/></inline-formula> is constructed with signal processing techniques that result in a slight distortion. But the final result is more natural than formant synthesis [<xref ref-type="bibr" rid="scirp.48303-ref1">1</xref>] . In most diphone synthesis systems, only <inline-formula><inline-graphic xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\58eabdd6-6153-4e3e-b794-4fbb81499a89.png" xlink:type="simple"/></inline-formula> and duration can be controlled, while intensity for each segment is not easy to control.</p><p>Emotions of speaker such as neutral, joy, boredom, anger, sadness, fear, and indignation are reflected in the duration and intonation of speech from a person [<xref ref-type="bibr" rid="scirp.48303-ref2">2</xref>] . By copying pitch and duration of the original utterances to a monotonous one, it can be proved that both factors are sufficient to express various emotions. Research results show that emotion can be expressed simply by manipulating pitch and duration. Intonation is a set of syntactic prosody inherent in the utterance sentence [<xref ref-type="bibr" rid="scirp.48303-ref2">2</xref>] . The intensity of sound is not manipulated.</p><p>Contour tones containing duration and pitch are calculated by converting an intonation plan into a sequence of prosodic tone and highly dependent on the model of the speaker [<xref ref-type="bibr" rid="scirp.48303-ref3">3</xref>] . Contour patterns derived from the speaker’s voice are applied at the time of speech synthesis.</p><p>Research on Indonesian speech synthesis which makes the coupling phoneme based speech synthesis has been done [<xref ref-type="bibr" rid="scirp.48303-ref4">4</xref>] . In general, the results can pronounce words in Indonesian quite fluently and can be understood by most listeners. However, speech synthesis has not produced intonation patterns (prosody) as the original speech [<xref ref-type="bibr" rid="scirp.48303-ref4">4</xref>] . Still there are some inaccuracies in the assembly of phonemes [<xref ref-type="bibr" rid="scirp.48303-ref4">4</xref>] .</p><p>In Indonesian, utterance position on word stress does not depend on the number of syllables, but the stress falls on the penultimate syllable [<xref ref-type="bibr" rid="scirp.48303-ref5">5</xref>] . Prosodic type is characterized by <inline-formula><inline-graphic xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\f5148571-be9a-4d8c-8aa6-a5e905e295a1.png" xlink:type="simple"/></inline-formula> and intensity. The pressure of nouns group on preverb position will be marked by the duration instead of the frequency. In this group, two peaks of <inline-formula><inline-graphic xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\a76c6bd5-8ce0-4ba7-b0a9-10ed059485b0.png" xlink:type="simple"/></inline-formula> are found.</p></sec><sec id="s2"><title>2. Speech Synthesis Model Using Pitch Pattern and Phrase Identification</title><p>Speech synthesis is built to converts the input text into speech. Conversion of text into speech considering a sentence structure, intonation patterns, as well as a normalization. The results are in the form of text that qualifies as input of voice generator.</p><p>Prosodic rules which determine the frequency, duration, and intonation was compiled from the results of studying materials about intonation in Indonesian [<xref ref-type="bibr" rid="scirp.48303-ref6">6</xref>] . This rule is encoded to ease the implementation later. Speech synthesis parameters are grouped into four categories, namely pitch, duration, quality and articulation. The pitch parameter determines the value of <inline-formula><inline-graphic xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\a3499102-c1ff-48a7-802b-e2218b83f9ba.png" xlink:type="simple"/></inline-formula> [<xref ref-type="bibr" rid="scirp.48303-ref7">7</xref>] . The duration parameter determines the duration of rhythm control and rate of speech. Usually the pitch and duration parameters are associated with the phenomenon of linguistic (words or phrases) [<xref ref-type="bibr" rid="scirp.48303-ref3">3</xref>] [<xref ref-type="bibr" rid="scirp.48303-ref8">8</xref>] . The sound quality of speech is a parameter that determines the overall sound quality. The articulation parameter determines clearance of speech. Based on these parameters the model of speech synthesis is composed and shown in <xref ref-type="fig" rid="fig1">Figure 1</xref>.</p><p>The text comes from standard input or the file is read by the text normalization module in order to obtain text which is free of numbers and signs so that the whole text can be pronounced. The output of a normalization module text-norm will be used by the selector pattern, syntax analyst, and synthesis module. The selector pattern module will produce a pattern based on text-norm and an existing pattern database. This module is responsible for the pitch of each phoneme. Syntax analyst module will generate sentence syntax structure of the text-norm with the phrases and words interval duration information.</p><p>The synthesis module combined text-norm from normalization module, phonemes with prosody information in it based on the pattern from pattern selector, and attributes of sentence structure from syntax analysts into phonemes and prosody form. This module is responsible for the articulation and the quality of speech. Finally, this form will be voiced by DSP module.</p><sec id="s2_1"><title>2.1. Normalization Module</title><p>The text is a phrase that came from the standard input or files that may be contains numbers or punctuation</p><fig id="fig1"><label>Figure 1</label><caption><p> The proposed model of speech synthesis</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\b0a7758c-e524-4684-9418-8a361fd22249.png"/></fig><p>marks that can be pronounced. If there are numbers then this module will convert it into text. For example, in the text there is a number “12345” then it will be converted to “dua belas ribu tiga ratus empat puluh lima”. Similarly, if there is an abbreviation, like “UGM” then it will be changed to “u-ge-em”. Symbols need to be changed too, e.g. “<inline-formula><inline-graphic xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\86a863fc-3f48-4c6c-abf9-6a40d6d772df.png" xlink:type="simple"/></inline-formula>”, “<inline-formula><inline-graphic xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\5d1554d2-1d97-4c59-8cb3-ba8da024bfd4.png" xlink:type="simple"/></inline-formula>”, or “<inline-formula><inline-graphic xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\f6163a46-7afa-43cb-b807-2d8b469eea2e.png" xlink:type="simple"/></inline-formula>”.</p></sec><sec id="s2_2"><title>2.2. Fundamental Frequency Pattern Module</title><p>Intonation patterns in the form of pairs of time <inline-formula><inline-graphic xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\dc19f887-ae11-4e60-a6f0-89bfc6ede00b.png" xlink:type="simple"/></inline-formula> and fundamental frequency <inline-formula><inline-graphic xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\2bf99a9d-2d31-4e44-929a-96ac617dffaa.png" xlink:type="simple"/></inline-formula> obtained by extraction <inline-formula><inline-graphic xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\a80b06bb-94da-4cd4-9270-ed6c1534393b.png" xlink:type="simple"/></inline-formula> of radio announcer voice. Based on the previously defined parameters the intonation pattern chose from a set of intonation patterns. The selection is done by paying attention on the pattern duration and length of the text. The duration of the speech text can be calculated to about sum of the duration of all phonemes. Results of pattern selector are <inline-formula><inline-graphic xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\b74c9528-d76e-4b60-8730-42aa1ce74357.png" xlink:type="simple"/></inline-formula> for time from 0 to<inline-formula><inline-graphic xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\2a784658-8378-4e9c-be19-6d5e64121068.png" xlink:type="simple"/></inline-formula>. The pattern have been adapted to the text by interpolation method.</p></sec><sec id="s2_3"><title>2.3. Syntactic Module</title><p>This module perform the text-parsing norm based standard grammatical structures. An important outcome of this module is the determination of the duration of phonemes and duration of pauses between phrases or between words.</p><p>Syntax analysis that will be used is Bahasa Indonesia structure rules [<xref ref-type="bibr" rid="scirp.48303-ref9">9</xref>] -[<xref ref-type="bibr" rid="scirp.48303-ref11">11</xref>] . The rules are converted into a context-free grammar structure (CFG). One notation is often used to write the CFG is Backus-Naur Form—BNF. Here are the example segments of Bahasa Indonesia BNF:</p><disp-formula id="scirp.48303-formula4329"><label>(1)</label><inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\962bfabc-0b19-42a8-b7f8-4c8843a0d236.png"/></disp-formula><disp-formula id="scirp.48303-formula4330"><label>(2)</label><inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\5d02be99-cbb7-4f46-aadb-5881326b9e69.png"/></disp-formula><disp-formula id="scirp.48303-formula4331"><label>(3)</label><inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\f83bee94-9a53-4517-995f-68f37d1d17ed.png"/></disp-formula><disp-formula id="scirp.48303-formula4332"><label>(4)</label><inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\183f8124-38a0-4ada-a0d9-fbf33c3692e7.png"/></disp-formula><disp-formula id="scirp.48303-formula4333"><label>(5)</label><inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\945c0a22-c962-4147-8a8c-add2771ad311.png"/></disp-formula><disp-formula id="scirp.48303-formula4334"><label>(6)</label><inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\7b7ad396-3160-4c3d-b518-b2218182cbbd.png"/></disp-formula><disp-formula id="scirp.48303-formula4335"><label>(7)</label><inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\9d939bb3-9dd4-451c-82e4-20b331f27a17.png"/></disp-formula><disp-formula id="scirp.48303-formula4336"><inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\6feaa3b7-8bb5-4911-abfe-7942c08f4fb8.png"/></disp-formula><disp-formula id="scirp.48303-formula4337"><label>(8)</label><inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\1a871758-25c5-4022-9658-98c0801d881b.png"/></disp-formula><p>Noun phrase (“frasa-nominal”) can be construct from noun (“nomina”) or noun phrase followed by another phrase. For example, noun phrase and verb phrase (“frasa-verbal”). In another segment verb phrase can be construct from verb (“verba”) and another phrase. Completed BNF for Bahasa Indonesia has been made.</p></sec><sec id="s2_4"><title>2.4. Synthesis Module</title><p>Sentence in the text is converted into the format of pho with attention to intonation patterns and phoneme duration and duration of pauses between phrases. Intonation patterns derived from the module selector pattern, while the duration of phonemes and pause duration is derived from syntax analyst module. The pho format is ready to be fed to the DSP module to be voiced.</p><p>From pattern selector module resulting array of phonemes, includes spaces, and pho notation for the corresponding phonem. On the other hand, analyst module generates duration of pauses between words from the input sentence. The length of this pause is then used to correct the lag length of the pattern selector module output. It is expected the end result is in accordance with the intonation patterns and also in accordance with the structure of the sentence.</p><p>Illustration of the results of output from the intonation pattern module, the sentence structure analyst module, and synthesis module are shown in <xref ref-type="fig" rid="fig2">Figure 2</xref>. Output of intonation patterns module focused on intonation while sentence analyst module more focused on the pause duration between words.</p></sec></sec><sec id="s3"><title>3. Results</title><p>Fist, a pattern that contents of series of <inline-formula><inline-graphic xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\c50d2a43-64b3-465d-a10d-6ad9948411b0.png" xlink:type="simple"/></inline-formula> is got from a recorded voice by Praat application. Then, the input text was normalized by a normalized module written in PHP. Its output is called as text-norm. The analyst syntax running on text-norm gets the sentence structure with duration of phrases and words. The synthesis module composed a series of phonemes and prosody from the norm-text, the <inline-formula><inline-graphic xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\63ee6b29-51ec-4bc7-bbe9-0054f3e113b6.png" xlink:type="simple"/></inline-formula> pattern, and the sentence structure. Finally, a voice will be generated by MBROLA application from this series of phonemes and prosody.</p><p>In this research some tools are used e.g. a Praat and an MBROLA applications. The Praat application is used to get <inline-formula><inline-graphic xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\a3af340e-9e86-4165-b97f-a5d8ebe635a9.png" xlink:type="simple"/></inline-formula> from voice files while MBROLA application used to generate voice from phonemes + prosody form. MBROLA is not a speech synthesis complete software. MBROLA requires input of phonemes and prosody in a form that matches to the MBROLA form called pho. That is why the text to be synthesized must be converted first into MBROLA form (pho). MBROLA made by TCTS Lab of Facult? &#233; Polytechnique de Mons Belgium [<xref ref-type="bibr" rid="scirp.48303-ref12">12</xref>] .</p><p><xref ref-type="table" rid="table1">Table 1</xref> shows the result of normalization module. Symbols, numbers, and abbreviation converted correctly. The punctuation marks: comma<inline-formula><inline-graphic xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\4223c60e-d0e1-404c-8d7f-32e994fc3728.png" xlink:type="simple"/></inline-formula>, period<inline-formula><inline-graphic xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\b5521869-dcba-4dac-83c8-7af34bfc6ecd.png" xlink:type="simple"/></inline-formula>, and hyphenation <inline-formula><inline-graphic xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\e8a27714-0029-43d1-849b-e4d833865a63.png" xlink:type="simple"/></inline-formula> are still unchanged.</p><p><inline-formula><inline-graphic xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\8c37cc59-48c6-44c2-9132-e1a85e4a6a1b.png" xlink:type="simple"/></inline-formula>of sentence “Empat LSM lembaga swadaya masyarakat diantaranya ICW indonesia corruption watch PSHK pusat studi hukum dan kebijakan menyampaikan aspirasi kepada panitia ad hock PAH empat DPD di</p><fig id="fig2"><label>Figure 2</label><caption><p> Relation between selector pattern and syntax analyst (sentence structure) modules. Top: result of intonation pattern, bottom: result of syntax analyst</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\00e9c394-92ff-4985-85c0-532daa0a0a0b.png"/></fig><p>komplek parlemen senayan Jakarta hari ini” is shown in <xref ref-type="table" rid="table2">Table 2</xref> but there are only partial results. A full result contained about 1200 data. The <inline-formula><inline-graphic xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\e01f0d97-770c-4bad-899c-1d168d7dd13b.png" xlink:type="simple"/></inline-formula> table obtained by Praat application [<xref ref-type="bibr" rid="scirp.48303-ref7">7</xref>] [<xref ref-type="bibr" rid="scirp.48303-ref13">13</xref>] .</p><p>If a sentence “DPD kemudian tetap melakukan uji kepatutan dan kelayakan” fed into the syntax analyst module, a sentence structure diagram obtained as shown in <xref ref-type="fig" rid="fig3">Figure 3</xref>.</p><p>Voice synthesis module in this study is part of the pho file processing module according to the rules of normalization followed the general tone combined with syntax analysts and intonation patterns. For comparison, the results will be presented waveform synthesis with flat intonation, intonation with the general rule, and intonation pattern.</p><table-wrap id="table1"  position="float"><object-id pub-id-type="pii">Table 1</object-id><label>Table 1</label><caption><p>. Result of normalization module for complete sentences</p></caption><table><thead><tr><th align="center" valign="middle" >Before</th><th align="center" valign="middle" >After</th></tr></thead><tbody><tr><td align="center" valign="middle" >Pemerintah menyatakan ekspor sepanjang tahun 2013 secara kumulatif melampaui target.</td><td align="center" valign="middle" >pemerintah menyatakan ekspor sepanjang tahun dua ribu tiga belas secara kumulatif melampaui target.</td></tr><tr><td align="center" valign="middle" >“Target ekspor 2013 itu 179 miliar dollar AS, dengan angka Desember, kumulatif kita mencapai 182,57 miliar dollar AS,” kata Wakil Menteri Perdagangan Bayu Krisnamurthi, Senin</td><td align="center" valign="middle" >“target ekspor dua ribu tiga belas itu seratus tujuh puluh sembilan miliar dollar a es, dengan angka desember, kumulatif kita mencapai seratus delapan puluh dua koma lima tujuh miliar dollar a es,” kata wakil menteri perdagangan bayu krisnamurthi, senin</td></tr><tr><td align="center" valign="middle" >Hal itu berdasarkan pengumuman Badan Pusat Statistik (BPS).</td><td align="center" valign="middle" >hal itu berdasarkan pengumuman badan pusat statistik be pe es</td></tr><tr><td align="center" valign="middle" >KJRI terus memantau kondisi kesehatan ybs., dengan melakukan pemeriksaan medis secara rutin ke rumah sakit.</td><td align="center" valign="middle" >ka je er i terus memantau kondisi kesehatan yang bersangkutan, dengan melakukan pemeriksaan medis secara rutin ke rumah sakit.</td></tr><tr><td align="center" valign="middle" >Berdasarkan hasil pemeriksaan, saat ini Sdri. Kokom dalam kondisi sehat, dan luka-lukanya telah berangsur-angsur pulih serta telah dapat berjalan secara normal, meskipun masih mengeluhkan nyeri dibagian kaki kirinya.</td><td align="center" valign="middle" >berdasarkan hasil pemeriksaan, saat ini Saudari kokom dalam kondisi sehat, dan luka-lukanya telah berangsur-angsur pulih serta telah dapat berjalan secara normal, meskipun masih mengeluhkan nyeri dibagian kaki kirinya.</td></tr></tbody></table></table-wrap><table-wrap id="table2"  position="float"><object-id pub-id-type="pii">Table 2</object-id><label>Table 2</label><caption><p>. Excerpts F<sub>0</sub> results for sentence</p></caption><table><thead><tr><th align="center" valign="middle" >Time (s)</th><th align="center" valign="middle" ><img src="htmlimages\2-3400350x\381bedee-a0fc-4f10-bd1e-6b235cdb7475.png" width="28.75" height="35" />(Hz)</th></tr></thead><tbody><tr><td align="center" valign="middle" >0.622125</td><td align="center" valign="middle" >138.447701</td></tr><tr><td align="center" valign="middle" >0.632125</td><td align="center" valign="middle" >134.561068</td></tr><tr><td align="center" valign="middle" >0.642125</td><td align="center" valign="middle" >--undefined--</td></tr><tr><td align="center" valign="middle" >0.652125</td><td align="center" valign="middle" >--undefined--</td></tr><tr><td align="center" valign="middle" >0.662125</td><td align="center" valign="middle" >--undefined--</td></tr><tr><td align="center" valign="middle" >0.672125</td><td align="center" valign="middle" >--undefined--</td></tr><tr><td align="center" valign="middle" >0.682125</td><td align="center" valign="middle" >--undefined--</td></tr><tr><td align="center" valign="middle" >0.692125</td><td align="center" valign="middle" >--undefined--</td></tr><tr><td align="center" valign="middle" >0.702125</td><td align="center" valign="middle" >--undefined--</td></tr><tr><td align="center" valign="middle" >0.712125</td><td align="center" valign="middle" >--undefined--</td></tr><tr><td align="center" valign="middle" >0.722125</td><td align="center" valign="middle" >151.699538</td></tr><tr><td align="center" valign="middle" >0.732125</td><td align="center" valign="middle" >149.879256</td></tr><tr><td align="center" valign="middle" >0.742125</td><td align="center" valign="middle" >149.277178</td></tr><tr><td align="center" valign="middle" >0.752125</td><td align="center" valign="middle" >149.884325</td></tr><tr><td align="center" valign="middle" >0.762125</td><td align="center" valign="middle" >151.686222</td></tr><tr><td align="center" valign="middle" >0.772125</td><td align="center" valign="middle" >153.601689</td></tr><tr><td align="center" valign="middle" >0.782125</td><td align="center" valign="middle" >154.488904</td></tr><tr><td align="center" valign="middle" >0.792125</td><td align="center" valign="middle" >152.781495</td></tr></tbody></table></table-wrap><fig id="fig3"><label>Figure 3</label><caption><p> Example of sentence structure</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\3e33fd62-d20e-4608-b512-088280b811f2.png"/></fig><p>Firstly, the sentence “Wimar tidak dapat mengikuti pemakaman Gus Dur di Jombang” is synthesized with flat intonation. The results of the graph <inline-formula><inline-graphic xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\2f6c23dd-7912-452e-8181-7d8228bdce87.png" xlink:type="simple"/></inline-formula> are shown in <xref ref-type="fig" rid="fig4">Figure 4</xref>. Secondly, for the same sentence but intonation synthesized according to the general rules, the results of the graph <inline-formula><inline-graphic xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\229147ab-3a01-4224-93c3-3d225572a0f8.png" xlink:type="simple"/></inline-formula> are shown in <xref ref-type="fig" rid="fig5">Figure 5</xref>. And, <xref ref-type="fig" rid="fig6">Figure 6</xref> shows the synthesis of the same as the second but improved by incorporating elements of intonation pattern. The selected intonation pattern has <inline-formula><inline-graphic xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\ee5e205f-fa63-4011-ba30-25a43f6148b0.png" xlink:type="simple"/></inline-formula> as shown in <xref ref-type="fig" rid="fig7">Figure 7</xref>.</p><p>Similarity checking beetween synthesis result and original voice can be done by quantity approach. In this study, the checking similarity uses the PESQ method [<xref ref-type="bibr" rid="scirp.48303-ref14">14</xref>] [<xref ref-type="bibr" rid="scirp.48303-ref15">15</xref>] . The result of PESQ for several sentences is shown in <xref ref-type="table" rid="table3">Table 3</xref>.</p></sec><sec id="s4"><title>4. Discussion</title><table-wrap id="table3"  position="float"><object-id pub-id-type="pii">Table 1</object-id><label>Table 1 shows that the integer number of year (2013) and the amount of money (179) can be converted correctly</label><caption><p>. Likewise fractions (182,57) can be converted to text correctly. Because 57 is a fraction, so it is converted into a “lima tujuh” (five seven) instead of “lima puluh tujuh” (fifty seven). The comma in 182,57 is decimal separator instead of thousand separator. Abbreviations and acronyms can be converted correctly. KJRI, BPS, and AS are converted to “ka je er i”, “be pe es”, and “a es”. The “ybs.” and “Sdri.” acronyms are converted to “yang bersangkutan” and “Saudari”</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink"  xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\4db4a8f5-cced-4308-8297-8cf8cde55317.png"/></table-wrap><p>A flat intonation was obtained from MBROLA when the frequency was set in one value for a long of speech</p><fig id="fig4"><label>Figure 4</label><caption><p> Graph F<sub>0</sub> with flat intonation</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\3e493e96-434b-4bad-bb90-05bf516dc177.png"/></fig><fig id="fig5"><label>Figure 5</label><caption><p> Graph F<sub>0</sub> with intonation according to general rules</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\99329b8f-6cb0-4541-b876-11689de3939a.png"/></fig><fig id="fig6"><label>Figure 6</label><caption><p> Graph F<sub>0</sub> with the general rules and intonation patterns</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\dcceedac-66ee-4cf6-bca8-fa7742bbc7f3.png"/></fig><p>synthesis. <xref ref-type="fig" rid="fig4">Figure 4</xref> shows the result of flat intonation graph when the frequency set to 115 Hz. Although several small peaks were shown in a flat intonation, it happened because of MBROLA characteristic.</p><p>In gerenal rules of intonation according to [<xref ref-type="bibr" rid="scirp.48303-ref5">5</xref>] the peaks of <inline-formula><inline-graphic xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\fd8e514e-ff51-4fc6-85c8-b82011c11dd4.png" xlink:type="simple"/></inline-formula> appeared at the syllables of nouns. For nouns before a verb <inline-formula><inline-graphic xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\65915c6d-7437-4447-aefd-31f5a902c3e5.png" xlink:type="simple"/></inline-formula> peak occur in a syllable before the last syllable called the penultimate. In a sentence “Wimar tidak bisa mengikuti pemakaman Gus Dur di Jombang” “Wimar” is a noun before the verb “mengikuti”. In that case <inline-formula><inline-graphic xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\85c24c1f-6823-4b74-a4c0-7d4ea15d94df.png" xlink:type="simple"/></inline-formula> peak occur in syllable “Wi” from the word “Wimar”. That is in the start of sentence. See the first seconds of graph in <xref ref-type="fig" rid="fig4">Figure 4</xref>. For nouns after a verb <inline-formula><inline-graphic xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\86eb5c33-a31f-429b-9e94-3953774da5e5.png" xlink:type="simple"/></inline-formula> peak occur in the first syllable of the first noun. In the sentence the first syllable of the first noun is “Gus”. This syllable is the <inline-formula><inline-graphic xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\f5e03422-738f-4d71-94e0-7b085768484a.png" xlink:type="simple"/></inline-formula> syllable of 19 syllables in the sentence. We can calculate that this syllable occur in about</p><disp-formula id="scirp.48303-formula4338"><label>(9)</label><inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\3de32eab-79c2-4cf7-a27d-f524a3516aac.png"/></disp-formula><fig id="fig7"><label>Figure 7</label><caption><p> Graph <img src="htmlimages\2-3400350x\7c40e1f6-ac86-44ed-9b95-cd6e5b2f6e1c.png" width="30" height="35" /> of selected pattern</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\133322f3-aa2f-43e5-a5b6-090f7d909220.png"/></fig><table-wrap id="table4"  position="float"><object-id pub-id-type="pii">Table 4</object-id><label>Table 3</label><caption><p>. Values of PESQ for several sentences</p></caption><table><thead><tr><th align="center" valign="middle" >Reference</th><th align="center" valign="middle" >Synthesis</th><th align="center" valign="middle" >PESQ</th></tr></thead><tbody><tr><td align="center" valign="middle" >kal101.wav</td><td align="center" valign="middle" >kal101-sin.wav</td><td align="center" valign="middle" >1.272</td></tr><tr><td align="center" valign="middle" >kal102a.wav</td><td align="center" valign="middle" >Kal102a-sin.wav</td><td align="center" valign="middle" >1.124</td></tr><tr><td align="center" valign="middle" >kal103a.wav</td><td align="center" valign="middle" >kal103a-sin.wav</td><td align="center" valign="middle" >1.067</td></tr><tr><td align="center" valign="middle" >kal104.wav</td><td align="center" valign="middle" >kal104a-sin.wav</td><td align="center" valign="middle" >1.048</td></tr><tr><td align="center" valign="middle" >kal105.wav</td><td align="center" valign="middle" >kal105-sin.wav</td><td align="center" valign="middle" >1.062</td></tr><tr><td align="center" valign="middle" >kal106.wav</td><td align="center" valign="middle" >kal106-sin.wav</td><td align="center" valign="middle" >1.037</td></tr><tr><td align="center" valign="middle" >kal107.wav</td><td align="center" valign="middle" >kal107-sin.wav</td><td align="center" valign="middle" >1.064</td></tr></tbody></table></table-wrap><p>This peak can be seen at the fourth second of the graph in <xref ref-type="fig" rid="fig5">Figure 5</xref>. After that the values of <inline-formula><inline-graphic xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\9a4bc9a2-e3f2-4d6c-8514-8d038868cbe0.png" xlink:type="simple"/></inline-formula> continues to fall until the end of the sentence.</p><p>The pattern has <inline-formula><inline-graphic xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\c43f9733-0a30-4a45-98e5-5a49383ce270.png" xlink:type="simple"/></inline-formula> graph as shown in <xref ref-type="fig" rid="fig7">Figure 7</xref>. It has two peaks that occur at <inline-formula><inline-graphic xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\6efd802b-35c7-4ef9-8717-a3bafb82274b.png" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\c2cc9250-4dc4-4e5f-9fdd-fc521bd6c119.png" xlink:type="simple"/></inline-formula> second from 17.5 s speech synthesis a long. We can calculate the segments of its value</p><disp-formula id="scirp.48303-formula4339"><label>(10)</label><inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\27cd2b96-058b-4a62-a5bc-ee0c4db436b2.png"/></disp-formula><disp-formula id="scirp.48303-formula4340"><label>(11)</label><inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\3449a296-dcd1-49a0-a84b-9be1eac49da9.png"/></disp-formula><p>This selected pattern then combined to speech voice that has <inline-formula><inline-graphic xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\42112c2d-7127-47be-b6cc-8277191523d1.png" xlink:type="simple"/></inline-formula> as <xref ref-type="fig" rid="fig5">Figure 5</xref> obtained a speech voice that has <inline-formula><inline-graphic xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\64148bdd-9399-4754-8f39-b1164d040bc2.png" xlink:type="simple"/></inline-formula> as <xref ref-type="fig" rid="fig6">Figure 6</xref>. The <inline-formula><inline-graphic xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\da5455fe-e65d-4640-80a9-ec3e8ca5555c.png" xlink:type="simple"/></inline-formula> peaks of the syntesis result can be predict from Equation (9) occur at</p><disp-formula id="scirp.48303-formula4341"><inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\fb166359-0b23-4842-a12d-9969b5f86226.png"/></disp-formula><p>This results is not match with the graph second peak in <xref ref-type="fig" rid="fig6">Figure 6</xref>. There is no peak in this value. If we used Equation (10) and (11) we get</p><disp-formula id="scirp.48303-formula4342"><inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\7b308620-1f84-4d7f-9e3c-50fe3e8551b4.png"/></disp-formula><p>and</p><disp-formula id="scirp.48303-formula4343"><inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\e19a6261-9094-4d82-89aa-083ed4c60484.png"/></disp-formula><p>There are peaks that match with peaks in <xref ref-type="fig" rid="fig6">Figure 6</xref>. The second peak can take a place the wrong result as Equation (9). Thus <xref ref-type="fig" rid="fig6">Figure 6</xref> is combination of <xref ref-type="fig" rid="fig5">Figure 5</xref> and <xref ref-type="fig" rid="fig7">Figure 7</xref>.</p><p>From <xref ref-type="table" rid="table3">Table 3</xref>, we get that the average value of PESQ is 1.096. This value can be converted to MOS-LQO scale by the following formula [<xref ref-type="bibr" rid="scirp.48303-ref14">14</xref>]</p><disp-formula id="scirp.48303-formula4344"><inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\f34e2560-f56a-4e65-abad-e73928a635c9.png"/></disp-formula><p>where <inline-formula><inline-graphic xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\1a1b1779-6bcd-46c7-b4a9-c2de97f4f6cc.png" xlink:type="simple"/></inline-formula> is MOS-LQO in the range 1.02 to 4.56, <inline-formula><inline-graphic xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\251bce47-e30f-4713-9140-9ef02feebb8f.png" xlink:type="simple"/></inline-formula>is PESQ value in the range <inline-formula><inline-graphic xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\5da23077-191d-4a81-9088-ab8e3ac35312.png" xlink:type="simple"/></inline-formula> to 4.5. The value of MOS-LQO is</p><disp-formula id="scirp.48303-formula4345"><inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\628ce75c-7d09-4112-abf7-c26c3c2795fb.png"/></disp-formula></sec><sec id="s5"><title>5. Conclusion</title><p>Speech synthesis model with a combination of general rules and intonation patterns of the fundamental frequency <inline-formula><inline-graphic xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\a9db7abc-5043-4a08-ba26-1c67002d8fba.png" xlink:type="simple"/></inline-formula> can be implemented. The peaks of <inline-formula><inline-graphic xlink:href="http://file.scirp.org/Html/htmlimages\2-3400350x\3f4c92a3-2184-4211-b81b-43267a576593.png" xlink:type="simple"/></inline-formula> determined by general rules or intonation patterns are dominant. Indonesian has intonation general rules, but should not be adhered strictly. Similarity test with the PESQ method shows that the synthesis results are about 1.18 at MOS-LQO scale. A speech synthesis based on the announcer voice intonation patterns is expected to be used as an alternative method of Indonesian speech synthesis refinement.</p></sec></body><back><ref-list><title>References</title><ref id="scirp.48303-ref1"><label>1</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>SCHRODER</surname><given-names> M. </given-names></name>,<etal>et al</etal>. (<year>2001</year>)<article-title>EMOTIONAL SPEECH SYNTHESIS: A REVIEW</article-title><source> PROCEEDINGS OF EUROSPEECH 2001</source><volume> 1</volume>,<fpage> 561</fpage>-<lpage>564</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.48303-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">VROOMEN, J., COLLIER, R. AND MOZZICONACCI, S. (1993) DURATION AND INTONATION IN EMOTIONAL SPEECH. PROCEEDINGS OF THE THIRD EUROPEAN CONFERENCE ON SPEECH COMMUNICATION AND TECHNOLOGY, BERLIN, 22-25 SEPTEMBER 1993, 577-580.</mixed-citation></ref><ref id="scirp.48303-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">CAMPBELL, W.N., ISARD, S., MONAGHAN, A.L.C. AND VERHOEVEN, J. (1990) DURATION, PITCH AND DIPHONES IN THE CSTR TTS SYSTEM. PROCEEDINGS OF THE INTERNATIONAL CONFERENCE ON SPOKEN LANGUAGE PROCESSING, KOBE, 1 JANUARY 1990, 825-828.</mixed-citation></ref><ref id="scirp.48303-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">TRITOASMORO, I.I. (2006) TEXT-TO-SPEECH BAHASA INDONESIA MENGGUNAKAN CONCATENATION SYNTHESIZER BERBASIS FONEM. SEMINAR NASIONAL SISTEM DAN INFORMATIKA, BALI, 17 NOVEMBER 2006, 171-176.</mixed-citation></ref><ref id="scirp.48303-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">LAKSMAN, M. (1995) REALISASI TEKANAN KATA DALAM BAHASA INDONESIA. PELLBA 8, PAGES 179{215. LEMBAGA BAHASA UNIKA ATMAJAYA, JAKARTA, 1995.</mixed-citation></ref><ref id="scirp.48303-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">HALIM, A. (1975) INTONATION IN RELATION TO SYNTAX IN BAHASA INDONESIA. PROYEK PENGEMBANGAN BAHASA DAN SASTRA INDONESIA DAN DAERAH, DEPARTEMEN PENDIDIKAN DAN KEBUDAYAAN.</mixed-citation></ref><ref id="scirp.48303-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">VAN LIESHOUT, P.P. (2003) PRAAT SHORT TUTORIAL. UNIVERSITY OF TORONTO, GRADUATE DEPARTMENT OF SPEECH-LANGUAGE PATHOLOGY, FACULTY OF MEDICINE, ORAL DYNAMICS LAB (ODL), TORONTO.</mixed-citation></ref><ref id="scirp.48303-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">HEUVEN, V.J.V. AND ZANTEN, E.V. (2007) PROSODY IN INDONESIAN LANGUAGES. LOT, UTRECHT.</mixed-citation></ref><ref id="scirp.48303-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">SAKRI, A. (1994) BANGUN KALIMAT BAHASA INDONESIA. 2ND EDITION, PENERBIT ITB, BANDUNG.</mixed-citation></ref><ref id="scirp.48303-ref10"><label>10</label><mixed-citation publication-type="other" xlink:type="simple">MCCUNE, K.M. (1985) THE INTERNAL STRUCTURE OF INDONESIAN ROOTS. NUMBER V. 2 IN THE INTERNAL STRUCTURE OF INDONESIAN ROOTS. BADAN PENYELENGGARA SERI NUSA, UNIVERSITAS KATOLIK INDONESIA ATMA JAYA, JAKARTA.</mixed-citation></ref><ref id="scirp.48303-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">VAMARASI, M.K. (1986) GRAMMATICAL RELATIONS IN BAHASA INDONESIA. CORNELL UNIVERSITY, ITHACA.</mixed-citation></ref><ref id="scirp.48303-ref12"><label>12</label><mixed-citation publication-type="other" xlink:type="simple">MBROLA, T. (2009) THE MBROLA HOME PAGE. HTTP://TCTS.FPMS.AC.BE/SYNTHESIS/MBROLA/</mixed-citation></ref><ref id="scirp.48303-ref13"><label>13</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>BORIL</surname><given-names> H. </given-names></name>,<name name-style="western"><surname> POLLÁK</surname><given-names> P. </given-names></name>,<etal>et al</etal>. (<year>2004</year>)<article-title>BORIL, H. AND POLLÁK, P.  DIRECT TIME DOMAIN FUNDAMENTAL FREQUENCY ESTIMATION OF SPEECH IN NOISY CONDITIONS</article-title><source> EUSIPCO 2004</source><volume> 2</volume>,<fpage> 1003</fpage>-<lpage>1006</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.48303-ref14"><label>14</label><mixed-citation publication-type="other" xlink:type="simple">ITU (2001) ITU-T RECOMMENDATION P.862: PERCEPTUAL EVALUATION OF SPEECH QUALITY (PESQ): AN OBJECTIVE METHOD FOR END-TO-END SPEECH QUALITY ASSESSMENT OF NARROW-BAND TELEPHONE NETWORKS AND SPEECH CODECS. TECHNICAL REPORT, ITU.</mixed-citation></ref><ref id="scirp.48303-ref15"><label>15</label><mixed-citation publication-type="other" xlink:type="simple">RIX, A.W., BEERENDS, J.G., HOLLIER, M.P. AND HEKSTRA, A.P. (2001) PERCEPTUAL EVALUATION OF SPEECH QUALITY (PESQ)— A NEW METHOD FOR SPEECH QUALITY ASSESSMENT OF TELEPHONE NETWORKS AND CODECS. IEEE INTERNATIONAL CONFERENCE ON ACOUSTICS, SPEECH, AND SIGNAL PROCESSING, SALT LAKE CITY, 7-11 MAY 2001, 749-752.</mixed-citation></ref></ref-list></back></article>