{"id":935,"date":"2012-07-27T16:47:38","date_gmt":"2012-07-27T13:47:38","guid":{"rendered":"http:\/\/ssrlab.by\/?p=935"},"modified":"2018-11-21T14:20:55","modified_gmt":"2018-11-21T11:20:55","slug":"proverka-sinteza-sluhom","status":"publish","type":"post","link":"https:\/\/ssrlab.by\/en\/935","title":{"rendered":"SB: \u201cSynthesis check by listening\u201d"},"content":{"rendered":"<p>Download (PDF, 242KB)<\/p>\n<div><\/div>\n<p style=\"text-align: justify;\"><img loading=\"lazy\" title=\"\u042e\u0440\u0438\u0439 \u0413\u0435\u0446\u0435\u0432\u0438\u0447 \u00ab\u0443\u0447\u0438\u0442\u00bb \u043a\u043e\u043c\u043f\u044c\u044e\u0442\u0435\u0440 \u0433\u043e\u0432\u043e\u0440\u0438\u0442\u044c.\" src=\"http:\/\/www.sb.by\/images\/articles\/12\/138\/komp2.jpg\" alt=\"\" width=\"402\" height=\"268\" align=\"left\" border=\"0\" hspace=\"6\" vspace=\"3\" \/><span style=\"font-weight: 400;\">To deceive the eyes is easier than ears. When Jura\u015b Hiecevi\u010d enters the sentence\u00a0<\/span><span style=\"font-weight: 400;\">\u00abMother washed dishes\u00bb<\/span><span style=\"font-weight: 400;\"> on the keyboard, &#8220;the talking head&#8221; immediately announces it on the monitor. The vision shows that facial expressions on the other side of the screen are absolutely reliable. But the ears still doubt the intonation. <\/span><span style=\"font-weight: 400;\">The differences between computer and natural speech are more evident while playing a large piece of text \u2013 there is not enough emotion. At first, it is difficult to understand anything, but when you set your mind to, the essence of what has been said is easily perceived.<\/span><\/p>\n<p style=\"text-align: justify;\"><span style=\"font-weight: 400;\">Nevertheless, the work on the creation of perfect, natural speech is one of the main problems to be solved today for scientists under the direction of Acting Head of speech synthesis and recognition laboratory of\u00a0<\/span><span style=\"font-weight: 400;\">United Institute of Informatics Problems, National Academy of Sciences of Belarus\u00a0<\/span><span style=\"font-weight: 400;\">Jura\u015b Hiecevi\u010d.<\/span><\/p>\n<p style=\"text-align: justify;\"><b>Intonation is good, but not perfect<\/b><\/p>\n<p style=\"text-align: justify;\"><span style=\"font-weight: 400;\">It would seem that talking computer and a mobile phone are not the latest things. Everyone can get programs which voice the text. However, commercial companies, adapting the idea of speech synthesis for applications in specific areas, don\u2019t get ahead of themselves. They use only tested scientific results.<\/span><\/p>\n<p style=\"text-align: justify;\"><span style=\"font-weight: 400;\">Jura\u015b Hiecevi\u010d notes that all synthesizers on the market are deficient of intonation \u2013 long sentences are pronounced, so that after one or two pages of text it is impossible to listen to. They have a relatively small number of intonation contours and rules for their application and therefore, inaccuracy is appeared in a voice variant.\u00a0<\/span><span style=\"font-weight: 400;\">As a rule, companies are waiting for new scientific publications and only then take the latest developments into service. Thus, for example, it happened with numbers processing and pronunciation in the text, which was rarity a few years ago. But today it is a common option.<\/span><\/p>\n<p style=\"text-align: justify;\"><span style=\"font-weight: 400;\">Text-to-speech synthesis is one of the most difficult areas with a great deal of work on it. For example, Jura\u015b Hiecevi\u010d is working out the mechanism so that the machine can understand and correctly read aloud the different combinations of numbers and letters, abbreviations, acronyms, automatically put the stress in new (unknown to the synthesizer) words.\u00a0<\/span><span style=\"font-weight: 400;\">Because not all people adhere to the rules for writing. His thesis is devoted to the linguistic processing of text-to-speech synthesizer: &#8220;We even can\u2019t put the stress to unknown surnames. How to teach the machine to look for such decisions? There is even more interesting task: homographs.\u00a0<\/span><span style=\"font-weight: 400;\">There are not so many homographs In Russian and Belarusian, about 10 thousand, but they spoil the picture! How a computer will understand correct word &#8220;<\/span><span style=\"font-weight: 400;\">\u043f\u0440\u0438\u043e\u0431\u0440\u0435\u0442\u0430\u0435\u0442 \u0432\u0441\u0435 \u0431\u041e\u043b\u044c\u0448\u0443\u044e \u043f\u043e\u043f\u0443\u043b\u044f\u0440\u043d\u043e\u0441\u0442\u044c \u0438\u043b\u0438 \u0431\u043e\u043b\u044c\u0448\u0423\u044e&#8221;?\u00a0<\/span><span style=\"font-weight: 400;\">I know triple homographs in Belarusian. For example,\u00a0<\/span><span style=\"font-weight: 400;\">\u00ab\u043f\u0440\u044b\u0433\u043e\u0436\u0430\u044f \u043a\u0430\u0437\u0430\u0447\u043a\u0430 \u0440\u0430\u0441\u043f\u0430\u0432\u044f\u043b\u0430 \u043a\u0430\u0437\u0430\u0447\u043a\u0443 \u0441\u0432\u0430\u0439\u043c\u0443 \u043a\u0430\u0437\u0430\u0447\u043a\u0443\u00bb<\/span><span style=\"font-weight: 400;\">&#8230; We have a system that looks for homographs, but sometimes we still face the fact that the machine is not able to perceive the meaning, context&#8221;. That is why the improvement of speech synthesis is a problem\u00a0of\u00a0the same level as the creation of artificial intelligence.<\/span><\/p>\n<p style=\"text-align: justify;\"><b>Word composition<\/b><\/p>\n<p style=\"text-align: justify;\"><span style=\"font-weight: 400;\">Since the youth innovation forum of the National Academy of Sciences where the project of Jura\u015b Hiecevi\u010d and Dzmicier Pakladok \u201cText-to-speech synthesizer for Russian and Belarusian for stationary and mobile platforms\u201d was considered the best, a lot of things happened. It was presented at a conference on artificial intelligence OSTIS-2012, participated in the innovation week, received a diploma at &#8220;TIBO-2012&#8221;. They are invited to exhibitions constantlY.\u00a0<\/span>After all, these young scientists have learned computer and mobile phone to speak Belarusian. Previously, there were no Belarusian synthesizers at all!<\/p>\n<p style=\"text-align: justify;\"><span style=\"font-weight: 400;\">To make a computer speak is a huge, painstaking work. You need to record the voice of a real person, then this record is decomposed in a special program which shows the smallest fluctuations in sound, cut into &#8220;details&#8221; \u2013 allophones (the smallest variations of phonemes) \u00a0because the same letter &#8220;a&#8221; in stressed and unstressed syllables is pronounced differently. As a result, the base contains thousands of allophones.\u00a0<\/span><span style=\"font-weight: 400;\">And then the algorithms are developed. They remove parts of words from the base, which are necessary to play, then join smallest parts to the word. It is important to note that the speaker doesn\u2019t have to record a large text. Scientists have developed a special well-balanced text for six minutes readings, which has all the necessary phonemes.<\/span><\/p>\n<p style=\"text-align: justify;\">Of course, the software that translates text files into sound files, should have a comprehensive dictionary and its replenishment system \u2013 the brainchild of Jura\u015b Hiecevi\u010d, operates more than two million words of the Russian and the Belarusian languages.<\/p>\n<p style=\"text-align: justify;\"><b>You don&#8217;t need iPhones<\/b><\/p>\n<p style=\"text-align: justify;\"><span style=\"font-weight: 400;\">40 years ago the first in the eastern European space who began to teach computers to pronounce typed text was Barys Labana\u016d, chief researcher at the speech synthesis and recognition Laboratory of\u00a0<\/span><span style=\"font-weight: 400;\">United Institute of Informatics Problems<\/span><span style=\"font-weight: 400;\">. He created the basis on which the speech synthesis is improved now in Belarus and even in Russia \u2013 by the way, for the most part by students of Barys Miefodzjevi\u010d. Jura\u015b Hiecevi\u010d is one of them. He pulls out an old mobile phone with the words: &#8220;I keep it for all not to think that our programs need iPhones. <\/span><span style=\"font-weight: 400;\">This is an experimental model of mobile speech-to-text synthesizer, made in our laboratory. It requires only two\u00a0megabytes of memory and therefore can work on the basic devices&#8221;. And a synthesized voice starts reading &#8220;Zorka Vieniera&#8221;. It can also voice text messages and caller&#8217;s name. You need only text!<\/span><\/p>\n<p style=\"text-align: justify;\"><span style=\"font-weight: 400;\">40 years ago the first in the eastern European space who began to teach computers to pronounce typed text was Barys Labana\u016d, chief researcher at the speech synthesis and recognition Laboratory of <\/span><span style=\"font-weight: 400;\">United Institute of Informatics Problems<\/span><span style=\"font-weight: 400;\">. He created the basis on which the speech synthesis is improved now in Belarus and even in Russia \u2013 by the way, for the most part by students of Barys Miefodzjevi\u010d. Jura\u015b Hiecevi\u010d is one of them. He pulls out an old mobile phone with the words: &#8220;I keep it for all not to think that our programs need iPhones.\u00a0<\/span><span style=\"font-weight: 400;\">This is an experimental model of mobile speech-to-text synthesizer, made in our laboratory. It requires only two megabytes of memory and therefore can work on the basic devices&#8221;. And a synthesized voice starts reading &#8220;Star Venera&#8221;. It can also voice text messages and caller&#8217;s name. You need only text!<\/span><\/p>\n<p style=\"text-align: justify;\"><span style=\"font-weight: 400;\">A computer system which creates audiobooks is also developed. Recently with students we voiced the textbook &#8220;Social Science&#8221; for 10th grade. It took only about a week. Students said that this was a really good practice that they have never had before.\u00a0<\/span><span style=\"font-weight: 400;\">&#8220;Talking Library&#8221; already exists. For example, it works in Molodechno school for children with visual impairments. In general, for those who have problems with vision, speech synthesis program is a godsend. Braille Books are expensive, not to mention the fact that you will not find the literary novelties among them.\u00a0<\/span><span style=\"font-weight: 400;\">Our program will transfer text into an audio version any product, the electronic version of which is in the network. Created program also is useful for those who need to learn how to speak, for example, after a stroke: &#8220;talking head&#8221; pronounces the word on the monitor, the mimics can be played even in slow motion with the possibility to imitate her.\u00a0<\/span><span style=\"font-weight: 400;\">And the latest development makes the synthesizer applicable for alerting services: it is enough to enter necessary information, and the voice will announce when and on which line the train is coming or what the next stop is at the trolley. Or one new product is a phone robot. It calls dozens of numbers of subscribers and reports on the debt, indicating the specific amount, as long as the data are\u00a0in the computer.<\/span><\/p>\n<p style=\"text-align: justify;\">The nearest plan of scientists is the creation of the Internet version of the text-to-speech synthesizer. It is likely to be that the first &#8220;talk&#8221; website will be the site of the National Library. Then, any visitor will be able to use voice search of the book. The entire text which gets over &#8220;mouse&#8221;, will be voiced, \u00a0whether it columns, tabs sections. In general, it is a great amount of practical applications which allow to use text-to-speech synthesizer for the education, rehabilitation, in the banking system, transport, housing and communal services. The matter depends on potential customers to point out all scientific achievements and estimate their benefits.<\/p>\n<p style=\"text-align: justify;\"><strong>You say, &#8220;the locomotive&#8221; the machine writes &#8220;milk&#8221;<\/strong><\/p>\n<p style=\"text-align: justify;\"><strong><span style=\"font-weight: 400;\">But the situation with the dream of writers and journalists is more complicated. It is a computer that would perceive voice and transfer it into text so that you can read poems and articles, pacing around the room. These programs are sold and advertised, but none of them is able to replace typing.\u00a0<\/span><span style=\"font-weight: 400;\">As a rule, more or less, they define only the voices of their creators. The common situation is like this: you say, &#8220;the locomotive&#8221; and machine writes &#8220;milk&#8221;. Jura\u015b Hiecevi\u010d explains the low efficiency of such programs: \u201cit is very difficult to single out the words from the speech stream and at the same time not to confuse anything\u201d. However, our scientists are looking for solutions.<\/span><\/strong><\/p>\n<p style=\"text-align: right;\"><span style=\"font-weight: 400;\">Authors: <a href=\"http:\/\/www.sb.by\/blog\/154194-yuliya-vasilishina\/\">Julija Vasili\u0161yna<\/a><\/span><\/p>\n<p style=\"text-align: right;\"><span style=\"font-weight: 400;\">Photo: Vitalij Hi\u013a<\/span><\/p>\n<p style=\"text-align: right;\"><span style=\"font-weight: 400;\">Date of publication: <\/span><span style=\"font-weight: 400;\">27.07.2012<\/span><\/p>\n<div>\n<div style=\"text-align: left;\">Source: <a href=\"http:\/\/www.sb.by\/print\/post\/134264\/\">http:\/\/www.sb.by\/print\/post\/134264\/<\/a><\/div>\n<div style=\"text-align: left;\"><\/div>\n<div style=\"text-align: left;\"><iframe src=\"\/\/docs.google.com\/viewer?url=http%3A%2F%2Fssrlab.by%2Fwp-content%2Fuploads%2F2014%2F02%2F%D0%9F%D0%BE%D1%80%D1%82%D0%B0%D0%BB-%D0%91%D0%B5%D0%BB%D0%B0%D1%80%D1%83%D1%81%D1%8C-%D0%A1%D0%B5%D0%B3%D0%BE%D0%B4%D0%BD%D1%8F-_-%D0%9E%D0%B1%D1%89%D0%B5%D1%81%D1%82%D0%B2%D0%BE-%D0%9F%D1%80%D0%BE%D0%B2%D0%B5%D1%80%D0%BA%D0%B0-%D1%81%D0%B8%D0%BD%D1%82%D0%B5%D0%B7%D0%B0-%D1%81%D0%BB%D1%83%D1%85%D0%BE%D0%BC.pdf&hl=en_US&embedded=true\" class=\"gde-frame\" style=\"width:100%; height:500px; border: none;\" scrolling=\"no\"><\/iframe>\n<p class=\"gde-text\"><a href=\"http:\/\/ssrlab.by\/wp-content\/uploads\/2014\/02\/\u041f\u043e\u0440\u0442\u0430\u043b-\u0411\u0435\u043b\u0430\u0440\u0443\u0441\u044c-\u0421\u0435\u0433\u043e\u0434\u043d\u044f-_-\u041e\u0431\u0449\u0435\u0441\u0442\u0432\u043e-\u041f\u0440\u043e\u0432\u0435\u0440\u043a\u0430-\u0441\u0438\u043d\u0442\u0435\u0437\u0430-\u0441\u043b\u0443\u0445\u043e\u043c.pdf\" class=\"gde-link\">Download (PDF, 242KB)Download (PDF, 242KB)<\/p>","protected":false},"excerpt":{"rendered":"<p>It is much easier to deceive eyes than ears. When Yuri Hetsevich types a sentence \u201c\u041c\u0430\u043c\u0430 \u043c\u044b\u043b\u0430 \u0440\u0430\u043c\u0443\u201d (\u201cMother was cleaning the window frame\u201d), and  this phrase is immediately sounded by the \u201ctalking head\u201d on the screen, our eyes agree: indeed, the facial gesture beyond the screen is absolutely precise. But our ears, however, are doubtful of intonation. When a large text passage is generated, the distinction between computer-generated and real speech is heard much stronger [\u2026]<\/p>\n<a class = \"excerpt\" href=\"https:\/\/ssrlab.by\/en\/935\">Read more...<\/a>","protected":false},"author":1,"featured_media":0,"comment_status":"closed","ping_status":"open","sticky":false,"template":"","format":"standard","meta":[],"categories":[268,109],"tags":[],"_links":{"self":[{"href":"https:\/\/ssrlab.by\/en\/wp-json\/wp\/v2\/posts\/935"}],"collection":[{"href":"https:\/\/ssrlab.by\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/ssrlab.by\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/ssrlab.by\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/ssrlab.by\/en\/wp-json\/wp\/v2\/comments?post=935"}],"version-history":[{"count":28,"href":"https:\/\/ssrlab.by\/en\/wp-json\/wp\/v2\/posts\/935\/revisions"}],"predecessor-version":[{"id":4633,"href":"https:\/\/ssrlab.by\/en\/wp-json\/wp\/v2\/posts\/935\/revisions\/4633"}],"wp:attachment":[{"href":"https:\/\/ssrlab.by\/en\/wp-json\/wp\/v2\/media?parent=935"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/ssrlab.by\/en\/wp-json\/wp\/v2\/categories?post=935"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/ssrlab.by\/en\/wp-json\/wp\/v2\/tags?post=935"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}