{"id":337,"date":"2024-07-29T23:35:18","date_gmt":"2024-07-29T21:35:18","guid":{"rendered":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/?p=337"},"modified":"2024-07-29T23:39:14","modified_gmt":"2024-07-29T21:39:14","slug":"annis-analysis-and-machine-translations","status":"publish","type":"post","link":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/2024\/07\/29\/annis-analysis-and-machine-translations\/","title":{"rendered":"ANNIS, analysis and machine translations"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\">Working with ANNIS was a lot easier than on spaCy and it has a pretty user-friendly interface. The query-builder function in the software helped in not only identifying the tendencies that machine annotations have towards non-English words but also in quantifying these tendencies. For example, while we had already established that POS tagging for non-English words is biased towards identifying them as porper nouns, through ANNIS we know that of 681 entries tagged&nbsp; \u201cForeign\u201c in our corpus, 420 are labelled with the POS tag PROPN, and 168 as NOUN.&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This also reflects how the non-English words have been used in the English text of the books in places of nouns in a sentence structure, thus also contributing to its recognition as a noun or proper noun.&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">As discussed in class, a variety of reasons underly this phenomenon. In an English sentence, replacing an English noun with a non-English one allows the author to preserve the sentence&#8217;s structure, making it more understandable for those who do not speak the language, and so the non-English words tend to be food, curse words, words that don\u2019t have a clear translation in English, or clothes.&nbsp; In contrast, using a foreign verb or adjective could change the grammar and potentially confuse readers who are not familiar with the syntactic rules of other languages.&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This in turn gives us an insight into the reception of the book, the market that it is written for, and the literary sociology around a text that is multilingual.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Our corpus consisted of several languages and it was interesting to note that these tendencies of POS tagging were applied to all non-English languages. This shows that the machine has been coded to essentially homogenise and annotate all languages that are not English in the same way, with no regard for the nuances and complexities that make languages so different from one another.&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong><em>Machine Translations<\/em><\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">While working on self-translations and machine translations for a multilingual text, we noticed that there is a lot of potential to develop software that can recognise context and more than two languages in a sentence.\u00a0The text I translated was Ocean Vuong\u2019s \u201cOn Earth We\u2019re Briefly Gorgeous\u201d which has English and Vietnamese, and manually translated as well as used Google Translate to convert the multilingual sentences to HIndi completely.\u00a0<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The machine translations often fell short in capturing the nuances and cultural context of languages, resulting in robotic and literal renditions that can distort meaning. For instance, in the provided translations from Ocean Vuong\u2019s &#8222;On Earth We\u2019re Briefly Gorgeous,&#8220; phrases like \u201c\u0110\u1eb9p qu\u00e1!\u201d were inaccurately translated to &#8222;\u0906\u092a \u0920\u0940\u0915 \u0939\u0948!&#8220; which means \u201cyou are alright,\u201d completely missing the expression of beauty. Similarly, &#8222;cream-colored orchid&#8220; was misinterpreted as &#8222;creamy,&#8220; suggesting a texture rather than color. Machine translations also struggled with the temporal context, as seen in \u201cthe cold winter week ahead of us,\u201d where a literal translation changed the sense of time to a spatial reference. short in capturing the nuances and cultural context of languages, resulting in robotic and literal renditions that can distort meaning.\u00a0<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The translation of multilingual sentences introduces additional challenges. In the text, Vietnamese phrases were awkwardly blended with English, resulting in jumbled and incorrect meanings, such as \u201cc\u00f3 \u0111u\u00f4i b\u00f2 kh\u00f4ng?\u201d becoming \u201c\u0915\u094d\u092f\u093e \u0906\u092a \u0916\u0941\u0936 \u0939\u0948\u0902?\u201d which translates to &#8222;are you happy?&#8220; instead of inquiring about oxtail availability. Machine translations lack the ability to discern these subtleties, leading to nonsensical and culturally insensitive outputs. In contrast, manual translations provide a more coherent and culturally aware interpretation, preserving the intended meaning and emotional depth of the original text.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Working with ANNIS was a lot easier than on spaCy and it has a pretty user-friendly interface. The query-builder function in the software helped in not only identifying the tendencies that machine annotations have towards non-English words but also in &hellip; <a href=\"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/2024\/07\/29\/annis-analysis-and-machine-translations\/\">Weiterlesen <span class=\"meta-nav\">&rarr;<\/span><\/a><\/p>\n","protected":false},"author":395,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1,2,4,3],"tags":[34,22,21,37,19,5,28,23,27],"class_list":["post-337","post","type-post","status-publish","format-standard","hentry","category-allgemein","category-blog-posts","category-student-entries","category-writing-across-languages","tag-annis","tag-digital-humanities","tag-literary-translation","tag-machine-translation","tag-multilingual","tag-multilingualism","tag-pos-tagging","tag-shared-experience","tag-student-entry"],"_links":{"self":[{"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/posts\/337","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/users\/395"}],"replies":[{"embeddable":true,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/comments?post=337"}],"version-history":[{"count":2,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/posts\/337\/revisions"}],"predecessor-version":[{"id":340,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/posts\/337\/revisions\/340"}],"wp:attachment":[{"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/media?parent=337"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/categories?post=337"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/tags?post=337"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}