{"id":280,"date":"2024-07-16T00:03:44","date_gmt":"2024-07-15T22:03:44","guid":{"rendered":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/?p=280"},"modified":"2024-07-16T00:03:54","modified_gmt":"2024-07-15T22:03:54","slug":"using-annis-as-a-tool-in-analysing-multilingual-sentences","status":"publish","type":"post","link":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/2024\/07\/16\/using-annis-as-a-tool-in-analysing-multilingual-sentences\/","title":{"rendered":"Using ANNIS as a Tool in Analysing Multilingual Sentences"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\">To me, using ANNIS to study and interpret the annotations of our collective corpus was fun. But only after figuring out the limits and possibilities that diffeI focused on one particular area. It was possible to discover different patterns in regards to the mistakes that occur when tagging non-English words. I specifically enjoyed analysing the POS-tagging of words that resembled English words but actually weren&#8217;t. <\/p>\n\n\n\n<h3 class=\"wp-block-heading\">How Non-English Words Are Tagged<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Across languages, there is an overlap in form, eventhough the meaning can vary drastically, words can look alike. I have addressed the English bias that can be found in the tagging of the multilingual passages in my first blog entry. But until using ANNIS to analyse the whole corpus I couldn&#8217;t estimate how often these mistakes occur nor what languages are affected. As our corpus shows, the given mistake is not uncommon, occurs across languages and across POS. By using the query <em>pos=\u200e&#8220;&#8230;.&#8220; &amp; isForeign=\u200e&#8220;True\u200e&#8220;<\/em> I could extract the following cases of words that seemed to be tagged according to English grammar.<\/p>\n\n\n\n<figure class=\"wp-block-table is-style-regular\"><table><tbody><tr><td><strong>POS (lemma)<\/strong><\/td><td><strong># &#8222;isForeign&#8220; but tagged according to English grammar<\/strong><\/td><\/tr><tr><td>PRON (me)<\/td><td>2<\/td><\/tr><tr><td>ADJ (mere)<\/td><td>2<\/td><\/tr><tr><td>Verb (do)<\/td><td>2<\/td><\/tr><tr><td>NOUN (soy)<\/td><td>1<\/td><\/tr><tr><td>ADV (so)<\/td><td>1<\/td><\/tr><tr><td>DET (a)<\/td><td>1<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Eventhough the tags seemed to consequently adhere to the English grammar when the word &#8222;looked&#8220; to be English, there was one exception. The word &#8222;won&#8220; was tagged as a noun and not a verb, as the software might have understood it. Instead, it seems that in this case even without a DET infront of it, the words was related to the currency &#8222;won&#8220;. <\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Implications for Further Research <\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">In order to successfully interpret the use of multilingual passages in literature, a correct tagging of words is neccessary. It matters, whether we try to determine what POS are most frequently non-English words etc. Questions arise, if these overlaps of languages are used by authors as a stylistic device or a challenge for the readers bias, or simply out of happenstance. Furthermore, the English dataset doesn&#8217;t suffice to annotate these multilingual texts. Moreover, these overlaps between languages might affect the machine translation of given passages and further highlight insufficencies or biases within AI\/ translations softwares. <\/p>\n","protected":false},"excerpt":{"rendered":"<p>To me, using ANNIS to study and interpret the annotations of our collective corpus was fun. But only after figuring out the limits and possibilities that diffeI focused on one particular area. It was possible to discover different patterns in &hellip; <a href=\"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/2024\/07\/16\/using-annis-as-a-tool-in-analysing-multilingual-sentences\/\">Weiterlesen <span class=\"meta-nav\">&rarr;<\/span><\/a><\/p>\n","protected":false},"author":393,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1,2,4,3],"tags":[25,26,22,5,28,27],"class_list":["post-280","post","type-post","status-publish","format-standard","hentry","category-allgemein","category-blog-posts","category-student-entries","category-writing-across-languages","tag-annotating","tag-annotation","tag-digital-humanities","tag-multilingualism","tag-pos-tagging","tag-student-entry"],"_links":{"self":[{"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/posts\/280","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/users\/393"}],"replies":[{"embeddable":true,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/comments?post=280"}],"version-history":[{"count":1,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/posts\/280\/revisions"}],"predecessor-version":[{"id":285,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/posts\/280\/revisions\/285"}],"wp:attachment":[{"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/media?parent=280"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/categories?post=280"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/tags?post=280"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}