{"id":295,"date":"2024-07-16T23:34:26","date_gmt":"2024-07-16T21:34:26","guid":{"rendered":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/?p=295"},"modified":"2024-07-16T23:34:36","modified_gmt":"2024-07-16T21:34:36","slug":"annis-much-potential-hindered-by-machine-annotation","status":"publish","type":"post","link":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/2024\/07\/16\/annis-much-potential-hindered-by-machine-annotation\/","title":{"rendered":"ANNIS &#8211; Much Potential, Hindered by Machine Annotation"},"content":{"rendered":"\n<h2 class=\"wp-block-heading\">Experience Using ANNIS<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">After getting used to the surface, I found ANNIS to be very user-friendly. The Query-Builder allows for a combination of all kinds of prompts. In theory, this could be used to find out how many non-English words in our corpus are proper nouns, nouns, adjectives, etc. Such information would be useful to find evidence for patterns found in post- and\/or multilingual novels, or might even lead to the discovery of patterns so far missed.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">However, the faulty machine annotation, which inevitably became part of our corpus, made work with the findings rather difficult. Firstly, most non-English words were simply classified as proper nouns. ANNIS underlines the absurdity of that number: If one looks for words which are both non-English and proper nouns, one receives over 400 matches. If one looks for words which are English and proper nouns, the number is far smaller.<\/p>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"389\" height=\"480\" src=\"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/07\/foreign-and-proper-noun-multilin.jpg\" alt=\"\" class=\"wp-image-296\" srcset=\"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/07\/foreign-and-proper-noun-multilin.jpg 389w, https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/07\/foreign-and-proper-noun-multilin-243x300.jpg 243w\" sizes=\"auto, (max-width: 389px) 100vw, 389px\" \/><figcaption class=\"wp-element-caption\">420 matches for non-English &#8222;proper nouns&#8220;<\/figcaption><\/figure>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"389\" height=\"482\" src=\"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/07\/not-foreign-and-proper-noun-multilin.jpg\" alt=\"\" class=\"wp-image-297\" srcset=\"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/07\/not-foreign-and-proper-noun-multilin.jpg 389w, https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/07\/not-foreign-and-proper-noun-multilin-242x300.jpg 242w\" sizes=\"auto, (max-width: 389px) 100vw, 389px\" \/><figcaption class=\"wp-element-caption\">149 matches for English proper nouns<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">The results above are already fishy. Something cannot be right here. Spoiler alert: We all know it is the preceding machine annotation. Playing around with ANNIS some more, I also found instances where the machine had been unable to annotate English words correctly, resulting in falsified ANNIS search results.<\/p>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"617\" height=\"175\" src=\"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/07\/non-foreign-prop-n-also-mistakes.jpg\" alt=\"\" class=\"wp-image-298\" srcset=\"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/07\/non-foreign-prop-n-also-mistakes.jpg 617w, https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/07\/non-foreign-prop-n-also-mistakes-300x85.jpg 300w\" sizes=\"auto, (max-width: 617px) 100vw, 617px\" \/><figcaption class=\"wp-element-caption\">Here, for instance, the adjective &#8222;teensy&#8220; is annotated as a proper noun.<\/figcaption><\/figure>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"629\" height=\"130\" src=\"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/07\/not-foreign-and-proper-noun-mistake-2.jpg\" alt=\"\" class=\"wp-image-299\" srcset=\"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/07\/not-foreign-and-proper-noun-mistake-2.jpg 629w, https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/07\/not-foreign-and-proper-noun-mistake-2-300x62.jpg 300w\" sizes=\"auto, (max-width: 629px) 100vw, 629px\" \/><figcaption class=\"wp-element-caption\">Here, a verb is annotated as a proper noun.<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Given that ANNIS can work with several languages at once, I would really like to see a manually annotated corpus of post- and\/or multilingual literatures. This experience strengthened my appreciation for corpora, but brought forward a lack of trust in machine annotation. I had expected at least the annotation of English words to be correct.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Our Little Experiment With (AI) Translation<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">We conducted a little experiment, translating our sentences ourselves and also having them translated by AI. Interestingly, the AI, such as DeepL and Google Translate, displayed tendencies similar to those of the machine annotation. Non-English words were combined with other nouns into compounds. For instance, &#8222;fallahi accent&#8220; (meaning: rural accent, the accent used by farmers) became &#8222;Fallahi-Akzent&#8220;. In other words, both machine annotation and AI translation tools have a tendency to deal with non-English words by interpreting them as nouns or proper nouns, rather than what they actually are.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Experience Using ANNIS After getting used to the surface, I found ANNIS to be very user-friendly. The Query-Builder allows for a combination of all kinds of prompts. In theory, this could be used to find out how many non-English words &hellip; <a href=\"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/2024\/07\/16\/annis-much-potential-hindered-by-machine-annotation\/\">Weiterlesen <span class=\"meta-nav\">&rarr;<\/span><\/a><\/p>\n","protected":false},"author":313,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1,2,4,3],"tags":[31,34,26,33,22,21,35,37,38,36,5,28],"class_list":["post-295","post","type-post","status-publish","format-standard","hentry","category-allgemein","category-blog-posts","category-student-entries","category-writing-across-languages","tag-anglocentrism","tag-annis","tag-annotation","tag-corpus","tag-digital-humanities","tag-literary-translation","tag-machine-annotation","tag-machine-translation","tag-machine-translation-bias","tag-manual-annotation","tag-multilingualism","tag-pos-tagging"],"_links":{"self":[{"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/posts\/295","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/users\/313"}],"replies":[{"embeddable":true,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/comments?post=295"}],"version-history":[{"count":1,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/posts\/295\/revisions"}],"predecessor-version":[{"id":301,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/posts\/295\/revisions\/301"}],"wp:attachment":[{"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/media?parent=295"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/categories?post=295"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/tags?post=295"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}