{"id":303,"date":"2024-07-17T09:10:59","date_gmt":"2024-07-17T07:10:59","guid":{"rendered":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/?p=303"},"modified":"2024-07-17T09:11:16","modified_gmt":"2024-07-17T07:11:16","slug":"working-with-annis-on-multilingual-sentences-experience-and-analysis","status":"publish","type":"post","link":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/2024\/07\/17\/working-with-annis-on-multilingual-sentences-experience-and-analysis\/","title":{"rendered":"Working with ANNIS on multilingual Sentences &#8211; Experience and Analysis"},"content":{"rendered":"\n<h2 class=\"wp-block-heading\">Set up and first impressions<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">I found working with ANNIS both fun and insightful. I especially appreciated that query results displayed annotations and sentence dependencies immediately, unlike in Google Collab, where multiple intermediate steps were needed. This intuitive interface made ANNIS very user-friendly in my experience. After overcoming some initial difficulties with setting up combined queries, it was actually quite fun to experiment with different search requests and observe the results.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">POS-tagging in ANNIS<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">As we have already established using Google Collab, most foreign-language words are recognized as proper nouns by machine annotation. The analysis with ANNIS confirms this impression: Out of 681 entries tagged as &#8222;Foreign&#8220; in our corpus, 420 are labelled with the POS tag PROPN. I was therefore particularly interested in the instances where other POS tags were assigned.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The category NOUN forms the second largest group in the corpus, with 168 entries in total. In many of these cases, the annotation process appears to identify misspellings of English words, as evidenced by the example &#8222;climinals,&#8220; correctly identified as a noun despite being categorized as foreign:<\/p>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"742\" height=\"163\" src=\"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/07\/image.png\" alt=\"\" class=\"wp-image-304\" srcset=\"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/07\/image.png 742w, https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/07\/image-300x66.png 300w\" sizes=\"auto, (max-width: 742px) 100vw, 742px\" \/><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\"> Additionally, I noticed that this category often includes words that are established foreign words in English, such as &#8222;malum,&#8220; &#8222;ennui,&#8220; and &#8222;hummus,&#8220; all of which are listed in the Oxford Dictionary.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The POS tag for adjectives is less common, appearing in only 26 entries. These entries often consist of foreign words introduced by an English determiner or occurring immediately before a word tagged as a noun. Interestingly, in one example, the English word &#8222;cadaverous&#8220; was mistakenly labeled as foreign but correctly identified as an adjective at the same time:<\/p>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"817\" height=\"163\" src=\"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/07\/image-1.png\" alt=\"\" class=\"wp-image-305\" srcset=\"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/07\/image-1.png 817w, https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/07\/image-1-300x60.png 300w, https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/07\/image-1-768x153.png 768w\" sizes=\"auto, (max-width: 817px) 100vw, 817px\" \/><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">The POS category VERB is also relatively rare, appearing in just 25 entries in our corpus. This tag assignment seems to primarily follow the V2 sentence structure of English, as the second word in multilingual sentences is most frequently assigned this category.<\/p>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"834\" height=\"162\" src=\"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/07\/image-2.png\" alt=\"\" class=\"wp-image-306\" srcset=\"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/07\/image-2.png 834w, https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/07\/image-2-300x58.png 300w, https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/07\/image-2-768x149.png 768w\" sizes=\"auto, (max-width: 834px) 100vw, 834px\" \/><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Overall, some of these observations are rather anecdotal, as many entries do not seem to follow a clear pattern in the machine annotation. It would be interesting to see what results a larger and possibly more refined corpus might yield in these cases.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Set up and first impressions I found working with ANNIS both fun and insightful. I especially appreciated that query results displayed annotations and sentence dependencies immediately, unlike in Google Collab, where multiple intermediate steps were needed. This intuitive interface made &hellip; <a href=\"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/2024\/07\/17\/working-with-annis-on-multilingual-sentences-experience-and-analysis\/\">Weiterlesen <span class=\"meta-nav\">&rarr;<\/span><\/a><\/p>\n","protected":false},"author":394,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1,2,4,3],"tags":[34,26,5,28,27],"class_list":["post-303","post","type-post","status-publish","format-standard","hentry","category-allgemein","category-blog-posts","category-student-entries","category-writing-across-languages","tag-annis","tag-annotation","tag-multilingualism","tag-pos-tagging","tag-student-entry"],"_links":{"self":[{"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/posts\/303","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/users\/394"}],"replies":[{"embeddable":true,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/comments?post=303"}],"version-history":[{"count":2,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/posts\/303\/revisions"}],"predecessor-version":[{"id":308,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/posts\/303\/revisions\/308"}],"wp:attachment":[{"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/media?parent=303"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/categories?post=303"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/tags?post=303"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}