{"id":290,"date":"2024-07-16T15:21:32","date_gmt":"2024-07-16T13:21:32","guid":{"rendered":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/?p=290"},"modified":"2024-07-16T15:23:55","modified_gmt":"2024-07-16T13:23:55","slug":"analysing-multilingual-sentences-with-annis","status":"publish","type":"post","link":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/2024\/07\/16\/analysing-multilingual-sentences-with-annis\/","title":{"rendered":"Analysing multilingual sentences with ANNIS"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\">For me, working with ANNIS was much more fun than annotating sentences in Google Collab. I liked actually being able to get some quantifications out of the sentences we annotated. Although in the end the corpus we uploaded on ANNIS was not that extensive, it was still very interesting to see the possibilities the program offers. <\/p>\n\n\n\n<p class=\"wp-block-paragraph\">During the course of this semester, we already noticed an extremely high error rate in the POS classification and dependency tagging of spaCY. I was interested to see, how high that rate is in numbers, now that we could do analysis of that sort with ANNIS. Because it was the end of the semester and my time and capacities were limited, I limited my research to four POS, which I thought were most interesting. <\/p>\n\n\n\n<p class=\"wp-block-paragraph\">With the help of the ANNIS Query Builder I looked for all proper nouns (PROPN), nouns (NOUN), verbs (VERB) and adjectives (ADJ) that were annotated (manually) as foreign. I then checked, how many of the POS tags were actually correct. Because my language skills are limited and it wasn&#8217;t always clear what language I was analysing, there were some words (or tokens) I couldn&#8217;t identify even through research. These are my findings: <\/p>\n\n\n\n<figure class=\"wp-block-table\"><table><tbody><tr><td>POS<\/td><td>Foreign words with that tag<\/td><td>classification considered correct<\/td><td>unsure about classification<\/td><\/tr><tr><td>PROPN<\/td><td>420<\/td><td>55<\/td><td>&#8211;<\/td><\/tr><tr><td>ADJ<\/td><td>26<\/td><td>4<\/td><td>2<\/td><\/tr><tr><td>NOUN<\/td><td>168<\/td><td>93<\/td><td>23<\/td><\/tr><tr><td>VERB<\/td><td>25<\/td><td>8<\/td><td>5<\/td><\/tr><\/tbody><\/table><figcaption class=\"wp-element-caption\">*I hope these numbers are correct. It is possible that I lost count or got confused at some point.<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">As you can see, the POS tagged the most is the proper noun, which was also our impression in previous sessions. It seems that labelling a token as PROPN is an easy solution for words which can&#8217;t be recognised by the machine. It is also the category with the most mistakes. <\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The POS with the least mistakes is the NOUN. This might be due to the fact that a noun is relatively easy to identify, especially if it is the only foreign word in an English sentence and maybe even accompanied by a determinant. This is probably also the reason, why authors of multilingual literary texts use relatively many foreign nouns: they are easy to discern for English readers. <\/p>\n\n\n\n<p class=\"wp-block-paragraph\">For ADJ and VERB, the number of findings is relatively low, so I&#8217;m not comfortable making any conclusions because of my &#8222;findings&#8220;. I suspect, that verbs and adjectives are a lot harder to identify than nouns and proper nouns, also because they usually change due to flexion and conjugation (at least in the languages that I know).<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">I know that our corpus is not yet big enough to come to definitive conclusions. It might also be biased, because we tried choosing a high variety of sentences out of every book. Still, it was quite interesting to see (if only on a small scale) what ANNIS can do. <\/p>\n","protected":false},"excerpt":{"rendered":"<p>For me, working with ANNIS was much more fun than annotating sentences in Google Collab. I liked actually being able to get some quantifications out of the sentences we annotated. Although in the end the corpus we uploaded on ANNIS &hellip; <a href=\"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/2024\/07\/16\/analysing-multilingual-sentences-with-annis\/\">Weiterlesen <span class=\"meta-nav\">&rarr;<\/span><\/a><\/p>\n","protected":false},"author":389,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1],"tags":[34,22,19,28,13,7],"class_list":["post-290","post","type-post","status-publish","format-standard","hentry","category-allgemein","tag-annis","tag-digital-humanities","tag-multilingual","tag-pos-tagging","tag-post-anglophone","tag-post-monolingual"],"_links":{"self":[{"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/posts\/290","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/users\/389"}],"replies":[{"embeddable":true,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/comments?post=290"}],"version-history":[{"count":3,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/posts\/290\/revisions"}],"predecessor-version":[{"id":293,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/posts\/290\/revisions\/293"}],"wp:attachment":[{"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/media?parent=290"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/categories?post=290"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/tags?post=290"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}