{"id":321,"date":"2024-07-17T11:36:01","date_gmt":"2024-07-17T09:36:01","guid":{"rendered":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/?p=321"},"modified":"2024-07-17T11:36:11","modified_gmt":"2024-07-17T09:36:11","slug":"the-proper-noun-problem-working-with-annis","status":"publish","type":"post","link":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/2024\/07\/17\/the-proper-noun-problem-working-with-annis\/","title":{"rendered":"The Proper Noun Problem &#8211; Working with ANNIS"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\">Working with the software ANNIS and the corpus consisting of our novels\u2019 example sentences, one of the predictions we had at the beginning of term came true: Machine annotation is indeed faulty and there is a tendency to classify non-English words as proper nouns. We already saw this tendency while working with Google Collab, but now ANNIS shows us that non-English words are indeed often analysed as such.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">To come to this conclusions, I pretty much just did a query for all foreign words with isForeign=&#8220;True&#8220;. There were 681 matches, so obviously quite a lot as this was our goal. Instead of looking through all matches, I decided to focus on one sentence, which starts with the the fourth result &#8222;Que siempre la lengua \u2026&#8220;.<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"128\" src=\"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/07\/Bildschirmfoto-2024-07-17-um-11.19.11-1024x128.png\" alt=\"\" class=\"wp-image-322\" srcset=\"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/07\/Bildschirmfoto-2024-07-17-um-11.19.11-1024x128.png 1024w, https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/07\/Bildschirmfoto-2024-07-17-um-11.19.11-300x38.png 300w, https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/07\/Bildschirmfoto-2024-07-17-um-11.19.11-768x96.png 768w, https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/07\/Bildschirmfoto-2024-07-17-um-11.19.11-1536x192.png 1536w, https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/07\/Bildschirmfoto-2024-07-17-um-11.19.11.png 2032w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">As can be seen, 13 out of 15 non-English words are annotated and analysed as a proper noun &#8211; which obviously does not make sense, as a sentence just cannot only be made up of proper nouns. Hence, this already confirms the prediction of a tendency to categorise non-English words as proper nouns.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">However, there are also two exceptions: siempre and la. Siempre is categorised as a verb, la as a determiner. Which is however, also not completely correct. &#8218;La&#8216; is a determiner as it is an article, however, &#8217;siempre&#8216; is &#8211; if my broken Spanish is anything to be sure of &#8211; an adverb, as it means always.\u00a0I do not really have an answer as to why these two words are exceptions, and especially not why one is still falsely while the other is correctly analysed.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Hence, most of our initial predictions did come true while working with ANNIS. This is obviously not a great result in terms of correctly annotating and tokenising sentences, however, in terms of emphasising the problem of an anglocentric view and an Anglocentrism in software developing, it is a very telling result.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Working with the software ANNIS and the corpus consisting of our novels\u2019 example sentences, one of the predictions we had at the beginning of term came true: Machine annotation is indeed faulty and there is a tendency to classify non-English &hellip; <a href=\"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/2024\/07\/17\/the-proper-noun-problem-working-with-annis\/\">Weiterlesen <span class=\"meta-nav\">&rarr;<\/span><\/a><\/p>\n","protected":false},"author":307,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1],"tags":[],"class_list":["post-321","post","type-post","status-publish","format-standard","hentry","category-allgemein"],"_links":{"self":[{"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/posts\/321","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/users\/307"}],"replies":[{"embeddable":true,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/comments?post=321"}],"version-history":[{"count":2,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/posts\/321\/revisions"}],"predecessor-version":[{"id":324,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/posts\/321\/revisions\/324"}],"wp:attachment":[{"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/media?parent=321"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/categories?post=321"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/tags?post=321"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}