{"id":277,"date":"2024-07-12T13:45:45","date_gmt":"2024-07-12T11:45:45","guid":{"rendered":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/?p=277"},"modified":"2024-07-12T13:47:12","modified_gmt":"2024-07-12T11:47:12","slug":"our-corpus","status":"publish","type":"post","link":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/2024\/07\/12\/our-corpus\/","title":{"rendered":"Our corpus"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\"><strong>Parts of speech<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Our corpus consists of 681 non-English and 3405 English words, meaning 4086 words in total. Here are some of the distributions as they were classified by the machine:<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table><tbody><tr><td><strong>Part of speech<\/strong><\/td><td><strong>Total amount of part of speech<\/strong><\/td><td><strong>Amount of non-English words<\/strong><\/td><td><strong>Amount of English words<\/strong><\/td><\/tr><tr><td>Proper nouns<\/td><td>989<\/td><td>420<\/td><td>569<\/td><\/tr><tr><td>Nouns<\/td><td>606<\/td><td>168<\/td><td>438<\/td><\/tr><tr><td>Verbs<\/td><td>408<\/td><td>25<\/td><td>383<\/td><\/tr><tr><td>Adjectives<\/td><td>184<\/td><td>26<\/td><td>158<\/td><\/tr><tr><td>Adverbs<\/td><td>105<\/td><td>10<\/td><td>95<\/td><\/tr><tr><td>Conjunctions<\/td><td>79<\/td><td>1<\/td><td>78<\/td><\/tr><tr><td>Determiners<\/td><td>258<\/td><td>3<\/td><td>255<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">The table shows the total amount of words belonging to part of speech category in the corpus and the amount of non-English and English words which belong to the respective part of speech. This shows that the machine usually classifies a word it does not know as a proper noun or maybe as a noun, but rarely as a conjunction, determiner, adjectives or adverbs.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">English determiners, conjunctions, adjectives, verbs and nouns were usually categorized correctly, while non-English determiners, conjunctions, adverbs, adjectives, verbs, nouns and English adverbs were usually categorized incorrectly. This clearly shows that the English words, except fort he English adverbs, are usually classified correctly, while the non-English words, together with the English adverbs, are usually classified incorrectly.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Dependency relations<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Regarding the dependency relations, a similar pattern is observable concerning to (in-)correct work by the machine: English words, except for the English adverbs have usually been assigned a correct dependency relation, while the non-English words, plus the English adverbs, have usually been assigned an incorrect dependency relation. It is also worth mentioning that some non-English words have not been assigned a dependency relation at all.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Parts of speech Our corpus consists of 681 non-English and 3405 English words, meaning 4086 words in total. Here are some of the distributions as they were classified by the machine: Part of speech Total amount of part of speech &hellip; <a href=\"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/2024\/07\/12\/our-corpus\/\">Weiterlesen <span class=\"meta-nav\">&rarr;<\/span><\/a><\/p>\n","protected":false},"author":391,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1],"tags":[33,22,21,5,23,27],"class_list":["post-277","post","type-post","status-publish","format-standard","hentry","category-allgemein","tag-corpus","tag-digital-humanities","tag-literary-translation","tag-multilingualism","tag-shared-experience","tag-student-entry"],"_links":{"self":[{"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/posts\/277","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/users\/391"}],"replies":[{"embeddable":true,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/comments?post=277"}],"version-history":[{"count":1,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/posts\/277\/revisions"}],"predecessor-version":[{"id":278,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/posts\/277\/revisions\/278"}],"wp:attachment":[{"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/media?parent=277"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/categories?post=277"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/tags?post=277"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}