{"id":226,"date":"2024-06-17T03:06:56","date_gmt":"2024-06-17T01:06:56","guid":{"rendered":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/?p=226"},"modified":"2024-06-17T03:07:14","modified_gmt":"2024-06-17T01:07:14","slug":"multi-lingual-annotations-in-python-with-a-concise-chinese-english-dictionary-for-lovers","status":"publish","type":"post","link":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/2024\/06\/17\/multi-lingual-annotations-in-python-with-a-concise-chinese-english-dictionary-for-lovers\/","title":{"rendered":"Multi-Lingual Annotations in Python with &#8222;A Concise Chinese-English Dictionary for Lovers&#8220;"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\">This was not my first time working with Python to figure out the tokens of sentences via computational means. However, my last time working with Python (in a linguistics BA seminar with Prof. Kevin Tang) has been some time back, so I have appreciated that working our example sentences through it has been made very easy.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Running the three example sentences given to us, I noticed that punctuations marking speech are often interpreted as part of words or compounds and are tagged as POS, which makes it difficult to check which tokens refer to which words.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Additionally, foreign words were always interpreted as proper nouns \u2013 even in the Spanish phrase <em>clarita del huevo<\/em>, in which <em>del<\/em> is clearly not a proper noun. I would have thought that maybe the Spanish language would be easier to interpret or perhaps translate than the Swahili sentence as it might be more commonly known, but Python does not do any translating and so struggles with anything that is not English. Dependencies can thus not be correctly determined: in the sentence <em>T\u00edas called me blanca, palida, clarita del huevo<\/em> the last three parts (<em>blanca, palida, clarita del huevo<\/em>) are a listing, all nouns or noun phrases are on equal standing and not dependent on each other, but Python marks <em>huevo<\/em> as a dependent of <em>palida<\/em>.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The given example with Swahili words in it cannot be properly tagged at all as \u2013 according to Python &#8211; <em>\u2018Ayaaaana! Haki ya Mungu &#8230; aieee!\u2019 <\/em><em>The threat-drenched contralto came from the bushes to the left of the mangroves. \u2018Aii, mwanangu, mbona wanitesa?\u2019 <\/em>is made up entirely of nouns or proper nouns. Only the English part could be identified correctly.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The English-Chinese example provides similar issues:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><em>In Chinese, it is the same word \u2018<\/em><em>\u5bb6<\/em><em>\u2019 (jia) for \u2018home\u2019 and \u2018family\u2019 and sometimes including \u2018house\u2019. To us, family is same thing as house, and this house is their only home too. \u2018<\/em><em>\u5bb6<\/em><em>\u2019, a roof on top, then some legs and arms inside.<\/em><\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">In this example, the Chinese <em>hanzi<\/em> are not computed as proper nouns. Python interprets \u5bb6 as a noun once and an adjective another time, marking it \u2013 quite nonsensically \u2013 as a dependent of <em>legs<\/em>.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Now looking at the novel I have been reading \u2013 <em>A Concise Chinese-English Dictionary for Lovers<\/em> by Xiaolu Guo \u2013 I looked for similar multilingual sentences to test in the programme. I couldn\u2019t find many such sentences, but tested the ones I did find:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><em>\u2018<\/em><em>\u77e5\u8bc6<\/em><em>\u2019 mean knowledge, \u2018<\/em><em>\u5206\u5b50<\/em><em>\u2019 mean molecule.<\/em><\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">\u77e5\u8bc6 is here correctly interpreted as a noun, although it is interpreted as one singular noun rather than a noun phrase consisting of a verb (\u77e5 \u2013 to know) and a noun (\u8bc6 \u2013 knowledge).<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Similar goes for\u5206\u5b50, interpreted as one noun rather than a noun phrase consisting of a verb (\u5206 \u2013 divide) and a noun (\u5b50 \u2013 son, child).<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The word \u201cmean\u201d in this sentence is the verb \u201cto mean\u201d, but not conjugated correctly because the narrator is not fluent in English yet and struggles with English grammar. Python thus tags it as an adjective. Accordingly, the dependencies turn out incorrect as well, as the <em>mean<\/em> should function as the head of its sentence. Instead, <em>knowledge<\/em> becomes that head upon which all other words depend.<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><em>\u5c41<\/em><em> <\/em><em>is fart in Chinese. It is the word made up from two parts. <\/em><em>\u5c38<\/em><em> <\/em><em>is a symbol of a body with tail, and underneath that <\/em><em>\u6bd4<\/em><em> <\/em><em>represent two legs. That means fart, a kind of Chi.<\/em><\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">In the 3<sup>rd<\/sup> sentence \u201crepresent\u201d is interpreted as a dependent of \u201cis\u201d \u2013 again, likely due to the improper grammar the narrator uses. The hanzi are not computed at all (\u5c41), marked as a proper noun (\u5c38) or a noun (\u6bd4).<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><em>Chi (<\/em><em>\u6c14<\/em><em>), everything to do with Chi is very important to us Chinese.<\/em><\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">This sentence starts out sort of elliptical. The phrase <em>very important to us Chinese<\/em> is tagged and interpreted correctly, but Python struggles with everything that comes before it, marking <em>Chi (<\/em><em>\u6c14<\/em><em>)<\/em> as a dependent of <em>is<\/em>. Evidently, Python cannot compute and interpret punctuation correctly, which in this case should indicated that the first word is separate from the sentence after the comma.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">In the next two examples I wanted to test how Python deals with incorrect English grammar but without the interference of non-English words, hypothesising that situations like the above <em>mean<\/em> would also occur here.<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><em>I feeling I can die for all kinds of situation in every second.<\/em><\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">In this case, coming from previous errors Python made due to incorrect grammar and conjugation, I thought that <em>feeling<\/em> might be interpreted as a noun because the auxiliary \u201cam\u201d of the progressive form is missing. Surprisingly, Python has no problem recognising <em>feeling<\/em> as the verb it is indeed supposed to be, marking it correctly as the head\/root of the entire sentence<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><em>I scared by cars because they seems coming from any possible directing.<\/em><\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">Similar as above, I wanted to test how Python deals with these grammatical errors (<em>seems<\/em> instead of <em>seem<\/em>; <em>directing<\/em> instead of <em>direction<\/em>) \u2013 again, surprisingly, all tokens were tagged and interpreted properly with all their dependencies. Even <em>directing<\/em> has been correctly identified as a noun instead of a progressive verb.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Evidently, annotating multi-lingual sentences correctly is not a possibility \u2013 at least not with the Python code we have been given. While the programme has no problem interpreting English sentences with incorrect grammar, it is thrown for a loop as soon as non-English words are introduced, which was a very interesting observation to make.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>This was not my first time working with Python to figure out the tokens of sentences via computational means. However, my last time working with Python (in a linguistics BA seminar with Prof. Kevin Tang) has been some time back, &hellip; <a href=\"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/2024\/06\/17\/multi-lingual-annotations-in-python-with-a-concise-chinese-english-dictionary-for-lovers\/\">Weiterlesen <span class=\"meta-nav\">&rarr;<\/span><\/a><\/p>\n","protected":false},"author":400,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1,2,4,3],"tags":[26,22,21,19,5,28],"class_list":["post-226","post","type-post","status-publish","format-standard","hentry","category-allgemein","category-blog-posts","category-student-entries","category-writing-across-languages","tag-annotation","tag-digital-humanities","tag-literary-translation","tag-multilingual","tag-multilingualism","tag-pos-tagging"],"_links":{"self":[{"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/posts\/226","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/users\/400"}],"replies":[{"embeddable":true,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/comments?post=226"}],"version-history":[{"count":1,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/posts\/226\/revisions"}],"predecessor-version":[{"id":227,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/posts\/226\/revisions\/227"}],"wp:attachment":[{"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/media?parent=226"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/categories?post=226"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/tags?post=226"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}