{"id":175,"date":"2024-06-09T23:31:51","date_gmt":"2024-06-09T21:31:51","guid":{"rendered":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/?p=175"},"modified":"2024-06-09T23:32:02","modified_gmt":"2024-06-09T21:32:02","slug":"conversion-and-annotation-of-susan-abulhawas-the-blue-between-sky-and-water-or-a-demonstration-of-software-failure-through-anglocentrism","status":"publish","type":"post","link":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/2024\/06\/09\/conversion-and-annotation-of-susan-abulhawas-the-blue-between-sky-and-water-or-a-demonstration-of-software-failure-through-anglocentrism\/","title":{"rendered":"Conversion and Annotation of Susan Abulhawa\u2019s &#8222;The Blue Between Sky and Water&#8220;, or A Demonstration of Software Failure through Anglocentrism?"},"content":{"rendered":"\n<h2 class=\"wp-block-heading\">Introduction<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Having focused on reading literature through a postcolonial studies lens as well as on Eurocentric bias in the field of linguistics throughout my studies, the attempt of using conversion and annotation tools on a postcolonial post-monolingual Anglophone novel seemed intriguing.  I was interested to see how it would deal with the novel I chose \u2013 Susan Abulhawa\u2019s <em>The Blue Between Sky and Water<\/em>. <\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The novel by the Palestinian-American writer and human rights activist mixes Palestinian Arabic with English in a few different ways, although some patterns can be observed: Food items, terms for relatives, and culture-specific terms are usually written in latinized Arabic. Terms are usually introduced in italics once, and then re-appear throughout the novel un-italicised. As the software works with raw text, these specificities were lost. Apart from nouns, however, the novel includes verbs, adjectives, and phrases here and there in Palestinian Arabic as well, sometimes translated in the next sentences or before that, and sometimes not.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Sentence Choice<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">I chose seven sentences that show variation in the mixing of languages. For instance, one sentence I picked only includes Arabic nouns which denote food items:<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0 \u201cOne of her brothers arrived and they all shared a late breakfast of eggs, potatoes, <em>za\u2019atar<\/em>, olive oil, olives, hummus, <em>fuul<\/em>, pickled vegetables, and warm fresh bread\u201d (174).<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Another contains only one adjective in Arabic:<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0 \u201c\u2018Who is there?\u2019 a woman\u2019s voice asked in Arabic and Nazmiyeh relaxed upon hearing the Palestinian <em>fallahi<\/em> accent\u201d (35).<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Yet another contains an entire phrase:<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0 \u201cThe woman with wilted breasts began to sob quietly as others consoled her and banished the devil with disapproving eyes at Nazmiyeh \u2013 <em>a\u2019ootho billah min al shaytan<\/em> \u2013 when a female soldier wheeled in a large box of clothes, and with a gesture of her hand, gave the naked women permission to get dressed\u201d (114).<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">In addition I chose other sentences which include nouns in different ways to compare how the software deals with them. This was useful, as will become clear in the analysis of the mistakes the software made in the annotation.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Problems and Technical Difficulties<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">In general, working with the Jupyter Notebook style interface via Google Colab and Python worked out without any greater issues. The only problem I noticed was that double quotation marks could not be used, as they are part of the code and thus confuse the interface. Single quotation marks were acceptable for the software, so I replaced double quotation marks with single ones.<\/p>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"472\" height=\"233\" src=\"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/06\/quotation-marks-dont-work-unless-they-are-single-ones.jpg\" alt=\"\" class=\"wp-image-180\" srcset=\"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/06\/quotation-marks-dont-work-unless-they-are-single-ones.jpg 472w, https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/06\/quotation-marks-dont-work-unless-they-are-single-ones-300x148.jpg 300w\" sizes=\"auto, (max-width: 472px) 100vw, 472px\" \/><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Google Colab, however, was not as user-friendly. I found it rather unintuitive and it often claimed I had too many tabs open at the same time. As other students recommended, clearing the history and waiting a few minutes seem to have solved that issue. It is, however, time-consuming.<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"548\" src=\"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/06\/problem-zu-viele-sitzungen-1-1024x548.jpg\" alt=\"\" class=\"wp-image-181\" srcset=\"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/06\/problem-zu-viele-sitzungen-1-1024x548.jpg 1024w, https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/06\/problem-zu-viele-sitzungen-1-300x161.jpg 300w, https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/06\/problem-zu-viele-sitzungen-1-768x411.jpg 768w, https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/06\/problem-zu-viele-sitzungen-1-1536x822.jpg 1536w, https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/06\/problem-zu-viele-sitzungen-1-500x268.jpg 500w, https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/06\/problem-zu-viele-sitzungen-1.jpg 1627w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><\/figure>\n\n\n\n<h2 class=\"wp-block-heading\">Mistakes in the Annotation<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Most of the time, Arabic words were labelled as PROPN, proper nouns. This was especially the case with the sentence including an entire phrase in Arabic. For instance, \u201ca\u2019ootho\u201d should be labelled as a verb, but was labelled as proper noun. \u201cbillah\u201d was labelled as one single proper noun, even though it should be labelled as a preposition in combination with a noun, or proper noun (Allah). In other sentences, the software labelled Arabic nouns as adverbs. One example is \u201cfuul\u201d, which is a type of bean. Yet other nouns, such as \u201cjomaa\u201d were incorrectly labelled as adjectives. So, while the most frequent mistake was to label any Arabic word as proper noun, the software was not consistent in its mislabeling throughout.<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"343\" src=\"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/06\/p-114-screenshot-wrong-labels-1024x343.jpg\" alt=\"\" class=\"wp-image-177\" srcset=\"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/06\/p-114-screenshot-wrong-labels-1024x343.jpg 1024w, https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/06\/p-114-screenshot-wrong-labels-300x100.jpg 300w, https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/06\/p-114-screenshot-wrong-labels-768x257.jpg 768w, https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/06\/p-114-screenshot-wrong-labels-500x167.jpg 500w, https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/06\/p-114-screenshot-wrong-labels.jpg 1449w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">The software was still able to locate the ROOT correctly, despite its confusion around the Arabic terms. Arabic words resembling English words were sometimes thought to be English words. For instance, \u201cUm\u201d (mother) was mistakenly labelled as an interjection.<\/p>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"443\" height=\"599\" src=\"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/06\/um-interjection.jpg\" alt=\"\" class=\"wp-image-179\" srcset=\"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/06\/um-interjection.jpg 443w, https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/06\/um-interjection-222x300.jpg 222w, https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/files\/2024\/06\/um-interjection-370x500.jpg 370w\" sizes=\"auto, (max-width: 443px) 100vw, 443px\" \/><figcaption class=\"wp-element-caption\">&#8222;Um&#8220; is mislabeled as an interjection.<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Overall, this was an interesting experience, but I was disappointed that the software was unable to deal with Arabic words to such an extent, even though it was expected. \u00a0<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Introduction Having focused on reading literature through a postcolonial studies lens as well as on Eurocentric bias in the field of linguistics throughout my studies, the attempt of using conversion and annotation tools on a postcolonial post-monolingual Anglophone novel seemed &hellip; <a href=\"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/2024\/06\/09\/conversion-and-annotation-of-susan-abulhawas-the-blue-between-sky-and-water-or-a-demonstration-of-software-failure-through-anglocentrism\/\">Weiterlesen <span class=\"meta-nav\">&rarr;<\/span><\/a><\/p>\n","protected":false},"author":313,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1,2,4,3],"tags":[31,25,26,24,22,5,28,7,27],"class_list":["post-175","post","type-post","status-publish","format-standard","hentry","category-allgemein","category-blog-posts","category-student-entries","category-writing-across-languages","tag-anglocentrism","tag-annotating","tag-annotation","tag-converting","tag-digital-humanities","tag-multilingualism","tag-pos-tagging","tag-post-monolingual","tag-student-entry"],"_links":{"self":[{"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/posts\/175","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/users\/313"}],"replies":[{"embeddable":true,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/comments?post=175"}],"version-history":[{"count":1,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/posts\/175\/revisions"}],"predecessor-version":[{"id":182,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/posts\/175\/revisions\/182"}],"wp:attachment":[{"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/media?parent=175"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/categories?post=175"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/blogs.phil.hhu.de\/writingacrosslanguages\/wp-json\/wp\/v2\/tags?post=175"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}