A Verified Arabic-IPA Mapping for Arabic Transcription Technology, Informed by Quranic Recitation, Traditional Arabic Linguistics, and Modern Phonetics

In this paper, we present a detailed mapping from the graphemes of Modern standard Arabic (MSA) to symbols from the international Phonetic Alphabet (IPA) for automated transcription of Arabic text. This mapping is distinctive in several ways. First, the corpus used in rule development is the full te...

Full description

Saved in:  
Bibliographic Details
Authors: Brierley, Clare (Author) ; Sawalha, Majdi (Author) ; Heselwood, Barry (Author) ; Atwell, Eric (Author)
Format: Electronic Article
Language:English
Check availability: HBZ Gateway
Journals Online & Print:
Drawer...
Fernleihe:Fernleihe für die Fachinformationsdienste
Published: Oxford University Press [2016]
In: Journal of Semitic studies
Year: 2016, Volume: 61, Issue: 1, Pages: 157-186
IxTheo Classification:BJ Islam
Online Access: Presumably Free Access
Volltext (Verlag)
Volltext (doi)
Description
Summary:In this paper, we present a detailed mapping from the graphemes of Modern standard Arabic (MSA) to symbols from the international Phonetic Alphabet (IPA) for automated transcription of Arabic text. This mapping is distinctive in several ways. First, the corpus used in rule development is the full text of the Qur’ān rendered in fully pointed MSA. Second, we validate our scheme via automatically-generated frequency distributions of Arabic letters and diacritics over the whole corpus to anticipate and disambiguate non-trivial, compound grapheme-to-phoneme events , thus reducing the number of letter-to-sound rules. Such difficult cases include: the definite article; the letters alif, wāw , and yā ’; the variant forms of hamza ; the tanwīn case mark; and words with special pronunciations. Finally, our mapping scheme is informed by theory and practice from medieval Arabic linguistics and traditional Quranic recitation or tajwīd ; we make a novel contribution with new translations for ancient terms which incorporate concepts familiar to modern phoneticians. Our principal objective in automating Arabic-IPA transcription is to generate phonemic citation forms of Arabic words to enhance Arabic dictionaries, to facilitate Arabic language learning, and for natural language engineering applications.
ISSN:1477-8556
Contains:Enthalten in: Journal of Semitic studies
Persistent identifiers:DOI: 10.1093/jss/fgv035