A Verified Arabic-IPA Mapping for Arabic Transcription Technology, Informed by Quranic Recitation, Traditional Arabic Linguistics, and Modern Phonetics
In this paper, we present a detailed mapping from the graphemes of Modern standard Arabic (MSA) to symbols from the international Phonetic Alphabet (IPA) for automated transcription of Arabic text. This mapping is distinctive in several ways. First, the corpus used in rule development is the full te...
Authors: | ; ; ; |
---|---|
Format: | Electronic Article |
Language: | English |
Check availability: | HBZ Gateway |
Journals Online & Print: | |
Fernleihe: | Fernleihe für die Fachinformationsdienste |
Published: |
Oxford University Press
[2016]
|
In: |
Journal of Semitic studies
Year: 2016, Volume: 61, Issue: 1, Pages: 157-186 |
IxTheo Classification: | BJ Islam |
Online Access: |
Presumably Free Access Volltext (Verlag) Volltext (doi) |
Summary: | In this paper, we present a detailed mapping from the graphemes of Modern standard Arabic (MSA) to symbols from the international Phonetic Alphabet (IPA) for automated transcription of Arabic text. This mapping is distinctive in several ways. First, the corpus used in rule development is the full text of the Qur’ān rendered in fully pointed MSA. Second, we validate our scheme via automatically-generated frequency distributions of Arabic letters and diacritics over the whole corpus to anticipate and disambiguate non-trivial, compound grapheme-to-phoneme events , thus reducing the number of letter-to-sound rules. Such difficult cases include: the definite article; the letters alif, wāw , and yā ’; the variant forms of hamza ; the tanwīn case mark; and words with special pronunciations. Finally, our mapping scheme is informed by theory and practice from medieval Arabic linguistics and traditional Quranic recitation or tajwīd ; we make a novel contribution with new translations for ancient terms which incorporate concepts familiar to modern phoneticians. Our principal objective in automating Arabic-IPA transcription is to generate phonemic citation forms of Arabic words to enhance Arabic dictionaries, to facilitate Arabic language learning, and for natural language engineering applications. |
---|---|
ISSN: | 1477-8556 |
Contains: | Enthalten in: Journal of Semitic studies
|
Persistent identifiers: | DOI: 10.1093/jss/fgv035 |