File Information

File: 05-lr/acl_arc_1_sum/cleansed_text/xml_by_section/abstr/06/p06-2114_abstr.xml

Size: 921 bytes

Last Modified: 2025-10-06 13:45:13

<?xml version="1.0" standalone="yes"?>
<Paper uid="P06-2114">
  <Title>Sinhala Grapheme-to-Phoneme Conversion and Rules for Schwa Epenthesis</Title>
  <Section position="2" start_page="0" end_page="0" type="abstr">
    <SectionTitle>
Abstract
</SectionTitle>
    <Paragraph position="0"> This paper describes an architecture to convert Sinhala Unicode text into phonemic specification of pronunciation. The study was mainly focused on disambiguating schwa-/\/ and /a/ vowel epenthesis for consonants, which is one of the significant problems found in Sinhala. This problem has been addressed by formulating a set of rules. The proposed set of rules was tested using 30,000 distinct words obtained from a corpus and compared with the same words manually transcribed to phonemes by an expert.</Paragraph>
    <Paragraph position="1"> The Grapheme-to-Phoneme (G2P) conversion model achieves 98 % accuracy.</Paragraph>
  </Section>
class="xml-element"></Paper>
Download Original XML