article · npj Heritage Science
This study introduces a machine learning framework to classify the four traditional oral noun reading styles of Ge’ez: Seyaf ( ), Wedaki ( ), Tetay ( ), and Tenesh ( ). Using a manually curated corpus of 4264 nouns, we propose a linguistically constrained feature extraction method utilizing character n-grams (1–12) to capture sub-word phonological cues. To mitigate data scarcity and class imbalance, we implemented character-level augmentation and fold-specific SMOTE. In repeated stratified cross-validation, Random forest achieved 96.73 ± 0.14% accuracy, outperforming Linear SVC (95.64 ± 0.25%). A Wilcoxon signed-rank test confirmed a statistically significant difference (W = 0.0000, p = 5.96 × 10 −8 ). The accuracy difference (Δ = 0.58%, 95% CI [0.47%, 0.69%]) indicated a large effect size (paired Cohen’s d = 2.22). This research establishes a quantitative computational baseline for low-resource classical languages, demonstrating that complex oral prosody can be predicted solely through structured orthographic features, supporting digital heritage preservation.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1038/s40494-026-02903-y
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.