Academic Journal

PROBABILISTIC SEMANTIC RECONSTRUCTION OF LOST PROTO INDO-EUROPEAN DIALECTS USING COMPUTATIONAL COMPARATIVE LINGUISTIC MODELING AND DEEP NEURAL ARCHIVING.

Bibliographic Details
Title: PROBABILISTIC SEMANTIC RECONSTRUCTION OF LOST PROTO INDO-EUROPEAN DIALECTS USING COMPUTATIONAL COMPARATIVE LINGUISTIC MODELING AND DEEP NEURAL ARCHIVING.
Authors: Tadjieva, Mastura1 tadjieva.mastura@mail.ru, Matniyazova, Zaynab2 matniyozova.zaynabjon@bsmi.uz, Djumayeva, Zarina3 zarinadijumayeva@gmail.com, Umarkhujaeva, Khilola4 umarkhujaevakhilola@gmail.com, Khoshimkhujaeva, Mokhirukh5 moxirux_xoshimxojayeva@tues.uz, Ankabayeva, Mohira6 mohira25674524@gmail.com, Yusupov, Otabek7 otabeksam@gmail.com
Source: Archives for Technical Sciences / Arhiv za Tehnicke Nauke. 2026, Issue 35, p49-59. 11p.
Subject Terms: *Reconstruction (Linguistics), *Computational linguistics, *Historical linguistics, *Artificial neural networks, *Bayesian analysis, *Phylogenetic models, *Indo-European languages
Abstract: Manual comparative methods have long been the main source for reconstructing Proto-Indo-European (PIE) dialects, with their weaknesses including fragmentary corpora, interpretive bias, and a lack of direct textual evidence. This paper introduces a probabilistic semantic reconstruction model that combines computational comparative linguistics and deep neural archiving to learn and reconstruct the dialectal variations that have been lost in PIE. A multilingual dataset of 12 Indo-European language branches and 18,742 cognate sets, with phonological, morphological, and semantic feature embeddings, was compiled and entered. An inverse phylogenetic inference model based on Bayesian inference and a transformer-based deep neural network trained on 4.6 million aligned lexical tokens was used to predict proto-forms and semantic shifts. When tested against known scholarly reconstructions, the proposed model achieved 86.3% accuracy in phonological reconstruction and 0.81 semantic consistency (cosine similarity metric). Cross-validation indicated a 14.7% decrease in reconstruction variance compared to traditional rule-based methods. Probabilistic confidence intervals (95% CI) also showed consistent predictions for high-frequency lexical roots, with posterior probabilities greater than 0.90 for the reconstructed forms (63%). Moreover, statistically significant divergence patterns (p < 0.01) were observed in the dialectal clustering analysis and were consistent with established Indo-European subgroup stratifications. The results show that probabilistic modelling with deep neural semantic archiving can significantly improve the reliability and interpretability of reconstruction. This framework offers a computational approach to historical linguistics that can be scaled and replicated. Also, it provides a new quantitative understanding of the evolution of proto-languages and dialect differentiation within the Indo-European family. [ABSTRACT FROM AUTHOR]
Database: Academic Search Index
Description
ISSN:18404855
DOI:10.70102/afts.2026.1835.049