Warning: Undefined array key "mm" in /www/wwwroot/www.ai-bt.com/si.php on line 10 Deprecated: trim(): Passing null to parameter #1 ($string) of type string is deprecated in /www/wwwroot/www.ai-bt.com/si.php on line 10 Improving the state-of-the-art in Thai semantic similarity using distributional semantics and ontological information.

Literature DB >> 33596220

Improving the state-of-the-art in Thai semantic similarity using distributional semantics and ontological information.

Ponrudee Netisopakul¹, Gerhard Wohlgenannt², Aleksei Pulich², Zar Zar Hlaing¹.

Abstract

Research into semantic similarity has a long history in lexical semantics, and it has applications in many natural language processing (NLP) tasks like word sense disambiguation or machine translation. The task of calculating semantic similarity is usually presented in the form of datasets which contain word pairs and a human-assigned similarity score. Algorithms are then evaluated by their ability to approximate the gold standard similarity scores. Many such datasets, with different characteristics, have been created for English language. Recently, four of those were transformed to Thai language versions, namely WordSim-353, SimLex-999, SemEval-2017-500, and R&G-65. Given those four datasets, in this work we aim to improve the previous baseline evaluations for Thai semantic similarity and solve challenges of unsegmented Asian languages (particularly the high fraction of out-of-vocabulary (OOV) dataset terms). To this end we apply and integrate different strategies to compute similarity, including traditional word-level embeddings, subword-unit embeddings, and ontological or hybrid sources like WordNet and ConceptNet. With our best model, which combines self-trained fastText subword embeddings with ConceptNet Numberbatch, we managed to raise the state-of-the-art, measured with the harmonic mean of Pearson on Spearman ρ, by a large margin from 0.356 to 0.688 for TH-WordSim-353, from 0.286 to 0.769 for TH-SemEval-500, from 0.397 to 0.717 for TH-SimLex-999, and from 0.505 to 0.901 for TWS-65.

Entities: CellLine Chemical Disease Gene Species

Mesh：

Year: 2021 PMID： 33596220 PMCID： PMC7888635 DOI： 10.1371/journal.pone.0246751

Source DB: PubMed Journal: PLoS One ISSN： 1932-6203 Impact factor: 3.240

1 in total

1. Distilling vector space model scores for the assessment of constructed responses with bifactor Inbuilt Rubric method and latent variables.

Authors: José Ángel Martínez-Huertas; Ricardo Olmos; Guillermo Jorge-Botana; José A León
Journal: Behav Res Methods Date: 2022-01-11

1 in total