Literature DB >> 31334764

Neural machine translation of clinical texts between long distance languages.

Xabier Soto1, Olatz Perez-de-Viñaspre1, Gorka Labaka1, Maite Oronoz1.   

Abstract

OBJECTIVE: To analyze techniques for machine translation of electronic health records (EHRs) between long distance languages, using Basque and Spanish as a reference. We studied distinct configurations of neural machine translation systems and used different methods to overcome the lack of a bilingual corpus of clinical texts or health records in Basque and Spanish.
MATERIALS AND METHODS: We trained recurrent neural networks on an out-of-domain corpus with different hyperparameter values. Subsequently, we used the optimal configuration to evaluate machine translation of EHR templates between Basque and Spanish, using manual translations of the Basque templates into Spanish as a standard. We successively added to the training corpus clinical resources, including a Spanish-Basque dictionary derived from resources built for the machine translation of the Spanish edition of SNOMED CT into Basque, artificial sentences in Spanish and Basque derived from frequently occurring relationships in SNOMED CT, and Spanish monolingual EHRs. Apart from calculating bilingual evaluation understudy (BLEU) values, we tested the performance in the clinical domain by human evaluation.
RESULTS: We achieved slight improvements from our reference system by tuning some hyperparameters using an out-of-domain bilingual corpus, obtaining 10.67 BLEU points for Basque-to-Spanish clinical domain translation. The inclusion of clinical terminology in Spanish and Basque and the application of the back-translation technique on monolingual EHRs significantly improved the performance, obtaining 21.59 BLEU points. This was confirmed by the human evaluation performed by 2 clinicians, ranking our machine translations close to the human translations. DISCUSSION: We showed that, even after optimizing the hyperparameters out-of-domain, the inclusion of available resources from the clinical domain and applied methods were beneficial for the described objective, managing to obtain adequate translations of EHR templates.
CONCLUSION: We have developed a system which is able to properly translate health record templates from Basque to Spanish without making use of any bilingual corpus of clinical texts or health records.
© The Author(s) 2019. Published by Oxford University Press on behalf of the American Medical Informatics Association. All rights reserved. For permissions, please email: journals.permissions@oup.com.

Entities:  

Keywords:  electronic health records; long distance languages; machine translation; natural language processing; neural networks

Mesh:

Year:  2019        PMID: 31334764      PMCID: PMC7647170          DOI: 10.1093/jamia/ocz110

Source DB:  PubMed          Journal:  J Am Med Inform Assoc        ISSN: 1067-5027            Impact factor:   4.497


  2 in total

Review 1.  Applications of natural language processing in ophthalmology: present and future.

Authors:  Jimmy S Chen; Sally L Baxter
Journal:  Front Med (Lausanne)       Date:  2022-08-08

2.  The Use of Automated Machine Translation to Translate Figurative Language in a Clinical Setting: Analysis of a Convenience Sample of Patients Drawn From a Randomized Controlled Trial.

Authors:  Hailee Tougas; Steven Chan; Tara Shahrvini; Alvaro Gonzalez; Ruth Chun Reyes; Michelle Burke Parish; Peter Yellowlees
Journal:  JMIR Ment Health       Date:  2022-09-06
  2 in total

北京卡尤迪生物科技股份有限公司 © 2022-2023.