Building an Italian-Chinese Parallel Corpus for Machine Translation from the Web

Tse, R.; Mirri, S.; Tang, S. -K.; Pau, G.; Salomoni, P.

doi:10.1145/3411170.3411258

In an increasingly globalized world, being able to understand texts in different languages (even more so in different alphabets and charsets) has become a necessity. This can be strategic even while moving and travelling across different countries, characterized by different languages. With this in mind, bilingual corpora become critical resources since they are the basis of every state-of-the-art automatic translation system; moreover, building a parallel corpus is usually a complex and very expensive operation. This paper describes an innovative approach we have defined and adopted to automatically build an Italian-Chinese parallel corpus, with the aim of using it for training an Italian-Chinese Neural Machine Translation. Our main idea is to scrape parallel texts from the Web: we defined a general pipeline, describing each specific step from the selection of the appropriate data sources to the sentence alignment method. A final evaluation was conducted to evaluate the goodness of our approach and its results show that 90% of the sentences were correctly aligned. The corpus we have obtained consists of more than 6,000 sentence pairs (Italian and Chinese), which are the basis for building a Machine Translation system.

Tse R., Mirri S., Tang S.-K., Pau G., Salomoni P. (2020). Building an Italian-Chinese Parallel Corpus for Machine Translation from the Web. ;2 Penn Plaza, Suite 701 : Association for Computing Machinery [10.1145/3411170.3411258].

Building an Italian-Chinese Parallel Corpus for Machine Translation from the Web

Tse R.;Mirri S.;Tang S. -K.;Pau G.;Salomoni P.

2020

Abstract

In an increasingly globalized world, being able to understand texts in different languages (even more so in different alphabets and charsets) has become a necessity. This can be strategic even while moving and travelling across different countries, characterized by different languages. With this in mind, bilingual corpora become critical resources since they are the basis of every state-of-the-art automatic translation system; moreover, building a parallel corpus is usually a complex and very expensive operation. This paper describes an innovative approach we have defined and adopted to automatically build an Italian-Chinese parallel corpus, with the aim of using it for training an Italian-Chinese Neural Machine Translation. Our main idea is to scrape parallel texts from the Web: we defined a general pipeline, describing each specific step from the selection of the appropriate data sources to the sentence alignment method. A final evaluation was conducted to evaluate the goodness of our approach and its results show that 90% of the sentences were correctly aligned. The corpus we have obtained consists of more than 6,000 sentence pairs (Italian and Chinese), which are the basis for building a Machine Translation system.

Scheda breve

Scheda completa

Scheda completa (DC)

	Anno
	
				2020
			
	Titolo del volume
	
				ACM International Conference Proceeding Series
			
	Pagina iniziale
	
				265
			
	Pagina finale
	
				268
			
	Codice DOI
	
				https://dx.doi.org/10.1145/3411170.3411258
			
	Citazione
	
				Tse R.,  Mirri S.,  Tang S.-K.,  Pau G.,  Salomoni P. (2020). Building an Italian-Chinese Parallel Corpus for Machine Translation from the Web. ;2 Penn Plaza, Suite 701 : Association for Computing Machinery [10.1145/3411170.3411258].
			
	Tutti gli autori
	
						Tse R.; Mirri S.; Tang S.-K.; Pau G.; Salomoni P.
					
	Appare nelle tipologie:
	
				4.01 Contributo in Atti di convegno

File in questo prodotto:

Eventuali allegati, non sono esposti

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11585/773799

Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni

ND

16

ND

ND

CRIS Current Research Information System