Logo image
A parallel corpus of Italian/German legal texts
Conference proceeding   Peer reviewed

A parallel corpus of Italian/German legal texts

Proceedings of the Second International Conference on Language Resources and Evaluation (LREC'00), pp.531-538
Language Resources and Evaluation Conference (Athens, 31/05/2000–02/06/2000)
2000
Handle:
https://hdl.handle.net/10863/53161

Abstract

This paper presents the creation of a parallel corpus of Italian and German legal documents which are translations of one another. The corpus, which contains approximately 5 mio. words, is primarily intended as a resource for (semi-)automatic terminology acquisition. The guidelines of the Corpus Encoding Standard have been applied for encoding structural information, segmentation information, and sentence alignment. Since the parallel texts have a one-to-one correspondence on the sentence level, building a perfect sentence alignment is rather straightforward. As a result of this the corpus constitutes also a valuable testbed for the evaluation of alignment algorithms. The paper discusses the intended use of the corpus, the various phases of corpus compilation, and basic statistics.
url
https://aclanthology.org/L00-1104/View
url
http://www.lrec-conf.org/proceedings/lrec2000/pdf/140.pdfView

Details

Metrics

1 Record Views
Logo image