Preparatory Work on Automatic Extraction of Bilingual Multi-Word Units from Parallel Corpora

Automatic extraction of bilingual Multi-Word Units is an important subject of research in the automatic bilingual corpus alignment field. There are many cases of single source words corresponding to target multi-word units. This paper presents an algorithm for the automatic alignment of single source words and target multi-word units from a sentence-aligned parallel spoken language corpus. On the other hand, the output can be also used to extract bilingual multi-word units. The problem with previous approaches is that the retrieval results mainly depend on the identification of suitable Bi-grams to initiate the iterative process. To extract multi-word units, this algorithm utilizes the normalized association score difference of multi target words corresponding to the same single source word, and then utilizes the average association score to align the single source words and target multi-word units. The algorithm is based on the Local Bests algorithm supplemented by two heuristic strategies: excluding words in a stop-list and preferring longer multi-word units.

並列關鍵字

bilingual alignment ； multiword unit ； translation lexicon ； average association score ； normalized association score difference

參考文獻

Chen, X. H.(1999).Automatic Analysis of Contemporary Chinese Using Visual C++.

Google Scholar

Fung, P.(1995).Proceedings of the 33rd Annual Meeting of the Association for Computational Linguistics.

Google Scholar

Patrick, Kenneth, K. K. W., K. W.(1990).Word Association Norms, Mutual Information and Lexicography.Computational Linguistics.16(1),22-29.

Google Scholar

Haruno, M.,Ikehara, S.,Yamazaki, Takio(1996).Proceedings on the 16th International Conference on Computational Linguistics(COLING96).

Google Scholar

Hiemstra, D.(1996).Using Statistical Methods to Create a Bilingual Dictionary.

Google Scholar

國際替代計量

Preparatory Work on Automatic Extraction of Bilingual Multi-Word Units from Parallel Corpora

全文下載

主題瀏覽