• Refine Query
  • Source
  • Publication year
  • to
  • Language
  • 1
  • 1
  • Tagged with
  • 1
  • 1
  • 1
  • 1
  • 1
  • 1
  • 1
  • 1
  • 1
  • 1
  • 1
  • 1
  • 1
  • 1
  • 1
  • About
  • The Global ETD Search service is a free service for researchers to find electronic theses and dissertations. This service is provided by the Networked Digital Library of Theses and Dissertations.
    Our metadata is collected from universities around the world. If you manage a university/consortium/country archive and want to be added, details can be found on the NDLTD website.
1

應用平行語料建構中文斷詞組件 / Applications of Parallel Corpora for Chinese Segmentation

王瑞平, Wang, Jui Ping Unknown Date (has links)
在本論文,我們建構一個基於中英平行語料的中文斷詞系統,並透過該系統對不同領域的語料斷詞。提供我們的系統不同領域的中英平行語料後,系統可以自動化地產生品質不錯的訓練語料,以節省透過人工斷詞方式取得訓練語料所耗費的時間、人力。 在產生訓練語料時,首先對中英平行語料中的所有中文句,透過查詢中文辭典的方式產生句子的各種斷詞組合,再利用英漢翻譯的資訊處理交集型歧異,將錯誤的斷詞組合去除。此外本研究從中英平行語料中擷取新的中英詞對與未知詞,並分別將其擴充至英漢辭典模組與中文辭典模組,以提升我們的系統之斷詞效能。 我們透過兩部分的實驗進行斷詞效能評估,而在實驗中會使用三種不同領域的實驗語料。在第一部分,我們以人工斷詞的測試語料進行斷詞效能評估。在第二部分,我們藉由漢英翻譯的翻譯品質間接地評估我們的系統之斷詞效能。由實驗結果顯示,我們的系統可以有一定的斷詞效能。 / In this paper, we construct a Chinese word segmentation system which based on Chinese-English Parallel Corpus to save time and manpower, and the corpora in different domains can be segmented by our system. By providing Chinese-English Parallel Corpus to our system, training corpus can be automatically produced by our system. Then segmentation model can be trained with the produced training corpus. We use Chinese translation of words in English parallel sentences to solve overlapping ambiguity. We extract translation pairs and unknown words from Chinese-English Parallel Corpus. In evaluation, two different experiments are conducted, and experimental data in three domains are used to evaluate segmentation performance in two experiments. In the first experiment, manually annotated Chinese sentences are used as testing data. In the second experiment, segmentation performance is indirectly indicated by translation quality. Experimental results show that our system achieves acceptable segmentation performance.

Page generated in 0.0276 seconds