在「應用聖經自然語言處理合作計畫」(Partnership for Applied Biblical NLP)中,我們正在定義一項關於詞彙切分(word segmentation)的「共享任務」(shared task)——基本上是一場比賽,看看哪款軟體能做得最好。我們的目標是開發出能更出色完成此任務的自然語言處理(NLP)軟體。我們正在考慮使用 Paratext 在 Wordanalyses.xml 中採用的格式,因此需要多種語言的高品質數據。
我已經擁有一些語言的這些檔案,這些語言是我至少以觀察者身份參與的專案,但我還需要更多,特別是那些在「逐字對照工具」(Interlinearizer)或「詞彙表」(Wordlist)中已進行大量形態學工作的語言。
我們對通用語言(languages of wider communication)以及形態複雜且代表不同類型語言的語言都感興趣。如果您有願意分享的數據,請告訴我!
[Email Removed]
In the Partnership for Applied Biblical NLP, we are defining a “shared task” for word segmentation - basically, a contest to see which software can do it best. The goal is to create NLP software that can do it better. We are considering using the format Paratext uses in Wordanalyses.xml, and we need high quality data for a variety of languages.
I have these files for some languages where I am at least an observer on the project, but I need more, particularly for languages where a lot of work has been done on the morphology in either the Interlinearizer or the Wordlist.
We are interested in both languages of wider communication and languages that are morphologically complex and represent different kinds of languages. If you have data you are willing to share, please let me know!
[Email Removed]
機器翻譯自 English