在“应用圣经自然语言处理伙伴关系”(Partnership for Applied Biblical NLP)中,我们正在定义一项关于词法切分(word segmentation)的“共享任务”——基本上,这是一场竞赛,旨在测试哪款软件能最好地完成这一任务。我们的目标是开发出能更出色地完成此任务的 NLP 软件。我们正在考虑使用 Paratext 在 Wordanalyses.xml 中采用的格式,并且我们需要多种语言的高质量数据。
我拥有一些语言的文件,在这些项目中我至少是观察者,但我需要更多,特别是那些在 Interlinearizer 或 Wordlist 中已对形态学(morphology)做了大量工作的语言。
我们对广泛使用的语言以及形态学复杂且代表不同语言类型的语言都感兴趣。如果您有愿意分享的数据,请告诉我!
[Email Removed]
In the Partnership for Applied Biblical NLP, we are defining a “shared task” for word segmentation - basically, a contest to see which software can do it best. The goal is to create NLP software that can do it better. We are considering using the format Paratext uses in Wordanalyses.xml, and we need high quality data for a variety of languages.
I have these files for some languages where I am at least an observer on the project, but I need more, particularly for languages where a lot of work has been done on the morphology in either the Interlinearizer or the Wordlist.
We are interested in both languages of wider communication and languages that are morphologically complex and represent different kinds of languages. If you have data you are willing to share, please let me know!
[Email Removed]
机器翻译自 English