हम जिस भाषा पर काम कर रहे हैं, उसमें कुछ उपसर्ग (affixes) और क्लिटिक्स (clitics) को मूल शब्द (stem) से अल्पविराम (hyphens) द्वारा अलग किया जाता है। यह अल्पविराम भाषा सेटिंग्स (language settings) में शब्द-मध्य बिंदु चिह्न (word-medial punctuation) के रूप में परिभाषित किया गया है।
अब, Morphology फ़ील्ड में Paratext कभी-कभी गलत अनुमान लगाता है क्योंकि वह ऐसे अल्पविरामों को नज़रअंदाज़ कर देता है। उदाहरण के लिए, यह ‹khøuwa-te› के लिए /khøuwat-e/ का अनुमान लगाता है। ‹khøuwat› और ‹-e› दोनों मौजूदा morphemes के रूप में हैं, लेकिन यहाँ अल्पविराम के कारण स्पष्ट होना चाहिए कि morphology ‹khøuwa› और ‹-te› होनी चाहिए, जो पहले से ही morphemes के रूप में परिभाषित हैं।
क्या कहीं कोई सेटिंग है जिसे मैं समायोजित कर सकूं ताकि Paratext Morphology guesser किसी भी ऐसे वर्ण श्रृंखला (character string) को एक morpheme के रूप में न माने जिसमें अल्पविराम हो?
अगर नहीं, तो मैं इसी प्रकार की सुविधा का सुझाव देने जा रहा हूँ।
धन्यवाद,
rfkvg
In the language we work some affixes and clitics are separated from the stem by hyphens. This hyphen is defined in the language settings as word-medial punctuation.
Now in the Morphology field Paratext sometimes makes wrong guesses because it ignores such hyphens. For example, it guesses the morphology of ‹khøuwa-te› as /khøuwat-e/. Both ‹khøuwat› and ‹-e› do exist as morphemes, but here it should be clear because of the hyphen that the morphology should be ‹khøuwa› and ‹-te›, which are also already defined as morphemes.
Is there a setting anywhere I could adjust so that the Paratext Morphology guesser will not treat any character string with a hyphen in it as one morpheme?
If not, I’m going to suggest such a feature.
Thanks,
rfkvg
English से मशीन-अनुवादित