0 वोट
1.5k व्यूज़

मेरे पास एक सहकर्मी है जिसके पास प्रकाशित न्यू टेस्टामेंट (NT) है, लेकिन कोई शब्दकोश नहीं। वे एक शब्दकोश बनाने में रुचि रखते हैं, लेकिन संभवतः 'भारी मेहनत' से बचना चाहते हैं। क्या ParaTExt शब्दसूची को FLEx में ले जाने का कोई तरीका है (जैसे, 'XML में निर्यात करें')? क्या यह अनुशंसित है? क्या कोई अन्य अनुशंसाएं हैं? किसी भी सलाह के लिए धन्यवाद।

English से मशीन-अनुवादित
Paratext में द्वारा (642 अंक) | 1.5k व्यूज़

8 उत्तर

+1 मत
सर्वोत्तम उत्तर

Paul - there are lots of thoughts on this. Some people say that you should
not build a dictionary off of translated text. Having said that you could
export the wordlist to xml and then try and import it into FLEx. However
this would give you a mess. You would have multiple different tenses and
forms. I’d suggest you look at the Rapid Word Collection site:
http://rapidwords.net/

anon848905
Americas Area Language Technology Coordinator
[Email Removed]
[Phone Removed]Office at JAARS)
[Phone Removed]Cell)
Skype name: anon848905

पुराना पोस्ट - इसकी मूल भाषा में दिखाया जा रहा है
द्वारा (9.9k अंक)

धन्यवाद anon848905। हम RWC के बारे में अच्छी तरह से जानते हैं (मैंने इस तरह से लगभग 4,000 शब्द एकत्र किए हैं और कुछ सहकर्मियों को इससे शुरू करने में भी मदद की है)। जिस विशेष सहकर्मी की बात हो रही है, उन्होंने उस भाषा में काम करने में बीस वर्षों से अधिक का समय बिताया है और वे शब्दरूपों (wordforms) को FLEx में ले जाने का एक तरीका खोज रहे हैं ताकि वे वहाँ उपलब्ध Bulk Edit टूल्स का उपयोग शुरू कर सकें। Rapid Word Collection Lite? मुझे लगता है कि मैं उन्हें 5,000 से अधिक मूल शब्दों (stems) को टाइप करने की कठिनाई से बचाने की कोशिश कर रहा हूँ। अभी यह वह बाधा है जो उन्हें आगे बढ़ने से रोकती है… मैं अन्य विकल्पों की तलाश जारी रखूंगा।

English से मशीन-अनुवादित

What if you extracted the word list from Paratext and created a text
document (or maybe 50 documents of 100 words each) and then gave each
document to FLEx and added to the dictionary only the wordforms that
belong there? (Depending on the language, you could take the time to set
up some parsing, so that “acted” would be parsed as “act” + “-ed” etc.)

पुराना पोस्ट - इसकी मूल भाषा में दिखाया जा रहा है

यहाँ मेरी ParaTExt शब्दसूची का एक अंश है जिसे XML में निर्यात किया गया है:

मुझे नहीं लगता कि इसे उपयोगी बनाने के लिए कुछ Find / Replace चालों को करना कठिन होगा:
खोजें: <item word=" -> \lx[space]
खोजें: spelling=“Correct” -> [कुछ नहीं]
हालाँकि, मुझे morphs के साथ क्या करना है, यह पता नहीं… शायद यह FLEx में किया जाना होगा।

मैं इसे उनके पास ले जाऊंगा और देखूंगा कि क्या यह वह शुरुआत है जिसकी वे तलाश कर रहे हैं। धन्यवाद।

English से मशीन-अनुवादित

हम्म… क्या इस फोरम में 'पेस्ट' नहीं किया जा सकता? मैंने इस विंडो में पाठ का एक खंड पेस्ट किया और यह रचना करते समय दिखाई दिया, लेकिन एक बार पोस्ट करने के बाद यह दिखाई नहीं देता…

English से मशीन-अनुवादित
0 वोट

आपने शायद वह पाठ पेस्ट किया था जिसे Markdown फॉर्मैटर ने विशेष रूप से व्याख्या किया था या जो अमान्य था और इसलिए हटा दिया गया था।
रचना करते समय प्रीव्यू पैनल का ध्यान रखना सुनिश्चित करें।

English से मशीन-अनुवादित
द्वारा [Expert]
(16.7k अंक)

मुझे फिर से कोशिश करने दें, प्रत्येक पंक्ति से प्रारंभिक < हटाकर:

item word=“abu” spelling=“Correct” morph=“abu” />
item word=“abua” spelling=“Correct” morph=“abu +a” />
item word=“abughami” spelling=“Correct” morph=“abu +ghami” />
item word=“abughamu” spelling=“Correct” morph=“abu +ghamu” />
item word=“abughinia” spelling=“Unknown” morph=“abu +ghini +a” />
item word=“abughita” spelling=“Correct” morph=“abu +ghita” />
item word=“abugho” spelling=“Correct” morph=“abu +gho” />
item word=“abui” spelling=“Correct” morph=“abu +i” />
item word=“abukoira” spelling=“Correct” morph=“abu +ko +ira” />
item word=“abukolu” spelling=“Unknown” morph=“abu +kolu” />
item word=“abura” spelling=“Correct” morph=“abu +ra” />
item word=“aburara” spelling=“Unknown” morph=“abu +ra +ra” />
item word=“abutughami” spelling=“Correct” morph=“abu +tu +ghami” />
item word=“abutuira” spelling=“Correct” morph=“abu +tu +ira” />
item word=“abuu” spelling=“Correct” morph=“abu +u” />

English से मशीन-अनुवादित
0 वोट

तो मैंने Find & Replace किया है ताकि < item word=" को \lx[space] में बदला जा सके। क्या कोई नियमित अभिव्यक्ति (regular expression) है जिसका उपयोग मैं शब्द के बाद के सभी को हटाने के लिए कर सकता हूँ, यानी, पहले " से अंत में > तक:

\lx abu" spelling=“Correct” morph=“abu” />

English से मशीन-अनुवादित
द्वारा (642 अंक)
0 वोट

निम्नलिखित काम करेगा:
खोजें: ".+?>
बदलें: (कुछ नहीं)

P.S. मैं नियमित अभिव्यक्तियों बनाने में मदद के लिए इस वेबसाइट की जोर-जोर से सिफारिश करता हूँ। यह सबसे अच्छी वेबसाइट है जिसे मैंने पाया है। :smile:

English से मशीन-अनुवादित
द्वारा [Expert]
(16.7k अंक)

बिल्कुल सही काम किया। धन्यवाद! मुझे लगता है कि यह मेरे सहकर्मियों को शुरू करने के लिए पर्याप्त होगा। :wink:

English से मशीन-अनुवादित
0 वोट

एक उपयोगिता (utility) है जो Paratext Interlinear से glosses को या तो एक SFM फ़ाइल के रूप में या एक Excel Spreadsheet के रूप में निकालती है। शायद उसका प्रयास करें। http://lingtransoft.info/apps/extract-paratext-interlinear-glosses

English से मशीन-अनुवादित
द्वारा [Expert]
(2.9k अंक)
0 वोट

मैंने Paratext Interlinearizer में glosses देखने के लिए एक और उपयोगिता पाई। http://lingtransoft.info/apps/glossy

English से मशीन-अनुवादित
द्वारा [Expert]
(2.9k अंक)

Glossy का उपयोग glosses को निकालने के लिए किया जा सकता है, जिस तरह से उन्हें FLEx में आयात किया जा सकता है, लेकिन इसका अभी-अभी इसका सेटअप नहीं है। यह इंटरैक्टिव शब्दकोश ब्राउज़िंग के लिए है। मुझे लगता है कि यदि लोग रुचि रखते हैं, तो मैं SFM फ़ाइल बनाने का एक तरीका तय कर सकता हूँ।

English से मशीन-अनुवादित

क्या Glossy को PT Interlineariser से glosses निकालने और उन्हें SFM फ़ाइल जैसी किसी चीज़ में आयात करने में कोई प्रगति हुई है? मैंने ParatextLexiconSetup प्रोग्राम पर lexicon.xml फ़ाइल को खींचने का प्रयास किया है। लेकिन जब मैं ParatextLexicon प्रोग्राम चलाता हूँ, तो मुझे पता नहीं कि 'SFM File' लेबल वाले स्लॉट में क्या दर्ज करना है। मुझे 'Convert Paratext Lexicon to SFM' शीर्षक वाला एक संवाद बॉक्स दिखाई देता है। फिर यह आपको दो फ़ील्डों के लिए फ़ाइल नाम दर्ज करने के लिए प्रॉम्प्ट करता है:
Paratext Lexicon, और
SMF File.
मैं आसानी से lexicon.xml फ़ाइल दर्ज कर सकता हूँ। लेकिन मुझे पता नहीं कि SFM फ़ाइल में क्या डालना है। क्या अलग-अलग SFM प्रारूप शैलियों की एक सूची है जिसमें से चुना जा सकता है?

English से मशीन-अनुवादित

anon088806,

SFM File वह नाम है जो चुनी गई lexicon.xml की आउटपुट फ़ाइल है। आप जो भी पूर्ण पथ नाम चाहें डाल सकते हैं। यह वांछित sfm फ़ाइल बन जाएगा।

kent_schroeder

English से मशीन-अनुवादित

नमस्ते kent_schroeder,

बहुत-बहुत धन्यवाद! मुझे पता ही नहीं था कि समाधान इतना सरल हो सकता है… बहुत आभारी हूँ! anon088806

English से मशीन-अनुवादित
0 वोट

लगता है कि हमें Paratext शब्दसूची फ़ाइल के साथ काम करने के लिए एक उपयोगिता की आवश्यकता हो सकती है। क्या कोई वहाँ है जो एक लिखने में रुचि रखता है?

English से मशीन-अनुवादित
द्वारा [Expert]
(2.9k अंक)

आप टूल से क्या करना चाहते हैं?

kent_schroeder
सॉफ्टवेयर डेवलपर/भाषा प्रौद्योगिकी सलाहकार
SIL – अफ्रीका-केंद्रित साझा सेवाएं

नैरोबी, केन्या

English से मशीन-अनुवादित

kent_schroeder, ऐसा लगता है कि वे Fieldworks में एक शब्दकोश बनाने के लिए शब्दसूची टूल से शब्द एकत्रित करना चाहते हैं। इंटरलिनियरेटर डेटा बेहतर है क्योंकि इसमें gloss भी है, लेकिन शब्दसूची टूल में संभवतः अधिक शब्द होंगे। यदि वे vernacular शब्दों को html फ़ाइल से हटा सकते हैं और \lx मार्कर के साथ चिह्नित कर सकते हैं, तो उस फ़ाइल को एक मौजूदा Flex डेटाबेस में आयात किया जा सकता है। उपयोगकर्ता को फिर प्रत्येक प्रविष्टि से गुजरना होगा और सतह रूप (surface form) को एक सच्चे लीक्सेम (lexeme) में समायोजित करना होगा और gloss और भाषा के भाग (part of speech) को स्वयं दर्ज करना होगा।

English से मशीन-अनुवादित

kent_schroeder, क्या आप अपने सॉफ्टवेयर को Biblicaltermsxyz.xml फ़ाइल को मानक प्रारूप में परिवर्तित करने की क्षमता दे सकते हैं? किसी अन्य पोस्ट में कोई ऐसा करना चाहता है।

English से मशीन-अनुवादित
0 वोट

यहाँ एक स्क्रिप्ट है जिसे मैंने python3 के लिए लिखा है (मैं Wasta Linux 18.04 का उपयोग करता हूँ)। ध्यान दें कि मैंने कोड में इनपुट और आउटपुट के लिए पथ हार्ड-कोड किया है…

#!/usr/bin/python3
# run with python3 convertLexicon2SFM.py3

# BEST TO IMPORT INTO TEXT AND WORDS - not into the dictionary itself.
#   -Import into Standard Format Words and Glosses

#
#  MODIFIED AND ONLY MILDLY TESTED!!!!  It only includes words that 1.) have glosses
# and 2.) are marked as correctly spelled and 3.) exist in the text.
#

import codecs
from lxml import etree
xmldoc = etree.parse("/home/justin/Desktop/Lexicon.xml")
#Paratext Wordlist Export as XML 
wordlist = etree.parse("/home/justin/Desktop/wordlist.xml")
outfile=codecs.open("/home/justin/Desktop/PT7_dictionary_py3.sfm", mode="w", encoding='utf-8')
outfile.write ("\_sh v3.0  400  MDF\n\_DateStampHasFourDigitYear\n\n")
#correctWords=spellings.getroot().findall("Status")
correctWords=wordlist.getroot().findall("item")
wordlistTotal=len(correctWords)
approvedWords=[]
#for index,word in reversed( list( enumerate(correctWords) ) ) :
for index,word in reversed( list( enumerate(correctWords) ) ) :
  if word.attrib['spelling'] == "Correct" :
  #if word.attrib['State'] == "W" :
    del correctWords[index]
    approvedWords.append(word.attrib['word'])
    #approvedWords.append(word.attrib['Word'])

itemList = xmldoc.getroot().findall("Entries/item")
for item in itemList :
  Lexeme=next( item.iter("Lexeme") )
  if (Lexeme.get('Type') == 'Word') and (Lexeme.get('Form') not in approvedWords) :
    print ( 'unused', Lexeme.get('Type'), Lexeme.get('Form'), "wordlist", wordlistTotal, "incorrect", len(correctWords) )
    continue
  print ( 'good', Lexeme.get('Type'), Lexeme.get('Form'), 'count', approvedWords.count(Lexeme.get('Form')), 'unglossed remaining', len(approvedWords) )
  if approvedWords.count(Lexeme.get('Form')) :
    approvedWords.remove( Lexeme.get('Form') )
  outfile.write ("\n\\lx ")
  if Lexeme.get("Type") == "Suffix" :
    outfile.write ("-")
    #outfile.write ("-", end='')
  outfile.write ("%s" % Lexeme.get("Form"))
  #outfile.write ("%s" % Lexeme.get("Form"), end='')
  if Lexeme.get("Type") == "Prefix" :
    outfile.write ("-")
    #outfile.write ("-", end='')
  outfile.write ("\n")
  outfile.write   ("\\co_Eng %s\n" % next( item.iter("Lexeme") ).get("Type"))
  entryList = item.iter("Gloss")
  sense=1
  for element in entryList :
    if element.get("Language") == "English" :
      outfile.write ("\\sn %s\n" % sense)
      outfile.write ("\\ge %s\n" %  element.text )
      sense+=1
    if element.get("Language") == "Korean" :
      outfile.write ("\\sn %s\n" % sense)
      outfile.write ("\\g_Kor %s\n" %  element.text )
      sense+=1

for Lexeme in approvedWords :
  outfile.write ("\n\\lx %s\n" % Lexeme)
  print ("adding unglossed words", len(approvedWords), Lexeme )

outfile.close()
English से मशीन-अनुवादित
द्वारा (105 अंक)
नमस्ते,

मैंने ऊपर दिया गया Python स्क्रिप्ट को सामान्य (general) बनाने के लिए संशोधित किया है, जिसमें शब्द सूची (word list) और लैक्सिकन (lexicon) (यदि उपलब्ध हो) के लिए एक प्रॉम्प्ट होगा और FLEx में आयात (import) करने के लिए एक SMF फ़ाइल का आउटपुट होगा।

#!/usr/bin/python3
"""
wordlist_to_sfm.py

Paratext शब्द सूची को SFM (Standard Format Markers) प्रारूप में निर्यात करता है
FLEx (FieldWorks Language Explorer) में आयात के लिए।

उपयोग:
    python wordlist_to_sfm.py

स्क्रिप्ट इनके लिए प्रॉम्प्ट देगी:
  - Paratext शब्द सूची XML फ़ाइल (उदाहरण के लिए, Paratext के शब्द सूची टूल से निर्यात की गई)
  - वैकल्पिक रूप से, अंग्रेज़ी ग्लॉस (glosses) प्रदान करने के लिए Paratext लैक्सिकन XML फ़ाइल
  - परिणामी .sfm फ़ाइल के लिए आउटपुट फ़ाइल पथ

डिफ़ॉल्ट रूप से, केवल वे शब्द निर्यात किए जाते हैं जिन्हें शब्द सूची में "Correct" के रूप में चिह्नित किया गया है।
नीचे दिए गए INCLUDE_* सेटिंग्स को संपादित करके इसे बदला जा सकता है।

आउटपुट प्रारूप MDF (Multi-Dictionary Formatter) है, जो FLEx द्वारा लैक्सिकन आयात के लिए उपयोग किया जाने वाला मानक SFM डायलैक्ट है।
"""

import codecs
import os
import xml.etree.ElementTree as etree

# --- Settings ----------------------------------------------------------------
# नियंत्रित करें कि निर्यात में किन वर्तनी श्रेणियों (spelling categories) को शामिल किया जाए।
# उस श्रेणी के शब्दों को शामिल करने के लिए True सेट करें, उन्हें बाहर रखने के लिए False सेट करें।

INCLUDE_CORRECT   = True   # वे शब्द जिन्हें अनुवादक ने स्वीकृत किया है
INCLUDE_UNKNOWN   = False  # अभी तक समीक्षा नहीं किए गए शब्द
INCLUDE_INCORRECT = False  # वर्तनी त्रुटियों के रूप में चिह्नित शब्द

# शब्द की कॉर्पस आवृत्ति (corpus frequency) को एक कमेंट फ़ील्ड (\co_count) के रूप में लिखने के लिए True सेट करें
INCLUDE_COUNT     = True
# -----------------------------------------------------------------------------

def prompt_path(prompt_text, default):
    """वैकल्पिक डिफ़ॉल्ट मान के साथ एक प्रॉम्प्ट दिखाएं; उपयोगकर्ता का इनपुट या डिफ़ॉल्ट लौटाएं।"""
    display = f" [{default}]" if default else ""
    value = input(f"{prompt_text}{display}: ").strip()
return value if value else default

def get_user_inputs():
    """निर्यात चलाने के लिए आवश्यक तीन फ़ाइल पथों के लिए उपयोगकर्ता से पूछें।"""
    print("Paratext Word List to SFM Exporter")
    print("-----------------------------------\n")

    # शब्द सूची आवश्यक है - फ़ाइल मिलने तक पूछते रहें
    wordlist_path = prompt_path("Word list file (XML)", "Notsi-WL.xml")
    while not os.path.exists(wordlist_path):
        print(f"  Cannot find '{wordlist_path}'. Please check the path and try again.")
        wordlist_path = prompt_path("Word list file (XML)", "Notsi-WL.xml")

    # लैक्सिकन वैकल्पिक है - यदि नहीं मिलता है तो चुपचाप छोड़ दें
    lexicon_path = prompt_path("Lexicon file for glosses (press Enter to skip)", "")
    if lexicon_path and not os.path.exists(lexicon_path):
        print(f"  Cannot find '{lexicon_path}'. Continuing without glosses.")
        lexicon_path = ""

    output_path = prompt_path("Output file name", "wordlist_for_flex.sfm")
    return wordlist_path, lexicon_path, output_path

def load_lexicon_glosses(lexicon_path):
    """
    Paratext लैक्सिकन XML फ़ाइल से अंग्रेज़ी ग्लॉस पढ़ें।

    लैक्सिम रूप (lexeme form) की कुंजी के साथ एक शब्दकोश लौटाता है, जहाँ प्रत्येक मान में लैक्सिम प्रकार (Word, Prefix, Suffix) और अंग्रेज़ी ग्लॉस स्ट्रिंग्स की सूची शामिल होती है।
    यदि कोई लैक्सिकन पथ प्रदान नहीं किया गया है, तो एक खाली शब्दकोश लौटाता है।
    """
    glosses = {}
    if not lexicon_path:
        print("No lexicon provided - entries will be exported without glosses.")
        return glosses

    print(f"Reading lexicon from: {lexicon_path}")
    with open(lexicon_path, "rb") as f:
        data = f.read()
    # किसी भी BOM विविधताओं को हटाएं (मानक UTF-8 BOM, या डबल-एन्कोडेड BOM)
    for bom in (b"\xef\xbb\xbf", b"\xc3\xaf\xc2\xbb\xc2\xbf"):
        if data.startswith(bom):
            data = data[len(bom):]
            break
    xmldoc = etree.fromstring(data)

    for item in xmldoc.findall("Entries/item"):
        lexeme = next(item.iter("Lexeme"), None)
        if lexeme is None:
            continue

        form     = lexeme.get("Form", "")
        lex_type = lexeme.get("Type", "Word")  # Word, Prefix, or Suffix

        # इस प्रविष्टि के लिए सभी अंग्रेज़ी ग्लॉस एकत्र करें (एक से अधिक अर्थ हो सकते हैं)
        entry_glosses = [
            gloss.text
            for gloss in item.iter("Gloss")
            if gloss.get("Language") == "en" and gloss.text
        ]

        if form and entry_glosses:
            glosses[form] = {"type": lex_type, "glosses": entry_glosses}

    print(f"  Found {len(glosses)} entries with English glosses.")
    return glosses

def build_approved_list(wordlist_path):
    """
    Paratext शब्द सूची XML पढ़ें और वे शब्द लौटाएं जो INCLUDE_* फ़िल्टर से गुजरते हैं।

    स्वीकृत शब्द रूपों की सूची और एक मेटाडेटा शब्दकोश लौटाता है जिसमें प्रत्येक शब्द की कॉर्पस गिनती, हाइफ़नेशन (hyphenation), और रूपांतरण (morphology) विभाजन शामिल है।
    """
    wordlist  = etree.parse(wordlist_path)
    all_items = wordlist.getroot().findall("item")
    print(f"Reading word list from: {wordlist_path}")
    print(f"  {len(all_items)} words found in list.")

    approved = []
    metadata = {}

    for item in all_items:
        spelling = item.attrib.get("spelling", "Unknown")
        word     = item.attrib.get("word", "")

        include = (
            (spelling == "Correct"   and INCLUDE_CORRECT)   or
            (spelling == "Unknown"   and INCLUDE_UNKNOWN)   or
            (spelling == "Incorrect" and INCLUDE_INCORRECT)
        )

        if include and word:
            approved.append(word)
            metadata[word] = {
                "count":          item.attrib.get("count", "0"),
                "hyphenation":    item.attrib.get("hyphenation", ""),
                "morphology":     item.attrib.get("morphology", ""),
                "morph_approved": item.attrib.get("morphologyApproved", "False"),
                "specificcase":   item.attrib.get("specificcase", ""),
            }

    print(f"  {len(approved)} words marked as correct and ready to export.")
    return approved, metadata

def write_sfm_entry(outfile, word, lex_type, glosses, meta, include_count):
    """
    आउटपुट फ़ाइल में एकल SFM लैक्सिकन प्रविष्टि लिखें।

    पालन किए गए MDF नियम:
      \\lx  - लैक्सिम (मुख्य शब्द); उपसर्गों को बाध्य पक्ष पर हाइफ़न मिलता है
      \\sn  - अर्थ संख्या (प्रत्येक ग्लॉस के लिए एक बार लिखा जाता है)
      \\ge  - उस अर्थ के लिए अंग्रेज़ी ग्लॉस
      \\mr  - रूपांतरण स्ट्रिंग (जब अनुवादक द्वारा स्वीकृत हो)
      \\co_* - मेटाडेटा के लिए कमेंट फ़ील्ड, जिनके लिए FLEx में कोई मानक मार्कर नहीं है
    """
    # मुख्य शब्द लिखें, उपसर्ग संलग्नता बिंदुओं को दिखाने के लिए हाइफ़न जोड़ें
    outfile.write("\n\\lx ")
    if lex_type == "Suffix":
        outfile.write("-")
    outfile.write(word)
    if lex_type == "Prefix":
        outfile.write("-")
    outfile.write("\n")

    # FLEx आयात के दौरान बने रहने के लिए उपसर्ग प्रकार को एक कमेंट के रूप में दर्ज करें
    if lex_type != "Word":
        outfile.write(f"\\co_type {lex_type}\n")

    # कॉर्पस आवृत्ति - शब्दकोश कार्य के दौरान प्रविष्टियों को प्राथमिकता देने के लिए उपयोगी
if include_count and meta:
        outfile.write(f"\\co_count {meta['count']}\n")

    # रूपांतरण: यदि विश्लेषण स्वीकृत हो गया है तो मानक \\mr मार्कर का उपयोग करें,
    # अन्यथा अस्वीकृत डेटा आयात करने से बचने के लिए इसे एक कमेंट के रूप में संग्रहीत करें
    if meta and meta["morphology"]:
        marker = "\\mr" if meta["morph_approved"] == "True" else "\\co_morph"
        outfile.write(f"{marker} {meta['morphology']}\n")

    # प्रत्येक ग्लॉस के लिए एक अर्थ ब्लॉक लिखें
    for sense_num, gloss_text in enumerate(glosses, start=1):
        outfile.write(f"\\sn {sense_num}\n")
        outfile.write(f"\\ge {gloss_text}\n")

def main():
    """
    मुख्य निर्यात रूटीन।

    दो-चरण दृष्टिकोण:
      चरण 1 - वे प्रविष्टियां लिखें जिनमें लैक्सिकन ग्लॉस हैं (पहले समृद्ध डेटा)
      चरण 2 - शेष स्वीकृत शब्द लिखें जिनमें कोई लैक्सिकन प्रविष्टि नहीं थी
    यह सुनिश्चित करता है कि ग्लॉस वाली प्रविष्टियां दोहराई न जाएं।
    """
    wordlist_path, lexicon_path, output_path = get_user_inputs()
    print()

    approved_words, metadata = build_approved_list(wordlist_path)
    lexicon_glosses           = load_lexicon_glosses(lexicon_path)

    # आउटपुट फ़ाइल को UTF-8 के रूप में खोलें और MDF हेडर लिखें
    outfile = codecs.open(output_path, mode="w", encoding="utf-8")
    outfile.write("\\_sh v3.0  400  MDF\n\\_DateStampHasFourDigitYear\n\n")

    remaining    = list(approved_words)  # चरण 1 के बाद अभी भी लिखे जाने वाले शब्द
    with_glosses = 0

    # चरण 1: वे प्रविष्टियां जो लैक्सिकन और स्वीकृत शब्द सूची दोनों में उपस्थित हैं
    for form, lex_data in lexicon_glosses.items():
        if form not in approved_words:
            print(f"  Skipping '{form}' (in lexicon but not marked as correct in word list)")
            continue
        if form in remaining:
            remaining.remove(form)
        write_sfm_entry(outfile, word=form, lex_type=lex_data["type"],
                        glosses=lex_data["glosses"], meta=metadata.get(form),
                        include_count=INCLUDE_COUNT)
        with_glosses += 1

    # चरण 2: वे स्वीकृत शब्द जिनमें कोई लैक्सिकन प्रविष्टि नहीं है - ग्लॉस के बिना निर्यात किया जाता है
    for word in remaining:
        write_sfm_entry(outfile, word=word, lex_type="Word", glosses=[],
                        meta=metadata.get(word), include_count=INCLUDE_COUNT)

    total = with_glosses + len(remaining)
    outfile.close()
    print(f"\nExport complete.")
    print(f"  Total entries written : {total}")
    print(f"  With glosses          : {with_glosses}")
    print(f"  Without glosses       : {len(remaining)}")
    print(f"  Output file           : {output_path}")

if __name__ == "__main__":
    main()
English से मशीन-अनुवादित

संबंधित प्रश्न

+1 मत
3 उत्तर 435 व्यूज़
Exporting the Wordlist to HTML is a nice feature, except that for many users including me, working with HTML is ... that are currently marked as spelled correctly in PT? anon101508
anon101508 117 पूछा गया मार्च 11, 2019
Paratext में
0 वोट
6 उत्तर 684 व्यूज़
यह आज FLEx सूची पर पोस्ट किया गया था। मेरे पास एक टिप्पणी है जिसे मैं इस उद्धृत प्रश्न के नीचे जोड़ूँगा: 12 अक्टूबर, 2021, ... लेकिन वे इस बग के कारण इसका उपयोग करने से बच रहे हैं?
bbryson 105 पूछा गया अक्टूबर 12, 2021
0 वोट
4 उत्तर 787 व्यूज़
मैं एक उपयोगकर्ता की Paratext Flex एकीकरण समस्या (Tone Sandhi से संबंधित) के समाधान में मदद करने की कोशिश कर रहा ... समस्या के खुलता है। Paratext प्रोजेक्ट 8 में है। कोई सुझाव?
MSEAIT_LT 478 पूछा गया मई 28, 2020
0 वोट
3 उत्तर 310 व्यूज़
I have a user who is trying to export from Paratext 8.0 to a FLEx text. If he does a copy/paste in Standard ... 444 | Papua New Guinea [Email Removed].pgmailto:[Email Removed].pg
SIL LSS PNG 411 पूछा गया नवंबर 15, 2017
Paratext में
0 वोट
1 उत्तर 187 व्यूज़
क्या Paratext 9.3 में यह संभव है: मेरे प्रोजेक्ट की सभी पुस्तकों की शब्दसूची (wordlist) को एक ऐसे फ़ाइल में प्रिंट करना ... नाम) को एक सूची में वर्णानुक्रम में प्रिंट कर सकता हूँ?
debbodaneejo 118 पूछा गया फ़रवरी 16, 2023
Paratext में
Welcome to Support Bible, where you can ask questions and receive answers from other members of the community.
But if we walk in the light, as he is in the light, we have fellowship with one another, and the blood of Jesus, his Son, purifies us from all sin.
1 John 1:7
3,046 प्रश्न
6,006 उत्तर
5,671 टिप्पणियाँ
2,027 उपयोगकर्ता