0 votes
1,5k vues

J'ai un collègue qui a un NT publié, mais pas de dictionnaire. Il est intéressé par la création d'un dictionnaire, mais aimerait éviter autant que possible le « travail lourd ». Existe-t-il un moyen de transférer la liste de mots de Paratext vers FLEx (par exemple, « Exporter en XML ») ? Est-ce recommandé ? Y a-t-il d'autres recommandations ? Merci pour vos conseils.

Traduit automatiquement depuis English
Paratext par (642 points) | 1,5k vues

8 Réponses

+1 vote
Meilleure réponse

Paul - il y a beaucoup de réflexions à ce sujet. Certains disent qu'il ne faut pas
construire un dictionnaire à partir de texte traduit. Cela dit, vous pourriez
exporter la liste de mots en XML et essayer de l'importer dans FLEx. Cependant,
ceci vous donnerait un désordre. Vous auriez plusieurs temps et
formes différents. Je vous suggère de consulter le site Rapid Word Collection :
http://rapidwords.net/

anon848905
Coordinateur des technologies linguistiques pour la région des Amériques
[Email Removed]
[Phone Removed]Bureau à JAARS)
[Phone Removed]Mobile)
Nom Skype : anon848905

Traduit automatiquement depuis English
par (9,9k points)

Merci anon848905. Nous connaissons bien le RWC (j'ai collecté environ 4 000 mots de cette manière et j'ai aidé quelques collègues à s'y mettre également). Les collègues en question ont passé plus de vingt ans à travailler dans cette langue et cherchent un moyen d'importer des formes de mots dans FLEx afin de pouvoir commencer à utiliser les outils d'édition en masse (Bulk Edit) disponibles là-bas. Rapid Word Collection Lite ? Je suppose que j'essaie de leur épargner la corvée de saisir ce qui sera probablement plus de 5 000 radicaux. Pour l'instant, c'est l'obstacle qui les empêche d'avancer… Je continuerai à chercher d'autres options.

Traduit automatiquement depuis English

Et si vous extrayiez la liste de mots de Paratext et créiez un document
texte (ou peut-être 50 documents de 100 mots chacun) et que vous donniez ensuite
chaque document à FLEx et n'ajoutiez au dictionnaire que les formes de mots qui
y appartiennent ? (Selon la langue, vous pourriez prendre le temps de
configurer une analyse morphologique, afin que « acted » soit analysé comme « act » + « -ed », etc.)

Traduit automatiquement depuis English

Voici un extrait de ma liste de mots Paratext exportée en XML :

Je ne pense pas que ce serait difficile de faire quelques passages Recherche / Remplacer pour le rendre utile :
Rechercher : <item word=" -> \lx[espace]
Rechercher : spelling=“Correct” -> [rien]
Je ne suis pas sûr de quoi faire avec les morphèmes, cependant… Peut-être que cela devrait être fait dans FLEx.

Je vais leur soumettre cela et voir si c'est le point de départ qu'ils recherchent. Merci.

Traduit automatiquement depuis English

Euh… Ne peut-on pas « coller » dans ce forum ? J'ai collé une section de texte dans cette fenêtre et elle s'est affichée pendant la composition, mais elle n'apparaît pas une fois que j'ai publié…

Traduit automatiquement depuis English
0 votes

Vous avez probablement collé du texte qui a été interprété d'une manière spéciale par le formatteur Markdown ou qui était invalide et a donc été supprimé.
Pensez à faire attention au panneau d'aperçu lors de la composition.

Traduit automatiquement depuis English
par [Expert]
(16,7k points)

Laissez-moi réessayer, en supprimant le < initial de chaque ligne :

item word=“abu” spelling=“Correct” morph=“abu” />
item word=“abua” spelling=“Correct” morph=“abu +a” />
item word=“abughami” spelling=“Correct” morph=“abu +ghami” />
item word=“abughamu” spelling=“Correct” morph=“abu +ghamu” />
item word=“abughinia” spelling=“Unknown” morph=“abu +ghini +a” />
item word=“abughita” spelling=“Correct” morph=“abu +ghita” />
item word=“abugho” spelling=“Correct” morph=“abu +gho” />
item word=“abui” spelling=“Correct” morph=“abu +i” />
item word=“abukoira” spelling=“Correct” morph=“abu +ko +ira” />
item word=“abukolu” spelling=“Unknown” morph=“abu +kolu” />
item word=“abura” spelling=“Correct” morph=“abu +ra” />
item word=“aburara” spelling=“Unknown” morph=“abu +ra +ra” />
item word=“abutughami” spelling=“Correct” morph=“abu +tu +ghami” />
item word=“abutuira” spelling=“Correct” morph=“abu +tu +ira” />
item word=“abuu” spelling=“Correct” morph=“abu +u” />

Traduit automatiquement depuis English
0 votes

J'ai donc effectué une opération Recherche & Remplacer pour changer < item word=" en \lx[espace]. Existe-t-il une expression régulière que je pourrais utiliser pour supprimer tout ce qui suit le mot, c'est-à-dire, de la première " à > à la fin :

\lx abu" spelling=“Correct” morph=“abu” />

Traduit automatiquement depuis English
par (642 points)
0 votes

Le suivant devrait fonctionner :
Rechercher : ".+?>
Remplacer : (rien)

P.-S. Je recommande vivement ce site web pour aider à créer des expressions régulières. C'est le meilleur que j'ai trouvé. :smile:

Traduit automatiquement depuis English
par [Expert]
(16,7k points)

Ça a parfaitement fonctionné. Merci ! Je pense que cela suffira pour permettre à mes collègues de commencer. :wink:

Traduit automatiquement depuis English
0 votes

Il existe un utilitaire qui extrait les gloses de l'Interlinéaire Paratext sous forme de fichier SFM ou de classeur Excel. Vous pourriez essayer. http://lingtransoft.info/apps/extract-paratext-interlinear-glosses

Traduit automatiquement depuis English
par [Expert]
(2,9k points)
0 votes

J'ai trouvé un autre utilitaire pour afficher les gloses dans l'Interlinéaire Paratext. http://lingtransoft.info/apps/glossy

Traduit automatiquement depuis English
par [Expert]
(2,9k points)

Glossy peut être utilisé pour extraire les gloses d'une manière qui pourrait être importée dans FLEx, mais ce n'est pas exactement configuré pour cela pour le moment. C'est pour la consultation interactive de lexiques. Je pense que je pourrais trouver un moyen de produire un fichier SFM si les gens étaient intéressés.

Traduit automatiquement depuis English

Y a-t-il eu des progrès pour faire en sorte que Glossy extrait les gloses de l'Interlinéaire PT et les importe dans quelque chose comme un fichier SFM ? J'ai essayé de faire glisser le fichier lexicon.xml sur le programme ParatextLexiconSetup. Mais lorsque j'exécute le programme ParatextLexicon, je ne sais pas quoi saisir dans le champ étiqueté « SFM File ». Je vois une boîte de dialogue intitulée « Convert Paratext Lexicon to SFM ». Ensuite, il vous demande de saisir des noms de fichiers pour deux champs :
Paratext Lexicon, et
SMF File.
Je peux facilement saisir le fichier lexicon.xml. Mais je ne sais pas quoi mettre dans le fichier SFM. Existe-t-il une liste de différents styles de formatage SFM à choisir ?

Traduit automatiquement depuis English

anon088806,

Le fichier SFM est le nom du fichier de sortie du lexicon.xml sélectionné. Vous pouvez mettre le nom de chemin complet que vous voulez. Cela deviendra le fichier sfm souhaité.

kent_schroeder

Traduit automatiquement depuis English

Bonjour kent_schroeder,

Merci beaucoup ! Je n'avais pas idée que la solution pouvait être aussi simple… Grandement apprécié ! anon088806

Traduit automatiquement depuis English
0 votes

Il semble que nous pourrions utiliser un utilitaire pour travailler avec le fichier de liste de mots de Paratext. Quelqu'un ici serait-il intéressé à en écrire un ?

Traduit automatiquement depuis English
par [Expert]
(2,9k points)

Que voulez-vous que l'outil fasse ?

kent_schroeder
Développeur logiciel / Consultant en technologies linguistiques
SIL – Services partagés axés sur l'Afrique

Nairobi, Kenya

Traduit automatiquement depuis English

kent_schroeder, il semble qu'ils veuillent récolter des mots à partir de l'outil de liste de mots pour créer un lexique dans Fieldworks. Les données de l'interlinéaire sont meilleures car elles incluent aussi la glose, mais l'outil de liste de mots aura probablement plus de mots. S'ils pouvaient extraire les mots vernaculaires du fichier html et les marquer avec le marqueur \lx, alors ce fichier pourrait être importé dans une base de données Flex existante. L'utilisateur devrait ensuite passer par chaque entrée et ajuster la forme de surface pour qu'elle soit un véritable lexème et saisir lui-même la glose et la classe grammaticale.

Traduit automatiquement depuis English

kent_schroeder, pouvez-vous donner à votre logiciel la capacité de convertir le fichier Biblicaltermsxyz.xml au format standard ? Quelqu'un dans un autre message souhaite le faire.

Traduit automatiquement depuis English
0 votes

Voici un script que j'ai écrit pour python3 (j'utilise Wasta Linux 18.04). Notez que j'ai codé en dur le chemin pour l'entrée et la sortie dans le code…

#!/usr/bin/python3
# run with python3 convertLexicon2SFM.py3

# BEST TO IMPORT INTO TEXT AND WORDS - not into the dictionary itself.
#   -Import into Standard Format Words and Glosses

#
#  MODIFIED AND ONLY MILDLY TESTED!!!!  It only includes words that 1.) have glosses
# and 2.) are marked as correctly spelled and 3.) exist in the text.
#

import codecs
from lxml import etree
xmldoc = etree.parse("/home/justin/Desktop/Lexicon.xml")
#Paratext Wordlist Export as XML 
wordlist = etree.parse("/home/justin/Desktop/wordlist.xml")
outfile=codecs.open("/home/justin/Desktop/PT7_dictionary_py3.sfm", mode="w", encoding='utf-8')
outfile.write ("\_sh v3.0  400  MDF\n\_DateStampHasFourDigitYear\n\n")
#correctWords=spellings.getroot().findall("Status")
correctWords=wordlist.getroot().findall("item")
wordlistTotal=len(correctWords)
approvedWords=[]
#for index,word in reversed( list( enumerate(correctWords) ) ) :
for index,word in reversed( list( enumerate(correctWords) ) ) :
  if word.attrib['spelling'] == "Correct" :
  #if word.attrib['State'] == "W" :
    del correctWords[index]
    approvedWords.append(word.attrib['word'])
    #approvedWords.append(word.attrib['Word'])

itemList = xmldoc.getroot().findall("Entries/item")
for item in itemList :
  Lexeme=next( item.iter("Lexeme") )
  if (Lexeme.get('Type') == 'Word') and (Lexeme.get('Form') not in approvedWords) :
    print ( 'unused', Lexeme.get('Type'), Lexeme.get('Form'), "wordlist", wordlistTotal, "incorrect", len(correctWords) )
    continue
  print ( 'good', Lexeme.get('Type'), Lexeme.get('Form'), 'count', approvedWords.count(Lexeme.get('Form')), 'unglossed remaining', len(approvedWords) )
  if approvedWords.count(Lexeme.get('Form')) :
    approvedWords.remove( Lexeme.get('Form') )
  outfile.write ("\n\\lx ")
  if Lexeme.get("Type") == "Suffix" :
    outfile.write ("-")
    #outfile.write ("-", end='')
  outfile.write ("%s" % Lexeme.get("Form"))
  #outfile.write ("%s" % Lexeme.get("Form"), end='')
  if Lexeme.get("Type") == "Prefix" :
    outfile.write ("-")
    #outfile.write ("-", end='')
  outfile.write ("\n")
  outfile.write   ("\\co_Eng %s\n" % next( item.iter("Lexeme") ).get("Type"))
  entryList = item.iter("Gloss")
  sense=1
  for element in entryList :
    if element.get("Language") == "English" :
      outfile.write ("\\sn %s\n" % sense)
      outfile.write ("\\ge %s\n" %  element.text )
      sense+=1
    if element.get("Language") == "Korean" :
      outfile.write ("\\sn %s\n" % sense)
      outfile.write ("\\g_Kor %s\n" %  element.text )
      sense+=1

for Lexeme in approvedWords :
  outfile.write ("\n\\lx %s\n" % Lexeme)
  print ("adding unglossed words", len(approvedWords), Lexeme )

outfile.close()
Traduit automatiquement depuis English
par (105 points)
Bonjour,

J'ai adapté le script Python ci-dessus pour le rendre générique : il demande maintenant la liste de mots et le lexique (s'il existe) en entrée, puis génère un fichier SMF à importer dans FLEx.

#!/usr/bin/python3
"""
wordlist_to_sfm.py

Exports a Paratext word list to SFM (Standard Format Markers) format
for import into FLEx (FieldWorks Language Explorer).

Usage:
    python wordlist_to_sfm.py

The script will prompt for:
  - A Paratext word list XML file (e.g. exported from Paratext's word list tool)
  - Optionally, a Paratext Lexicon XML file to supply English glosses
  - An output file path for the resulting .sfm file

Only words marked as "Correct" in the word list are exported by default.
This can be changed by editing the INCLUDE_* settings below.

Output format is MDF (Multi-Dictionary Formatter), the standard SFM dialect used by FLEx for lexicon import.
"""

import codecs
import os
import xml.etree.ElementTree as etree

# --- Settings ----------------------------------------------------------------
# Control which spelling categories are included in the export.
# Set to True to include words in that category, False to exclude them.

INCLUDE_CORRECT   = True   # Words the translator has approved
INCLUDE_UNKNOWN   = False  # Words not yet reviewed
INCLUDE_INCORRECT = False  # Words flagged as misspellings

# Set to True to write the word's corpus frequency as a comment field (\co_count)
INCLUDE_COUNT     = True
# -----------------------------------------------------------------------------

def prompt_path(prompt_text, default):
    """Show a prompt with an optional default value; return the user's input or the default."""
    display = f" [{default}]" if default else ""
    value = input(f"{prompt_text}{display}: ").strip()
return value if value else default

def get_user_inputs():
    """Ask the user for the three file paths needed to run the export."""
    print("Paratext Word List to SFM Exporter")
    print("-----------------------------------\n")

    # Word list is required - keep asking until the file is found
    wordlist_path = prompt_path("Word list file (XML)", "Notsi-WL.xml")
    while not os.path.exists(wordlist_path):
        print(f"  Cannot find '{wordlist_path}'. Please check the path and try again.")
        wordlist_path = prompt_path("Word list file (XML)", "Notsi-WL.xml")

    # Lexicon is optional - skip silently if not found
    lexicon_path = prompt_path("Lexicon file for glosses (press Enter to skip)", "")
    if lexicon_path and not os.path.exists(lexicon_path):
        print(f"  Cannot find '{lexicon_path}'. Continuing without glosses.")
        lexicon_path = ""

    output_path = prompt_path("Output file name", "wordlist_for_flex.sfm")
    return wordlist_path, lexicon_path, output_path

def load_lexicon_glosses(lexicon_path):
    """
    Read English glosses from a Paratext Lexicon XML file.

    Returns a dictionary keyed by lexeme form, where each value contains the lexeme type (Word, Prefix, Suffix) and a list of English gloss strings.
    Returns an empty dictionary if no lexicon path is provided.
    """
    glosses = {}
    if not lexicon_path:
        print("No lexicon provided - entries will be exported without glosses.")
        return glosses

    print(f"Reading lexicon from: {lexicon_path}")
    with open(lexicon_path, "rb") as f:
        data = f.read()
    # Strip any BOM variants (standard UTF-8 BOM, or double-encoded BOM)
    for bom in (b"\xef\xbb\xbf", b"\xc3\xaf\xc2\xbb\xc2\xbf"):
        if data.startswith(bom):
            data = data[len(bom):]
            break
    xmldoc = etree.fromstring(data)

    for item in xmldoc.findall("Entries/item"):
        lexeme = next(item.iter("Lexeme"), None)
        if lexeme is None:
            continue

        form     = lexeme.get("Form", "")
        lex_type = lexeme.get("Type", "Word")  # Word, Prefix, or Suffix

        # Collect all English glosses for this entry (there may be more than one sense)
        entry_glosses = [
            gloss.text
            for gloss in item.iter("Gloss")
            if gloss.get("Language") == "en" and gloss.text
        ]

        if form and entry_glosses:
            glosses[form] = {"type": lex_type, "glosses": entry_glosses}

    print(f"  Found {len(glosses)} entries with English glosses.")
    return glosses

def build_approved_list(wordlist_path):
    """
    Read the Paratext word list XML and return words that pass the INCLUDE_* filters.

    Returns a list of approved word forms and a metadata dictionary containing each word's corpus count, hyphenation, and morphology breakdown.
    """
    wordlist  = etree.parse(wordlist_path)
    all_items = wordlist.getroot().findall("item")
    print(f"Reading word list from: {wordlist_path}")
    print(f"  {len(all_items)} words found in list.")

    approved = []
    metadata = {}

    for item in all_items:
        spelling = item.attrib.get("spelling", "Unknown")
        word     = item.attrib.get("word", "")

        include = (
            (spelling == "Correct"   and INCLUDE_CORRECT)   or
            (spelling == "Unknown"   and INCLUDE_UNKNOWN)   or
            (spelling == "Incorrect" and INCLUDE_INCORRECT)
        )

        if include and word:
            approved.append(word)
            metadata[word] = {
                "count":          item.attrib.get("count", "0"),
                "hyphenation":    item.attrib.get("hyphenation", ""),
                "morphology":     item.attrib.get("morphology", ""),
                "morph_approved": item.attrib.get("morphologyApproved", "False"),
                "specificcase":   item.attrib.get("specificcase", ""),
            }

    print(f"  {len(approved)} words marked as correct and ready to export.")
    return approved, metadata

def write_sfm_entry(outfile, word, lex_type, glosses, meta, include_count):
    """
    Write a single SFM lexicon entry to the output file.

    MDF conventions followed:
      \\lx  - lexeme (headword); affixes get a hyphen on the bound side
      \\sn  - sense number (written once per gloss)
      \\ge  - English gloss for that sense
      \\mr  - morphology string (when approved by the translator)
      \\co_* - comment fields for metadata FLEx doesn't have a standard marker for
    """
    # Write the headword, adding hyphens to show affix attachment points
    outfile.write("\n\\lx ")
    if lex_type == "Suffix":
        outfile.write("-")
    outfile.write(word)
    if lex_type == "Prefix":
        outfile.write("-")
    outfile.write("\n")

    # Record affix type as a comment so it survives the FLEx import
    if lex_type != "Word":
        outfile.write(f"\\co_type {lex_type}\n")

    # Corpus frequency - useful for prioritising entries during dictionary work
if include_count and meta:
        outfile.write(f"\\co_count {meta['count']}\n")

    # Morphology: use the standard \\mr marker if the analysis has been approved,
    # otherwise store it as a comment to avoid importing unverified data
    if meta and meta["morphology"]:
        marker = "\\mr" if meta["morph_approved"] == "True" else "\\co_morph"
        outfile.write(f"{marker} {meta['morphology']}\n")

    # Write one sense block per gloss
    for sense_num, gloss_text in enumerate(glosses, start=1):
        outfile.write(f"\\sn {sense_num}\n")
        outfile.write(f"\\ge {gloss_text}\n")

def main():
    """
    Main export routine.

    Two-pass approach:
      Pass 1 - write entries that have lexicon glosses (richer data first)
      Pass 2 - write remaining approved words that had no lexicon entry
    This ensures glossed entries are not duplicated.
    """
    wordlist_path, lexicon_path, output_path = get_user_inputs()
    print()

    approved_words, metadata = build_approved_list(wordlist_path)
    lexicon_glosses           = load_lexicon_glosses(lexicon_path)

    # Open output file as UTF-8 and write the MDF header
    outfile = codecs.open(output_path, mode="w", encoding="utf-8")
    outfile.write("\\_sh v3.0  400  MDF\n\\_DateStampHasFourDigitYear\n\n")

    remaining    = list(approved_words)  # Words still to be written after pass 1
    with_glosses = 0

    # Pass 1: entries that appear in both the lexicon and the approved word list
    for form, lex_data in lexicon_glosses.items():
        if form not in approved_words:
            print(f"  Skipping '{form}' (in lexicon but not marked as correct in word list)")
            continue
        if form in remaining:
            remaining.remove(form)
        write_sfm_entry(outfile, word=form, lex_type=lex_data["type"],
                        glosses=lex_data["glosses"], meta=metadata.get(form),
                        include_count=INCLUDE_COUNT)
        with_glosses += 1

    # Pass 2: approved words with no lexicon entry - exported without glosses
    for word in remaining:
        write_sfm_entry(outfile, word=word, lex_type="Word", glosses=[],
                        meta=metadata.get(word), include_count=INCLUDE_COUNT)

    total = with_glosses + len(remaining)
    outfile.close()
    print(f"\nExport complete.")
    print(f"  Total entries written : {total}")
    print(f"  With glosses          : {with_glosses}")
    print(f"  Without glosses       : {len(remaining)}")
    print(f"  Output file           : {output_path}")

if __name__ == "__main__":
    main()
Traduit automatiquement depuis English

Questions connexes

+1 vote
3 réponses 435 vues
Exporting the Wordlist to HTML is a nice feature, except that for many users including me, working with HTML is ... that are currently marked as spelled correctly in PT? anon101508
anon101508 117 posée mars 11, 2019
0 votes
6 réponses 684 vues
Ceci a été publié sur la liste de diffusion FLEx aujourd'hui. J'ai un commentaire que j'ajouterai ci-dessous cette question ... mais ils évitent de l'utiliser à cause de ce bug ?
bbryson 105 posée oct. 12, 2021
0 votes
4 réponses 787 vues
J'essaie d'aider un utilisateur avec un problème d'intégration Paratext Flex (lié au Tone Sandhi) et après avoir installé les ... . Le projet Paratext est en version 8. Des idées ?
MSEAIT_LT 478 posée mai 28, 2020
0 votes
3 réponses 310 vues
I have a user who is trying to export from Paratext 8.0 to a FLEx text. If he does a copy/paste in Standard ... 444 | Papua New Guinea [Email Removed].pgmailto:[Email Removed].pg
SIL LSS PNG 411 posée nov. 15, 2017
0 votes
1 réponse 187 vues
Dans Paratext 9.3, est-il possible de : imprimer la liste de mots de tous les livres de mon projet dans un ... lieux et de personnes) dans une seule liste en ordre alphabétique ?
debbodaneejo 118 posée févr. 16, 2023
Welcome to Support Bible, where you can ask questions and receive answers from other members of the community.
I appeal to you, brothers and sisters, in the name of our Lord Jesus Christ, that all of you agree with one another in what you say and that there be no divisions among you, but that you be perfectly united in mind and thought.
1 Corinthians 1:10
3,046 questions
6,006 réponses
5,671 commentaires
2,027 utilisateurs