0
1.5k 次瀏覽

我有一位同事已經出版了新約,但沒有字典。他們有興趣製作一本,但希望盡量減少繁重的基礎工作。有沒有辦法將 Paratext 詞彙表匯入 FLEx(例如「匯出為 XML」)?這樣做是否被推薦?還有其他建議嗎?感謝任何建議。

機器翻譯自 English
Paratext (642 點) 提出 | 1.5k 次瀏覽

8 個回答

+1
最佳回答

Paul - there are lots of thoughts on this. Some people say that you should
not build a dictionary off of translated text. Having said that you could
export the wordlist to xml and then try and import it into FLEx. However
this would give you a mess. You would have multiple different tenses and
forms. I’d suggest you look at the Rapid Word Collection site:
http://rapidwords.net/

anon848905
Americas Area Language Technology Coordinator
[Email Removed]
[Phone Removed]Office at JAARS)
[Phone Removed]Cell)
Skype name: anon848905

舊文章 - 以原文顯示
(9.9k 點) 提出

謝謝 anon848905。我們對 RWC 非常熟悉(我已經用這種方式收集了約 4,000 個詞彙,並幫助幾位同事開始使用它)。具體提到的這些同事在該語言領域工作超過二十年,他們正在尋找一種將詞形匯入 FLEx 的方法,以便開始使用那裡可用的批量編輯(Bulk Edit)工具。Rapid Word Collection Lite?我想我試圖讓他們免去輸入可能超過 5,000 個詞根的枯燥工作。目前這是阻礙他們前進的障礙……我還會繼續尋找其他選項。

機器翻譯自 English

如果從 Paratext 提取字詞清單,建立文字
文件(或建立 50 份各含 100 個字詞的文件),然後將每份
文件交給 FLEx,並只將屬於該詞典的字形
加入詞典,會怎樣呢?(視語言而定,您可以花時間
設定一些解析規則,以便將 “acted” 解析為 “act” + “-ed” 等。)

機器翻譯自 English

這是我從 Paratext 詞彙表匯出為 XML 的節選:

我認為進行幾次查找/替換(Find / Replace)操作使其變得實用並不困難:
查找: <item word=" -> \lx[space]
查找: spelling=“Correct” -> [無]
不過,我不確定該怎麼辦 morphs……也許這必須在 FLEx 中完成。

我會將此發送給他們,看看這是否是他們尋找的起點。謝謝。

機器翻譯自 English

嗯……能不能在這個論壇上「貼上」?我在這個窗口貼上了一段文本,在編輯時顯示出來,但一旦我發布後就不見了……

機器翻譯自 English
0

你可能貼上了被 Markdown 格式化程序以特殊方式解釋的文本,或者是無效的文本,因此被移除了。
編輯時請務必注意預覽窗格。

機器翻譯自 English
[Expert]
(16.7k 點) 提出

讓我再試一次,移除每行開頭的 <:

item word=“abu” spelling=“Correct” morph=“abu” />
item word=“abua” spelling=“Correct” morph=“abu +a” />
item word=“abughami” spelling=“Correct” morph=“abu +ghami” />
item word=“abughamu” spelling=“Correct” morph=“abu +ghamu” />
item word=“abughinia” spelling=“Unknown” morph=“abu +ghini +a” />
item word=“abughita” spelling=“Correct” morph=“abu +ghita” />
item word=“abugho” spelling=“Correct” morph=“abu +gho” />
item word=“abui” spelling=“Correct” morph=“abu +i” />
item word=“abukoira” spelling=“Correct” morph=“abu +ko +ira” />
item word=“abukolu” spelling=“Unknown” morph=“abu +kolu” />
item word=“abura” spelling=“Correct” morph=“abu +ra” />
item word=“aburara” spelling=“Unknown” morph=“abu +ra +ra” />
item word=“abutughami” spelling=“Correct” morph=“abu +tu +ghami” />
item word=“abutuira” spelling=“Correct” morph=“abu +tu +ira” />
item word=“abuu” spelling=“Correct” morph=“abu +u” />

機器翻譯自 English
0

所以我已經進行了查找和替換,將 < item word=" 更改為 \lx[space]。我可以使用什麼正則表達式來去除詞彙之後的所有內容,即從第一個 " 到結尾的 >:

\lx abu" spelling=“Correct” morph=“abu” />

機器翻譯自 English
(642 點) 提出
0

以下應該有效:
查找: ".+?>
替換: (無)

附註:我強烈推薦 這個網站 來幫助創建正則表達式。這是我找到的最好的。 :smile:

機器翻譯自 English
[Expert]
(16.7k 點) 提出

完美運作。謝謝!我想這足以讓我的同事開始工作。 :wink:

機器翻譯自 English
0

有一個工具可以從 Paratext 逐字對照(Interlinear)中提取釋義,作為 SFM 文件或 Excel 電子表格。可以試試那個。 http://lingtransoft.info/apps/extract-paratext-interlinear-glosses

機器翻譯自 English
[Expert]
(2.9k 點) 提出
0

我找到了另一個用於查看 Paratext 逐字對照工具(Interlinearizer)中釋義的工具。 http://lingtransoft.info/apps/glossy

機器翻譯自 English
[Expert]
(2.9k 點) 提出

Glossy 可以用於提取釋義,以便匯入 FLEx,但目前它並非完全為此目的而設置。它是用於交互式詞典瀏覽的。我想如果人們有興趣,我可以想出一種方法來生成 SFM 文件。

機器翻譯自 English

在讓 Glossy 從 PT 逐字對照工具(Interlineariser)中提取釋義並匯入類似 SFM 文件方面,有任何進展嗎?我嘗試將文件 lexicon.xml 拖放到 ParatextLexiconSetup 程序上。但是當我運行 ParatextLexicon 程序時,我不知道該在標有 ‘SFM File’ 的欄位中輸入什麼。我看到一個標題為 ‘Convert Paratext Lexicon to SFM’ 的對話框。然後它提示你為兩個欄位輸入文件名:
Paratext Lexicon,以及
SMF File.
我可以輕鬆輸入 lexicon.xml 文件。但我不知道 SFM 文件該填什麼。是否有一個不同的 SFM 格式樣式列表可供選擇?

機器翻譯自 English

anon088806,

SFM File 是選定的 lexicon.xml 的輸出文件名。你可以輸入任何你想要的完整路徑名。這將成為所需的 sfm 文件。

kent_schroeder

機器翻譯自 English

嗨 kent_schroeder,

非常感謝!我完全沒想到解決方案可以這麼簡單……非常感激! anon088806

機器翻譯自 English
0

看來我們需要一個工具來處理 Paratext 詞彙表文件。那裡有人有興趣寫一個嗎?

機器翻譯自 English
[Expert]
(2.9k 點) 提出

你希望這個工具做什麼?

kent_schroeder
軟件開發人員/語言技術顧問
SIL – 非洲重點共享服務

肯尼亞,內羅畢

機器翻譯自 English

kent_schroeder,看來他們希望從詞彙表工具中收集詞彙,以便在 Fieldworks 中製作詞典。逐字對照數據更好,因為它還有釋義,但詞彙表工具很可能會有更多的詞彙。如果他們能將土著語言詞彙從 html 文件中剝離並標記為 \lx 標記,那麼該文件就可以匯入現有的 Flex 數據庫。然後用戶必須檢查每個條目,將表面形式調整為真正的詞根,並自行輸入釋義和詞性。

機器翻譯自 English

kent_schroeder,你能讓你的軟件具備將 Biblicaltermsxyz.xml 文件轉換為標準格式的能力嗎?另一個帖子中有人希望這樣做。

機器翻譯自 English
0

這是我為 python3 編寫的一個腳本(我使用 Wasta Linux 18.04)。注意,我在代碼中硬編碼了輸入和輸出路徑……

#!/usr/bin/python3
# run with python3 convertLexicon2SFM.py3

# BEST TO IMPORT INTO TEXT AND WORDS - not into the dictionary itself.
#   -Import into Standard Format Words and Glosses

#
#  MODIFIED AND ONLY MILDLY TESTED!!!!  It only includes words that 1.) have glosses
# and 2.) are marked as correctly spelled and 3.) exist in the text.
#

import codecs
from lxml import etree
xmldoc = etree.parse("/home/justin/Desktop/Lexicon.xml")
#Paratext Wordlist Export as XML 
wordlist = etree.parse("/home/justin/Desktop/wordlist.xml")
outfile=codecs.open("/home/justin/Desktop/PT7_dictionary_py3.sfm", mode="w", encoding='utf-8')
outfile.write ("\_sh v3.0  400  MDF\n\_DateStampHasFourDigitYear\n\n")
#correctWords=spellings.getroot().findall("Status")
correctWords=wordlist.getroot().findall("item")
wordlistTotal=len(correctWords)
approvedWords=[]
#for index,word in reversed( list( enumerate(correctWords) ) ) :
for index,word in reversed( list( enumerate(correctWords) ) ) :
  if word.attrib['spelling'] == "Correct" :
  #if word.attrib['State'] == "W" :
    del correctWords[index]
    approvedWords.append(word.attrib['word'])
    #approvedWords.append(word.attrib['Word'])

itemList = xmldoc.getroot().findall("Entries/item")
for item in itemList :
  Lexeme=next( item.iter("Lexeme") )
  if (Lexeme.get('Type') == 'Word') and (Lexeme.get('Form') not in approvedWords) :
    print ( 'unused', Lexeme.get('Type'), Lexeme.get('Form'), "wordlist", wordlistTotal, "incorrect", len(correctWords) )
    continue
  print ( 'good', Lexeme.get('Type'), Lexeme.get('Form'), 'count', approvedWords.count(Lexeme.get('Form')), 'unglossed remaining', len(approvedWords) )
  if approvedWords.count(Lexeme.get('Form')) :
    approvedWords.remove( Lexeme.get('Form') )
  outfile.write ("\n\\lx ")
  if Lexeme.get("Type") == "Suffix" :
    outfile.write ("-")
    #outfile.write ("-", end='')
  outfile.write ("%s" % Lexeme.get("Form"))
  #outfile.write ("%s" % Lexeme.get("Form"), end='')
  if Lexeme.get("Type") == "Prefix" :
    outfile.write ("-")
    #outfile.write ("-", end='')
  outfile.write ("\n")
  outfile.write   ("\\co_Eng %s\n" % next( item.iter("Lexeme") ).get("Type"))
  entryList = item.iter("Gloss")
  sense=1
  for element in entryList :
    if element.get("Language") == "English" :
      outfile.write ("\\sn %s\n" % sense)
      outfile.write ("\\ge %s\n" %  element.text )
      sense+=1
    if element.get("Language") == "Korean" :
      outfile.write ("\\sn %s\n" % sense)
      outfile.write ("\\g_Kor %s\n" %  element.text )
      sense+=1

for Lexeme in approvedWords :
  outfile.write ("\n\\lx %s\n" % Lexeme)
  print ("adding unglossed words", len(approvedWords), Lexeme )

outfile.close()
機器翻譯自 English
(105 點) 提出
問候,

我將上述 Python 腳本進行通用化處理,使其會提示輸入詞表以及(如果有的話)詞典,並輸出一個 SFM 檔案以匯入 FLEx。

#!/usr/bin/python3
"""
wordlist_to_sfm.py

將 Paratext 詞表匯出為 SFM(標準格式標記)格式,
以便匯入 FLEx(FieldWorks Language Explorer)。

用法:
    python wordlist_to_sfm.py

該腳本將提示輸入:
  - Paratext 詞表 XML 檔案(例如從 Paratext 的詞表工具匯出)
  - 可選的 Paratext 詞典 XML 檔案,用於提供英文釋義
  - 生成的 .sfm 檔案的輸出路徑

預設情況下,只匯出在詞表中标記為「正確」的詞。
這可以通過編輯下面的 INCLUDE_* 設定來更改。

輸出格式為 MDF(多詞典格式化器),這是 FLEx 用於詞典匯入的標準 SFM 方言。
"""

import codecs
import os
import xml.etree.ElementTree as etree

# --- 設定 ----------------------------------------------------------------
# 控制哪些拼寫類別包含在匯出中。
# 設為 True 以包含該類別中的詞,設為 False 以排除它們。

INCLUDE_CORRECT   = True   # 翻譯者已批准的詞
INCLUDE_UNKNOWN   = False  # 尚未審核的詞
INCLUDE_INCORRECT = False  # 被標記為拼寫錯誤的詞

# 設為 True 以將詞的語料庫頻率寫為註解欄位(\co_count)
INCLUDE_COUNT     = True
# -----------------------------------------------------------------------------

def prompt_path(prompt_text, default):
    """顯示帶有可選預設值的提示;返回用戶輸入或預設值。"""
    display = f" [{default}]" if default else ""
    value = input(f"{prompt_text}{display}: ").strip()
return value if value else default

def get_user_inputs():
    """詢問用戶運行匯出所需的三個檔案路徑。"""
    print("Paratext 詞表至 SFM 匯出器")
    print("-----------------------------------\n")

    # 詞表是必需的 - 持續詢問直到找到檔案
    wordlist_path = prompt_path("詞表檔案 (XML)", "Notsi-WL.xml")
    while not os.path.exists(wordlist_path):
        print(f"  找不到 '{wordlist_path}'。請檢查路徑並重試。")
        wordlist_path = prompt_path("詞表檔案 (XML)", "Notsi-WL.xml")

    # 詞典是可選的 - 如果找不到則靜默跳過
    lexicon_path = prompt_path("用於釋義的詞典檔案(按 Enter 鍵跳過)", "")
    if lexicon_path and not os.path.exists(lexicon_path):
        print(f"  找不到 '{lexicon_path}'。繼續執行,但不包含釋義。")
        lexicon_path = ""

    output_path = prompt_path("輸出檔案名稱", "wordlist_for_flex.sfm")
    return wordlist_path, lexicon_path, output_path

def load_lexicon_glosses(lexicon_path):
    """
    從 Paratext 詞典 XML 檔案中讀取英文釋義。

    返回一個以詞素形式為鍵的字典,其中每個值包含詞素類型(Word, Prefix, Suffix)和一個英文釋義字串列表。
    如果未提供詞典路徑,則返回空字典。
    """
    glosses = {}
    if not lexicon_path:
        print("未提供詞典 - 條目將在不包含釋義的情況下匯出。")
        return glosses

    print(f"正在從以下位置讀取詞典:{lexicon_path}")
    with open(lexicon_path, "rb") as f:
        data = f.read()
    # 去除任何 BOM 變體(標準 UTF-8 BOM,或雙重編碼的 BOM)
    for bom in (b"\xef\xbb\xbf", b"\xc3\xaf\xc2\xbb\xc2\xbf"):
        if data.startswith(bom):
            data = data[len(bom):]
            break
    xmldoc = etree.fromstring(data)

    for item in xmldoc.findall("Entries/item"):
        lexeme = next(item.iter("Lexeme"), None)
        if lexeme is None:
            continue

        form     = lexeme.get("Form", "")
        lex_type = lexeme.get("Type", "Word")  # Word, Prefix, 或 Suffix

        # 收集此條目的所有英文釋義(可能有多個義項)
        entry_glosses = [
            gloss.text
            for gloss in item.iter("Gloss")
            if gloss.get("Language") == "en" and gloss.text
        ]

        if form and entry_glosses:
            glosses[form] = {"type": lex_type, "glosses": entry_glosses}

    print(f"  找到 {len(glosses)} 個帶有英文釋義的條目。")
    return glosses

def build_approved_list(wordlist_path):
    """
    讀取 Paratext 詞表 XML 並返回通過 INCLUDE_* 過濾器的詞。

    返回一個已批准詞形式的列表,以及一個包含每個詞的語料庫計數、連字符和形態學分解的元數據字典。
    """
    wordlist  = etree.parse(wordlist_path)
    all_items = wordlist.getroot().findall("item")
    print(f"正在從以下位置讀取詞表:{wordlist_path}")
    print(f"  在列表中發現 {len(all_items)} 個詞。")

    approved = []
    metadata = {}

    for item in all_items:
        spelling = item.attrib.get("spelling", "Unknown")
        word     = item.attrib.get("word", "")

        include = (
            (spelling == "Correct"   and INCLUDE_CORRECT)   or
            (spelling == "Unknown"   and INCLUDE_UNKNOWN)   or
            (spelling == "Incorrect" and INCLUDE_INCORRECT)
        )

        if include and word:
            approved.append(word)
            metadata[word] = {
                "count":          item.attrib.get("count", "0"),
                "hyphenation":    item.attrib.get("hyphenation", ""),
                "morphology":     item.attrib.get("morphology", ""),
                "morph_approved": item.attrib.get("morphologyApproved", "False"),
                "specificcase":   item.attrib.get("specificcase", ""),
            }

    print(f"  {len(approved)} 個詞被標記為正確並準備匯出。")
    return approved, metadata

def write_sfm_entry(outfile, word, lex_type, glosses, meta, include_count):
    """
    將單個 SFM 詞典條目寫入輸出檔案。

    遵循的 MDF 慣例:
      \\lx  - 詞素(標題詞);詞綴在結合側加連字符
      \\sn  - 義項編號(每個釋義寫一次)
      \\ge  - 該義項的英文釋義
      \\mr  - 形態學字串(當被翻譯者批准時)
      \\co_* - 用於 FLEx 沒有標準標記的元數據的註解欄位
    """
    # 寫入標題詞,添加連字符以顯示詞綴結合點
    outfile.write("\n\\lx ")
    if lex_type == "Suffix":
        outfile.write("-")
    outfile.write(word)
    if lex_type == "Prefix":
        outfile.write("-")
    outfile.write("\n")

    # 將詞綴類型記錄為註解,以便在 FLEx 匯入後保留
    if lex_type != "Word":
        outfile.write(f"\\co_type {lex_type}\n")

    # 語料庫頻率 - 對於在詞典工作期間優先處理條目很有用
if include_count and meta:
        outfile.write(f"\\co_count {meta['count']}\n")

    # 形態學:如果分析已獲批准,則使用標準 \\mr 標記,
    # 否則將其存儲為註解,以避免匯入未經驗證的數據
    if meta and meta["morphology"]:
        marker = "\\mr" if meta["morph_approved"] == "True" else "\\co_morph"
        outfile.write(f"{marker} {meta['morphology']}\n")

    # 為每個釋義寫入一個義項塊
    for sense_num, gloss_text in enumerate(glosses, start=1):
        outfile.write(f"\\sn {sense_num}\n")
        outfile.write(f"\\ge {gloss_text}\n")

def main():
    """
    主要匯出例程。

    兩遍處理方法:
      第一遍 - 寫入具有詞典釋義的條目(先寫入更豐富的數據)
      第二遍 - 寫入剩餘的已批准詞,這些詞沒有詞典條目
    這確保了帶釋義的條目不會重複。
    """
    wordlist_path, lexicon_path, output_path = get_user_inputs()
    print()

    approved_words, metadata = build_approved_list(wordlist_path)
    lexicon_glosses           = load_lexicon_glosses(lexicon_path)

    # 以 UTF-8 打開輸出檔案並寫入 MDF 標頭
    outfile = codecs.open(output_path, mode="w", encoding="utf-8")
    outfile.write("\\_sh v3.0  400  MDF\n\\_DateStampHasFourDigitYear\n\n")

    remaining    = list(approved_words)  # 第一遍之後仍需寫入的詞
    with_glosses = 0

    # 第一遍:同時出現在詞典和已批准詞表中的條目
    for form, lex_data in lexicon_glosses.items():
        if form not in approved_words:
            print(f"  跳過 '{form}'(在詞典中但未在詞表中标記為正確)")
            continue
        if form in remaining:
            remaining.remove(form)
        write_sfm_entry(outfile, word=form, lex_type=lex_data["type"],
                        glosses=lex_data["glosses"], meta=metadata.get(form),
                        include_count=INCLUDE_COUNT)
        with_glosses += 1

    # 第二遍:沒有詞典條目的已批准詞 - 在不包含釋義的情況下匯出
    for word in remaining:
        write_sfm_entry(outfile, word=word, lex_type="Word", glosses=[],
                        meta=metadata.get(word), include_count=INCLUDE_COUNT)

    total = with_glosses + len(remaining)
    outfile.close()
    print(f"\nExport 完成。")
    print(f"  寫入的總條目數 : {total}")
    print(f"  帶有釋義的      : {with_glosses}")
    print(f"  沒有釋義的      : {len(remaining)}")
    print(f"  輸出檔案        : {output_path}")

if __name__ == "__main__":
    main()
機器翻譯自 English

相關問題

+1
3 個回答 435 次瀏覽
Exporting the Wordlist to HTML is a nice feature, except that for many users including me, working with HTML is ... that are currently marked as spelled correctly in PT? anon101508
anon101508 117 提出 已提問 3月 11, 2019
0
6 個回答 682 次瀏覽
這是今天發佈在 FLEx 清單上的內容 我有一個評論,將添加在這個引用的問題下方: 2021 年 10 月 12 日,下午 3:29,jeffh [Email Removed] 寫道: 有人能解釋一下 Paratext 和 FLEx 之間的目前連線嗎?我知 ... 逐字對譯資料庫),他們會「丟失數據」嗎? 有多少人想要使用此功能,但因為這個錯誤而避免使用它?
bbryson 105 提出 已提問 10月 12, 2021
0
4 個回答 786 次瀏覽
我正在協助一位使用者處理 Paratext 與 Flex 的整合問題(與聲調變調 Tone Sandhi 有關),在將 FLEx 和 Paratext 專案都安裝到我的電腦後,我遇到了一個 Paratext 的問題 當我嘗試開啟 Paratext ... FLEx 專案的版本是 8.1.3,可以正常開啟 Paratext 專案的版本是 8 有什麼建議嗎?
MSEAIT_LT 478 提出 已提問 5月 28, 2020
0
3 個回答 310 次瀏覽
I have a user who is trying to export from Paratext 8.0 to a FLEx text. If he does a copy/paste in Standard ... 444 | Papua New Guinea [Email Removed].pgmailto:[Email Removed].pg
SIL LSS PNG 411 提出 已提問 11月 15, 2017
0
1 個回答 184 次瀏覽
在 Paratext 9.3 中,是否可能: 將我專案中所有書卷的詞彙表列印到一個檔案中,以便我按字母順序排列? 將所有專有名詞(地名和人名)按字母順序列印到一個清單中?
debbodaneejo 118 提出 已提問 2月 16, 2023
Welcome to Support Bible, where you can ask questions and receive answers from other members of the community.
Finally, all of you, be like-minded, be sympathetic, love one another, be compassionate and humble.
1 Peter 3:8
3,045 個問題
6,005 個回答
5,671 則評論
2,026 位使用者