幾年前,我記得看過一份 Paratext 通訊,裡面宣布了一項新功能,可協助使用者視覺化文字中的 Unicode 碼位(codepoints)。內容大意是:「我們知道,通常有多種 Unicode 碼位的組合可以用來表示某個特定字元。通常,字元或詞彙的組成方式很大程度上取決於譯者使用的鍵盤以及他們按下的特定鍵。如果您在搜尋結果中找不到詞彙,或在嘗試執行正規表示式(regex)操作時遇到困難,可以使用我們為您製作的這個工具,來視覺化每個詞彙和字元底層的表示方式。」
我熟悉使用 Alt+X 來查看字元的碼位,也熟悉 Tools>Checking Inventories>Characters Inventory,但這些功能都不會顯示包含碼位的完整詞彙——這就是我在該公告的截圖中記得看到的內容。
有人知道可以在哪裡找到這份公告或它所描述的功能嗎?我搜尋了一圈,但找不到。
我最終的目標是更好地理解不同的字元編碼方式和詞彙表示法如何影響語言專案。無論是透過閱讀相關文件,還是與該領域的專家交談。
謝謝!
Years ago I remember seeing a Paratext newsletter with an announcement about a new feature that helps users visualize the Unicode codepoints in the text. It said something like “We know that there is often more than one combination of Unicode codepoints that can be used to represent a given character. Often, the way the character or word is composed depends a lot on the keyboard the translator is using and the particular keystrokes he used. If you’re having trouble finding words in search results or when trying to perform regex operations, you can use this tool we made for you to visualize how each word and character is represented underlyingly.”
I’m familiar with Alt+X to see the codepoints for a character, and also with Tools>Checking Inventories>Characters Inventory, but neither of these show entire words with their codepoints- that’s what I remember seeing in the printscreen in the announcement.
Does anyone know where to find this announcement or the feature it describes? I’ve searched around and can’t find it.
My ultimate goal is to better understand how different ways of encoding characters and representing words affect language projects. Be it through looking through accompanying documentation or talking with experts on the topic.
Thank you!
機器翻譯自 English