几年前,我记得看到过一份 Paratext 通讯,其中宣布了一项新功能,可以帮助用户可视化文本中的 Unicode 码点。内容大致是:“我们知道,通常有多种 Unicode 码点组合可以用来表示同一个字符。通常,字符或单词的构成方式很大程度上取决于翻译者使用的键盘以及他们使用的特定按键。如果您在搜索结果中查找单词或尝试执行正则表达式操作时遇到困难,可以使用我们为您制作的这个工具来可视化每个单词和字符底层的表示方式。”
我熟悉使用 Alt+X 来查看字符的码点,也熟悉 Tools>Checking Inventories>Characters Inventory,但这两个功能都不能显示带有码点的完整单词——我记得在公告的截图中看到过这样的功能。
有人知道在哪里可以找到这份公告或它描述的功能吗?我到处搜索过,但找不到。
我的最终目标是更好地理解不同的字符编码和词语表示方式如何影响语言项目。无论是通过查阅相关文档,还是与相关领域的专家交谈。
谢谢!
Years ago I remember seeing a Paratext newsletter with an announcement about a new feature that helps users visualize the Unicode codepoints in the text. It said something like “We know that there is often more than one combination of Unicode codepoints that can be used to represent a given character. Often, the way the character or word is composed depends a lot on the keyboard the translator is using and the particular keystrokes he used. If you’re having trouble finding words in search results or when trying to perform regex operations, you can use this tool we made for you to visualize how each word and character is represented underlyingly.”
I’m familiar with Alt+X to see the codepoints for a character, and also with Tools>Checking Inventories>Characters Inventory, but neither of these show entire words with their codepoints- that’s what I remember seeing in the printscreen in the announcement.
Does anyone know where to find this announcement or the feature it describes? I’ve searched around and can’t find it.
My ultimate goal is to better understand how different ways of encoding characters and representing words affect language projects. Be it through looking through accompanying documentation or talking with experts on the topic.
Thank you!
机器翻译自 English