0
305 次浏览

Reading about Unicode equivalence here and on SIL’s NRSI pages here and also here on Wikipedia it would appear that Unicode software can choose to work as NFC or NFD but is supposed to treat all sequences as equivalent and can transform whatever underlying form there is at will. Given the behaviour of the character inventory in listing the character codes (and sequences of characters if you tick the Combinations box), I would like to know how Paratext deals with Unicode normalisation internally, in Find/Replace, in checks/inventories, the word list and regarding output.

Some projects I have seen end up with a mixture of combined forms (single code point for a complex character, such as ‘a’ with a tilde above it) and the separate forms (‘a’ + combining tilde). Visually identical words then are treated as distinct words in the Wordlist (7.5 at least).

Does Paratext try to avoid normalising its text? If so, for what reasons? It would just be nice to know what the software is trying to do, so that any input in terms of keyboarding systems, autocorrect etc and output (apps etc) can be made most helpfully.

较早的帖子 - 以原始语言显示
Paratext (510 分) 发布 | 305 次浏览

3 个回答

+1
最佳答案

Yes, in 7.5 and before, Paratext did not work well with non-normalized text. We tried to fix much of these types of problems in 7.6 so that the Wordlist doesn’t show multiple words that are the same, Find/Replace finds stuff correctly, Biblical Terms don’t fail to find renderings, etc.
In addition, 7.6 also includes a new option for new projects to select the normalization mode for the entire project so that the files are always saved in a certain normalized format. We don’t allow projects to change normalization mid-way because it will likely create massive merge conflicts during S/R (If I recall correctly, you can use Convert Project to change the normalization of a project mid-way if it’s important).

So, short answer is that Paratext 7.6+ accepts the input as-is from the keyboard and keeps it that way for any existing projects. The tools have been updated to handle non-normalized text much better. Any new projects can also select a normalization to have all text be normalized the same to disk.

EDIT: Aww, @John+Wickberg beat me. :disappointed:

较早的帖子 - 以原始语言显示
[Expert]
(16.7k 分) 发布

已重新显示
+1

Prior to Paratext 7.6, Paratext just saved whatever the user typed - there
was no normalization of the data.

In Paratext 7.6, we added a value for normalization on the Advanced tab of
the Project Properties and Settings dialog. The default value for new
projects is NFC, but that can be changed to NFD or None. The value for
existing projects can’t be changed, so any project created before Paratext
7.6 will be left as None.

We at first were just going to go with NFC, but found that NFD was required
for some projects to display correctly. The None option was later added
since even NFD does some ordering that was causing problems for some
projects.

If NFC or NFD is used, any typed input will be normalized before it is used.

The Tools > Advanced > Convert Project command can be used to create a copy
of a project with a new normalization setting.

John+Wickberg

较早的帖子 - 以原始语言显示
[Administrator]
(3.4k 分) 发布

已重新显示
0

Adding keywords for search:
composed decomposed

较早的帖子 - 以原始语言显示
(1.4k 分) 发布
已重新显示

相关问题

0
1 个回答 237 次浏览
你好, 我想了解一下 Paratext 在保存 自动保存以及发送/接收(Send/Receive)方面的具体行为 我尝试在这个网站和 Paratext 帮助文档中查找答案,但没有找到相关信息 如果我编辑了某个经文,在未保存的情况下点击发送/接 ... 根据一些快速测试,看起来 PT 在发送/接收(SR)和关闭之前会自动保存,但我希望确认一下 非常感谢!
alex_larkin 379 发布 提问于 四月 3, 2021
0
0 个回答 189 次浏览
This week I discovered that Alt-X does work for Unicode characters where the code is longer than 4 hex digits. The ... desired character when typing two periods: ..-->\ud804\udf
[Expert]
sewhite
3.3k 发布
提问于 八月 16, 2019
+1
0 个回答 160 次浏览
Some Unicode characters have a code longer than 4 hex digits. For instance David Rowe was trying to help someone ... on that character gives the UTF16 coding equivalent for it.
[Expert]
sewhite
3.3k 发布
提问于 三月 18, 2019
0
2 个回答 389 次浏览
We are working with a language that has long vowels which we can't write with Odia letters. There are some books printed ... type it neither we know the Unicode of it. please help
anon480013 162 发布 提问于 十月 1, 2018
0
4 个回答 573 次浏览
我使用 比较版本 (Compare versions)工具来查看项目文本中的更改 它很有用 然而,我注意到这种奇怪的滚动行为 如果我在 比较版本 工具中点击一个单词或选择文本,窗口就会滚动 这是出乎意料的行为,使得该工具更难使用 我花了数秒钟( ... 需要滚动才能显示的新文本来发生的 我该如何阻止在 比较版本 工具中仅通过点击单词就发生滚动? 谢谢帮助!
bit 495 发布 提问于 十二月 27, 2021
Welcome to Support Bible, where you can ask questions and receive answers from other members of the community.
And let us consider how we may spur one another on toward love and good deeds, not giving up meeting together, as some are in the habit of doing, but encouraging one another—and all the more as you see the Day approaching.
Hebrews 10:24-25
3,046 个问题
6,006 个回答
5,671 条评论
2,026 位用户