+1 票
571 次浏览

If you use the Latin alphabet with diacritics over characters, you may encounter a puzzling phenomenon in the Wordlist. You may have multiple entries of what appears to be the same word.
This happens because the Unicode standard has two ways of entering many letters with diacritics. For example, the é character can be the Unicode character 00E9, or it can be two characters, 0065 followed by 0301 (lower case e followed by combining diacritic acute). If this is the case, words containing 00E9 will not be identical with words containing 0065 + 0301, even if all the other letters are the same.

You can verify this by looking at the character inventory. This example is of a test project with two t characters, one *00E9 and one 0065 0301. You can see that it breaks down the difference between the é characters.

image

The Unicode standard does say that these sequences should be treated as identical by the software, but many programs have not managed to do that yet. Paratext does not, Microsoft Word does not either.

If you have this issue, you can fix it by doing a search and replace, for instance searching for one kind of é and replacing it with the other kind. To keep it from occurring again, translators working on the same project should either use the same keyboard file, or at least keyboard files that output the same character sequence for letters with diacritics.

较早的帖子 - 以原始语言显示
Paratext 中 由 [Expert]
(3.3k 分) 发布

已重新显示 | 571 次浏览

4 个回答

0 票
最佳答案

In Paratext 7.6 (still in development), new projects will be default use
"Normal Form Composed" when saving text data - this will prevent these
differences in text that are usually caused by using different keyboards.
The “Normal Form Decomposed” can be used instead - this may be required to
make text display correctly in some fonts.

There is also a Convert Project tool in Paratext 7.6 that can be used to
change all text in a project to either NFC or NFD. The Convert Project tool
does create a new project in the process, but project history is preserved.

较早的帖子 - 以原始语言显示
由 [Administrator]
(3.4k 分) 发布
0 票

The problem we’re into is that accent marks are being classified on their own or being grouped with the character that follows rather than with the character they modify.

image

What should we be doing differently?

较早的帖子 - 以原始语言显示
由 (448 分) 发布

Dear drwww,

Go to Project/Language Settings and choose the Other Characters Tab. There
you will see a little tick box which says, “Non-standard diacritics follow
base character”. Tick this, and your problems should go away.

You should also make sure that the alphabetic characters tab is filled in
properly according to the “More Help” information provided on the side
slideout. You may have problems with Paratext accepting the capital form of
some of your combining diacritic characters. If you do, just put in the
lower case form(s) of the offending character(s) with a space between them.
The capitalization will be properly predicted.

Also, you should make sure to view the Characters Inventory with “Show
combinations” ticked. It won’t look right unless the tick box in the
language settings is chosen first, but it is important for making sure that
you are approving your characters properly, joined, rather than separate
from their diacritics.

Blessings,

Shegnada J.

Language Technology & Publishing Coordinator, Nigeria

GPS Text Processing Specialist

Wycliffe

+[Phone Removed]

Skype:Shegnada.James.

较早的帖子 - 以原始语言显示
0 票

This is actually fixed in 7.6. Even if the project is not fully normalized in the way @John+Wickberg suggests, the Wordlist will still only show one entry for these two normalized cases.

较早的帖子 - 以原始语言显示
由 [Expert]
(16.7k 分) 发布
0 票

For additional experiences and other discussions, see (in chronological order):

  • What is Paratext's behaviour for composing/decomposing unicode characters/sequences?
  • Matches Not Found in Word List - Composed and Decomposed (whole conversation)
  • Guide: Checking > Characters Inventory (second post)

Adding keywords for search:
normalisation normalization

较早的帖子 - 以原始语言显示
由 (1.4k 分) 发布
已重新显示

相关问题

0 票
0 个回答 68 次浏览
在我看来, 空白和隐藏字符 (WHC)选项将对我们的项目产生重大影响 它带来了一些切实的好处,但如果用于格式设置目的,也有被滥用的风险 我们需要仔细思考,以确保正确使用它 [已编辑] 参见 John Wickberg 在相关帖子中的回答 ... 时,自动将它们转换为 u00A0 可能是一个好主意 我期待听到其他人的评论和观察,也许还能得到开发人员的回复
Kent Spielmann 1.8k 发布 提问于 十一月 19, 2024
0 票
1 个回答 274 次浏览
Following migration of our project to PT8, the work I had previously done in an interlinear back-translation into NT Greek appears ... . What do I need to do to fix these problems?
anon084127 105 发布 提问于 十月 27, 2017
0 票
5 个回答 916 次浏览
你好,我正在使用最新版本的 Paratext WIN10 和 PTXPrint Satere (SAT) 语言使用鼻化元音 我安装了一种特殊字体来转换为鼻化元音 直撇号 (') 用于表示喉塞音 当我使用 PTXPrint 打印《创世记》或《出埃及 ... 了字符和标点符号清单,发现了一些问题,但我尝试使用 PTXPrint 打印时无法识别字符的问题仍然存在?
anon165192 116 发布 提问于 二月 21, 2023
0 票
2 个回答 572 次浏览
我的文本中有类似 3 000 000 的数字。 字符清单检查显示一个错误:“initial zero non valid.”(初始零无效)。但初始零并没有显示在清单中或其他任何地方,因此我无法将其标记为有效, 我该如何解决这个问题?
anon057624 136 发布 提问于 九月 6, 2021
0 票
6 个回答 657 次浏览
我试图将我的项目(Manya - mzj)中的分解字符更改为组合字符 但是当我尝试在 项目属性 (Project Properties)部分应用此选项时,系统不允许我选择 确定 (Okay) Shot51194 1246 32.6 ... 1162 24.9 KB Shot12292 1062 171 KB 关于如何解决这些问题,大家有什么想法吗?
anon055151 176 发布 提问于 三月 22, 2021
Welcome to Support Bible, where you can ask questions and receive answers from other members of the community.
Make every effort to keep the unity of the Spirit through the bond of peace.
Ephesians 4:3
3,060 个问题
6,021 个回答
5,689 条评论
2,035 位用户