+1 票
568 次瀏覽

If you use the Latin alphabet with diacritics over characters, you may encounter a puzzling phenomenon in the Wordlist. You may have multiple entries of what appears to be the same word.
This happens because the Unicode standard has two ways of entering many letters with diacritics. For example, the é character can be the Unicode character 00E9, or it can be two characters, 0065 followed by 0301 (lower case e followed by combining diacritic acute). If this is the case, words containing 00E9 will not be identical with words containing 0065 + 0301, even if all the other letters are the same.

You can verify this by looking at the character inventory. This example is of a test project with two t characters, one *00E9 and one 0065 0301. You can see that it breaks down the difference between the é characters.

image

The Unicode standard does say that these sequences should be treated as identical by the software, but many programs have not managed to do that yet. Paratext does not, Microsoft Word does not either.

If you have this issue, you can fix it by doing a search and replace, for instance searching for one kind of é and replacing it with the other kind. To keep it from occurring again, translators working on the same project should either use the same keyboard file, or at least keyboard files that output the same character sequence for letters with diacritics.

舊文章 - 以原文顯示
Paratext 由 [Expert]
(3.3k 點) 提出

已重新顯示 | 568 次瀏覽

4 個回答

0 票
最佳回答

In Paratext 7.6 (still in development), new projects will be default use
"Normal Form Composed" when saving text data - this will prevent these
differences in text that are usually caused by using different keyboards.
The “Normal Form Decomposed” can be used instead - this may be required to
make text display correctly in some fonts.

There is also a Convert Project tool in Paratext 7.6 that can be used to
change all text in a project to either NFC or NFD. The Convert Project tool
does create a new project in the process, but project history is preserved.

舊文章 - 以原文顯示
由 [Administrator]
(3.4k 點) 提出
0 票

The problem we’re into is that accent marks are being classified on their own or being grouped with the character that follows rather than with the character they modify.

image

What should we be doing differently?

舊文章 - 以原文顯示
由 (448 點) 提出

Dear drwww,

Go to Project/Language Settings and choose the Other Characters Tab. There
you will see a little tick box which says, “Non-standard diacritics follow
base character”. Tick this, and your problems should go away.

You should also make sure that the alphabetic characters tab is filled in
properly according to the “More Help” information provided on the side
slideout. You may have problems with Paratext accepting the capital form of
some of your combining diacritic characters. If you do, just put in the
lower case form(s) of the offending character(s) with a space between them.
The capitalization will be properly predicted.

Also, you should make sure to view the Characters Inventory with “Show
combinations” ticked. It won’t look right unless the tick box in the
language settings is chosen first, but it is important for making sure that
you are approving your characters properly, joined, rather than separate
from their diacritics.

Blessings,

Shegnada J.

Language Technology & Publishing Coordinator, Nigeria

GPS Text Processing Specialist

Wycliffe

+[Phone Removed]

Skype:Shegnada.James.

舊文章 - 以原文顯示
0 票

This is actually fixed in 7.6. Even if the project is not fully normalized in the way @John+Wickberg suggests, the Wordlist will still only show one entry for these two normalized cases.

舊文章 - 以原文顯示
由 [Expert]
(16.7k 點) 提出
0 票

For additional experiences and other discussions, see (in chronological order):

  • What is Paratext's behaviour for composing/decomposing unicode characters/sequences?
  • Matches Not Found in Word List - Composed and Decomposed (whole conversation)
  • Guide: Checking > Characters Inventory (second post)

Adding keywords for search:
normalisation normalization

舊文章 - 以原文顯示
由 (1.4k 點) 提出
已重新顯示

相關問題

0 票
0 個回答 68 次瀏覽
我認為空白與隱藏字元(WHC)選項將對我們的專案產生重大影響 它帶來了一些實際的好處,但如果用於格式設定目的,也有被濫用的風險 我們需要仔細思考,以確保能妥善使用此功能 [已編輯] 請參閱 John Wickberg 在相關貼文中提供的 ... 時,自動將它們轉換為 u00A0 可能也是個好主意 我期待聽到其他人的評論和觀察,也許還能收到開發人員的回覆
Kent Spielmann 1.8k 提出 已提問 11月 19, 2024
0 票
1 個回答 274 次瀏覽
Following migration of our project to PT8, the work I had previously done in an interlinear back-translation into NT Greek appears ... . What do I need to do to fix these problems?
anon084127 105 提出 已提問 10月 27, 2017
0 票
5 個回答 915 次瀏覽
你好,我使用的是最新版本的 Paratext WIN10 和 PTXPrint Satere (SAT) 語言使用鼻化元音 我安裝了一個特殊字體來轉換為鼻化元音 直撇號 (') 用於表示喉塞音 當我使用 PTXPrint 列印《創世記》或《出埃及 ... 和標點符號清單,發現有一些問題,但我嘗試使用 PTXPrint 列印時出現的無法識別字符的問題仍然存在?
anon165192 116 提出 已提問 2月 21, 2023
0 票
2 個回答 572 次瀏覽
我的文字中有像 3 000 000 這樣的數字。 字元清單檢查顯示一個錯誤:「initial zero non valid.」(開頭零無效)。但開頭的零在清單中或其他地方都沒有顯示出來,所以我無法將其標記為有效, 我該如何處理這個問題?
anon057624 136 提出 已提問 9月 6, 2021
0 票
6 個回答 657 次瀏覽
我試圖將專案(Manya - mzj)中的分解字元(decomposed characters)變更為組合格式(composed characters) 但當我嘗試在「專案屬性」(Project Properties)部分套用此選項時,系統不讓我點選 ... 1162 24.9 KB Shot12292 1062 171 KB 有什麼想法可以解決這些問題嗎?
anon055151 176 提出 已提問 3月 22, 2021
Welcome to Support Bible, where you can ask questions and receive answers from other members of the community.
If anyone destroys God’s temple, God will destroy that person; for God’s temple is sacred, and you together are that temple.
1 Corinthians 3:17
3,060 個問題
6,021 個回答
5,687 則評論
2,035 位使用者