0 suara
303 tampilan

Reading about Unicode equivalence here and on SIL’s NRSI pages here and also here on Wikipedia it would appear that Unicode software can choose to work as NFC or NFD but is supposed to treat all sequences as equivalent and can transform whatever underlying form there is at will. Given the behaviour of the character inventory in listing the character codes (and sequences of characters if you tick the Combinations box), I would like to know how Paratext deals with Unicode normalisation internally, in Find/Replace, in checks/inventories, the word list and regarding output.

Some projects I have seen end up with a mixture of combined forms (single code point for a complex character, such as ‘a’ with a tilde above it) and the separate forms (‘a’ + combining tilde). Visually identical words then are treated as distinct words in the Wordlist (7.5 at least).

Does Paratext try to avoid normalising its text? If so, for what reasons? It would just be nice to know what the software is trying to do, so that any input in terms of keyboarding systems, autocorrect etc and output (apps etc) can be made most helpfully.

Postingan lama - ditampilkan dalam bahasa aslinya
Paratext oleh (510 poin) | 303 tampilan

3 Jawaban

+1 suara
Jawaban terbaik

Yes, in 7.5 and before, Paratext did not work well with non-normalized text. We tried to fix much of these types of problems in 7.6 so that the Wordlist doesn’t show multiple words that are the same, Find/Replace finds stuff correctly, Biblical Terms don’t fail to find renderings, etc.
In addition, 7.6 also includes a new option for new projects to select the normalization mode for the entire project so that the files are always saved in a certain normalized format. We don’t allow projects to change normalization mid-way because it will likely create massive merge conflicts during S/R (If I recall correctly, you can use Convert Project to change the normalization of a project mid-way if it’s important).

So, short answer is that Paratext 7.6+ accepts the input as-is from the keyboard and keeps it that way for any existing projects. The tools have been updated to handle non-normalized text much better. Any new projects can also select a normalization to have all text be normalized the same to disk.

EDIT: Aww, @John+Wickberg beat me. :disappointed:

Postingan lama - ditampilkan dalam bahasa aslinya
oleh [Expert]
(16,7k poin)

ditampilkan kembali
+1 suara

Prior to Paratext 7.6, Paratext just saved whatever the user typed - there
was no normalization of the data.

In Paratext 7.6, we added a value for normalization on the Advanced tab of
the Project Properties and Settings dialog. The default value for new
projects is NFC, but that can be changed to NFD or None. The value for
existing projects can’t be changed, so any project created before Paratext
7.6 will be left as None.

We at first were just going to go with NFC, but found that NFD was required
for some projects to display correctly. The None option was later added
since even NFD does some ordering that was causing problems for some
projects.

If NFC or NFD is used, any typed input will be normalized before it is used.

The Tools > Advanced > Convert Project command can be used to create a copy
of a project with a new normalization setting.

John+Wickberg

Postingan lama - ditampilkan dalam bahasa aslinya
oleh [Administrator]
(3,4k poin)

ditampilkan kembali
0 suara

Adding keywords for search:
composed decomposed

Postingan lama - ditampilkan dalam bahasa aslinya
oleh (1,4k poin)
ditampilkan kembali

Pertanyaan terkait

0 suara
1 jawaban 237 tampilan
Halo, Saya ingin bertanya tentang beberapa detail perilaku Paratext terkait penyimpanan (saving), penyimpanan otomatis ( ... ingin bertanya untuk memastikan. Terima kasih banyak!
alex_larkin 379 bertanya Apr 3, 2021
0 suara
0 jawaban 189 tampilan
This week I discovered that Alt-X does work for Unicode characters where the code is longer than 4 hex digits. The ... desired character when typing two periods: ..-->\ud804\udf
[Expert]
sewhite
3,3k
bertanya Agu 16, 2019
+1 suara
0 jawaban 160 tampilan
Some Unicode characters have a code longer than 4 hex digits. For instance David Rowe was trying to help someone ... on that character gives the UTF16 coding equivalent for it.
[Expert]
sewhite
3,3k
bertanya Mar 18, 2019
0 suara
2 jawaban 389 tampilan
We are working with a language that has long vowels which we can't write with Odia letters. There are some books printed ... type it neither we know the Unicode of it. please help
anon480013 162 bertanya Okt 1, 2018
0 suara
4 jawaban 573 tampilan
Saya menggunakan alat Compare versions untuk melihat perubahan dalam teks proyek. Alat ini berguna. Namun, saya ... di alat Compare Versions ? Terima kasih atas bantuannya!
bit 495 bertanya Des 27, 2021
Welcome to Support Bible, where you can ask questions and receive answers from other members of the community.
I urge you, brothers and sisters, to watch out for those who cause divisions and put obstacles in your way that are contrary to the teaching you have learned. Keep away from them.
Romans 16:17
3,046 pertanyaan
6,006 jawaban
5,671 komentar
2,026 pengguna