Alguns anos atrás, lembro de ter visto um boletim do Paratext com um anúncio sobre um novo recurso que ajuda os usuários a visualizar os pontos de código Unicode no texto. Dizia algo como: “Sabemos que muitas vezes há mais de uma combinação de pontos de código Unicode que pode ser usada para representar um determinado caractere. Frequentemente, a forma como o caractere ou a palavra é composta depende muito do teclado que o tradutor está usando e das teclas específicas que ele pressionou. Se você está tendo dificuldade em encontrar palavras nos resultados de pesquisa ou ao tentar executar operações de expressão regular, pode usar esta ferramenta que criamos para você visualizar como cada palavra e caractere é representado subjacentemente.”
Estou familiarizado com Alt+X para ver os pontos de código de um caractere, e também com Ferramentas>Inventários de Verificação>Inventário de Caracteres, mas nenhum desses mostra palavras inteiras com seus pontos de código - é isso que lembro de ter visto no printscreen do anúncio.
Alguém sabe onde encontrar esse anúncio ou o recurso que ele descreve? Procurei por aí e não consigo encontrar.
Meu objetivo final é compreender melhor como diferentes formas de codificação de caracteres e representação de palavras afetam projetos de linguagem. Seja através da consulta à documentação acompanhante ou conversando com especialistas no assunto.
Obrigado!
Years ago I remember seeing a Paratext newsletter with an announcement about a new feature that helps users visualize the Unicode codepoints in the text. It said something like “We know that there is often more than one combination of Unicode codepoints that can be used to represent a given character. Often, the way the character or word is composed depends a lot on the keyboard the translator is using and the particular keystrokes he used. If you’re having trouble finding words in search results or when trying to perform regex operations, you can use this tool we made for you to visualize how each word and character is represented underlyingly.”
I’m familiar with Alt+X to see the codepoints for a character, and also with Tools>Checking Inventories>Characters Inventory, but neither of these show entire words with their codepoints- that’s what I remember seeing in the printscreen in the announcement.
Does anyone know where to find this announcement or the feature it describes? I’ve searched around and can’t find it.
My ultimate goal is to better understand how different ways of encoding characters and representing words affect language projects. Be it through looking through accompanying documentation or talking with experts on the topic.
Thank you!
Tradução automática de English