Latin diacritics PDP

Tapani Tarvainen ncsg at TAPANI.TARVAINEN.INFO
Wed Mar 26 09:54:46 EET 2025


Dear all,

The Latin Diacritics PDP Working Group just concluded its first
meeting (the kick-off meeting in Seattle apparently doesn't count).

The discussion was mainly about scope, which has some perhaps
surprising limitations:

(1) The "diacritic" is used with Unicode definition, which differs
from how linguistiscs use it, but more importantly it excludes
letters used in various Latin-based alphabets that technically
aren't formed by adding a diacritic mark to an ASCII letter.

Notably, that excludes ø, even though it looks like it combines o
and /, it isn't technicall done like that in Unicode, and ligatures
like æ and œ. As æ and ø are Norwegian/Danish equivalent to ä and ö
(also used in other languages, e.g., Icelandic uses æ and ö).
Likewise ŋ is excluded even though it looks like n with a tail.

I suspect Danes and Norwegians &c won't be happy about this.

Still it doesn't need to be a big problem, if the WG can come up
with good, generalizable rules (if we have a rule on how to deal
with ä and ö, same could be used with æ and ø), but we may need
another PDP (or several...) to deal with it.

It was suggested, however, that a (partial) list of such
potentially problematic letters be compiled for reference.
If anybody is interested in suggesting such, feel free to
contact me.

(2) Only two strings will be considered. This sounds odd,
given that there are even proper names that differ only
in diacritics, e.g., Sjoberg, Sjöberg, Sjóberg (also Sjøberg,
but that's excluded as per (1) above).

(3) The only case considered is where a single applicant
is applying for two strings. It won't be possible to
give the ASCII version and the diacritic version to
different applicants (or at least this WG cannot
consider that possibility).

The WG has asked for early input from various parties,
and the issue should be on GNSO council's agenda in April.

I'm not sure to which extent any of those issues could
still be changed. Probably easiest would be (2), it
wasn't quite as explictly hard-coded as the rest.

-- 
Tapani Tarvainen



More information about the Ncsg-discuss mailing list