Unicode’s weirdest characters
Most people think of Unicode as the system that makes emojis, accented letters, and different alphabets work across devices.
But Unicode also includes characters you cannot see, characters that control how text behaves, unusual spaces, and symbols that can change what appears on screen without adding anything obvious to the sentence.
Some are essential for typography and different writing systems. Others are responsible for the occasional moment when copied text suddenly behaves in a way that makes no sense. Here are some of the strangest.
The space you cannot see
The Zero Width Space does exactly what its name suggests. It has no visible width.
Unlike a normal space, you cannot look at a sentence and immediately tell that it is there. Its purpose is to provide a possible place for text to break without adding visible spacing between the surrounding characters. Unicode defines U+200B as a character with no visible glyph or width.
That makes it useful in certain writing systems and text processing situations. It can also make copied text surprisingly difficult to troubleshoot.
Two strings can look identical on screen while one contains an invisible character that the other does not.
The invisible character behind some emojis
The Zero Width Joiner, U+200D, is another character you normally never see.
Its job is to influence how neighbouring characters are displayed together. It is also used in many emoji sequences where several characters combine into what appears to be one emoji. This means one visible emoji can actually be made up of several Unicode characters connected by invisible joiners.
So what looks like one character on screen is not always one Unicode character underneath.
The space that refuses to separate
A No Break Space looks much like an ordinary space.
The difference appears when text reaches the end of a line. A normal space can provide a place for the text to wrap. A No Break Space keeps the surrounding text together instead.
This is useful for things such as names, numbers, units, and other text that should not be split across separate lines.
Unicode also includes a Narrow No Break Space, which performs a similar job while taking up less visual room. Apparently one type of space was never going to be enough.
The hyphen that waits for the right moment
The Soft Hyphen is one of Unicode’s more unusual formatting characters. Most of the time, it is invisible.
It marks a possible place where a word may break if the text needs to continue on another line. In Latin text, a visible hyphen may appear when that break is actually used.
In simple terms, it is a hyphen that can sit quietly inside a word until the layout needs it. Useful for typography.
Slightly confusing when copied text starts behaving differently between systems.
The character that can turn text around
Unicode also contains characters that control text direction. One of the more unusual examples is Right to Left Override, U+202E.
This directional formatting character changes how following text is displayed within its scope.
Characters like this exist because Unicode needs to support writing systems that run in different directions, including languages written from right to left.
They are legitimate and necessary. But because the character itself is invisible while the surrounding text changes direction, it can produce some very strange looking results if it appears unexpectedly.
The invisible switch behind emoji style
Some Unicode characters can appear either as regular text symbols or in emoji style. That is where Variation Selector 16, U+FE0F, comes in.
It is an invisible character used with supported base characters to request emoji presentation rather than text presentation.
Its partner, Variation Selector 15, can request text presentation instead. This means two sequences can contain the same visible base character but display differently because one includes an invisible variation selector.
It is another good example of why counting what you see on screen does not always tell you how many characters are actually present.
Unicode even has a visible space
The Ogham Space Mark, U+1680, is an unusual member of the Unicode space family.
In traditional Ogham typography, it is generally rendered with a visible horizontal line rather than appearing completely blank. Some fonts may render it differently. It comes from the Ogham writing system and shows just how broad Unicode really is.
Unicode is not only designed for modern Latin text, emojis, and common symbols. It also represents historical scripts and writing conventions from around the world.
Why strange Unicode characters can affect SMS
For SMS, unusual Unicode characters are more than interesting trivia.
Standard SMS commonly uses the GSM 7 character set. If a message contains a character that GSM 7 does not support, the message may switch to Unicode encoding instead.
LINK Mobility uses these limits in its SMS guidance and SMS length calculator.
This means an unusual or invisible character copied into a marketing message can potentially change the encoding without the person writing it noticing immediately.
Emojis and smart punctuation are more common examples, but invisible Unicode characters show why checking the final message can be useful before sending SMS at scale. Sometimes the weirdest character in the message is the invisible one.
Did you find the article and topic interesting?
If you would like to explore the subject further, discuss ideas, or understand how it could apply to your business, we are here to continue the conversation.