Language: English

tools / hex

hex/

Convert text to hexadecimal and back. Choose UTF-8, GBK, UTF-16 and more; the output can use a separator, uppercase and a 0x or \x prefix.

Text

hex

FAQ

How many bytes is one Chinese character in hex, and why do other tools give a different result?

It depends on the charset. In UTF-8 a common Chinese character takes 3 bytes, for example 中 is e4 b8 ad; in GBK and GB2312 it takes 2 bytes, and 中 is d6 d0; an emoji takes 4 bytes in UTF-8. English letters and digits are 1 byte in both UTF-8 and GBK, so they come out the same. When the result differs from another tool, check which charset that tool uses. This page defaults to UTF-8.

Hex to text says the bytes are not valid text, or the result is garbled. What should I do?

It means the bytes were not encoded with the current charset, so try another one. For Chinese text the usual suspects are UTF-8 and GBK (older software and database exports often use GBK); hex in UTF-16 usually has many 00 bytes. This tool never shows undecodable bytes as garbage. The one exception is ISO-8859-1: every byte maps to some character, so it never errors, but Chinese text turns into garbled characters. GBK and GB2312 are both decoded with the browser's built-in GB18030 decoder, which is more lenient than Python or Java, so GB2312 also decodes characters that only exist in GBK. If the bytes were never text in the first place (an image, encrypted data, a hash), no charset will decode them.

Which formats can hex to text read?

Upper and lower case both work. Bytes can be separated by spaces, line breaks, tabs, colons, commas or hyphens, or not separated at all, and each byte may start with 0x or \x, so 0x48,0x65 and \x48\x65 decode as they are. Each group between two separators must have an even number of digits, so write 0x01 rather than 0x1. Any character other than 0-9 and a-f also gives an error. Both errors point to the character position. Forms like U+4E2D and %E4%B8%AD are not recognized; use the URL tool on this site for URL encoding.

How do I get formats like 0x48,0x65 or \x48\x65?

Set the separator to "," and the prefix to "0x" to get 0x48,0x65, which you can drop into a C or Java byte array. Set the separator to "None" and the prefix to "\x" to get \x48\x65, the byte notation used in Python and C strings. The prefix is added before every byte, so "None" plus 0x gives 0x480x65, which is rarely useful. "Uppercase" only affects a-f; the x in 0x stays lowercase.

Why does the UTF-16 result start with feff, and why does it differ from other tools?

UTF-16 follows Java: it starts with FE FF (a byte order mark) and is big-endian, so 中 becomes fe ff 4e 2d. UTF-16BE without the mark is 4e 2d; copy the result and delete the first two bytes. UTF-32 is big-endian with no mark, so 中 is 00 00 4e 2d. When decoding, UTF-16 accepts a leading FE FF (big-endian) or FF FE (little-endian) and assumes big-endian when there is no mark.

Why does Chinese text give an error with ASCII or ISO-8859-1?

ASCII covers only 0 to 127, and ISO-8859-1 stops at U+00FF (Western European letters with accents, no Chinese). Anything outside that range gives an error such as "The character "中" cannot be encoded in ASCII" instead of being silently replaced by ? the way Java's getBytes does. GBK and GB2312 behave the same way: emoji and rare characters missing from the code table give an error. Switch to UTF-8 or GB18030 to convert them.

Are spaces and line breaks in my input converted?

In text to hex, yes. Every character in the input box is converted, including leading and trailing spaces and line breaks: a space is 20 and a line break is 0a. The browser's text box normalizes line breaks to \n, so a line break in multi-line text becomes 0a rather than 0d 0a; edit the result yourself if you need Windows line endings. In hex to text, whitespace in the hex is ignored and only the bytes count, and a 0d 0a in the result is also shown as a single line break.

Is what I enter uploaded?

No. The conversion runs in your browser, makes no network requests and saves nothing.