What is character encoding?
Character encoding is the rule a browser or crawler uses to convert the raw bytes of a web page into letters, numbers, and symbols. If the page is written in one encoding and read in another, text turns into garbage: café instead of café, or question marks and boxes in place of quotes and dashes. This garbled output is called mojibake.
UTF-8 is the standard encoding for the web. It covers every character in Unicode, from accented Latin letters to Chinese, Arabic, and emoji, and the HTML standard requires it for new documents. The overwhelming majority of websites use it.
How it works
A page can declare its encoding in two places. The server can send it in the HTTP Content-Type header, as text/html; charset=utf-8, and the HTML can declare it with <meta charset="utf-8">. If they disagree, the HTTP header wins.
The meta tag must appear within the first 1024 bytes of the document so the browser finds it before it starts interpreting text. Put it as the first element inside <head>, before the <title>, because a title with non-ASCII characters could otherwise be read in the wrong encoding.
Declaring UTF-8 is only half of it. The file itself, the database, and any text inserted into the page must actually be UTF-8. A correct declaration on top of bytes stored as Windows-1252 still produces broken characters.
Example
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="utf-8">
<title>Café Menu: Espresso, Pour-Over, and Cold Brew</title>
</head>
And the matching response header:
Content-Type: text/html; charset=utf-8
The older form, <meta http-equiv="Content-Type" content="text/html; charset=utf-8">, still works but the short form is standard.
In WordPress
WordPress stores the encoding in the blog_charset option, which defaults to UTF-8, and themes print it with <meta charset="<?php bloginfo( 'charset' ); ?>">. The database should use utf8mb4, which WordPress has used by default since version 4.2 and which is needed for emoji. Encoding problems on WordPress usually come from old databases migrated with the wrong character set, or from content pasted from other systems. The hydrogenseo.com bulk audit flags pages with a missing or non-UTF-8 declaration.
Common mistakes
- No charset declaration, leaving browsers to guess.
- Declaring the charset after the title or deep in the head.
- Header and meta tag disagreeing, for example a server sending ISO-8859-1.
- Migrating a database without converting it, producing mojibake across old posts.
Common questions
Does character encoding affect SEO?
Indirectly. Search engines can usually detect encoding, but garbled text in titles, snippets, or content looks broken to users and can be indexed incorrectly.
Should I use UTF-8 or something else?
Use UTF-8. It supports every language, it is required by the HTML standard for new documents, and there is no practical reason to choose another encoding for a new site.