The html-lang-valid rule flags a lang attribute on the <html> element whose value is not a valid language tag – for example en_GB with an underscore, or the word english. Software cannot read it, so the page is treated as if it had no language at all. The fix is a well-formed tag such as lang="en-GB".
What the rule means
WCAG 3.1.1 Language of Page requires that software can determine the page's main language. A lang value only does that if it follows BCP 47, the standard for language tags: a primary language code from the IANA registry, optionally followed by subtags separated by hyphens – en, en-GB, de-AT, zh-Hant.
axe-core checks the primary subtag – the part before the first hyphen – against its list of known language codes. So:
en_GB,de_DE,english,eng: the first part is not a known code, and the rule fails.en-GB,EN-gb,de: pass. Case does not matter.- An empty
lang=""is not checked here; a missing or empty value is undefined.
What the rule does not check: whether the tag names the language the page is really written in (a German page with lang="en" passes), and whether the region part makes sense (en-EN passes, although there is no region "EN").
Who is affected
The same people as with a missing language: blind and partially sighted people using a screen reader, whose software picks its voice and pronunciation from the page language. An unreadable tag leaves the screen reader on its default voice, so the text may be read with the wrong pronunciation rules. Browser translation, hyphenation and spell checking in form fields also rely on a tag they can parse.
Why the check fails
- CMS locales copied straight into the template. Many systems store the locale as
de_DEoren_US; printed intolangunchanged, the underscore makes it invalid. - Language names instead of codes, such as
lang="english"orlang="deutsch". - Template placeholders that were never replaced, such as
lang="{{ locale }}"orlang="LANG". - Three-letter codes such as
lang="eng"orlang="ger". Where a language has a two-letter code, BCP 47 uses it:en,de. - Country codes used as language codes, such as
lang="dk"for Danish (correct:da) orlang="jp"for Japanese (correct:ja).lang="se"for Swedish is worse:seis Northern Sami, so the rule passes and the page is still labelled with the wrong language.
How to fix it
- Find where the template prints
langand where its value comes from – a CMS setting, a locale variable, a translation plugin. - Convert locale identifiers to BCP 47: replace the underscore with a hyphen (
de_DE→de-DE), or use only the language part (de). - Look up unusual codes in the IANA language subtag registry – a language code is not the same as a country code.
<!-- Before: a CMS locale with an underscore is not a language tag -->
<!doctype html>
<html lang="en_GB">
<head>
<meta charset="utf-8">
<title>Opening hours – Example Library</title>
</head>
<body>
<main><h1>Opening hours</h1></main>
</body>
</html>
<!-- After: a valid BCP 47 tag, with a hyphen -->
<!doctype html>
<html lang="en-GB">
<head>
<meta charset="utf-8">
<title>Opening hours – Example Library</title>
</head>
<body>
<main><h1>Opening hours</h1></main>
</body>
</html>
In PHP, str_replace('_', '-', $locale) does the conversion. Many frameworks already offer a helper that returns the tag in the right form – use it rather than the raw locale.
How to test it manually
- Open the page source or run
document.documentElement.langin the browser console. - Check the format: letters, hyphens, no underscores, no spaces, no full language names.
- Look up the first part in the IANA language subtag registry, or paste the tag into the free W3C Internationalization Checker, which reports malformed tags.
- Check that the tag matches the text on the page – the rule cannot tell
defromenon the wrong page. - On a multilingual site, repeat for each language version.
Related WCAG criterion
3.1.1 Language of Page, Level A. Related: 3.1.2 Language of Parts for passages in another language, where the same tag format applies.
How Reviseberg reports it
Reviseberg runs html-lang-valid on every crawled page at 1280px. Because the value usually comes from one template or one CMS setting, the finding tends to appear on every page of a language version at once. The issue list shows the rule with its severity, WCAG criterion 3.1.1 and level, the number of pages, and the points you get back by fixing it. The detail view shows the selector, the HTML snippet with the invalid value and every affected page – so a multilingual site shows at a glance which language version is affected. You can mark a finding as "ignore", "can't fix" or "false positive" with a reason; every decision is logged.
Whether the tag names the right language is not something a crawl can decide. Check that by hand, as described above.
Related rules
- undefined – the
<html>element has nolangat all. - undefined – the same check for
langon elements inside the page. - undefined – another page-level fact set in the template.
Check the language tags across your whole site – get a free scan