Website and documentation: casoon.github.io/html-conform
A Rust library for HTML5 specification conformance checking — validated against the Nu Html Checker (vnu)'s differential test corpus (0 false positives, 99.98 % accuracy across 4,655 fixtures), without a JVM, subprocesses, or HTTP network requests. Embeddable directly into any Rust application, CLI, or web service.
html-conform is validated continuously against the official W3C/vnu differential test suite (4,655 test fixtures vendored from validator/validator@388cb36).
- Precision Floor: 0 False Positives (
BASELINE_FALSE_POSITIVE = 0) — zero false alarms across all 4,655 test cases. - Accuracy: 99.98 % overall accuracy across the entire corpus (3,745 True Positives, 909 True Negatives).
- Residual False Negatives: 1 / 4,655, deliberate rather than unimplemented: flagging a mistyped property name inside
<style>needs a real CSS parser plus the full CSS property registry (vnu delegates this to a vendored W3C CSS Validator; a single corpus fixture is too little evidence for that surface area). It is a documented, deliberate limitation, not a silent gap. (The 2D table cell grid, formerly the largest remaining gap, landed assrc/table_integrity.rs; tree-construction-error tracking landed viahtml5-parser0.3.0 and was extended in 0.4.0 to every token the parser ignores and to misnested formatting elements; it still covers a subset of the spec's tree-construction errors, see What's not covered.) - One measured deviation from the corpus, by design: the missing-
langwarning. vnu's own test runner setsnu.validator.checker.ignoreMissingLang=trueglobally and flips it tofalseonly for the single fixture whose filename containsmissing-lang, so 752 expected-clean fixtures have nolangattribute merely because the check was switched off when their expectations were recorded. Real, production vnu warns on all of them, and so doeshtml-conform— the differential test corrects for the harness artifact in its comparison (tests/differential.rs'shas_findings_for_comparison) instead of the checker shipping less than vnu.
html-conform combines six independent finding sources into a single, unified CheckReport:
- HTML5 Tree Construction (
parser.html5) — Spec-compliant, error-tolerant tree parsing viahtml5-parser. Emits tokenizer, DOCTYPE, and tree-construction parse findings with line, column, and byte offset locations. Tokenizer errors are complete; tree-construction errors are a subset (see What's not covered). - Grammar & Content Model (
schema.html5) — Validation against the full vendored W3C RELAX NG schema (relax-ng), including SVG 1.1 and MathML 3 subtrees. - Custom Datatype Micro-Syntaxes (
w:*) — Full spec-compliant datatype validation for 50 custom W3C attribute microsyntaxes (w:image-candidate-stringsforsrcset,w:content-security-policy,w:media-query,w:datetime,w:iri-ref, BCP 47 language tags, etc.). - Schematron Co-Constraints (
rules/*.sch) — High-precision assertion rules viaschematron-engineandxpath-eval(ARIA 1.2 constraints, structural HTML restrictions, heading hierarchy, link/script attribute combinations). - Script & CSP Validation (
scripts.import-map,scripts.speculation-rules,csp.meta-enforcement) — Dedicated JSON validation for<script type="importmap">/<script type="speculationrules">contents, and Content Security Policy (<meta http-equiv="Content-Security-Policy">) enforcement against inline scripts/styles viacsp-parse. - Table Cell Grid (
tables.integrity) — Lays every table out over itscolspan/rowspanvalues to detect overlapping cells, cells spanning past the end of their row group, and columns that no cell ever begins in — a stateful 2D grid walk that the declarative XPath 1.0 rule layer cannot express.
The corpus numbers above only measure what the vendored fixtures exercise. Known gaps that real documents hit:
- Some tree-construction parse errors.
html5-parser0.4.0 records a parse error wherever it ignores a token (a stray</div>,</li>or</h2>, a second<body>,<td>outside a table, …) and when the adoption agency algorithm repairs misnested formatting elements (<b><i>…</b></i>). It still records nothing for the spec's parse errors on tokens it keeps: an end tag that closes an element while other elements inside it are still open (<div><span></div>), content after</html>, and the obsolete frameset modes. vnu reports these; the repaired tree carries no trace of them, sohtml-conformcannot detect them afterwards. - CSS inside
<style>— see the one residual false negative above.
The comparison with vnu lists the full set of differences.
Add html-conform to your Cargo.toml:
[dependencies]
html-conform = "0.2.0"use html_conform::check;
fn main() {
let html = r#"<!DOCTYPE html>
<html lang="en">
<head><title>Test Document</title></head>
<body><p>Hello world</p></body>
</html>"#;
let report = check(html).expect("checker execution succeeded");
for finding in &report.findings {
println!(
"[{:?}] {} at line {:?}: {}",
finding.severity, finding.rule_id, finding.location, finding.message
);
}
if report.has_errors() {
eprintln!("Document has conformance errors!");
}
}use html_conform::{CheckOptions, check_with_options};
let options = CheckOptions {
include_parse_errors: false, // exclude tokenizer/parser diagnostics
};
let report = check_with_options(html, options).unwrap(); HTML Source String / Document
│
▼
┌─────────────────────┐
│ 1. html5-parser │ WHATWG Tree Construction
└─────────────────────┘
│
┌──────────────────┼──────────────────┐
▼ ▼ ▼
┌──────────────────┐ ┌───────────────┐ ┌──────────────────┐
│ 2. relax-ng │ │ 3. Schematron │ │ 4. JSON / CSP │
│ (Schema & │ │ (Co-Con- │ │ (Import-Maps, │
│ Datatypes) │ │ straints) │ │ Speculation, │
└──────────────────┘ └───────────────┘ │ CSP) │
│ │ └──────────────────┘
└──────────────────┼──────────────────┘
│
▼
┌───────────────────────┐
│ CheckReport │
│ Vec<Finding> │
└───────────────────────┘
html-conform enforces quality through structured maintenance loops:
- Loop A (Schema Sync): Mechanical updates when W3C RELAX NG schemas or vnu upstream specifications update (
xtask/vendor-corpus.sh). - Loop B (Assertion Refinement Loop): Iterative refinement of Schematron rules and datatype checkers against the 4,655-fixture differential test suite, strictly maintaining the 0 False Positive floor.
- Loop C (Real-World Sanity Check): Supplementary, manual/non-CI signal that fetches real websites and diffs
html-conform's findings against a locally-run Nu Html Checker (vnu) jar (xtask/fetch-real-world.sh+xtask/compare-real-world.sh, seextask/README.md). Not part of the pinned 4,655-fixture corpus baseline, and not expected to hit 0 false positives — real pages carry their own unrelated markup errors.
- Fuzzing:
fuzz/is a standalonecargo-fuzzcrate (requires a nightly toolchain) that feeds arbitrary bytes through the wholecheck()pipeline. Local dev tool only — not run in CI:(thecargo +nightly fuzz run check fuzz/corpus/check tests/corpus/html
tests/corpus/htmlargument seeds the fuzzer from the existing corpus without copying it; crashes land infuzz/artifacts/check/). - Benchmarks:
benches/check.rs(criterion) timescheck()against a small typical page and a large, table-heavy real-world page fromtests/corpus/. CI only compiles the benchmarks (cargo bench --no-run) to catch bit-rot; run them locally for actual numbers:cargo bench
- Code: MIT License — see
LICENSEorLICENSES/MIT.txt. This is a REUSE-compliant project. - Third-Party & Vendored Assets: the RELAX NG schema and test corpus (
schema/,tests/corpus/) are vendored fromvalidator/validatorunder the MIT License, not authored by this project — seeTHIRD-PARTY-NOTICES.md.