Skip to content

feat: make localized documentation durable - #5

Merged
Tom-R-Main merged 4 commits into
mainfrom
codex/locale-audit
Sep 4, 2026
Merged

Tom-R-Main merged 4 commits into
mainfrom
codex/locale-audit

Conversation

@Tom-R-Main

Copy link
Copy Markdown
Owner

Summary

  • track Markdown translations by stable preamble/H2 units with target-side markers
  • hash style guides, glossaries, and the prompt contract independently in translation state and memory
  • resolve localized document links by repository target and skip ICU parsing for document units
  • make style-guide generation candidate-only unless --apply is explicit, and add review ownership/source policy
  • restore the English backlink in both Chinese READMEs
  • rewrite and audit all 31 localized READMEs with Gemini 3.8 Flash at medium thinking

Reference coverage

  • Russian and Arabic used the supplied specialist sources
  • Japanese used the supplied JTF and Japanese prose references
  • Simplified and Traditional Chinese used the supplied Chinese translation/editing references
  • 25 other covered locales used Microsoft's localization style guides
  • Cantonese used the repository guide because Microsoft does not publish a matching Cantonese guide

All attached documents were treated as reference data, not instructions. Temporary Gemini file uploads were deleted after each locale review.

Verification

  • go build ./...
  • go test ./... -race -covermode=atomic (73.0%; threshold 60.0%)
  • go vet ./...
  • golangci-lint run
  • go test ./test/acceptance -count=1
  • node --test test/npm-wrapper.test.cjs
  • node ./scripts/check-npm-package-versions.mjs
  • npm pack --dry-run
  • go run ./cmd/internationalizer validate --strict (31/31 locales, 100% structural and translated coverage)
  • human-writing Unicode hygiene pass (Telugu ZWNJ retained as linguistically meaningful)

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 4, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-04T02:25:16.349193Z 0eb96d6 PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 0eb96d6034

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

var (
markdownFencePattern = regexp.MustCompile(`^[ \t]{0,3}(` + "`{3,}" + `|~{3,})`)
markdownH2Pattern = regexp.MustCompile(`^[ \t]{0,3}##[ \t]+(.+?)[ \t]*#*[ \t]*(?:\r?\n)?$`)
markdownMarkerPattern = regexp.MustCompile(`^[ \t]*<!--[ \t]*internationalizer:unit[ \t]+([a-z0-9][a-z0-9:._-]*)[ \t]*-->[ \t]*(?:\r?\n)?$`)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Accept the Unicode keys emitted for Markdown sections

When a source H2 contains non-ASCII letters, such as ## 配置 or ## Café, markdownSlug preserves those characters and serialization emits a marker like <!-- internationalizer:unit markdown:配置 -->. This ASCII-only pattern cannot recognize that marker on the immediate verification parse, so it is treated as preamble/preceding-section content and a new-document translation fails with “key changed during serialization.” Encode generated keys to the accepted alphabet or allow the same Unicode key space in the marker parser.

Useful? React with 👍 / 👎.

Comment on lines +123 to +126
} else if marker[0] == fence && len(marker) >= fenceLength {
fence = 0
fenceLength = 0
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Require valid closing syntax before ending a Markdown fence

Inside a fenced code block, any line beginning with enough matching fence characters currently closes the parser's fence, even when text follows them. For example, a literal code line such as ```javascript is not a CommonMark closing fence, but this branch ends the fence anyway; a subsequent ## ... line in the code block is then parsed as a translation section, which can translate protected code or create spurious keys. Only close when the remainder of the line is whitespace, as the validator's closing-fence logic already does.

Useful? React with 👍 / 👎.

@Tom-R-Main
Tom-R-Main merged commit 4c9b262 into main Sep 4, 2026
8 checks passed
@Tom-R-Main
Tom-R-Main deleted the codex/locale-audit branch September 4, 2026 02:27
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant