Skip to content

[M10 correctness] Harden LSP URI, UTF-16, and document lifecycle boundaries #74

Description

@Teakowa

Parent: #27

Depends on: #72, #73

Goal

Make the LSP protocol boundary correct for real standard clients by centralizing URI/path conversion, UTF-16 position handling, and document lifecycle behavior instead of relying on ASCII/hand-built URI assumptions.

Context

M10 currently advertises UTF-16 and has basic non-BMP coverage, but several code paths still mix Unicode scalar counts, UTF-8 byte slicing, and LSP UTF-16 offsets. File URI handling is also largely manual (file:// stripping/concatenation), which is insufficient for percent-encoded paths, spaces, Unicode filenames, and platform-specific filesystem paths. Lifecycle validation is incomplete relative to the declared M10 contract.

Scope

  • Establish one shared, well-tested boundary for LSP UTF-16 position/range ↔ Rust/source/compiler position conversion.
  • Remove any direct UTF-16-offset-as-UTF-8-byte-index slicing in completion/member-context and related paths.
  • Ensure full-document and source edit ranges use UTF-16 units, including trailing-line/non-BMP cases.
  • Replace manual file:// stripping/concatenation with robust URI ↔ filesystem path conversion appropriate for lsp_types::Uri/standard file URIs.
  • Cover percent encoding, spaces, Unicode paths, normalization, and supported Windows drive-path behavior.
  • Define and validate document lifecycle behavior for didOpen/didChange/didSave/didClose.
  • Ensure close/save behavior is explicit even where the correct behavior is intentionally minimal/no-op.
  • Keep the LSP adapter thin; protocol conversion utilities must not duplicate compiler semantics.

Non-goals

  • Shipping editor extensions.
  • Adding new language-service features.
  • Introducing asynchronous request infrastructure solely for this issue.
  • Expanding non-file URI support beyond evidence-backed Wright use cases.

Acceptance criteria

  • No LSP UTF-16 position is directly used as a Rust string byte index.
  • Completion/member access, hover/navigation, semantic tokens, diagnostics, and edit ranges have non-ASCII/non-BMP regression coverage where they consume or emit positions.
  • A full-document rename/edit ending on a line containing non-BMP characters produces a correct LSP UTF-16 end position.
  • File URIs containing spaces/percent encoding and Unicode filenames resolve to the intended filesystem paths and round-trip to valid LSP URIs.
  • Supported platform path cases are tested without hand-built malformed file URIs.
  • didOpen/didChange/didSave/didClose behavior is protocol-tested; didClose correctly retires diagnostics/state as defined by [M10 correctness] Make diagnostics source-aware and dependency-reactive #72.
  • Existing cross-file navigation/rename and overlay tests continue to pass through the hardened URI/position boundary.
  • The standard LSP protocol adapter remains semantically thin.

Planning notes

Do this after #72 and #73 so the semantic/project contracts are stable before hardening their protocol encoding. Prefer a small centralized conversion layer over feature-specific fixes.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions