Skip to content

Reduce noise in AttributeReadingPerformanceTests scaling assertion - #133496

Merged
krwq merged 2 commits into
mainfrom
fix-attribute-perf-test-noise
Sep 11, 2026
Merged

krwq merged 2 commits into
mainfrom
fix-attribute-perf-test-noise

Conversation

@krwq

@krwq krwq commented Sep 9, 2026 •

Copy link
Copy Markdown
Member

Fixes the intermittent CI failure tracked in #131973 by making the scaling-ratio assertion in AttributeReadingPerformanceTests resilient to CI wall-clock jitter, without changing what regression it detects and without increasing test runtime.

The flake

The test asserts largeTime / smallTime <= 10 * MaxRatioMultiplier. A recent failure:

Scaling ratio 41.4x exceeded 40.0x limit (input grew 10x). Small (2000): 23 ms, Large (20000): 953 ms.

That's about as marginal as a failure can get — a single ~1–2 ms wobble on the small run flips the result:

small time ratio (large=953 ms) verdict
23 ms 41.4x ❌ fail
24 ms 39.7x ✅ pass
25 ms 38.1x ✅ pass

Same failure has been logged against four unrelated reader subclasses (XmlReaderCreate, XmlNodeReader, XPathNavigatorReader, WrappedReader), including readers that don't even go through the XmlTextReaderImpl duplicate-check path being stressed by the test (they read from a pre-parsed XmlDocument / XPathDocument). That's the fingerprint of harness noise, not a product regression.

Why the lower bound is the fix

smallTime is the denominator of the scaling ratio, so its noise dominates the assertion. When SmallN = 2_000:

  • The _BelowThreshold variants (attrCount < 64) do very little parser work — measured wall time is ~5 ms locally and ~20 ms on the affected CI agent.
  • At that scale, ordinary CI jitter (GC pause, scheduler blip, tiered-JIT recompilation) is a large fraction of the baseline. A 2 ms wobble on a 20 ms baseline is ±10 % on the denominator, which turns straight into 10 % ratio inflation against the 40x cap.
  • The large side is not the problem — it's hundreds of ms to a few seconds, so jitter there is a rounding error.

So the correction is to lift the small-run effective denominator out of the noise floor.

Approach: baseline floor (no runtime cost)

Instead of scaling up SmallN/LargeN (which would double or triple test wall time), this PR introduces a 30 ms baseline floor on the divisor:

private const long MinBaselineMs = 30;
...
long baseline = Math.Max(smallTime, MinBaselineMs);
double actualRatio = (double)largeTime / baseline;
  • If smallTime >= 30 ms, the assertion behaves exactly as before.
  • If smallTime < 30 ms (the noise regime), the divisor is clamped to 30, absorbing sub-30 ms jitter without letting it inflate the ratio.

30 ms is deliberately just above the observed noise floor (the flaking runs measured ~20–25 ms). It's small enough that a true O(N²) regression still blows past the 50x cap easily.

Together with a slightly relaxed MaxRatioMultiplier (4 → 5, giving a 50x cap) this eliminates the observed flake mode.

Regression detection is preserved

Prior-art measurements on the pre-fix (#130968) code path — the very code these tests were written to guard — showed scaling ratios of 75–81x for 10× input growth on the long-URI variants (the actual O(N²) signature we care about). With the 30 ms floor and SmallN = 2_000, a regression at that scale would compute as largeTime / 30ms, i.e. still well above 50x. Detection stays intact.

Sensitivity table (limit = 50x, input growth = 10x):

smallTime largeTime raw ratio floored ratio (÷ max(s, 30)) verdict
5 ms 50 ms 10x 1.7x ✅
23 ms 953 ms 41x 32x ✅ (was ❌ before)
30 ms 300 ms 10x 10x ✅ (floor inert)
100 ms 1000 ms 10x 10x ✅ (floor inert)
20 ms 1600 ms 80x 53x ❌ regression caught
30 ms 2400 ms 80x 80x ❌ regression caught

Other changes

  • IsNotCoreClrInterpreter → IsNotInterpreter — skip on the Mono interpreter too. That runner also produces highly variable timings on this workload.
  • Error message now reports both the raw smallTime and the effective baseline used, so a future failure clearly shows whether the floor kicked in.
  • Warmup comment added to the pre-measurement read.

Change (constants + a Math.Max)

src/libraries/System.Private.Xml/tests/Misc/AttributeReadingPerformanceTests.cs:

-        private const double MaxRatioMultiplier = 4;
+        private const double MaxRatioMultiplier = 5;
+        // Floor for the small-side divisor: below this, wall-clock noise dominates the reading.
+        private const long MinBaselineMs = 30;
...
-            double actualRatio = (double)largeTime / Math.Max(smallTime, 1);
+            long baseline = Math.Max(smallTime, MinBaselineMs);
+            double actualRatio = (double)largeTime / baseline;

SmallN / LargeN are unchanged from main, so total test wall time is unchanged.

Fixes #131973.

Note

This pull request was drafted with the help of AI. Please review before merging.

The scaling-ratio assertion in AssertLinearScaling was intermittently failing on CI (issue #131973) because SmallN=2000 produced measurements as small as ~5-23 ms. At that scale, ordinary CI jitter (GC pauses, scheduler blips, tiered-JIT recompilation) is a large fraction of the baseline, and small denominators inflate the ratio. A single ~20 ms wobble in the small run flipped a healthy 30x reading into a 41x failure against the 40x cap.

Bumping SmallN from 2000 to 4000 (and LargeN from 20000 to 40000 to preserve the 10x growth ratio) roughly doubles the small-run baseline into the tens of milliseconds so jitter becomes a small percentage of the measurement rather than the whole thing. Nudging MaxRatioMultiplier from 4 to 5 adds a small extra CI headroom for any residual noise; quadratic regressions still trigger easily (prior pre-fix data measured 75-81x on 10x growth on the affected code paths).

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot AI lite review requested due to automatic review settings September 9, 2026 13:23
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 3 pipeline(s).
13 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟢 Approval recommended

The change is limited to test constants (no product code impact) and the updated values remain internally consistent (preserving the 10× size ratio while adjusting the allowed scaling bound).

Pull request overview

This PR adjusts the input sizes and scaling tolerance used by AttributeReadingPerformanceTests so the linear-scaling assertion is less sensitive to wall-clock jitter in CI, while keeping the same overall “large vs small” ratio model.

Changes:

  • Increase the “small” input size from 2,000 → 4,000 to move the denominator further out of the noise floor.
  • Increase the “large” input size from 20,000 → 40,000 to preserve the 10× SizeRatio.
  • Increase MaxRatioMultiplier from 4 → 5 (raising the cap from 40× → 50×).
File summaries
File Description
src/libraries/System.Private.Xml/tests/Misc/AttributeReadingPerformanceTests.cs Updates scaling-test constants (SmallN/LargeN and max ratio multiplier) to reduce intermittent failures from timing jitter.
Review details
  • Files reviewed: 1/1 changed files
  • Comments generated: 0
  • Review effort level: Lite

@adamsitnik adamsitnik left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM!

@krwq

krwq commented Sep 10, 2026

Copy link
Copy Markdown
Member Author

seems there is still some issue, I'm guessing overall timeout since I increased test run. Will take a look in a min

- SmallN/LargeN back to 2_000/20_000 (no runtime increase vs main)

- MinBaselineMs=30 floors the small-side divisor so wall-clock noise on fast readings doesn't inflate the ratio

- Switch conditional from IsNotCoreClrInterpreter to IsNotInterpreter so Mono interpreter is also skipped

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot AI review requested due to automatic review settings September 10, 2026 15:24

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

The new absolute baseline clamp (MinBaselineMs) can significantly weaken regression detection on fast machines and the PR description does not match the actual code changes (e.g., SmallN/LargeN not updated).

Once you've addressed the issues Copilot identified, you can request another Copilot review.

Review details
  • Files reviewed: 1/1 changed files
  • Comments generated: 2
  • Review effort level: Lite

@krwq

krwq commented Sep 11, 2026

Copy link
Copy Markdown
Member Author

seems like current failures are known. Fix slightly differs from original to reduce time we run the test for the lower bound I'm assuming some minimum time - if it's below threshold I treat it as noise. I picked 30ms. 12ms is what I treat as measurement error from personal experience this is slightly above double that. Additionally I kept slightly more relaxed multiplier. This should still detect regressions but hopefully not cause issues.

@krwq
krwq merged commit f8bd182 into main Sep 11, 2026
80 of 83 checks passed
@krwq
krwq deleted the fix-attribute-perf-test-noise branch September 11, 2026 10:33
@dotnet-milestone-bot dotnet-milestone-bot Bot added this to the 12.0-preview1 milestone Sep 14, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

4 participants