perf(block): precompute OTEL for chunker hot paths - #2236
Merged
Merged
Conversation
levb
marked this pull request as ready for review
March 26, 2026 21:32
Replace per-call `attribute.String(...)` allocation in every
`Slice`/`ReadAt`/`runFetch` with precomputed `metric.MeasurementOption`
values built once at package init.
- `Stopwatch.end`: build one `attribute.NewSet` and reuse across all three
instrument calls (histogram, sum, count) instead of three separate
`metric.WithAttributes(kv...)` allocations.
- New `PrecomputeAttrs(kv...)`: builds a reusable `MeasurementOption`.
- New `Stopwatch.Record(ctx, total, precomputedAttrs)`: zero per-call
attribute allocation alternative to `Success`/`Failure`.
- Exported `Success`/`Failure` attribute vars for use with `PrecomputeAttrs`.
- Add `precomputedAttrs` struct with all Slice/fetch attribute combos.
- Package-level `chunkerAttrs` var built once at init.
- All `timer.Success(ctx, n, attribute.String(...))` calls replaced with
`timer.Record(ctx, n, a.successFromCache)` etc.
FullFetchChunker.Slice on cache hit (64 MiB cache, 4K blocks, 4 MiB chunks):
```
│ baseline │ precomputed attrs │
│ sec/op │ sec/op vs base │
ChunkerSlice_CacheHit-16 20.57µ ± 2% 19.85µ ± 3% -3.49% (p=0.009 n=6)
│ baseline │ precomputed attrs │
│ B/op │ B/op vs base │
ChunkerSlice_CacheHit-16 9.211Ki ± 0% 8.047Ki ± 0% -12.64% (p=0.002 n=6)
│ baseline │ precomputed attrs │
│ allocs/op │ allocs/op vs base │
ChunkerSlice_CacheHit-16 16.000 ± 0% 4.000 ± 0% -75.00% (p=0.002 n=6)
```
75% fewer allocations (16 → 4) and 13% less memory per Slice call on the
NBD page fault hot path.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
levb
force-pushed
the
lev-telemetry-precompute
branch
from
March 26, 2026 21:39
1cbb5ad to
155ef71
Compare
dobrac
requested changes
Mar 27, 2026
- renamed `Record` -> `RecordRaw` - dropped the `a` alias - refactored `end()` to call `RecordRaw`
dobrac
approved these changes
Mar 27, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Replace per-call
attribute.String(...)allocation in everySlice/ReadAt/runFetchwith precomputedmetric.MeasurementOptionvalues built once at package init.Stopwatch.end: build oneattribute.NewSetand reuse across all three instrument calls (histogram, sum, count) instead of three separatemetric.WithAttributes(kv...)allocations.New
PrecomputeAttrs(kv...): builds a reusableMeasurementOption.New
Stopwatch.Record(ctx, total, precomputedAttrs): zero per-call attribute allocation alternative toSuccess/Failure.Exported
Success/Failureattribute vars for use withPrecomputeAttrs.Add
precomputedAttrsstruct with all Slice/fetch attribute combos.Package-level
chunkerAttrsvar built once at init.All
timer.Success(ctx, n, attribute.String(...))calls replaced withtimer.Record(ctx, n, a.successFromCache)etc.FullFetchChunker.Slice on cache hit (64 MiB cache, 4K blocks, 4 MiB chunks):
75% fewer allocations (16 → 4) and 13% less memory per Slice call on the NBD page fault hot path.