Conversation
|
/workflow-approve |
|
The Darwin/macOS e2e job is failing on a fixture mismatch, Then commit the updated |
Apple Silicon does not implement IOPMCopyCPUPowerStatus, so fetchCPUPowerStatus returns kIOReturnNotFound there. Update returned that error straight away, which aborted the collector before updateTemperatures ran, so no temperature metrics were collected at all. The error was also not ErrNoData, so it was logged at error level on every scrape. On an M5 Pro, Update emitted 0 metrics and failed, while updateTemperatures on its own returned 52 temperature metrics. Treat kIOReturnNotFound as a system that does not report CPU power status: skip the three CPU power metrics, log at debug level and carry on to the temperature sensors. Any other non-success return code is still returned as an error. Systems that do report CPU power status are unaffected. This does not make the CPU power metrics available on Apple Silicon, since the underlying API provides no data. It only stops their absence from suppressing the temperature metrics. Adds a regression test covering the case. Signed-off-by: SaiPisey2 <piseysai0202@gmail.com>
8d1e633 to
907bf2f
Compare
|
Thanks for the review. Pushed, but regenerating the fixture turned up two things worth flagging rather than just committing the output. The fixture change itself I did not commit a regenerated file. Running
Worth knowing before anyone else follows the same instruction. The Darwin fixture path is built with I also added (Unrelated and left alone: the A real bug the fixture regeneration exposed This is the part I would not have caught otherwise. Once temperatures are actually collected, the collector emits duplicate label sets: and the scrape logs A sensor is identified only by its IOHID Since there is no property that makes them addressable, I report the first reading per name and count the rest in a debug line. The endpoint now returns 200 with 17 unique series and no gather errors. CI never sees this, since the runner has no sensors, but every Apple Silicon machine would have. Verified locally: Happy to split the two |
|
Thanks for the fixes. Let's split the On the duplicate sensors: dropping readings loses data. #3646 hit the same colliding-label problem for hwmon and fixed it by disambiguating the label instead (suffix it, falling back to something always unique when the first suffix still collides). Take a look there for the approach. |
Collecting temperatures on Apple Silicon surfaced a second problem that was previously unreachable, because Update returned before updateTemperatures ever ran. A sensor is labelled with its IOHID "Product" property, which is not unique. On an M5 Pro, 52 temperature-reporting services carry only 17 distinct product names, so the registry rejects the colliding samples and logs an error on every scrape. The collisions are of two kinds. Six "gas gauge battery" services are distinct sensors that merely share a name; they have distinct LocationIDs. The "PMU tdie*" services come in threes that share both the name and the LocationID, and report readings that differ by a few tenths of a degree, so they are separate sensing elements rather than the same one listed repeatedly. Dropping either kind loses readings. Following the approach taken for hwmon in prometheus#3646, read all sensors first, then qualify a name shared by several services with its location, falling back to the service registry ID when the location does not separate them either. Registry IDs are unique per service, which the OS guarantees and which held for every service observed here. Names that do not collide are left untouched. The endpoint now reports all 52 readings with unique label sets and no gather errors, where before the change 35 of them were rejected. Also update collector/fixtures/e2e-output-darwin.txt for node_scrape_collector_success{collector="thermal"}, which is now 1 because the collector no longer fails. resolveSensorNames is kept free of cgo so that the labelling rules are covered by table-driven tests, which matters because the CI runner reports no thermal sensors at all. Signed-off-by: SaiPisey2 <piseysai0202@gmail.com>
907bf2f to
1d0ecba
Compare
|
Both done. The You were right that dropping loses data — more than I realised before looking properly. Probing the services directly, the collisions turn out to be two different things:
So both kinds carry real readings, and the earlier "first one wins" would have thrown away 35 of 52. Following #3646: read all sensors first, then qualify a colliding name with its location, falling back to the registry ID when the location does not separate them either. Names that do not collide are untouched.
Result on the same machine, where before the change 35 of 52 samples were rejected: One deliberate choice worth flagging: the label resolution lives in Registry IDs are not stable across reboots, so those labels will change if the machine restarts. That is the same trade-off as the Verified: |
|
/workflow-approve |
|
Thanks for running the workflows. The Darwin e2e failure is the runner rather than this change — the diff is only: That runner has a Worth noting the same class of problem is why the fixture needs |
|
@nicolastakashi anything you need from me here? The only failure is the Darwin e2e picking up the runner's extra |
The spec said two lines (enable the thermal collector, widen the keep regex) would get the throttle ratios fleet-wide on the pinned 1.8.2. Measured on 2026-09-23, that is wrong. The darwin thermal collector asks macOS for the CPU power status before reading any sensor and returns when none is recorded, which is the normal state of a healthy machine, so 1.8.2 and the latest release, 1.12.1, both emit zero series. Apple Silicon does not implement that API at all, so the ratios never exist on this fleet. The fix is upstream and unmerged (prometheus/node_exporter#3767); built from source it reported 45 sensors on an M3 Pro. Mac temperatures therefore wait on that release, the bench test reads die temperatures from a node_exporter built with the fix instead of ratios that do not exist, and the scope no longer promises the two lines. The x86 nodes' open question is answered: they join the cluster, the edge node reads through its egress Service after #13546, and the switches report through the Omada controller after #13545. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…mal pressure No released node_exporter reports a temperature on a healthy Apple silicon host. Its darwin thermal collector asks macOS for the CPU power status first and returns before reading any sensor when none is recorded, which is the normal state of a machine that has never throttled. Measured on 1.8.2 (the fleet's version) and 1.12.1 (the latest): zero thermal series. The upstream fix, prometheus/node_exporter#3767, is unmerged, and shipping a patched build would mean building node_exporter with cgo and the Apple SDK on a Mac, since the operator image cross-builds everything from Linux. node_exporter also has no fan or power reading on macOS, and a dead fan is the failure the rack cares most about. infra/macos-host-sensors is a small binary that calls IOKit and the SMC through purego rather than cgo, so it cross-builds with CGO_ENABLED=0 in the operator image like tart-kubelet and the log shipper. It samples once and writes node_exporter's textfile collector a file with: - macos_sensor_temperature_celsius{sensor}, the HID temperature services, hottest reading per name (Apple silicon lists each die sensor three times), so a name is one stable series across reboots - macos_fan_speed_rpm, macos_fan_target_rpm and macos_fan_max_rpm per fan; a fan well below its target is failing - macos_system_power_watts (SMC PSTR) - macos_thermal_pressure_level, Apple's own 0-4 throttling summary The bootstrap installs it as a launchd job that runs every 30 seconds, and node_exporter now runs its textfile collector over /var/lib/tuist-host-sensors. The readings ride the existing tuist-macos-node-exporter scrape: no port, no egress Service, no tailnet grant. The step follows the log shipper's pattern: re-signed in place, the plist compared before reloading, convergent so that a missing binary uninstalls the job and its stale readings, run last on both the bootstrap and drift paths, and hashed (script and binary) so the drift loop delivers every change. The provider takes --host-sensors-binary-path, which the chart passes beside node_exporter's under macosFleet.tailscale.enabled. The macOS scrape keeps the new families and node_textfile_mtime_seconds (the sample's age). macos_* is outside the families the non-production filter drops, so every mini keeps it in every environment; rack minis also keep the textfile series through the rack_host exemption. This rolls to every Mac mini in every environment on the next drift pass, production included: node_exporter restarts with the textfile collector on, and the job is installed. Validation, on an Apple silicon Mac: - the binary built exactly as the image builds it (darwin/arm64, CGO_ENABLED=0) read 46 temperature services as 21 names, two fans, power and thermal pressure, and exited 0 - node_exporter 1.8.2 served the file through its textfile collector: 29 series, node_textfile_scrape_error 0 - the eight served metric families pass the rendered macOS keep regex, and the replayed destination filter keeps them for staging rack minis and production - go test and go vet for macos-host-sensors on darwin and linux, and for macos-host-bootstrap (new tests for the install, uninstall, gating, hash and the node_exporter flags; the per-host and hash-installer guards updated) and the provider's controllers - helm template for staging and production renders the flag; the rendered Alloy configs parse with alloy fmt; the workflow gains a test job that vets and cross-builds the darwin reader Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Addresses #2906.
Apple Silicon does not implement
IOPMCopyCPUPowerStatus, sofetchCPUPowerStatusgets backkIOReturnNotFound.Updatereturned that error straight away, which meant it never reachedupdateTemperatures, so no temperature metrics were collected at all. The error also isn'tErrNoData, so it was logged at error level on every scrape.Measured on an M5 Pro (darwin/arm64) against master:
52 usable temperature sensors on that machine, none of them exported, because an expected and unrelated condition aborted the collector.
Change
Treat
kIOReturnNotFoundas "this system does not report CPU power status" rather than a failure: skip the three CPU power metrics, log at debug level, and continue on to the temperature sensors. Any other non-success return code is still returned as an error, exactly as before, so systems that do report CPU power status are unaffected.After the change, on the same machine,
Update()returns no error and emits all 52 temperature metrics.Scope
This does not make the CPU power metrics appear on Apple Silicon and does not resolve #2218 — the underlying API provides no data there, as @rexagod already established on #2906. It only stops their absence from suppressing the temperature metrics. I've left
Closes #2906out of the commit deliberately, since whether that issue is fully answered by this is a maintainer call.Testing
Added
collector/thermal_darwin_test.go, which fails on master:and passes with the change. Also verified locally:
go build ./...go vet ./collector/gofmt -lclean on both touched filesgo test ./collector/full packagego build -tags notherm ./collector/