Fix puck activation after firmware updates - #44
Conversation
…dpoint # Conflicts: # src/main.cpp # src/steam/SteamController.cpp
The shared-access fallback fixed a real problem — something in Windows holds a compatible write handle on the live puck collection even with Steam closed — but applied unconditionally it also swallowed the signal the Steam handoff depends on. A failed exclusive claim is precisely how the app detects that Steam is holding the controller: game mode failing to activate is what escalates TryAcquireController to the device cycle that takes the handle back. With the fallback always succeeding, the app would instead report success and drive the controller alongside Steam, both sending competing lizard-mode feature reports, and would never cycle. None of the PR's validation would catch it, since every check was run with Steam closed. Gate it on whether steam.exe is running. TrayApp already watches that for the auto modes, so it now pushes presence into ControllerManager from the WM_STEAMSTATE handler — outside ApplySteamState, which returns early in Manual mode, because this decision matters in every mode. Also cheapen the state-report probe. It costs the full 250ms timeout on an empty puck slot and runs per slot, while the acquire path retries in a burst of eleven — three empty slots would have blocked the UI thread for roughly eight seconds. Slots found silent are skipped for 1.5s, so the burst probes each at most once. And WaitForStateReport now stops when a read returns instantly rather than timing out, instead of spinning on a dead handle for the rest of the window. The skip message no longer says "puck", since the gate applies to every transport. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
@danbosscher Thanks for getting this up so quickly. You truly are the "valve broke the app again" hero that we needed. Also, the In order to respect your time, and the fact that I muddled everything by pushing new fixes before seeing this, I took the time to rebase and get this ready to land. **1. I added support for BT connectiones between now and then. Two conflicts:
2. Gated the shared-access fallback on Steam not running. This is one part I'd like your eyes on, because it deviates the most from what you designed here. What you saw here is true, something in Windows does hold a compatible write handle on the live puck collection even with Steam closed. But applied unconditionally, it swallows a signal the app depends on elsewhere: a failed exclusive claim is how we detect that Steam is holding the controller. Game mode failing to activate is what escalates That doesn't show up in validation unless you're also testing handing off to Steam. So I also cheapened the state-report probe. It costs the full 250ms on an empty slot and runs per slot, and the acquire path retries in a burst of eleven — on a four-slot puck that's roughly eight seconds of blocked UI thread. Silent slots are now skipped for 1.5s so the burst probes each at most once, and |
|
1 > Awesome! 2> I think design wise this is 100% right, only try and take over when Steam is idle. However, on my machine, the Windows handle persisted even though Steam was idling - I'll try a few different state combinations and circle back |
|
heya @ddeverill - on my machine, since this firmware update, I can't get the Steam Controller working reliably until steam.exe is closed. I don't have other folks in my network with a Steam Controller + Windows to test this with unfortunately. I tested in the new Halo w/ your changes in, but ran into several issues regardless of Steam’s controller settings: Steamless first, then Steam: one D-pad press moved two menu entries because both systems produced input. If we want to enable steam.exe to keep running, it looks like we'd have to do something similar to HidHide and hide the controller from Steam: https://github.com/nefarius/HidHide And it looks like that uses a kernel driver... bit of a messy path. May be better to offer integration instead or follow a different path. I'll keep looking around a bit for options. |
|
There is a -'nojoy' commandline option when launching Steam. Alternatively, Steam lets you blacklist the steam controller, so that could be an on/off toggle. However, that's done in config.vdf and persisted across Steam restarts. |
|
I have not been able to find a way to get the steam controller working while steam.exe and steamless coexist normally. I have to use -nojoy, blacklisting or closing steam on my machine or I get double inputs or no inputs at all. I wonder if this is an issue on my machine only, but I've not been able to find an issue on my system. I'm just going to implement every possible option, list them under a Debug menu, and ask for user feedback, defaulting to this device blacklist, given that other projects went that route too. I'm also updating the uninstaller so we don't leave people's config.vdf touched. Have a look and let me know which direction you want to go in? |
|
(codex notes): Consolidated updateThis PR has grown beyond the original puck activation fix, so here is the current state in one place. 1. Puck/firmware activation fix
2. Steam coexistence redesignThe main finding from testing is that "could not get an exclusive handle" is no longer reliable proof that Steam owns the controller. Depending on startup order, treating it that way caused either duplicate input or a virtual Xbox pad that received no input. The new behavior therefore fails closed: Steamless takes control only when Steam is absent or there is current-session proof that Steam cannot see the physical controller. If that proof disappears, Steamless immediately yields before both mappers can emit input. The Debug: Steam coexistence menu now exposes all three workable policies:
Hovering each option explains its tradeoffs. Selecting one releases the controller, stops Steam, applies the configuration, and relaunches Steam. If a Steam game is running, the app asks before closing it. The original-support option additionally exposes only the handoff settings that apply to it:
For safety, blacklist mode is not trusted from 3. Toggle and configuration lifecycle
Before its first change, Steamless records which managed Valve PIDs and whether Uninstall now:
4. Transport, sleep, and diagnostics
|
d690e1b to
b189bf3
Compare
The Steam coexistence work (strategy menu, config.vdf blacklist, -nojoy, Steam restart, uninstall cleanup, tray toggle changes) is being split into its own PR so the puck fix — which is the P0 here, since the receiver is broken today — can land on its own. Nothing is lost: those commits remain in this branch's history and can be cherry-picked onto a fresh branch for the follow-up. This is a revert of the tree, not of the work. Reverts the branch tree to a532804, the last state reviewed and built. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Carried over from the coexistence branch as the one piece of it that is self-contained: 0x1305 alongside the Proteus puck at 0x1304, and the same path token treated as a dongle transport. Purely additive — it only affects users who own that receiver and get nothing today. Untested here; no Nereid hardware on hand. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
Hey Dan! Thanks for all the context here, and oh boy thank you Codex for the summary at the end :) Made it much easier to understand what's going on and the changes that were made. The coexistence investigation here is super important, and will likely take some time and back and forth to get to the bottom of what we want to do here. The work you've done to get our options up and running and get some actual on-handset feedback is great and I don't want to lose it. We need to get to a resolution here. But, I also want to be sure we're fixing the puck issue that you've already fixed at the beginning while we hash out what's going on. That's why I want to pull that extra work out, land the Puck fix, and then kick off our conversation in a fresh PR that we can hack on. Here's some behind-the-scenes information on why I want to take some extra time here. One of the guiding principles I had for this when I was working on it was that I wanted to be sure that we're doing right by Steam. This meant not doing anything to hide the controller from the system, or possibly get into a state where I'm either using HIDHide or hiding controllers from Steam and then if SteamlessController crashes or fails, a less tech-savvy customer now has a controller they can't recover easily. I've tried to make sure I'm using the default ejection/handling systems built into Windows so things are recoverable and standard. I worry that some of the options we've got going forward will cause me to compromise a bit here, and want to take the appropriate time to weigh the options, but also understand what might be happening better. With all that said, I really appreciate your efforts to try to mitigate those concerns (one good example here is blocking uninstall until we can confirm we put Steam back how we found it). I'm going to pull in the Nereid forward into the original feedback and take that. I'll run a quick QA pass on the bits as they exist, and spin up a new release to help folks that might have a puck that doesn't work. ==Will also include this in the next PR for the coexistence issue== Some thoughts on the path forward for the coexistence conversation.
Ok, now actionable feedback :) Before we kick off a redesign around the issues you're seeing with unavailability, we need to rule out one thing I spent a couple of days in the last PR I submitted (for BT support) on the reclaim path and found two independent bugs tha made the device cycle a silent no-op. My scheduled task pointed at a stale helper that didn't know the current PIDs, and 'DIF_PROPERTYCHANGE' reports success while not actually restarting a devnode that has an open handle. Both of those issues should be fixed on 'main' now; cycling uses disable/enable, and there's a 'cycle.log' that distinguishes 'CYCLED" from "FELL BACK TO RESTART". Since you're running SteamlessController from a build folder, you might have no registered task at all, which means that every reclaim attempt on your machine reported success and didn't actually do anything. This was an issue I had when I was building the last patch and it gave me similar behavior to what you're reporting. On my end, with current 'main', a cycle broke Steam's handle over BT and we were back in action (I don't have a tilde on this keyboard)1.2 seconds later. |
Three defects surfaced by QA against the puck. The battery payload is [chargeState, percent] on every transport, not just Bluetooth. Dongle reports read 1/92 and 1/91 — a constant 1 (CHARGE_STATE_DISCHARGING) followed by the real level — so a 92% battery was being reported, and fed to the DS4 battery level, as 1%. Game mode could not tell an idle receiver from a contested one. A puck publishes a slot interface whether or not a controller is paired into it, so with the controller switched off every slot is silent, which the acquire path read as "blocked" and answered with three device cycles. That restarted the receiver's devnodes repeatedly, logged a misleading ownership error, and on a machine without a registered cycle task would also throw a UAC prompt at someone who had merely not turned their controller on yet. EnableGameMode now reports NoActiveController separately, and the acquire path waits for the wake-up device-change event instead of cycling. Finally, the app announced failure while the last cycle was still landing. Observed: gave up at 20:12:24.478, arrival at 20:12:25.396, game mode enabled at 20:12:25.415 — the balloon fired a second before success. A cycle on a four-interface receiver outruns the 2.5s retry timer once the helper's disable/enable pause is included. The timer goes to 4s, and the verdict is deferred behind a grace timer that stays quiet if the controller arrives after all. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Acquisition worked and then destroyed itself. A cycle is asynchronous — fired through Task Scheduler, with the helper pausing a second between disable and enable — and every device arrival it produces re-enters TryAcquireController. Finding the claim still blocked, that fired another cycle on top of the one already running. From the log: cycle 1 at 20:22:52.515, cycle 2 at 20:22:54.857 (2.3s later, triggered by cycle 1's own arrivals rather than the 4s timer), game mode enabled at 20:22:55.355 from cycle 1 — then cycle 2 landed and removed the device at 20:22:56.367, disabling what had just succeeded. It recovered at 20:22:57.537, but from the outside this reads as the app flapping and never taking the controller. Put a floor under the spacing: no new cycle within 4s of the last, just re-arm the retry timer and let the one in flight finish. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Regression from the previous commit. Skipping the device cycle when no slot is emitting state was right, but it left nothing to retry: the code assumed a controller switching on would arrive as a device-change event. It doesn't. A receiver publishes all of its slot interfaces permanently, whether or not a controller is paired into one, so waking a controller creates and removes no HID interface and Windows says nothing. Measured: 88 seconds of complete log silence after the controller was switched on, with the app sitting there having skipped all four silent slots. Previously the cycling loop retried often enough to stumble across the wake, which is why this never showed. Arm a poll instead. The probe is 80ms rather than 250ms because a live slot streams continuously — 80ms is around twenty reports on the dongle — so polling four silent slots costs ~320ms every two seconds while idle, and stops the moment a controller appears. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Reclaiming took 9.1s on a puck, 6.6s of which was the cycle helper. CycleDevices restarted each interface in turn and each paid its own one-second settle, so a four-slot receiver paid it four times. Disable the whole set, settle once, enable the whole set — and drop the settle to 400ms, since the one second was a conservative guess rather than a measured requirement. Callers get a per-device outcome so the log can still tell CYCLED from FELL BACK TO RESTART, and now also from LEFT DISABLED. Also fixes a multi-controller bug in the wake poll. EnableGameMode reports Enabled if any slot came up, and the acquire path then killed the poll outright — so with player one active, player two switching on would never be noticed, because a receiver raises no device event for a slot becoming occupied. Keep polling while any slot is still inactive, at a slower cadence than when we have nothing: each poll briefly blocks the UI thread per empty slot, and this case is opportunistic rather than something a user is waiting on. Untested against multiple controllers — no hardware here for it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Batching the cycle was wrong on both counts it was justified by. It saved nothing. The settle was not the dominant cost — the SetupDiCallClassInstaller disable calls are, and they are paid per device either way. Batched, the whole cycle still took ~5s. Worse, it broke the reclaim. Cycling sequentially happens to bring the occupied slot back last and alone, giving the app an uncontested shot at it: at 20:41 the live slot re-enabled at .690 and game mode came up at .735, winning by 45ms. Released together, all four interfaces arrive at once, the app spends its time probing the empty ones, and Steam claims the live one first. Three cycles in a row failed that way, ending in the give-up alert, with a flood of Steam connect/disconnect toasts along the way. Keeps the 400ms settle (down from 1s) and the per-device outcome reporting, both of which stand on their own. Real speed-up here is to cycle only the contested interface rather than all of them, which also fixes reclaiming one controller disrupting the others. That needs the helper to learn which path to target, so it is its own change. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A reclaim restarted every Valve device present. On a four-slot receiver that meant three empty slots torn down for nothing, and in a multi- controller setup it dropped controllers Steam had never touched — one person quitting a game briefly killed everyone's pad. The app already knows which interface was refused, so it now names it. The scheduled task runs the helper with no arguments and cannot be given any, so the paths are handed over through a file the helper consumes and deletes; an empty or stale handover still means "cycle everything", which keeps a helper run by hand working as before. A handover older than a minute is discarded so a crashed run cannot steer a later cycle. Targets are captured before ReleaseDevices, which clears the slots. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Three attempts to make the reclaim faster all made it worse, so this returns the tree to 1260656 — the state verified end to end on the puck, where the reclaim took 9.1s but succeeded on the first cycle with no spurious failure alert. Reverted: - batching the cycle (saved nothing, and losing the sequential order cost the race against Steam) - the 400ms settle and per-device outcome reporting that rode along with it - cycling only the contested interface Also reverted, and worth re-doing on its own: keeping the wake poll armed while another controller is active. That fixed a real multi-controller gap — player two switching on is never noticed, because a receiver raises no device event for a slot becoming occupied — but it was never verified, and it does not belong in a change meant to restore known-good behaviour. The 9.1s reclaim is dominated by SetupDiCallClassInstaller disabling four interfaces in turn. Making that faster means not cycling all four, which is exactly what failed here; it needs to be worked out with hardware rather than reasoned about. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Summary
FF00/0002management endpoint from controller enumeration.SteamProbeendpoint diagnostics and fix its/WXbuild warnings.Root cause
The puck exposes four
FF00/0001controller slots plus anFF00/0002management interface. v1.8 opened every interface, treated failure to obtain exclusive write access as fatal, and sent the lizard-mode feature command before reading from the slot.On the current puck firmware, empty slots do not emit state reports and reject the command with
ERROR_GEN_FAILURE(31). The live slot accepts the command after it has emitted a0x42or0x45state report. Windows also keeps a compatible write handle on the live controller collection, producingERROR_SHARING_VIOLATION(32) for an exclusive reopen even with Steam closed and Valve's Xbox Enhanced Features driver uninstalled.Validation
VID_28DE&PID_1304; live slot wasmi_02, reporting0x42/54 bytes.mi_02, and recordsGAMEMODE: enabled.ERROR_DEVICE_NOT_CONNECTED(1167).Fixes #40