Both MCP hosts (Darling and Lite) register one shared analysis service. When a pass is already running, AnalyzeAsync returns an empty list, and the tool then reads the running pass's state. The second caller gets "No significant findings. All metrics are within normal ranges" for a server that was never analyzed, with the other pass's coverage and fact count attached. MCP clients often call tools in parallel, and the window is the whole length of the first pass. Scheduled analysis is not affected, so alerts are not suppressed.
Expected
Each call analyzes its own server, whether or not another call is running.
Where
- Darling host:
Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpHostService.cs:530 (one shared instance) and :694 (Stateless = true, no gate). Tools: DarlingMcpTools.cs:79,82,101,135,148-177.
- Darling service:
Darling/PerformanceMonitor.Darling.Analysis/DarlingAnalysisService.cs:328-331 returns an empty list before the try when busy. The busy check is not atomic.
- Lite:
Lite/Mcp/McpHostService.cs:73,95, Lite/Analysis/AnalysisService.cs:183-184, Lite/Mcp/McpAnalysisTools.cs:45-143.
- The code already notes the hazard on a sibling path:
DarlingAnalysisService.cs:699-707 and Lite AnalysisService.cs:507-515.
- Not affected: the worker (a new service per pass,
DarlingWorker.cs:8453), Lite's scheduler and Recommendations tab, and the web host, which never calls AnalyzeAsync (DarlingWebEndpoints.cs:110).
Line numbers are as of 1248a67.
Fix
Tests that must fail before the fix
- Two concurrent
analyze_server calls for different servers each return their own server's result. Today the second returns the all-clear.
- Update
McpServiceParameterDiSeatCensusTests and McpToolLatencyRecordingTests.cs:74 on purpose, for the new registration.
Both MCP hosts (Darling and Lite) register one shared analysis service. When a pass is already running,
AnalyzeAsyncreturns an empty list, and the tool then reads the running pass's state. The second caller gets "No significant findings. All metrics are within normal ranges" for a server that was never analyzed, with the other pass's coverage and fact count attached. MCP clients often call tools in parallel, and the window is the whole length of the first pass. Scheduled analysis is not affected, so alerts are not suppressed.Expected
Each call analyzes its own server, whether or not another call is running.
Where
Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpHostService.cs:530(one shared instance) and:694(Stateless = true, no gate). Tools:DarlingMcpTools.cs:79,82,101,135,148-177.Darling/PerformanceMonitor.Darling.Analysis/DarlingAnalysisService.cs:328-331returns an empty list before thetrywhen busy. The busy check is not atomic.Lite/Mcp/McpHostService.cs:73,95,Lite/Analysis/AnalysisService.cs:183-184,Lite/Mcp/McpAnalysisTools.cs:45-143.DarlingAnalysisService.cs:699-707and LiteAnalysisService.cs:507-515.DarlingWorker.cs:8453), Lite's scheduler and Recommendations tab, and the web host, which never callsAnalyzeAsync(DarlingWebEndpoints.cs:110).Line numbers are as of 1248a67.
Fix
DarlingMcpHostService.cs:530andLite/Mcp/McpHostService.cs:73, the way the worker already does it. Keep passing the sharedBaselineCache(Every scheduled analysis pass recomputes every 30-day baseline: a fresh DarlingAnalysisService per pass never reuses (or shares with MCP) the bucket cache #3941) so baselines are not recomputed. With one instance per call, no semaphore is needed.Tests that must fail before the fix
analyze_servercalls for different servers each return their own server's result. Today the second returns the all-clear.McpServiceParameterDiSeatCensusTestsandMcpToolLatencyRecordingTests.cs:74on purpose, for the new registration.