Repository navigation
get_ag_health caps its fleet-wide answer, most severe first, and says when it cut (#4471) - #4474
Merged
Merged
Conversation
Adds a limit parameter to get_ag_health (default 11, matching get_analysis_findings' pattern), orders groups worst severity first then by the largest lag/queue magnitude, and adds groups_returned / groups_total / groups_truncated (+ note) to the envelope.
erikdarlingdata
marked this pull request as ready for review
September 27, 2026 17:15
This was referenced Sep 27, 2026
get_ag_health returns ~266 KB fleet-wide with no cap: too large for an MCP client's first call
#4471
Closed
erikdarlingdata
added a commit
that referenced
this pull request
Sep 28, 2026
…, not a fixed group count (#4568) get_ag_health fills its fleet-wide answer to the 32 KB budget by size, not a fixed group count. Refs #4471, #4474. - DarlingAgReader.Build walks the availability groups most-severe-first, serializing each group once and adding its bytes to the envelope's, and stops before the group that would cross McpResponseBudget.DefaultBytes. At least one group always comes back. - The response actually returned is then re-measured, note included, and the last group dropped (the note rebuilt) while it is over the budget. - The default of 11 groups stays as an upper bound, an explicit limit is still honored, and groups_returned, groups_total, groups_truncated and the note stay truthful; the note says whether the size budget or the limit cut the page. - The tool and parameter descriptions say so, within the tools-list budget. - Tests (DarlingAgReaderTests): 42 groups of 10 databases each come back within 32 KB and truncated (280,142 bytes on the code before the change); 5 small groups all come back; one oversized group returns alone, not truncated; an explicit small limit is exact; and groups that fit only without the note are cut to fit with it. The live test checks that the serialized response fits the budget.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #4471.
Why
get_ag_health's only parameter wasserver_name. A fleet-wide call (the tool's own description invites omitting it) returned every (reporting server, AG) group with no cap: a production fleet measured 265,794 characters, over an MCP client's typical per-result token limit, and got spilled to a file or refused outright.get_analysis_findingsalready carries the pattern this tool lacked — alimitwith a stated default, most-severe-first ordering, and a truncation flag.What changes
get_ag_health(and the sharedDarlingAgReader.GetAgHealthAsync/Build) gained alimitparameter, default 11, on the unit a caller reasons about: groups (one monitored server's view of one AG).secondary_lag_secondsor queue depth (KB) anywhere in the group. Lag and queue depth stay un-banded, by design (see the reader's own doc comment) — this tie-break only decides which same-severity groups a cut drops first, it never changes a group'sseverity.groups_returned,groups_total,groups_truncated(+groups_truncated_notewhen true), naming the same trioget_analysis_findingsuses for its ownfindings_truncated.availability_group_count/distinct_ag_count/worst_severitystill describe the WHOLE scope, not just the returned page./api/read/get_ag_healthdispatch row and the API catalog gained thelimitparameter to match;/api/ag(the dashboard's own read) is unchanged and stays uncapped — it renders its own page, not an MCP client's result buffer.get_ag_healthis Darling-only, per its own doc comment), so no Lite change was needed.Test plan
Live Postgres pins in
DarlingMcpAgToolsTests.cs(container rig, roledarlingon a Timescale image):AgHealth_DefaultLimitCapsMostSevereFirstAndFlagsTruncation_AgainstDevPostgres(new): seeds 42 HEALTHY single-replica groups plus one CRITICAL group (a suspended database) planted LAST by both insertion order and name, then asserts the default call (nolimitargument) returns exactly 11 groups,groups_returned/groups_total/groups_truncated/groups_truncated_noteare correct, and the critical group is first. Also assertslimit >= totalclears the flag.DarlingMcpAgToolsLivePostgresTests/DarlingMcpAgToolsSurfaceTestspins pass unchanged (the surface pin's expected parameter list was updated toserver_name,limit).McpToolsListBudgetTests(5/5),McpToolGuideHeadsAgStoreTests(5/5),DocCommentHygieneTests(77/77),DarlingWebEndpointsTests(70/70),AvailabilityGroupCountReadTests,DarlingAgStatesReaderLiveEqualityTests,DarlingAgStatesReaderTests,PgTargetMcpSurfaceTests,DarlingCoreToolProfileTests— all green.origin/dev(6f50c73b0): a throwaway test file, compiled clean on dev (it callsGetAgHealthwith nolimitargument, since dev's signature has none), seeded 15 groups on one server and called the tool with defaults — failed at runtime asserting<= 11groups returned andgroups_truncated == true, because dev returns all 15 uncapped with no such field. Removed after confirming.groups.Sort(CompareGroups)— the new severity-first pin went RED. (2) removed thelimitcut (pagedGroups/groupsTruncatedstayed unconditional) — the same pin went RED (no truncation, 43 groups back). Both reverted; the pin is GREEN again on the current branch (7/7 in the class).Sizes (43-server / 3-AG-per-server / 2-replica / 3-database fixture, ~2.7 KB/group on this shape, via a throwaway harness against the built DLLs):
limitset past the total, i.e. the old, unbounded behavior on this fixture): 348,187 characters.limit=11): 30,115 bytes — under the shared 32 KB budget.limit=12measured 32,812 bytes, over it, which is why 11 is the largest default that clears the budget on this shape.CHANGELOG entry
SECTION: Changed
ENTRY:
get_ag_healthreturns at most 11 Availability Group views by default, the least healthy first, and says when more exist ([get_ag_health caps its fleet-wide answer, most severe first, and says when it cut (#4471) #4474]) - a fleet-wide call on a large fleet returned hundreds of kilobytes; scope byserver_nameor raiselimitfor the rest.REF:
[get_ag_health caps its fleet-wide answer, most severe first, and says when it cut (#4471) #4474]: get_ag_health caps its fleet-wide answer, most severe first, and says when it cut (#4471) #4474