Skip to content

Updates_20261001 - The Nightmare Before Compaction - #899

Merged
erikdarlingdata merged 72 commits into
mainfrom
dev
Sep 29, 2026
Merged

erikdarlingdata merged 72 commits into
mainfrom
dev

Conversation

@erikdarlingdata

Copy link
Copy Markdown
Owner

Updates_20261001: The Nightmare Before Compaction.

sp_QueryStoreCleanup got most of the new features this month. A long removal can now run in two sessions, the procedure can compact Query Store's internal tables afterwards, and report mode returns a summary. Most of the other procedures got bug fixes from a round of reviews.

Changes in behavior

sp_QueryStoreCleanup

This procedure is not in the combined installer. Install it from its own folder.

sp_QuickieStore

sp_QueryReproBuilder

sp_QuickieCache

sp_HealthParser

sp_HumanEvents

sp_PerfCheck

sp_PressureDetector

  • The wait stats results show max_wait_time_ms. The procedure adds the column to an existing _Waits log table.

Docs and CI


All eleven procedures move to .10, version date 20261001. That includes sp_QueryStoreCleanup, which moves from 1.6 to 1.10. (#898)

🤖 Generated with Claude Code

https://claude.ai/code/session_01Dr8NyHme4WMWdwyz99UxWq

erikdarlingdata and others added 30 commits September 3, 2026 13:16
Move darlingdata-release and review-sql out of the personal
~/.claude/skills directory into .claude/skills so they travel with the
repo to other machines and to cloud sessions, which do not read the
local personal skills directory.

.gitignore now un-ignores .claude/skills only. settings.local.json,
commands/, scripts/, worktrees/, and the findings notes stay ignored.

darlingdata-release searched under an absolute C:\GitHub\DarlingData
path; it now searches from the repository root.

review-sql reads the T-SQL style guide from CLAUDE.md, which is tracked
in this repo, so the skill resolves the guide on any clone.

Neither skill touches Install-All/DarlingData.sql, which stays CI-owned.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EbYty5EYQwd4YGTDcJfEkE
Add project skills for release prep and T-SQL style review
Surfaces sys.dm_os_wait_stats.max_wait_time_ms between avg_ms_per_wait
and percent_signal_waits in both the cumulative and @sample_seconds
result sets, and in the _Waits logging table.

In sample mode the value only shows when a new high-water mark was
recorded during the window, since a cumulative max cannot be
delta'd (same reasoning as the window-local percent_signal_waits).

Existing _Waits log tables from older versions get the column added
via a conditional ALTER so logging keeps working after upgrade.

Tested on sql2016, sql2017, sql2022, and sql2025: install, cumulative
and sample runs, fresh log table creation, old-schema upgrade path,
and the full 10-proc harness (40/40 install, 40/40 execute). sql2019
was unavailable.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013RUm9uRhhfQ6aKaTJd3t7f
- actions/checkout v4 -> v7 in sql-tests, dev-installall-artifact,
  and build-sqlfile (persist-credentials default is unchanged, so the
  app-token commit-and-push in build-sqlfile still works)
- actions/create-github-app-token v2 -> v3 in build-sqlfile: the v3
  breaking changes are Node 24 (needs a current runner; ubuntu-latest
  is) and proxy env handling (not used here); the app-id/private-key
  inputs are unchanged
- claude.yml / claude-code-review.yml deliberately left at checkout@v6:
  a claude-*.yml edit on dev breaks the review gate until main catches
  up, so those bump as a coordinated dev+main pair at release time

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UbaAfAZPVqXbZVPy79AL81
Quarterly maintenance 2026-09: CI action pin bumps
…ts in sp_QueryStoreCleanup

The removal filter protected queries with a forced plan from deletion
but not queries with a forced Query Store hint (sys.query_store_query_hints,
SQL Server 2022+), so cleanup could delete a query and its DBA-forced hint
along with it. Added the same NOT EXISTS protection everywhere is_forced_plan
was already checked (Steps 2/3 dedupe and the shared @removal_filters
builder), gated on the catalog view's existence so it's a no-op pre-2022.

Also added HASH JOIN to OPTION(RECOMPILE) on both #removals inserts,
matching the hint sp_QuickieStore already uses for the same kind of join
against these catalog views, whose cardinality estimates are unreliable.

Verified against real SQL2022 (hint forced via sp_query_store_set_hints,
confirmed excluded from removal; forced-plan protection still holds
alongside it) and SQL2019 (confirmed the generated dynamic SQL never
references sys.query_store_query_hints there).

Fixes #865, Fixes #866

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0138XqNy4wNJhgzQgQjXFoKV
…hints-and-hash-join

Protect forced Query Store hints and add HASH JOIN to #removals inserts
…on in sp_HealthParser

The blocking-report XML extraction chain expected an <event> wrapper that
no longer existed after an earlier .query() call had already narrowed the
XML down to <blocked-process-report>, so #blocked/#blocking could return
zero rows; the activity column's own exist() check was separately dead
for the same reason. waits_by_duration (#td) grouped by its raw metric
columns instead of aggregating like its sibling #tc, and a downstream
ROW_NUMBER dedupe with no time bucket in its partition key silently
dropped or duplicated rows across buckets.

Closes #868

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0138XqNy4wNJhgzQgQjXFoKV
…_QueryStoreCleanup

@min_age_days accepted negative or zero values, which pushed @age_cutoff
into the future and silently disabled age-based removal protection
instead of erroring, unlike every other parameter in this proc. Also
drops a redundant @ToTal variable that never diverged from
@removal_count, and fixes five IF/AND alignment inconsistencies against
the style already used elsewhere in the same file.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0138XqNy4wNJhgzQgQjXFoKV
…g-and-wait-duration-bugs

Fix broken blocking-report extraction and waits_by_duration aggregation in sp_HealthParser
…review-findings

Add @min_age_days validation and clean up minor review findings in sp_QueryStoreCleanup
@cleanup = 1 was falling through to the ELSE branch, creating and
starting the throwaway XE session, then immediately GOTOing to
cleanup to tear it back down. Skip creation whenever @cleanup = 1,
regardless of @keep_alive, matching the RAISERROR comment that was
already there describing the intended behavior.

@session_id_filter and @username_filter were being appended to
@session_filter_blocking, but sqlserver.session_id/username are
global fields that reflect the system session in a blocked_process_report
(is_system = 1) context, not the actual blocked/blocking sessions.
Both filters were silent no-ops for blocking captures, so they're
removed from that filter concatenation.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0138XqNy4wNJhgzQgQjXFoKV
…-and-blocking-filter

Fix sp_HumanEvents cleanup-ordering and blocking-filter no-op bugs
…entsBlockViewer

- Add the missing comma between blocked_process_report_xml and the
  PRIMARY KEY CLUSTERED constraint in the @log_to_table = 1 dynamic
  CREATE TABLE. Without it, SQL Server silently drops the primary key
  and creates a heap with zero indexes instead of throwing a syntax
  error (confirmed by running the CREATE TABLE directly).

- In the system_health (#blocking_xml_sh) blocked/blocking extraction,
  replace the two-step OUTER APPLY chain
  (.nodes('/event') then .nodes('//blocked-process-report/...'))
  with a single OUTER APPLY straight against the real document root,
  matching the fix already applied for the same bug class in
  sp_HealthParser (PR #869). Also replace the now-guaranteed-true
  bd.exist()/bg.exist() activity CASE with the plain literal
  'blocked'/'blocking', same as sp_HealthParser.

- Add the missing ELSE to the activity CASE in the #blocked/#blocking
  (non-system-health) construction so activity is never NULL when the
  EXISTS check is false.

- Remove the redundant hardcoded ' > ' in the blocking_tree ELSE
  branch so 'blocked' rows don't get one extra arrow versus 'blocking'
  rows at the same blocking_level.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0138XqNy4wNJhgzQgQjXFoKV
Several joins and EXISTS checks compared plan_id/query_id/replica_group_id
alone against temp tables that accumulate rows from every database when
@get_all_databases = 1. Since these IDs restart from 1 in each database's
Query Store, a value from one database can collide with an unrelated row
from a different database and silently corrupt results (wrong wait totals
in regression aggregates, wrong has_query_feedback/has_query_store_hints/
has_plan_variants flags, wrong tuning recommendation/hint/variant/feedback
rows, wrong AG replica/plan-forcing-location matches). Each site now also
compares database_id, either against the sibling table's own database_id
column or against the current per-database loop variable when the other
side is a raw catalog-view alias scoped to one database already.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0138XqNy4wNJhgzQgQjXFoKV
@new is set once, early in the procedure, from the SQL Server version
(2017+ or Azure), and an existing guard already forces @sort_order back to
'cpu' (and @sort_order_is_a_wait to 0) whenever @new = 0 and the requested
sort order was log/tempdb/total log/total tempdb/wait-based. Since @new
never resets to 0 afterward and @sort_order/@sort_order_is_a_wait are never
reassigned again, every later
'CASE WHEN @new = 1 THEN X ELSE Y END' wrapper around those same sort
orders always evaluates to X by the time it runs; the ELSE branch is dead.
Collapsed each of these down to just the THEN expression, including one
spot where an inner CASE's two branches were already identical text.

Also removed a redundant .query('.') call immediately before .nodes('x')
in the @string_split_ints/@string_split_strings templates: querying the
whole XML document and then shredding it is the same as shredding it
directly.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0138XqNy4wNJhgzQgQjXFoKV
…is safe

Confirmed via live blocked_process_report captures (independent unrelated
pairs and a genuine A-blocks-B-blocks-C chain) that a single report always
has exactly one blocked-process element, so the two independent APPLYs
against #blocked/#blocking never produce a spurious cross join. Only
blocking-process repeats (one victim, multiple simultaneous blockers),
which is a safe 1xN fan-out.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0138XqNy4wNJhgzQgQjXFoKV
…ewer-review-findings

Fix sp_HumanEventsBlockViewer PK loss, activity NULL, and display bugs
…atabase-joins

Fix sp_QuickieStore cross-database join/EXISTS bugs and dead code
- UNC-path drive_location branch required 4 leading backslashes
  instead of 2, so it never matched a real UNC path
  (\server\share\...) and always fell through to the ELSE branch.
  Fixing that WHEN also exposed a second, previously dead-code bug
  one line below: the SUBSTRING/CHARINDEX that builds the "UNC: "
  label searched for two consecutive backslashes instead of one,
  which no realistic UNC path contains after the leading \, so it
  threw "Invalid length parameter" once the WHEN above started
  matching. Fixed both so UNC-hosted database files report a real
  per-server drive_location instead of erroring or collapsing to \.

- Query Store Suboptimal Configuration (check 7012) reported the
  wrong reason for sizes between 1000-1023 MB: its details CASE
  checked max_storage_size_mb < 1024 while the governing WHERE used
  < 1000, so a row selected by the WHERE for some other reason (e.g.
  capture mode NONE) in that gap wrongly got the size-related message
  instead of the real, more serious problem. Aligned the CASE
  threshold with the WHERE.

- Extremely Large Auto-Growth Setting (check 7104) had no type_desc
  filter in its WHERE, so a FILESTREAM (or other non-ROWS/LOG) file
  with a qualifying growth setting was selected but its details CASE
  only has ROWS/LOG branches, so the concatenated details silently
  came back NULL. Added the same ROWS/LOG filter its sibling checks
  (7101, 7102) already use.


Claude-Session: https://claude.ai/code/session_0138XqNy4wNJhgzQgQjXFoKV

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
… a removal

sp_query_store_remove_query serializes on an exclusive lock, so a large
removal runs one query at a time. Two sessions in the same order gain
nothing: they keep trying the same query_id, and the loser still waits
for the lock before failing.

@sort_direction (ASC or DESC, default ASC) sets the removal order by
query_id, so two sessions can work one list from opposite ends. Each
removal now checks that the query still exists first and counts it as
skipped if not, so neither session retries the other's half after they
meet. The finish message reports the skipped count.

Tested on a Query Store with about 800,000 queries: one session removed
14.6 queries a second, two sessions from opposite ends 20.2, and two in
the same order 14.7. Two procedure runs, ASC and DESC at once, removed
all 3,338 targets with one failure, at the crossing.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…p-sort-direction

sp_QueryStoreCleanup: add @sort_direction so two sessions can split a removal
…dge case in sp_QueryReproBuilder (#877)

* Fix wait-stats aggregation, SET DATEFORMAT mapping, and SET-options edge case in sp_QueryReproBuilder

The #query_store_wait_stats population had the same bug sp_QuickieStore
was recently fixed for: a TOP (5) ORDER BY avg_query_wait_time_ms DESC
inside the CROSS APPLY dropped wait categories ranked 6th+ per interval
before the outer GROUP BY ran, and avg/min/max_query_wait_time_ms used
SUM() instead of AVG()/MIN()/MAX(). Mirrors sp_QuickieStore's fix exactly.

The SET DATEFORMAT mapping from sys.query_context_settings.date_format
was off by one (coded as 0-5, real values are 1-6 per empirical testing),
so every generated repro script emitted the wrong SET DATEFORMAT line.

If a captured connection had all 7 tracked SET options off, context_settings
was an empty string (not NULL), which ISNULL didn't catch, producing the
literal invalid line "SET ON;". Guarded with an explicit empty-string check.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0138XqNy4wNJhgzQgQjXFoKV

* Fix NULL-safety regression in sp_QueryReproBuilder SET-options CASE

The empty-string fix for qsrs.context_settings replaced an ISNULL wrapper
with a CASE that only checked for ''. When context_settings is actually
NULL (every @query_plan_xml-driven case, since there's no real Query
Store connection behind a synthetic plan), the CASE fell through to the
ELSE branch, REPLACE(NULL, ...) produced NULL, and that NULL poisoned the
entire generated repro script through string concatenation. Caught by CI
failing all 4 SQL Server checks on the prior commit. Add back the NULL
check alongside the empty-string check.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0138XqNy4wNJhgzQgQjXFoKV

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
…QuickieCache (#876)

PERCENT_RANK() for cpu_pctl/duration_pctl/reads_pctl/writes_pctl/grant_pctl/
spills_pctl/executions_pctl was computed after the WHERE EXISTS filter down
to #candidates, so every query was ranked only against the other candidates
instead of the full workload the adjacent comment says it measures. Moved
the ranking into a CTE over the unfiltered #query_stats, then filter to
#candidates afterward.

@find_duplicate_plans never applied @start_date/@end_date, unlike every
other mode in this proc (including the sibling @find_single_use_plans fix
already landed for the same gap). Added the same filter.


Claude-Session: https://claude.ai/code/session_0138XqNy4wNJhgzQgQjXFoKV

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
A start date after the end date reached xp_readerrorlog unchanged, which
silently returns zero rows instead of erroring. That collapsed the archive
filter to one log and produced a clean-looking, error-free empty result -
exactly the "looks healthy but never actually searched" failure mode this
proc's own comments guard against everywhere else. Every other malformed
parameter combination in this proc already gets an explicit RAISERROR;
this one didn't.


Claude-Session: https://claude.ai/code/session_0138XqNy4wNJhgzQgQjXFoKV

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
…#879)

Statement-type rows built object_name via QUOTENAME on both the schema
and object parts ([dbo].[Proc]). The Procedure/Function/Trigger path a
few hundred lines down built the same column with plain concatenation
(dbo.Proc), for the same underlying object. Wrapped both parts in
QUOTENAME there too so the column is formatted the same way regardless
of which query_type produced the row.


Claude-Session: https://claude.ai/code/session_0138XqNy4wNJhgzQgQjXFoKV

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
query_hash is computed from normalized query text alone and is not
database-qualified. Grouping by query_hash only meant the same query
text running in two unrelated databases got merged into one row: an
arbitrary database_name (MAX(dbid), picked for no meaningful reason),
with plan_count/cpu/cache-size summed across both databases as if they
were one workload. It also produced false positives: 1 plan in DB A
plus 1 plan in DB B (neither a real duplicate) summed to plan_count=2
and passed the "> 1" duplicate-plans filter.

Group by (query_hash, database) instead, so each row is one database's
actual duplicate-plan situation. Verified live: the same query text run
in two scratch databases now produces two correctly-labeled rows
instead of one blended one.

Design call confirmed with Erik: split by database rather than listing
all involved databases in one row or leaving the existing behavior as-is.


Claude-Session: https://claude.ai/code/session_0138XqNy4wNJhgzQgQjXFoKV

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
…rd (#881)

* Fix HumanEvents_Queries_np view name and add a blocking-CTE cycle guard

@skip_plans = 1 inserted the no-plans Queries view into #view_check under
the wrong literal (HumanEvents_Queries instead of _np), so the view loop's
schema-qualify REPLACE left a stray _np outside the brackets and view
creation failed with Msg 102. Also ports BlockViewer's recursive-CTE cycle
guard (c01543b) into the live blocking hierarchy CTE, which had no
MAXRECURSION or cycle check, plus two ELSE-arm parity fixes, a help-text
row for the new view, and a regression test for the view-name bug.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0138XqNy4wNJhgzQgQjXFoKV

* Correct two CTE comments ported from sp_HumanEventsBlockViewer

sp_HumanEvents never had MAXRECURSION 0, so the "reverted from
MAXRECURSION 0" note described BlockViewer's history, not this file's.
The old OPTION(RECOMPILE) already used the default limit of 100, which
is why the unguarded cycle failed with Msg 530. Say that instead.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0138XqNy4wNJhgzQgQjXFoKV

* Plain-English pass on the new tests README text

Split the long sentences, drop the double-hyphen dashes and the
semicolon, and give the skip_plans regression case its own paragraph.
No facts changed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0138XqNy4wNJhgzQgQjXFoKV

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
…uery

@report_only = 1 used to return one row per query_id, which for a big
cleanup is hundreds of thousands of rows and no counts. It now returns:

1. One summary row: queries, distinct query hashes, query texts, plans
   and plan hashes, queries in modules, oldest and newest last execution,
   and the share of all Query Store queries.
2. One row per query_hash, biggest first, with the same counts, the
   module name, and a sample query_id with the first 200 characters of
   its text.

Each Query Store view is staged in its own temp table before anything is
joined. @debug = 1 still lists every query_id.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The LIKE over query_sql_text now compares
UPPER(query_sql_text) COLLATE Latin1_General_100_BIN2 against upper-cased
patterns when the database collation ignores case, and plain BIN2 when it
does not, so matches are unchanged. On a 745K-query Query Store the custom
search went from 73 s to 9 s, and the default system and maintenance
search from about 140 s to 28 s, with identical matches.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
erikdarlingdata and others added 29 commits September 27, 2026 21:30
- sp_QuickieCache: widen object_name to nvarchar(517) so long
  [schema].[object] names fit (#879)
- sp_HealthParser: waits by count and by duration take MAX of the
  since-startup snapshots instead of SUM
- sp_HumanEventsBlockViewer: restore the ' > ' prefix on the
  "is blocked by" rows (#872)
- sp_HumanEventsBlockViewer, sp_HumanEvents: drop the unreachable ELSE
  labels on activity (#872, #881)
- sp_HumanEvents: say when blocking ignores @session_id and @username (#871)
- sp_QuickieStore, sp_QueryReproBuilder: correct the wait stats comment (#877)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dr8NyHme4WMWdwyz99UxWq
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dr8NyHme4WMWdwyz99UxWq
Fix the remaining review findings from September 24 to 26
Report mode printed 'Null value is eliminated by an aggregate' when a PSP
parent (NULL last_execution_time) was on the removal list. This commit
moves the #report_groups MIN/MAX into a pre-filtered derived table. The
summary SELECT at lines 1395-1396 still needs the same treatment.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dr8NyHme4WMWdwyz99UxWq
The summary SELECT still had bare MIN/MAX(rq.last_execution_time), which
prints "Null value is eliminated by an aggregate" when a PSP parent (NULL
last_execution_time) is on the removal list. This moves both aggregates
into scalar subqueries filtered to non-NULL rows, matching the plans =
subquery style, the same fix the groups insert already had.

The harness now also checks that the summary's oldest_last_execution and
newest_last_execution match MIN/MAX(last_execution_time) read straight from
sys.query_store_query, so a fix that silences the warning but changes the
values still fails.

Adds the CI step that runs this harness (39 checks against dev before this
fix, one warning failure only; 75 clean against the finished branch), and a
README documenting what it covers and the two fixture problems: a
monitoring tool's queries landing in Query Store, and Query Store reporting
READ_WRITE before it captures anything.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dr8NyHme4WMWdwyz99UxWq
The polling in query_store_on() only confirms the ALTER DATABASE has taken
effect. The actual fix for a plan compiled during that async gap is the
per-attempt plan cache clear in dupes_workload() and the PSP workload, so
document both.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dr8NyHme4WMWdwyz99UxWq
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dr8NyHme4WMWdwyz99UxWq
…arness

sp_QueryStoreCleanup: assertion harness, and #883 NULL warning fix
…y compat level

sp_QuickieStore and sp_QueryReproBuilder:
- Wait stats come from every runtime stats row the other metrics use,
  not just each plan's latest interval. The min_query_wait_time_ms > 0
  filter is replaced by total_query_wait_time_ms > 0, because Query
  Store can report a minimum of 0 for a category with wait time.
- #query_store_query and #query_store_query_text read distinct IDs,
  so a query with several plans (or rows) is not repeated.

sp_QuickieStore regression mode:
- Outer parentheses around the ORed time periods in the runtime insert
  and both wait sorts. Without them the current period lost the plan
  filter and counts and totals were multiplied.
- LAST_VALUE windows are partitioned by time period.
- Wait stats carry the time period; top_waits and the expert wait sets
  match on it, and the by-query set shows it (not when logging).
- #query_store_plan reads distinct plan IDs.

sp_QueryStoreCleanup: Step 4 builds one single-EXISTS list per dedupe
strategy, joined with UNION ALL, so OPTION(RECOMPILE, HASH JOIN)
compiles at every compat level. The compat level check is gone.

Tests: a regression and wait window group for the QuickieStore harness
(209), a waitwindow case for the ReproBuilder harness (275), and the
QueryStoreCleanup compat test at every supported level (100 on 2025).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dr8NyHme4WMWdwyz99UxWq
…h-join

Wait stats over the whole window, regression mode fixes, HASH JOIN at every compat level
The totals come from the CI run on #890.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dr8NyHme4WMWdwyz99UxWq
The waiter ran its query once any X key lock was granted in the
fixture database. Another session's lock (Query Store's own writes,
for example) could satisfy that before the holder locked its row, and
the query then ran without waiting. SQL Server 2025 CI hit this on
#891: "slot A: rm_q1 waited on a lock (lock wait added=0 ms)".

The check now names the fixture table. The dance also retries up to
three times, and a holder that exits nonzero counts as an error.

Reproduced locally with a decoy X key lock and a holder that starts
1 second late: the old check recorded 0 ms, the new one 2,083 ms.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dr8NyHme4WMWdwyz99UxWq
Fix a race in the lock-wait test fixture; per-version totals in the QSC README
In regression mode each plan has a wait row per period, but the
WaitStatsByQuery log table had no period column, so the logged rows
could not be told apart. The table now has
from_regression_baseline_time_period, the same column the RuntimeStats
log table already has, and the regression mode insert fills it.

Tables created by earlier versions get the column added first, the same
upgrade path sp_PressureDetector uses for its _Waits table.

Harness: a RegressionLog group logs regression mode twice, into new
tables and into an old-shape WaitStatsByQuery table, and checks each
logged wait row's period and rm_q1's Lock wait per period against Query
Store. 215 checks.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dr8NyHme4WMWdwyz99UxWq
Label logged regression-mode wait rows with their period
With a text filter, Steps 2 and 3 find duplicate hashes among the
queries that match it, but Step 4 removed every query with those
hashes, including queries the filter never matched. A custom filter on
one marker could remove an unmarked copy of the same statement.

Step 4's dedupe branch now keeps only queries whose text matches the
filter. With @cleanup_targets = 'none' there is no text filter, and
every copy of a duplicated hash is still removed.

Harness: the duplicate fixture gets C1T, C1's statement under its own
marker, so it shares C1's query_hash and plan hash but never matches
the custom filter. The existing custom-filter counts and the DESC
removal now fail if it is swept in, and three new checks compare the
debug removal list with the expected groups. The dev copy fails 12
checks; this one passes 103 on SQL Server 2025.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dr8NyHme4WMWdwyz99UxWq
The expert mode wait stats total result set summed every column across
plans, so avg, min, max, and last were sums. In regression mode it also
merged the two time periods into one row.

Per wait category, and per period in regression mode, the columns are now:
- total: the sum across plans (unchanged)
- avg: the average of the plan averages, as elsewhere in the procedure
- min and max: the lowest and highest plan values
- last: the value from the plan with the most recent wait in the category
  (latest interval, then latest execution, then plan_id), following
  5c7d979, which made last_ columns report the actual lasts

WaitStatsTotal log tables get a from_regression_baseline_time_period
column, with the same upgrade path as WaitStatsByQuery.

The harness adds rm_q3, a second query that waits on the same lock, after
rm_q1 in slot A and before it in slot B. New checks compare the totals with
the by-query rows and with Query Store, in both runs and in the log tables.
227 assertions pass on SQL Server 2025. The dev procedure fails 8 of the new
checks, and a MAX-of-lasts variant fails the 2 last-value checks.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dr8NyHme4WMWdwyz99UxWq
Keep sp_QueryStoreCleanup dedupe removals inside the text filter
The review of #894 read the plan and query joins in the totals CROSS APPLY as dead weight. They keep the totals over the same rows that wait stats by query shows, which needs them for object_name.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dr8NyHme4WMWdwyz99UxWq
Make sp_QuickieStore's wait stats total report real avg, min, max, and last values
…nd run it in CI

A pre-release run of every sp_QuickieCache mode on SQL Server 2016 to 2025 passed, but NULL parameters misbehaved: @top = NULL failed with Msg 1014, and a NULL @impact_threshold or @minimum_execution_count returned nothing. NULL parameters now take their defaults, as in sp_QuickieStore.

@database_name = N'master' returned nothing, because @ignore_system_databases = 1 by default. A system database named in @database_name is now searched.

CI installed sp_QuickieCache through the bundle but never checked or ran it. The existence check, the help test, and the smoke test now include it, with its modes and NULL parameters.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dr8NyHme4WMWdwyz99UxWq
…nd-ci

sp_QuickieCache: NULL parameters take their defaults, a named system database is searched, and CI runs it
The root parameter tables lacked nine sp_PerfCheck thresholds,
sp_QueryStoreCleanup's @sort_direction and @compact_tables, and
sp_QuickieStore's @find_parameter_sensitive. Each new row copies the
procedure's own README. The @slow_read_ms and @slow_write_ms rows use
the help text's "(High at 5x this value)" in both READMEs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dr8NyHme4WMWdwyz99UxWq
…-params

Root README: list the 12 missing parameters
The nine threshold parameters keep 0 and default only NULL and
negative values (the >= 0 check), but @help said "any positive".

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dr8NyHme4WMWdwyz99UxWq
…thresholds

sp_PerfCheck: help says the thresholds accept 0
The ten combined-install procs move to .10, and sp_QueryStoreCleanup
moves from 1.6 to 1.10 because it changed this cycle. Version date
20261001 for all eleven.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dr8NyHme4WMWdwyz99UxWq
Release style audit: the comment on @find_duplicate_plans' date filter
in sp_QuickieCache, and the APPLY comment at both sp_HumanEventsBlockViewer
sites, started their text on the /* line. The text is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dr8NyHme4WMWdwyz99UxWq
Updates_20261001 release prep: version bumps
@erikdarlingdata
erikdarlingdata merged commit 25670ab into main Sep 29, 2026
10 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant