Skip to content

Store cumulative VACUUM statistics in the collector and expose them through vacuum_stats extension - #2079

Open
Alena0704 wants to merge 18 commits into
apache:REL_2_STABLEfrom
Alena0704:vacuum_reporting_rel2
Open

Alena0704 wants to merge 18 commits into
apache:REL_2_STABLEfrom
Alena0704:vacuum_reporting_rel2

Conversation

@Alena0704

Copy link
Copy Markdown
Collaborator

What does this PR do?

Make VACUUM work available to SQL so administrators can compare maintenance
results across runs, relations and segments. Counter differences help find
relations that consume most vacuum time, retain dead tuples, repeatedly
lose visibility-map markings, or need aggressive maintenance.

This PR#2078 includes the instrumentation branch
through a dedicated merge commit. The six following commits add reporting
and the extension. Review only this PR's additions.
When targeting REL_2_STABLE, the full diff also contains the first PR
until its commits are merged into that branch.

  • Accumulate tuple removal, remaining dead tuples, page work and freeze-age
    counts per relation and database. Keep index work out of database totals
    where the owning table already includes it.
  • Report AO row/column compaction, exact truncated bytes, remaining hidden
    tuples and the latest segment-file metadata count, including index work.
  • Count completed heap vacuums that entered wraparound failsafe and heap
    vacuums interrupted by ERROR, including cancellation.
  • Add the vacuum_stats extension with local pg_stat_vacuum_* views and
    coordinator/segment gp_stat_vacuum_* views identified by gp_segment_id.
    Expose work counters alongside timing and visibility-map statistics.
  • Add relation and database vacuum-statistics reset functions that preserve
    ordinary access counters, maintenance counts/timestamps and ANALYZE times.

Type of Change

  • New feature

Impact

Extended work counters require track_vacuum_statistics = on on the
coordinator and segments, followed by a restart. The setting is off by
default; disabled tracking omits storage for those counters. Timing,
visibility-map clearings, failsafe and interruption counts follow
track_counts independently of this setting.

Install SQL access with CREATE EXTENSION vacuum_stats. All new functions
and views belong to the extension; the system catalog and its version are
unchanged. The saved statistics format changes, so older saved statistics
are discarded on startup.

Most values are cumulative observations, not current relation contents.
total_file_segs is the latest AO snapshot. Summed segment time is accumulated
work, not the wall-clock duration of a distributed VACUUM.

Checklist

@Alena0704 Alena0704 added the type: Enhancement New feature or request, ideas label Oct 5, 2026
@Alena0704
Alena0704 requested a review from leborchuk October 5, 2026 16:20
nathan-bossart and others added 18 commits October 6, 2026 17:06
This function is used in both vacuum and analyze code paths, and a
follow-up commit will require distinguishing between the two.  This
commit forces callers to specify whether they are in a vacuum or
analyze path, but it does not use that information for anything
yet.

Author: Nathan Bossart <nathandbossart@gmail.com>
Co-authored-by: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
Discussion: https://postgr.es/m/ZmaXmWDL829fzAVX%40ip-10-97-1-34.eu-west-3.compute.internal

Backported from PostgreSQL 18 (commit e5b0b0c)
as a prerequisite of the vacuum statistics series.  Cloudberry adaptation:
the append-optimized and AOCS sample-row acquisition and compaction call
vacuum_delay_point() too; the former pass is_analyze = true, the latter
false.

(cherry picked from commit e07fe98)
(cherry picked from commit ffe16da)
Introduce the process-local VacuumDelayTime accumulator and measure time
spent in cost-based VACUUM delays, in microseconds. Gate the measurement
with track_cost_delay_timing, disabled by default. Consumers take
differences around their operations to attribute the measured delay.

Use is_analyze to exclude ANALYZE sampling and statistics computation from
the vacuum delay accumulator. Cost-based throttling and interrupt checks
continue to apply to both operations.

Port the delay measurement from PostgreSQL 18 without its progress-view
changes, so this commit needs no catalog update. Omit the progress slots,
progress-increment helper and parallel progress reporting. The accumulator
is process-local and Cloudberry does not enable parallel vacuum. Register
the GUC in the unsynchronized list.

Author: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
Co-authored-by: Nathan Bossart <nathandbossart@gmail.com>
Co-authored-by: Alena Rybakina <alenka.rybakina@gmail.com>
Discussion: https://postgr.es/m/ZmaXmWDL829fzAVX%40ip-10-97-1-34.eu-west-3.compute.internal

Backported from PostgreSQL 18 (commit bb8dff9).
(cherry picked from commit 79c8bbc)
Collect visible_page_marks_cleared and frozen_page_marks_cleared.
Count transitions of heap visibility-map bits from
set to clear during normal backend activity. Further modifications of
a page whose bit is already clear do not increment that counter. Count
clearings even if the modifying transaction subsequently rolls back.

Comparing the all-visible clearing rate with DML volume helps identify
modifications spread over pages vacuum previously marked all-visible.
Index-only scans must visit the heap for those pages until visibility
is re-established. The all-frozen counter shows how often pages lose
the property that allows aggressive vacuum to skip them. Clearing this
bit does not unfreeze tuples that are already frozen.

Adapt the v44 VM stability patch to Cloudberry's statistics collector.
Deliver the counters with ordinary table-statistics messages and
accumulate both relation and database totals. Collection follows
track_counts. These are DML-driven visibility-map transitions rather than
work performed by VACUUM. SQL access and TAP coverage follow in the
separate vacuum_stats extension commit; this commit does not change the
system catalog.

Count the cleared bits in visibilitymap_clear(), which already takes
a Relation on this branch. Skip relations without a statistics entry,
including fake relations used during recovery.

Authors: Alena Rybakina <lena.ribackina@yandex.ru>,
         Andrei Lepikhov <lepihov@gmail.com> (@danolivo),
         Andrei Zubkov <zubkov@moonset.ru> (@zubkov-andrei)
Based-on: https://www.postgresql.org/message-id/attachment/204710/v44-0009-Track-table-VM-stability.patch

Co-authored-by: Andrei Lepikhov <lepihov@gmail.com>
Co-authored-by: Andrei Zubkov <zubkov@moonset.ru>
Assemble backend-local tuple and page measurements and take per-call
snapshots of the existing index results. Access methods accumulate results
across passes, so retain each pass's delta without counting earlier work
again.

Keep SP-GiST's newly deleted page count cumulative within one vacuum,
without counting pages that were already empty on entry. Assigning all
deleted pages to pages_newly_deleted would count empty pages again on
subsequent scans.
Introduce dead_pages in LVRelState and PgStat_VacuumStats. Increment it
for each heap page that retains dead tuples after pruning. For index
cleanup, derive the count from deleted pages not yet reusable.

Collector delivery and SQL exposure follow in the reporting commit.

Include these per-operation measurements in VACUUM VERBOSE output.
Introduce pages_frozen in LVRelState and PgStat_VacuumStats. Increment
the counter when vacuum executes freezing work on a heap page and
include it in the backend-local measurements. Collector reporting
and SQL exposure follow separately.

Include these per-operation measurements in VACUUM VERBOSE output.
Introduce pages_all_visible in LVRelState and PgStat_VacuumStats.
Count visibility-map updates for empty pages, scanned pages and pages
revisited during heap cleanup, testing the all-visible bit when needed.

Collector delivery and SQL exposure follow in the reporting commit.

Include these per-operation measurements in VACUUM VERBOSE output.

Related to vm_new_visible_pages in v40-0006, "Extended vacuum statistics:
visibility-map page transitions for tables". That patch reports existing
PostgreSQL VM transition counters; here the all-visible measurement is
added to Cloudberry's PostgreSQL 14 vacuum paths. The patch's separate
all-frozen and combined visible/frozen counters are not included here.

Related-to: https://www.postgresql.org/message-id/attachment/199749/v40-0006-Extended-vacuum-statistics-visibility-map-page-trans.patch
Introduce freeze_age_vacuum_count and preserve the freeze-age
decision before DISABLE_PAGE_SKIPPING can make the scan aggressive.
Record one for a run driven by relfrozenxid or relminmxid, and zero
when aggressive scanning is forced only by DISABLE_PAGE_SKIPPING.
Include VACUUM FREEZE, which sets the freeze-age thresholds to zero.

This introduces the measurement; collector accumulation and SQL
exposure follow in the reporting commit.

Include these per-operation measurements in VACUUM VERBOSE output.
Preserve remaining dead tuples from post-cleanup and accumulate elapsed
and cost-delay time across AO vacuum phases. Measure index cleanup time
as well. Widen existing compaction tuple and byte accumulators to int64,
and keep the reclaimed-block conversion 64-bit as well.

Collector reporting and SQL-level AO tests follow in a separate commit.

Include these per-operation measurements in VACUUM VERBOSE output.

Measure cleanup-only index scans when there are no obsolete AO segments.
…alyze

This commit adds four fields to the statistics of relations, aggregating
the amount of time spent for each operation on a relation:
- total_vacuum_time, for manual vacuum.
- total_autovacuum_time, for vacuum done by the autovacuum daemon.
- total_analyze_time, for manual analyze.
- total_autoanalyze_time, for analyze done by the autovacuum daemon.

This gives users the option to derive the average time spent for these
operations with the help of the related "count" fields.

Bump PGSTAT_FILE_FORMAT_ID for the additions in PgStat_StatTabEntry.
SQL access follows in the separate vacuum_stats extension commit; the
system catalog is unchanged.

Author: Sami Imseih
Reviewed-by: Bertrand Drouvot, Michael Paquier
Discussion: https://postgr.es/m/CAA5RZ0uVOGBYmPEeGF2d1B_67tgNjKx_bKDuL+oUftuoz+=Y1g@mail.gmail.com

(cherry picked from commit 30a6ed0)

(cherry picked from commit 6a99f9d)
(cherry picked from commit 5f76c5e)

Adaptation for this branch: carry elapsed time in UDP VACUUM/ANALYZE
messages and initialize the new counters in both relation-entry creation
paths. Keep microsecond precision internally and expose milliseconds.
Time AO vacuum from the first phase executed by this worker and pass that recorded
start time when reporting the final phase.
The per-index autovacuum log line reports page counts but not the number
of index entries removed. Add tuples_removed accumulated by the index's
bulkdelete passes to that line:

  index "t_pkey": tuples: 500 removed; pages: 30 in total, ...

Port v44-0001 to Cloudberry. On this PG14 base the summary is used by
autovacuum; manual VACUUM VERBOSE already reports removed index row
versions from lazy_cleanup_one_index(). Preserve that existing output.
This uses the access method's existing counter and adds no new statistics
collection or dependency on track_vacuum_statistics.

Suggested by Bharath Rupireddy (@BRupireddy2).

Based-on: https://www.postgresql.org/message-id/attachment/204702/v44-0001-Report-per-index-removed-tuples-in-vacuum-instru.patch
Add cumulative elapsed and delay time for indexes and databases, and
delay time for tables. Keep manual VACUUM and autovacuum totals separate.
Database totals use table timings, which already include index work.

Adapt v44-0002 to Cloudberry's PostgreSQL 14 UDP collector, including AO
vacuum. Store times in microseconds and report delay in VACUUM VERBOSE
and autovacuum logs. Collection follows track_counts; delay measurement
also requires track_cost_delay_timing and excludes ANALYZE calls.

SQL access is provided separately through vacuum_stats. Bump the
statistics-file format for the new fields; the system catalog is unchanged.

Based-on: https://www.postgresql.org/message-id/attachment/204703/v44-0002-Track-vacuum-times-for-indexes-and-databases-and.patch
(cherry picked from commit 77a3ffa32d07d8a2947bf7e8dbb8d342ce05549d)

Co-authored-by: Andrei Lepikhov <lepihov@gmail.com>
Co-authored-by: Andrei Zubkov <zubkov@moonset.ru>
Port the collection part of v44-0003 to Cloudberry's PG14 statistics
collector. Add vacuum_failsafe_count to relation and database statistics.
SQL access and TAP coverage follow in the separate vacuum_stats extension
commit; built-in functions, system views and the catalog version are unchanged.

Failsafe disables cost-based delay and skips index vacuuming and heap
truncation to prioritize freezing old XIDs. Count each completed heap
vacuum that entered this mode, separately from vacuums made aggressive
by the freeze-age threshold. A failsafe event can help identify tables
whose regular vacuum schedule or freeze settings need attention.

Pass vacrel->failsafe_active through the regular UDP VACUUM report;
accumulate the relation and database counters together. Index-pass
reports cannot increment these totals. AO reports false because the AO
parent has no heap wraparound failsafe; its auxiliary heap relations
are accounted for by their own vacuums. Collection follows track_counts.

The existing failsafe warning identifies the affected run in the server
log. Bump the statistics-file format version for the accumulated counts.

Based-on: https://www.postgresql.org/message-id/attachment/204704/v44-0003-Count-wraparound-failsafe-vacuums-in-pg_stat-vie.patch
Co-authored-by: Andrei Lepikhov <lepihov@gmail.com>
Co-authored-by: Andrei Zubkov <zubkov@moonset.ru>
Add vacuum_interrupt_count to database statistics for ordinary heap
vacuums interrupted by ERROR while the per-relation vacuum error callback is
installed. This includes query cancellation. Such a run never reaches
the normal end-of-vacuum statistics report. Lower-severity messages
carrying vacuum context must not increment the counter.

Port v44-0004 to Cloudberry's PG14 collector. Keep the error callback
limited to incrementing process-local counters; geterrlevel() distinguishes
ERROR from other reports. Include the pending counts in ordinary TABSTAT
messages from pgstat_report_stat(), outside the transaction and error path.
Send shared-relation errors to the InvalidOid database entry and ensure
pending errors are flushed even when there are no table entries to send.

Collection follows track_counts and uses ordinary database statistics.
Preserve counts across clean restart and clear them with ordinary database
reset. Bump the statistics-file format version and adjust TABSTAT payload
capacity for its new field. SQL access and TAP coverage follow in the
separate vacuum_stats extension commit, without changing the catalog.

Document the scope: this does not count VACUUM FULL, errors before the
heap callback is installed, or AO parent compaction. Auxiliary heap
relations are counted by their own vacuums.

The existing ERROR report identifies the failed operation in the log;
count it without logging an additional message from the error callback.

Original patch reviewed by Vlada Pogozheskaya
<v.pogozheskaya@postgrespro.ru>.

Based-on: https://www.postgresql.org/message-id/attachment/204705/v44-0004-Count-vacuums-interrupted-by-errors-in-pg_stat_d.patch
Co-authored-by: Andrei Lepikhov <lepihov@gmail.com>
Co-authored-by: Andrei Zubkov <zubkov@moonset.ru>
…storage

Deliver backend measurements through VACSTATS messages and accumulate
relation and database counters in the statistics collector. Reuse the
existing tuple-removal and heap/index page results; page, VM, timing and
failsafe measurements were introduced separately. Keep index work out of
database totals where it is already included in table work, and report
per-pass index deltas without counting earlier passes again. Retain that
work even when a later cleanup is skipped.

Add track_vacuum_statistics as a postmaster setting, off by default.
Allocate extended vacuum counters with ordinary relation/database entries
only when enabled; omit their storage from collector and snapshot hashes
when disabled. Keep VM clearings, maintenance times, failsafe and error
counts in ordinary entries governed by track_counts. Tag saved statistics
so changing the startup setting preserves ordinary counters while disabling
tracking discards the extended fields. Configure each node before restart
instead of synchronizing the setting through dispatched sessions.

Tie the counters and statistics snapshots to their owning relation/database
entries and update the statistics-file format. AO reporting follows in a
separate commit using the same collector and startup setting.

SQL functions, views, documentation and TAP tests follow in the separate
vacuum_stats extension commit. The system catalog is unchanged.

Cloudberry adaptation of the PostgreSQL extended vacuum statistics
collector. Use the PG14 UDP collector instead of v44's report hook and
custom statistics kinds. The startup GUC and optional storage are
Cloudberry additions.

Based-on: https://www.postgresql.org/message-id/flat/cb305107-5935-4c34-9847-6ff0fef89f06%40yandex.ru
Co-authored-by: Andrei Lepikhov <lepihov@gmail.com>
Co-authored-by: Andrei Zubkov <zubkov@moonset.ru>
Report existing compaction tuple/byte counts and remaining dead tuples.
Connect AO index-pass elapsed and cost-delay time to the ordinary counters,
including cleanup without obsolete segments or AM page results. Include
obsolete index entries for relocated rows and exclude index work from
database totals. AO itself has no heap VM or wraparound failsafe.

Record exact bytes truncated from heap and AO table files in relation and
database statistics. Record the AO segment metadata count after the latest
vacuum as total_file_segs, replacing this state on each report. Update the
statistics-file format for the added fields. Include truncated bytes and
segment counts in AO VERBOSE output.

Build and send extended AO table/index reports only when the startup
track_vacuum_statistics setting is enabled. Keep ordinary index timing
outside that guard so it continues to follow track_counts, including
cleanup-only scans. VERBOSE measurements remain available with tracking off.

SQL access, documentation and tests for AO row and column tables follow
in the separate vacuum_stats extension commit. The system catalog is unchanged.
Add SQL getters and local pg_stat_vacuum_tables, pg_stat_vacuum_indexes and
pg_stat_vacuum_database views for the previously introduced statistics.
Read VM clearings, maintenance times, failsafe and interruption counts from
ordinary collector entries, and work counters from the optional vacuum
statistics storage. Keep all getters in the extension library and schema;
no built-in OIDs, system views or catalog-version changes are required.

Expose the timing backport's separate table ANALYZE and autoanalyze totals
alongside manual/autovacuum elapsed and delay times. VACUUM delay totals
exclude calls classified as ANALYZE. ANALYZE elapsed time has its own
counters; there is no separate ANALYZE delay counter. Test standalone
ANALYZE and VACUUM ANALYZE with cost-based throttling enabled.

Include datid zero for shared-relation statistics. Add coordinator
and segment wrappers and gp_stat_vacuum_* views with gp_segment_id. Suppress
the segment branch in utility mode so local rows appear exactly once.
Resolve extension getters explicitly in the extension's installation schema.

Document counter meanings and their use in evaluating vacuum effectiveness.
Add TAP coverage for heap and index work, AO row/column compaction, VM
transitions, manual and automatic timing, failsafe, canceled vacuums, startup
allocation and persistence. Adapt the v44 scenarios to the PG14 UDP
collector and wait for observable reports instead of using fixed sleeps.

Cloudberry adaptation of the PostgreSQL extended vacuum statistics SQL
interface. Read counters already collected by the backend, including
timing, failsafe and VM statistics, through the extension without changing
Cloudberry's system catalog.

Based-on: https://www.postgresql.org/message-id/flat/cb305107-5935-4c34-9847-6ff0fef89f06%40yandex.ru
Add extension functions that reset extended vacuum counters, VM clearings,
manual/autovacuum elapsed and cost-delay times, and vacuum_failsafe_count.
Preserve ordinary row, access and maintenance counters, timestamps and
ANALYZE times. Relation resets leave other relations and database totals
intact; database resets exclude shared catalogs. Provide segment dispatch
wrappers and restrict execution through SQL grants.

Distinguish database-wide resets explicitly in collector messages.
Reject OID zero in the relation overload and route shared relations and
indexes only to the shared collector entry. NULL relation arguments do
nothing. Reset ordinary vacuum fields even when extended tracking is off.

Cover reset boundaries, invalid and NULL OIDs, shared catalogs and
indexes, AO segment snapshots and database totals in TAP. Synchronize
with collector-report wait helpers.

Check extension installation in a non-default schema, unchanged system
statistics views and built-in functions, shared-relation rows, preservation
of ANALYZE times on vacuum reset, and collector statistics across extension
removal and reinstallation.
@Alena0704
Alena0704 force-pushed the vacuum_reporting_rel2 branch from 79f6c1c to e02222e Compare October 6, 2026 14:06
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

type: Enhancement New feature or request, ideas

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants