Repository navigation
[Proposal] Apache Cloudberry as PostgreSQL 19 Extensions #2065
Replies: 3 comments
|
Hi! Thank you for the attention to our project ) The article https://github.com/igor-suhorukov/cloudberry/blob/extension_postgresql_19/pg19/doc/greenplum-without-the-fork.md looks more broader and has more interesting details, if you don't mind I'll answer mainly to ideas from you original post rather than technical description here. Because some aspects need discussion at first and only then implementation.
Ease rebasement - what we constantly bear in mind. And the story we usually share with postgres community. You could get postgres, add here PAX - and now you have analytical data storage format. And the same with other new components. There are not so much of them comparing with original greenplum innovations. But now it's just the matter of time. There will be more and more components, I believe ) It will be great to evolve old approaches and make them extension-based. For example, as you described - make ao and aoco tables as extensions. We could start with some specific functionality and then step-by-step goes to the situation where you have postgresql fork extended with various MPP functions. Not based on postgres our own MPP database. The difference of course in the number of changes in a core part (and conflicts need close attention) and code reusing. The idealistic (I think it's too simple to be true) picture is you have a set of extensions and combining them could easily make your own specific database. How to achieve this - if you are really interested in it and ready to participate, let's continue discussion (good example is #1683 ), volunteer project developers and gradually improve our codebase. LLM will help us, but not replace, we still need to clearly understand what we are doing.
Kernel rebasing is great but as you have written, we rebase kernel not to the sole kernel version but to achieve something - new functionality or better performance. Good example is clickbench - new kernels could get you better performance. But I need to say modern PG kernel is not enough. PG is just (not fully describe current situation but it's too hard to express it succinctly) not good enough to beat clickhouse in clickbench. Not because clickbench purposely was written by clickhouse developers to beat all other competitors (but because of that, too), but mainly because clickhouse constantly compact (sort) data in background and have many others good improvements. We also could improve our group by facilities and be comparable in some aspects with clickhouse. It's possible, but not easy. if you interested - feel free to create discussion. I believe we could do it. Also we could create our own benchmark to beat all the competitors in it.
You mentioned other Postgres-related projects. We could not only compete with them but also get good approaches from them. |
|
Hi @igor-suhorukov, Thanks for this impressive experiment. It shows something many of us assumed was impossible: distributed snapshots, 2PC and shared reader/writer transactions running on a small set of dormant hooks, with ORCA compiled unmodified behind That said, I don't think the community should adopt the prototype as it is, for five reasons:
My suggestions:
|
|
@yjhjstz @leborchuk thank you for sharing your expert opinion! @yjhjstz thanks, point 2 is fair, so I took it as a test plan. Today the port runs 60 of the 262 tests in isolation2_schedule, plus 14 of the 15 parallel retrieve cursor tests. These are the ones on distributed transactions, snapshots, locks, GDD, FTS Among them, reader_waits_for_lock reaches XactAdoptTransactionState(), under ORCA. distributed_snapshot, distributedlog-bug and gdd/concurrent_update reach the hidden-commit path. None of Cloudberry's tests reaches a segment waiting on a prepared part. Cloudberry's coordinator PANICs rather than end a transaction with a part still prepared, so its tests never get Added after your comment:
Full run on that commit: isolation2 80/80 in both passes, greenplum_schedule 392/392 and 394/394. Still to do: the other 200 isolation2 tests (133 of them append-optimized) and latency and concurrency benchmarks (your point 3). |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Proposers
Igor Suhorukov (@igor-suhorukov)
Proposal Status
Under Discussion
Abstract
Updated 2026-10-08: what has been built on the port since the first version, a vectorized executor and an Arrow Flight SQL endpoint (Motivation, Implementation), and current figures: 24 core commits, 744 test files, the progress of
isolation2. The first comments refer to the version of 2026-09-30.I would like to discuss running Apache Cloudberry on PostgreSQL 19 as a set of extensions on top of 24 dormant core hooks, instead of a forked kernel. A working prototype exists; it is my personal experiment, not an ASF release. Without the modules, the patched server is measurably vanilla PostgreSQL 19: the same PostgreSQL tests pass, the ABI is unchanged, catalogs are byte-identical, and the hooks add less than 0.2% to executed instructions. With them, it is a Greenplum-style MPP engine with plan slices, Motions, two-phase commit and distributed snapshots, running unmodified ORCA and stock PostGIS and pgvector. ORCA plans all 121 TPC-H and TPC-DS queries at SF1 with answers matching DuckDB's, and 744 of Cloudberry's 1,180 test files pass so far.
The port gives users more than a newer PostgreSQL. With Cloudberry's components as modules, new capabilities arrive as further extensions. pg_vexec, built on the port without a single new core patch, adds vectorized execution that ORCA plans by cost, PAX and
ao_columntables read and written by column, Arrow frames between segments, and an Arrow Flight SQL endpoint for BI and ML clients.Claude Code wrote the code, the port over 12 days and 629 commits and pg_vexec over three days; I set the goals and signed off every change to the PostgreSQL core.
Motivation
Greenplum and Cloudberry have always followed PostgreSQL a few releases behind: Greenplum 6 shipped on PostgreSQL 9.4, Greenplum 7 on 12, Cloudberry 2.x on 14, and
mainreached 16.9 in May 2026. The reason is structural. Cloudberry edits 882 PostgreSQL backend and header files, so every new PostgreSQL major is a large merge. #1095 estimated the 14 → 16 upgrade at 4–5 months, three of them for fixing regression tests, and noted that PostgreSQL 14, the base of Cloudberry 2.x, reaches end of life in 2026. Meanwhile PostgreSQL 17 and 18 have shipped, and 19 is close to release.In #1877, the project named close compatibility with modern PostgreSQL as one of its goals. The prototype tests whether that goal can become cheap to sustain. If Cloudberry runs as extensions, the next PostgreSQL major means rebasing a small patch series instead of merging a fork. Users get a current PostgreSQL, including PostgreSQL 19 features such as eager aggregation and parallel autovacuum. They also get stock extensions: upstream PostGIS 3.7 instead of a PostGIS fork, and pgvector 0.8.6 built unpatched, which is relevant to #2043.
More than a newer PostgreSQL. A current base is the first gain, not the only one. As modules, Cloudberry's planner, storage and interconnect can be combined with other extensions, so new capabilities arrive as extensions rather than as more changes to a fork. Apache Cloudberry's code has no vectorized executor: only traces of a closed-source engine remain, a
create_vectorization_planflag the planner always passes as false, aWindowHashAggnode with no executor behind it, and a PAX adapter underVEC_BUILDthat no longer compiles. pg_vexec was built on top of the port in three days of implementation (write-up), and it brings Cloudberry users:CCostModelVec) and its translator builds them; PostgreSQL's planner gets them through its path hooks. No finished plan is rewritten, and a fallback to PostgreSQL's expression evaluator inside the same node keeps PostgreSQL's semantics.ao_columnhand batches to the executor and take them back on insert.count,min,max,sumandavgwithout conditions come from PAX's statistics:count(*)over 10 million rows takes 0.33 ms instead of 265.cdbhashthat reproducesGpHashSegmentbit for bit, and a newshminterconnect passes them through shared memory between the segments of one host.vexec_flightserves results as Arrow record batches straight from the executor, andadbc_ingest()loads DataFrames into tables, by column into PAX andao_column, with the standard ADBC Flight SQL and Flight SQL JDBC drivers. Logins are checked againstpg_hba.confand statements run through portals, so privileges, RLS andpg_stat_statementsbehave as they do over pgwire, which stays as it is.These are first measurements, and partial ones; every answer was checked against DuckDB's. On one node, on ClickBench at 10 million rows, vexec in auto mode takes 19–24% off PostgreSQL's planner and 16–19% off ORCA by geometric mean, and the best queries run 8.8–9.5× faster, though a few got slower; Flight SQL delivers a million
lineitemrows 4.7× faster than binary COPY and loads PAX 1.65× faster than COPY CSV. On four segments TPC-H Q1 runs 1.8–2.8× faster, but the first TPC run, with only scans and aggregates vectorized, averaged about 1×; the cluster's measurements with vectorized hash joins and Arrow frames through Motions are next.None of it needed a new core patch: vexec also runs on vanilla PostgreSQL 19, and
gp_orcanow builds for it withoutgp_core(-Dorca_single_node), so ORCA plans vector nodes on a stock server too. It relies on what PostgreSQL added after 16:heap_prepare_pagescan(17),explain_per_plan_hook(18), andplanner_setup_hook,planner_shutdown_hookand the batched visibility checkHeapTupleSatisfiesMVCCBatch(19), none of which Cloudberry'smainhas. In a fork, each such capability would be one more set of kernel changes to carry across every PostgreSQL major.The prototype compiles Cloudberry's own sources in place, so it builds on the community's work rather than replacing it.
Implementation
Details: Greenplum Without the Fork. Code: core series, port, build instructions.
The ground rule: a server with the core patches but without Cloudberry loaded must be indistinguishable from vanilla PostgreSQL 19, and this is measured.
Core series: 24 commits, 49 files, +1,622/−41 lines. Existing extension points come first: hooks, table AMs, CustomScan, custom WAL resource managers, background workers, security labels and FDWs. Where they are not enough, a new hook is a NULL-by-default function pointer, a new exported function, a registration list or an off-by-default flag. No existing signature changes, no struct gains a field, and no catalog, WAL, page or protocol format changes.
Vanilla checks compare Docker images built from the same PostgreSQL commit with and without the series:
meson testresults;abidiff: 34 symbols added, 0 changed;Cloudberry's tree is untouched. The whole port is one new directory,
pg19/, plus a release workflow in.github/workflows. Cloudberry's files compile where they lie, so merges from Cloudberry keep applying, and a file that must change becomes a copy that says what changed. ORCA's 920 core files compile unmodified behindplanner_hook, and PAX compiles in place with 57 of its 91 files unchanged.Modules:
gp_core: dispatcher, Motions, interconnect, 2PC, distributed snapshots, GDD, FTS;gp_orca;gp_sql: Cloudberry syntax rewritten into PostgreSQL 19's grammar;gp_ao,pax;gp_exttable,pxf_fdw,gpcloud,datalake_fdw;gp_resource,gp_security,gp_matview,gp_task,diskquota;gpftsand gpMgmt's tools fromgpinitsystemtogpexpand.Built on the port since: pg_vexec, in its own repository:
vexec(about 43,000 lines of C),vexec_flight(7,900) and the kernel packs for pgvector and PostGIS (1,000). vexec builds with PGXS both for vanilla PostgreSQL 19 and for the port. Its side of the port is module code inpg19/, about 13,600 lines with theshmtransport and PAX's copies: gp_orca's API for vector nodes andCCostModelVec, gp_core's API 1.15 (a Motion rebuilt by another module, the distribution's hash functions,squelch_subtree(), figures in EXPLAIN ANALYZE), and PAX's and gp_ao's batch readers and sinks. vexec calls none of the core patches.What changes for users:
gp.*, althoughgpconfigaccepts the old names;Rollout/Adoption Plan
Nothing here affects Cloudberry 2.x or the 3.0 plans: the prototype lives in my fork and changes no lines in Cloudberry's files. I would like to ask the community for four things:
pg19/. It is one directory that tracks Cloudberry's sources without editing them. If the direction looks right, would the project consider hosting it, for example as an experimental branch or a separate repository?isolation2_schedulenow runs 113 of its 262 tests in both passes (60 when this was first posted); its 133 append-optimized tests and the 95 skipped tests ofgreenplum_scheduleare next.gp.mpp_plannerswitch would allow side-by-side comparison, and I would value the view of the planner experts here.For users, migration would be logical, similar to the cluster-to-cluster copy planned for 2.0 → 3.0 in #1095. Separately, four of the custom hooks (table AM registry, smgr file events, parser hook and OID hook) are general enough to propose to PostgreSQL upstream. Support from Cloudberry developers as their real-world user would strengthen that case.
Are you willing to submit a PR?
All reactions