What Made Postgres 18 Faster: Larger Reads, Not Async I/O
We moved SnoutData Cloud to Postgres 18 and wanted our own number for its best-known new feature, asynchronous I/O. On a host shaped like our fleet's, one query got much faster: a cold bitmap heap scan took 2.38 seconds on 18 against 10.33 on 17. But turning asynchronous I/O off did not slow it down, a sequential scan ran at the same speed on both versions, and a vacuum was slower on 18. This note is the method, every run, and what the speedup is actually made of.
The question
Postgres 18 can keep several reads in flight instead of waiting for each one in turn. With io_method=worker, the setting we run, a small pool of I/O processes does the reading while the query works. The case for it is strongest on network storage, where each read waits on a round trip, and on a cold cache, which is where a project that has just woken from object storage starts. Both describe SnoutData Cloud, so we measured it on the hardware a project actually runs on, instead of quoting someone else's benchmark.
How we measured
- The host: an EC2 t4g.medium (Graviton2, 2 vCPUs, 4 GiB), Debian 13, rootless Podman 5.4.2, nothing else running. That is the shape of our fleet's hosts. The data sat on a separate 100 GB gp3 volume (3,000 IOPS and 125 MiB/s, its defaults), mounted where production keeps pod data.
- The pods: our own published images,
:17(Postgres 17.11) and:18(18.6), started through their normal entrypoint with a Plus project's limits: 1 CPU, 1,024 MB of memory,shared_buffers128 MB,work_mem4 MB. Only one pod ran at a time. - Four configurations, each with its own copy of the data: 17 as shipped; 17 with
effective_io_concurrencyandmaintenance_io_concurrencyraised to 16, which are 18's new defaults; 18 as shipped (io_method=worker, 3 workers); and 18 withio_method=sync, which turns asynchronous I/O off and keeps everything else about 18. Every other setting matched, and we recorded them from each server. - The data: pgbench's tables at scale 200, generated the way
pgbench -idoes, except thatabalancewas filled so an index on it is selective (pgbench leaves it zero).pgbench_accountsis 20 million rows, 2,561 MB of heap and 3,119 MB with its two indexes, well over the pod's memory. - Cold means cold: before every timed run, every pod was stopped, the host's page cache was dropped, the one pod under test was started, and the cache was dropped again. The four configurations took turns within each round, so anything that drifted during the run (see the burst, below) drifted for all four alike.
- Five rounds of each test, the median reported with the spread beside it, except where a section below says why it shows something else. Times are psql's own
\timing, from a client inside the pod.
The three tests:
-- a sequential scan of the whole table (index plans switched off, see below)
select count(*), sum(abalance) from pgbench_accounts;
-- a bitmap heap scan touching 10% of the rows (2,000,191 of 20,000,000)
select count(*), sum(bid) from pgbench_accounts where abalance between 0 and 9836;
-- a vacuum, after updating a contiguous 5% of the rows (1,000,000)
update pgbench_accounts set abalance = abalance + 1
where aid between 10000001 and 11000000; -- a new slice each round
vacuum (verbose) pgbench_accounts;Two plan notes, because each one would otherwise have measured the wrong thing. With an index on abalance, every version answers the first query with an index-only scan, which reads about 190 MB of index instead of 2.5 GB of table; our first attempt timed that by mistake. So the first query runs with index plans switched off, on every version, and its plan is a Parallel Seq Scan on all four. The second query's rows sit in runs of about 100 heap pages scattered across the table, but the planner estimates them as spread over every page and prefers a sequential scan, so it runs with enable_seqscan and enable_indexscan off, on every version. Its plan is a Parallel Bitmap Heap Scan reading the same 34,439 blocks on all four.
Results, cold cache
| Median seconds, and range | 17 | 17, io concurrency 16 | 18 (worker) | 18 (sync) |
|---|---|---|---|---|
| Sequential scan, whole table | 21.0520.03 to 21.07 | 21.0520.03 to 21.07 | 21.0520.03 to 21.06 | 21.0820.03 to 21.09 |
| Bitmap heap scan, 10% of rows | 10.339.83 to 10.62 | 11.5310.53 to 11.56 | 2.382.00 to 2.44 | 2.111.72 to 2.17 |
The low end of each scan's range is the first round, before the first update grew the table by 5%; every round after it read the same, slightly larger table on all four.
Two things in that table, and a third in the next one.
The bitmap heap scan is 4.3 times faster on 18, and asynchronous I/O is not why. With it switched off, 18 was a little faster still. And giving 17 the two settings 18 raised did not help it; it made it 12% slower.
The sequential scan is identical on all four, to within 0.2%. The volume's own counters show it reading 2,770 MiB in those 21 seconds, about 130 MiB a second, which is the volume's throughput limit (125 MiB/s, as provisioned). Nothing a database does can read faster than the disk delivers, and 17 already issued its sequential reads in large requests.
Vacuum is slower on 18 here
Each round updated a different contiguous million rows on all four copies, then timed a cold vacuum (verbose). Every round is shown, because the conditions were not the same in all of them: round 1 was the first vacuum to see pages the earlier setup had left without their visibility bits, and the host's burst allowance (below) ran out partway through round 4, so in that round 18 with io_method=sync ran after it was gone and the other three before.
| Cold vacuum, seconds | 17 | 17, io concurrency 16 | 18 (worker) | 18 (sync) |
|---|---|---|---|---|
| Round 1 (first vacuum after setup) | 46.9 | 46.5 | 67.8 | 69.4 |
| Round 2 | 9.3 | 9.6 | 12.2 | 11.5 |
| Round 3 | 9.6 | 9.6 | 11.8 | 12.1 |
| Round 4 (burst ran out before the last) | 9.8 | 9.8 | 12.2 | 38.6 |
| Round 5 (burst used up for all four) | 33.4 | 35.3 | 39.5 | 40.2 |
In every round run under the same conditions, 18 took 18% to 45% longer than 17, and io_method made no consistent difference. The vacuum logs say where the time went. Both versions scanned the same 10% of the table, read the same pages and dirtied the same 59,342 of them. But 18 wrote 319 MB of WAL per vacuum against 17's 45 MB, with 42,126 full page images against 24,905, and froze all 1,000,000 rows the update had just written, where 17 froze about 50,000.
That is the pattern data checksums produce. With checksums on, the first change to a page after a checkpoint, even just a hint bit, writes a full image of the page to the WAL, and a vacuum that is already logging a full image of a page freezes it at the same time, because that is nearly free then. 18 turns checksums on by default; our 17 image had them off. We did not run 17 with checksums on to confirm it, so we call this consistent with checksums rather than caused by them. It is a trade we would make again: corruption reported as an error is worth some vacuum time, and rows frozen now are rows a later anti-wraparound vacuum does not have to touch.
What the speedup is made of
The bitmap scan reads the same 34,439 blocks on every version. What changed is how many requests the disk sees for them. We read the volume's own counters before and after one cold run of each (device-wide, so they include the kernel's readahead and the index):
| One cold bitmap heap scan | 17 | 17, io concurrency 16 | 18 (worker) | 18 (sync) |
|---|---|---|---|---|
| Reads the volume served | 52,295 | 63,479 | 10,185 | 8,467 |
| Average read size | 11 KB | 9 KB | 61 KB | 69 KB |
17 asked the volume for 52,295 reads averaging 11 KB; 18 asked for 10,185 averaging 61 KB, five times fewer and more than five times larger, and with asynchronous I/O off it asked for fewer still. The sequential scan, for comparison, was already about 124 KB a read on 17 (about 22,900 reads on all four), which is why it did not move.
That is a change in 18 that has nothing to do with io_method: bitmap heap scans now read through the same streaming read path that sequential scans already used in 17, and that path merges neighbouring blocks into one larger read. Our rows sit in runs of about 100 pages, so there were plenty of neighbours to merge. A bitmap scan whose matches are scattered one per page would have little to merge, and we would expect a smaller gain there; we did not measure that case.
Asynchronous I/O found nothing to add on this host. On network storage every request costs a round trip, and 18 had already cut the number of requests; keeping the remaining ones in flight at once made no measurable difference, and with io_method=worker the scan took 13% longer than with sync. We have not pinned down why. One thing we can say about the setup: a Plus pod has one CPU, and the I/O workers share it with the query.
The burst, and a busy host
A t4g.medium can push data to its volumes faster than its baseline for a limited time, and it spends a credit balance to do it. The run above started with that balance full and watched it drain from 99% to zero, mostly spent by the vacuums. A host that has used it up is held to about 43 MB/s, a third of what the volume itself allows. So once the balance read zero, we ran the same scans again, five rounds, and kept the four in which it stayed there for all four configurations.
| Burst used up, median seconds (rounds 2 to 5) | 17 | 17, io concurrency 16 | 18 (worker) | 18 (sync) |
|---|---|---|---|---|
| Sequential scan | 66.50 | 60.15 | 59.54 | 66.50 |
| Bitmap heap scan | 20.33 | 20.88 | 14.00 | 13.30 |
Once the host is held to its baseline, the bitmap heap scan is still faster on 18, but by 1.45 times, not 4.3. The sequential scan settled at two speeds, about 60 and 66.5 seconds, and the two followed each pod's place in the round (the first and last ran at 66.5, the middle two at about 60, every round), not its version. 66.5 seconds is 2.8 GB at about 43 MB/s. A host whose projects are mostly asleep, which is most of ours, keeps its balance; a host doing heavy I/O for long enough lives in this table.
Warm cache, for honesty
The same three once more, each straight after a run that had read the same data, with the burst used up. The sequential scan is not really warm: 2.8 GB does not stay cached for a pod limited to 1 GB of memory, so it read from disk again, at the same speed on all four. The bitmap scan's 282 MB does stay cached, and from memory all four take under a second. The vacuum, run straight after its update, is 32% slower on 18 again.
| One warm run, seconds | 17 | 17, io concurrency 16 | 18 (worker) | 18 (sync) |
|---|---|---|---|---|
| Sequential scan | 64.82 | 64.86 | 64.82 | 64.88 |
| Bitmap heap scan | 0.83 | 0.80 | 0.87 | 0.84 |
| Vacuum after a 5% update | 25.5 | 26.1 | 33.7 | 33.7 |
One cost we noticed
Loading the data, which is not one of the tests, took 250 seconds on 18 and 205 on 17, once each; building the primary key alone took 79 seconds against 58. 18 also does more work per page, because its data checksums are on by default and 17's were off, and every page read above was checksummed on 18 and not on 17. We have not separated that from the rest of the load's difference, and with one run each we report it as an observation, not a result.
What this means for a project on SnoutData Cloud
- A query that reads clustered rows through a bitmap scan, which is common for range filters on a column that roughly follows insert order, can read several times faster from a cold cache on 18. That is the case that matters most on a platform where an idle project sleeps and wakes with nothing cached.
- A full-table scan is bounded by the disk on this hardware, on any version, and 18 does not change that. A vacuum is slower on 18, by 18% to 45% here, which as far as we can tell is the price of its checksums.
- We keep
io_method=worker, the setting 18 ships with. On this hardware it made no difference to the sequential scan or the vacuum and cost 13% on the bitmap scan, against a 4.3 times gain from the rest of 18. We have not measured it on a pod with more than one CPU.
We did our best to make this fair. We also made mistakes on the way, and each one was caught by a number that looked wrong: our first sequential scan was an index-only scan on every version, and our first vacuum test rewrote the whole table each round and spent most of the host's burst allowance. Both were redone. If you think we got something wrong, write to [email protected] and we will re-run it and publish the correction.
Every timed run, the plans, the vacuum logs and the scripts that produced them are kept with the run. The engineering side of the move, including the three things in it that would have bitten us, is in the blog post.