Transactions & connectionsstaff🔒 Pro concept

Connection scaling and GetSnapshotData (the PG14 dense-array snapshot rewrite)

Simple terms

You have heard that thousands of idle connections wreck snapshot performance for the whole cluster. Half of that is out of date. Every query starts by taking a snapshot of which transactions are still running. Before Postgres 14, building that snapshot got more expensive with every connection open, idle or not, so the folklore was earned. From 14 onward the cost tracks how many transactions are actually active. Idle connections are close to free for readers now. They still cost you, just not here.

You might be asked

Production wisdom says 'idle-in-transaction connections destroy snapshot performance for the whole cluster' and 'thousands of idle connections make GetSnapshotData O(N) and crush you'. On a modern PostgreSQL (14+), which half of that is still true and which half is folklore? Walk me through what GetSnapshotData actually scans, what makes a snapshot cheap to take, and prove where the real wall is when you scale connections.

TopicConnections / snapshots / ProcArray / GetSnapshotData / scaling
PostgreSQL14, 15, 16, 17, 18
Tools usedpgbench (read-only -S sweep, median-of-3), pg_stat_activity wait-event sampling, /proc/<pid>/status VmRSS, a two-phase settle-polled holder harness (600 no-xid vs 600 live-xid backends), source navigation @REL_17_10, PG 17.10 on 8 logical cores
Last reviewed2026-06-19

Pro concept

Full answer and evidence depth sit behind Pro

You have the question and a plain-English lead. Pro unlocks the short answer, what the docs say, what the code does, the labeled lab evidence or reproduction protocol, how to use it under pressure, and the references.

ProFull answer + evidence depth

Unlock the full breakdown for Connection scaling and GetSnapshotData (the PG14 dense-array snapshot rewrite)

You have the interview question and a plain-English lead. Pro opens the short answer, the manual walkthrough, the source decode, and a 56-line evidence section that states whether it is raw output, a captured run summary, or a protocol to run yourself, plus how to use it under pressure.
  • How you'd answer it, the full senior-level short answer
  • What the docs say, the manual's actual wording, with the citations
  • What the code does, the mechanism decoded from PostgreSQL's own source at a pinned tag
  • Proof from a real run, labeled as raw output, a captured run summary, or a run-it-yourself protocol
  • Using it under pressure, the situation, the call you'd make, what goes wrong, and what people get wrong
  • Version notes, what changed across PostgreSQL releases
  • 5 interviewer follow-ups for this concept

Card required. Cancel before day 7 and you are not charged.

Compare plans

How this was verified

The open teaser is the question and a plain-English lead. The short answer, the manual walkthrough, the source decode, the lab evidence, and how to use it under pressure unlock with Pro. Evidence is labeled as raw output, a captured run summary, or a run-it-yourself protocol.

Connected

Where this concept connects

How this concept links across the library, the interview questions that test it, its plain-English glossary definition, and the guided pathways it belongs to. Open the full map to explore further.

Open in the interactive map →