Skip to content

Renew config message TTLs at most once an hour per swarm - #2225

Merged
mpretty-cyro merged 3 commits into
session-foundation:devfrom
mpretty-cyro:fix/config-ttl-extension-cooldown
Sep 28, 2026
Merged

mpretty-cyro merged 3 commits into
session-foundation:devfrom
mpretty-cyro:fix/config-ttl-extension-cooldown

Conversation

@mpretty-cyro

@mpretty-cyro mpretty-cyro commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Every poll asked the swarm to extend the TTL of every active config message. That is a write on every storage node holding those messages, every few seconds, from every client, and service nodes see real disk I/O from it. The extension goes to weeks from now, so doing it once an hour loses nothing.

The change

A poll now skips the config TTL extension for a swarm whose last extension succeeded within the last hour.

  • Per swarm. The user swarm and each group swarm have their own cooldown, so one group's renewal never suppresses another's or the user's own.
  • Success only. A failed extension leaves the cooldown untouched, so the next poll retries it. A failure that started the cooldown would leave the configs un-renewed while looking handled, and repeated failures could let them age out of the swarm.
  • In memory, one hour, constant. No persistence or migration. A restart re-arms it, so the first poll after launch always extends.
  • A throttled poll sends no expire request at all, rather than sending one and ignoring the answer.

Disappearing-message expiry uses a separate path and is unchanged.

Android specifics

  • ConfigTtlExtensionThrottle (@Singleton, monotonic TimeSource) holds the per-swarm timestamps. Poller and GroupPoller wrap their existing AlterTtlApi extend in extendIfDue(swarmPubKeyHex) { … }.
  • The cooldown starts only when the closure returns normally. With batching, SnodeApiBatcher hands each sub-request its own handleResponse result, and AbstractSnodeApi throws on non-2xx. So a 500 on just the expire, or a failure of the whole batch, throws inside the closure and nothing is recorded. The due check does not suspend, so it doesn't push the extend outside the 100ms batching window.

Verification

ConfigTtlExtensionThrottleTest, 8 tests using TestTimeSource: two polls inside the window send one; the window is still closed just before an hour; a poll after the window sends another; a failed, repeatedly failing, or cancelled extension doesn't start the cooldown; per-swarm independence. 8/8 on testPlayDebugUnitTest. With the timestamp set before extend() instead of after, exactly the 4 failure-path tests go red.

Same change in the other clients

Every poll asked the swarm to extend the TTL of every active config
message, which is a write on each storage node holding them, every few
seconds, from every client. Service nodes see real disk I/O from it.

The extension is now skipped for a swarm whose last extension succeeded
within the hour. The user swarm and each group swarm are tracked
separately. A failed extension leaves the cooldown untouched so the next
poll retries it. State is in memory only, so the first poll after launch
always extends.
@mpretty-cyro
mpretty-cyro merged commit 16782ae into session-foundation:dev Sep 28, 2026
5 checks passed
@mpretty-cyro
mpretty-cyro deleted the fix/config-ttl-extension-cooldown branch September 28, 2026 06:24
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant