Commit Graph
100 Commits
Author SHA1 Message Date
Jannis Braun 3988c5823a feat(shared): add no_recipient reason + undeliverable response field (#18)
Additive protocol extension. No consumers yet — follow-up commits wire
the new bucket into the relay endpoint, sendCallRelay, sendFederatedCallStart,
and the toast copy.
2026-04-24 19:12:54 +02:00
Jannis Braun a68eab3edc Merge branch 'feat/remote-participant-host-unreachable' 2026-04-24 01:09:35 +02:00
Jannis Braun 5c94bc7669 refactor(server): remove unused admin_reset PeerDeactivationReason
The admin reset endpoint (DELETE-pattern gated on peer.status !== 'needs_attention')
doesn't transition status — it deletes the row of an already-deactivated peer.
onPeerDeactivated already fired at the earlier needs_attention transition, so
the reset site correctly has no hook. The enum value was defensive-unused; per
project principles (no backwards-compat shims, no placeholders) drop it.
2026-04-24 01:07:47 +02:00
Jannis Braun 7c0dd33123 docs(systems): document host_unreachable phase + onPeerDeactivated + sentinel 2026-04-24 00:54:46 +02:00
Jannis Braun 5e509cf3df feat(web): host_unreachable phase copy for dm_call_undeliverable (TDD) 2026-04-24 00:51:30 +02:00
Jannis Braun 314df6c5a1 feat(server): 30s federated-call sentinel worker (TDD) 2026-04-24 00:49:04 +02:00
Jannis Braun 71445c5f27 feat(server): wire onPeerDeactivated at admin revoke/reset sites
Admin-revoke endpoint (DELETE /api/federation/peers/:id): fires
onPeerDeactivated(id, 'admin_revoked') after the status write to
'revoked', evicting any in-flight federated calls for the now-revoked
peer.

Admin-reset endpoint (POST /api/federation/peers/:id/reset): hook
SKIPPED. The reset endpoint is guarded to only run when status is
already 'needs_attention' (active peers are rejected at the boundary
with a 400). Because the peer was already deactivated before reset is
called, onPeerDeactivated was already fired at the active→needs_attention
transition. The reset deletes the row entirely rather than writing a new
status; it does not represent a transition OUT OF active, so wiring it
here would be a semantic error — double-evicting an already-deactivated
peer.
2026-04-24 00:45:02 +02:00
Jannis Braun 8e6639648e feat(server): wire onPeerDeactivated on performHandshake 403 rejection 2026-04-24 00:43:32 +02:00
Jannis Braun 743fdcac97 feat(server): wire onPeerDeactivated at federationWorker peer-deactivation sites 2026-04-24 00:42:20 +02:00
Jannis Braun 0a8949fbfb feat(server): onPeerDeactivated utility mirrors onPeerActivated (TDD) 2026-04-24 00:40:42 +02:00
Jannis Braun 3b61380a1e feat(server): ConnectionManager.evictFederatedCallsForHost (TDD) 2026-04-24 00:37:44 +02:00
Jannis Braun 942619f422 feat(shared): add 'host_unreachable' phase to DmCallPhase 2026-04-24 00:34:05 +02:00
Jannis Braun 726341f8f8 Merge branch 'feat/call-state-machine-hardening' 2026-04-23 23:44:49 +02:00
Jannis Braun 1e59e7012c fix(server): scope accept-rollback terminal to the acceptor only
Code-review catch: the Path-2 accept-rollback previously emitted
dm_call_undeliverable { terminal: true } via sendToFederatedCallUsers,
which broadcasts to every ringedUserIds entry. In a group DM this
would prematurely tear down non-accepting ringees whose own accept /
reject / timeout paths should govern their state. Switch to
sendToUser(acceptorId) so only the acting user gets the terminal
signal. Reorder the clearFederatedCall to happen before the emit so a
concurrent end-handler sees a cleared entry (clearFederatedCall is
idempotent). Spec updated, test extended to assert the scoping with a
two-ringee group-DM fixture.
2026-04-23 23:38:31 +02:00
Jannis Braun 1719e6d580 docs(systems): document federation call state machine hardening 2026-04-23 23:23:06 +02:00
Jannis Braun 6a5b02b1a0 feat(web): phase-aware dm_call_undeliverable toast copy (TDD) 2026-04-23 23:21:40 +02:00
Jannis Braun f6252b8ce1 feat(server): fan dm_call_end out on host ring timeout 2026-04-23 23:19:42 +02:00
Jannis Braun 6cd7728f3e feat(server): fanOutCallEvent returns failures; surface to host-side caller 2026-04-23 23:14:50 +02:00
Jannis Braun 3cb80d110d feat(server): aggregate Path-1 call fan-out failures and surface to originator 2026-04-23 23:12:55 +02:00
Jannis Braun 13241345de feat(server): surface dm_call_end relay failure to originator (TDD) 2026-04-23 23:11:20 +02:00
Jannis Braun 5170d316ba feat(server): surface dm_call_reject relay failure to rejector (TDD) 2026-04-23 23:10:18 +02:00
Jannis Braun 07c5b0e7de feat(server): surface dm_call_accept relay failure to acceptor (TDD) 2026-04-23 23:09:18 +02:00
Jannis Braun 06e1ed92e7 refactor(server): add buildFailureFromResult + CallFanoutFailure types 2026-04-23 23:07:07 +02:00
Jannis Braun 5e9b353124 feat(shared): add DmCallPhase to dm_call_undeliverable event 2026-04-23 23:06:16 +02:00
Jannis Braun fe01b94876 Merge branch 'fix/dev-db-rename-drift'
Close backlog #30: second-pass heal that DROPs and rebuilds empty
drifted tables to reconcile pre-rename column drift on old pre-drizzle
dev DBs (companion to #29). `healInitialSchemaDrift` handles
addable-column drift; this handles the NOT-NULL-no-default case that
heal-by-ALTER can't resolve — e.g. federation_outbox's
`message_id`/`dm_channel_id` that were renamed to
`entity_id`/`context_id` in the pre-drizzle manual migrate.ts.

Rebuild target is the current-migration-state snapshot (derived from
__drizzle_migrations ↔ _journal.json by `when` timestamp), not the
latest on disk — rebuilding forward would duplicate-column with
pending ALTER TABLE ADD COLUMN statements drizzle is about to run.

Empty-table gate preserves data. Missing-column gate (not extra-only)
avoids touching unused leftover columns from pre-drizzle manual
migrations that no current code reads.

Verified against three simulated scenarios (fresh / pre-drizzle
drift / production-like) plus live boot on the actual affected dev
DB: server binds :3005 cleanly, no outbox worker SQLITE_ERROR ticks,
migrations complete silently. Server tests 110/110, web 131/131.

No migration files changed. Deployed instances unaffected. Skip
redeploy.
2026-04-23 03:20:08 +02:00
Jannis Braun 06c538b013 fix(server): rebuild empty drifted tables to heal pre-rename column drift
Companion to #29. `healInitialSchemaDrift` can only ADD columns, so it
skips NOT NULL-without-default columns like `federation_outbox.entity_id`
/ `context_id` — which on some old pre-drizzle dev DBs carry the
pre-rename names `message_id` / `dm_channel_id` instead. The tables
load but the outbox worker fails every tick with "no such column:
federation_outbox.context_id" once the server is up.

Adds healRenamedColumns() — a second pass that runs right after
`healInitialSchemaDrift`. For each table whose physical column set is
*missing* columns declared by the current-migration-state snapshot AND
which holds zero rows, it DROPs the table and rebuilds it from the
snapshot's JSON: columns, defaults, foreign keys, composite PKs,
unique constraints, indexes.

Key design decisions:

- **Target is the current-migration-state snapshot, not the latest on
  disk.** The current state is determined by the highest
  `__drizzle_migrations.created_at` matched against `_journal.json`'s
  `when` timestamps (with backward walk for idx values that lack a
  snapshot, like the hand-written 0002). Rebuilding to a *future*
  snapshot would introduce columns that drizzle's migrator is about
  to add via ALTER TABLE ADD COLUMN, causing duplicate-column errors.
  Rebuilding to the *current* snapshot preserves the invariant that
  drizzle's pending migrations can run cleanly afterwards.

- **Missing-column gate, not extra-column.** Extra columns alone don't
  break anything at runtime (the ORM ignores them); they're leftover
  from pre-drizzle manual migrations and might matter to the operator.
  Missing columns DO break runtime queries, so only those trigger
  rebuild.

- **Empty-table gate.** Non-empty tables log a warning and skip —
  data preservation wins over heal, and this path should only ever
  hit a pre-drizzle dev DB that never exercised the affected tables
  in the first place.

- **Transactional rebuild.** DROP + CREATE + index reinstatement wrap
  in a single `db.transaction()` so a partial rebuild rolls back.

Verified against three scenarios via in-memory simulation:
(A) fresh install — heal no-op, drizzle creates everything; (B) pre-
drizzle dev DB with fed_outbox/fed_mutation_log rename drift —
tables rebuilt to 0000 snapshot, drizzle then applies 0001–0004
successfully to reach the current target schema; (C) post-migration-
correct (production-like) — heal no-op, drizzle no-op, schema
unchanged. Live boot on my actual dev DB: migrations complete
silently, server binds :3005, no outbox worker errors. Server tests
110/110, web 131/131, typecheck clean.

No migration files changed. Deployed Pi+VM instances are unaffected
(their schema matches the snapshot exactly — heal won't touch
anything).

Closes backlog #30.
2026-04-23 03:19:56 +02:00
Jannis Braun 02a8b27d2b Merge branch 'fix/joinspace-tests-and-dev-db-drift'
Two small items from the S2S DM unification backlog.

#28 — 4 stale JoinSpace test assertions (test-only fix)
  Assertions from the old JoinServer component drifted when the file was
  renamed (fc06e25) without being updated: placeholder text, submit
  button now disables-when-empty (so the 'Invite code is required' error
  path is unreachable from the rendered form), and joinByCode signature
  gained a second `origin` argument. Updated each assertion to match
  current behavior. No code change. 131/131 web tests now pass.

#29 — Local dev DB missing federation_peers.remote_max_upload_size
  Root cause: baselineExistingInstall trusts that any pre-existing
  install's schema matches 0000_initial. That breaks for dev DBs from
  the pre-drizzle manual migrate.ts system that didn't ran every
  idempotent ALTER — 0000 gets marked applied without its columns
  actually existing, and a later migration that recreates the table
  (0004_cooing_black_knight) crashes. New healInitialSchemaDrift()
  walks 0000_snapshot.json, and for each existing table ADDs any
  declared-but-missing columns before drizzle's migrate() runs.
  Idempotent; production instances are no-op.

Remaining dev-DB drift flagged for a follow-up: federation_outbox and
federation_mutation_log retain pre-rename column names (message_id /
dm_channel_id) that the old manual system renamed to entity_id /
context_id. Both tables are empty on affected DBs, but the clean fix
requires DROP + RECREATE with index reinstatement, beyond #29's
scope. Outbox worker emits SQLITE_ERROR ticks post-boot on unfixed
dev DBs; separate item.

No migration files changed. Deployed instances are unaffected by
either commit (#28 test-only, #29 heals only DBs missing columns —
production DBs aren't missing any). Skip redeploy.
2026-04-23 03:00:49 +02:00
Jannis Braun e703e29f8a fix(server): heal 0000-baseline schema drift after baselining existing installs
baselineExistingInstall marks 0000_initial as applied when it detects
pre-existing tables, on the assumption the install's schema matches the
0000 baseline. That assumption is false for dev DBs created under the
pre-drizzle manual migrate.ts system that skipped or never ran some of
its idempotent ALTER TABLE steps — for example the b9e4c65 migration
that added federation_peers.remote_max_upload_size. On such DBs, 0000
is marked done without the column actually existing, and a later
migration that recreates the table (0004_cooing_black_knight)
subsequently crashes with "no such column: remote_max_upload_size"
while building its __new_federation_peers SELECT.

Adds healInitialSchemaDrift(): walks every table in 0000_snapshot.json,
and for each table that already exists, ADDs any columns the snapshot
declares but the physical table is missing. Runs immediately after
baselining, before drizzle's migrate() — so later migrations find the
schema they expect. Columns that SQLite's ALTER TABLE ADD COLUMN can't
safely express (PRIMARY KEY; NOT NULL without a default) are skipped
with a warning rather than corrupting data.

Idempotent: on fresh installs and correctly-migrated DBs every column
is already present, so the loop is a no-op. Production Pi+VM instances
are unaffected.

Verification: local dev DB that previously crashed on 0004 now boots
cleanly — federation_peers gained remote_max_upload_size, nonce_supported,
pending_hmac_secret, secret_rotation_at, secret_rotated_at, and
auto_rotate_interval_days; __drizzle_migrations advanced from 4 to 5
entries; server binds :3005. 110/110 server tests + 131/131 web tests
still pass.

Not covered: a deeper drift on federation_outbox /
federation_mutation_log where the physical tables retain pre-rename
column names (message_id / dm_channel_id) instead of the current
entity_id / context_id. Heal skips those (NOT NULL without default)
and the outbox worker emits SQLITE_ERROR ticks post-boot. Both tables
are empty on affected dev DBs, but a clean fix requires DROP +
RECREATE with index reinstatement which is out of #29's stated scope
("column existing"). Flagged for a follow-up.

Closes backlog #29.
2026-04-23 03:00:30 +02:00
Jannis Braun 086158511c test(join-space): update stale assertions to match current UI and API
Four assertions drifted from the current JoinSpace modal, carried over
when the file was renamed from the old JoinServer component in fc06e25
without being updated:

- Placeholder was expanded to cover URL-form invite input
  ('e.g. abc123' → 'e.g. abc123 or https://instance.com/join/abc123').
- 'shows validation error when submitting empty code' asserted a code
  path that no longer exists: the submit button is now disabled when
  the trimmed input is empty (JoinSpace.tsx line 166), so clicking it
  is a no-op and the 'Invite code is required' error from the parser
  is unreachable from the rendered form. Replaced with an assertion
  that the button is disabled while the input is empty — the actual
  validation UX.
- joinByCode signature took on a second `origin` argument during the
  S2S DM unification + federated-join work (spaceStore.ts line 69).
  parseInviteInput returns { code, origin: undefined } for a bare
  code, so the call is `joinByCode('my-invite-code', undefined)`.
  Assertion updated to match exactly.

No code behavior change — tests now reflect actual behavior, which
was already correct and deployed. Closes backlog #28.
2026-04-23 02:52:00 +02:00
Jannis Braun 7220f4cbcd Merge branch 'fix/cross-store-resolver-tdz'
Close backlog #27: extract cross-store resolvers into a neutral utility
(packages/web/src/utils/crossStoreResolvers.ts) to break a TDZ cycle
between spaceStore and instanceStore. instanceStore's top-level
setResolver calls used to race with spaceStore's `let _getApiForOrigin`
declaration when the module graph was entered from instanceStore
(JoinSpaceModal → useInstanceStore), crashing with "Cannot access
'_getApiForOrigin' before initialization" and preventing
InviteModal.test.tsx and JoinSpace.test.tsx from loading.

Moves the three resolver lets + setters + pure getters + the
WS-populated user-ID cache into the utility; spaceStore re-exports the
public surface; instanceStore imports the setters directly from the
utility (re-exports do not resolve at module-init time under vite-ssr
in the cycle). authStore-using wrappers (resolveUserOrigin,
getLayoutHomeOrigin, getMyUserIdForOrigin) stay in spaceStore but
delegate to the utility.

Also adds the AudioManager mock to InviteModal/JoinSpace test files
(established pattern) so their suites can load.

Verification: typecheck clean, 127/131 web tests pass (up from 121/121;
+6 newly unlocked), server 90/90 unchanged. The 4 remaining JoinSpace
failures are pre-existing stale UI-text assertions (placeholder
expanded, submit button now disable-when-empty) — unrelated to this
work, made visible only because the suite loads now.

Spec: docs/systems/client-federation.md updated.
Smoke test: vite dev bundle serves spaceStore + crossStoreResolvers
clean; full live E2E blocked by a pre-existing local DB-migration
error unrelated to this client-side refactor (reproduces on main).
2026-04-23 02:46:37 +02:00
Jannis Braun 4d2e50b55b fix(web): extract cross-store resolvers into neutral utility to break TDZ
instanceStore registers three resolver functions at module load —
setApiForOriginResolver, setUserIdForOriginResolver,
setOriginFromHostnameResolver — whose backing `let` bindings used to
live in spaceStore. When the module graph was entered from
instanceStore (e.g. JoinSpaceModal importing useInstanceStore) the
order became spaceStore → chatStore → useWebSocket → socialStore →
instanceStore (top-level setter call) while spaceStore was still
paused on its line-8 chatStore import, so the backing `let` had not
been reached yet and the setter crashed with
`Cannot access '_getApiForOrigin' before initialization`. This left
InviteModal.test.tsx and JoinSpace.test.tsx unable to even load their
suites once AudioManager was mocked away.

Move the three `let` bindings, their setters, their pure getters, plus
the WS-populated user-ID cache (`_myUserIdByOrigin`, setMyUserIdForOrigin,
getCachedUserIdForOrigin, clearMyUserIdCache) into
`packages/web/src/utils/crossStoreResolvers.ts`. The utility imports
nothing from `./stores/*`, so no back-edge exists. spaceStore re-exports
the public surface for backward compatibility with the many existing
import sites; instanceStore imports the setters directly from the
utility (the in-cycle re-export path does not resolve at module-init
time under vite-ssr, so a direct import is required for the top-level
setter calls).

spaceStore's remaining wrappers (resolveUserOrigin, getLayoutHomeOrigin,
getMyUserIdForOrigin) stay where they are — they combine the utility's
pure lookups with authStore state — but now delegate to the utility.

Also adds the AudioManager mock to InviteModal.test.tsx and
JoinSpace.test.tsx so their suites actually load (same pattern already
used in 5 other test files). Net test-suite result: 127/131 pass (up
from 121/121 — +6 newly unlockable). The 4 remaining JoinSpace
failures are pre-existing stale UI-text assertions (the placeholder was
expanded and the submit button was made disable-when-empty) made
visible by the suite now loading; they're orthogonal to this change
and handed back for a separate triage.

Closes backlog #27.
2026-04-23 02:46:19 +02:00
Jannis Braun 15d2a3f163 Merge branch 'fix/pre-existing-test-failures'
Triage and fix of 12 pre-existing test failures on main (flagged by
the #10 DM origin failover merge, commit ff39ab0).

Two root causes:

1. Node 20+'s built-in localStorage/sessionStorage stub shadows jsdom's
   working implementation because vitest's populateGlobal doesn't
   overwrite globals outside its known allow-list. Any zustand persist
   store threw "storage.setItem is not a function". Polyfilled with an
   in-memory Storage in src/test/setup.ts. Resolves 11 of 12 failures
   (10 keybindStore + 1 FriendsPage toast).

2. FriendsPage "Message button" DM test carried a stale two-arg
   assertion that predated the 2026-04-01 federation refactor (commit
   7f3ca4e) which dropped the second argument from addDmChannel.
   Dropped the trailing '' so the assertion matches current behavior.

Hand-backs (not touched on this branch):
- InviteModal.test.tsx and JoinSpace.test.tsx still fail to LOAD (not
  in the 12 tests but flagged by #10's merge note). After stubbing
  AudioManager a second blocker surfaces: TDZ error on _getApiForOrigin
  in spaceStore.ts:961, caused by a circular-import init order between
  spaceStore and instanceStore (via socialStore → useWebSocket →
  voiceStore). Federation-adjacent — filed as backlog item for a
  structural fix.

Server: 90/90 unchanged.  Web: 121/121 pass (was 109/121), 2 suite
loads still failing (tracked).
2026-04-23 02:29:20 +02:00
Jannis Braun ae5bdaa338 test(friends): drop stale origin arg from addDmChannel assertion
The DM "Message button" test asserted addDmChannel was called with two
arguments — the channel and an empty-string origin — but the assertion
has been stale since commit 7f3ca4e ("route DM creation to home instance
with federated identity", 2026-04-01). That refactor made FriendsPage
always route DM creation through the home api client and dropped the
second argument from the addDmChannel call because the remote friend's
instanceOrigin no longer applies — home-created DMs don't need a
channelOriginMap entry (lookups default to '' for missing keys; remote-
delivered DMs still get their origin tagged by useWebSocket).

The two-arg assertion was introduced on 2026-03-25 (commit 277b69a)
against an intermediate form of the code that was later rewritten. Drop
the trailing '' so the assertion matches the current, intentional
one-arg call.
2026-04-23 02:28:57 +02:00
Jannis Braun 9fbe07571e test(web): polyfill localStorage/sessionStorage in jsdom env
Node 20+ ships a built-in localStorage/sessionStorage stub on globalThis
that throws "storage.setItem is not a function" unless Node is launched
with --localstorage-file=PATH. Vitest's jsdom env only overwrites a
fixed allow-list of globals, and neither storage is in that list — so
Node's broken stub shadows jsdom's working implementation and crashes
any code using zustand's persist middleware.

Override both globals in the test setup with an in-memory Storage
implementation. Resolves 10 keybindStore failures plus 1 FriendsPage
toast failure (all symptoms of the same root cause).

The remaining FriendsPage DM assertion failure is unrelated and is
left for follow-up triage as it touches federation DM-creation
behavior.
2026-04-23 02:23:05 +02:00
Jannis Braun ff39ab077d Merge branch 'feat/dm-origin-failover'
Client-side DM origin failover on WS disconnect (backlog #10).

When a remote instance's WebSocket drops mid-session, every DM pinned to
that origin is re-keyed to a connected sibling that mirrors the same
federated DM via S2S replication. Covers the channel-ID-per-origin
reality (each instance assigns its own local Snowflake; only federatedId
is shared) by keeping a `dmAlternatives: Map<federatedId, Map<origin,
localChannelId>>` on spaceStore, populated by every `ready` payload
regardless of dedup outcome. On transition, `rekeyDmChannel` atomically
renames the DM across spaceStore (dmChannels / channelOriginMap /
channelLastMessageIds / dmAlternatives), chatStore (messages / hasMore /
scrollPositions / channelAccessTimes / typingUsers / readStates /
unreadChannels / currentChannelId), and the URL (history.replaceState
when viewing the rekeyed DM). Triggers: setInstanceStatus on
connected→disconnected|error, and disconnectInstance / forceRemoveEntry
before removeInstanceSpaces. Voice state (activeDmCall / outgoingCall /
incomingCall) is intentionally not rewritten — LiveKit rooms can't
migrate across origins.

As an in-scope adjacent fix (§3.11 of the spec), `dm_message_created`
now consults `dmAlternatives` before its legacy 2-member-identity
fallback via a new `resolveDmChannelId(rawId)` helper. This closes a
pre-existing phantom-sidebar-entry bug for group DMs in multi-instance
sessions and handles post-failover routing when the reconnected
original origin's WS still addresses the DM by its old local id.

Live verification on Pi+VM (youruser@nova.ddns.net with Orbit.Backspace
as remote): (1) baseline — DM pinned to home, WS drop of remote is a
no-op, reconnect clean, no flap; (2) forced-rekey path — WebSocket
construction delayed on wss://nova.ddns.net via a client-side patch
so orbit's `ready` arrived first, pinning the Nova DM to orbit with
orbit's local id. Stopping the orbit container triggered
failoverDmOriginsFromDisconnected; URL auto-swapped from
`/channels/@me/<orbit-local-id>` to `/channels/@me/<nova-local-id>`
via history.replaceState, the chat view re-fetched from nova via the
new primary id, and subsequent message sends routed to nova. Restart
of orbit left the pin on nova — no re-home flap (§3.6). Group-DM
phantom fix covered by unit tests (9 in dmOriginFailover.test.ts plus
contract test); not exercised live because it requires concurrent
delivery from the non-primary origin's WS, which the forced-rekey
session didn't naturally produce.

Design: internal notes
Plan:   internal notes

Pre-existing test failures on main (keybindStore, FriendsPage,
InviteModal, JoinSpace — 12 tests) are unchanged by this branch.
2026-04-23 02:04:21 +02:00
Jannis Braun de0b6c2a42 docs(federation): document DM origin failover mechanism
New subsection under client-federation.md's origin-aware routing section
covering dmAlternatives, failoverDmOriginsFromDisconnected, rekey flow,
trigger points, the intentional cache-flush trade-off, voice-out-of-scope,
no-re-home policy, and the WS routing contract. dm-system.md gets a
one-line cross-reference from the Client routing bullet.
2026-04-23 01:31:22 +02:00
Jannis Braun b1844e126f feat(federation): dmAlternatives fallback in dm_message_created
Adds a new resolution step before the legacy 2-member-identity fallback:
if the event's dmChannelId is an alternate-origin local id for a DM
whose primary is in dmChannels, route the message to the primary via
resolveDmChannelId. Covers 1-on-1 AND group DMs uniformly — closes a
pre-existing phantom-sidebar-entry bug for group DMs in multi-instance
sessions and handles post-failover routing when the reconnected original
origin's WS still addresses the DM by its old local id.
2026-04-23 01:29:28 +02:00
Jannis Braun cd1f5c2b64 feat(federation): trigger DM failover on user-initiated disconnect
disconnectInstance and forceRemoveEntry now run
failoverDmOriginsFromDisconnected BEFORE removeInstanceSpaces so any DM
with a connected sibling survives the disconnect via rekey; only DMs
without alternatives are cleared alongside the rest of the instance.
Switched setInstanceStatus to the same static import (dmOriginFailover
lazily reads store state, so no import cycle).
2026-04-23 01:25:42 +02:00
Jannis Braun 1a23871367 feat(federation): trigger DM failover on setInstanceStatus transition
When an instance transitions from 'connected' to 'disconnected' or
'error', fire failoverDmOriginsFromDisconnected for that origin. Dynamic
import preserves the circular-dep-safe resolver pattern used elsewhere
in instanceStore. Fire-and-forget; the failover utility reads fresh
state at call time.
2026-04-23 01:22:06 +02:00
Jannis Braun d393a870c2 feat(federation): dmOriginFailover utility (rekey + failover)
failoverDmOriginsFromDisconnected(origin) walks pinned DMs and re-keys
them to a connected sibling origin's local channel id (via dmAlternatives
federatedId lookup). Preference: home first, then any connected remote in
insertion order. rekeyDmChannel performs the atomic rename across
spaceStore (dmChannels / channelOriginMap / channelLastMessageIds /
dmAlternatives), chatStore (via rekeyChannelState), and the URL (via
history.replaceState when viewing the rekeyed DM). Voice state is
intentionally untouched — LiveKit sessions can't migrate across origins.
Old origin's local id is retained in dmAlternatives for possible later
fail-back without another ready round-trip.
2026-04-23 01:13:48 +02:00
Jannis Braun 678790b88b feat(federation): resolveDmChannelId for alternate-origin DM ids
Resolves any raw DM channel ID (primary or alternate-origin local ID)
to its primary dmChannels entry via dmAlternatives federatedId lookup.
Returns null for unknown IDs. Used by the dm_message_created handler
in a later commit to prevent phantom sidebar entries from alternate-
origin deliveries (closes a pre-existing group-DM bug and supports
post-failover routing).
2026-04-23 01:06:46 +02:00
Jannis Braun d66932362a feat(chat): rekeyChannelState moves channel state from oldId to newId
Deletes every channel-keyed entry under oldId (messages, hasMore,
typingUsers, readStates, channelAccessTimes, scrollPositions) without
seeding newId — subscribers refetch naturally from the new origin.
Transfers unreadChannels membership only if oldId was already unread
(mirror state, don't over-badge). Updates currentChannelId if it
matched oldId. Groundwork for DM origin failover rekey.
2026-04-23 01:04:42 +02:00
Jannis Braun e7430f1a54 feat(federation): prune dmAlternatives on removeInstanceSpaces
Drops the given origin from every inner (origin→localId) map; removes
the outer federatedId entry when its inner map becomes empty. Keeps the
store from accumulating stale origin references across long sessions
with connect/disconnect churn.
2026-04-23 01:02:05 +02:00
Jannis Braun 088fd40834 feat(federation): record DM origin alternatives in spaceStore
Every DM arriving in a ready payload with a federatedId now gets its
(origin, localChannelId) pair recorded in dmAlternatives, regardless of
whether the dedup pass kept this copy in dmChannels. Enables client-side
DM origin failover: when the primary origin drops, we can look up an
alternate origin's local channel ID for the same federated DM.

Prep for #10 (DM origin failover on disconnect).
2026-04-23 00:58:37 +02:00
Jannis Braun f1827872aa Merge branch 'fix/outbox-duplicate-terminal'
Outbox worker treats `reason: 'duplicate'` as effectively-accepted:
deletes the entry rather than retaining for retry. Duplicate is a
terminal signal (the peer already has the message — retrying will
fail identically forever until TTL expires). Logged at info level
to distinguish from retained-for-retry warnings.

Other rejection reasons (attribution_mismatch, processing_error,
etc.) stay on the retry path; treating additional reasons as
terminal is deferred until observed accumulating.

Discovered during post-#10b deploy verification: Pi had a stuck
outbox entry retrying a message VM already had, logged every
outbox tick. This patch fixes that class of issue.
2026-04-23 00:12:13 +02:00
Jannis Braun 6ff983b46c fix(federation): treat duplicate rejection as terminal in outbox worker
Duplicate rejection means the peer already has the message (e.g.,
delivered earlier via outbox AND pulled via sync in the same
window). Retrying will fail identically forever until TTL expires.

Before this patch: duplicate-rejected outbox entries were retained
with attempts++ and exponential backoff, creating log noise and
outbox bloat for up to 30 days.

After: duplicate-rejected entityIds join the terminal set alongside
accepted ones and are deleted from the outbox. Logged at info level
('outbox entry removed (terminal)') to distinguish from warn-level
transient-rejection retries.

Other rejection reasons (attribution_mismatch, processing_error,
etc.) stay on the retry path; some may also be terminal but are
deferred until observed accumulating.
2026-04-23 00:10:34 +02:00
Jannis Braun 3252868152 Merge branch 'fix/sync-poison-pill-skip'
Closes #25 (poison-pill event blocks sync forever).

syncPeerMutationLog now replays incoming events one-by-one inside
the pagination loop with try/catch per event. On exception:
console.error with event type, messageId, timestamp, peer origin,
and error message; continue with next event. lastSyncedAt advances
normally at end of all three passes, so a single failing event no
longer blocks catch-up forever.

Final 'replayed N events' log line reports '(K skipped due to
errors)' suffix when K > 0. Trade-off documented in
docs/systems/federation.md: forward progress prioritized over
strict at-least-once delivery.
2026-04-22 01:46:49 +02:00
Jannis Braun 15e42a7cc1 fix(federation): per-event fault isolation in syncPeerMutationLog (#25)
Replace the batch-level processRelayEvents call with a per-event
loop wrapped in try/catch. On exception: log event type, messageId,
timestamp, peer origin, and the error message; continue to the next
event.

Previously, a single poison-pill event (e.g., UNIQUE conflict from
a malformed relay payload) would throw, be caught by the outer
try/catch, and block lastSyncedAt from advancing — causing every
future activation to retry the same broken window indefinitely.

The final 'replayed N events' log line now reports '(K skipped due
to errors)' when K > 0, surfacing the count to operators. Individual
event failures are logged via console.error with enough context to
debug or replay manually.

Trade-off documented in docs/systems/federation.md: forward progress
of the sync pipeline takes priority over strict at-least-once
delivery. An event that fails to process is lost to the receiver
unless replayed manually.
2026-04-22 01:45:14 +02:00
Jannis Braun 091988c718 Merge branch 'feat/peer-activation-recovery'
Closes #10b (S2S outbox sync recovery after peer state transitions).

Unifies three related bugs under a single on-peer-activation handler:
- Stranded outbox backoff after unreachable→active recovery
- Runtime sync only firing at startup (not on runtime peer re-creation)
- Silent enqueue failure for awaiting_approval / needs_attention peers

Plus:
- Mutation log coverage extended to dm_close/reopen, read_state_update,
  profile_update, and file_rejected (previously bypassed appendMutationLog)
- /api/federation/sync response builder gains serializers for those 5
  event types and a new contextType='profile' branch
- queueOutboxEvent rewritten with explicit per-status handling, mid-call
  race catch, and compile-time exhaustiveness check
- Pre-existing gap fixed: ensurePeered now handles needs_attention
  explicitly instead of falling through to auto-healing handshake

Verified end-to-end on Pi + VM across all three manual integration
scenarios (unreachable recovery, awaiting_approval drain, post-Reset
catch-up via mutation log).
2026-04-22 01:35:32 +02:00
Jannis Braun 911c7e3479 fix(federation): ensurePeered must not auto-heal needs_attention peers
Discovered during live verification of #10b Scenario 3: the switch
in ensurePeered had no case for needs_attention, so it fell through
to performHandshake. Because the /peer/accept idempotent-200-no-update
safeguard covers needs_attention on the inbound side, the remote
returned 200 without writing the new secret, and performHandshake
transitioned the local peer to 'active' on the 200 response — auto-
healing a state that requires admin intervention.

Affected paths: sendCallRelay non-blocking warm-up (used by typing
relay); any future caller of ensurePeered on a needs_attention peer.
Not affected: resolvePendingPeers (already filters on status='pending').

Fix: explicit case 'needs_attention' returning { status: 'rejected',
error }. Caller observes the rejection and does not advance state.
2026-04-22 01:27:07 +02:00
Jannis Braun d5d56db254 docs(federation): document peer-activation recovery
Adds 'Peer Activation Recovery' subsection with call-site roster,
peer-state x outbox-enqueue x recovery matrix, mutation log
coverage table, /api/federation/sync contextType filter values,
and a Known Issues note about the poison-pill edge case.

Also updates stale references to runInitialSyncForNewPeers (removed
in commit 02a1ed7) to point at the unified startupBootstrapSync
path.
2026-04-22 00:57:57 +02:00
Jannis Braun 3fe7500ff0 feat(federation): /sync serializers for 5 new event types
Adds contextType='profile' query branch. Extends the DM-pass
mutation_type IN-clause to include dm_close, dm_reopen,
read_state_update, file_rejected (events with no associated
dm_messages row). Builds channelFederatedIdMap for O(1)
federatedId resolution. Adds serializer branches for all five
new event types (dm_close/dm_reopen, read_state_update,
file_rejected, profile_update) that emit FederationRelayEvent
objects compatible with the existing inbound processors.

Inbound processors (processDmCloseEvent, processDmReopenEvent,
processReadStateUpdateEvent, processProfileUpdateEvent,
processFileRejectedEvent) already exist; sync replay feeds
events through processRelayEvents without any new receive-side
code.
2026-04-22 00:53:46 +02:00
Jannis Braun a23e02339e feat(federation): capture dm_close/reopen/read_state/profile/file_rejected in mutation log
Four event types previously bypassed appendMutationLog, making
them unrecoverable via /api/federation/sync after peer inactivity:
  - queueDmCloseRelay (dm_close, dm_reopen)
  - queueReadStateRelay (read_state_update)
  - handleSizeRejection in federationWorker (file_rejected)
  - profile PATCH route (profile_update) — two call sites,
    one appendMutationLog per profile change (not per target origin)

The /api/federation/sync response builder is extended to
serialize these event types in the next task.
2026-04-22 00:48:49 +02:00
Jannis Braun 250596c0f6 feat(federation): wire onPeerActivated into 8 transition sites
Every code location that sets federation_peers.status='active'
now invokes onPeerActivated(peerId, reason). HTTP handler sites
use fire-and-forget (.catch(log)) so the response isn't blocked
by sync-pull pagination. The worker-internal health-check site
awaits the handler since the tick is already async.

Sites: /peer/initiate, /peer/accept (4 branches), /approval-
requests/:id/approve, health check recovery, ensurePeered/
performHandshake.
2026-04-22 00:39:58 +02:00
Jannis Braun 57d7ca66d3 fix(federation): explicit per-status handling in queueOutboxEvent
Replaces the silent UNIQUE-swallow placeholder branch. Each peer
status has an explicit branch:
  active/pending/unreachable: race-catch — re-fetch peer row and
    enqueue via matchedPeers. Previously skipped silently, losing
    real-time delivery under asymmetric failure.
  awaiting_approval/needs_attention/rejected/revoked: drop with
    logged reason. Mutation log still captures; sync-pull on
    activation replays.
  default: exhaustiveness check (no 'as never' cast) — TypeScript
    enforces that every status value is handled explicitly.
2026-04-22 00:34:43 +02:00
Jannis Braun ae035eba9b fix(federation): remove dead processRelayEvents import
Missed in 02a1ed7. The new sync-pull path in federationPeerActivation.ts
uses a dynamic import of processRelayEvents from routes/federation.js;
the static import in federationWorker.ts is no longer used after
runInitialSyncForNewPeers deletion.
2026-04-22 00:30:54 +02:00
Jannis Braun 02a1ed73f4 refactor(federation): replace runInitialSyncForNewPeers with startupBootstrapSync
The per-peer sync body is now syncPeerMutationLog (in the new
peer-activation module), invoked via onPeerActivated. The startup
path scans for status='active' AND lastSyncedAt=0 and calls the
unified handler for each — same trigger condition as before, unified
code path with runtime transitions.
2026-04-22 00:28:18 +02:00
Jannis Braun cca2245cdf feat(federation): implement onPeerActivated with dedup
Two-invariant handler: resetOutboxBackoff + syncPeerMutationLog.
In-flight map keyed by peerId coalesces concurrent activations —
a second call for a peer whose activation is still running shares
the same promise. Errors are swallowed and logged — the handler
never throws so fire-and-forget callers at HTTP handler sites
are safe.
2026-04-22 00:24:57 +02:00
Jannis Braun 39c43032e8 fix(federation): address Task 3 review — pagination test + polish
- Add pagination-advance test: verifies since=checkpoint on second
  iteration within a pass, and each pass re-seeds since from
  peer.lastSyncedAt (not carried from prior pass).
- Eliminate four peer! non-null assertions by capturing the narrowed
  value in activePeer after the guard.
- Tighten bodyObj type from Record<string, unknown> to a local
  SyncRequestBody type alias.
- Drop the no-op federationRelayEnabled UPDATE in test beforeEach
  (default is already 1 per baseline migration).
2026-04-22 00:21:34 +02:00
Jannis Braun 5c1b42938e feat(federation): implement syncPeerMutationLog
Three-pass pull-sync (dm, friend, profile) from peer's
/api/federation/sync endpoint, paginated. Seeds sinceTimestamp from
peer.lastSyncedAt so a recovered peer pulls only the delta. Updates
lastSyncedAt to Date.now() on full success; leaves it untouched on
transient failure so the next activation retries the same window.
Replaces the body of the soon-to-be-removed runInitialSyncForNewPeers.
2026-04-22 00:15:48 +02:00
Jannis Braun fd37e0c604 feat(federation): implement resetOutboxBackoff
Unconditionally resets nextRetryAt=now and attempts=0 for all
outbox entries of the given peer. No WHERE filter on nextRetryAt —
resetting attempts=0 on already-eligible rows is the correctness fix:
without it, a previously-failed entry keeps stale attempts, and its
next failure uses BACKOFF_SCHEDULE_MS[attempts] (5min to 24h) on a
peer that just recovered.
2026-04-22 00:10:27 +02:00
Jannis Braun 14a96efa58 feat(federation): scaffold peer-activation recovery module
Empty stubs for onPeerActivated, resetOutboxBackoff, syncPeerMutationLog,
and startupBootstrapSync. Functions are filled in by subsequent tasks
following TDD cycles.
2026-04-22 00:06:17 +02:00
Jannis Braun 24c1d5f83d Merge branch 'fix/peer-initiate-202' 2026-04-21 22:44:44 +02:00
Jannis Braun 531104fecc fix(federation): handle 202 in admin /peer/initiate handshake
/peer/initiate checked `response.ok` to decide whether to activate the
local peer. `response.ok` is true for the full 2xx range, so a remote
that returned 202 (queued for admin approval — autoAcceptPeering off
on their side) caused the local peer to flip to `active` while the
remote had us `awaiting_approval`. The split only self-healed when
the remote admin approved and pushed us an `awaiting_approval → active`
override via the peer_approval_requests inbound path.

The auto-peer flow in federationPeering.ts:performHandshake already
had the correct 202 branch: set local status to awaiting_approval,
broadcast federation_peers_changed, surface a pending outcome. Mirror
it here:

- Check response.status === 202 BEFORE the !response.ok branch so the
  fall-through can't reach the activation code.
- Transition local peer to awaiting_approval (not active).
- Broadcast federation_peers_changed so other admin tabs refresh.
- Return 202 with the sanitized peer so the client observes the
  queued state distinctly from both success and failure.

Also added the missing federation_peers_changed broadcast on the
activation (200) path for parity with every other peer-state-change
site in the codebase — it was a pre-existing drift that would leave
sibling admin tabs stale after an initiate. Pattern-aligned with
federationPeering.ts:160 and the rest of routes/federation.ts.

Docs: expanded Phase 1 bullets in docs/systems/federation.md to cover
the 200 / 202 / other non-2xx / network-error branches explicitly and
reference the mirrored auto-peer branch.

Verified: pnpm -r typecheck clean (shared + server), vitest 70/70
pass.

Closes #21 from S2S DM unification backlog.
2026-04-21 22:41:44 +02:00
Jannis Braun 43f0c40685 Merge branch 'feat/federation-cleanup-sweep' 2026-04-21 22:36:02 +02:00
Jannis Braun 95c9666213 build(server): drop stale project reference to @backspace/shared
The three typecheck failures in ws/events.ts for DmCallUndeliverable{Failure,Reason}
and 'dm_call_undeliverable' looked like missing exports from @backspace/shared,
but the types are fully defined and exported in packages/shared/src/types.ts
(lines 361, 367, 429). The real cause was architectural.

packages/server/tsconfig.json declared:

    "references": [{ "path": "../shared" }]

TypeScript project references make the dependent project's typecheck consume
the referenced project's *build output* (dist/*.d.ts), not its sources. The
shared package's dist is gitignored, is not rebuilt by any script before
`pnpm -r typecheck` or `pnpm --filter @backspace/server typecheck`, and the
referenced-project-stale failure mode surfaces as TS6305 in the server, or
(when the dist is present but older) as "no exported member" for types
that were added after the last shared build. The #16 merge (which added
DmCallUndeliverableFailure/Reason and the dm_call_undeliverable event
variant) landed the types in src but the local dist was never refreshed,
so server's typecheck started failing against the stale .d.ts.

packages/web/tsconfig.json already resolved @backspace/shared via
`moduleResolution: "bundler"` + the package.json `exports` field, which
points directly at ./src/types.ts. That path has no build-ordering
dependency, never goes stale, and already typecheck-passes cleanly.

Fix: remove the server-side project reference so server matches web's
bundler-style resolution. Server still emits its own dist on `tsc` (its
rootDir confines emission to its own src/); the runtime already reads
TS directly via tsx, so nothing in the dev or prod run path changes.
The Dockerfile/root build scripts that build shared explicitly are also
unaffected.

Verified: `pnpm -r typecheck` passes for shared and server; standalone
`pnpm exec tsc --noEmit` in packages/web passes; `pnpm build` completes
all three packages.

Fixes #24 (pre-existing typecheck failure on main).
2026-04-21 22:27:51 +02:00
Jannis Braun 4b398cf45a types(web): unify FederationPeer with shared type
packages/web/src/api/client.ts declared a local FederationPeer that had
drifted from @backspace/shared: it loosened `status` to `string` (losing
the exhaustive 7-value union) and widened `consecutiveFailures` and
`lastSyncedAt` to `number | null`. The latter two are spurious — the
server never returns null for either — and `status: string` defeated
the compiler's ability to flag a missed case when `rejected`,
`awaiting_approval`, or `needs_attention` were added over the course
of the auto-peering / approval-queue / outbox-auth-failure-recovery
work.

Replace the local interface with a re-export of the shared type. All
three status switches in FederationPanel.tsx (peerStatusColor,
peerStatusDotColor, peerStatusLabel) and the StatusFilter union were
already exhaustive over the 7 values, so no behaviour change is
needed — the re-export just pins the compile-time contract.

web tsc --noEmit is clean after the swap.

Follow-up #23 from S2S DM unification backlog.
2026-04-21 22:21:47 +02:00
Jannis Braun 010aa7f7ca fix(schema): normalize federation_peers.consecutive_failures to NOT NULL
The column was `integer DEFAULT 0` (nullable) since the initial schema.
Counters should not be nullable — the semantics are a count, not an
optional measurement. `consecutive_auth_failures` (added later) was
correctly declared NOT NULL; tightening `consecutive_failures` to match
removes the drift and eliminates the "|null" burden everywhere the value
is read.

SQLite does not support in-place ALTER … SET NOT NULL, so drizzle-kit
cannot auto-generate this. The manual migration uses the standard
SQLite recreate pattern (new table + INSERT SELECT + DROP + RENAME +
recreate index) under `PRAGMA defer_foreign_keys = ON` so the existing
federation_outbox → federation_peers FK survives the swap. The COPY
step coalesces any hypothetical NULL to 0 defensively; live probes on
both test instances (nova, orbit) showed zero NULL rows so no
actual backfill is required.

Verified by applying the full migration chain against a copy of the VM's
live DB: column ends as `notnull=1 dflt=0`, the peer row is preserved,
the unique index on origin is recreated, NULL inserts are rejected, and
`PRAGMA foreign_key_check` reports no violations.

Server `SanitizedPeer.consecutiveFailures` tightened to `number` to
match the new drizzle inference and the shared `FederationPeer` shape.

Follow-up #22 from S2S DM unification backlog.
2026-04-21 22:21:06 +02:00
Jannis Braun 0d74d1d112 perf(federation-worker): tighten health-check cadence to 15 min
HEALTH_CHECK_INTERVAL_MS was 1 h, but ROTATION_GRACE_PERIOD_MS is 15 min.
Phase skew between two peers' health-check ticks could stretch rotation
finalization desync up to ~1 h, during which signatures from the already-
finalized side verify against the other side's primary-only secret (grace
has expired; verifyPeerSignature stops trying the pending secret). With
AUTH_FAILURE_THRESHOLD = 5 and the existing backoff schedule, this
occasionally tripped legitimate rotations into needs_attention.

Setting the interval to 15 min (= ROTATION_GRACE_PERIOD_MS) guarantees a
finalization tick fires within one grace window on each side, so the
cross-verification window where one peer signs with NEW while the other
still treats NEW as pending cannot outlast the grace period.

Per-tick cost is negligible for the worker's steady state: the only
network fetches are per-active-peer /peer/rotate calls when the 90-day
rotation interval hits (rare) and per-unreachable-peer /instance/info
health pings (bounded by outage count). Going lower than 15 min would
reduce the residual desync but increase tick overhead with diminishing
returns; 15 min is the grace-period-aligned value that the original spec
("runs hourly") deviated from without justification.

Follow-up #20 from S2S DM unification backlog; reduces #19 false-positive
rate (outbox auth-failure transition) on legitimate rotations.
2026-04-21 22:16:09 +02:00
Jannis Braun 9400189a8d Merge branch 'feat/outbox-auth-failure-recovery'
Replace the federation outbox worker's 401/403 wipe-and-rehandshake loop
with bounded retry (AUTH_FAILURE_THRESHOLD=5, ~21.5 min backoff window) and
a new `needs_attention` peer state. Surfaces persistent HMAC desync to
admins via a first-class 'Reset peering' action instead of the prior
silent loop.

Closes backlog item #19. Security invariants verified on live infra
(Pi+VM):

- hmac_secret is NEVER wiped in response to a network-observed 401/403
  (Task 5 removes the wipe; Task 7 extends the /peer/accept idempotent-
  200-no-update safeguard to cover needs_attention peers).
- Auth failures increment only consecutive_auth_failures, never the
  network counter consecutive_failures (Task 5 splits
  handleOutboxDeliveryFailure → applyOutboxEntryBackoff).
- Transition occurs at exactly 5 consecutive 401/403 responses; below
  threshold, entries get backoff but state is preserved; above, peer
  flips to needs_attention, affected users get federation_peer_rejected
  WS with 'Federation trust broken — admin must reset peering'.
- /peer/accept safeguard confirmed against attacker curl probe on the
  live Pi instance while in needs_attention — forged-secret request
  returned 200-no-update, local hmac_secret unchanged.
- Legitimate rotation (Scenario B) does not false-positive: both sides
  capture pending secret, auth_failures stays at 0, DM delivers cleanly
  during grace period.

Four follow-up backlog items discovered during the work:
#20 health-check cadence tightening (15-min grace vs 1-hour tick)
#21 /peer/initiate 202 handling
#22 consecutive_failures nullability normalization
#23 unify client-side FederationPeer with shared type

Spec: internal notes
Plan: internal notes
2026-04-21 22:04:31 +02:00
Jannis Braun e236d0730e docs(systems): document needs_attention state and Reset peering action 2026-04-21 21:11:17 +02:00
Jannis Braun 64d231d782 audit(federation): verify no parallel hmac-wipe-on-401 paths
Task 11 due-diligence audit for #19. Checked all signed-fetch sites in
packages/server/src for 4xx-branch mutations of federationPeers.hmacSecret
or federationPeers.status:

- sendCallRelay (federationOutbox.ts): on 4xx returns post_failed; on
  5xx/network returns peer_transient_failure. No peer-state mutation.
- sendTypingRelay (federationOutbox.ts): delegates to sendCallRelay with
  peeringTimeoutMs:0 (fire-and-forget). No peer-state mutation.
- cleanupExpiredApprovalRequests (storageJanitor.ts): sends denial,
  only deletes peerApprovalRequests row on success. No federationPeers
  mutation.
- DELETE /identity (users.ts): per-origin cleanup; only deletes local
  userFederationRegistry on success, never touches federationPeers.
- denyApprovalRequest (federation.ts): requires 2xx from remote before
  inserting/updating a rejected peer row. Admin-driven, not wipe.
- POST /peers/:id/rotate (federation.ts): mutates pendingHmacSecret only
  on 2xx; returns 502 on 4xx without state change.
- Auto-rotation in federationWorker.ts: same 2xx-gated pattern as manual
  rotate.
- Unreachable-recovery health check: only promotes to active on 2xx.
- ensurePeered/performHandshake (federationPeering.ts): on 403 with
  PEERING_REQUIRES_APPROVAL sets status='rejected' (explicit, not a
  HMAC-mismatch wipe); on other 4xx/5xx only deletes the row if it was
  a freshly created placeholder (existingPeerId falsy). Pre-existing
  peers are untouched.

Only federationWorker.ts:269 mutates HMAC-related state in response to
401/403, and that path was rewritten in Task 5 to use
evaluateAuthFailure and transition to needs_attention. No additional
handlers require the bounded-retry refactor.
2026-04-21 21:07:24 +02:00
Jannis Braun 5a1e354ae1 feat(federation-ui): add needs_attention pill and Reset peering action
- peerStatusLabel/Color/DotColor gain a 'needs_attention' case (rose).
- StatusFilter row gains 'Needs Attention' toggle.
- PeerRow hides Rotate/Revoke and shows 'Reset Peering' when status is
  needs_attention, plus an Auth Failures stat.
- Parent panel routes 'reset' through a ConfirmDialog (danger variant)
  that spells out the destructive nature and the out-of-band re-peer step.
- Client FederationPeer interface gains consecutiveAuthFailures (Task 2
  extended the shared type but the web client's local mirror was stale).

Codifies the manual 'delete both sides, re-peer' workaround as a
first-class admin action.
2026-04-21 21:02:49 +02:00
Jannis Braun 8a084b0652 feat(api): add federation.resetPeer client method 2026-04-21 20:58:59 +02:00
Jannis Braun 0c2864a3d2 feat(federation): add POST /api/federation/peers/:id/reset
Admin-only endpoint for recovering from needs_attention. Deletes the
local peer row; FK cascade removes queued outbox entries. Gated to
peers in needs_attention to prevent accidental resets of healthy
peerings (use /peers/:id for revoke on active peers).

Also extends the Task 8.5 test mock of '../db/index.js' to re-export
`schema`. federation.ts imports `schema` from the re-export alongside
`getDb`; the previous mock only exposed `getDb`, causing the route
handler to blow up with 500s before reaching any assertion. This is
a scaffolding fix — no test assertions were changed.
2026-04-21 20:56:26 +02:00
Jannis Braun 54f32657e5 test(federation): add route-level tests for POST /peers/:id/reset
Four cases from the spec's testing strategy: 404 on missing peer, 400 on
wrong status, 403 for non-admin, and successful delete including FK
cascade of queued outbox entries. Introduces a minimal Fastify-inject
harness for route testing — previously the codebase had only pure-function
unit tests under utils/.

Tests intentionally FAIL at this commit — Task 8 will add the handler and
close the loop.
2026-04-21 20:52:14 +02:00
Jannis Braun ceaa08b4bd fix(federation): extend /peer/accept safeguard to cover needs_attention
Unauthenticated /peer/accept must not overwrite hmac_secret for peers in
needs_attention, same as active. needs_attention means 'auth trust broke
and we don't know why' — letting an unauthenticated request flip it back
would reintroduce a path for silent HMAC rotation via the outbox-401 loop
the rest of #19 closes. Legitimate recovery is the admin 'Reset peering'
action (next task).
2026-04-21 20:49:19 +02:00
Jannis Braun 48dbe32a69 fix(federation-worker): reset consecutive_auth_failures on successful delivery
Pairs with the new 401/403 handler — a 2xx relay confirms HMAC trust is
healthy so the counter should clear. Mirrors the existing
consecutive_failures reset for network-layer health.
2026-04-21 20:46:52 +02:00
Jannis Braun 012e489bc7 fix(federation-worker): auth failures must not increment consecutive_failures
Code review of the previous commit found that the backoff branch of the
new 401/403 handler delegated to handleOutboxDeliveryFailure, which
double-dips by also incrementing consecutive_failures (the network-layer
counter that drives the 'unreachable' transition at threshold 10). Per the
design spec §State Machine Changes → Reset logic, auth failures must
increment consecutive_auth_failures ONLY.

Split handleOutboxDeliveryFailure into:
- applyOutboxEntryBackoff: just the per-entry backoff update (safe to call
  from the auth-failure path)
- handleOutboxDeliveryFailure: entry backoff + peer's consecutive_failures
  bump (network-error path only)

Also adds a console.warn to the backoff branch so operators can diagnose
clock-skew and rotation-grace incidents before the peer hits the terminal
threshold.

Part of backlog #19.
2026-04-21 20:45:33 +02:00
Jannis Braun e5afd376d2 fix(federation-worker): replace 401/403 wipe-and-rehandshake with bounded retry
The previous handler (commit ce33ccf + its 403 extension) wiped hmac_secret
and reset peer status to 'pending' on any 401/403 from an active peer. This
collapsed three distinct failure modes — transient clock skew, legitimate
split-brain, active MITM attempt — into "silently establish new trust
immediately." The remote's /peer/accept idempotent-200-no-update safeguard
then prevented the re-handshake from actually working, producing a 1-req/sec
loop observed during backlog #16 verification.

New behavior: increment consecutive_auth_failures, apply backoff to outbox
entries. At AUTH_FAILURE_THRESHOLD (5) transition to needs_attention,
preserve hmac_secret, surface delivery-impossible to affected users,
notify admins. Secret is NEVER wiped in response to a network-observed
401/403.

Part of backlog #19.
2026-04-21 20:39:37 +02:00
Jannis Braun 695ea0849d refactor(federation-worker): extract buildContextMapForPeer helper
Pure refactor — will be reused by the needs_attention transition handler.
No behavior change.
2026-04-21 20:35:51 +02:00
Jannis Braun 617d71ab4b feat(federation): add evaluateAuthFailure decision function
Pure function deciding whether the next 401/403 from an active peer
triggers backoff or a transition to needs_attention. Threshold = 5,
corresponding to ~21.5 min of the existing BACKOFF_SCHEDULE_MS.
2026-04-21 20:32:38 +02:00
Jannis Braun 1df25737b6 types: align FederationPeer union with actual server statuses
Adds 'rejected', 'awaiting_approval', and 'needs_attention' to the status
union, plus the consecutiveAuthFailures / autoRotateIntervalDays /
secretRotatedAt / rotationInProgress fields the UI already reads.
2026-04-21 20:29:14 +02:00
Jannis Braun 0d3343ab6d feat(schema): add consecutive_auth_failures column on federation_peers
Tracks HMAC-failure count separately from consecutive_failures (network
errors). Auth failures and network failures have different resolution
paths; mixing them would let a single successful retry after a network
blip mask real secret desync.

Part of backlog #19 — outbox auth-failure recovery.

Note: drizzle-kit generated the 0003 SQL with an unexpected CREATE TABLE
for peer_approval_requests because the 0002 migration was authored
manually without a corresponding 0002_snapshot.json (see 14041e9). The
generated SQL has been trimmed to the single intended ALTER TABLE. The
regenerated 0003_snapshot.json correctly reflects the full current
schema, so future migrations will diff cleanly.
2026-04-21 18:40:39 +02:00
Jannis Braun 85668d4da8 Merge branch 'feat/call-relay-auto-peering'
Closes backlog #16. Implements sendCallRelay/sendTypingRelay auto-peering
and caller-facing dm_call_undeliverable failure surface.

See internal notes
and internal notes
for the full design + implementation plan.
2026-04-21 16:27:50 +02:00
Jannis Braun 9cdc5921d9 docs: clarify that livekit_unavailable is emitted from sendFederatedCallStart, not sendCallRelay 2026-04-21 14:08:31 +02:00
Jannis Braun 740dae298d docs: document call-relay auto-peering and dm_call_undeliverable surface 2026-04-21 14:05:59 +02:00
Jannis Braun 6a7d1fb38c feat(web): handle dm_call_undeliverable — toast + tear down outgoing call on terminal 2026-04-21 14:02:08 +02:00
Jannis Braun 44a44163af polish(server): consolidate federation_peers queries, narrow targetedPeers map, use return values from Promise.all to drop non-null assertion 2026-04-21 14:00:29 +02:00
Jannis Braun f83c2af357 feat(server): aggregate call-start failures into dm_call_undeliverable
sendFederatedCallStart now collects per-targeted-peer results and emits
a single dm_call_undeliverable event to the caller when any targeted
peer relay fails. Destroys the local ring room when no plausible
recipient remains (no targeted success + no connected local ringee).

LiveKit pre-flight also emits via this path with reason
'livekit_unavailable' instead of a silent console.warn, closing the
60s hang for unconfigured instances.

Guards against phantom toasts when the caller cancels mid-race by
checking getRoom() before emitting.
2026-04-21 13:53:34 +02:00
Jannis Braun 53483d6981 polish(server): align sendCallRelay timeout message with codebase convention; use .then on sendTypingRelay fire-and-forget 2026-04-21 13:49:20 +02:00
Jannis Braun 21f220739c feat(server): sendCallRelay auto-peers on demand, typing passes peeringTimeoutMs:0
sendCallRelay now returns CallRelayResult with a typed reason on failure.
When the peer is not already active (or unreachable), runs a racePeering
against CALL_PEERING_TIMEOUT_MS (3s). Background handshake is not aborted
on race loss — next attempt succeeds.

sendTypingRelay passes peeringTimeoutMs:0 so typing never blocks on a
handshake; instead a warm-up ensurePeered runs in the background for any
non-active peer so the NEXT relay (message, call, or typing) benefits.
2026-04-21 13:45:29 +02:00
Jannis Braun 4ddb09edf1 fix(server): racePeering normalizes handshake rejections and only warns on timeout win
Addresses code review on b22a7bd: (1) a rejected handshake now returns
{ status: 'failed', error } instead of throwing, keeping the structured
contract; (2) the "background handshake" warn only fires when the
timeout arm wins — not when the handshake is itself the race winner by
rejection. Timing tests migrated to vi.useFakeTimers for determinism.
Regression test added for the handshake-wins-by-rejection case.
2026-04-21 13:41:44 +02:00
Jannis Braun b22a7bd0c6 feat(server): add racePeering helper with tests
Exports `racePeering(origin, timeoutMs, ensurePeeredFn?)` that races
`ensurePeered` against a deadline. On timeout, the background handshake
continues (warming the peer for the next attempt) and a warn-logged
.catch() prevents unhandledRejection. Injectable `ensurePeeredFn` param
enables full DI in tests without mocking module internals.
2026-04-21 13:37:30 +02:00
Jannis Braun 1be544bbd5 feat(shared): add dm_call_undeliverable event type 2026-04-21 13:33:52 +02:00
Jannis Braun 852e3657f9 fix: dedup federation membership events by (sourceInstance, messageId)
processMemberAddEvent, processMemberRemoveEvent, and processOwnershipTransferEvent inserted system messages unconditionally. Outbox retries and initial-sync replays (triggered whenever an admin re-approves a peering request, which recreates the peer row with lastSyncedAt=0) duplicated the system message on every delivery. Each new snowflake ID exceeded the user's last_read_message_id, flipping the channel back to unread after every deploy.

Processors now short-circuit on a matching (source_instance, source_message_id) row, and persist those fields when inserting. processMemberAddEvent emits the tagged system message in both bootstrap and incremental paths so bootstrap replays don't fall through and insert a second one; the bootstrap's dm_channel_created broadcast carries that message as lastMessage so sidebar previews and unread anchors agree across instances.
2026-04-21 01:06:43 +02:00
Jannis Braun aeebf79feb fix: federated friend request routed to wrong user with same name
When two instances each have a native user with the same username, the
Add Friend search card for the federated one sent its request to the
local namesake instead of the intended remote user.

Root cause: `isNative = !homeUserId` in socialStore's searchUsers and
loadFriends dedup. The server backfills native users' homeUserId to
their own id so federation tier-1 lookups succeed, so `homeUserId` is
set on natives too. Only `homeInstance` distinguishes native (null)
from replicated stubs. With the wrong check, no entry was ever "native"
and the home-origin stub of the remote user was kept over the true
native record — leaving `_instanceOrigin=''`, which caused the Send
button handler to drop the domain suffix and POST to the home API,
where "nova" resolved to a completely different local user.

Also fixes loadRequests dedup to prefer the target-native record so the
search card correctly flips to "Request Pending" after sending.
2026-04-21 00:36:15 +02:00
Jannis Braun 3d8709d20a feat: real-time Federation panel updates via WS events
Added federation_peers_changed (no-payload signal) broadcast from every
peer state mutation, and federation_approval_request_received when a new
approval request is queued. Client subscribes via onFederationPeersChanged
callback registry. FederationPanel and PendingApprovals debounce-refetch
on any event. sendToAdmins helper broadcasts only to admin users.
2026-04-20 18:28:10 +02:00
Jannis Braun 6afad97bd1 fix: revert outgoing peering blocks — autoAcceptPeering only gates incoming
autoAcceptPeering means 'don't accept peering initiated by others', not
'don't initiate peering ourselves'. Two checks were incorrectly blocking
outgoing peering when auto-accept was off:

1. ensurePeered() refused to auto-initiate — reverted. When a local user
   sends a DM, the server should initiate peering. The remote's
   peer/accept decides whether to accept or queue.

2. queueOutboxEvent() refused to create placeholders — reverted. The
   outbox needs placeholders to queue entries. Without them, DM relay
   silently fails.
2026-04-20 18:08:05 +02:00
Jannis Braun 165fda44a3 fix: accept incoming handshake for awaiting_approval peers to break approval ping-pong
When both instances have autoAcceptPeering off, the approval flow
ping-ponged indefinitely. Admin A approves → handshakes to B → B
queues (202) → A's peer becomes awaiting_approval. Admin B approves →
handshakes to A → but A's gate only matched 'pending', not
'awaiting_approval', so it re-queued instead of accepting.

Now the gate matches both 'pending' and 'awaiting_approval'. When the
second admin approves and handshakes back, the first instance recognizes
its admin already approved and accepts — completing the peering.
2026-04-20 18:01:47 +02:00