Commit Graph
1661 Commits
Author SHA1 Message Date
Jannis Braun 0d74d1d112 perf(federation-worker): tighten health-check cadence to 15 min
HEALTH_CHECK_INTERVAL_MS was 1 h, but ROTATION_GRACE_PERIOD_MS is 15 min.
Phase skew between two peers' health-check ticks could stretch rotation
finalization desync up to ~1 h, during which signatures from the already-
finalized side verify against the other side's primary-only secret (grace
has expired; verifyPeerSignature stops trying the pending secret). With
AUTH_FAILURE_THRESHOLD = 5 and the existing backoff schedule, this
occasionally tripped legitimate rotations into needs_attention.

Setting the interval to 15 min (= ROTATION_GRACE_PERIOD_MS) guarantees a
finalization tick fires within one grace window on each side, so the
cross-verification window where one peer signs with NEW while the other
still treats NEW as pending cannot outlast the grace period.

Per-tick cost is negligible for the worker's steady state: the only
network fetches are per-active-peer /peer/rotate calls when the 90-day
rotation interval hits (rare) and per-unreachable-peer /instance/info
health pings (bounded by outage count). Going lower than 15 min would
reduce the residual desync but increase tick overhead with diminishing
returns; 15 min is the grace-period-aligned value that the original spec
("runs hourly") deviated from without justification.

Follow-up #20 from S2S DM unification backlog; reduces #19 false-positive
rate (outbox auth-failure transition) on legitimate rotations.
2026-04-21 22:16:09 +02:00
Jannis Braun 9400189a8d Merge branch 'feat/outbox-auth-failure-recovery'
Replace the federation outbox worker's 401/403 wipe-and-rehandshake loop
with bounded retry (AUTH_FAILURE_THRESHOLD=5, ~21.5 min backoff window) and
a new `needs_attention` peer state. Surfaces persistent HMAC desync to
admins via a first-class 'Reset peering' action instead of the prior
silent loop.

Closes backlog item #19. Security invariants verified on live infra
(Pi+VM):

- hmac_secret is NEVER wiped in response to a network-observed 401/403
  (Task 5 removes the wipe; Task 7 extends the /peer/accept idempotent-
  200-no-update safeguard to cover needs_attention peers).
- Auth failures increment only consecutive_auth_failures, never the
  network counter consecutive_failures (Task 5 splits
  handleOutboxDeliveryFailure → applyOutboxEntryBackoff).
- Transition occurs at exactly 5 consecutive 401/403 responses; below
  threshold, entries get backoff but state is preserved; above, peer
  flips to needs_attention, affected users get federation_peer_rejected
  WS with 'Federation trust broken — admin must reset peering'.
- /peer/accept safeguard confirmed against attacker curl probe on the
  live Pi instance while in needs_attention — forged-secret request
  returned 200-no-update, local hmac_secret unchanged.
- Legitimate rotation (Scenario B) does not false-positive: both sides
  capture pending secret, auth_failures stays at 0, DM delivers cleanly
  during grace period.

Four follow-up backlog items discovered during the work:
#20 health-check cadence tightening (15-min grace vs 1-hour tick)
#21 /peer/initiate 202 handling
#22 consecutive_failures nullability normalization
#23 unify client-side FederationPeer with shared type

Spec: internal notes
Plan: internal notes
2026-04-21 22:04:31 +02:00
Jannis Braun e236d0730e docs(systems): document needs_attention state and Reset peering action 2026-04-21 21:11:17 +02:00
Jannis Braun 64d231d782 audit(federation): verify no parallel hmac-wipe-on-401 paths
Task 11 due-diligence audit for #19. Checked all signed-fetch sites in
packages/server/src for 4xx-branch mutations of federationPeers.hmacSecret
or federationPeers.status:

- sendCallRelay (federationOutbox.ts): on 4xx returns post_failed; on
  5xx/network returns peer_transient_failure. No peer-state mutation.
- sendTypingRelay (federationOutbox.ts): delegates to sendCallRelay with
  peeringTimeoutMs:0 (fire-and-forget). No peer-state mutation.
- cleanupExpiredApprovalRequests (storageJanitor.ts): sends denial,
  only deletes peerApprovalRequests row on success. No federationPeers
  mutation.
- DELETE /identity (users.ts): per-origin cleanup; only deletes local
  userFederationRegistry on success, never touches federationPeers.
- denyApprovalRequest (federation.ts): requires 2xx from remote before
  inserting/updating a rejected peer row. Admin-driven, not wipe.
- POST /peers/:id/rotate (federation.ts): mutates pendingHmacSecret only
  on 2xx; returns 502 on 4xx without state change.
- Auto-rotation in federationWorker.ts: same 2xx-gated pattern as manual
  rotate.
- Unreachable-recovery health check: only promotes to active on 2xx.
- ensurePeered/performHandshake (federationPeering.ts): on 403 with
  PEERING_REQUIRES_APPROVAL sets status='rejected' (explicit, not a
  HMAC-mismatch wipe); on other 4xx/5xx only deletes the row if it was
  a freshly created placeholder (existingPeerId falsy). Pre-existing
  peers are untouched.

Only federationWorker.ts:269 mutates HMAC-related state in response to
401/403, and that path was rewritten in Task 5 to use
evaluateAuthFailure and transition to needs_attention. No additional
handlers require the bounded-retry refactor.
2026-04-21 21:07:24 +02:00
Jannis Braun 5a1e354ae1 feat(federation-ui): add needs_attention pill and Reset peering action
- peerStatusLabel/Color/DotColor gain a 'needs_attention' case (rose).
- StatusFilter row gains 'Needs Attention' toggle.
- PeerRow hides Rotate/Revoke and shows 'Reset Peering' when status is
  needs_attention, plus an Auth Failures stat.
- Parent panel routes 'reset' through a ConfirmDialog (danger variant)
  that spells out the destructive nature and the out-of-band re-peer step.
- Client FederationPeer interface gains consecutiveAuthFailures (Task 2
  extended the shared type but the web client's local mirror was stale).

Codifies the manual 'delete both sides, re-peer' workaround as a
first-class admin action.
2026-04-21 21:02:49 +02:00
Jannis Braun 8a084b0652 feat(api): add federation.resetPeer client method 2026-04-21 20:58:59 +02:00
Jannis Braun 0c2864a3d2 feat(federation): add POST /api/federation/peers/:id/reset
Admin-only endpoint for recovering from needs_attention. Deletes the
local peer row; FK cascade removes queued outbox entries. Gated to
peers in needs_attention to prevent accidental resets of healthy
peerings (use /peers/:id for revoke on active peers).

Also extends the Task 8.5 test mock of '../db/index.js' to re-export
`schema`. federation.ts imports `schema` from the re-export alongside
`getDb`; the previous mock only exposed `getDb`, causing the route
handler to blow up with 500s before reaching any assertion. This is
a scaffolding fix — no test assertions were changed.
2026-04-21 20:56:26 +02:00
Jannis Braun 54f32657e5 test(federation): add route-level tests for POST /peers/:id/reset
Four cases from the spec's testing strategy: 404 on missing peer, 400 on
wrong status, 403 for non-admin, and successful delete including FK
cascade of queued outbox entries. Introduces a minimal Fastify-inject
harness for route testing — previously the codebase had only pure-function
unit tests under utils/.

Tests intentionally FAIL at this commit — Task 8 will add the handler and
close the loop.
2026-04-21 20:52:14 +02:00
Jannis Braun ceaa08b4bd fix(federation): extend /peer/accept safeguard to cover needs_attention
Unauthenticated /peer/accept must not overwrite hmac_secret for peers in
needs_attention, same as active. needs_attention means 'auth trust broke
and we don't know why' — letting an unauthenticated request flip it back
would reintroduce a path for silent HMAC rotation via the outbox-401 loop
the rest of #19 closes. Legitimate recovery is the admin 'Reset peering'
action (next task).
2026-04-21 20:49:19 +02:00
Jannis Braun 48dbe32a69 fix(federation-worker): reset consecutive_auth_failures on successful delivery
Pairs with the new 401/403 handler — a 2xx relay confirms HMAC trust is
healthy so the counter should clear. Mirrors the existing
consecutive_failures reset for network-layer health.
2026-04-21 20:46:52 +02:00
Jannis Braun 012e489bc7 fix(federation-worker): auth failures must not increment consecutive_failures
Code review of the previous commit found that the backoff branch of the
new 401/403 handler delegated to handleOutboxDeliveryFailure, which
double-dips by also incrementing consecutive_failures (the network-layer
counter that drives the 'unreachable' transition at threshold 10). Per the
design spec §State Machine Changes → Reset logic, auth failures must
increment consecutive_auth_failures ONLY.

Split handleOutboxDeliveryFailure into:
- applyOutboxEntryBackoff: just the per-entry backoff update (safe to call
  from the auth-failure path)
- handleOutboxDeliveryFailure: entry backoff + peer's consecutive_failures
  bump (network-error path only)

Also adds a console.warn to the backoff branch so operators can diagnose
clock-skew and rotation-grace incidents before the peer hits the terminal
threshold.

Part of backlog #19.
2026-04-21 20:45:33 +02:00
Jannis Braun e5afd376d2 fix(federation-worker): replace 401/403 wipe-and-rehandshake with bounded retry
The previous handler (commit ce33ccf + its 403 extension) wiped hmac_secret
and reset peer status to 'pending' on any 401/403 from an active peer. This
collapsed three distinct failure modes — transient clock skew, legitimate
split-brain, active MITM attempt — into "silently establish new trust
immediately." The remote's /peer/accept idempotent-200-no-update safeguard
then prevented the re-handshake from actually working, producing a 1-req/sec
loop observed during backlog #16 verification.

New behavior: increment consecutive_auth_failures, apply backoff to outbox
entries. At AUTH_FAILURE_THRESHOLD (5) transition to needs_attention,
preserve hmac_secret, surface delivery-impossible to affected users,
notify admins. Secret is NEVER wiped in response to a network-observed
401/403.

Part of backlog #19.
2026-04-21 20:39:37 +02:00
Jannis Braun 695ea0849d refactor(federation-worker): extract buildContextMapForPeer helper
Pure refactor — will be reused by the needs_attention transition handler.
No behavior change.
2026-04-21 20:35:51 +02:00
Jannis Braun 617d71ab4b feat(federation): add evaluateAuthFailure decision function
Pure function deciding whether the next 401/403 from an active peer
triggers backoff or a transition to needs_attention. Threshold = 5,
corresponding to ~21.5 min of the existing BACKOFF_SCHEDULE_MS.
2026-04-21 20:32:38 +02:00
Jannis Braun 1df25737b6 types: align FederationPeer union with actual server statuses
Adds 'rejected', 'awaiting_approval', and 'needs_attention' to the status
union, plus the consecutiveAuthFailures / autoRotateIntervalDays /
secretRotatedAt / rotationInProgress fields the UI already reads.
2026-04-21 20:29:14 +02:00
Jannis Braun 0d3343ab6d feat(schema): add consecutive_auth_failures column on federation_peers
Tracks HMAC-failure count separately from consecutive_failures (network
errors). Auth failures and network failures have different resolution
paths; mixing them would let a single successful retry after a network
blip mask real secret desync.

Part of backlog #19 — outbox auth-failure recovery.

Note: drizzle-kit generated the 0003 SQL with an unexpected CREATE TABLE
for peer_approval_requests because the 0002 migration was authored
manually without a corresponding 0002_snapshot.json (see 14041e9). The
generated SQL has been trimmed to the single intended ALTER TABLE. The
regenerated 0003_snapshot.json correctly reflects the full current
schema, so future migrations will diff cleanly.
2026-04-21 18:40:39 +02:00
Jannis Braun 85668d4da8 Merge branch 'feat/call-relay-auto-peering'
Closes backlog #16. Implements sendCallRelay/sendTypingRelay auto-peering
and caller-facing dm_call_undeliverable failure surface.

See internal notes
and internal notes
for the full design + implementation plan.
2026-04-21 16:27:50 +02:00
Jannis Braun 9cdc5921d9 docs: clarify that livekit_unavailable is emitted from sendFederatedCallStart, not sendCallRelay 2026-04-21 14:08:31 +02:00
Jannis Braun 740dae298d docs: document call-relay auto-peering and dm_call_undeliverable surface 2026-04-21 14:05:59 +02:00
Jannis Braun 6a7d1fb38c feat(web): handle dm_call_undeliverable — toast + tear down outgoing call on terminal 2026-04-21 14:02:08 +02:00
Jannis Braun 44a44163af polish(server): consolidate federation_peers queries, narrow targetedPeers map, use return values from Promise.all to drop non-null assertion 2026-04-21 14:00:29 +02:00
Jannis Braun f83c2af357 feat(server): aggregate call-start failures into dm_call_undeliverable
sendFederatedCallStart now collects per-targeted-peer results and emits
a single dm_call_undeliverable event to the caller when any targeted
peer relay fails. Destroys the local ring room when no plausible
recipient remains (no targeted success + no connected local ringee).

LiveKit pre-flight also emits via this path with reason
'livekit_unavailable' instead of a silent console.warn, closing the
60s hang for unconfigured instances.

Guards against phantom toasts when the caller cancels mid-race by
checking getRoom() before emitting.
2026-04-21 13:53:34 +02:00
Jannis Braun 53483d6981 polish(server): align sendCallRelay timeout message with codebase convention; use .then on sendTypingRelay fire-and-forget 2026-04-21 13:49:20 +02:00
Jannis Braun 21f220739c feat(server): sendCallRelay auto-peers on demand, typing passes peeringTimeoutMs:0
sendCallRelay now returns CallRelayResult with a typed reason on failure.
When the peer is not already active (or unreachable), runs a racePeering
against CALL_PEERING_TIMEOUT_MS (3s). Background handshake is not aborted
on race loss — next attempt succeeds.

sendTypingRelay passes peeringTimeoutMs:0 so typing never blocks on a
handshake; instead a warm-up ensurePeered runs in the background for any
non-active peer so the NEXT relay (message, call, or typing) benefits.
2026-04-21 13:45:29 +02:00
Jannis Braun 4ddb09edf1 fix(server): racePeering normalizes handshake rejections and only warns on timeout win
Addresses code review on b22a7bd: (1) a rejected handshake now returns
{ status: 'failed', error } instead of throwing, keeping the structured
contract; (2) the "background handshake" warn only fires when the
timeout arm wins — not when the handshake is itself the race winner by
rejection. Timing tests migrated to vi.useFakeTimers for determinism.
Regression test added for the handshake-wins-by-rejection case.
2026-04-21 13:41:44 +02:00
Jannis Braun b22a7bd0c6 feat(server): add racePeering helper with tests
Exports `racePeering(origin, timeoutMs, ensurePeeredFn?)` that races
`ensurePeered` against a deadline. On timeout, the background handshake
continues (warming the peer for the next attempt) and a warn-logged
.catch() prevents unhandledRejection. Injectable `ensurePeeredFn` param
enables full DI in tests without mocking module internals.
2026-04-21 13:37:30 +02:00
Jannis Braun 1be544bbd5 feat(shared): add dm_call_undeliverable event type 2026-04-21 13:33:52 +02:00
Jannis Braun 852e3657f9 fix: dedup federation membership events by (sourceInstance, messageId)
processMemberAddEvent, processMemberRemoveEvent, and processOwnershipTransferEvent inserted system messages unconditionally. Outbox retries and initial-sync replays (triggered whenever an admin re-approves a peering request, which recreates the peer row with lastSyncedAt=0) duplicated the system message on every delivery. Each new snowflake ID exceeded the user's last_read_message_id, flipping the channel back to unread after every deploy.

Processors now short-circuit on a matching (source_instance, source_message_id) row, and persist those fields when inserting. processMemberAddEvent emits the tagged system message in both bootstrap and incremental paths so bootstrap replays don't fall through and insert a second one; the bootstrap's dm_channel_created broadcast carries that message as lastMessage so sidebar previews and unread anchors agree across instances.
2026-04-21 01:06:43 +02:00
Jannis Braun aeebf79feb fix: federated friend request routed to wrong user with same name
When two instances each have a native user with the same username, the
Add Friend search card for the federated one sent its request to the
local namesake instead of the intended remote user.

Root cause: `isNative = !homeUserId` in socialStore's searchUsers and
loadFriends dedup. The server backfills native users' homeUserId to
their own id so federation tier-1 lookups succeed, so `homeUserId` is
set on natives too. Only `homeInstance` distinguishes native (null)
from replicated stubs. With the wrong check, no entry was ever "native"
and the home-origin stub of the remote user was kept over the true
native record — leaving `_instanceOrigin=''`, which caused the Send
button handler to drop the domain suffix and POST to the home API,
where "nova" resolved to a completely different local user.

Also fixes loadRequests dedup to prefer the target-native record so the
search card correctly flips to "Request Pending" after sending.
2026-04-21 00:36:15 +02:00
Jannis Braun 3d8709d20a feat: real-time Federation panel updates via WS events
Added federation_peers_changed (no-payload signal) broadcast from every
peer state mutation, and federation_approval_request_received when a new
approval request is queued. Client subscribes via onFederationPeersChanged
callback registry. FederationPanel and PendingApprovals debounce-refetch
on any event. sendToAdmins helper broadcasts only to admin users.
2026-04-20 18:28:10 +02:00
Jannis Braun 6afad97bd1 fix: revert outgoing peering blocks — autoAcceptPeering only gates incoming
autoAcceptPeering means 'don't accept peering initiated by others', not
'don't initiate peering ourselves'. Two checks were incorrectly blocking
outgoing peering when auto-accept was off:

1. ensurePeered() refused to auto-initiate — reverted. When a local user
   sends a DM, the server should initiate peering. The remote's
   peer/accept decides whether to accept or queue.

2. queueOutboxEvent() refused to create placeholders — reverted. The
   outbox needs placeholders to queue entries. Without them, DM relay
   silently fails.
2026-04-20 18:08:05 +02:00
Jannis Braun 165fda44a3 fix: accept incoming handshake for awaiting_approval peers to break approval ping-pong
When both instances have autoAcceptPeering off, the approval flow
ping-ponged indefinitely. Admin A approves → handshakes to B → B
queues (202) → A's peer becomes awaiting_approval. Admin B approves →
handshakes to A → but A's gate only matched 'pending', not
'awaiting_approval', so it re-queued instead of accepting.

Now the gate matches both 'pending' and 'awaiting_approval'. When the
second admin approves and handshakes back, the first instance recognizes
its admin already approved and accepts — completing the peering.
2026-04-20 18:01:47 +02:00
Jannis Braun 072858cbbb fix: multiple federation peering bugs
1. queueOutboxEvent no longer creates pending peer placeholders when
   autoAcceptPeering is disabled — prevents bypassing the admin's
   peering control

2. Approval endpoint checks for 202 before response.ok — when the
   remote also has autoAcceptPeering off, sets peer to awaiting_approval
   instead of incorrectly activating it

3. awaiting_approval status added to Federation panel UI — status label,
   colors, filter options so these peers are visible and manageable
2026-04-20 17:54:00 +02:00
Jannis Braun b40c57f227 fix: call resolvePendingPeers before early return in processOutboxTick
When all peers are pending (no active peers with outbox entries),
processOutboxTick returned early at line 141 before reaching
resolvePendingPeers at line 302. Pending peers were never resolved
because the only code path to resolvePendingPeers was after the
active-peer delivery loop — which never ran.
2026-04-20 17:41:34 +02:00
Jannis Braun 34fe9115b9 fix: compute target origins for 1-on-1 DMs so pending peers are created
getGroupDmTargetOrigins() returned undefined for 1-on-1 DMs, which
queueOutboxEvent() treated as 'broadcast to all existing peers'. When
no peers existed, nothing was queued and no handshake was ever triggered.
Now always computes target origins from DM participants so the pending
placeholder creation path runs, enabling ensurePeered() → peer/accept
→ approval queue flow.
2026-04-20 17:15:22 +02:00
Jannis Braun 0aec716d4c fix: gate all client DM events on active S2S peer status
The client's direct WS connection to remote instances (via Connections)
delivered DM events independently of S2S peering. Added activePeerOrigins
allowlist to ready payload — all DM event handlers now silently drop
events from non-home origins without an active peer. This prevents
notifications, sounds, previews, typing indicators, calls, and channel
updates from instances where peering was revoked or never established.
2026-04-20 17:05:57 +02:00
Jannis Braun 83b682d501 fix: block auto-peering initiation when autoAcceptPeering is disabled
ensurePeered() now checks the local autoAcceptPeering setting before
initiating new peering. When disabled, only admin-explicit peer/initiate
and approval-request approve bypass this check. Closes the bypass where
client peer/ensure or outbox worker could auto-initiate outward peering
even when the admin intended to control all peering.
2026-04-20 16:51:44 +02:00
Jannis Braun 8be30dc95f fix: check 202 before response.ok so queued approval isn't treated as accepted 2026-04-20 16:41:56 +02:00
Jannis Braun e4e0d0d1f1 fix: handle 403 (inactive peer) alongside 401 for stale peer re-handshake 2026-04-20 16:32:24 +02:00
Jannis Braun ce33ccf69e fix: reset stale peer to pending on 401 so ensurePeered re-handshakes 2026-04-20 16:21:53 +02:00
Jannis Braun ea8f786c2b docs: document pending peering approval queue, new endpoints, and awaiting_approval status 2026-04-20 15:16:07 +02:00
Jannis Braun b0d3ee93b0 fix: restore approvalCount state variable in FederationPanel 2026-04-20 15:11:47 +02:00
Jannis Braun 9e3c411e80 feat: badge count on Federation tab for pending approval requests 2026-04-20 15:10:56 +02:00
Jannis Braun 975ef93cc0 feat: pending approval requests section in Federation panel 2026-04-20 15:10:07 +02:00
Jannis Braun 39028ae56f feat: two-variant DM unreachable indicator for rejected vs awaiting_approval 2026-04-20 15:06:48 +02:00
Jannis Braun c3fe1bc9d4 feat: track awaitingApprovalPeerOrigins and show login toast for pending approvals 2026-04-20 15:06:00 +02:00
Jannis Braun 12bb11e9de feat: add approval request API methods to client 2026-04-20 15:05:03 +02:00
Jannis Braun 33c45e3184 feat: add awaitingApprovalPeerOrigins and pendingApprovalCount to ready payload 2026-04-20 15:03:29 +02:00
Jannis Braun 5f50ffc5f2 feat: janitor expiry for peer approval requests with signed denial 2026-04-20 15:01:31 +02:00
Jannis Braun 1920324469 feat: add admin approval-request endpoints (list, approve, deny) 2026-04-20 14:59:52 +02:00