Commit Graph
149 Commits
Author SHA1 Message Date
Jannis Braun 50bc510e8d feat(federation): outbound gate in ensurePeered with required intent argument
- New 'admin_required' EnsurePeeredResult variant.
- Required intent argument (no optional default) prevents future
  user-initiated callers from silently getting system behavior.
- Gate runs only when no peer row exists; toggling autoAccept later
  does not retroactively gate established peer rows.
- queueOutboundApproval helper upserts parent + subscriber rows and
  broadcasts to admins + user.
- Schema enum-narrows direction ('inbound' | 'outbound') and
  notifications.kind ('approved' | 'denied' | 'expired') at the column
  level; drizzle generate confirmed no migration delta.
- Existing call sites now fail typecheck — fixed in Task 5.
2026-04-26 21:18:43 +02:00
Jannis Braun b687bfe721 feat(federation): schema for outbound peering gate (direction column, subscribers, notifications) 2026-04-26 21:02:27 +02:00
Jannis Braun 058eb992e4 feat(federation): verify approval token before awaiting_approval → active
Receiver-side defense — closes the trust-bypass class the cheap fix
(4533e36) cannot cover. The /peer/accept handler's awaiting_approval
branch now requires an approvalToken matching the one stored on the
local peer row before promoting to active. Without a match:
- autoAccept=0 falls through to queueApprovalRequest (no bypass; new
  approval-request queued, existing awaiting_approval row untouched).
- autoAccept=1 falls back to permissive promotion (no regression vs
  prior behavior, since autoAccept=1 would accept any inbound regardless).

Successful match also deletes any stale approval-request row for the
origin to prevent debris accumulation from prior bypass attempts.

Refactor: queueing path extracted to a top-level queueApprovalRequest()
helper so both the no-existing-peer case and the awaiting_approval
mismatch fallback share one implementation. Helper generates a fresh
single-use token on every call.

Adds 'accept_awaiting_approval_fallback' variant to PeerActivationReason
to distinguish the autoAccept=1 fallback path from the verified path
in onPeerActivated audit logs.

Spec §3.5, §3.6.

Test counts: 289 → 301 (+12).
2026-04-26 11:48:33 +02:00
Jannis Braun 9e078c44ba feat(federation): performHandshake captures + clears approval token
- 202 response: parse approvalToken from body and store on the local
  federation_peers row alongside status='awaiting_approval'. Legacy
  receivers that omit the field result in null stored token, handled
  gracefully by the verification logic landing in subsequent commits.
- 200 response: clear approvalToken in the same UPDATE that sets
  status='active' (single-use consumption per spec §3.2).

Test coverage: 4 tests in federationPeering.approvalToken.test.ts covering
JSON token, missing token (legacy), non-JSON body, and 200 clear after
prior token storage.
2026-04-26 11:45:16 +02:00
Jannis Braun 11c5a2bf06 fix(federation): refuse outbound handshake when inbound approval pending
Closes the auto-reconnect trust-bypass: any code path calling
ensurePeered(remote) on an instance with autoAcceptPeering=0 could
previously bypass the admin gate by initiating a fresh handshake to the
remote, which the remote then accepted against its existing
awaiting_approval row.

The trigger surfaced was stores/instanceStore.ts:1010 — the silent
.catch(() => {}) auto-reconnect that fires for any user with the
remote in their replicatedInstances (commonly: any admin). Anyone with
that profile reloading their session activated peering on both sides
without any admin approval action.

Surgical fix: ensurePeered now returns rejected when an unresolved
inbound peer_approval_requests row exists for the target origin. The
legitimate admin-approve flow (routes/federation.ts:1089) does not call
ensurePeered; it deletes the approval-request and does its own direct
fetch to /peer/accept, so this check does not block legitimate approvals.

The receiver-side trust assumption at routes/federation.ts:619-645
(awaiting_approval branch in /peer/accept) still has the same flaw
— an adversarial peer that knows the timing could re-handshake at the
right moment to flip the receiver to active. That deeper trust-model
rework is plan-grade work tracked at internal notes
2026-04-26-peer-handshake-trust-model.md.
2026-04-26 00:12:42 +02:00
Jannis Braun 0677c21ab9 test(federation): buildFriendContextId determinism (cross-instance invariant)
The friend contextId must be byte-identical on both peers for a given
canonical pair regardless of argument order. Initial-sync backfill keys
on contextId; if the two sides computed different values for the same
friendship, missed events would never reconcile.
2026-04-25 22:24:04 +02:00
Jannis Braun 08872f1e08 feat(federation): rollbackFriendRequestCreate + side-effect registration 2026-04-25 21:51:04 +02:00
Jannis Braun 9e3417485c feat(federation-worker): treat receiver-ack 4xx reasons as terminal + invoke rollback 2026-04-25 21:46:28 +02:00
Jannis Braun be3eb5284e feat(federation): permanent-failure callback registry 2026-04-25 21:42:57 +02:00
Jannis Braun 03737c0955 feat(federation): resolveOriginFromHostname helper 2026-04-25 21:41:25 +02:00
Jannis Braun fbe2fb370b feat(federation): lookupRemoteUser outbound HMAC helper 2026-04-25 21:37:41 +02:00
Jannis Braun 1b98b052a1 chore(federation): T1 review polish — top-of-file import + undefined test case 2026-04-25 21:18:11 +02:00
Jannis Braun 1b63ca538e feat(federation): normalizeOriginForCompare helper 2026-04-25 21:14:58 +02:00
Jannis Braun 18d6b0acfa fix(federation): persist remote instance_name when ensurePeered/performHandshake succeeds
Initiator side of the bidirectional handshake exchange. /peer/accept now returns instanceName in the response body (prior commit); performHandshake parses it and persists alongside status='active'. Tolerates missing field (older peers) and non-JSON bodies.
2026-04-25 00:45:06 +02:00
Jannis Braun 2dbd2b9b9f test(server): add undeliverable:[] to processRelayEvents mock (#18)
Follow-up from Task 2 code review. Keeps the test mock aligned with
the widened return type even though vi.mock doesn't structurally
typecheck the factory.
2026-04-24 21:24:18 +02:00
Jannis Braun 26a4925032 feat(server): reclassify undeliverable targeted-peer as no_recipient failure (#18)
sendFederatedCallStart now treats a 200-with-undeliverable-messageId as
a peer-failure instead of unconditional success. Feeds the existing
failures[] array and terminal-determination machinery from #16.
New sendFederatedCallStartForTest export mirrors the existing
handleDm*ForTest pattern. TDD — three tests cover single-peer terminal
no_recipient, group-DM mixed delivered+undeliverable non-terminal, and
the happy-path (empty undeliverable → no event).

Also hardens sendCallRelay's response parse: validates undeliverable
is an Array and entries are well-shaped, logs protocol drift at warn/debug
rather than silently falling back to old-peer semantics.
2026-04-24 21:11:11 +02:00
Jannis Braun 7d2137b6d4 feat(server): sendCallRelay surfaces undeliverable messageIds (#18)
CallRelayResult success arm gains undeliverable: string[]. sendCallRelay
parses FederationRelayResponse.undeliverable (when present) and returns
the messageIds so sendFederatedCallStart can reclassify per-peer results.
Old peers that omit the field → empty array → today's behavior.
TDD — three tests cover old-peer, new-peer-with-undeliverable, and 5xx paths.
2026-04-24 21:05:28 +02:00
Jannis Braun 5c94bc7669 refactor(server): remove unused admin_reset PeerDeactivationReason
The admin reset endpoint (DELETE-pattern gated on peer.status !== 'needs_attention')
doesn't transition status — it deletes the row of an already-deactivated peer.
onPeerDeactivated already fired at the earlier needs_attention transition, so
the reset site correctly has no hook. The enum value was defensive-unused; per
project principles (no backwards-compat shims, no placeholders) drop it.
2026-04-24 01:07:47 +02:00
Jannis Braun 314df6c5a1 feat(server): 30s federated-call sentinel worker (TDD) 2026-04-24 00:49:04 +02:00
Jannis Braun 8e6639648e feat(server): wire onPeerDeactivated on performHandshake 403 rejection 2026-04-24 00:43:32 +02:00
Jannis Braun 743fdcac97 feat(server): wire onPeerDeactivated at federationWorker peer-deactivation sites 2026-04-24 00:42:20 +02:00
Jannis Braun 0a8949fbfb feat(server): onPeerDeactivated utility mirrors onPeerActivated (TDD) 2026-04-24 00:40:42 +02:00
Jannis Braun f6252b8ce1 feat(server): fan dm_call_end out on host ring timeout 2026-04-23 23:19:42 +02:00
Jannis Braun 6ff983b46c fix(federation): treat duplicate rejection as terminal in outbox worker
Duplicate rejection means the peer already has the message (e.g.,
delivered earlier via outbox AND pulled via sync in the same
window). Retrying will fail identically forever until TTL expires.

Before this patch: duplicate-rejected outbox entries were retained
with attempts++ and exponential backoff, creating log noise and
outbox bloat for up to 30 days.

After: duplicate-rejected entityIds join the terminal set alongside
accepted ones and are deleted from the outbox. Logged at info level
('outbox entry removed (terminal)') to distinguish from warn-level
transient-rejection retries.

Other rejection reasons (attribution_mismatch, processing_error,
etc.) stay on the retry path; some may also be terminal but are
deferred until observed accumulating.
2026-04-23 00:10:34 +02:00
Jannis Braun 15e42a7cc1 fix(federation): per-event fault isolation in syncPeerMutationLog (#25)
Replace the batch-level processRelayEvents call with a per-event
loop wrapped in try/catch. On exception: log event type, messageId,
timestamp, peer origin, and the error message; continue to the next
event.

Previously, a single poison-pill event (e.g., UNIQUE conflict from
a malformed relay payload) would throw, be caught by the outer
try/catch, and block lastSyncedAt from advancing — causing every
future activation to retry the same broken window indefinitely.

The final 'replayed N events' log line now reports '(K skipped due
to errors)' when K > 0, surfacing the count to operators. Individual
event failures are logged via console.error with enough context to
debug or replay manually.

Trade-off documented in docs/systems/federation.md: forward progress
of the sync pipeline takes priority over strict at-least-once
delivery. An event that fails to process is lost to the receiver
unless replayed manually.
2026-04-22 01:45:14 +02:00
Jannis Braun 911c7e3479 fix(federation): ensurePeered must not auto-heal needs_attention peers
Discovered during live verification of #10b Scenario 3: the switch
in ensurePeered had no case for needs_attention, so it fell through
to performHandshake. Because the /peer/accept idempotent-200-no-update
safeguard covers needs_attention on the inbound side, the remote
returned 200 without writing the new secret, and performHandshake
transitioned the local peer to 'active' on the 200 response — auto-
healing a state that requires admin intervention.

Affected paths: sendCallRelay non-blocking warm-up (used by typing
relay); any future caller of ensurePeered on a needs_attention peer.
Not affected: resolvePendingPeers (already filters on status='pending').

Fix: explicit case 'needs_attention' returning { status: 'rejected',
error }. Caller observes the rejection and does not advance state.
2026-04-22 01:27:07 +02:00
Jannis Braun a23e02339e feat(federation): capture dm_close/reopen/read_state/profile/file_rejected in mutation log
Four event types previously bypassed appendMutationLog, making
them unrecoverable via /api/federation/sync after peer inactivity:
  - queueDmCloseRelay (dm_close, dm_reopen)
  - queueReadStateRelay (read_state_update)
  - handleSizeRejection in federationWorker (file_rejected)
  - profile PATCH route (profile_update) — two call sites,
    one appendMutationLog per profile change (not per target origin)

The /api/federation/sync response builder is extended to
serialize these event types in the next task.
2026-04-22 00:48:49 +02:00
Jannis Braun 250596c0f6 feat(federation): wire onPeerActivated into 8 transition sites
Every code location that sets federation_peers.status='active'
now invokes onPeerActivated(peerId, reason). HTTP handler sites
use fire-and-forget (.catch(log)) so the response isn't blocked
by sync-pull pagination. The worker-internal health-check site
awaits the handler since the tick is already async.

Sites: /peer/initiate, /peer/accept (4 branches), /approval-
requests/:id/approve, health check recovery, ensurePeered/
performHandshake.
2026-04-22 00:39:58 +02:00
Jannis Braun 57d7ca66d3 fix(federation): explicit per-status handling in queueOutboxEvent
Replaces the silent UNIQUE-swallow placeholder branch. Each peer
status has an explicit branch:
  active/pending/unreachable: race-catch — re-fetch peer row and
    enqueue via matchedPeers. Previously skipped silently, losing
    real-time delivery under asymmetric failure.
  awaiting_approval/needs_attention/rejected/revoked: drop with
    logged reason. Mutation log still captures; sync-pull on
    activation replays.
  default: exhaustiveness check (no 'as never' cast) — TypeScript
    enforces that every status value is handled explicitly.
2026-04-22 00:34:43 +02:00
Jannis Braun ae035eba9b fix(federation): remove dead processRelayEvents import
Missed in 02a1ed7. The new sync-pull path in federationPeerActivation.ts
uses a dynamic import of processRelayEvents from routes/federation.js;
the static import in federationWorker.ts is no longer used after
runInitialSyncForNewPeers deletion.
2026-04-22 00:30:54 +02:00
Jannis Braun 02a1ed73f4 refactor(federation): replace runInitialSyncForNewPeers with startupBootstrapSync
The per-peer sync body is now syncPeerMutationLog (in the new
peer-activation module), invoked via onPeerActivated. The startup
path scans for status='active' AND lastSyncedAt=0 and calls the
unified handler for each — same trigger condition as before, unified
code path with runtime transitions.
2026-04-22 00:28:18 +02:00
Jannis Braun cca2245cdf feat(federation): implement onPeerActivated with dedup
Two-invariant handler: resetOutboxBackoff + syncPeerMutationLog.
In-flight map keyed by peerId coalesces concurrent activations —
a second call for a peer whose activation is still running shares
the same promise. Errors are swallowed and logged — the handler
never throws so fire-and-forget callers at HTTP handler sites
are safe.
2026-04-22 00:24:57 +02:00
Jannis Braun 39c43032e8 fix(federation): address Task 3 review — pagination test + polish
- Add pagination-advance test: verifies since=checkpoint on second
  iteration within a pass, and each pass re-seeds since from
  peer.lastSyncedAt (not carried from prior pass).
- Eliminate four peer! non-null assertions by capturing the narrowed
  value in activePeer after the guard.
- Tighten bodyObj type from Record<string, unknown> to a local
  SyncRequestBody type alias.
- Drop the no-op federationRelayEnabled UPDATE in test beforeEach
  (default is already 1 per baseline migration).
2026-04-22 00:21:34 +02:00
Jannis Braun 5c1b42938e feat(federation): implement syncPeerMutationLog
Three-pass pull-sync (dm, friend, profile) from peer's
/api/federation/sync endpoint, paginated. Seeds sinceTimestamp from
peer.lastSyncedAt so a recovered peer pulls only the delta. Updates
lastSyncedAt to Date.now() on full success; leaves it untouched on
transient failure so the next activation retries the same window.
Replaces the body of the soon-to-be-removed runInitialSyncForNewPeers.
2026-04-22 00:15:48 +02:00
Jannis Braun fd37e0c604 feat(federation): implement resetOutboxBackoff
Unconditionally resets nextRetryAt=now and attempts=0 for all
outbox entries of the given peer. No WHERE filter on nextRetryAt —
resetting attempts=0 on already-eligible rows is the correctness fix:
without it, a previously-failed entry keeps stale attempts, and its
next failure uses BACKOFF_SCHEDULE_MS[attempts] (5min to 24h) on a
peer that just recovered.
2026-04-22 00:10:27 +02:00
Jannis Braun 14a96efa58 feat(federation): scaffold peer-activation recovery module
Empty stubs for onPeerActivated, resetOutboxBackoff, syncPeerMutationLog,
and startupBootstrapSync. Functions are filled in by subsequent tasks
following TDD cycles.
2026-04-22 00:06:17 +02:00
Jannis Braun 0d74d1d112 perf(federation-worker): tighten health-check cadence to 15 min
HEALTH_CHECK_INTERVAL_MS was 1 h, but ROTATION_GRACE_PERIOD_MS is 15 min.
Phase skew between two peers' health-check ticks could stretch rotation
finalization desync up to ~1 h, during which signatures from the already-
finalized side verify against the other side's primary-only secret (grace
has expired; verifyPeerSignature stops trying the pending secret). With
AUTH_FAILURE_THRESHOLD = 5 and the existing backoff schedule, this
occasionally tripped legitimate rotations into needs_attention.

Setting the interval to 15 min (= ROTATION_GRACE_PERIOD_MS) guarantees a
finalization tick fires within one grace window on each side, so the
cross-verification window where one peer signs with NEW while the other
still treats NEW as pending cannot outlast the grace period.

Per-tick cost is negligible for the worker's steady state: the only
network fetches are per-active-peer /peer/rotate calls when the 90-day
rotation interval hits (rare) and per-unreachable-peer /instance/info
health pings (bounded by outage count). Going lower than 15 min would
reduce the residual desync but increase tick overhead with diminishing
returns; 15 min is the grace-period-aligned value that the original spec
("runs hourly") deviated from without justification.

Follow-up #20 from S2S DM unification backlog; reduces #19 false-positive
rate (outbox auth-failure transition) on legitimate rotations.
2026-04-21 22:16:09 +02:00
Jannis Braun 48dbe32a69 fix(federation-worker): reset consecutive_auth_failures on successful delivery
Pairs with the new 401/403 handler — a 2xx relay confirms HMAC trust is
healthy so the counter should clear. Mirrors the existing
consecutive_failures reset for network-layer health.
2026-04-21 20:46:52 +02:00
Jannis Braun 012e489bc7 fix(federation-worker): auth failures must not increment consecutive_failures
Code review of the previous commit found that the backoff branch of the
new 401/403 handler delegated to handleOutboxDeliveryFailure, which
double-dips by also incrementing consecutive_failures (the network-layer
counter that drives the 'unreachable' transition at threshold 10). Per the
design spec §State Machine Changes → Reset logic, auth failures must
increment consecutive_auth_failures ONLY.

Split handleOutboxDeliveryFailure into:
- applyOutboxEntryBackoff: just the per-entry backoff update (safe to call
  from the auth-failure path)
- handleOutboxDeliveryFailure: entry backoff + peer's consecutive_failures
  bump (network-error path only)

Also adds a console.warn to the backoff branch so operators can diagnose
clock-skew and rotation-grace incidents before the peer hits the terminal
threshold.

Part of backlog #19.
2026-04-21 20:45:33 +02:00
Jannis Braun e5afd376d2 fix(federation-worker): replace 401/403 wipe-and-rehandshake with bounded retry
The previous handler (commit ce33ccf + its 403 extension) wiped hmac_secret
and reset peer status to 'pending' on any 401/403 from an active peer. This
collapsed three distinct failure modes — transient clock skew, legitimate
split-brain, active MITM attempt — into "silently establish new trust
immediately." The remote's /peer/accept idempotent-200-no-update safeguard
then prevented the re-handshake from actually working, producing a 1-req/sec
loop observed during backlog #16 verification.

New behavior: increment consecutive_auth_failures, apply backoff to outbox
entries. At AUTH_FAILURE_THRESHOLD (5) transition to needs_attention,
preserve hmac_secret, surface delivery-impossible to affected users,
notify admins. Secret is NEVER wiped in response to a network-observed
401/403.

Part of backlog #19.
2026-04-21 20:39:37 +02:00
Jannis Braun 695ea0849d refactor(federation-worker): extract buildContextMapForPeer helper
Pure refactor — will be reused by the needs_attention transition handler.
No behavior change.
2026-04-21 20:35:51 +02:00
Jannis Braun 617d71ab4b feat(federation): add evaluateAuthFailure decision function
Pure function deciding whether the next 401/403 from an active peer
triggers backoff or a transition to needs_attention. Threshold = 5,
corresponding to ~21.5 min of the existing BACKOFF_SCHEDULE_MS.
2026-04-21 20:32:38 +02:00
Jannis Braun 53483d6981 polish(server): align sendCallRelay timeout message with codebase convention; use .then on sendTypingRelay fire-and-forget 2026-04-21 13:49:20 +02:00
Jannis Braun 21f220739c feat(server): sendCallRelay auto-peers on demand, typing passes peeringTimeoutMs:0
sendCallRelay now returns CallRelayResult with a typed reason on failure.
When the peer is not already active (or unreachable), runs a racePeering
against CALL_PEERING_TIMEOUT_MS (3s). Background handshake is not aborted
on race loss — next attempt succeeds.

sendTypingRelay passes peeringTimeoutMs:0 so typing never blocks on a
handshake; instead a warm-up ensurePeered runs in the background for any
non-active peer so the NEXT relay (message, call, or typing) benefits.
2026-04-21 13:45:29 +02:00
Jannis Braun 4ddb09edf1 fix(server): racePeering normalizes handshake rejections and only warns on timeout win
Addresses code review on b22a7bd: (1) a rejected handshake now returns
{ status: 'failed', error } instead of throwing, keeping the structured
contract; (2) the "background handshake" warn only fires when the
timeout arm wins — not when the handshake is itself the race winner by
rejection. Timing tests migrated to vi.useFakeTimers for determinism.
Regression test added for the handshake-wins-by-rejection case.
2026-04-21 13:41:44 +02:00
Jannis Braun b22a7bd0c6 feat(server): add racePeering helper with tests
Exports `racePeering(origin, timeoutMs, ensurePeeredFn?)` that races
`ensurePeered` against a deadline. On timeout, the background handshake
continues (warming the peer for the next attempt) and a warn-logged
.catch() prevents unhandledRejection. Injectable `ensurePeeredFn` param
enables full DI in tests without mocking module internals.
2026-04-21 13:37:30 +02:00
Jannis Braun 3d8709d20a feat: real-time Federation panel updates via WS events
Added federation_peers_changed (no-payload signal) broadcast from every
peer state mutation, and federation_approval_request_received when a new
approval request is queued. Client subscribes via onFederationPeersChanged
callback registry. FederationPanel and PendingApprovals debounce-refetch
on any event. sendToAdmins helper broadcasts only to admin users.
2026-04-20 18:28:10 +02:00
Jannis Braun 6afad97bd1 fix: revert outgoing peering blocks — autoAcceptPeering only gates incoming
autoAcceptPeering means 'don't accept peering initiated by others', not
'don't initiate peering ourselves'. Two checks were incorrectly blocking
outgoing peering when auto-accept was off:

1. ensurePeered() refused to auto-initiate — reverted. When a local user
   sends a DM, the server should initiate peering. The remote's
   peer/accept decides whether to accept or queue.

2. queueOutboxEvent() refused to create placeholders — reverted. The
   outbox needs placeholders to queue entries. Without them, DM relay
   silently fails.
2026-04-20 18:08:05 +02:00
Jannis Braun 072858cbbb fix: multiple federation peering bugs
1. queueOutboxEvent no longer creates pending peer placeholders when
   autoAcceptPeering is disabled — prevents bypassing the admin's
   peering control

2. Approval endpoint checks for 202 before response.ok — when the
   remote also has autoAcceptPeering off, sets peer to awaiting_approval
   instead of incorrectly activating it

3. awaiting_approval status added to Federation panel UI — status label,
   colors, filter options so these peers are visible and manageable
2026-04-20 17:54:00 +02:00
Jannis Braun b40c57f227 fix: call resolvePendingPeers before early return in processOutboxTick
When all peers are pending (no active peers with outbox entries),
processOutboxTick returned early at line 141 before reaching
resolvePendingPeers at line 302. Pending peers were never resolved
because the only code path to resolvePendingPeers was after the
active-peer delivery loop — which never ran.
2026-04-20 17:41:34 +02:00