Files
backspace/docs/systems/federation.md
T
Jannis Braun a3a7527c9e chore: add system docs, specs, and misc updates from other sessions
- Add complete docs/systems/ reference (18 system docs)
- Add federation relay status doc and prior spec/plan docs
- Remove superseded docs/federation-dm-s2s.md (replaced by docs/systems/federation.md)
- CLAUDE.md updates
- Minor fixes in social.ts, types.ts, AddDmMemberModal, NewDmModal, UserSettings
2026-03-31 03:40:34 +02:00

46 KiB

Federation System

Source files:

  • packages/server/src/routes/federation.ts -- API endpoints (peer handshake, relay, sync) + all inbound event processors + identity resolution functions
  • packages/server/src/utils/federationAuth.ts -- HMAC signing, verification, header parsing, getOurOrigin()
  • packages/server/src/utils/federationOutbox.ts -- Event queuing, coalescing, relay payload construction, mutation log, participant/target resolution
  • packages/server/src/utils/federationWorker.ts -- Background workers: outbox delivery, file download, health check, janitor, initial sync
  • packages/server/src/utils/storageJanitor.ts -- Federation GC: outbox expiry, mutation log retention, file queue cleanup, DM channel purge
  • packages/server/src/routes/social.ts -- Friend request/accept/cancel/remove endpoints that queue federation events
  • packages/server/src/routes/dm.ts -- DM REST endpoints that queue federation events (message relay, group lifecycle)
  • packages/server/src/ws/events.ts -- WebSocket event handlers that queue DM message/reaction relay events
  • packages/web/src/utils/profileSync.ts -- Client-side profile sync via LWW timestamps (not S2S relay)
  • packages/web/src/utils/identity.ts -- Client-side federated identity resolution helpers

DB tables: federation_peers, federation_outbox, federation_file_queue, federation_mutation_log, plus users (identity), dm_channels/dm_members/dm_messages (DM federation), friends/friend_requests (friend federation), attachments (file replication). See docs/systems/database.md for full schemas.


Architecture Overview

Backspace federation is peer-to-peer with no central authority. Each instance maintains its own copy of all data. Peers exchange real-time events for DMs and friendships via a signed relay protocol.

Canonical identity: The (homeUserId, homeInstance) pair is globally unique. Local users have homeInstance = NULL and homeUserId = NULL. Federated users are represented as replicated user stubs -- minimal user records with passwordHash = '!federation-replicated' (bcrypt never produces this value, so login is impossible).

Trust model: Symmetric shared-secret HMAC. Both peers share the same 256-bit secret. Events are attributed to users by homeUserId + homeInstance in the payload, with authority checks verifying the source instance matches the claimed origin of the acting user.


1. Peer Handshake & Discovery

2-Phase Flow

Phase 1 -- Initiate (POST /api/federation/peer/initiate)

  • Auth: JWT + admin role required
  • Validates remoteOrigin is a well-formed HTTP(S) URL via validateOrigin()
  • Prevents self-peering (localOrigin === remoteOrigin)
  • Handles existing peers: active -> return 200, pending -> return 409, revoked -> delete and re-initiate
  • Generates HMAC secret: generateHmacSecret() -> randomBytes(32).toString('hex') (256-bit)
  • Generates challenge: randomBytes(16).toString('hex') (128-bit, currently unused by acceptor)
  • Creates local peer record with status='pending'
  • POSTs to {remoteOrigin}/api/federation/peer/accept with { sourceOrigin, challenge, hmacSecret }
  • Timeout: 10 seconds (AbortSignal.timeout)
  • On remote acceptance: updates local peer to status='active', sets lastSeenAt
  • On failure: deletes pending peer, returns 502 (network error) or 504 (timeout)

Phase 2 -- Accept (POST /api/federation/peer/accept)

  • Auth: none (first contact -- no JWT, no HMAC)
  • Rate-limited: 10 requests per minute per IP (in-memory sliding window, buckets cleaned every 60s)
  • Validates sourceOrigin, challenge, and hmacSecret from body
  • Handles existing peers: active -> return 200 (idempotent), revoked -> return 403, pending -> update with new secret and activate
  • New peer: creates record with provided hmacSecret, sets status='active'
  • Returns { accepted: true } on success

Secret Storage

Both instances store the same HMAC secret. The initiating instance generates it and sends it in the accept request. There is no secret rotation mechanism -- the secret persists until the peer is revoked and re-initiated.

Peer Status Lifecycle

                     initiate
  (none) ──────────► pending ──────────► active
                                           │
                    10+ consecutive         │ delivery failures
                    failures               ▼
                                       unreachable
                                           │
                    health check OK        │
                                           ▼
                                        active
                                           │
                    admin revoke           │
                                           ▼
                                        revoked ──► (delete) ──► re-initiate
Status Outbox delivery Health check Relay accepts Re-initiation
active Yes No Yes No (returns existing)
pending No No No No (returns 409)
unreachable No (entries wait) Yes (1h interval) Yes (resets to active) No
revoked No (entries purged) No No (returns 403) Yes (old record deleted)

PEER_UNREACHABLE_THRESHOLD

Defined in federationWorker.ts:45 as 10. After 10 consecutive delivery failures for a peer, the worker sets status = 'unreachable'. The health check worker (1h interval) pings GET /api/instance/info on unreachable peers and reverts to active on success.

Admin Endpoints

Endpoint Method Auth Purpose
/api/federation/peer/initiate POST JWT + admin Start peering handshake
/api/federation/peer/accept POST None (rate-limited) Accept incoming handshake
/api/federation/peers GET JWT + admin List all peers (secret excluded)
/api/federation/peers/:id DELETE JWT + admin Revoke peer, purge outbox

2. HMAC Request Authentication

Signing Format

HMAC-SHA256(secret, "${timestamp}.${requestBody}")

Where timestamp is Date.now() (Unix milliseconds) and requestBody is the JSON string.

HTTP Headers

Header Format Example
X-Federation-Signature sha256=<hex> sha256=a1b2c3...
X-Federation-Origin Full URL https://nova.ddns.net
X-Federation-Timestamp Unix ms string 1711619400000
Content-Type application/json --

Verification (federationAuth.ts:verifySignature)

  1. Validate inputs: reject empty/missing body, signature, or secret
  2. Timestamp window: Math.abs(Date.now() - timestamp) <= maxAgeMs (default 15 minutes)
  3. Recompute: HMAC-SHA256(secret, "${timestamp}.${body}")
  4. Constant-time comparison: crypto.timingSafeEqual on hex-decoded buffers
  5. Length check: mismatched buffer lengths are rejected before timingSafeEqual

Replay Attack Prevention

The 15-minute timestamp window prevents replaying old requests. However, there is no nonce or sequence number -- a valid request can be replayed within the 15-minute window. See Known Issues.

Inbound Verification Flow (POST /api/federation/relay)

  1. parseFederationHeaders() extracts origin, timestamp, signature from headers
  2. Look up peer by origin in federation_peers -- must exist and be status = 'active'
  3. Re-serialize request body to JSON: JSON.stringify(request.body)
  4. verifySignature(bodyString, signature, peer.hmacSecret, timestamp) -- reject if false

Important: The body is re-serialized server-side. This means Fastify's JSON parsing and re-stringification must produce identical output to the sender's JSON.stringify. In practice this works because both sides use standard JSON.stringify with no custom replacers.


3. Identity Resolution

Functions

resolveLocalUser(homeUserId, db) -- federation.ts:878

  • Read-only lookup. Returns undefined if not found.
  • Matches: (users.homeUserId = homeUserId) OR (users.id = homeUserId AND homeInstance IS NULL)
  • Excludes deleted users (isDeleted = 0)
  • When multiple candidates exist: prefers the one with homeUserId set (replicated stub) over a local ID match
  • Use when: Optional lookups where null is acceptable (member_remove, reaction processing, friend_remove)

resolveOrCreateReplicatedUser(homeUserId, homeInstance, db) -- federation.ts:912

  • Calls resolveLocalUser first. If found, returns it.
  • If not found, creates a stub with:
    • username: {homeUserId}@{domain} (domain extracted from homeInstance URL)
    • passwordHash: '!federation-replicated'
    • homeInstance: the full URL passed in
    • homeUserId: the remote user's home ID
    • Collision-safe: appends _1, _2, ..., _10 suffix if username exists; after 10 attempts, uses _<random hex>
  • Use when: You MUST have a valid user ID (setting ownerId, inserting dm_members, creating messages)

hydrateReplicatedUserProfile(user, profile, db) -- federation.ts:2041

  • Updates replicated stubs only (homeInstance must be set)
  • Only updates null/empty fields (preserves manually-set local values)
  • Exception: avatar/banner are overwritten if the current value is a bare filename (not an absolute URL)
  • Resolves bare filenames to {homeInstance}/api/uploads/{filename} absolute URLs
  • Sets displayName from profile.displayName || profile.username -- ensures federated users show a human-readable name instead of user@instance

Critical Rule

Any code path that sets ownerId, creates a dm_members row, or inserts a message MUST use resolveOrCreateReplicatedUser. Using resolveLocalUser with a ?? null fallback has caused data corruption (see Known Issues: ownerId nulling).

Origin Normalization

Two formats exist in the database:

Location Format Example
users.home_instance Bare domain OR full URL nova.ddns.net or https://nova.ddns.net
federation_peers.origin Full URL https://nova.ddns.net
getOurOrigin() return Full URL https://orbit.ddns.net
resolveOrCreateReplicatedUser stores Full URL (passed through) https://nova.ddns.net
Auth registration stores Bare domain nova.ddns.net

The inconsistency exists because:

  • resolveOrCreateReplicatedUser stores homeInstance as-is from the relay event (full URL)
  • The auth registration path (/api/auth/register with homeInstance param) validates as bare domain only (regex: /^[a-zA-Z0-9._-]+$/)
  • Relay event payloads populate homeInstance from getOurOrigin() (full URL) or from user.homeInstance || getOurOrigin() (which falls back to full URL)

Normalization pattern used in code:

const normalized = homeInstance.startsWith('http') ? homeInstance : `https://${homeInstance}`;

Locations where normalization is applied:

  • getGroupDmTargetOrigins() (federationOutbox.ts:294) -- normalizes before comparing to ourOrigin
  • dm.ts:655 -- isLocalMember broadcast filter checks both formats
  • dm.ts:743 -- normalizes target homeInstance before peer origin comparison

Locations with potential mismatch (see Known Issues):

  • federation.ts:1278 -- memberUser?.homeInstance === sourceInstance -- compares stored homeInstance (possibly bare domain) against sourceInstance (full URL from relay request header)
  • federationOutbox.ts:376-379 -- getFriendEventTargets compares fromHomeInstance against ourOrigin without normalization. The passed values come from user.homeInstance || domainOrigin where domainOrigin = getOurOrigin(). If homeInstance is a bare domain, homeInstance !== ourOrigin is true, so the bare domain gets added to targets, but queueOutboxEvent then fails to match it against federation_peers.origin
  • federationWorker.ts:424 -- user.homeInstance === ourOrigin in handleSizeRejection. Bare domain homeInstance won't match, potentially including a user in affectedUserIds who shouldn't be (minor).
  • federation.ts:2388 -- from.homeInstance === ourOrigin in processFriendAddEvent determines which user is "local" for broadcasting. The from.homeInstance comes from the relay event payload, which should be a full URL, so this comparison works correctly in practice.
  • federation.ts:2447 -- same pattern in processFriendRemoveEvent

4. DM Message Relay

1-on-1 DMs

Outbound (origin instance):

  1. Message created via REST (POST /api/dm/:id/messages) or WS (dm_message_create)
  2. queueDmRelay(message, channelId, 'create') called from dm.ts / events.ts
  3. buildRelayPayload() constructs the message portion with homeUserId, homeInstance, content, replyToId, editedAt, createdAt
  4. getDmParticipants(channelId) resolves all members to (homeUserId, homeInstance) pairs with profile snapshots
  5. getGroupDmTargetOrigins(channelId) returns undefined (no owner -> broadcast to all)
  6. queueOutboxEvent(messageId, channelId, 'create', payload, undefined) -> queued to ALL active peers

Inbound (receiving instance -- processCreateEvent):

  1. Validate: event.message and event.participants (>= 2) required
  2. Dedup: check (sourceInstance, sourceMessageId) -- reject if exists
  3. Resolve ALL participants via resolveOrCreateReplicatedUser, hydrate profiles
  4. No event.federatedId -> 1-on-1 path
  5. Compute deterministic federatedId = SHA256(sorted([homeUserIdA, homeUserIdB])).slice(0, 32)
  6. findOrCreateDmChannel(federatedId, [localUserA.id, localUserB.id], db):
    • Find by federatedId in dm_channels
    • If exists: ensure both users are members (idempotent insert)
    • If not: create channel with federatedId, add both members
  7. Insert dm_messages with sourceInstance and sourceMessageId
  8. Process attachments (see File Replication)
  9. Broadcast dm_message_created to local members, skipping members whose homeInstance === sourceInstance (they already have the original)

Group DMs

Outbound (origin instance): Same as 1-on-1 except:

  • getGroupDmTargetOrigins(channelId) returns a list of peer origins that have at least one participant
  • Normalizes homeInstance to full URL before comparison
  • queueOutboxEvent receives targetPeerOrigins and only queues to those peers
  • Payload includes federatedId (random UUID assigned at channel creation)

Inbound (receiving instance -- processCreateEvent):

  1. event.federatedId present -> group DM path
  2. Find channel by federatedId -- must already exist (bootstrapped by prior member_add)
  3. If not found -> reject with channel_not_found
  4. Insert message, broadcast to local members

Federated ID Generation (federationOutbox.ts:computeFederatedId)

// 1-on-1: deterministic 32-char hex hash
const sorted = [homeUserIdA, homeUserIdB].sort();
return sha256(sorted.join(':')).slice(0, 32);

// Group: random 36-char UUID with dashes
return crypto.randomUUID();

The format difference (32-char hash vs 36-char UUID) is used by the self-healing migration to detect channel type independently of owner_id.

Message Deduplication

Every relayed message is stored with:

  • source_instance: the relay request's sourceInstance header value
  • source_message_id: the event.messageId (original message ID on source instance)

The (source_instance, source_message_id) pair is checked before insertion. Duplicates are rejected with reason 'duplicate'. A unique partial index enforces this at the DB level: idx_dm_messages_source_unique ON dm_messages(source_instance, source_message_id) WHERE source_instance IS NOT NULL.


5. Outbox & Relay Pipeline

Event Queuing (federationOutbox.ts:queueOutboxEvent)

Trigger (API/WS handler)
  -> isFederationRelayEnabled()? No -> return silently
  -> Fetch active peers from federation_peers
  -> Filter to targetPeerOrigins (if specified) -- EXACT string match against peer.origin
  -> If zero peers match -> return silently (KNOWN ISSUE: silent event dropping)
  -> For each peer, in a transaction:
      -> Check for existing outbox entry by (peerId, entityId)
      -> COALESCE:
          - delete + existing create -> delete both (net: never relayed)
          - update + existing create -> update payload, keep 'create' eventType
          - update + existing update -> update payload and eventType
          - no existing -> insert new entry
      -> TTL: now + (relayTtlDays * 86400000)

Coalescing Rules (per-peer, per-entity)

Incoming Existing Result
delete create Entry removed (message was never relayed)
update create Payload updated, keeps create type (peer gets full message)
update update Payload updated, type becomes latest
delete update Payload updated, type becomes delete
any none New entry inserted

Outbox Delivery Worker (federationWorker.ts:processOutboxTick)

Interval: 10 seconds (OUTBOX_INTERVAL_MS) Batch size: 50 (OUTBOX_BATCH_LIMIT) Timeout: 30 seconds per request (OUTBOX_FETCH_TIMEOUT_MS)

  1. Query entries where nextRetryAt <= now joined with active peers, ordered by createdAt ASC, limit 50
  2. Group by peer
  3. For each peer, reconstruct FederationRelayEvent[] from stored payloads:
    • Parse JSON payload
    • Copy fields: federatedId, participants, message, reactions, reaction, membership, ownership, group, friendship, file_rejected fields
    • Set eventType, contextType, messageId, dmChannelId, encryptionVersion, timestamp
  4. Build FederationRelayRequest with version: 1, sourceInstance: ourOrigin
  5. Sign with buildFederationHeaders(body, peerHmacSecret, ourOrigin)
  6. POST to {peerOrigin}/api/federation/relay
  7. On success (200):
    • Delete accepted entries from outbox (matched by entityId -> outboxId)
    • Log rejected entries (remain in outbox for retry)
    • Store result.maxUploadSize on peer record
    • Update peer: lastSeenAt = now, consecutiveFailures = 0
  8. On failure (non-200 or network error):
    • handleOutboxDeliveryFailure():
      • Increment attempts per entry, compute nextRetryAt = now + backoff
      • Increment peer consecutiveFailures, set lastFailureAt
      • If consecutiveFailures >= PEER_UNREACHABLE_THRESHOLD (10) -> mark peer unreachable

Retry Backoff Schedule

Attempt Delay
1 30 seconds
2 1 minute
3 5 minutes
4 15 minutes
5 1 hour
6 6 hours
7+ 24 hours (cap)

Relay Request/Response Format

Request:

interface FederationRelayRequest {
  version: 1;
  sourceInstance: string;          // Full URL, e.g., "https://nova.ddns.net"
  events: FederationRelayEvent[];  // Max 50 per batch
}

Response:

interface FederationRelayResponse {
  accepted: string[];              // messageIds successfully processed
  rejected: Array<{
    messageId: string;
    reason: string;                // e.g., 'duplicate', 'unknown_message', 'missing_participants'
  }>;
  maxUploadSize: number;           // This instance's max upload size in bytes
}

Inbound Relay Dispatch (POST /api/federation/relay)

Body limit: 10 MB. Max 50 events per batch.

eventType Processor contextType
create processCreateEvent dm
update processUpdateEvent dm
delete processDeleteEvent dm
reaction_add processReactionAddEvent dm
reaction_remove processReactionRemoveEvent dm
member_add processMemberAddEvent dm
member_remove processMemberRemoveEvent dm
ownership_transfer processOwnershipTransferEvent dm
friend_request_create processFriendRequestCreateEvent friend
friend_request_update processFriendRequestUpdateEvent friend
friend_request_cancel processFriendRequestCancelEvent friend
friend_add processFriendAddEvent friend
friend_remove processFriendRemoveEvent friend
file_rejected processFileRejectedEvent dm

After processing all events, the relay endpoint updates the peer's lastSeenAt and resets consecutiveFailures, then returns accepted/rejected arrays plus maxUploadSize.


6. Group DM Lifecycle over Federation

member_add (processMemberAddEvent -- federation.ts:1618)

Required fields: event.federatedId, event.membership.user

Two paths:

Bootstrap path (channel does not exist locally by federatedId):

  1. Requires event.group metadata (owner + full member roster)
  2. Creates dm_channels row with federatedId, ownerId (resolved via resolveOrCreateReplicatedUser), ownerHomeUserId, ownerHomeInstance
  3. Adds ALL roster members from event.group.members (each resolved via resolveOrCreateReplicatedUser)
  4. Sends dm_channel_created to local-only members (home instance matches getOurOrigin(), with normalization for bare domain)
  5. Sets bootstrapped = true to skip redundant system messages and member_add broadcasts below

Incremental path (channel already exists):

  1. Validates authority: sourceInstance === channel.ownerHomeInstance (only owner's instance can add)
  2. Cancels soft-delete if channel was pending GC
  3. Resolves added user via resolveOrCreateReplicatedUser
  4. Enforces max 10 members
  5. Inserts dm_members row (idempotent -- skip if exists)
  6. Inserts system message, broadcasts dm_member_added to local WebSocket clients

member_remove (processMemberRemoveEvent -- federation.ts:1825)

  1. Find channel by federatedId -- if not found, accept idempotently
  2. Validate authority: owner's instance for kicks (reason !== 'leave'), any instance for self-leave
  3. Resolve user via resolveLocalUser -- if not found, accept idempotently
  4. Insert system message (before deletion, so broadcast includes the leaving user)
  5. Delete dm_members row, clean up read_states
  6. Broadcast dm_member_removed to remaining local members
  7. If zero members remain -> soft-delete channel (deletedAt = now)

ownership_transfer (processOwnershipTransferEvent -- federation.ts:1938)

  1. Find channel by federatedId -- if not found, accept idempotently
  2. Validate authority: sourceInstance === channel.ownerHomeInstance
  3. Resolve new owner via resolveOrCreateReplicatedUser (never resolveLocalUser -- must guarantee valid ID)
  4. Update dm_channels: ownerId, ownerHomeUserId, ownerHomeInstance
  5. Broadcast dm_owner_updated WebSocket event
  6. Insert system message with previous owner as actor

Local-Only Broadcast Principle

Users connected to multiple instances must see each DM channel exactly once (from their home instance). All structural broadcasts (dm_channel_created, system messages) filter to local members only:

const isLocalMember = (u: { homeInstance?: string | null }) =>
  !u.homeInstance || !domainOrigin ||
  u.homeInstance === domainOrigin ||
  `https://${u.homeInstance}` === domainOrigin;

Does NOT apply to: Regular DM messages (dm_message_created for user messages). These broadcast to all local dm_members regardless of home instance.

System Messages

System messages (type = 'system' in dm_messages) are instance-local -- they are NOT relayed via federation. Each instance creates its own when processing events.

Event Content JSON Actor (userId)
member_added {event, targetUserId, targetDisplayName} User who added them
member_removed {event, targetUserId, targetDisplayName, reason} User who left/was removed
owner_changed {event, newOwnerId, newOwnerDisplayName} Previous owner

Outbound Queuing (Origin Instance -- dm.ts)

When a group DM is created or modified locally, the origin instance queues federation events:

Group DM creation (POST /api/dm/group):

  • Iterates each remote target user (those with homeInstance !== domainOrigin)
  • Builds a member_add event per remote user, carrying the full roster in event.group
  • Computes finalTargets by starting from getGroupDmTargetOrigins() and adding the new member's normalized homeInstance
  • Calls appendMutationLog + queueOutboxEvent per event

Add member to existing group (POST /api/dm/:id/members):

  • Same structure as creation -- builds member_add with full group metadata
  • Normalizes new member's homeInstance to full URL before including in targets

Leave group (DELETE /api/dm/:id/members):

  • Computes fedTargetOrigins before deleting the member (so the leaving user's peer is still included)
  • Queues member_remove event with reason: 'leave'

Ownership transfer (PATCH /api/dm/:id):

  • Queues ownership_transfer event with previousOwner and newOwner

7. File Replication

Outbound (origin instance)

When queueDmRelay constructs the relay payload, each attachment gets a sourceUrl:

sourceUrl: `${getOurOrigin()}/api/uploads/${attachment.filename}`

Inbound (receiving instance -- processCreateEvent)

  1. For each attachment in event.message.attachments:
    • SSRF check: isUrlFromPeer(sourceUrl, peerOrigin) -- hostname of sourceUrl must match peer origin hostname
    • Create attachments row with filename = sourceUrl (remote URL as interim filename)
    • Queue federation_file_queue entry with status = 'pending', expiresAt = now + 30 days
  2. Initial WebSocket broadcast uses sourceUrl directly (frontend's AttachmentRenderer detects http prefix)

File Download Worker (federationWorker.ts:processFileQueueEntry)

Interval: 30 seconds. Batch: 5 files. Timeout: 60 seconds per download.

  1. SSRF protection: validate sourceUrl hostname matches peerOrigin hostname
  2. Pre-download size check against maxUploadSizeBytes from instance settings
  3. Download via fetch with streaming pipeline to disk (Readable.fromWeb -> fs.createWriteStream)
  4. Post-download size verification (defense in depth)
  5. Generate thumbnail via sharp (same as local upload flow)
  6. Update attachments row: filename = localFilename, size, thumbnailFilename
  7. Fallback: if no existing attachment row was found (legacy queue entry), insert a new one
  8. Mark file queue entry as completed with targetFilename
  9. Broadcast dm_message_updated to refresh client-side attachment display

Size Rejection Flow (handleSizeRejection)

When a file exceeds the local instance's size limit:

  1. Mark file queue entry as rejected with reason = 'size_limit_exceeded'
  2. Update local attachment: federationStatus = 'remote', federationMeta = source info JSON
  3. Determine affected local users (native to this instance -- !user.homeInstance || user.homeInstance === ourOrigin)
  4. Queue file_rejected reverse relay event to the sender's instance (sourceInstance)
  5. Broadcast dm_message_updated locally so clients see the 'remote' badge

Inbound file_rejected (processFileRejectedEvent -- federation.ts:2458)

When the origin instance receives a file_rejected event:

  1. Find local message by event.messageId (the original local message ID)
  2. Match attachment by sourceFilename or fallback to single attachment
  3. Resolve affectedUserIds (homeUserIds) to local replicated user stubs
  4. Merge rejection info into federationMeta (accumulates from multiple peers)
  5. Set federationStatus = 'remote_partial'
  6. Broadcast dm_message_updated + targeted federation_file_rejected toast to message author

Federation Status on Attachments

Status Meaning
null Local upload, no federation involvement
'local' Successfully downloaded from peer
'remote' Rejected (size limit), federationMeta has source instance info
'remote_partial' Rejected by some peers, federationMeta has per-user rejection array

File Download Retry

Uses the same backoff schedule as outbox delivery. Max attempts: 10 (MAX_FILE_ATTEMPTS). After exceeding max attempts: status = 'failed', rejectionReason = 'max_attempts_exceeded'.


8. Friend Relay

Event Flow (social.ts)

User Action Federation Event Authority Check
Send friend request friend_request_create from.homeInstance === sourceInstance
Accept/decline request friend_request_update to.homeInstance === sourceInstance
Cancel outgoing request friend_request_cancel from.homeInstance === sourceInstance
Accept creates friendship friend_add to.homeInstance === sourceInstance
Remove friend friend_remove Either side's instance

Target Resolution (getFriendEventTargets)

Computes which peer origins need the event. Compares fromHomeInstance and toHomeInstance against getOurOrigin(). Known issue: no normalization -- bare domain homeInstance will not match full URL ourOrigin, causing the bare domain to be passed as a target. However, queueOutboxEvent then fails to match it against federation_peers.origin (full URL), silently dropping the event.

Context ID

Friend events use a deterministic context ID: friend:${sorted[homeUserIdA, homeUserIdB].join(':')}.

Outbound Payload Construction

Each friend endpoint builds a FederationRelayEvent with:

  • contextType: 'friend'
  • friendship payload containing from and to as FederationRelayParticipant objects
  • fromProfile and/or toProfile snapshots (FederationRelayProfileSnapshot)
  • entityId formatted as friend_req:${sorted_ids}:${timestamp} (for requests) or friend_remove:${sorted_ids}:${timestamp}

The full event payload is stored in both appendMutationLog (for sync) and queueOutboxEvent (for delivery).

Inbound Processing

processFriendRequestCreateEvent (federation.ts:2082):

  • Authority check: from.homeInstance !== sourceInstance -> reject
  • Resolve sender via resolveOrCreateReplicatedUser + hydrate profile
  • Resolve recipient via resolveLocalUser (must be native to this instance)
  • Idempotency: if already friends or pending request exists, accept as no-op
  • Create friend_requests row, broadcast friend_request_received to recipient

processFriendRequestUpdateEvent (federation.ts:2178):

  • Authority check: to.homeInstance !== sourceInstance -> reject
  • Resolve sender (original requester) via resolveLocalUser (must exist locally)
  • Resolve recipient (acceptor/decliner) via resolveOrCreateReplicatedUser
  • Find pending request, update status
  • Broadcast friend_request_accepted or friend_request_declined to the original sender

processFriendRequestCancelEvent (federation.ts:2254):

  • Authority check: from.homeInstance !== sourceInstance -> reject
  • Both users must exist locally. If not, accept idempotently.
  • Delete the pending friend request. Broadcast friend_request_cancelled to recipient.

processFriendAddEvent (federation.ts:2318):

  • Authority check: to.homeInstance !== sourceInstance -> reject
  • Resolve both users via resolveOrCreateReplicatedUser + hydrate profiles
  • Insert friends row (idempotent)
  • Auto-resolve any pending friend_requests to 'accepted' (handles out-of-order delivery)
  • Determine which user is local (from.homeInstance === ourOrigin) and broadcast friend_request_accepted

processFriendRemoveEvent (federation.ts:2404):

  • Authority check: either from.homeInstance or to.homeInstance must be sourceInstance
  • Both users resolved via resolveLocalUser. If not found, accept idempotently.
  • Delete friends row in both directions
  • Determine local user (whose homeInstance is NOT the source) and broadcast friend_removed

9. Profile Sync

Profile sync uses two mechanisms that operate independently:

S2S Profile Hydration (Server-side)

When relay events carry FederationRelayProfileSnapshot data:

  • processCreateEvent: hydrates participant profiles on message relay
  • processFriendRequestCreateEvent / processFriendAddEvent: hydrates friend profiles

hydrateReplicatedUserProfile only fills null/empty fields. Avatar/banner are overwritten only if the current value is not an absolute URL (catches stale bare filenames).

Client-side LWW Sync (profileSync.ts)

Operates via the web client, not S2S relay:

  • On connect to a remote instance, compares profileUpdatedAt timestamps
  • If home is newer: pushes profile to remote (re-uploads avatar/banner)
  • If remote is newer: pulls from remote to home, then relays to all other remotes
  • Incremental: syncProfileUpdateToRemotes pushes partial updates after local profile edits

This is a client-driven mechanism -- it only runs when a user is actively connected to multiple instances. It does not use the relay pipeline or outbox.


10. Reaction Relay

Outbound

Reactions are queued by WS event handlers in events.ts:

  • dm_reaction_add -> queueOutboxEvent(reactionId, channelId, 'reaction_add', payload, targetOrigins)
  • dm_reaction_remove -> queueOutboxEvent(messageId, channelId, 'reaction_remove', payload, targetOrigins)

Payload includes userId, homeUserId, emoji, createdAt, plus messageId and messageHomeInstance for cross-instance message resolution.

The mutation log entry for reactions stores a simpler payload (no messageId/messageHomeInstance), while the outbox entry carries the full reaction payload including those fields.

Inbound

processReactionAddEvent (federation.ts:1480):

  1. Resolve message via resolveLocalDmMessage(canonicalMessageId, messageHomeInstance, sourceInstance, db):
    • If messageHomeInstance === getOurOrigin() -> find by local ID (the message originated here)
    • Otherwise -> find by (messageHomeInstance || sourceInstance, canonicalMessageId) tracking -- uses messageHomeInstance when available (correct origin in 3-instance relay), falls back to sourceInstance
  2. Resolve reacting user via resolveLocalUser (must already exist)
  3. Dedup: check existing reaction by (dmMessageId, userId, emoji)
  4. Insert dm_reactions, broadcast reaction_added to local clients

processReactionRemoveEvent (federation.ts:1561):

  • Same resolution logic
  • Delete matching reaction, broadcast reaction_removed if changes > 0

11. Initial Sync

runInitialSyncForNewPeers() (federationWorker.ts:739)

Triggered once at server startup (async, non-blocking). Finds peers with status = 'active' and lastSyncedAt = 0.

For each unsynced peer:

  1. DM sync pass: Paginate through POST {peerOrigin}/api/federation/sync with sinceTimestamp = 0, limit = 100
  2. Self-POST: Relay received events by POSTing to {ourOrigin}/api/federation/relay -- this routes through the standard inbound processing
  3. Friend sync pass: Same pagination with contextType: 'friend'
  4. Update lastSyncedAt = Date.now() after completion
  5. On failure: don't update lastSyncedAt -- retried on next startup

Sync Endpoint (POST /api/federation/sync)

HMAC-authenticated. Returns events from the federation_mutation_log.

Request:

{ sinceTimestamp: number, dmChannelId?: string, federatedId?: string, contextType?: 'dm'|'friend', limit?: 1-500 }

Response:

{ events: FederationRelayEvent[], hasMore: boolean, checkpoint: number }

DM sync:

  • Queries all dm_channels with non-null federated_id (not soft-deleted)
  • Joins federation_mutation_log with dm_messages to reconstruct events
  • Only returns locally-created messages (source_instance IS NULL via the LEFT JOIN)
  • Handles delete mutations separately (message rows don't exist for deletes)
  • For create/update: fetches current message state from DB, builds full relay event with attachments and participants
  • Membership/friend mutations store the full event payload in the mutation log, so they are returned directly

Friend sync:

  • Queries federation_mutation_log WHERE context_type = 'friend'
  • Returns stored payloads directly (friend events carry their complete data)

Known Bug: DNS Hairpin Self-POST

runInitialSyncForNewPeers POSTs to {ourOrigin}/api/federation/relay where ourOrigin = getOurOrigin(). In production, ourOrigin is https://{DOMAIN}, e.g., https://nova.ddns.net. This means the server makes an HTTP request to itself through the public DNS and reverse proxy (Caddy). This works but:

  • Adds unnecessary network round-trip latency
  • Fails if DNS hairpin is not supported by the network
  • Fails if the server is behind NAT without hairpin NAT configured

A direct function call to the relay processing logic would be more robust.


12. DM Calls over Federation

DM calls use LiveKit for WebRTC signaling and media transport. The call lifecycle is managed entirely via WebSocket events (dm_call_start, dm_call_accept, dm_call_reject, dm_call_end in ws/events.ts).

Current state: DM calls do NOT work across federated instances.

The call state machine is local to a single server instance -- there is no federation relay for call events. When user A on instance 1 calls user B on instance 2:

  • The dm_call_incoming event is sent via connectionManager.sendToUser(targetUser.id, ...) which only broadcasts to WebSocket connections on the local instance
  • User B's replicated stub exists on instance 1, but user B is connected via WebSocket to instance 2
  • The call event is never delivered

LiveKit tokens are also instance-local (/api/livekit/token requires JWT auth for the local instance).


13. Background Workers

All workers are started by startFederationWorkers() on server boot and stopped by stopFederationWorkers() on shutdown. Each worker uses setTimeout chains (not setInterval) with abort controllers for graceful shutdown.

Worker Interval Batch Timeout Source
Outbox delivery 10s 50 30s processOutboxTick
File download 30s 5 60s processFileQueueTick
Health check 1h all unreachable 10s processHealthCheckTick
Janitor 1h -- -- runFederationJanitor (sync)
Initial sync Once at startup -- 30s per page runInitialSyncForNewPeers

Janitor Cleanup (storageJanitor.ts:runFederationJanitor)

Target Condition Retention
federation_outbox expiresAt < now Configurable via federationRelayTtlDays (default 30)
federation_mutation_log mutatedAt < (now - 90 days) 90 days
federation_file_queue (completed) createdAt < (now - 7 days) 7 days
federation_file_queue (any) expiresAt < now 30 days (set at queue time)
dm_channels (soft-deleted) deletedAt < (now - 24h) 24-hour grace period

DM channel hard-delete cascades: reactions, embeds, attachments (DB rows + disk files), messages, members, outbox entries, mutation log entries, file queue entries.


14. Settings Cache

federationOutbox.ts caches federationRelayEnabled and federationRelayTtlDays from instance_settings for 30 seconds (CACHE_TTL_MS). This prevents repeated DB reads on every message send. The cache is invalidated by TTL only -- there is no explicit cache bust on settings change.

Relevant settings in instance_settings:

Column Default Purpose
federation_relay_enabled 1 Master toggle for all federation relay
federation_relay_ttl_days 30 Outbox entry TTL
max_upload_size_bytes null (uses config.maxUploadSize) File download size limit

15. Client-Side Identity Helpers (identity.ts)

The frontend needs to resolve federated identities for display purposes:

parseFederatedUsername(username) -- splits "youruser@nova.ddns.net" into {baseName: "youruser", domain: "nova.ddns.net"}.

isSelf(user, homeUser) -- determines if a user object is the logged-in user or their replicated stub. Uses cascading checks: same ID, known self-ID set, homeInstance + baseName match.

canonicalUserMatch(a, b) -- federation-safe check for whether two user objects represent the same person. Cascades through: same local ID, homeUserId cross-match, username + homeInstance fallback.

resolveDisplayIdentity(user, homeUser) -- returns homeUser for display if user is a replicated stub of homeUser, enabling consistent avatars and display names across instances.

Cross-instance self-ID registry: registerSelfId(id) / clearSelfIds() track all Snowflake IDs belonging to the current user across connected instances, populated from WS ready events.


16. Self-Healing Migrations (migrate.ts)

The migration system includes several data integrity checks that run on every server startup:

Group DM ownerId repair: Detects group DMs with UUID-format federated_id (length 36, matches ________-____-____-____-____________) but NULL owner_id. Restores the owner from the first remaining member or from owner_home_user_id/owner_home_instance. Root cause: a bug in processOwnershipTransferEvent (fixed in cd7aff0) could set ownerId = NULL via resolveLocalUser fallback.

Federated ID backfill: Finds 1-on-1 DM channels without federated_id, computes deterministic SHA-256 hash from home user IDs, and sets it. Also detects relay-created duplicate channels with the same federated_id and merges messages into the oldest channel.

Duplicate channel merge: Finds federated_id values appearing on multiple channels and merges them into the oldest, moving messages, members, and cleaning up the duplicates.

Mutation log backfill: If the federation_mutation_log table exists but is empty, populates it with create entries for all existing DM messages where source_instance IS NULL (locally-created messages).


Known Issues

1. Origin Format Inconsistency (PARTIALLY FIXED)

Root cause: users.home_instance stores both bare domains (from auth registration: nova.ddns.net) and full URLs (from resolveOrCreateReplicatedUser: https://nova.ddns.net). federation_peers.origin and getOurOrigin() always use full URLs.

Fixed locations:

  • getGroupDmTargetOrigins() normalizes before comparison (federationOutbox.ts:294)
  • isLocalMember in dm.ts:655 checks both formats
  • DM member_add target resolution in dm.ts:743 normalizes

Remaining unpatched comparisons:

  • federation.ts:1278 -- memberUser?.homeInstance === sourceInstance in processCreateEvent. If the member's homeInstance is a bare domain and sourceInstance is a full URL, this comparison fails. Result: the member receives the message even though they should be skipped (minor -- causes duplicate delivery, not data loss).
  • getFriendEventTargets() (federationOutbox.ts:376-379) -- compares fromHomeInstance/toHomeInstance against getOurOrigin() without normalization. When a local user's homeInstance is null, the fallback domainOrigin (full URL) is used, which works correctly. But when a user has homeInstance stored as a bare domain and that domain is this instance (e.g., an old replicated stub), the comparison bareDomain !== fullUrl evaluates to true, incorrectly including the local instance as a target, which then silently drops in queueOutboxEvent.
  • federationWorker.ts:424 -- user.homeInstance === ourOrigin in handleSizeRejection. Bare domain homeInstance won't match, potentially including a user in affectedUserIds who shouldn't be (minor).

2. Duplicate User Stubs

The same remote user can have multiple replicated records. Code paths that call resolveOrCreateReplicatedUser:

  • processCreateEvent (for each participant)
  • processMemberAddEvent (bootstrap roster + incremental add + owner + addedBy)
  • processOwnershipTransferEvent (new owner)
  • processFriendRequestCreateEvent (sender)
  • processFriendRequestUpdateEvent (recipient)
  • processFriendAddEvent (both users)

resolveLocalUser (called first by resolveOrCreateReplicatedUser) matches on homeUserId OR (id = homeUserId AND homeInstance IS NULL). If a user was created via auth registration (with homeInstance as bare domain) and later via relay (with homeInstance as full URL), resolveLocalUser may not find the first record if the IDs differ. The collision-safe username suffix ensures the insert succeeds, but now two stubs exist for the same person.

3. Silent Failures in queueOutboxEvent

queueOutboxEvent returns silently (no error, no log) when:

  • Federation relay is disabled
  • Zero active peers exist
  • targetPeerOrigins filter produces zero matches (origin format mismatch)

The third case is the most dangerous -- it looks like the event was queued but nothing was actually written. This has been the root cause of events silently disappearing for group DMs and friend events.

4. Trust Model Analysis

Threat Mitigation Gap
Peer impersonation X-Federation-Origin is verified against federation_peers.origin An attacker who compromises the HMAC secret can impersonate the peer
User attribution fraud Authority checks: e.g., from.homeInstance !== sourceInstance rejects events where the acting user doesn't belong to the source instance The check is string equality on homeInstance from the payload, which the sender controls. A malicious peer could claim any user belongs to them by setting homeInstance to their own origin.
Event flooding Outbox batches limited to 50 events. /api/federation/peer/accept rate-limited to 10/min. No rate limit on /api/federation/relay itself. A peer could send unlimited relay requests.
Replay attacks 15-minute timestamp window No nonce -- valid requests can be replayed within the window
Message content manipulation None A compromised peer can forge message content attributed to any user on their instance

5. DNS Hairpin Self-POST Bug

runInitialSyncForNewPeers (federationWorker.ts:792-797) POSTs received sync events to {ourOrigin}/api/federation/relay via public DNS. This adds unnecessary latency and fails when DNS hairpin is not configured. The function should call the relay processing logic directly instead of making an HTTP request to itself.

6. DM Calls Do Not Work over Federation

See section 12. The call state machine is entirely local to a single server instance. No federation relay exists for call events (dm_call_start, dm_call_incoming, dm_call_accept, dm_call_reject, dm_call_end).