- Add complete docs/systems/ reference (18 system docs) - Add federation relay status doc and prior spec/plan docs - Remove superseded docs/federation-dm-s2s.md (replaced by docs/systems/federation.md) - CLAUDE.md updates - Minor fixes in social.ts, types.ts, AddDmMemberModal, NewDmModal, UserSettings
46 KiB
Federation System
Source files:
packages/server/src/routes/federation.ts-- API endpoints (peer handshake, relay, sync) + all inbound event processors + identity resolution functionspackages/server/src/utils/federationAuth.ts-- HMAC signing, verification, header parsing,getOurOrigin()packages/server/src/utils/federationOutbox.ts-- Event queuing, coalescing, relay payload construction, mutation log, participant/target resolutionpackages/server/src/utils/federationWorker.ts-- Background workers: outbox delivery, file download, health check, janitor, initial syncpackages/server/src/utils/storageJanitor.ts-- Federation GC: outbox expiry, mutation log retention, file queue cleanup, DM channel purgepackages/server/src/routes/social.ts-- Friend request/accept/cancel/remove endpoints that queue federation eventspackages/server/src/routes/dm.ts-- DM REST endpoints that queue federation events (message relay, group lifecycle)packages/server/src/ws/events.ts-- WebSocket event handlers that queue DM message/reaction relay eventspackages/web/src/utils/profileSync.ts-- Client-side profile sync via LWW timestamps (not S2S relay)packages/web/src/utils/identity.ts-- Client-side federated identity resolution helpers
DB tables: federation_peers, federation_outbox, federation_file_queue, federation_mutation_log, plus users (identity), dm_channels/dm_members/dm_messages (DM federation), friends/friend_requests (friend federation), attachments (file replication).
See docs/systems/database.md for full schemas.
Architecture Overview
Backspace federation is peer-to-peer with no central authority. Each instance maintains its own copy of all data. Peers exchange real-time events for DMs and friendships via a signed relay protocol.
Canonical identity: The (homeUserId, homeInstance) pair is globally unique. Local users have homeInstance = NULL and homeUserId = NULL. Federated users are represented as replicated user stubs -- minimal user records with passwordHash = '!federation-replicated' (bcrypt never produces this value, so login is impossible).
Trust model: Symmetric shared-secret HMAC. Both peers share the same 256-bit secret. Events are attributed to users by homeUserId + homeInstance in the payload, with authority checks verifying the source instance matches the claimed origin of the acting user.
1. Peer Handshake & Discovery
2-Phase Flow
Phase 1 -- Initiate (POST /api/federation/peer/initiate)
- Auth: JWT + admin role required
- Validates
remoteOriginis a well-formed HTTP(S) URL viavalidateOrigin() - Prevents self-peering (
localOrigin === remoteOrigin) - Handles existing peers: active -> return 200, pending -> return 409, revoked -> delete and re-initiate
- Generates HMAC secret:
generateHmacSecret()->randomBytes(32).toString('hex')(256-bit) - Generates challenge:
randomBytes(16).toString('hex')(128-bit, currently unused by acceptor) - Creates local peer record with
status='pending' - POSTs to
{remoteOrigin}/api/federation/peer/acceptwith{ sourceOrigin, challenge, hmacSecret } - Timeout: 10 seconds (
AbortSignal.timeout) - On remote acceptance: updates local peer to
status='active', setslastSeenAt - On failure: deletes pending peer, returns 502 (network error) or 504 (timeout)
Phase 2 -- Accept (POST /api/federation/peer/accept)
- Auth: none (first contact -- no JWT, no HMAC)
- Rate-limited: 10 requests per minute per IP (in-memory sliding window, buckets cleaned every 60s)
- Validates
sourceOrigin,challenge, andhmacSecretfrom body - Handles existing peers: active -> return 200 (idempotent), revoked -> return 403, pending -> update with new secret and activate
- New peer: creates record with provided
hmacSecret, setsstatus='active' - Returns
{ accepted: true }on success
Secret Storage
Both instances store the same HMAC secret. The initiating instance generates it and sends it in the accept request. There is no secret rotation mechanism -- the secret persists until the peer is revoked and re-initiated.
Peer Status Lifecycle
initiate
(none) ──────────► pending ──────────► active
│
10+ consecutive │ delivery failures
failures ▼
unreachable
│
health check OK │
▼
active
│
admin revoke │
▼
revoked ──► (delete) ──► re-initiate
| Status | Outbox delivery | Health check | Relay accepts | Re-initiation |
|---|---|---|---|---|
active |
Yes | No | Yes | No (returns existing) |
pending |
No | No | No | No (returns 409) |
unreachable |
No (entries wait) | Yes (1h interval) | Yes (resets to active) | No |
revoked |
No (entries purged) | No | No (returns 403) | Yes (old record deleted) |
PEER_UNREACHABLE_THRESHOLD
Defined in federationWorker.ts:45 as 10. After 10 consecutive delivery failures for a peer, the worker sets status = 'unreachable'. The health check worker (1h interval) pings GET /api/instance/info on unreachable peers and reverts to active on success.
Admin Endpoints
| Endpoint | Method | Auth | Purpose |
|---|---|---|---|
/api/federation/peer/initiate |
POST | JWT + admin | Start peering handshake |
/api/federation/peer/accept |
POST | None (rate-limited) | Accept incoming handshake |
/api/federation/peers |
GET | JWT + admin | List all peers (secret excluded) |
/api/federation/peers/:id |
DELETE | JWT + admin | Revoke peer, purge outbox |
2. HMAC Request Authentication
Signing Format
HMAC-SHA256(secret, "${timestamp}.${requestBody}")
Where timestamp is Date.now() (Unix milliseconds) and requestBody is the JSON string.
HTTP Headers
| Header | Format | Example |
|---|---|---|
X-Federation-Signature |
sha256=<hex> |
sha256=a1b2c3... |
X-Federation-Origin |
Full URL | https://nova.ddns.net |
X-Federation-Timestamp |
Unix ms string | 1711619400000 |
Content-Type |
application/json |
-- |
Verification (federationAuth.ts:verifySignature)
- Validate inputs: reject empty/missing body, signature, or secret
- Timestamp window:
Math.abs(Date.now() - timestamp) <= maxAgeMs(default 15 minutes) - Recompute:
HMAC-SHA256(secret, "${timestamp}.${body}") - Constant-time comparison:
crypto.timingSafeEqualon hex-decoded buffers - Length check: mismatched buffer lengths are rejected before
timingSafeEqual
Replay Attack Prevention
The 15-minute timestamp window prevents replaying old requests. However, there is no nonce or sequence number -- a valid request can be replayed within the 15-minute window. See Known Issues.
Inbound Verification Flow (POST /api/federation/relay)
parseFederationHeaders()extracts origin, timestamp, signature from headers- Look up peer by
origininfederation_peers-- must exist and bestatus = 'active' - Re-serialize request body to JSON:
JSON.stringify(request.body) verifySignature(bodyString, signature, peer.hmacSecret, timestamp)-- reject if false
Important: The body is re-serialized server-side. This means Fastify's JSON parsing and re-stringification must produce identical output to the sender's JSON.stringify. In practice this works because both sides use standard JSON.stringify with no custom replacers.
3. Identity Resolution
Functions
resolveLocalUser(homeUserId, db) -- federation.ts:878
- Read-only lookup. Returns
undefinedif not found. - Matches:
(users.homeUserId = homeUserId)OR(users.id = homeUserId AND homeInstance IS NULL) - Excludes deleted users (
isDeleted = 0) - When multiple candidates exist: prefers the one with
homeUserIdset (replicated stub) over a local ID match - Use when: Optional lookups where null is acceptable (member_remove, reaction processing, friend_remove)
resolveOrCreateReplicatedUser(homeUserId, homeInstance, db) -- federation.ts:912
- Calls
resolveLocalUserfirst. If found, returns it. - If not found, creates a stub with:
username:{homeUserId}@{domain}(domain extracted from homeInstance URL)passwordHash:'!federation-replicated'homeInstance: the full URL passed inhomeUserId: the remote user's home ID- Collision-safe: appends
_1,_2, ...,_10suffix if username exists; after 10 attempts, uses_<random hex>
- Use when: You MUST have a valid user ID (setting
ownerId, insertingdm_members, creating messages)
hydrateReplicatedUserProfile(user, profile, db) -- federation.ts:2041
- Updates replicated stubs only (
homeInstancemust be set) - Only updates null/empty fields (preserves manually-set local values)
- Exception: avatar/banner are overwritten if the current value is a bare filename (not an absolute URL)
- Resolves bare filenames to
{homeInstance}/api/uploads/{filename}absolute URLs - Sets
displayNamefromprofile.displayName || profile.username-- ensures federated users show a human-readable name instead ofuser@instance
Critical Rule
Any code path that sets ownerId, creates a dm_members row, or inserts a message MUST use resolveOrCreateReplicatedUser. Using resolveLocalUser with a ?? null fallback has caused data corruption (see Known Issues: ownerId nulling).
Origin Normalization
Two formats exist in the database:
| Location | Format | Example |
|---|---|---|
users.home_instance |
Bare domain OR full URL | nova.ddns.net or https://nova.ddns.net |
federation_peers.origin |
Full URL | https://nova.ddns.net |
getOurOrigin() return |
Full URL | https://orbit.ddns.net |
resolveOrCreateReplicatedUser stores |
Full URL (passed through) | https://nova.ddns.net |
| Auth registration stores | Bare domain | nova.ddns.net |
The inconsistency exists because:
resolveOrCreateReplicatedUserstoreshomeInstanceas-is from the relay event (full URL)- The auth registration path (
/api/auth/registerwithhomeInstanceparam) validates as bare domain only (regex:/^[a-zA-Z0-9._-]+$/) - Relay event payloads populate
homeInstancefromgetOurOrigin()(full URL) or fromuser.homeInstance || getOurOrigin()(which falls back to full URL)
Normalization pattern used in code:
const normalized = homeInstance.startsWith('http') ? homeInstance : `https://${homeInstance}`;
Locations where normalization is applied:
getGroupDmTargetOrigins()(federationOutbox.ts:294) -- normalizes before comparing toourOrigindm.ts:655--isLocalMemberbroadcast filter checks both formatsdm.ts:743-- normalizes target homeInstance before peer origin comparison
Locations with potential mismatch (see Known Issues):
federation.ts:1278--memberUser?.homeInstance === sourceInstance-- compares stored homeInstance (possibly bare domain) againstsourceInstance(full URL from relay request header)federationOutbox.ts:376-379--getFriendEventTargetscomparesfromHomeInstanceagainstourOriginwithout normalization. The passed values come fromuser.homeInstance || domainOriginwheredomainOrigin = getOurOrigin(). IfhomeInstanceis a bare domain,homeInstance !== ourOriginis true, so the bare domain gets added to targets, butqueueOutboxEventthen fails to match it againstfederation_peers.originfederationWorker.ts:424--user.homeInstance === ourOrigininhandleSizeRejection. Bare domain homeInstance won't match, potentially including a user inaffectedUserIdswho shouldn't be (minor).federation.ts:2388--from.homeInstance === ourOrigininprocessFriendAddEventdetermines which user is "local" for broadcasting. Thefrom.homeInstancecomes from the relay event payload, which should be a full URL, so this comparison works correctly in practice.federation.ts:2447-- same pattern inprocessFriendRemoveEvent
4. DM Message Relay
1-on-1 DMs
Outbound (origin instance):
- Message created via REST (
POST /api/dm/:id/messages) or WS (dm_message_create) queueDmRelay(message, channelId, 'create')called fromdm.ts/events.tsbuildRelayPayload()constructs the message portion withhomeUserId,homeInstance,content,replyToId,editedAt,createdAtgetDmParticipants(channelId)resolves all members to(homeUserId, homeInstance)pairs with profile snapshotsgetGroupDmTargetOrigins(channelId)returnsundefined(no owner -> broadcast to all)queueOutboxEvent(messageId, channelId, 'create', payload, undefined)-> queued to ALL active peers
Inbound (receiving instance -- processCreateEvent):
- Validate:
event.messageandevent.participants(>= 2) required - Dedup: check
(sourceInstance, sourceMessageId)-- reject if exists - Resolve ALL participants via
resolveOrCreateReplicatedUser, hydrate profiles - No
event.federatedId-> 1-on-1 path - Compute deterministic
federatedId = SHA256(sorted([homeUserIdA, homeUserIdB])).slice(0, 32) findOrCreateDmChannel(federatedId, [localUserA.id, localUserB.id], db):- Find by
federatedIdindm_channels - If exists: ensure both users are members (idempotent insert)
- If not: create channel with
federatedId, add both members
- Find by
- Insert
dm_messageswithsourceInstanceandsourceMessageId - Process attachments (see File Replication)
- Broadcast
dm_message_createdto local members, skipping members whosehomeInstance === sourceInstance(they already have the original)
Group DMs
Outbound (origin instance): Same as 1-on-1 except:
getGroupDmTargetOrigins(channelId)returns a list of peer origins that have at least one participant- Normalizes
homeInstanceto full URL before comparison queueOutboxEventreceivestargetPeerOriginsand only queues to those peers- Payload includes
federatedId(random UUID assigned at channel creation)
Inbound (receiving instance -- processCreateEvent):
event.federatedIdpresent -> group DM path- Find channel by
federatedId-- must already exist (bootstrapped by priormember_add) - If not found -> reject with
channel_not_found - Insert message, broadcast to local members
Federated ID Generation (federationOutbox.ts:computeFederatedId)
// 1-on-1: deterministic 32-char hex hash
const sorted = [homeUserIdA, homeUserIdB].sort();
return sha256(sorted.join(':')).slice(0, 32);
// Group: random 36-char UUID with dashes
return crypto.randomUUID();
The format difference (32-char hash vs 36-char UUID) is used by the self-healing migration to detect channel type independently of owner_id.
Message Deduplication
Every relayed message is stored with:
source_instance: the relay request'ssourceInstanceheader valuesource_message_id: theevent.messageId(original message ID on source instance)
The (source_instance, source_message_id) pair is checked before insertion. Duplicates are rejected with reason 'duplicate'. A unique partial index enforces this at the DB level: idx_dm_messages_source_unique ON dm_messages(source_instance, source_message_id) WHERE source_instance IS NOT NULL.
5. Outbox & Relay Pipeline
Event Queuing (federationOutbox.ts:queueOutboxEvent)
Trigger (API/WS handler)
-> isFederationRelayEnabled()? No -> return silently
-> Fetch active peers from federation_peers
-> Filter to targetPeerOrigins (if specified) -- EXACT string match against peer.origin
-> If zero peers match -> return silently (KNOWN ISSUE: silent event dropping)
-> For each peer, in a transaction:
-> Check for existing outbox entry by (peerId, entityId)
-> COALESCE:
- delete + existing create -> delete both (net: never relayed)
- update + existing create -> update payload, keep 'create' eventType
- update + existing update -> update payload and eventType
- no existing -> insert new entry
-> TTL: now + (relayTtlDays * 86400000)
Coalescing Rules (per-peer, per-entity)
| Incoming | Existing | Result |
|---|---|---|
delete |
create |
Entry removed (message was never relayed) |
update |
create |
Payload updated, keeps create type (peer gets full message) |
update |
update |
Payload updated, type becomes latest |
delete |
update |
Payload updated, type becomes delete |
| any | none | New entry inserted |
Outbox Delivery Worker (federationWorker.ts:processOutboxTick)
Interval: 10 seconds (OUTBOX_INTERVAL_MS)
Batch size: 50 (OUTBOX_BATCH_LIMIT)
Timeout: 30 seconds per request (OUTBOX_FETCH_TIMEOUT_MS)
- Query entries where
nextRetryAt <= nowjoined with active peers, ordered bycreatedAt ASC, limit 50 - Group by peer
- For each peer, reconstruct
FederationRelayEvent[]from stored payloads:- Parse JSON payload
- Copy fields:
federatedId,participants,message,reactions,reaction,membership,ownership,group,friendship, file_rejected fields - Set
eventType,contextType,messageId,dmChannelId,encryptionVersion,timestamp
- Build
FederationRelayRequestwithversion: 1,sourceInstance: ourOrigin - Sign with
buildFederationHeaders(body, peerHmacSecret, ourOrigin) - POST to
{peerOrigin}/api/federation/relay - On success (200):
- Delete accepted entries from outbox (matched by
entityId->outboxId) - Log rejected entries (remain in outbox for retry)
- Store
result.maxUploadSizeon peer record - Update peer:
lastSeenAt = now,consecutiveFailures = 0
- Delete accepted entries from outbox (matched by
- On failure (non-200 or network error):
handleOutboxDeliveryFailure():- Increment
attemptsper entry, computenextRetryAt = now + backoff - Increment peer
consecutiveFailures, setlastFailureAt - If
consecutiveFailures >= PEER_UNREACHABLE_THRESHOLD (10)-> mark peerunreachable
- Increment
Retry Backoff Schedule
| Attempt | Delay |
|---|---|
| 1 | 30 seconds |
| 2 | 1 minute |
| 3 | 5 minutes |
| 4 | 15 minutes |
| 5 | 1 hour |
| 6 | 6 hours |
| 7+ | 24 hours (cap) |
Relay Request/Response Format
Request:
interface FederationRelayRequest {
version: 1;
sourceInstance: string; // Full URL, e.g., "https://nova.ddns.net"
events: FederationRelayEvent[]; // Max 50 per batch
}
Response:
interface FederationRelayResponse {
accepted: string[]; // messageIds successfully processed
rejected: Array<{
messageId: string;
reason: string; // e.g., 'duplicate', 'unknown_message', 'missing_participants'
}>;
maxUploadSize: number; // This instance's max upload size in bytes
}
Inbound Relay Dispatch (POST /api/federation/relay)
Body limit: 10 MB. Max 50 events per batch.
| eventType | Processor | contextType |
|---|---|---|
create |
processCreateEvent |
dm |
update |
processUpdateEvent |
dm |
delete |
processDeleteEvent |
dm |
reaction_add |
processReactionAddEvent |
dm |
reaction_remove |
processReactionRemoveEvent |
dm |
member_add |
processMemberAddEvent |
dm |
member_remove |
processMemberRemoveEvent |
dm |
ownership_transfer |
processOwnershipTransferEvent |
dm |
friend_request_create |
processFriendRequestCreateEvent |
friend |
friend_request_update |
processFriendRequestUpdateEvent |
friend |
friend_request_cancel |
processFriendRequestCancelEvent |
friend |
friend_add |
processFriendAddEvent |
friend |
friend_remove |
processFriendRemoveEvent |
friend |
file_rejected |
processFileRejectedEvent |
dm |
After processing all events, the relay endpoint updates the peer's lastSeenAt and resets consecutiveFailures, then returns accepted/rejected arrays plus maxUploadSize.
6. Group DM Lifecycle over Federation
member_add (processMemberAddEvent -- federation.ts:1618)
Required fields: event.federatedId, event.membership.user
Two paths:
Bootstrap path (channel does not exist locally by federatedId):
- Requires
event.groupmetadata (owner + full member roster) - Creates
dm_channelsrow withfederatedId,ownerId(resolved viaresolveOrCreateReplicatedUser),ownerHomeUserId,ownerHomeInstance - Adds ALL roster members from
event.group.members(each resolved viaresolveOrCreateReplicatedUser) - Sends
dm_channel_createdto local-only members (home instance matchesgetOurOrigin(), with normalization for bare domain) - Sets
bootstrapped = trueto skip redundant system messages and member_add broadcasts below
Incremental path (channel already exists):
- Validates authority:
sourceInstance === channel.ownerHomeInstance(only owner's instance can add) - Cancels soft-delete if channel was pending GC
- Resolves added user via
resolveOrCreateReplicatedUser - Enforces max 10 members
- Inserts
dm_membersrow (idempotent -- skip if exists) - Inserts system message, broadcasts
dm_member_addedto local WebSocket clients
member_remove (processMemberRemoveEvent -- federation.ts:1825)
- Find channel by
federatedId-- if not found, accept idempotently - Validate authority: owner's instance for kicks (
reason !== 'leave'), any instance for self-leave - Resolve user via
resolveLocalUser-- if not found, accept idempotently - Insert system message (before deletion, so broadcast includes the leaving user)
- Delete
dm_membersrow, clean upread_states - Broadcast
dm_member_removedto remaining local members - If zero members remain -> soft-delete channel (
deletedAt = now)
ownership_transfer (processOwnershipTransferEvent -- federation.ts:1938)
- Find channel by
federatedId-- if not found, accept idempotently - Validate authority:
sourceInstance === channel.ownerHomeInstance - Resolve new owner via
resolveOrCreateReplicatedUser(neverresolveLocalUser-- must guarantee valid ID) - Update
dm_channels:ownerId,ownerHomeUserId,ownerHomeInstance - Broadcast
dm_owner_updatedWebSocket event - Insert system message with previous owner as actor
Local-Only Broadcast Principle
Users connected to multiple instances must see each DM channel exactly once (from their home instance). All structural broadcasts (dm_channel_created, system messages) filter to local members only:
const isLocalMember = (u: { homeInstance?: string | null }) =>
!u.homeInstance || !domainOrigin ||
u.homeInstance === domainOrigin ||
`https://${u.homeInstance}` === domainOrigin;
Does NOT apply to: Regular DM messages (dm_message_created for user messages). These broadcast to all local dm_members regardless of home instance.
System Messages
System messages (type = 'system' in dm_messages) are instance-local -- they are NOT relayed via federation. Each instance creates its own when processing events.
| Event | Content JSON | Actor (userId) |
|---|---|---|
member_added |
{event, targetUserId, targetDisplayName} |
User who added them |
member_removed |
{event, targetUserId, targetDisplayName, reason} |
User who left/was removed |
owner_changed |
{event, newOwnerId, newOwnerDisplayName} |
Previous owner |
Outbound Queuing (Origin Instance -- dm.ts)
When a group DM is created or modified locally, the origin instance queues federation events:
Group DM creation (POST /api/dm/group):
- Iterates each remote target user (those with
homeInstance !== domainOrigin) - Builds a
member_addevent per remote user, carrying the full roster inevent.group - Computes
finalTargetsby starting fromgetGroupDmTargetOrigins()and adding the new member's normalized homeInstance - Calls
appendMutationLog+queueOutboxEventper event
Add member to existing group (POST /api/dm/:id/members):
- Same structure as creation -- builds
member_addwith full group metadata - Normalizes new member's homeInstance to full URL before including in targets
Leave group (DELETE /api/dm/:id/members):
- Computes
fedTargetOriginsbefore deleting the member (so the leaving user's peer is still included) - Queues
member_removeevent withreason: 'leave'
Ownership transfer (PATCH /api/dm/:id):
- Queues
ownership_transferevent withpreviousOwnerandnewOwner
7. File Replication
Outbound (origin instance)
When queueDmRelay constructs the relay payload, each attachment gets a sourceUrl:
sourceUrl: `${getOurOrigin()}/api/uploads/${attachment.filename}`
Inbound (receiving instance -- processCreateEvent)
- For each attachment in
event.message.attachments:- SSRF check:
isUrlFromPeer(sourceUrl, peerOrigin)-- hostname of sourceUrl must match peer origin hostname - Create
attachmentsrow withfilename = sourceUrl(remote URL as interim filename) - Queue
federation_file_queueentry withstatus = 'pending',expiresAt = now + 30 days
- SSRF check:
- Initial WebSocket broadcast uses sourceUrl directly (frontend's
AttachmentRendererdetectshttpprefix)
File Download Worker (federationWorker.ts:processFileQueueEntry)
Interval: 30 seconds. Batch: 5 files. Timeout: 60 seconds per download.
- SSRF protection: validate sourceUrl hostname matches peerOrigin hostname
- Pre-download size check against
maxUploadSizeBytesfrom instance settings - Download via
fetchwith streaming pipeline to disk (Readable.fromWeb->fs.createWriteStream) - Post-download size verification (defense in depth)
- Generate thumbnail via
sharp(same as local upload flow) - Update
attachmentsrow:filename = localFilename,size,thumbnailFilename - Fallback: if no existing attachment row was found (legacy queue entry), insert a new one
- Mark file queue entry as
completedwithtargetFilename - Broadcast
dm_message_updatedto refresh client-side attachment display
Size Rejection Flow (handleSizeRejection)
When a file exceeds the local instance's size limit:
- Mark file queue entry as
rejectedwithreason = 'size_limit_exceeded' - Update local attachment:
federationStatus = 'remote',federationMeta= source info JSON - Determine affected local users (native to this instance --
!user.homeInstance || user.homeInstance === ourOrigin) - Queue
file_rejectedreverse relay event to the sender's instance (sourceInstance) - Broadcast
dm_message_updatedlocally so clients see the 'remote' badge
Inbound file_rejected (processFileRejectedEvent -- federation.ts:2458)
When the origin instance receives a file_rejected event:
- Find local message by
event.messageId(the original local message ID) - Match attachment by
sourceFilenameor fallback to single attachment - Resolve
affectedUserIds(homeUserIds) to local replicated user stubs - Merge rejection info into
federationMeta(accumulates from multiple peers) - Set
federationStatus = 'remote_partial' - Broadcast
dm_message_updated+ targetedfederation_file_rejectedtoast to message author
Federation Status on Attachments
| Status | Meaning |
|---|---|
null |
Local upload, no federation involvement |
'local' |
Successfully downloaded from peer |
'remote' |
Rejected (size limit), federationMeta has source instance info |
'remote_partial' |
Rejected by some peers, federationMeta has per-user rejection array |
File Download Retry
Uses the same backoff schedule as outbox delivery. Max attempts: 10 (MAX_FILE_ATTEMPTS). After exceeding max attempts: status = 'failed', rejectionReason = 'max_attempts_exceeded'.
8. Friend Relay
Event Flow (social.ts)
| User Action | Federation Event | Authority Check |
|---|---|---|
| Send friend request | friend_request_create |
from.homeInstance === sourceInstance |
| Accept/decline request | friend_request_update |
to.homeInstance === sourceInstance |
| Cancel outgoing request | friend_request_cancel |
from.homeInstance === sourceInstance |
| Accept creates friendship | friend_add |
to.homeInstance === sourceInstance |
| Remove friend | friend_remove |
Either side's instance |
Target Resolution (getFriendEventTargets)
Computes which peer origins need the event. Compares fromHomeInstance and toHomeInstance against getOurOrigin(). Known issue: no normalization -- bare domain homeInstance will not match full URL ourOrigin, causing the bare domain to be passed as a target. However, queueOutboxEvent then fails to match it against federation_peers.origin (full URL), silently dropping the event.
Context ID
Friend events use a deterministic context ID: friend:${sorted[homeUserIdA, homeUserIdB].join(':')}.
Outbound Payload Construction
Each friend endpoint builds a FederationRelayEvent with:
contextType: 'friend'friendshippayload containingfromandtoasFederationRelayParticipantobjectsfromProfileand/ortoProfilesnapshots (FederationRelayProfileSnapshot)entityIdformatted asfriend_req:${sorted_ids}:${timestamp}(for requests) orfriend_remove:${sorted_ids}:${timestamp}
The full event payload is stored in both appendMutationLog (for sync) and queueOutboxEvent (for delivery).
Inbound Processing
processFriendRequestCreateEvent (federation.ts:2082):
- Authority check:
from.homeInstance !== sourceInstance-> reject - Resolve sender via
resolveOrCreateReplicatedUser+ hydrate profile - Resolve recipient via
resolveLocalUser(must be native to this instance) - Idempotency: if already friends or pending request exists, accept as no-op
- Create
friend_requestsrow, broadcastfriend_request_receivedto recipient
processFriendRequestUpdateEvent (federation.ts:2178):
- Authority check:
to.homeInstance !== sourceInstance-> reject - Resolve sender (original requester) via
resolveLocalUser(must exist locally) - Resolve recipient (acceptor/decliner) via
resolveOrCreateReplicatedUser - Find pending request, update status
- Broadcast
friend_request_acceptedorfriend_request_declinedto the original sender
processFriendRequestCancelEvent (federation.ts:2254):
- Authority check:
from.homeInstance !== sourceInstance-> reject - Both users must exist locally. If not, accept idempotently.
- Delete the pending friend request. Broadcast
friend_request_cancelledto recipient.
processFriendAddEvent (federation.ts:2318):
- Authority check:
to.homeInstance !== sourceInstance-> reject - Resolve both users via
resolveOrCreateReplicatedUser+ hydrate profiles - Insert
friendsrow (idempotent) - Auto-resolve any pending
friend_requeststo'accepted'(handles out-of-order delivery) - Determine which user is local (
from.homeInstance === ourOrigin) and broadcastfriend_request_accepted
processFriendRemoveEvent (federation.ts:2404):
- Authority check: either
from.homeInstanceorto.homeInstancemust besourceInstance - Both users resolved via
resolveLocalUser. If not found, accept idempotently. - Delete
friendsrow in both directions - Determine local user (whose
homeInstanceis NOT the source) and broadcastfriend_removed
9. Profile Sync
Profile sync uses two mechanisms that operate independently:
S2S Profile Hydration (Server-side)
When relay events carry FederationRelayProfileSnapshot data:
processCreateEvent: hydrates participant profiles on message relayprocessFriendRequestCreateEvent/processFriendAddEvent: hydrates friend profiles
hydrateReplicatedUserProfile only fills null/empty fields. Avatar/banner are overwritten only if the current value is not an absolute URL (catches stale bare filenames).
Client-side LWW Sync (profileSync.ts)
Operates via the web client, not S2S relay:
- On connect to a remote instance, compares
profileUpdatedAttimestamps - If home is newer: pushes profile to remote (re-uploads avatar/banner)
- If remote is newer: pulls from remote to home, then relays to all other remotes
- Incremental:
syncProfileUpdateToRemotespushes partial updates after local profile edits
This is a client-driven mechanism -- it only runs when a user is actively connected to multiple instances. It does not use the relay pipeline or outbox.
10. Reaction Relay
Outbound
Reactions are queued by WS event handlers in events.ts:
dm_reaction_add->queueOutboxEvent(reactionId, channelId, 'reaction_add', payload, targetOrigins)dm_reaction_remove->queueOutboxEvent(messageId, channelId, 'reaction_remove', payload, targetOrigins)
Payload includes userId, homeUserId, emoji, createdAt, plus messageId and messageHomeInstance for cross-instance message resolution.
The mutation log entry for reactions stores a simpler payload (no messageId/messageHomeInstance), while the outbox entry carries the full reaction payload including those fields.
Inbound
processReactionAddEvent (federation.ts:1480):
- Resolve message via
resolveLocalDmMessage(canonicalMessageId, messageHomeInstance, sourceInstance, db):- If
messageHomeInstance === getOurOrigin()-> find by local ID (the message originated here) - Otherwise -> find by
(messageHomeInstance || sourceInstance, canonicalMessageId)tracking -- usesmessageHomeInstancewhen available (correct origin in 3-instance relay), falls back tosourceInstance
- If
- Resolve reacting user via
resolveLocalUser(must already exist) - Dedup: check existing reaction by
(dmMessageId, userId, emoji) - Insert
dm_reactions, broadcastreaction_addedto local clients
processReactionRemoveEvent (federation.ts:1561):
- Same resolution logic
- Delete matching reaction, broadcast
reaction_removedif changes > 0
11. Initial Sync
runInitialSyncForNewPeers() (federationWorker.ts:739)
Triggered once at server startup (async, non-blocking). Finds peers with status = 'active' and lastSyncedAt = 0.
For each unsynced peer:
- DM sync pass: Paginate through
POST {peerOrigin}/api/federation/syncwithsinceTimestamp = 0,limit = 100 - Self-POST: Relay received events by POSTing to
{ourOrigin}/api/federation/relay-- this routes through the standard inbound processing - Friend sync pass: Same pagination with
contextType: 'friend' - Update
lastSyncedAt = Date.now()after completion - On failure: don't update
lastSyncedAt-- retried on next startup
Sync Endpoint (POST /api/federation/sync)
HMAC-authenticated. Returns events from the federation_mutation_log.
Request:
{ sinceTimestamp: number, dmChannelId?: string, federatedId?: string, contextType?: 'dm'|'friend', limit?: 1-500 }
Response:
{ events: FederationRelayEvent[], hasMore: boolean, checkpoint: number }
DM sync:
- Queries all
dm_channelswith non-nullfederated_id(not soft-deleted) - Joins
federation_mutation_logwithdm_messagesto reconstruct events - Only returns locally-created messages (
source_instance IS NULLvia the LEFT JOIN) - Handles delete mutations separately (message rows don't exist for deletes)
- For create/update: fetches current message state from DB, builds full relay event with attachments and participants
- Membership/friend mutations store the full event payload in the mutation log, so they are returned directly
Friend sync:
- Queries
federation_mutation_log WHERE context_type = 'friend' - Returns stored payloads directly (friend events carry their complete data)
Known Bug: DNS Hairpin Self-POST
runInitialSyncForNewPeers POSTs to {ourOrigin}/api/federation/relay where ourOrigin = getOurOrigin(). In production, ourOrigin is https://{DOMAIN}, e.g., https://nova.ddns.net. This means the server makes an HTTP request to itself through the public DNS and reverse proxy (Caddy). This works but:
- Adds unnecessary network round-trip latency
- Fails if DNS hairpin is not supported by the network
- Fails if the server is behind NAT without hairpin NAT configured
A direct function call to the relay processing logic would be more robust.
12. DM Calls over Federation
DM calls use LiveKit for WebRTC signaling and media transport. The call lifecycle is managed entirely via WebSocket events (dm_call_start, dm_call_accept, dm_call_reject, dm_call_end in ws/events.ts).
Current state: DM calls do NOT work across federated instances.
The call state machine is local to a single server instance -- there is no federation relay for call events. When user A on instance 1 calls user B on instance 2:
- The
dm_call_incomingevent is sent viaconnectionManager.sendToUser(targetUser.id, ...)which only broadcasts to WebSocket connections on the local instance - User B's replicated stub exists on instance 1, but user B is connected via WebSocket to instance 2
- The call event is never delivered
LiveKit tokens are also instance-local (/api/livekit/token requires JWT auth for the local instance).
13. Background Workers
All workers are started by startFederationWorkers() on server boot and stopped by stopFederationWorkers() on shutdown. Each worker uses setTimeout chains (not setInterval) with abort controllers for graceful shutdown.
| Worker | Interval | Batch | Timeout | Source |
|---|---|---|---|---|
| Outbox delivery | 10s | 50 | 30s | processOutboxTick |
| File download | 30s | 5 | 60s | processFileQueueTick |
| Health check | 1h | all unreachable | 10s | processHealthCheckTick |
| Janitor | 1h | -- | -- | runFederationJanitor (sync) |
| Initial sync | Once at startup | -- | 30s per page | runInitialSyncForNewPeers |
Janitor Cleanup (storageJanitor.ts:runFederationJanitor)
| Target | Condition | Retention |
|---|---|---|
federation_outbox |
expiresAt < now |
Configurable via federationRelayTtlDays (default 30) |
federation_mutation_log |
mutatedAt < (now - 90 days) |
90 days |
federation_file_queue (completed) |
createdAt < (now - 7 days) |
7 days |
federation_file_queue (any) |
expiresAt < now |
30 days (set at queue time) |
dm_channels (soft-deleted) |
deletedAt < (now - 24h) |
24-hour grace period |
DM channel hard-delete cascades: reactions, embeds, attachments (DB rows + disk files), messages, members, outbox entries, mutation log entries, file queue entries.
14. Settings Cache
federationOutbox.ts caches federationRelayEnabled and federationRelayTtlDays from instance_settings for 30 seconds (CACHE_TTL_MS). This prevents repeated DB reads on every message send. The cache is invalidated by TTL only -- there is no explicit cache bust on settings change.
Relevant settings in instance_settings:
| Column | Default | Purpose |
|---|---|---|
federation_relay_enabled |
1 | Master toggle for all federation relay |
federation_relay_ttl_days |
30 | Outbox entry TTL |
max_upload_size_bytes |
null (uses config.maxUploadSize) |
File download size limit |
15. Client-Side Identity Helpers (identity.ts)
The frontend needs to resolve federated identities for display purposes:
parseFederatedUsername(username) -- splits "youruser@nova.ddns.net" into {baseName: "youruser", domain: "nova.ddns.net"}.
isSelf(user, homeUser) -- determines if a user object is the logged-in user or their replicated stub. Uses cascading checks: same ID, known self-ID set, homeInstance + baseName match.
canonicalUserMatch(a, b) -- federation-safe check for whether two user objects represent the same person. Cascades through: same local ID, homeUserId cross-match, username + homeInstance fallback.
resolveDisplayIdentity(user, homeUser) -- returns homeUser for display if user is a replicated stub of homeUser, enabling consistent avatars and display names across instances.
Cross-instance self-ID registry: registerSelfId(id) / clearSelfIds() track all Snowflake IDs belonging to the current user across connected instances, populated from WS ready events.
16. Self-Healing Migrations (migrate.ts)
The migration system includes several data integrity checks that run on every server startup:
Group DM ownerId repair:
Detects group DMs with UUID-format federated_id (length 36, matches ________-____-____-____-____________) but NULL owner_id. Restores the owner from the first remaining member or from owner_home_user_id/owner_home_instance. Root cause: a bug in processOwnershipTransferEvent (fixed in cd7aff0) could set ownerId = NULL via resolveLocalUser fallback.
Federated ID backfill:
Finds 1-on-1 DM channels without federated_id, computes deterministic SHA-256 hash from home user IDs, and sets it. Also detects relay-created duplicate channels with the same federated_id and merges messages into the oldest channel.
Duplicate channel merge:
Finds federated_id values appearing on multiple channels and merges them into the oldest, moving messages, members, and cleaning up the duplicates.
Mutation log backfill:
If the federation_mutation_log table exists but is empty, populates it with create entries for all existing DM messages where source_instance IS NULL (locally-created messages).
Known Issues
1. Origin Format Inconsistency (PARTIALLY FIXED)
Root cause: users.home_instance stores both bare domains (from auth registration: nova.ddns.net) and full URLs (from resolveOrCreateReplicatedUser: https://nova.ddns.net). federation_peers.origin and getOurOrigin() always use full URLs.
Fixed locations:
getGroupDmTargetOrigins()normalizes before comparison (federationOutbox.ts:294)isLocalMemberindm.ts:655checks both formats- DM member_add target resolution in
dm.ts:743normalizes
Remaining unpatched comparisons:
federation.ts:1278--memberUser?.homeInstance === sourceInstanceinprocessCreateEvent. If the member'shomeInstanceis a bare domain andsourceInstanceis a full URL, this comparison fails. Result: the member receives the message even though they should be skipped (minor -- causes duplicate delivery, not data loss).getFriendEventTargets()(federationOutbox.ts:376-379) -- comparesfromHomeInstance/toHomeInstanceagainstgetOurOrigin()without normalization. When a local user'shomeInstanceis null, the fallbackdomainOrigin(full URL) is used, which works correctly. But when a user hashomeInstancestored as a bare domain and that domain is this instance (e.g., an old replicated stub), the comparisonbareDomain !== fullUrlevaluates to true, incorrectly including the local instance as a target, which then silently drops inqueueOutboxEvent.federationWorker.ts:424--user.homeInstance === ourOrigininhandleSizeRejection. Bare domain homeInstance won't match, potentially including a user inaffectedUserIdswho shouldn't be (minor).
2. Duplicate User Stubs
The same remote user can have multiple replicated records. Code paths that call resolveOrCreateReplicatedUser:
processCreateEvent(for each participant)processMemberAddEvent(bootstrap roster + incremental add + owner + addedBy)processOwnershipTransferEvent(new owner)processFriendRequestCreateEvent(sender)processFriendRequestUpdateEvent(recipient)processFriendAddEvent(both users)
resolveLocalUser (called first by resolveOrCreateReplicatedUser) matches on homeUserId OR (id = homeUserId AND homeInstance IS NULL). If a user was created via auth registration (with homeInstance as bare domain) and later via relay (with homeInstance as full URL), resolveLocalUser may not find the first record if the IDs differ. The collision-safe username suffix ensures the insert succeeds, but now two stubs exist for the same person.
3. Silent Failures in queueOutboxEvent
queueOutboxEvent returns silently (no error, no log) when:
- Federation relay is disabled
- Zero active peers exist
targetPeerOriginsfilter produces zero matches (origin format mismatch)
The third case is the most dangerous -- it looks like the event was queued but nothing was actually written. This has been the root cause of events silently disappearing for group DMs and friend events.
4. Trust Model Analysis
| Threat | Mitigation | Gap |
|---|---|---|
| Peer impersonation | X-Federation-Origin is verified against federation_peers.origin |
An attacker who compromises the HMAC secret can impersonate the peer |
| User attribution fraud | Authority checks: e.g., from.homeInstance !== sourceInstance rejects events where the acting user doesn't belong to the source instance |
The check is string equality on homeInstance from the payload, which the sender controls. A malicious peer could claim any user belongs to them by setting homeInstance to their own origin. |
| Event flooding | Outbox batches limited to 50 events. /api/federation/peer/accept rate-limited to 10/min. |
No rate limit on /api/federation/relay itself. A peer could send unlimited relay requests. |
| Replay attacks | 15-minute timestamp window | No nonce -- valid requests can be replayed within the window |
| Message content manipulation | None | A compromised peer can forge message content attributed to any user on their instance |
5. DNS Hairpin Self-POST Bug
runInitialSyncForNewPeers (federationWorker.ts:792-797) POSTs received sync events to {ourOrigin}/api/federation/relay via public DNS. This adds unnecessary latency and fails when DNS hairpin is not configured. The function should call the relay processing logic directly instead of making an HTTP request to itself.
6. DM Calls Do Not Work over Federation
See section 12. The call state machine is entirely local to a single server instance. No federation relay exists for call events (dm_call_start, dm_call_incoming, dm_call_accept, dm_call_reject, dm_call_end).