Commit Graph
100 Commits
Author SHA1 Message Date
Jannis Braun 2d8ecd0131 feat(federation): add approval_token columns to peer schema
Nullable text columns on federation_peers and peer_approval_requests.
Existing rows degrade gracefully (NULL token) per spec §5; the verification
logic landing in subsequent commits routes legacy null-token state through
the existing autoAccept gate.
2026-04-26 11:44:06 +02:00
Jannis Braun 4533e36e51 Merge branch 'fix/peer-handshake-trust-bypass' 2026-04-26 00:16:10 +02:00
Jannis Braun 11c5a2bf06 fix(federation): refuse outbound handshake when inbound approval pending
Closes the auto-reconnect trust-bypass: any code path calling
ensurePeered(remote) on an instance with autoAcceptPeering=0 could
previously bypass the admin gate by initiating a fresh handshake to the
remote, which the remote then accepted against its existing
awaiting_approval row.

The trigger surfaced was stores/instanceStore.ts:1010 — the silent
.catch(() => {}) auto-reconnect that fires for any user with the
remote in their replicatedInstances (commonly: any admin). Anyone with
that profile reloading their session activated peering on both sides
without any admin approval action.

Surgical fix: ensurePeered now returns rejected when an unresolved
inbound peer_approval_requests row exists for the target origin. The
legitimate admin-approve flow (routes/federation.ts:1089) does not call
ensurePeered; it deletes the approval-request and does its own direct
fetch to /peer/accept, so this check does not block legitimate approvals.

The receiver-side trust assumption at routes/federation.ts:619-645
(awaiting_approval branch in /peer/accept) still has the same flaw
— an adversarial peer that knows the timing could re-handshake at the
right moment to flip the receiver to active. That deeper trust-model
rework is plan-grade work tracked at internal notes
2026-04-26-peer-handshake-trust-model.md.
2026-04-26 00:12:42 +02:00
Jannis Braun d5c9cd297e Merge branch 'feat/s2s-friend-add' 2026-04-26 00:01:26 +02:00
Jannis Braun e15661e1bc docs(systems): S2S friend-add — social/client-federation/federation/api
Reflects what shipped on feat/s2s-friend-add (T1-T22 verified live):
- social.md: rewrite §6 outbound friend_request_create flow (sender's
  home is now the queueing instance for native users); §8 sendFriendRequest
  collapsed to single home-API call; §12 drops ConnectInstanceModal
  trigger; new "Failure Handling" subsection covers rollback path;
  relayMessageId schema note added.
- client-federation.md: split paragraph clarifying friend/DM = S2S,
  spaces = client-federated; §1 clarifies federated accounts are now
  spaces-only; new API-client error contract subsection (err.message
  carries the code, not err.body — caught + fixed in fe969a7).
- federation.md: endpoints table + S2S User Lookup subsection;
  TERMINAL_REJECTION_REASONS + permanent-failure callback registry;
  ghost-row note (rollback errors are best-effort).
- api.md: POST /api/social/requests new error-code table; new
  federation lookup route entry.
2026-04-25 23:33:13 +02:00
Jannis Braun fe969a7d96 fix(web): friend-add error toasts surface raw codes — read err.message, not err.body
The T19/T20 catch blocks looked for an `.body` property on thrown errors
to extract the structured error code. The shared API client (api/client.ts:298)
actually throws `new Error(body.error)` — the code lives in `err.message`,
and there's no `.body` attached.

Live E2E (T22 scenario 2) caught this: typing alice@orbit against an
awaiting_approval peer surfaced the raw code 'peer_pending_approval' as
the toast text instead of the human-readable mapServerErrorToMessage
output. Same defect would have hit every server-error toast on both
FriendsPage (AddFriend + UserDiscoverCard) and UserProfileModal.

Catch blocks now use err.message as both the code and the fallback text;
the inline comment points at the API client throw site so the contract
is documented at the consumer.
2026-04-25 23:17:19 +02:00
Jannis Braun 4a939b743f refactor(web): UserProfileModal — toast on server errors, drop ConnectInstanceModal triggers
Same pattern as FriendsPage cleanup (T19): the friend-add catch no longer
branches on the deleted InstanceNotConnectedError / InstanceDisconnectedError;
the ConnectInstanceModal trigger is removed since the server handles
all routing/peering. Errors surface via toast using mapServerErrorToMessage.

This restores the web package to a compileable state.
2026-04-25 22:35:52 +02:00
Jannis Braun 7309f44de5 refactor(web): FriendsPage — toast on server errors, drop ConnectInstanceModal triggers
Removes try/catch on the deleted InstanceNotConnectedError/Disconnected
classes (T17). Server now returns structured error codes; client maps
them to human-readable toasts via the new mapServerErrorToMessage helper.

The friend-add flow no longer triggers ConnectInstanceModal — the server
handles all routing/peering/lookup. The modal itself stays for Connections
settings and space-join flows.
2026-04-25 22:33:26 +02:00
Jannis Braun 9d3f75b33c feat(web): WS handlers for friend_request_sent + friend_request_relay_failed
Multi-tab sync: friend_request_sent appends the new outbound request to
socialStore (deduped by id+origin), no toast.

Async rollback: friend_request_relay_failed removes the row by id and
surfaces a warning toast with the target handle and reason. Wires the
client side of the rollback hook from T10.
2026-04-25 22:29:45 +02:00
Jannis Braun b28bf6646d refactor(socialStore): collapse sendFriendRequest; delete federation error classes
Server now handles all parsing/routing/peering/lookup (T11-T14). Client
sends the trimmed username verbatim to /api/social/requests and surfaces
server errors via toast (added in T18-T20).

Note: FriendsPage.tsx and UserProfileModal.tsx will fail to compile
until T19 and T20 remove their now-dead try/catch blocks for the
deleted error classes. TypeScript catches it; pnpm dev will not start
until those tasks land.
2026-04-25 22:26:58 +02:00
Jannis Braun 0677c21ab9 test(federation): buildFriendContextId determinism (cross-instance invariant)
The friend contextId must be byte-identical on both peers for a given
canonical pair regardless of argument order. Initial-sync backfill keys
on contextId; if the two sides computed different values for the same
friendship, missed events would never reconcile.
2026-04-25 22:24:04 +02:00
Jannis Braun bd70950de7 test(federation): receiver accepts home-queued friend_request_create 2026-04-25 22:23:04 +02:00
Jannis Braun d88d369a82 fix(social): friend_request_sent local broadcast must carry target profile
The local branch was reusing friendRequestPayload (built for
friend_request_received with user=sender) for the sender-side
friend_request_sent broadcast. Tab B on the sender would render
the sender's own avatar where the target's should appear.
Match the federated branch — sent broadcast carries target.
2026-04-25 22:18:47 +02:00
Jannis Braun 1ad1fd1b0c feat(social): broadcast friend_request_sent on local request creation 2026-04-25 22:15:57 +02:00
Jannis Braun bf13e2a223 feat(social): federated branch — authority + self-friend + idempotency
Adds direction-aware idempotency and already-friends checks to
handleFederatedFriendRequest (after stub hydration, before transaction),
plus 6 tests covering authority defense, bare-host normalization, cannot_friend_self,
already_friends, same-direction idempotent 200, and opposite-direction 409.
2026-04-25 22:12:08 +02:00
Jannis Braun 72f4b170f6 test(social): federated branch peer-status + lookup-failure mappings 2026-04-25 22:09:42 +02:00
Jannis Braun 0574a10ff3 feat(social): federated branch for POST /api/social/requests (happy path)
Refactors the POST handler into handleLocalFriendRequest + handleFederatedFriendRequest helpers. The federated branch resolves the target domain, ensures peering, looks up the remote user, creates/hydrates a replicated stub, and writes a transactional (friend_requests + mutation_log + outbox) event with relayMessageId set. Also exports hydrateReplicatedUserProfile from federation.ts and adds the T11 happy-path test.
2026-04-25 22:04:36 +02:00
Jannis Braun 08872f1e08 feat(federation): rollbackFriendRequestCreate + side-effect registration 2026-04-25 21:51:04 +02:00
Jannis Braun 9e3417485c feat(federation-worker): treat receiver-ack 4xx reasons as terminal + invoke rollback 2026-04-25 21:46:28 +02:00
Jannis Braun be3eb5284e feat(federation): permanent-failure callback registry 2026-04-25 21:42:57 +02:00
Jannis Braun 03737c0955 feat(federation): resolveOriginFromHostname helper 2026-04-25 21:41:25 +02:00
Jannis Braun fbe2fb370b feat(federation): lookupRemoteUser outbound HMAC helper 2026-04-25 21:37:41 +02:00
Jannis Braun 03cef0e6b3 feat(federation): POST /api/federation/users/lookup endpoint 2026-04-25 21:32:49 +02:00
Jannis Braun 069a1525ea feat(federation): per-peer rate limiter for user lookup (60/min) 2026-04-25 21:28:44 +02:00
Jannis Braun bd9657c5d2 fix(db): commit drizzle meta journal/snapshot for 0001 migration
The T3 commit (e4086fe) generated the SQL migration but missed the
drizzle meta files that track migration state. Without these, future
db:generate runs would re-emit or skew migration ordering.
2026-04-25 21:22:56 +02:00
Jannis Braun e4086fecff feat(db): add relay_message_id column to friend_requests
Nullable indexed column tracking the entityId of the originating relay
event for federated friend requests. Used by the rollback hook in
federationRollback.ts to locate and delete rows when a federated
friend_request_create is permanently rejected (spec §5).
2026-04-25 21:22:25 +02:00
Jannis Braun 1a9dc547c4 feat(shared): types for federation user lookup + WS events
Add FederationUserLookupRequest, FederationUserLookupProfile, and
FederationUserLookupResponse for the peer lookup endpoint contract.
Add friend_request_sent and friend_request_relay_failed ServerEvent
variants alongside existing friend_request_* cases.
2026-04-25 21:19:40 +02:00
Jannis Braun 1b98b052a1 chore(federation): T1 review polish — top-of-file import + undefined test case 2026-04-25 21:18:11 +02:00
Jannis Braun 1b63ca538e feat(federation): normalizeOriginForCompare helper 2026-04-25 21:14:58 +02:00
Jannis Braun 39a542a1ec docs(systems): update social.md for search filter + Direct-Add changes
Sections 3 and 12 now describe the discover-equivalent filter set
applied to /api/social/search and the widened Direct-Add row
contract (always-visible for well-formed input, resolved-form
display, server-side username normalization).
2026-04-25 19:35:37 +02:00
Jannis Braun 92a8a6a727 chore(friends): tighten Direct-Add gate consistency
Per code-review: directAddDisplay now reads directAt === -1
instead of re-deriving includes('@'); add a one-line comment on
showDirectAdd so the predicate's intent is obvious at first read.
2026-04-25 19:34:29 +02:00
Jannis Braun da98d78489 feat(friends): always-visible Direct-Add row with resolved-form display
The Send-Friend-Request action row in the Add Friend tab previously
appeared only when the typed query contained a non-edge @, leaving
no way to fire a blind request for a bare local handle. Widen the
gate to allow non-empty bare handles, keep the malformed @ shapes
(@, @bob, bob@) hidden. When the typed query has no @, display the
resolved form <query>@<window.location.host> so the user sees which
instance the request will hit. Submission string is unchanged.

Updates FriendsPage.test.tsx: inverts the now-stale 'does not show
Direct Add row for plain usernames' test into the new positive
assertion, and adds a separate test for the malformed @ shapes.
2026-04-25 19:32:03 +02:00
Jannis Braun 781e293cf7 chore(social-client): drop redundant comment, harden defineProperty
Per code-review:
- Drop the inline comment in sendFriendRequest; the commit message
  for the prior commit already covers the why and CLAUDE.md prefers
  no comments when the code is self-explanatory.
- Add writable: true to the window.location defineProperty in the
  test so re-firing beforeEach across jsdom version drift is safe.
2026-04-25 19:29:06 +02:00
Jannis Braun fba1f0b87d fix(social-client): lowercase parsed @domain in sendFriendRequest
Hostnames are case-insensitive (RFC 4343), and both right-hand
sides of the routing comparisons (window.location.host and
URL.host) are already canonical lowercase. The user-typed domain
substring was compared with strict ===, so ORBIT.ddns.net
failed to match an existing connected peer and popped a spurious
Connect Instance modal. Normalize at parse time.
2026-04-25 19:20:07 +02:00
Jannis Braun b5b48e407e test(social): note the shared-state harness contract
Per code-review suggestion: add a brief comment explaining why
module-level sqlite/testDb/app reassignment works (mock getter
closes over the current binding) and what would break it
(top-level it, .concurrent describe). Prevents a future foot-gun.
2026-04-25 19:17:50 +02:00
Jannis Braun 5f5590a3c9 fix(social): normalize username lookup on POST /api/social/requests
Registration canonicalizes usernames to lowercase (auth.ts:32),
but the friend-request endpoint compared with strict eq() against
raw user input, so 'Bob' returned 404 even when 'bob' existed.
Trim and lowercase before lookup, matching the rest of the auth
boundary. Empty-after-trim now returns 400 (was: 404).

Tests appended to social.test.ts as a second describe block
sharing the harness from Task 1.
2026-04-25 19:14:31 +02:00
Jannis Braun e99e0e7862 fix(social): tighten comment on federated-stub filter
Drop the file-path reference per code-review suggestion (paths in
comments rot); keep the load-bearing why — substring of the domain
matches every stub because stubs are stored as <homeUserId>@<domain>.
2026-04-25 19:12:40 +02:00
Jannis Braun b37d3bfa29 fix(social): apply discover-equivalent filters to /api/social/search
Tombstoned users, replicated federated stubs, and users with
discoverable=0 were all surfacing in Add Friend search results.
Add the three WHERE filters that /api/social/discover already
applies. Federated users continue to be surfaced via the
client-side cross-instance fan-out in socialStore.searchUsers.

New tests: social.test.ts covers all five filter cases plus
existing self-exclusion and displayName-match behaviours.
2026-04-25 18:42:05 +02:00
Jannis Braun 9aa40c0304 docs: refresh stale embed renderer descriptions after Task 1
Final-review reviewer flagged two minor staleness items:
- embeds.md §10 ImageEmbed bullets still described the pre-Task-1
  shape (no wrapper, no aspect-ratio). Replaced with the actual
  current shape, with an explicit pointer to the Dimension
  reservation contract section that explains why the dims-null
  branch deliberately has no fallback.
- message-list.md said VideoEmbed uses "padding-bottom" without
  noting the direct-video branch uses aspectRatio. Now describes
  both branches explicitly.

No code changes; both are documentation-only touch-ups.
2026-04-25 13:15:14 +02:00
Jannis Braun 56830972f4 docs(embeds): document bidirectional dimension reservation contract 2026-04-25 13:12:10 +02:00
Jannis Braun d659637930 docs(message-list): note smooth-scroll exclusion; point sentinel comment at subsystem doc
Reviewer caught two small gaps after Task 3:
- Effect A's smooth-scroll path on new messages is intentionally NOT
  instrumented with the sentinel (the animation lands asynchronously
  across frames; no intermediate scrollTop is worth pinning to). The
  doc now records this so the reader's intuition matches the code.
- The sentinel-branch comment in MessageList.tsx pointed at "spec §2",
  which is the planning doc rather than the durable subsystem spec.
  Pointed at docs/systems/message-list.md instead.
2026-04-25 13:11:05 +02:00
Jannis Braun 4b040114ed docs: add docs/systems/message-list.md subsystem spec 2026-04-25 13:06:44 +02:00
Jannis Braun d18845acb5 fix(web): close handleScroll race that disabled auto-bottom on cold-cache image load
Tracks the post-clamp scrollTop of every programmatic scroll-to-bottom in
lastProgrammaticBottomScrollRef. handleScroll skips the at-bottom flip and
re-pins when the event's scrollTop matches the sentinel — i.e., the event
was queued by our own command and layout grew underneath. User scrolls
break the match (scrollTop changes) and flow through the normal path.

This complements the 2026-03-25 race-fix (which gated auxiliary effects on
isAtBottomRef) by also preventing handleScroll from flipping that ref to
false based on a post-growth distance measurement of our own scroll.
2026-04-25 13:02:01 +02:00
Jannis Braun 84ec8e04b7 fix(web): restore known-dimension reservation in ImageEmbed without letterbox fallback
Restores the dimension reservation reverted in dae6f2d, scoped to
the embed.width && embed.height case only. No fallback aspect-ratio
when dims are null — that was the source of the dark letterbox bars
on Tenor/Klipy GIFs and OG-less images that triggered the revert.

The server-side probe in embedResolver.ts:196 already populates dims
for all image-type embeds; this change makes the client honor them.
2026-04-25 12:53:11 +02:00
Jannis Braun 7c6065c52b Merge branch 'feat/prelaunch-cleanup-bundle' 2026-04-25 11:22:58 +02:00
Jannis Braun f0d6bf2ed9 fix(federation): persist remote instance_name from approval-requests/:id/approve handshake response
Third initiator path that calls remote /peer/accept. Mirrors performHandshake (auto-peer) and /peer/initiate (admin-initiate) — same try/catch parse, same null-or-non-empty-string guard. Caught in final review of #33; same root cause as Bug #1, bundled rather than fragmented to a new backlog item.
2026-04-25 11:18:12 +02:00
Jannis Braun 4aced5654b docs(federation): document bidirectional instanceName exchange during peer handshake 2026-04-25 01:01:04 +02:00
Jannis Braun a9bf5aeb53 fix(dm): include federatedId in POST /api/dm idempotent existing-DM response
Fresh-create returned {id, ownerId, federatedId, createdAt, members, lastMessage}; the existing-DM path returned the same shape minus federatedId. Inconsistency was a footgun for any future feature reading federatedId from this response — fresh-create tests would pass while idempotent path would break. One-line addition to the result builder.
2026-04-25 00:59:16 +02:00
Jannis Braun 618056659e fix(federation): persist remote instance_name from /peer/initiate handshake response
Mirrors the previous performHandshake fix for the admin-initiated path. /peer/initiate now parses the remote's instanceName from the /peer/accept response body and writes it alongside status='active'.
2026-04-25 00:48:42 +02:00
Jannis Braun 18d6b0acfa fix(federation): persist remote instance_name when ensurePeered/performHandshake succeeds
Initiator side of the bidirectional handshake exchange. /peer/accept now returns instanceName in the response body (prior commit); performHandshake parses it and persists alongside status='active'. Tolerates missing field (older peers) and non-JSON bodies.
2026-04-25 00:45:06 +02:00
Jannis Braun eba16e16a3 fix(federation): include own instanceName in /peer/accept response body
Bidirectional handshake exchange. Today the responder learns the initiator's instance name from request body but the initiator never learns the responder's. Adding {instanceName} to the response body lets the initiator persist it on its side (next commit). Field is optional so older peers omitting it cause no ill effect.
2026-04-25 00:42:11 +02:00
Jannis Braun 7fe9476e57 fix(federation): persist peer instance_name on /peer/accept activation paths
Previously /peer/accept read body.instanceName only when queueing for admin approval. The four paths that mutate federation_peers (rejected→active override, awaiting_approval→active, pending→active, new-peer create) all wrote status='active' without persisting instance_name. Result: every peer established via direct handshake had instance_name = NULL forever. Anywhere peerLabel was rendered fell back to origin hostname.

Active/needs_attention idempotent early-return path deliberately left alone — same security posture that already refuses to overwrite hmac_secret from unauthenticated requests on already-active peers.

Existing live NULL rows are repaired post-deploy via manual UPDATE statements (see plan).
2026-04-25 00:35:42 +02:00
Jannis Braun 940a0f7ba9 Merge branch 'feat/migration-squash-phase-2'
Squashes drizzle migrations 0000-0004 into a single baseline
(0000_lethal_wildside.sql) and deletes the three migration adapter
functions that existed only because of pre-squash intermediate states:
baselineExistingInstall, healInitialSchemaDrift, healRenamedColumns.

__drizzle_migrations surgery completed and verified live on Pi and VM
before this merge. Both instances boot cleanly against the single
baseline; cross-instance DM delivery verified.

Closes backlog #31 Phase 2.
2026-04-24 23:57:41 +02:00
Jannis Braun 798aa22690 docs(systems): rewrite database.md migration workflow section after migration squash
Line 4 pointed at a stale `runMigrations()` name and omitted the drizzle-kit
generate step entirely. Replace with a description of the real workflow:
`pnpm db:generate` produces SQL from `schema.ts`, `initDatabase()` runs
`drizzle.migrate()` + `ensureDefaults()` on startup. Note the 2026-04-24
baseline squash for future readers.

Delete the "Migration flags (internal)" line — those flags belonged to
pre-drizzle data-fix migrations deleted wholesale in 3acaea2 (2026-04-09)
and are historical trivia with no present referent.

Refs backlog #31 Phase 2.
2026-04-24 23:38:47 +02:00
Jannis Braun ab5e7c8839 refactor(server): delete baseline + heal functions now unreachable under single-baseline history
baselineExistingInstall, healInitialSchemaDrift, and healRenamedColumns
existed only because of pre-squash intermediate states: dev DBs drifted
against drizzle-kit's assumed 0000 baseline, or carried pre-rename
columns from the old manual migration system. Under a single squashed
baseline there are no intermediate states to drift against, and these
functions are unreachable.

initDatabase() is now: pragmas -> drizzle() -> migrate() -> ensureDefaults().
ensureDefaults stays (idempotent startup seeding for settings row, worker
ID, first-admin promotion).

Refs backlog #31 Phase 2.
2026-04-24 23:33:35 +02:00
Jannis Braun 707d3217f1 refactor(server): squash drizzle migrations 0000-0004 into single baseline
Generated by pnpm db:generate against the current schema.ts — replaces
five historical migrations (0000_initial, 0001_clear_earthquake,
0002_peer_approval_requests, 0003_classy_loki, 0004_cooing_black_knight)
with one baseline that matches the schema shape all five produced together.

Pi + VM __drizzle_migrations rows will be rewritten per the Phase 2
surgery procedure before this deploys to either instance; DBs already
match the new baseline, no DDL runs. Phase 1 audit verified no still-present
bugs depend on any of the squashed migrations.

Refs backlog #31 Phase 2.
2026-04-24 23:26:20 +02:00
Jannis Braun a181b1ad10 Merge branch 'feat/remote-200-no-recipient' 2026-04-24 22:35:39 +02:00
Jannis Braun f6de72d556 chore(verify): #18 live Pi↔VM scenarios A-D pass
Harness: /tmp/scenario18-harness.mjs (WS-level assertion, pattern
matches #17's /tmp/call-test-harness.mjs).

- A (Pi→VM logged-out callee, Path A zero-ringee):
    terminal dm_call_undeliverable reason=no_recipient peerOrigin=VM
    elapsed 212ms (budget 2s, pre-fix 60s)
- B (VM→Pi logged-out callee, symmetric):
    elapsed 155ms, peerOrigin=Pi
- C (Path B, DM deleted on VM):
    relay hits Path B after DB delete, Bob offline,
    elapsed 140ms, reason=no_recipient
- D (group DM with online member, non-regression):
    VM accepted (Bob rung), no toast on caller
    Bob's dm_call_incoming arrived in 144ms

All payloads correct: terminal:true, phase:'start', failures[0].reason
'no_recipient', correct peerOrigin. peerLabel empty because both
instances' federation_peers.instance_name is NULL — pre-existing state,
toast code already falls back to origin hostname (not a #18 concern).

Cannot merge from agent per plan Task 10 Step 7 — coordinator's call.
2026-04-24 21:41:34 +02:00
Jannis Braun 2dbd2b9b9f test(server): add undeliverable:[] to processRelayEvents mock (#18)
Follow-up from Task 2 code review. Keeps the test mock aligned with
the widened return type even though vi.mock doesn't structurally
typecheck the factory.
2026-04-24 21:24:18 +02:00
Jannis Braun 34622a4290 docs(systems): #18 review followups — bullet formatting + host_unreachable phase
dm-system.md: fold the voice.md cross-reference into the paragraph so it
renders as part of the Cross-instance access explanation instead of an
orphaned bullet.

websocket.md: add 'host_unreachable' to the dm_call_undeliverable phase
union — stale since #32 was merged (docs drift noted in Task 8 review).
2026-04-24 21:22:21 +02:00
Jannis Braun 6edf02cb33 docs(systems): document three-way ack classification + no_recipient (#18)
federation.md — new undeliverable bucket subsection with three-way
classification table, Path A/B semantics, and wire backward-compat note.
voice.md — no_recipient row in failure-surface table.
websocket.md — DmCallUndeliverableReason union updated to include no_recipient.
dm-system.md — cross-reference to voice.md for no_recipient reason.
2026-04-24 21:19:34 +02:00
Jannis Braun b16ece93b8 feat(web): no_recipient toast copy arm (#18)
buildCallUndeliverableToast renders "{peerLabel} couldn't ring anyone."
for the single-failure terminal case; multi-failure + non-terminal paths
fall through to existing lines (which already fold the new reason in by
peer label). TDD — four new assertions.
2026-04-24 21:15:21 +02:00
Jannis Braun 26a4925032 feat(server): reclassify undeliverable targeted-peer as no_recipient failure (#18)
sendFederatedCallStart now treats a 200-with-undeliverable-messageId as
a peer-failure instead of unconditional success. Feeds the existing
failures[] array and terminal-determination machinery from #16.
New sendFederatedCallStartForTest export mirrors the existing
handleDm*ForTest pattern. TDD — three tests cover single-peer terminal
no_recipient, group-DM mixed delivered+undeliverable non-terminal, and
the happy-path (empty undeliverable → no event).

Also hardens sendCallRelay's response parse: validates undeliverable
is an Array and entries are well-shaped, logs protocol drift at warn/debug
rather than silently falling back to old-peer semantics.
2026-04-24 21:11:11 +02:00
Jannis Braun 7d2137b6d4 feat(server): sendCallRelay surfaces undeliverable messageIds (#18)
CallRelayResult success arm gains undeliverable: string[]. sendCallRelay
parses FederationRelayResponse.undeliverable (when present) and returns
the messageIds so sendFederatedCallStart can reclassify per-peer results.
Old peers that omit the field → empty array → today's behavior.
TDD — three tests cover old-peer, new-peer-with-undeliverable, and 5xx paths.
2026-04-24 21:05:28 +02:00
Jannis Braun 7533d8ca14 feat(server): Path A connection gate + undeliverable on zero ringee (#18)
processDmCallStartEvent Path A now skips offline local members (matching
Path B's pre-existing per-member check) and pushes undeliverable when no
member could be rung, instead of creating a stranded FederatedCallEntry.
TDD — two new tests cover zero-online and mixed-online cases.
2026-04-24 21:01:19 +02:00
Jannis Braun 792634c2cf feat(server): Path B zero-match → undeliverable ack (#18)
processDmCallStartEvent Path B no longer silently accepts when no local
participant is reachable. Pushes {messageId, reason: 'no_recipient'} to
the undeliverable ack bucket so the caller can surface fast-fail.
TDD — test asserts undeliverable push + no FederatedCallEntry.
2026-04-24 19:34:34 +02:00
Jannis Braun 0057cb4d42 refactor(server): thread undeliverable collector through processRelayEvents (#18)
Additive plumbing. No behavior change — every existing event-type path
continues to push to accepted/rejected only. Response serializes the new
bucket only when non-empty (byte-identical responses in the normal case).
Tasks 3-4 add actual undeliverable pushes for dm_call_start paths.
2026-04-24 19:17:27 +02:00
Jannis Braun 3988c5823a feat(shared): add no_recipient reason + undeliverable response field (#18)
Additive protocol extension. No consumers yet — follow-up commits wire
the new bucket into the relay endpoint, sendCallRelay, sendFederatedCallStart,
and the toast copy.
2026-04-24 19:12:54 +02:00
Jannis Braun a68eab3edc Merge branch 'feat/remote-participant-host-unreachable' 2026-04-24 01:09:35 +02:00
Jannis Braun 5c94bc7669 refactor(server): remove unused admin_reset PeerDeactivationReason
The admin reset endpoint (DELETE-pattern gated on peer.status !== 'needs_attention')
doesn't transition status — it deletes the row of an already-deactivated peer.
onPeerDeactivated already fired at the earlier needs_attention transition, so
the reset site correctly has no hook. The enum value was defensive-unused; per
project principles (no backwards-compat shims, no placeholders) drop it.
2026-04-24 01:07:47 +02:00
Jannis Braun 7c0dd33123 docs(systems): document host_unreachable phase + onPeerDeactivated + sentinel 2026-04-24 00:54:46 +02:00
Jannis Braun 5e509cf3df feat(web): host_unreachable phase copy for dm_call_undeliverable (TDD) 2026-04-24 00:51:30 +02:00
Jannis Braun 314df6c5a1 feat(server): 30s federated-call sentinel worker (TDD) 2026-04-24 00:49:04 +02:00
Jannis Braun 71445c5f27 feat(server): wire onPeerDeactivated at admin revoke/reset sites
Admin-revoke endpoint (DELETE /api/federation/peers/:id): fires
onPeerDeactivated(id, 'admin_revoked') after the status write to
'revoked', evicting any in-flight federated calls for the now-revoked
peer.

Admin-reset endpoint (POST /api/federation/peers/:id/reset): hook
SKIPPED. The reset endpoint is guarded to only run when status is
already 'needs_attention' (active peers are rejected at the boundary
with a 400). Because the peer was already deactivated before reset is
called, onPeerDeactivated was already fired at the active→needs_attention
transition. The reset deletes the row entirely rather than writing a new
status; it does not represent a transition OUT OF active, so wiring it
here would be a semantic error — double-evicting an already-deactivated
peer.
2026-04-24 00:45:02 +02:00
Jannis Braun 8e6639648e feat(server): wire onPeerDeactivated on performHandshake 403 rejection 2026-04-24 00:43:32 +02:00
Jannis Braun 743fdcac97 feat(server): wire onPeerDeactivated at federationWorker peer-deactivation sites 2026-04-24 00:42:20 +02:00
Jannis Braun 0a8949fbfb feat(server): onPeerDeactivated utility mirrors onPeerActivated (TDD) 2026-04-24 00:40:42 +02:00
Jannis Braun 3b61380a1e feat(server): ConnectionManager.evictFederatedCallsForHost (TDD) 2026-04-24 00:37:44 +02:00
Jannis Braun 942619f422 feat(shared): add 'host_unreachable' phase to DmCallPhase 2026-04-24 00:34:05 +02:00
Jannis Braun 726341f8f8 Merge branch 'feat/call-state-machine-hardening' 2026-04-23 23:44:49 +02:00
Jannis Braun 1e59e7012c fix(server): scope accept-rollback terminal to the acceptor only
Code-review catch: the Path-2 accept-rollback previously emitted
dm_call_undeliverable { terminal: true } via sendToFederatedCallUsers,
which broadcasts to every ringedUserIds entry. In a group DM this
would prematurely tear down non-accepting ringees whose own accept /
reject / timeout paths should govern their state. Switch to
sendToUser(acceptorId) so only the acting user gets the terminal
signal. Reorder the clearFederatedCall to happen before the emit so a
concurrent end-handler sees a cleared entry (clearFederatedCall is
idempotent). Spec updated, test extended to assert the scoping with a
two-ringee group-DM fixture.
2026-04-23 23:38:31 +02:00
Jannis Braun 1719e6d580 docs(systems): document federation call state machine hardening 2026-04-23 23:23:06 +02:00
Jannis Braun 6a5b02b1a0 feat(web): phase-aware dm_call_undeliverable toast copy (TDD) 2026-04-23 23:21:40 +02:00
Jannis Braun f6252b8ce1 feat(server): fan dm_call_end out on host ring timeout 2026-04-23 23:19:42 +02:00
Jannis Braun 6cd7728f3e feat(server): fanOutCallEvent returns failures; surface to host-side caller 2026-04-23 23:14:50 +02:00
Jannis Braun 3cb80d110d feat(server): aggregate Path-1 call fan-out failures and surface to originator 2026-04-23 23:12:55 +02:00
Jannis Braun 13241345de feat(server): surface dm_call_end relay failure to originator (TDD) 2026-04-23 23:11:20 +02:00
Jannis Braun 5170d316ba feat(server): surface dm_call_reject relay failure to rejector (TDD) 2026-04-23 23:10:18 +02:00
Jannis Braun 07c5b0e7de feat(server): surface dm_call_accept relay failure to acceptor (TDD) 2026-04-23 23:09:18 +02:00
Jannis Braun 06e1ed92e7 refactor(server): add buildFailureFromResult + CallFanoutFailure types 2026-04-23 23:07:07 +02:00
Jannis Braun 5e9b353124 feat(shared): add DmCallPhase to dm_call_undeliverable event 2026-04-23 23:06:16 +02:00
Jannis Braun fe01b94876 Merge branch 'fix/dev-db-rename-drift'
Close backlog #30: second-pass heal that DROPs and rebuilds empty
drifted tables to reconcile pre-rename column drift on old pre-drizzle
dev DBs (companion to #29). `healInitialSchemaDrift` handles
addable-column drift; this handles the NOT-NULL-no-default case that
heal-by-ALTER can't resolve — e.g. federation_outbox's
`message_id`/`dm_channel_id` that were renamed to
`entity_id`/`context_id` in the pre-drizzle manual migrate.ts.

Rebuild target is the current-migration-state snapshot (derived from
__drizzle_migrations ↔ _journal.json by `when` timestamp), not the
latest on disk — rebuilding forward would duplicate-column with
pending ALTER TABLE ADD COLUMN statements drizzle is about to run.

Empty-table gate preserves data. Missing-column gate (not extra-only)
avoids touching unused leftover columns from pre-drizzle manual
migrations that no current code reads.

Verified against three simulated scenarios (fresh / pre-drizzle
drift / production-like) plus live boot on the actual affected dev
DB: server binds :3005 cleanly, no outbox worker SQLITE_ERROR ticks,
migrations complete silently. Server tests 110/110, web 131/131.

No migration files changed. Deployed instances unaffected. Skip
redeploy.
2026-04-23 03:20:08 +02:00
Jannis Braun 06c538b013 fix(server): rebuild empty drifted tables to heal pre-rename column drift
Companion to #29. `healInitialSchemaDrift` can only ADD columns, so it
skips NOT NULL-without-default columns like `federation_outbox.entity_id`
/ `context_id` — which on some old pre-drizzle dev DBs carry the
pre-rename names `message_id` / `dm_channel_id` instead. The tables
load but the outbox worker fails every tick with "no such column:
federation_outbox.context_id" once the server is up.

Adds healRenamedColumns() — a second pass that runs right after
`healInitialSchemaDrift`. For each table whose physical column set is
*missing* columns declared by the current-migration-state snapshot AND
which holds zero rows, it DROPs the table and rebuilds it from the
snapshot's JSON: columns, defaults, foreign keys, composite PKs,
unique constraints, indexes.

Key design decisions:

- **Target is the current-migration-state snapshot, not the latest on
  disk.** The current state is determined by the highest
  `__drizzle_migrations.created_at` matched against `_journal.json`'s
  `when` timestamps (with backward walk for idx values that lack a
  snapshot, like the hand-written 0002). Rebuilding to a *future*
  snapshot would introduce columns that drizzle's migrator is about
  to add via ALTER TABLE ADD COLUMN, causing duplicate-column errors.
  Rebuilding to the *current* snapshot preserves the invariant that
  drizzle's pending migrations can run cleanly afterwards.

- **Missing-column gate, not extra-column.** Extra columns alone don't
  break anything at runtime (the ORM ignores them); they're leftover
  from pre-drizzle manual migrations and might matter to the operator.
  Missing columns DO break runtime queries, so only those trigger
  rebuild.

- **Empty-table gate.** Non-empty tables log a warning and skip —
  data preservation wins over heal, and this path should only ever
  hit a pre-drizzle dev DB that never exercised the affected tables
  in the first place.

- **Transactional rebuild.** DROP + CREATE + index reinstatement wrap
  in a single `db.transaction()` so a partial rebuild rolls back.

Verified against three scenarios via in-memory simulation:
(A) fresh install — heal no-op, drizzle creates everything; (B) pre-
drizzle dev DB with fed_outbox/fed_mutation_log rename drift —
tables rebuilt to 0000 snapshot, drizzle then applies 0001–0004
successfully to reach the current target schema; (C) post-migration-
correct (production-like) — heal no-op, drizzle no-op, schema
unchanged. Live boot on my actual dev DB: migrations complete
silently, server binds :3005, no outbox worker errors. Server tests
110/110, web 131/131, typecheck clean.

No migration files changed. Deployed Pi+VM instances are unaffected
(their schema matches the snapshot exactly — heal won't touch
anything).

Closes backlog #30.
2026-04-23 03:19:56 +02:00
Jannis Braun 02a8b27d2b Merge branch 'fix/joinspace-tests-and-dev-db-drift'
Two small items from the S2S DM unification backlog.

#28 — 4 stale JoinSpace test assertions (test-only fix)
  Assertions from the old JoinServer component drifted when the file was
  renamed (fc06e25) without being updated: placeholder text, submit
  button now disables-when-empty (so the 'Invite code is required' error
  path is unreachable from the rendered form), and joinByCode signature
  gained a second `origin` argument. Updated each assertion to match
  current behavior. No code change. 131/131 web tests now pass.

#29 — Local dev DB missing federation_peers.remote_max_upload_size
  Root cause: baselineExistingInstall trusts that any pre-existing
  install's schema matches 0000_initial. That breaks for dev DBs from
  the pre-drizzle manual migrate.ts system that didn't ran every
  idempotent ALTER — 0000 gets marked applied without its columns
  actually existing, and a later migration that recreates the table
  (0004_cooing_black_knight) crashes. New healInitialSchemaDrift()
  walks 0000_snapshot.json, and for each existing table ADDs any
  declared-but-missing columns before drizzle's migrate() runs.
  Idempotent; production instances are no-op.

Remaining dev-DB drift flagged for a follow-up: federation_outbox and
federation_mutation_log retain pre-rename column names (message_id /
dm_channel_id) that the old manual system renamed to entity_id /
context_id. Both tables are empty on affected DBs, but the clean fix
requires DROP + RECREATE with index reinstatement, beyond #29's
scope. Outbox worker emits SQLITE_ERROR ticks post-boot on unfixed
dev DBs; separate item.

No migration files changed. Deployed instances are unaffected by
either commit (#28 test-only, #29 heals only DBs missing columns —
production DBs aren't missing any). Skip redeploy.
2026-04-23 03:00:49 +02:00
Jannis Braun e703e29f8a fix(server): heal 0000-baseline schema drift after baselining existing installs
baselineExistingInstall marks 0000_initial as applied when it detects
pre-existing tables, on the assumption the install's schema matches the
0000 baseline. That assumption is false for dev DBs created under the
pre-drizzle manual migrate.ts system that skipped or never ran some of
its idempotent ALTER TABLE steps — for example the b9e4c65 migration
that added federation_peers.remote_max_upload_size. On such DBs, 0000
is marked done without the column actually existing, and a later
migration that recreates the table (0004_cooing_black_knight)
subsequently crashes with "no such column: remote_max_upload_size"
while building its __new_federation_peers SELECT.

Adds healInitialSchemaDrift(): walks every table in 0000_snapshot.json,
and for each table that already exists, ADDs any columns the snapshot
declares but the physical table is missing. Runs immediately after
baselining, before drizzle's migrate() — so later migrations find the
schema they expect. Columns that SQLite's ALTER TABLE ADD COLUMN can't
safely express (PRIMARY KEY; NOT NULL without a default) are skipped
with a warning rather than corrupting data.

Idempotent: on fresh installs and correctly-migrated DBs every column
is already present, so the loop is a no-op. Production Pi+VM instances
are unaffected.

Verification: local dev DB that previously crashed on 0004 now boots
cleanly — federation_peers gained remote_max_upload_size, nonce_supported,
pending_hmac_secret, secret_rotation_at, secret_rotated_at, and
auto_rotate_interval_days; __drizzle_migrations advanced from 4 to 5
entries; server binds :3005. 110/110 server tests + 131/131 web tests
still pass.

Not covered: a deeper drift on federation_outbox /
federation_mutation_log where the physical tables retain pre-rename
column names (message_id / dm_channel_id) instead of the current
entity_id / context_id. Heal skips those (NOT NULL without default)
and the outbox worker emits SQLITE_ERROR ticks post-boot. Both tables
are empty on affected dev DBs, but a clean fix requires DROP +
RECREATE with index reinstatement which is out of #29's stated scope
("column existing"). Flagged for a follow-up.

Closes backlog #29.
2026-04-23 03:00:30 +02:00
Jannis Braun 086158511c test(join-space): update stale assertions to match current UI and API
Four assertions drifted from the current JoinSpace modal, carried over
when the file was renamed from the old JoinServer component in fc06e25
without being updated:

- Placeholder was expanded to cover URL-form invite input
  ('e.g. abc123' → 'e.g. abc123 or https://instance.com/join/abc123').
- 'shows validation error when submitting empty code' asserted a code
  path that no longer exists: the submit button is now disabled when
  the trimmed input is empty (JoinSpace.tsx line 166), so clicking it
  is a no-op and the 'Invite code is required' error from the parser
  is unreachable from the rendered form. Replaced with an assertion
  that the button is disabled while the input is empty — the actual
  validation UX.
- joinByCode signature took on a second `origin` argument during the
  S2S DM unification + federated-join work (spaceStore.ts line 69).
  parseInviteInput returns { code, origin: undefined } for a bare
  code, so the call is `joinByCode('my-invite-code', undefined)`.
  Assertion updated to match exactly.

No code behavior change — tests now reflect actual behavior, which
was already correct and deployed. Closes backlog #28.
2026-04-23 02:52:00 +02:00
Jannis Braun 7220f4cbcd Merge branch 'fix/cross-store-resolver-tdz'
Close backlog #27: extract cross-store resolvers into a neutral utility
(packages/web/src/utils/crossStoreResolvers.ts) to break a TDZ cycle
between spaceStore and instanceStore. instanceStore's top-level
setResolver calls used to race with spaceStore's `let _getApiForOrigin`
declaration when the module graph was entered from instanceStore
(JoinSpaceModal → useInstanceStore), crashing with "Cannot access
'_getApiForOrigin' before initialization" and preventing
InviteModal.test.tsx and JoinSpace.test.tsx from loading.

Moves the three resolver lets + setters + pure getters + the
WS-populated user-ID cache into the utility; spaceStore re-exports the
public surface; instanceStore imports the setters directly from the
utility (re-exports do not resolve at module-init time under vite-ssr
in the cycle). authStore-using wrappers (resolveUserOrigin,
getLayoutHomeOrigin, getMyUserIdForOrigin) stay in spaceStore but
delegate to the utility.

Also adds the AudioManager mock to InviteModal/JoinSpace test files
(established pattern) so their suites can load.

Verification: typecheck clean, 127/131 web tests pass (up from 121/121;
+6 newly unlocked), server 90/90 unchanged. The 4 remaining JoinSpace
failures are pre-existing stale UI-text assertions (placeholder
expanded, submit button now disable-when-empty) — unrelated to this
work, made visible only because the suite loads now.

Spec: docs/systems/client-federation.md updated.
Smoke test: vite dev bundle serves spaceStore + crossStoreResolvers
clean; full live E2E blocked by a pre-existing local DB-migration
error unrelated to this client-side refactor (reproduces on main).
2026-04-23 02:46:37 +02:00
Jannis Braun 4d2e50b55b fix(web): extract cross-store resolvers into neutral utility to break TDZ
instanceStore registers three resolver functions at module load —
setApiForOriginResolver, setUserIdForOriginResolver,
setOriginFromHostnameResolver — whose backing `let` bindings used to
live in spaceStore. When the module graph was entered from
instanceStore (e.g. JoinSpaceModal importing useInstanceStore) the
order became spaceStore → chatStore → useWebSocket → socialStore →
instanceStore (top-level setter call) while spaceStore was still
paused on its line-8 chatStore import, so the backing `let` had not
been reached yet and the setter crashed with
`Cannot access '_getApiForOrigin' before initialization`. This left
InviteModal.test.tsx and JoinSpace.test.tsx unable to even load their
suites once AudioManager was mocked away.

Move the three `let` bindings, their setters, their pure getters, plus
the WS-populated user-ID cache (`_myUserIdByOrigin`, setMyUserIdForOrigin,
getCachedUserIdForOrigin, clearMyUserIdCache) into
`packages/web/src/utils/crossStoreResolvers.ts`. The utility imports
nothing from `./stores/*`, so no back-edge exists. spaceStore re-exports
the public surface for backward compatibility with the many existing
import sites; instanceStore imports the setters directly from the
utility (the in-cycle re-export path does not resolve at module-init
time under vite-ssr, so a direct import is required for the top-level
setter calls).

spaceStore's remaining wrappers (resolveUserOrigin, getLayoutHomeOrigin,
getMyUserIdForOrigin) stay where they are — they combine the utility's
pure lookups with authStore state — but now delegate to the utility.

Also adds the AudioManager mock to InviteModal.test.tsx and
JoinSpace.test.tsx so their suites actually load (same pattern already
used in 5 other test files). Net test-suite result: 127/131 pass (up
from 121/121 — +6 newly unlockable). The 4 remaining JoinSpace
failures are pre-existing stale UI-text assertions (the placeholder was
expanded and the submit button was made disable-when-empty) made
visible by the suite now loading; they're orthogonal to this change
and handed back for a separate triage.

Closes backlog #27.
2026-04-23 02:46:19 +02:00
Jannis Braun 15d2a3f163 Merge branch 'fix/pre-existing-test-failures'
Triage and fix of 12 pre-existing test failures on main (flagged by
the #10 DM origin failover merge, commit ff39ab0).

Two root causes:

1. Node 20+'s built-in localStorage/sessionStorage stub shadows jsdom's
   working implementation because vitest's populateGlobal doesn't
   overwrite globals outside its known allow-list. Any zustand persist
   store threw "storage.setItem is not a function". Polyfilled with an
   in-memory Storage in src/test/setup.ts. Resolves 11 of 12 failures
   (10 keybindStore + 1 FriendsPage toast).

2. FriendsPage "Message button" DM test carried a stale two-arg
   assertion that predated the 2026-04-01 federation refactor (commit
   7f3ca4e) which dropped the second argument from addDmChannel.
   Dropped the trailing '' so the assertion matches current behavior.

Hand-backs (not touched on this branch):
- InviteModal.test.tsx and JoinSpace.test.tsx still fail to LOAD (not
  in the 12 tests but flagged by #10's merge note). After stubbing
  AudioManager a second blocker surfaces: TDZ error on _getApiForOrigin
  in spaceStore.ts:961, caused by a circular-import init order between
  spaceStore and instanceStore (via socialStore → useWebSocket →
  voiceStore). Federation-adjacent — filed as backlog item for a
  structural fix.

Server: 90/90 unchanged.  Web: 121/121 pass (was 109/121), 2 suite
loads still failing (tracked).
2026-04-23 02:29:20 +02:00
Jannis Braun ae5bdaa338 test(friends): drop stale origin arg from addDmChannel assertion
The DM "Message button" test asserted addDmChannel was called with two
arguments — the channel and an empty-string origin — but the assertion
has been stale since commit 7f3ca4e ("route DM creation to home instance
with federated identity", 2026-04-01). That refactor made FriendsPage
always route DM creation through the home api client and dropped the
second argument from the addDmChannel call because the remote friend's
instanceOrigin no longer applies — home-created DMs don't need a
channelOriginMap entry (lookups default to '' for missing keys; remote-
delivered DMs still get their origin tagged by useWebSocket).

The two-arg assertion was introduced on 2026-03-25 (commit 277b69a)
against an intermediate form of the code that was later rewritten. Drop
the trailing '' so the assertion matches the current, intentional
one-arg call.
2026-04-23 02:28:57 +02:00