Skip to content

Fix findings from the Crew merge QA campaign - #377

Merged
Broccolito merged 365 commits into
mainfrom
claude/crew-qa-fixes-2026-09-27
Sep 29, 2026
Merged

Broccolito merged 365 commits into
mainfrom
claude/crew-qa-fixes-2026-09-27

Conversation

@Broccolito

@Broccolito Broccolito commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Fixes for everything the Crew merge QA campaign (crew-merge-qa-2026-09-27) found after PR #366 merged Crew into main: a static audit of the merged code, two live QA runs on disposable AWS fleets with synthetic data only, and three fix waves, each re-verified live. The record of the whole campaign is docs/history/biorouter-crew/evidence/merge-qa-2026-09-27.md.

  • No P0 anywhere. No confidentiality boundary broke and no acknowledged data was lost in either live run.
  • Found: 59 static items (10 P1, 49 P2); 113 live findings in run 1 plus the security lane's SEC-2..SEC-9, which triage turned into 117 new wave-2 items (8 P1, 109 P2); 85 wave-3 items (1 P1, 84 P2) from run 2.
  • Fixed: all 59 static items (57 in wave 1, CROSSCUT-6 and DOCS-3 in wave 2), all 119 wave-2 items, and all 85 wave-3 items.
  • Final live round on 341df2d06 (the last code commit): 80 pass, 3 partial, 1 not reproduced, 1 not run; no P0 or P1 regression.
  • After the final round, hosted CI found four problems the live and local runs could not: 92094da45 makes the exported app launcher list its token's hex digits, because macOS's /bin/bash 3.2 under a UTF-8 locale let [!0-9a-f] accept an upper-case token, and runs a launcher test from a file (Git Bash misparsed it on the Windows command line); 5443d717b fixes two Crew tests that failed on a runner with no keyring and writes the Claude Code and Codex fake CLIs from a child process to avoid ETXTBSY; 6b37a4f06 makes four ChatCrewAccessBar tests wait for the composer hold the bar publishes a render later (two failed once on a loaded runner). None of these was re-checked on the fleet.
  • Docs: d94624565 adds the campaign's evidence record and links it from the history indexes; 1ced37e53 brings the configure transcripts, BIOROUTER_MAX_TURNS, the non-private-model notice, the landing providers and Crew pages and three Crew manual lines in line with the final behaviour. check-crew-manual (53/53), the docs placement lint and a relative-link check of every touched page pass.

Fleets and lanes

Run 1 (merged build ead0498d6) Run 2 and final round (42e5391a, then 341df2d06)
Fleet crewqa-20260927T0650Z, torn down and verified 2026-09-27 23:09Z crewqa-20260928T0400Z, supervisor-enforced deadline 2026-09-29 12:00Z
Hosts Ubuntu 24.04 lab server, Debian 11 lab server (glibc 2.31), Ubuntu bastion, AlmaLinux 9.8 Slurm cluster (login + 2 nodes, SELinux Enforcing, NFS /home) reached by ProxyJump the same shape
People 12 no-sudo accounts, one with password/PAM login, one in-scope adversary 15 (adds three novice testers)
Lanes setup-hpc, setup-foreign, setup-chen, security, docs-walk, messaging, agents, providers, resilience, files setup-chen, setup-foreign, setup-hpc, messaging, agents, files, providers, resilience, cli-docs, ux-novice, hpc; then fix3-recheck

Every agent turn used the private versa_azure / gpt-5.5-2026-04-24 model. Run 1's 134 tasks: 81 unaided, 33 friction, 10 failed, 9 n/a, 1 blocked.

Findings fixed, by severity

P1 (19 items, all fixed):

  • Static: BROKER-1 (one member could freeze all writes), BROKER-2 (unpaged snapshot lockout), BROKER-3 (/tmp runtime-path squat), BROKER-4 (workspace-wide attachment quota), RENDERER-1 (untrusted names shown and offered as Save defaults), CLI-1 (connections remove with no confirmation), CLI-2 (single-entry vault passphrase), PROVIDERS-1 (Crew chat on Llama Server refused mid-reply), PROVIDERS-2 (Quick/Deep effort broke Crew turns), CROSSCUT-1 (per-process Crew scope).
  • Wave 2: W2-BRK-1 (wrapped invitation token), W2-BRK-3 (full disk wedged every member), W2-DMN-1 (headless node needs a vault, opaque 400), W2-DMN-15 and W2-PRV-10 (max turns 0 stopped every chat), W2-HRD-1 (loopback port let another account drive an Agent Drafter app), W2-HRD-2 (world-readable secrets.yaml), W2-SHL-1 (app could not recover after a daemon restart).
  • Wave 3: T3-SH-1 (non-private-model notice came only after the first turn).

P2 (all fixed):

  • Static (49): BROKER-5, BROKER-6; DAEMON-1..7; RENDERER-2..6; CLI-3..20; DOCS-1..7; PROVIDERS-3..6; CROSSCUT-2, 4..8.
  • Wave 2 (111): W2-BRK-2, 4..8; W2-DMN-2..14; W2-HRD-3..7; W2-CLI-1..18; W2-UIC-1..15; W2-UIW-1..20; W2-SHL-2..9; W2-PRV-1..9, 11..16; W2-DOC-1..8; W2-STR-1; plus CROSSCUT-6 and DOCS-3.
  • Wave 3 (84): T3-BE-1..18; T3-CLI-1..18; T3-UI-1..31; T3-SH-2..12; T3-DOC-1..6.

The evidence page maps every item to its commits.

Tests added

About 413 new Rust #[test] functions and 641 new Vitest/node cases, and 51 new test files, including:

  • Broker contracts: crates/biorouter-crew/tests/quota_contract.rs, storage_fault_contract.rs, presence_contract.rs, runtime_contract.rs, plus additions to journal_fault_contract.rs, name_rules_contract.rs and join_contract.rs.
  • Daemon and server: crates/biorouter/tests/crew_provider_admission.rs, crates/biorouter-server/tests/crew_context_exits.rs, and additions to the crew_* route suites.
  • CLI: output and refusal-code tests in crates/biorouter-cli/src/commands/crew/, and configure checks pinned to the desktop's mode and Max turns copy.
  • Desktop: Crew timeline, composer, sidebar, dialogs, access pane, transfers and daemon-restart tests; provider settings and Max turns tests.
  • Docs: scripts/check-crew-manual.mjs rules (each shown failing on the old wording) and the docs/crew dash guard.

Final live verdicts

  • Run 2 (stage 2 on 42e5391a): every lane re-verified its wave-1 and wave-2 items live. Headline tallies: agents 35 pass, 2 partial, 4 not testable live; cli-docs 42, 6, 2; messaging 27, 4, 1; providers 19, 4, 1; files 14, 2, 1; hpc 8, 2, 1; ux-novice 22, 3, 1; setup-chen 11, 1; setup-foreign 11, 1; resilience 21, 4; setup-hpc all pass except SETUPHPC-F3's by-design residual. The partials and new findings became wave 3.
  • Final round (fix3-recheck on 341df2d06): 80 pass, 3 partial (T3-BE-5, T3-BE-13, T3-UI-18), 1 not reproduced (T3-SH-9, the Azure content filter would not trip), 1 not run (T3-UI-22, needs native macOS sheets CDP cannot drive).

Security-relevant changes for human review

CLAUDE.md asks for a person to review auth, permission, credential and concurrency code. Each change below had agent review only:

  • Broker quotas, headroom and retention (BROKER-1/2/4, W2-BRK-3/7): dfa33e47f, 270029505, 79bbed1df, 12bde6dba, 6c55eb9cd, 69b49ef20.
  • Broker runtime path and decoys (BROKER-3/5/6): 25b677397, 597dad5a7, 643b9fdfd.
  • Daemon loopback port and Agent Drafter app tokens (W2-HRD-1, 21 commits including 1827ff3c4, 8e0b93391, b49d3bb3a); secrets.yaml 0600 (71a22bc95).
  • Provider key checks and destination keys (W2-PRV-2, T3-SH-3/5/7): 8465f38b1, 81af85c37, cd51f7e62, c132cd0c8, acde3fe3b, 10183bfb3, fe1c876dd, 5e8070941, c20e4b063, b8d8dc9aa, d643a981d.
  • Crew scope, provider admission and grant stores (CROSSCUT-1/7/8, DAEMON-*): dceeb9eff, 47c021cda, 95b81e17f, d236fb9ee.
  • Institution, public-model and privacy-mode gates (W2-DMN-9, T3-BE-4, T3-BE-5): 035ab9639, 1249ece66, 7609429c1, 55f5702b5, 42e9526b0, 2264810c9.
  • Untrusted names and text (RENDERER-1/2/3, W2-UIC-2, W2-HRD-4..7, T3-BE-10/11/12, T3-UI-21): 5c9a13fc9, 68a6e17a8, 99301a85f, 765ecbdf8, 38c0f4236, 92a0b3cf3.
  • Daemon restart and reattach (W2-SHL-1, T3-SH-10): 3ff0b4aee, 73cc7e101, 473dd2819.
  • Credential store classification and vault (W2-DMN-1, CLI-2): 7f44f4750, c9a740f79, fb6dc7717.
  • Exported app launcher token check (found by CI): 92094da45.
  • Sign in window typeahead (T3-BE-17): c7a75c476. Disclosure before a public model's first turn (T3-SH-1): d2a1d9c9c. Declassify refusal of a Crew chat (T3-BE-18): 91ac108d8.

Left open

  • T3-BE-13 / T3-UI-18 (P2): a member whose app idles on Home, and who writes nothing, may not learn that the server stopped saving; the app's bar changes only after a window focus event.
  • T3-BE-5 (P2): --expected-mode public is still accepted on a connection whose own mode is Public while the workspace is Private for everyone. It fails safe: the content stays restricted.
  • Not verified live: T3-UI-22 (native sheets), T3-SH-9 (content filter), the packaged-build reattach prompt after a daemon restart (harness limit).
  • Final-round observations (P2, not yet triaged): grants list shows a grant ended by a personal settings change as Active; run --resume of such a chat prints the app's wording; a JSON error names a channel by ID; SSH refusals embed Code:/Details: lines in JSON error; the host's files status for a storage-paused upload says "Ask the host"; the member broker-down copy says "within a few minutes" where 13 to 16 s was measured; Configure Versa API Azure opens with focus on Remove; Escape from Switch models opened from the composer leaves focus on the page; the configuration editor's Done discards an unsaved field; connect prints "Status: Connected" for an unjoined connection.
  • Earlier leftovers no later wave closed: a link whose public-looking host resolves to a private address is refused without saying why (RENDERER-3 residual); snapshot member lists are bounded by per-member caps, not paged, and channels list/teams list do not say when a listing was cut short (BROKER-2 residual); an Agent Drafter app cookie is not port-scoped, so on a shared host a system browser could send it to another account's loopback server (W2-HRD-1 residual, needs HTTPS or a per-request proof), and it is not rotated on redemption; the secret guard does not cover <state>/apps-daemon/ records or <state>/app-launch/ pages; the new browser opener path ran on macOS only.
  • By design or deferred: the daemon keeps its loopback listener and a secretless /status (documented); toast position left as the user decided; reply, direct messages, mute and leave are feature requests.

🤖 Generated with Claude Code

…he umask

With the keyring disabled, secrets.yaml was written with a bare
std::fs::write, so under umask 022 it was -rw-r--r-- in 0755 folders and
every provider key was readable by any account that could traverse the
home (PROV-F3). On unix it is now staged 0600 (set on the descriptor),
synced and renamed into place, a missing folder is created 0700, and a
looser existing store is tightened before the write. A secrets.yaml kept
behind a link is followed and stays a link.
… shared daemon restarts

The local proxy pins the daemon instance it verified and refuses every
request once that instance is gone or replaced, which is right, but nothing
ever attached it to the new one: after any daemon restart (a crash,
'crew daemon stop', a CLI that started a new daemon) every chat and Crew
action failed with 'reconnect explicitly' until the app was relaunched, and
no control reconnected (R-1).

The proxy now reports the loss (gone or replaced, once per change) and
answers with a daemon_restarted JSON refusal in plain words. The main
process asks once in a native prompt ('Biorouter's background service
restarted. Reconnect?'); on Reconnect it rediscovers the profile's daemon
exactly as a launch does, asks for that instance's approval secret (never
reading one from disk), or starts a new daemon with a new secret when none
answers, and retargets the SAME proxy after verifying the new instance's
identity and that it accepts the secret. Every window keeps its
BIOROUTER_API_HOST, so nothing is reloaded and no draft is lost. The proxy
still never follows a different instance by itself.

After 'Not Now' or a failed reconnect, the sidebar shows a standing notice
with Reconnect and Quit and reopen (the existing restartApp), start-chat
failures say the same, and a new window asks again instead of failing its
readiness check and quitting the app.
…request

A person's signed request on a connection already counted for them now only
moves their last-request time, instead of cloning their principal ID into two
map lookups on every request.
…y, and check member-only channel IDs

Api::resolve ended in a plain bail!, so a JSON error for an unknown or
ambiguous name had no code (DW-10). It now keeps the resolver's own code,
unknown_name or ambiguous_name. 'You're not in a channel called #x' read as
a membership problem for a typo or a renamed channel (M12); it is now 'No
channel you're in is called #x.' (and the same for teams).

A channel ID skipped resolution, so an ID from another workspace reached
the daemon's 'refresh' refusal (AG-F13). Commands only a member can run
(history, search, watch, send, mark-read, files, tasks start, grants grant)
now check an ID against the person's channels when the snapshot lists them
all, and refuse it in the words a name gets. A partial snapshot, and host
commands that may act on a channel the host is not in, leave it to the broker.
…access binds and as whom it posts

Chat access offered an enabled Allow for a model the workspace would
refuse, showed the refusal only after the click in the daemon's words
('...the model's resolved affiliation'), and the refusal then stayed as a
red banner over #general once the pane closed (SF-F5, SF-F4). Its consent
named no model and no workspace, said nothing of the first access fixing
them, and said the chat posts as the person rather than as their agent
(AG-F1).

- useChatModel reads the chat's bound model from this window's chat
  store and follows its binding announcements; the pane runs the same
  institution and public-model checks Ask my agent runs, disables Allow
  with the reason, and words the daemon's refusals as Ask my agent does
  (by their new codes, or an older daemon's words).
- The consent names the workspace, the model with its tier chip, 'The
  first access fixes this chat's workspace, channel and model.', and
  posts 'as Alice Chen (@alice)'s agent'.
- Leaving Chat access (closing the pane or switching mode) dismisses its
  error, as leaving Ask my agent already did (W2-UIW-14, W2-UIW-16).
…, and settle a post in doubt

The composer printed the broker's code: text after Couldn't send. for every
refused post (QA M1, M5, R-2, FILES-F9). composer/sendFailure.ts reads the
refusal as the dialogs do (dialogs/refusals.ts) and gives one sentence each:
the 64 KB limit, a full workspace, a storage failure (addressed to the host,
or telling the host to restart), a file or path shared in another channel, an
archived channel, and a lost bridge. A message over 64 KB of UTF-8 (or twice
that escaped) is caught before it is sent, and kept.

A send failure now belongs to its channel and draft (QA M5, R-4): it carries
its destination, shows only in that channel's composer, goes at the draft's
next edit, is put aside with the draft and comes back with it, and a failure
of the link is dismissed when the connection verifies again.

A post whose outcome is unknown (crew_outcome_unknown, or the older daemon's
'outcome may be unknown' text) is no longer 'Couldn't send': the composer says
it is checking, then that it was sent once the message is in the channel, or
that it could not confirm it after the channel was read again, keeping the
words and their key. Each send records its outcome and the message id the
broker answered with, so a deduplicated resend is delivered at once instead of
leaving a Sending ghost (QA R-4).

The text box stays writable while a post is on its way, and success takes out
only what was sent, so words typed after Enter are kept (QA M8). A draft put
back gets its caret at the end (QA M9). Several dropped files leave a red note
naming the one Crew took, until closed (DW-18).
…t line breaks

A quoted-printable body wraps a long line with '=' before each break, and a
paste of the raw text keeps it. The host's token has no padding, so an '='
ending one of the token's runs is dropped when the runs are joined.
…ing an empty queue

The broker leaves pending_joins out of a non-host's snapshot, and the CLI
read the missing list as empty, so a member was told 'No one is waiting to
join.' with exit 0, as a host with an empty queue is (DW-13). A missing list
now refuses: a member is told only the host can see it (crew_host_required),
and a host whose server leaves it out is told the server doesn't support
joining by name (crew_join_by_name_unsupported).
…n nothing is remembered

channelForTeam fell back to the team's first open channel in snapshot order
and teamForView to the snapshot's first team, and the snapshot lists both by
their random IDs, so Bob, Henry and Mallory each landed on #random instead of
where the host's welcome is (setup F6). A remembered or current choice still
wins; otherwise the team's general_channel_id opens while it is open, and the
team shown is the first by name.
Crew downloads are written by biorouterd, not a download manager, and
carried no quarantine mark, so Gatekeeper never checked an app inside a
shared .zip or .dmg (FILES-F8). Before the .part file is renamed into
place, macOS sets com.apple.quarantine on the open descriptor and
Windows writes a Zone.Identifier stream naming the Internet zone, so the
mark arrives with the name. Best effort with a warning; the daemon's own
receipts are not marked.
… chat-access link

On /crew?sessionId=... the app sidebar sees the same path and only announces a
same-route reset, and nothing in Crew listened, so the chat's connect note and
its Chat access pane stayed until the person went Home and back (SF-F5). The
controller now hears useSameRouteReset('/crew'): it closes a Chat access pane
and replaces the location with plain /crew, which drops the link's chat and
its route-state intent. It is a replace within Crew, not a navigation to it,
so the navigation census is unchanged; a source guard pins the wiring.
… word the institution refusal

Connection settings had no client check on Identity file, and the
daemon's 'Identity file must be an absolute path' rendered only as the
dialog's one note at the end of the scroll body, below the view, with the
field unmarked and nothing focused, so Save looked like it did nothing.

- Identity file and Remote work folder are checked as typed (an absolute
  path; a leading ~ refused with the same sentence), marked invalid with
  the note under the field, and a submit with Advanced closed opens it on
  the field.
- The daemon's field refusals (identity file, jump hosts, login, name)
  and its institution refusal are said under their field, which is
  marked invalid and focused; editing the value takes the refusal away.
  The institution refusal names both institutions instead of 'Crew
  aliases have different institutions...'.
- A dialog's one error note scrolls itself into view (W2-UIW-13, and the
  Connection settings part of W2-UIW-14).
…names are unavailable offline

files watch streamed each receipt with default options, so its rows said
'download from this channel' where files status said '#methods', and
--show-ids was ignored on them (DW-07). The rows now use the snapshot's
names, read once, and --show-ids.

When the snapshot could not be read, names() fell back to an empty
directory, so offline grants list and files pending printed 'this channel'
on every row (AG-F17, R-8). Such a row now says 'a channel in WORKSPACE
(names unavailable while disconnected)' with the channel's ID, the one
thing that tells two rows apart.
…ost first, and say when a list is partial

The broker sends teams, channels and principals in the order of their
random IDs, and members, channels list and workspace show printed them as
they came, so two teams' channels interleaved and #general could be fourth
(M17, F6, SF-F10). Channels now group by team (teams by name), each team's
#general first and the rest by name; people come host first, then you, then
by the name their row shows, as the desktop's peopleInOrder does. Applied to
text and to the JSON arrays of members, teams list and channels list.

A bounded snapshot lists only as many teams and channels as fit, counting
the rest in totals; the lists now say 'Showing N of your M channels.' when
some are left out (a wave-1 follow-up).
…its model is fixed, and going private is said first

W2-PRV-6: an unsent chat has no session, so its model chip opened the
new-chats scope and a pick rewrote BIOROUTER_PROVIDER/MODEL for every
window. BaseChat now provides a pending-model scope to a chat with no
session: Switch models offers 'this chat' with the same 'Also use for new
chats' box, holds the pick, the chip names it, and the first send binds it
through the per-chat update_provider right after /agent/start, before the
message goes anywhere. A refused bind sends nothing and hands the text back.
Home keeps its explicit new-chats scope.

W2-PRV-15: a Crew chat's model is fixed by its grant. Switch models now says
so up front ('This chat's model is fixed by its Crew access. Start a new
chat to use another model.', the daemon's own sentence) and disables every
control that could change it. A crew_model_fixed refusal is shown in the
dialog as that sentence instead of a '<provider>/<model> failed' toast plus
'then try again'.

W2-PRV-7: switching a public chat to a private model now says, before the
switch, that the chat becomes private after the next message and how to make
it public again from History. The barred-row reason in a private chat names
History > Make public, and the comment claiming declassification did not
exist is corrected.
…who is online

members add printed 'Added. @crew_carol can now see #general.', naming the
person by username alone and omitting the team, though every team has a
#general (M11). It now says 'Added "Carol Nguyen" (@crew_carol) to Chen Lab.
They can now see #general.' for a team, and for channels names the team in
front of a channel whose name another of the person's teams shares.

Member and people rows append '· online' for the people the broker reports
in online_principal_ids (W2-BRK-6), and members JSON carries is_online; an
older broker reports no presence and none is shown (M18).
…say what Retry found

In the chat's Crew bar, an offline revoke showed 'Stopped on this
device. Crew confirms it with the workspace when it reconnects.' and,
beneath it, 'Crew access to #jobs was removed, so this chat can't
continue', with Retry, Start a new chat and Grant access again for one
state (AG-F11). A Retry that got the same 503 while still offline
re-rendered the identical sentence, so nothing visibly happened and the
alert was not announced again (AG-F12).

- The lapsed-and-unconfirmed bar shows one note naming the workspace,
  with Retry and Start a new chat; Grant access again waits for the
  connection it needs.
- RevokeResultNote, shared by the bar, the Chat access pane and the
  access list, counts a Retry that comes back unconfirmed while offline
  and then says 'Still can't reach {workspace} · checked just now.
  Connect to confirm it now.', mounted afresh so it is announced, with a
  Connect control where the surface has one (W2-UIW-18).
Crew now opens the first team by name when nothing is chosen (setup F6), and
the controller tests' second team, Imaging, sorted before Lab, so eighteen
tests about drafts and observers opened the wrong team. The fixture is renamed
Microscopy Imaging; what those tests pin is unchanged.
… of deleting it

Losing access to a channel cleared the composer and forgot any draft kept
for it, and a draft kept for a channel the person had moved away from went
without a word (QA M10). The words are now offered once, in a note above the
message box (or on the no-channel screen) with Copy draft and Dismiss. They
live only in that note's controller entry, bounded to five, are filtered to
their connection, are dropped on dismiss, and are never put back into any
composer. forgetConnectionDrafts returns what it forgot so the quiet case is
announced the same way. Privacy and access clears still clear without an
offer.
… dock, and notify while the person is elsewhere

A mention or a new message was invisible until the person opened Crew: the
Crew item had no count, nothing set a dock badge, and nothing in Crew sent
a notification (M2).

AppSidebar now mounts a watcher (components/crew/attention) that reads each
connected workspace's snapshot every 10 s with the person's proof, sums its
unread map for a badge on the Crew item ('99+' past 99, 'Crew, N unread' to
assistive technology) and reports it to the main process, which puts the
largest any window reports on the dock. When a channel's count rises while
the window is not showing Crew in front, it reads only that channel's new
messages (at most 20, which marks nothing read) and asks for a notification:
'Alice Chen mentioned you in #general' for a mention of @you outside code,
otherwise '3 new messages in chen-lab'. The main process validates it, keys
its limit by the ids, shows at most one per channel a minute across every
window and none while another window of the app is in front; clicking it
brings the window forward and opens Crew on that channel. A snapshot read
every 10 s costs a twentieth of a live observation and holds none of the
daemon's observer slots.
… add files show

A message carries only its attachments' IDs, and history printed '2
attachments [attachment IDs a, b]', so the manual's recipe (take the ID from
history --show-ids, then files download ID) could not pick the right file
when a message had two (DW-17). history, search and watch now look each ID
up with blob.status (once per command, at most 40), print 'Attachment:
counts.csv (55 KB)' beside its ID, and add attachment_details {id: {name,
size, media_type}} to the JSON beside the unchanged attachments array.
'files show ID' shows a file's name, size, type and channel.

The observer callback is now async so watch can look names up between
frames.
…s by naming the path, and pin the codes the picker words

security-hardening landed the download refusal codes (99301a8). The two
the picker words itself, crew_file_name_hidden and crew_folder_shared, now
match the daemon's sentences word for word, held by a test that reads
local_files.rs. crew_destination_is_folder, crew_destination_exists and
crew_file_is_program name the path they refuse, so the picker shows those in
the daemon's words instead of a rebuilt sentence of its own.
… 'Use other provider' back to the chat

A signed-in Codex showed Ready in the catalog but was missing from a chat's
Switch models list: the daemon reports a coding agent unconfigured until
'Use <agent>' saves its command key. The picker now reads the shared
coding-agent status (probed once per renderer, never polled) and offers a
signed-in agent that was never connected; choosing it saves the same command
key before the switch.

'Use other provider' carried nothing, so the catalog's Back went to
Settings > Models, and 'Use Codex' wrote BIOROUTER_PROVIDER globally before
any model was chosen, then opened the model dialog for new chats. The
picker now passes the chat, its tier and the screen to return to; the
catalog route returns there and scopes its model step to that chat; and
connect() writes only the command key, leaving the provider and model to
the model step, written together to the scope picked there.
…ribe a connection the same way in both places

- The privacy chip's institution, cut to 'stanf...' in the 240px column,
  carries its full words as a title while it is cut (no tooltip while
  it fits, as Q2-17 wants).
- The workspace menu's fingerprint line wraps between its digit groups
  instead of showing 13 of 16 digits a host reads out, and the menu's
  title, host and 'Signed in as' lines carry their full words as titles.
- The popover's 'Your connection' row is the same one badge Settings >
  Privacy draws, with the connection's own institution, and both
  Institution rows show the institution in force (the workspace's, else
  the connection's) saying whose it is.
- Settings > Privacy's 'what this changes' line sits under the 'Make my
  connection public...' row it describes, not under Institution
  (W2-UIW-15).
…ow a grant's scope in crew context

'biorouter session --shared-daemon' wrote the raw SessionBinding JSON to
stderr in text mode (AG-F7). Text now says 'Chat ID is ready
(provider/model).'; the JSON formats keep the binding.

'crew context' printed only the manifest's messages, newest first, never
the grant's channels that its help promises. Text now leads with 'Access:
#dest · also reads #a, #b' (the channel it posts in comes from the grants
list), then the messages oldest first. JSON is the manifest as before.
…r channels they name

The per-channel limit keyed by ids let a busy workspace, or a window naming
channel after channel, put one notification per channel on screen at once.
…s and privacy show

status, connections show and privacy show printed the connection's own
setting under the bare label 'Privacy', so after 'privacy set-personal
public' in a workspace that is Private for everyone the CLI said 'Public'
while the app said 'Private' (SF-F1). They now read the workspace's policy
for a connected Public connection and lead with the effective privacy
('Privacy: Private (okafor-lab is Private for everyone)'), with the
connection's own setting labelled 'Your connection'. Offline they say the
workspace's setting can't be checked. JSON carries effective_mode (null
when unknown), workspace_mode and workspace_name.

'privacy set-personal public' in such a workspace still asks for the typed
name, says in its question that nothing changes while the workspace is
Private for everyone, and says so again after saving.
…ort a replacement once, and bound attention reads

Review of the R-1 and M2 changes: an instance that answered again while the
'Reconnect?' prompt was open was ignored, so 'Not Now' left every window
saying the service restarted; concurrent requests that all met a replaced
instance each reported it; and the attention watcher's reads had no
deadline, so one hung daemon request froze the Crew badge for good (now
15 s each).
…mpt reads it

The sign-in runs `ssh -M -N` in a PTY whose line discipline had ECHO on.
OpenSSH turns echo off only while readpassphrase reads, so text typed during
PAM's delay after a wrong password was echoed in clear onto its own line, and
stayed in the window's scrollback and xterm's accessibility text.

ECHO and ECHONL are now turned off on the terminal before ssh starts (through
its own device, opened without making it a controlling terminal), so what
readpassphrase restores is off too and typed-ahead text is never drawn. The
one prompt that shows its answer, OpenSSH's host-key yes/no question, is drawn
by the daemon: while that prompt is the last line ssh wrote, printable input
is written back to the window and an erase takes one character back. What
reaches ssh is unchanged.

T3-BE-17.
… every CrewError field against the OpenAPI schema

A save refused because its institution is not the workspace's own (T3-BE-4)
answers crew_institution_mismatch with connection_institution,
workspace_institution and workspace. REFUSAL_FIELDS did not list
workspace_institution, so --output-format json dropped the one field that
names the institution to retry with.

REFUSAL_FIELDS is a hand-kept allowlist, so the next CrewError field would
have been dropped the same way. The fields the CLI writes itself are now
CLI_WRITTEN_FIELDS, and a test reads CrewError's properties from
ui/desktop/openapi.json (whose freshness CI already checks) and fails when a
property is in neither list, or when either list names a property CrewError
does not have. The JSON failure test gains the save refusal.
…every way on, and pin CrewError's fields to JSON output

T3-CLI-4: the daemon answers crew_connection_exists for an existing
connection only when a setting it compares differs (same_settings: server
login, port, key file, jump host, mode, institution); a matching save
succeeds. The terminal still said "Run ... join to finish joining", which
never said nothing was saved and pointed at the connection the person had
just asked to change. join-invitation --mode private --institution ucsf over
a saved Public connection printed "You'll join it as Private (UCSF)" and
then advice that connected the Public one.

The refusal now says the workspace is already saved as NAME with a
different server login, port, key file, jump host or privacy, so nothing
was saved, and lists a command for each way on: connections show to
compare, privacy set-personal with the privacy the invitation would save
(the preview's mode and institution, quoted for the shell and escaped),
connections update FILE for the route, connections remove and save again
while it has not joined, and join only to use it as it is. JSON keeps the
daemon's code and connection_id.

T3-CLI-10's tests: a test reads CrewError's properties from
ui/desktop/openapi.json and fails on any field neither copied into a JSON
failure nor written by the CLI, and the JSON failure test gains the
workspace_institution save refusal.
…ect waits

The immediate re-dial is still under way when the bridge is first retired,
and "in a moment" is true then; the test now waits until the schedule records
its long gap, so it no longer races the failed dial under load.

T3-BE-16.
…saving one no chat can start on

unknown_provider_refusal trimmed the value before comparing it with the
registry, but upsert_config stores the value as sent and the factory looks the
name up untrimmed. So 'openai ' passed the check, was saved, and every new chat
then failed with Unknown provider: the failure T3-SH-7 exists to close.

The name is now compared exactly as it will be stored. A padded name whose
trimmed form is registered is refused with a sentence that says to remove the
spaces, rather than rewritten: every gate before this check judged the value as
sent, and the value written must be the value they judged. The unit test that
pinned ' versa_azure ' as accepted now pins it (and 'openai ', a tab and a
newline) as refused, and the keyed-daemon route test shows nothing is written.
refresh_custom_providers removed the custom entries under one write lock and
loaded them again under a second, so a new chat started between the two found
its custom provider missing (Unknown provider). /config/upsert now re-reads the
custom providers whenever it is sent a provider name it does not know, which
made that window reachable from an ordinary settings write.

The files are now read and parsed before the lock is taken, and the old entries
are replaced under a single write lock. A failed read leaves the registry as it
did before: the custom entries removed and nothing half-loaded. A source guard
pins one write() in the function and the read before the lock.
…hange's reconnect has no verified view

Workspace settings stays open through the refresh its own Make private causes
(SF2-N7), but the daemon's policy_changed end says clear: true, which drops the
last verified copy with the live one. Until the view verifies again the dialog
had no snapshot, the people directory named nobody as the host, and the host
read "Hosted by Unknown member", "No one else has joined … yet.", "Only the
host can invite new people.", "Only the host can change …'s privacy." and an
institution attributed to the connection alone, with the Workspace row blank.

The dialog does not keep its own copy of the snapshot, because the clear: true
contract says nothing verified stays on screen. Instead every tab says
Checking… in place of what the snapshot decides (who hosts, who is in, the
workspace's own mode and institution), the panels are aria-busy, and the saved
connection's own rows (server, Your connection) stay as they are. The host's
view comes back when the view verifies. survivesRefresh and useDialogView now
say which path keeps the last verified copy and which does not.

Tests: a render test drives the real controller through Make private, a held
re-observation and a policy_changed end, and checks the host never reads a
member's words and gets the host controls back; a dialog test pins each tab's
checking state. Both fail against the previous dialog.
…o could be stale

Review of T3-BE-5: an expectation of Private met only because the workspace's
last signed hello said it was Private for everyone could go stale if the host
allowed Public since. A post that required Private is now told to the
workspace as Private, so it is restricted whatever the workspace's mode is
now; and a task or grant that required Private is judged again against the
snapshot admission reads, whose policy epoch the workspace checks when the run
is created, so one that allows Public now is refused before any run exists.
What the desktop sends (the connection's own mode) is told as before.

Also moves the institution checks of a save into saved_institution and the
worker's answer bookkeeping into note_worker_answer, with the worker method
list as a constant, to keep both functions within clippy's line budget, and
puts require_mode's doc comment back on it.
A redundant closure in connect_recorded, and a string slice in
bounded_detail, which now keeps whole characters up to the limit.
…me to what the drop flow words

A file dropped or pasted into a channel is shared by shareDroppedFile, whose
refusals are worded by crewFileRefusal and otherwise get the general note
"Crew couldn't take …". It has no case for crew_file_name_invisible yet, so a
drag gets that note while the paperclip and the command line show the
daemon's sentence. New rule drop-refusals reads crewFileRefusal's cases and
fails either way: while the code is unworded the row must say what a drop
shows, and once a case words it the row must drop that note and quote the
drop flow's sentence whole. The hidden-character family now also reads the
drop flow's copy, so the page quotes one sentence for both. The row now says
paste as well as drag, and says it plainly.
… or a fall-through label

drop-refusals took the copy a code's case returns only when the return
followed the label directly, so a case written as a block, or reached through
a label that falls through, would have read as unreadable. It now takes the
one crewShareCopy entry the arm uses, past any fall-through labels.
Review of T3-BE-8: a remote path may be 4096 bytes and the line names up to
32 of them, which whole could outgrow the post. Past 120 characters a path is
shown by its end, where the file name is, after an ellipsis.
…he connection instead of deciding it as a loss

The daemon disconnects, saves and connects again inside one PATCH
(CrewManager::update). The old observer wakes inside it and ends with
policy_changed; the 300 ms re-observation then read the record as
disconnected, met no transport (observation_refused) and was decided as a
lost connection: observationFailure closed Workspace settings and flashed an
error until the PATCH returned.

While a save that began on a connected connection is on its way, an end a
reconnect explains is now left to the save, in the observer and in the loss
handler alike. Nothing verified stays on screen, nothing is decided, the
status reads Updating… and the screen Checking rather than Offline, and the
offline follow waits. Once the last save is back, an end left to it is
observed again unless the save's own refresh already did. Answers about
access or the person are still shown at once. A save of a connection that
was not connected leaves nothing to the save.
…pace used it

Tries while the workspace server is down spend the schedule's time budget
without taking gaps from the growing list. A failure of another kind after
them (the host stopped the server, then its computer left the network) took
every gap still listed, each cut to the zero time left, and dialled back to
back until the list ran out: about 12 SSH spawns in well under a second under
the default timing, breaking the rule that the daemon's own dials are always
spaced (W2-DMN-6). The schedule now ends in every branch once no time is left.

The new test injects timing where one steady wait uses up the whole budget and
the next try is unreachable: 4 dials with the fix, 36 without it.
…orce, and begin a Private upload as Private

T3-BE-5 (review): POST /crew/files still compared an expected mode with the
connection's own mode alone, so 'biorouter crew --expected-mode private files
upload' from a personal Public connection in a workspace that is Private for
everyone was refused as 'Your connection is Public, but this request required
Private' while status and privacy show said Private. The selection is now
judged by the rule every other door uses (require_expected_mode, against the
workspace's mode from its last signed hello).

An upload accepted only because that hello said Private could go out
unrestricted if the workspace has allowed Public since, because remote() told
blob.begin the connection's own mode. An upload's selection that required
Private now marks its receipt requires_private (persisted; a resume never
clears it, and may add it), and blob.begin asks for Private, which the daemon
holds against the workspace's latest hello: refused if it now allows Public,
otherwise told to the workspace as Private so the attachment is restricted.
An upload that required Private also never adds a part to an attachment the
workspace does not restrict, such as one a resume finds begun earlier. Either
refusal ends the transfer failed with the refusal's sentence, since a
reselection would meet it again. A start replay must ask for the same privacy;
receipts that did not require Private keep their intent digest.

The OpenAPI spec gains Receipt.requires_private, and the POST /crew/files
description of crew_mode_mismatch now says what is compared.
…per to clear too_many_lines

The provider-name check and its custom-provider re-read grew upsert_config
to 107 counted lines, over clippy::too_many_lines (100), which is not
baselined for this handler. The check now lives in provider_name_refusal,
called at the same point: after every privacy gate, before the write.
The source guard now also asserts the helper asks the registry, re-reads
the custom providers, and returns the second answer.
… final live round

Adds docs/history/biorouter-crew/evidence/merge-qa-2026-09-27.md: both
fleets' shapes, the lanes of run 1 and run 2, task outcomes, findings by
severity (59 static, 113 live leading to 117 wave-2 items, 90 run-2
findings leading to 85 wave-3 items), every item mapped to the commits
that fixed it, the final round's live verdicts on 341df2d, the
security results and what is still open. Text only; the raw evidence
stays local. The history indexes link it.
- configure transcripts (permission modes, workflows, quickstart) use the
  menu labels and mode names configure now shows, and troubleshooting
  quotes its new save prompt
- BIOROUTER_MAX_TURNS takes a whole number of at least 1; anything else
  is ignored with a warning
- the non-private-model notice also appears when a public model is picked
  for an unsent chat, and sits above the Public tab's API providers
- the landing providers page says configure saves nothing until the
  provider accepts the key, and the Crew page lists the stopped-saving bar
- the Crew manual says status reads Not joined yet, that grants list can
  still show a grant ended by the person's own settings change as Active,
  and that the stopped-saving bar refreshes when the window comes to the
  front
…uses upper case

open_url checked the one-time token with *[!0-9a-f]*. macOS's /bin/bash
3.2 reads that range by the locale's collation under a UTF-8 locale and
accepts upper-case letters, so an upper-case token passed the check and
reached the opener. Hosted CI on macos-26 caught it; it reproduces locally
with /bin/bash and LANG=en_US.UTF-8. The digits and letters are now listed.

The syntax test also runs its find_biorouterd check from a file instead of
bash -c: on Windows the launcher crossed the command line, and Git Bash's
argument parsing reported an unmatched quote in text that bash -n of the
same file on disk accepted.
…m a child process

- crew::tests asserted that a request past the privacy guards carried no
  typed refusal at all. On a Linux runner the fixture's credential lookup
  answers crew_credential_store_unavailable, which is typed, so the tests
  failed while the guards behaved. They now assert that no privacy guard
  (mode, public model, institution) stopped the request.
- The claude_code and codex fake CLIs were written with fs::write. A test
  on another thread that forks while that handle is open gives its child a
  copy, and exec of the file then fails with ETXTBSY (seen on ubuntu CI).
  A child sh now writes and chmods the file, so this process never holds
  a write handle to it.
… or pressing Enter

The bar publishes the chat's hold from an effect after it renders, so on a
loaded CI runner the hold landed a render after 'blocked' turned true: one
test read it as 'none', and one pressed Enter before it existed and waited
out its 5 s timeout for a toast. Each Enter-after-block test now waits for
the hold first.
@Broccolito
Broccolito merged commit 2ca4a7e into main Sep 29, 2026
18 checks passed
@Broccolito
Broccolito deleted the claude/crew-qa-fixes-2026-09-27 branch September 29, 2026 08:33
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant