How a Background Thread Silently Ate Your Login
We shipped an account switcher in v0.5.4. It worked great — unless you first signed in with a GitHub account that didn't have Copilot access. When you switched to the right account, the UI stayed stuck on "Waiting for authorization..." forever. The fix was a six-line architectural change, but the bug exposed a pattern worth documenting.
The setup: two callers, one resource
The brainstem's device code login flow has two consumers. A background
thread (_bg_poll_loop) polls GitHub so the token gets
captured even if the browser disconnects. And the client
polls POST /login/poll every 5 seconds so the UI updates when
auth completes.
Both callers invoked the same function: poll_device_code().
That function, on success, saves the token and clears the
_pending_login state. Whoever gets there first wins the token.
The loser finds _pending_login empty, returns None,
and tells the client {"status": "pending"}. Forever.
Why it only showed up on account switch
On a normal first login, the race exists but is harmless — either caller
reports success and the UI dismisses. On an account switch, the first
(wrong) account goes through the full device code flow successfully at
the GitHub level. GitHub grants a token. But the Copilot token exchange
fails because the account has no Copilot license. The background thread
catches this error silently, and the client is left polling an empty
_pending_login dict.
The second login attempt (correct account) hits the same race. If the background thread wins — which it reliably does because it's already polling at the right interval — the client never sees the result.
The fix: single-writer pattern
The solution is to stop having two callers compete for the same resource.
We introduced a shared _login_result dict. The background
thread is now the sole caller of
poll_device_code(). When it gets a result — success or
failure — it writes to _login_result.
The /login/poll endpoint no longer calls
poll_device_code() at all. It just reads
_login_result. Python's GIL makes dict assignment atomic,
so no locks are needed.
_login_result = {} # Written by bg thread only
def _bg_poll_loop():
token = poll_device_code()
if token:
try:
get_copilot_token()
_login_result = {"status": "ok", ...}
except NO_COPILOT_ACCESS:
_login_result = {"status": "error", ...}
@app.route("/login/poll")
def login_poll():
if _login_result: # bg thread wrote something
return jsonify(_login_result)
if not _pending_login: # no flow in progress
return jsonify({"status": "expired"})
return jsonify({"status": "pending"}) # still waiting
Bonus fixes
- NO_COPILOT_ACCESS surfaces to the UI. Previously swallowed — the user saw "Authenticated!" then got errors on first chat. Now the login overlay shows "username doesn't have Copilot access" with "Switch account" and "Sign up for Copilot" links.
- Client poll has a timeout. 180 attempts at 5-second intervals (15 minutes, matching GitHub's device code expiry). No more infinite loops.
- Stale Copilot cache cleared on new flow. Starting a
fresh device code now wipes
.copilot_sessionand the in-memory cache, so a previous account's session can't bleed through.
The lesson
Background threads that "help" by doing the same work as a request handler create races that only show up under specific timing. The thread was added to capture tokens when the browser disconnects — a real need. But it should have been the only writer from day one, with the HTTP endpoint as a passive reader. Adding a background optimization to an existing request path requires rethinking ownership, not just adding another caller.