How a Background Thread Silently Ate Your Login

We shipped an account switcher in v0.5.4. It worked great — unless you first signed in with a GitHub account that didn't have Copilot access. When you switched to the right account, the UI stayed stuck on "Waiting for authorization..." forever. The fix was a six-line architectural change, but the bug exposed a pattern worth documenting.

The setup: two callers, one resource

The brainstem's device code login flow has two consumers. A background thread (_bg_poll_loop) polls GitHub so the token gets captured even if the browser disconnects. And the client polls POST /login/poll every 5 seconds so the UI updates when auth completes.

Both callers invoked the same function: poll_device_code(). That function, on success, saves the token and clears the _pending_login state. Whoever gets there first wins the token. The loser finds _pending_login empty, returns None, and tells the client {"status": "pending"}. Forever.

Background thread Client poll (/login/poll) │ │ ├── poll_device_code() ────┐ │ │ ✓ got token │ │ │ ✓ cleared state ───────┤ │ │ │ ├── poll_device_code() │ │ │ ✗ state empty → None │ │ │ returns {"pending"} │ │ │ │ │ ├── ...forever ▼ ▼ ▼

Why it only showed up on account switch

On a normal first login, the race exists but is harmless — either caller reports success and the UI dismisses. On an account switch, the first (wrong) account goes through the full device code flow successfully at the GitHub level. GitHub grants a token. But the Copilot token exchange fails because the account has no Copilot license. The background thread catches this error silently, and the client is left polling an empty _pending_login dict.

The second login attempt (correct account) hits the same race. If the background thread wins — which it reliably does because it's already polling at the right interval — the client never sees the result.

The fix: single-writer pattern

The solution is to stop having two callers compete for the same resource. We introduced a shared _login_result dict. The background thread is now the sole caller of poll_device_code(). When it gets a result — success or failure — it writes to _login_result.

The /login/poll endpoint no longer calls poll_device_code() at all. It just reads _login_result. Python's GIL makes dict assignment atomic, so no locks are needed.

_login_result = {}  # Written by bg thread only

def _bg_poll_loop():
    token = poll_device_code()
    if token:
        try:
            get_copilot_token()
            _login_result = {"status": "ok", ...}
        except NO_COPILOT_ACCESS:
            _login_result = {"status": "error", ...}

@app.route("/login/poll")
def login_poll():
    if _login_result:          # bg thread wrote something
        return jsonify(_login_result)
    if not _pending_login:     # no flow in progress
        return jsonify({"status": "expired"})
    return jsonify({"status": "pending"})  # still waiting
Pattern: When a background thread and a request handler both need the result of the same operation, don't let them race to call it. Have one writer and N readers. The writer owns the function call; the readers check a shared result.

Bonus fixes

  • NO_COPILOT_ACCESS surfaces to the UI. Previously swallowed — the user saw "Authenticated!" then got errors on first chat. Now the login overlay shows "username doesn't have Copilot access" with "Switch account" and "Sign up for Copilot" links.
  • Client poll has a timeout. 180 attempts at 5-second intervals (15 minutes, matching GitHub's device code expiry). No more infinite loops.
  • Stale Copilot cache cleared on new flow. Starting a fresh device code now wipes .copilot_session and the in-memory cache, so a previous account's session can't bleed through.

The lesson

Background threads that "help" by doing the same work as a request handler create races that only show up under specific timing. The thread was added to capture tokens when the browser disconnects — a real need. But it should have been the only writer from day one, with the HTTP endpoint as a passive reader. Adding a background optimization to an existing request path requires rethinking ownership, not just adding another caller.

Why We Killed Remote Agents (For Now)

The brainstem shipped with a feature that felt magical: paste a GitHub repo URL, toggle agents on, and they'd hot-load into your running server. No restart, no file copying. It worked by fetching manifest.json from the repo, downloading individual *_agent.py files, shimming their imports, and injecting them into the running Python process.

We removed all of it in v0.1.0. Here's why.

The complexity cost

Remote agent loading touched almost every layer of the system. It required URL normalization (handling github.com/, owner/repo, GitHub Pages URLs), manifest fetching (with three fallback strategies), file downloads, sys.path manipulation, sys.modules shimming, auto-pip-install on import failures, persistent config in .repos.json, restore-on-startup logic, and four HTTP endpoints for the UI to manage it all.

That's a lot of surface area for a feature whose primary value — making agents available — can be solved by just dropping a .py file into a folder.

The statelessness argument

The brainstem is designed to deploy as an Azure Function. Azure Functions are stateless by design — each invocation starts clean. Caching agents in memory and hot-loading from remote repos assumes a long-lived process. That assumption breaks in production.

By making agent loading stateless (fresh discovery every call, no cache, no remote state), we made the brainstem behave identically whether it's running locally on Flask or deployed as a serverless function. Same code path everywhere.

Design principle: If it works differently in dev vs prod, it's a bug in the architecture, not a feature.

What stays

The import shims (_register_shims) still exist. They're valuable for a different reason: agents written for the Azure deployment import utils.azure_file_storage, and the shims redirect those imports to local_storage.py so the same agent code runs locally without modification. That's a portability feature, not a remote-loading feature.

The auto-pip-install logic also stays. If a local agent imports beautifulsoup4 and it's not installed, the brainstem installs it and retries. That makes onboarding frictionless — drop in an agent, the brainstem figures out the deps.

The path forward

Remote agents will likely return, but as a first-class packaging system rather than a runtime hot-loader. Think: brainstem install github.com/org/agents that clones the repo into your local agents folder, resolves dependencies, and you're done. Install-time, not runtime. Explicit, not magic.

agents/ ├── hello_agent.py ← local, loads automatically ├── my_custom_agent.py ← local, loads automatically ├── context_memory_agent.py ← local, loads automatically └── experimental/ └── converter_agent.py ← excluded from auto-discovery

Version Tracking with a Plain Text File

The brainstem installs via a one-liner: curl ... | bash. That same one-liner should also handle upgrades. The question is: how does the installer know whether to upgrade?

The approach

A single file — rapp_brainstem/VERSION — contains the semver string. That's it. No package registry, no GitHub releases API, no git describe parsing.

$ cat rapp_brainstem/VERSION
0.1.0

The installer does two things:

  • Reads the local VERSION file from ~/.brainstem/src/rapp_brainstem/VERSION
  • Fetches the remote VERSION from the raw GitHub URL

If they match, print "Already up to date" and exit. If remote is newer, proceed with the full install flow (git pull, pip install, CLI wrapper update). The comparison is a simple semver walk — split on dots, compare integers left to right.

Why not git-based detection?

We could compare commit SHAs or use git fetch + git rev-list to detect ahead/behind. But the installer runs before cloning on a fresh install. We need a mechanism that works with a single HTTP request against a raw file URL, even when there's no local git repo yet.

Why not GitHub Releases API?

The GitHub API requires authentication for higher rate limits and adds a JSON parsing dependency. A raw file on GitHub Pages is cacheable, fast, and works with a bare curl. The VERSION file is also readable by Python (brainstem.py reads it at startup), the shell installers, and humans — one file serves every consumer.

Bump process: Edit rapp_brainstem/VERSION, commit, push. The next time any user runs the one-liner, they get the update. That's the whole release process.

Exposed in the API

The version is available at GET /version and included in the GET /health response. The startup banner prints it too. Everything reads from the same VERSION file — there's exactly one place to update.

$ curl -s localhost:7071/health | python3 -m json.tool
{
    "status": "ok",
    "version": "0.1.0",
    "model": "gpt-4o",
    "agents": ["HelloAgent", "ContextMemoryAgent"],
    "copilot": "✓",
    ...
}