Troubleshooting¶
Fixes for the snags people hit most often. If none of these help, see Help & community.
The agent can't connect to a tensor server¶
Biopb resolves the server in this order:
- an explicit
--serverflag, BIOPB_TENSOR_URL(withBIOPB_TENSOR_TOKEN),- your control plane's own answer — it knows the host, the port, and whether it's TLS,
- the default
grpc://localhost:8815.
A stray BIOPB_TENSOR_URL is the usual culprit. It sits above your control plane in that
list, so if it is set anywhere in your environment — a shell profile, a launcher script, an
older setup — it silently wins, and your agent talks to that address instead of the server your
control plane manages. Check it, and clear it if you didn't mean it:
If it comes back in a new terminal, it is set in your shell profile (~/.bashrc, ~/.zshrc)
or in the Windows environment variables — remove it there too.
Because it bypasses the control plane, biopb deliberately won't attach your local credential to
it, so a BIOPB_TENSOR_URL with no matching BIOPB_TENSOR_TOKEN fails to authenticate even
when the address is correct.
To reach a server someone else runs, don't set this variable
Add the server as a source instead, so your own biopb proxies it and caches what you read. See Connecting to a server someone else runs.
Testing a server directly. biopb tensor query --server <url> checks the connection on its
own, without launching an agent. Use grpcs:// for a TLS server and grpc:// for a plaintext
one — a scheme mismatch looks just like an unreachable server.
A local server won't auto-start¶
The data plane is started by the control plane, so if nothing comes up, the control plane is usually the thing that isn't running. Check it:
If it's down, start it with biopb control start (or biopb dashboard). If the biopb
command isn't on your PATH at all, install the full
biopb system.
"port is held by a process the control did not start"¶
The control plane only manages a data plane it started itself, so it refuses to adopt one it
finds already on the port — that's what makes biopb control stop a complete teardown. Usually
something else is holding the port: an orphaned data plane, or another user's login session.
Find the owner (lsof -i :8815, or netstat -ano | findstr 8815 on Windows) and stop it, then
start the control plane again. Or move your whole deployment out of the way with --base-port.
A server starts but then fails¶
When auto-start fails, the browser shows the underlying cause inline, and the full server
output is written to ~/.local/share/biopb/log/ — check there for the real error, or read
it from the dashboard at http://127.0.0.1:8813/logs. Common
causes:
Port already in use¶
Most likely on a shared machine or HPC node where another user already holds the default gRPC
port (8815). Either:
- use the server that's already running instead of starting a second one — add it as a source (Connecting to a server someone else runs), or
-
move your whole deployment to free ports with one number:
All three listeners shift together (control = base+3, sidecar = base+4, gRPC = base+5), and the control plane publishes where it landed so your other commands follow it. In the containerized / HPC server the same convention is spelled
BIOPB_BASE_PORT=9000.
Each user gets a private on-disk cache (%TEMP%\biopb-cache-<username> on Windows,
<tmp>/biopb-cache-<uid> elsewhere) with its own lock, so multiple users running their own
server on the same node don't collide. See A workstation you
share for the full per-user setup.
Server started but not reachable in time¶
Startup exceeded the timeout. Check the server log for the real error, then try connecting again once it's up.
Biopb is slow to start on Windows¶
Windows Defender rescans biopb's DLLs and compiled Python files on every launch, which usually dominates the first-start wait. Biopb can add a Defender exclusion for its own install trees:
biopb quick-start # show the current status
biopb quick-start --enable # add the exclusions
biopb quick-start --disable # remove them again
It needs administrator rights, so expect one UAC prompt, and it's fully reversible. This is Windows-only — the command does nothing on Linux or macOS.
An image looks scrambled or "ghosted" in the viewer¶
Once in a while a layer renders wrong: pixels from a layer you already closed bleed through, the image looks scrambled or mixed with unrelated data, or a stale frame "sticks" and won't update when you change the slice, zoom, or 2D/3D view.
Your data is fine. This is a display glitch in napari's GPU canvas, not a problem with your pixels — what the data server sent and the layer is holding is unchanged. It tends to show up after a session that has added and removed a lot of layers (napari doesn't always clear the old ones off the GPU cleanly, so a leftover image keeps drawing underneath the current one), and older or virtualized graphics drivers make it more likely.
To clear it:
- Restart the viewer — the reliable fix. Close and reopen the napari window, or ask your agent to restart its session. This rebuilds the canvas from scratch and clears the leftovers.
- Quicker things to try first (not guaranteed): toggle between 2D and 3D (press
2or3), resize the window, or remove and re-add the affected layer. If the scramble survives, fall back to a restart.
Re-add the layer to confirm — if it now looks right, the earlier frame was a display artifact, not your data. Keeping your graphics drivers up to date, and not bulk-adding then removing large numbers of layers in a single session, makes it happen less often.
Algorithm server problems¶
- No GPU / CUDA errors — confirm an NVIDIA driver (>= 525) and the NVIDIA Container
Toolkit
are installed, and that you launched the container with
--gpus=all. - Connection refused — check the published port (
-p 50051:50051) and that the agent is pointed at the right host and port.
See Algorithm servers for the full run reference.
Still stuck?¶
Open an issue on GitHub with the error message you saw, or ask on the image.sc forum. See Help & community.