Data (tensor) servers¶
A data server — a tensor server — is what stands between your microscopy files and everything else in biopb. It reads whatever format your microscope wrote and serves it as uniform, lazy, chunked arrays, so your agent can work with a 200 GB dataset on a laptop with 16 GB of RAM.
The default install already runs one for you, on your own machine. If you work alone on data that lives on your own computer, there is nothing on this page you need to do — skip to Working with your agent and come back when your setup outgrows it.
Which setup is yours?¶
Each recipe below is complete on its own. Find the one that matches and follow only that one.
| Your situation | Recipe |
|---|---|
| You share a workstation with other people | A workstation you share |
| Your lab has one store of data and several people analyzing it | One data server for the lab |
| Someone else already runs the server; you just need to reach it | Connecting to a server someone else runs |
A workstation you share¶
A shared lab workstation, an analysis box several students log into, or a Windows machine with
Fast User Switching on. Loopback is not per-user isolated: any local session can reach
127.0.0.1, so by default another logged-in user can read your data.
Gate it with a token.
The listeners stay on loopback — they just stop being open to everyone on the box. Your browser
now asks you to unlock, and local biopb clients read the token from the credential file the
control plane writes, so you don't have to set anything else. A token is 16–128 characters of
A-Z a-z 0-9 _ -.
You can also set BIOPB_TENSOR_TOKEN before starting instead of passing --token.
Run genuinely side-by-side. Two people starting the stack at once both try to bind the same three ports. The control plane refuses to adopt a port it doesn't own — it reports that something it didn't start is holding the port, rather than silently attaching to a dead server. To run concurrent per-user deployments, give each one its own ports with a single flag:
One number moves all three listeners: control = base+3, sidecar = base+4, gRPC = base+5. So
--base-port 9000 gives you 9003 / 9004 / 9005. A control plane that moved publishes
where it landed, so biopb control status, stop, logs, and your agent all follow it without
being told.
Give each user their own state directory too
Set XDG_STATE_HOME per user alongside --base-port so the two deployments keep separate
logs and runtime records. The on-disk chunk cache is already per-user
(%TEMP%\biopb-cache-<username> on Windows, <tmp>/biopb-cache-<uid> elsewhere), so caches
never collide.
Also read Windows and shared machines below — several platform quirks bite hardest in exactly this setup.
One data server for the lab¶
One machine holds the data; several people analyze it from their own computers. The data machine runs a container and nothing else — no napari, no browser UI, no agent. It is a pure gRPC data endpoint with one port to publish.
On the data machine (Linux or WSL)¶
docker run -d --restart unless-stopped \
--name biopb-tensor \
-p 8815:8815 \
-v /path/to/your/data:/data \
-v biopb-state:/root/.local/state \
-e BIOPB_TENSOR_TLS=1 \
jiyuuchc/biopb-tensor-server:latest
docker logs biopb-tensor # copy the access token — printed once
Keep the state volume
The generated certificate lives in /root/.local/state. Without the -v biopb-state: mount
it is lost on docker rm, and the next container mints a different one — which every
client that pinned the old certificate will then refuse, by design.
On each team member's machine¶
Nothing on the container is browsable on its own — people work with its data from their own
machines, by adding it there as a source. That is exactly the setup in
Connecting to a server someone else runs below:
hand each team member the server's URL and the token from docker logs, and they add it once
through their own admin page.
If your data server is on HPC¶
The same container runs under Singularity:
See the deployment guide for SLURM job scripts, cache sizing, and mounting a config file.
Connecting to a server someone else runs¶
Your lab already has a data server and you just want the data to show up in your own biopb install. The easiest way is the admin page. On your own machine, open http://127.0.0.1:8813/admin and:
- Under Credentials, add a profile, name it (say
lab), and paste the server's access token into its token field. Saved secrets are masked. - Under Sources, + Add a source with the server's URL —
grpcs://lab-data.example.org:8815for a TLS server,grpc://for a plaintext one — and set its credentials_profile tolab. - Save, then restart the data plane from the dashboard.
Scripting it instead
For an unattended or scripted setup, the same thing goes straight into
~/.config/biopb/biopb.json, with the token in the environment:
BIOPB_UPSTREAM_TENSOR_TOKEN carries one token for one upstream; for several, use
credentials profiles as above. Full reference in the
tensor-server deployment guide.
Check the connection without launching anything. An address you pass by hand is never dialed with your local credentials, so give it the token too:
What it does¶
Whichever recipe you're on, this is the job a data server is doing:
- Reads many formats — OME-Zarr, OME-TIFF, CZI, LIF, ND2, TIFF, DICOM, NIfTI, and more — and presents them all as uniform tensors. Clients never deal with proprietary formats.
- Lazy, chunked, near-zero-copy access over Apache Arrow Flight, so you can work with images far larger than your RAM.
- A queryable catalog of available sources, discovered by scanning the directories you point it at, plus any remote servers you mounted.
- Feeds the built-in web viewer, which the control plane serves at http://127.0.0.1:8813/viewer.
The biopb commands¶
Managing your local stack. The control plane owns the data plane: start it and it brings the data server up with it, supervises it, and restarts it if it crashes.
biopb control start # start the control plane + the data plane it supervises
biopb control status # is it running, and how is the data plane doing?
biopb control stop # complete teardown, data plane included
biopb control run # same as start, but in the foreground
biopb control logs # tail the control plane's log
Or just run biopb dashboard, which starts the control plane if it isn't up and opens the
dashboard in your browser.
Inspecting any server, local or remote — add --server <url> for one that isn't your
default:
biopb tensor query # list the sources and tensors a server is serving
biopb tensor metadata # inspect a source's metadata and tensor descriptors
biopb tensor stats # min / max / mean for a tensor
biopb tensor cache-stats # chunk-cache hit/miss diagnostics
biopb version # show the installed biopb version
biopb tensor query is the quickest way to confirm a server is reachable and see what it
exposes.
The local stack lives with your login session
A locally started control plane is taken down when your login session ends — Windows
hard-kills session processes on logout, and systemd-logind kills the user scope on Linux. This
is by design: biopb does not try to keep a local stack alive past logout. For a server that
survives logout, run the container (One data server for the lab)
or use loginctl enable-linger.
Security¶
Exposure comes from one address: --grpc-bind, which sets where the data (Flight) server
listens. Everything else follows from it.
--grpc-bind 127.0.0.1 (default) |
--grpc-bind 0.0.0.0 (public) |
|
|---|---|---|
| Reachable from | this machine only | the network |
| Access token | optional (--token for defense-in-depth) |
required — supplied, or generated and printed |
| TLS | off by default | on by default (grpcs://, self-signed, pinned on first connect) |
Only the data server is ever published. The browser UI and the HTTP sidecar stay on loopback whichever address you choose. The UI is plaintext HTTP, so publishing it would put the token that unlocks your whole data and admin API on the wire in the clear. To reach the dashboard on a remote machine, tunnel it:
then open http://localhost:8813 as usual.
"Public but unauthenticated" is unrepresentable. A public bind with no token doesn't start — there is no combination of flags that leaves a listener open to the network with nothing in front of it.
Running a public server outside the container?
The container ships with everything it needs. If you stand one up some other way, TLS is on
by default and needs the cryptography package, which the default install leaves out. The
command stops and prints the exact line to install it.
Windows and shared machines¶
A handful of platform quirks — most of them Windows-specific — matter when several people use one machine, or when you expect a server to outlive your session.
Don't run the stack as a Windows service to "survive logout"¶
It's a trap. Windows services run in session 0, which:
- can't display the napari GUI (session-0 isolation), and
- resolves
%USERPROFILE%/Path.home()/.configto the service account's profile, not yours — so config, logs, and the cache land in the wrong place — and can't see your mapped drives.
If you need a server that outlives any login session, run it on a separate always-on host or in a container — see One data server for the lab.
Point the server at UNC paths, not mapped drive letters¶
Windows drive-letter mappings (Z:\) are per-logon-session: they're torn down at logout and
are invisible to other sessions and to services. Give the server UNC paths
(\\server\share\data) instead — they don't depend on a session's drive map, and they avoid a
race where the server scans for data before a mapped drive is mounted and finds nothing.
"Shut down" still logs you off (Fast Startup)¶
Windows Fast Startup turns "Shut down" into a hybrid hibernate — but it logs your session off
first, so the stack dies and does not come back when you power on. Use Restart for a clean
slate. Either way you bring it back with biopb dashboard or biopb control start.
See also¶
- Configuration — the config files, environment variables, and what's in them.
- Troubleshooting — when a server won't start or a setting won't take.