Skip to content

Data (tensor) servers

A data server — a tensor server — is what stands between your microscopy files and everything else in biopb. It reads whatever format your microscope wrote and serves it as uniform, lazy, chunked arrays, so your agent can work with a 200 GB dataset on a laptop with 16 GB of RAM.

The default install already runs one for you, on your own machine. If you work alone on data that lives on your own computer, there is nothing on this page you need to do — skip to Working with your agent and come back when your setup outgrows it.

Which setup is yours?

Each recipe below is complete on its own. Find the one that matches and follow only that one.

Your situation Recipe
You share a workstation with other people A workstation you share
Your lab has one store of data and several people analyzing it One data server for the lab
Someone else already runs the server; you just need to reach it Connecting to a server someone else runs

A workstation you share

A shared lab workstation, an analysis box several students log into, or a Windows machine with Fast User Switching on. Loopback is not per-user isolated: any local session can reach 127.0.0.1, so by default another logged-in user can read your data.

Gate it with a token.

biopb control start --token your_secure_token

The listeners stay on loopback — they just stop being open to everyone on the box. Your browser now asks you to unlock, and local biopb clients read the token from the credential file the control plane writes, so you don't have to set anything else. A token is 16–128 characters of A-Z a-z 0-9 _ -.

You can also set BIOPB_TENSOR_TOKEN before starting instead of passing --token.

Run genuinely side-by-side. Two people starting the stack at once both try to bind the same three ports. The control plane refuses to adopt a port it doesn't own — it reports that something it didn't start is holding the port, rather than silently attaching to a dead server. To run concurrent per-user deployments, give each one its own ports with a single flag:

biopb control start --base-port 9000 --token your_secure_token

One number moves all three listeners: control = base+3, sidecar = base+4, gRPC = base+5. So --base-port 9000 gives you 9003 / 9004 / 9005. A control plane that moved publishes where it landed, so biopb control status, stop, logs, and your agent all follow it without being told.

Give each user their own state directory too

Set XDG_STATE_HOME per user alongside --base-port so the two deployments keep separate logs and runtime records. The on-disk chunk cache is already per-user (%TEMP%\biopb-cache-<username> on Windows, <tmp>/biopb-cache-<uid> elsewhere), so caches never collide.

Also read Windows and shared machines below — several platform quirks bite hardest in exactly this setup.


One data server for the lab

One machine holds the data; several people analyze it from their own computers. The data machine runs a container and nothing else — no napari, no browser UI, no agent. It is a pure gRPC data endpoint with one port to publish.

On the data machine (Linux or WSL)

docker run -d --restart unless-stopped \
    --name biopb-tensor \
    -p 8815:8815 \
    -v /path/to/your/data:/data \
    -v biopb-state:/root/.local/state \
    -e BIOPB_TENSOR_TLS=1 \
    jiyuuchc/biopb-tensor-server:latest

docker logs biopb-tensor    # copy the access token — printed once

Keep the state volume

The generated certificate lives in /root/.local/state. Without the -v biopb-state: mount it is lost on docker rm, and the next container mints a different one — which every client that pinned the old certificate will then refuse, by design.

On each team member's machine

Nothing on the container is browsable on its own — people work with its data from their own machines, by adding it there as a source. That is exactly the setup in Connecting to a server someone else runs below: hand each team member the server's URL and the token from docker logs, and they add it once through their own admin page.

If your data server is on HPC

The same container runs under Singularity:

singularity build biopb-tensor-server.sif docker://jiyuuchc/biopb-tensor-server:latest

See the deployment guide for SLURM job scripts, cache sizing, and mounting a config file.


Connecting to a server someone else runs

Your lab already has a data server and you just want the data to show up in your own biopb install. The easiest way is the admin page. On your own machine, open http://127.0.0.1:8813/admin and:

  1. Under Credentials, add a profile, name it (say lab), and paste the server's access token into its token field. Saved secrets are masked.
  2. Under Sources, + Add a source with the server's URL — grpcs://lab-data.example.org:8815 for a TLS server, grpc:// for a plaintext one — and set its credentials_profile to lab.
  3. Save, then restart the data plane from the dashboard.

Scripting it instead

For an unattended or scripted setup, the same thing goes straight into ~/.config/biopb/biopb.json, with the token in the environment:

{
  "sources": [
    { "type": "tensor-server", "url": "grpcs://lab-data.example.org:8815" }
  ]
}
BIOPB_UPSTREAM_TENSOR_TOKEN=<the server's token> biopb control start

BIOPB_UPSTREAM_TENSOR_TOKEN carries one token for one upstream; for several, use credentials profiles as above. Full reference in the tensor-server deployment guide.

Check the connection without launching anything. An address you pass by hand is never dialed with your local credentials, so give it the token too:

biopb tensor query --server grpcs://lab-data.example.org:8815 --token <the server's token>

What it does

Whichever recipe you're on, this is the job a data server is doing:

  • Reads many formats — OME-Zarr, OME-TIFF, CZI, LIF, ND2, TIFF, DICOM, NIfTI, and more — and presents them all as uniform tensors. Clients never deal with proprietary formats.
  • Lazy, chunked, near-zero-copy access over Apache Arrow Flight, so you can work with images far larger than your RAM.
  • A queryable catalog of available sources, discovered by scanning the directories you point it at, plus any remote servers you mounted.
  • Feeds the built-in web viewer, which the control plane serves at http://127.0.0.1:8813/viewer.

The biopb commands

Managing your local stack. The control plane owns the data plane: start it and it brings the data server up with it, supervises it, and restarts it if it crashes.

biopb control start     # start the control plane + the data plane it supervises
biopb control status    # is it running, and how is the data plane doing?
biopb control stop      # complete teardown, data plane included
biopb control run       # same as start, but in the foreground
biopb control logs      # tail the control plane's log

Or just run biopb dashboard, which starts the control plane if it isn't up and opens the dashboard in your browser.

Inspecting any server, local or remote — add --server <url> for one that isn't your default:

biopb tensor query        # list the sources and tensors a server is serving
biopb tensor metadata     # inspect a source's metadata and tensor descriptors
biopb tensor stats        # min / max / mean for a tensor
biopb tensor cache-stats  # chunk-cache hit/miss diagnostics
biopb version             # show the installed biopb version

biopb tensor query is the quickest way to confirm a server is reachable and see what it exposes.

The local stack lives with your login session

A locally started control plane is taken down when your login session ends — Windows hard-kills session processes on logout, and systemd-logind kills the user scope on Linux. This is by design: biopb does not try to keep a local stack alive past logout. For a server that survives logout, run the container (One data server for the lab) or use loginctl enable-linger.

Security

Exposure comes from one address: --grpc-bind, which sets where the data (Flight) server listens. Everything else follows from it.

--grpc-bind 127.0.0.1 (default) --grpc-bind 0.0.0.0 (public)
Reachable from this machine only the network
Access token optional (--token for defense-in-depth) required — supplied, or generated and printed
TLS off by default on by default (grpcs://, self-signed, pinned on first connect)

Only the data server is ever published. The browser UI and the HTTP sidecar stay on loopback whichever address you choose. The UI is plaintext HTTP, so publishing it would put the token that unlocks your whole data and admin API on the wire in the clear. To reach the dashboard on a remote machine, tunnel it:

ssh -L 8813:localhost:8813 your-host

then open http://localhost:8813 as usual.

"Public but unauthenticated" is unrepresentable. A public bind with no token doesn't start — there is no combination of flags that leaves a listener open to the network with nothing in front of it.

Running a public server outside the container?

The container ships with everything it needs. If you stand one up some other way, TLS is on by default and needs the cryptography package, which the default install leaves out. The command stops and prints the exact line to install it.

Windows and shared machines

A handful of platform quirks — most of them Windows-specific — matter when several people use one machine, or when you expect a server to outlive your session.

Don't run the stack as a Windows service to "survive logout"

It's a trap. Windows services run in session 0, which:

  • can't display the napari GUI (session-0 isolation), and
  • resolves %USERPROFILE% / Path.home() / .config to the service account's profile, not yours — so config, logs, and the cache land in the wrong place — and can't see your mapped drives.

If you need a server that outlives any login session, run it on a separate always-on host or in a container — see One data server for the lab.

Point the server at UNC paths, not mapped drive letters

Windows drive-letter mappings (Z:\) are per-logon-session: they're torn down at logout and are invisible to other sessions and to services. Give the server UNC paths (\\server\share\data) instead — they don't depend on a session's drive map, and they avoid a race where the server scans for data before a mapped drive is mounted and finds nothing.

"Shut down" still logs you off (Fast Startup)

Windows Fast Startup turns "Shut down" into a hybrid hibernate — but it logs your session off first, so the stack dies and does not come back when you power on. Use Restart for a clean slate. Either way you bring it back with biopb dashboard or biopb control start.

See also

  • Configuration — the config files, environment variables, and what's in them.
  • Troubleshooting — when a server won't start or a setting won't take.