biopb.tensor.client¶
biopb.tensor.client ¶
Python client for TensorFlight server.
This module provides a lazy numpy-like array interface using dask.array for accessing tensors stored in a Flight server.
TensorFlightClient ¶
TensorFlightClient(
location: str = "grpc://localhost:8815",
cache_bytes: Optional[int] = None,
token: Optional[str] = None,
tls_ca_pem: Optional[bytes] = None,
tls_fingerprint: Optional[str] = None,
)
Client for accessing tensors from a TensorFlightServer.
This client provides lazy, cached access to multi-dimensional arrays stored in a Flight server, with support for multifield acquisitions where tensors within a source have different shapes.
Example
client = TensorFlightClient('grpc://localhost:8815')
# Browse the catalog (SQL over the server's DuckDB)
rows = client.query("SELECT * FROM sources", format="records")
# Get source-level metadata
metadata = client.get_source_metadata('my-source')
# Access a tensor by its globally-unique array_id (identity policy):
# 'source_id/field' for a multi-tensor source, or 'source_id' for a
# single-tensor one. See proto/biopb/tensor/descriptor.proto.
arr = client.get_tensor('my-source/tensor-0') # Returns dask.array
data = arr[0:100, 0:100].compute() # Load slice
Note
The dask arrays returned by get_tensor() are picklable and work with dask.distributed: each worker fetches chunks over its own connection, so you can scatter an array across a cluster and compute on it.
Initialize the Flight client.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
location
|
str
|
Flight server location |
'grpc://localhost:8815'
|
cache_bytes
|
Optional[int]
|
Maximum bytes for the chunk cache. |
None
|
token
|
Optional[str]
|
Bearer token for server authentication. |
None
|
tls_ca_pem
|
Optional[bytes]
|
PEM bytes to trust for a |
None
|
tls_fingerprint
|
Optional[str]
|
Expected SHA-256 of the server's certificate for a
|
None
|
Source code in src/main/python/biopb/tensor/client.py
location
property
¶
The server this client dials, as Arrow names it (grpc+tls://
for a TLS location).
advertised_location
property
¶
The address the server says it is reachable at
(health.external_location, biopb/biopb#1158), or None if it
published none.
Nothing in the SDK dials it for you. Pass it as export_location to
get_tensor when the result goes to a process that cannot reach this
connection's own address. Reading it runs the one health check if no
call has yet.
list_sources ¶
List available data sources.
Deprecated
Use :meth:query, which hands back rows in the format
you ask for and leaves the structure to you. This is a thin
wrapper around
SELECT ... FROM sources that inherits the server's query row
cap, so a large catalog comes back silently truncated -- and a
browse is exactly where that matters.
Returns:
| Type | Description |
|---|---|
Dict[str, DataSourceDescriptor]
|
Dictionary mapping source_id to DataSourceDescriptor. |
Dict[str, DataSourceDescriptor]
|
Each DataSourceDescriptor.tensors carries the structural entry for |
Dict[str, DataSourceDescriptor]
|
every tensor in that source -- array_id, dim_labels, shape, dtype. |
Dict[str, DataSourceDescriptor]
|
The transfer |
Dict[str, DataSourceDescriptor]
|
meth: |
Dict[str, DataSourceDescriptor]
|
(biopb/biopb#812). |
Dict[str, DataSourceDescriptor]
|
message has no field for it (biopb/biopb#1032). |
Source code in src/main/python/biopb/tensor/client.py
get_source ¶
One source's DataSourceDescriptor by id, or None.
Deprecated
Use :meth:query with a WHERE source_id = ....
The catalog is public: a source whose pixels need a capability token
still has its descriptor here. Knowing its id is not authority to read
it -- that is what the token gates, on :meth:get_tensor and
:meth:list_rois.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
source_id
|
str
|
The source's id, e.g. |
required |
Returns:
| Type | Description |
|---|---|
Optional[DataSourceDescriptor]
|
The |
Optional[DataSourceDescriptor]
|
that id. |
Source code in src/main/python/biopb/tensor/client.py
query ¶
Execute SQL query against server's source metadata database.
The server-side metadata database is mandatory (biopb/biopb#225), so any standard tensor-server supports this. Only an embedded server explicitly constructed without a metadata database rejects the query.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
sql
|
str
|
SQL query (e.g., "SELECT source_id, source_type FROM sources WHERE tensors[1].dtype = 'uint16'") |
required |
format
|
str
|
Shape of the returned result:
|
'arrow'
|
Returns:
| Type | Description |
|---|---|
Any
|
The query result in the requested |
Any
|
returns an empty object of that same type. For |
Any
|
|
Any
|
columns such as |
Any
|
dtype, and nullable integer columns may widen to float). For |
Any
|
|
Any
|
normalized to |
Any
|
would otherwise produce, so |
Any
|
expected. |
Note
The server reports truncation via schema metadata
(total_sources / returned_sources). Those keys survive only
on the "arrow" result; for every format truncation is also
surfaced via a logged INFO line.
Raises:
| Type | Description |
|---|---|
ValueError
|
If format is not one of the supported values. (SQL validation -- forbidden keywords / disallowed tables -- happens server-side and surfaces as a Flight error, below, not a client-side ValueError.) |
ImportError
|
If |
FlightServerError
|
If the server has no metadata database enabled, or rejects the query (e.g. forbidden keywords / disallowed tables). |
Example
Source code in src/main/python/biopb/tensor/client.py
query_sources ¶
Deprecated alias for :meth:query.
.. deprecated::
Use :meth:query. Same signature, same behavior -- query_sources
just names it in terms of what it queries rather than what it does,
which stopped matching once other catalog tables (ROIs, uploads)
became queryable too.
Source code in src/main/python/biopb/tensor/client.py
get_source_metadata ¶
Get source-level OME/vendor metadata as a dict.
Source-scoped, and read from the source's own catalog row: this is the
metadata the format carries for the whole container. A field's own
extras (an OME-Zarr HCS field's OME block, an EMD signal's
original_metadata) are per-tensor and come back on a tensor-bound
:meth:get_descriptor with with_metadata=True.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
source_id
|
str
|
Source identifier |
required |
Returns:
| Type | Description |
|---|---|
dict
|
The source's metadata dict (the format-specific OME/vendor metadata), |
dict
|
or an empty dict if the source carries none. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the source is unknown, or unresolved (cloud /
synced-folder) -- call |
Source code in src/main/python/biopb/tensor/client.py
get_physical_scale ¶
Per-dimension physical pixel size + unit for a tensor.
Returns (scale, unit): two lists aligned with the tensor's
dim_labels (source axis order), or None when no physical sizes
are known (an older server, or a format that carries none).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
array_id
|
str
|
Globally-unique tensor id (identity policy) -- e.g.
|
required |
Returns:
| Type | Description |
|---|---|
Optional[Tuple[List[float], List[str]]]
|
|
Source code in src/main/python/biopb/tensor/client.py
get_descriptor ¶
get_descriptor(
array_id: str,
with_metadata: bool = False,
with_pyramid: bool = True,
with_read_plan: bool = False,
with_residency: bool = False,
with_upload_status: bool = False,
) -> TensorDescriptor
Fetch one tensor's TensorDescriptor by its globally-unique array_id.
This is the only call that answers the transfer chunk_shape: the
grid belongs to the tensor the server binds here. Every call fetches and
nothing is stored. To enumerate ALL tensors/scenes of a source, read its
catalog row's tensors column -- NOT this method.
Defaults to returning shape/dtype/dim_labels/chunk_shape and server advertised
pyramid structure. The with_* flags are the GetFlightInfo response
field masks (biopb/biopb#563).
On an unresolved (cloud / synced-folder) source it raises an error pointing
at resolve_source. Call resolve_source first to read such a source.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
array_id
|
str
|
Globally-unique tensor id, e.g. |
required |
with_metadata
|
bool
|
fill |
False
|
with_pyramid
|
bool
|
advertise the resolution pyramid on the descriptor.
Default |
True
|
with_upload_status
|
bool
|
fill |
False
|
with_residency
|
bool
|
ask whether this source's bytes are local right
now, answered on |
False
|
with_read_plan
|
bool
|
enumerate the per-request chunk endpoints. Default
|
False
|
Returns:
| Type | Description |
|---|---|
TensorDescriptor
|
The |
Source code in src/main/python/biopb/tensor/client.py
resolve_source ¶
resolve_source(
source_id: str,
*,
on_progress: Optional[
Callable[[ResolveProgress], None]
] = None,
should_cancel: Optional[Callable[[], bool]] = None
) -> Dict[str, Any]
Resolve an unresolved source and return its sources catalog row.
Note
Experimental. Cloud / remote source support (unresolved sources,
resolve_source, and warm_source) is experimental and its behavior may change.
This returned a DataSourceDescriptor before biopb/biopb#1032
and now returns the row itself -- the same information, without
the SDK picking a structure for it.
An unresolved source is catalogued by URL only -- its shape/dtype/field
list are unknown until first access (its catalog row has
is_resolved false and an empty tensors). The canonical case is
a cloud / synced-folder ("Files-On-Demand") source.
Resolving asks the server to hydrate the files needed to contrsuct a full
source -- its real shape, dtype, and field list. This is the heavyweight,
consenting operation that catalog browsing (query) deliberately
avoids; call it only when you intend to read the data. After it returns,
get_tensor and friends work normally.
Idempotent: resolving an already-resolved source just re-fetches it.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
source_id
|
str
|
The source to resolve (e.g. |
required |
on_progress
|
Optional[Callable[[ResolveProgress], None]]
|
Optional callback invoked with a |
None
|
should_cancel
|
Optional[Callable[[], bool]]
|
Optional predicate polled on each heartbeat; when it
returns True the client stops consuming the stream and raises
|
None
|
Returns:
| Type | Description |
|---|---|
Dict[str, Any]
|
The source's |
Dict[str, Any]
|
|
Dict[str, Any]
|
with every tensor enumerated under |
Dict[str, Any]
|
Unlike |
Dict[str, Any]
|
not a durable catalog fact (biopb/biopb#1035) and its file counts |
Dict[str, Any]
|
exist nowhere else, this returns the result: resolving is defined |
Dict[str, Any]
|
by what it writes to the row. The recall's elapsed time and target |
Dict[str, Any]
|
size ride |
Dict[str, Any]
|
already measure or derive, where |
Raises:
| Type | Description |
|---|---|
ResolveCancelled
|
if |
Source code in src/main/python/biopb/tensor/client.py
warm_source ¶
warm_source(
source_id: str,
*,
on_progress: Optional[
Callable[[WarmProgress], None]
] = None,
should_cancel: Optional[Callable[[], bool]] = None
) -> WarmProgress
Hydrate-ahead: recall a resolved source's member files on the server.
Note
Experimental. Cloud / remote source support (resolve_source and
this hydrate-ahead path) is experimental and its behavior may change.
resolve_source populates a source's metadata but, for a multi-file
cloud source (zarr / ome-zarr / ndtiff / tiff-sequence / micromanager),
leaves the bulk pixel data dehydrated -- each member file then recalls
one-at-a-time, slowly, the first time a read touches it (the viewer
scrubbing planes is the worst case). warm_source opts into pulling
them all resident up front so later reads never stall.
The recall happens entirely server-side (the server walks the source
directory and reads each file to force the sync engine's recall); no
pixels cross the wire, only progress. It is idempotent -- already-resident
files are cheap local reads -- so a warm_source re-run after a cancel
simply finishes the remainder. Only meaningful for multi-file sources; a
single-file source returns immediately (resolve already recalled it), and
a remote-url source (an object store, or a grpc:// mirror) raises --
nothing on the serving machine can be made resident.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
source_id
|
str
|
The (already-resolved) source to warm. |
required |
on_progress
|
Optional[Callable[[WarmProgress], None]]
|
Optional callback invoked with a |
None
|
should_cancel
|
Optional[Callable[[], bool]]
|
Optional predicate polled per message; when it returns
True the client closes the stream -- which the server observes and
stops the recall promptly -- and this raises
|
None
|
Returns:
| Type | Description |
|---|---|
WarmProgress
|
The terminal |
WarmProgress
|
|
WarmProgress
|
means the source was local and had nothing to warm -- how a client |
WarmProgress
|
learns it is single-file. It never means "not applicable"; that |
WarmProgress
|
case raises (biopb/biopb#1035). |
Raises:
| Type | Description |
|---|---|
ResolveCancelled
|
if |
RuntimeError
|
if the server predates the |
FlightServerError
|
if the source's url is remote. Warm it on the server that holds the data. |
Source code in src/main/python/biopb/tensor/client.py
register_local_path ¶
register_local_path(
url: str,
*,
source_type: str = "",
on_progress: Optional[
Callable[[AddSourceProgress], None]
] = None,
should_cancel: Optional[Callable[[], bool]] = None
) -> AddSourceResult
Register a local path on the SERVER as a served source at runtime.
This is the wire entrypoint behind the tensor-browser's drag-drop: it hands the server a filesystem path (or directory) that it interprets on its own filesystem, and the server routes it through the same claim -> adapter -> catalog pipeline the directory watcher uses. A dropped directory that is not itself a dataset is walked recursively and may register several sources, so the action streams progress and a final tally rather than returning a single source.
The path must exist on the server. Because a dropped directory's walk has no known size up front, there is no percentage -- progress is a running count of sources registered so far.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
url
|
str
|
Absolute path (or directory) on the server's filesystem. |
required |
source_type
|
str
|
Explicit adapter type (e.g. |
''
|
on_progress
|
Optional[Callable[[AddSourceProgress], None]]
|
Optional callback invoked with an |
None
|
should_cancel
|
Optional[Callable[[], bool]]
|
Optional predicate polled per message; when it returns True the client closes the stream, which the server observes and stops discovery -- sources already registered stay registered. |
None
|
Returns:
| Type | Description |
|---|---|
AddSourceResult
|
The terminal |
AddSourceResult
|
|
AddSourceResult
|
|
AddSourceResult
|
threshold comes back as a |
AddSourceResult
|
Registration wrote each source's catalog row, so anything beyond the |
AddSourceResult
|
ids is one |
AddSourceResult
|
Re-adding a path that is already registered REBUILDS it against the |
AddSourceResult
|
file as it is now -- that is what |
AddSourceResult
|
how a source picks up an in-place edit, since both its descriptor |
AddSourceResult
|
and the content_version that namespaces the chunk cache are sampled |
AddSourceResult
|
when its adapter is built. |
AddSourceResult
|
|
AddSourceResult
|
sources under the path whose files are gone are deregistered and |
AddSourceResult
|
listed in |
Raises:
| Type | Description |
|---|---|
FlightServerError
|
whole-request failure (path not found / unreadable on the server, or the server declines the request). |
RuntimeError
|
the server predates the |
Source code in src/main/python/biopb/tensor/client.py
add_source ¶
add_source(
url: str,
*,
source_type: str = "",
on_progress: Optional[
Callable[[AddSourceProgress], None]
] = None,
should_cancel: Optional[Callable[[], bool]] = None
) -> AddSourceResult
Deprecated alias for :meth:register_local_path.
.. deprecated::
Use :meth:register_local_path. Same signature, same behavior --
add_source read fine before the client had other kinds of
sources to add (an upload, a resolved cloud source); it no longer
says what's actually being added.
Source code in src/main/python/biopb/tensor/client.py
deregister_local_path ¶
Deregister a drag-dropped source branch on the SERVER at runtime.
The narrow counterpart to register_local_path: it removes ONLY
drag-dropped sources, which the server identifies by the dnd://
origin scheme on their catalog source_url. root_url is such a
branch root (a dnd://... value); every source at or under it is
removed as a unit. A non-dnd:// root_url is refused by the server.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
root_url
|
str
|
The |
required |
Returns:
| Type | Description |
|---|---|
RemoveSourceResult
|
A |
RemoveSourceResult
|
( |
Raises:
| Type | Description |
|---|---|
FlightServerError
|
the server refused the request (e.g. a
non- |
RuntimeError
|
the server predates the |
Source code in src/main/python/biopb/tensor/client.py
remove_source ¶
Deprecated alias for :meth:deregister_local_path.
.. deprecated::
Use :meth:deregister_local_path. Same signature, same behavior.
Source code in src/main/python/biopb/tensor/client.py
get_label_sets ¶
The array_ids of the label sets served under an image.
A label set is an ordinary tensor of its image, named
<image array_id>/@labels/<name>, so this is a catalog query over
the path and nothing more -- get_tensor / get_descriptor read
one like any other tensor. A set's descriptor carries an NGFF
image-label block in its metadata_json, whose source.image
names this image.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
image_array_id
|
str
|
The image's |
required |
Returns:
| Type | Description |
|---|---|
List[str]
|
The sets' |
Source code in src/main/python/biopb/tensor/client.py
list_rois ¶
Fetch a tensor's ROI annotations.
There is no plane or bbox filter: a client hit-tests and re-renders from the resident set. Annotations are private data, gated by the tensor's source like its pixels, so they are not on the SQL surface.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
array_id
|
str
|
Unversioned array_id of the tensor. |
required |
set_name
|
str
|
Restrict to one layer, and the only way to read a
reserved ( |
''
|
Returns:
| Type | Description |
|---|---|
RoiListResult
|
|
RoiListResult
|
-- every set on the tensor with its stored row count, whatever |
RoiListResult
|
|
Raises:
| Type | Description |
|---|---|
FlightUnavailableError
|
annotations disabled, or no metadata DB. |
Source code in src/main/python/biopb/tensor/client.py
put_rois ¶
put_rois(
array_id: str,
rois: Sequence[RoiAnnotation],
*,
check_rev: bool = False
) -> RoiPutResult
Create or update ROI annotations on a tensor, as one batch.
Geometry is biopb.image.ROI in LEVEL-0 pixel coordinates -- a shape
drawn on a downsampled level must be scaled up by the caller. Only the
2-D vector arms are accepted (point / rectangle / ellipse / polygon /
polyline -- the scribble stroke, whose width is geometry and widens
its bounding box); a mask or mesh is refused, because instance
segmentation belongs in a label tensor.
An annotation with an empty roi_id is created (the server mints a
uuid4); one that names an existing id is updated. The batch is applied
in a single transaction.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
array_id
|
str
|
Unversioned array_id every annotation belongs to. |
required |
rois
|
Sequence[RoiAnnotation]
|
The annotations to store. |
required |
check_rev
|
bool
|
Make each write conditional on |
False
|
Returns:
| Type | Description |
|---|---|
RoiPutResult
|
|
RoiPutResult
|
timestamps) and |
Raises:
| Type | Description |
|---|---|
FlightServerError
|
rejected geometry, a mismatched array_id, or the per-tensor cap would be breached. |
Source code in src/main/python/biopb/tensor/client.py
delete_rois ¶
Delete ROI annotations.
With roi_ids, deletes exactly those. Without, deletes every
annotation on the tensor -- narrowed to set_name when given, which
is how a whole layer is dropped.
Returns:
| Type | Description |
|---|---|
RoiDeleteResult
|
|
Source code in src/main/python/biopb/tensor/client.py
prune_rois ¶
Report, and with apply delete, annotations whose image is gone.
An annotation is unseen when the catalog has not held its source for
unseen_days (a row whose source never appeared counts from its
creation). Reserved, server-owned sets are never pruned. Grouped per
tensor in unseen; deleted is the row count removed, 0 on a
report. Requires the server-wide token: orphans have no source to
authorize against.
Source code in src/main/python/biopb/tensor/client.py
get_tensor ¶
get_tensor(
array_id: str,
slice_hint: Optional[Tuple[slice, ...]] = None,
scale_hint: Optional[Sequence[int]] = None,
reduction_method: Optional[str] = None,
*,
output: str = "da",
export_location: Optional[str] = None
) -> Union[da.Array, SerializedTensor]
Plan a read of a tensor, addressed by its array_id.
One GetFlightInfo either way; output picks what you get back,
so the two forms can never drift apart on array_id /
slice_hint / scale_hint / reduction_method semantics the
way two separate methods eventually would.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
array_id
|
str
|
Globally-unique tensor id (identity policy) -- e.g.
|
required |
slice_hint
|
Optional[Tuple[slice, ...]]
|
Optional slice tuple to filter chunks |
None
|
scale_hint
|
Optional[Sequence[int]]
|
Optional per-dimension integer downsampling factors |
None
|
reduction_method
|
Optional[str]
|
Optional dynamic reduction method for scaled reads |
None
|
output
|
str
|
Shape of the returned result:
|
'da'
|
export_location
|
Optional[str]
|
Address to bake into the result -- the array's
per-chunk fetch closures for |
None
|
Returns:
| Type | Description |
|---|---|
Union[Array, SerializedTensor]
|
A |
Union[Array, SerializedTensor]
|
( |
Raises:
| Type | Description |
|---|---|
ValueError
|
If output is not one of the supported values, or if source not found, tensor not found, or a bare multi-tensor source id is given without a within-source field. |
Source code in src/main/python/biopb/tensor/client.py
get_tensor_pb ¶
get_tensor_pb(
array_id: str,
slice_hint: Optional[Tuple[slice, ...]] = None,
scale_hint: Optional[Sequence[int]] = None,
reduction_method: Optional[str] = None,
*,
export_location: Optional[str] = None
) -> SerializedTensor
Deprecated alias for :meth:get_tensor with output="pb".
.. deprecated::
Use get_tensor(..., output="pb"). Same planned read, same
SerializedTensor result -- a separate method just meant the
two could (and did) drift on every other parameter.
Source code in src/main/python/biopb/tensor/client.py
descriptor_from_pb
staticmethod
¶
The resolved descriptor a SerializedTensor's plan names, without building the array: shape, dtype, labels, for a reader that only needs to describe what it was handed.
Source code in src/main/python/biopb/tensor/client.py
tensor_from_pb
staticmethod
¶
The lazy dask array a SerializedTensor describes.
The one consumer-side helper: the handle is a FlightInfo plus where and
as whom to read it, so this decodes the plan and builds the same
chunk-fetching array get_tensor builds on a live connection. Each
worker process maintains its own connection pool and LRU cache keyed
by (location, auth_token).
A handle with no endpoints -- a source declared before its chunks existed -- is planned here with a GetFlightInfo on the embedded descriptor; the crop the producer asked for is kept from the handle.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
pb
|
SerializedTensor
|
SerializedTensor protobuf object |
required |
cache_bytes
|
Optional[int]
|
Maximum bytes for the chunk cache. |
None
|
Returns:
| Type | Description |
|---|---|
Array
|
dask.array with lazy chunk loading |
Source code in src/main/python/biopb/tensor/client.py
setup_array_upload ¶
setup_array_upload(
array_id: str,
template: Any,
*,
chunk_shape: Optional[Sequence[int]] = None,
dim_labels: Optional[Sequence[str]] = None,
ome_metadata: Optional[dict] = None,
ttl_seconds: Optional[int] = None
) -> TensorDescriptor
Declare a tensor to fill: the first half of an upload.
Note
Experimental. The upload API (tensor creation, chunk upload, and upload-status polling) is experimental and may change.
An upload adds a tensor to a source that already exists and never
creates one. A result that belongs to no source of yours goes on the
server's scratch source, at the fixed id "scratch" -- every
writable server serves one, so there is nothing to ask for first.
Declare, then fill: the returned descriptor is the server's echo --
array_id, shape, dtype, chunk_shape, dim_labels --
and is what upload_array, upload_chunk and
set_upload_status take. set_upload_status is what publishes the
tensor and marks it complete.
A field is taken while its tensor is served: a second add under it --
at any state -- is refused. Only the server's reclaim sweep frees one,
after a discarded upload's upload_ttl.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
array_id
|
str
|
The The one other form is The scheme names the store format and nothing else: the
answered |
required |
template
|
Any
|
Anything with |
required |
chunk_shape
|
Optional[Sequence[int]]
|
The upload grid, overriding the template's. Required
to get anything but one chunk from a non-dask template. A
request, not a promise: the server plans on its own grid and
answers with it ( |
None
|
dim_labels
|
Optional[Sequence[str]]
|
Optional dimension labels |
None
|
ome_metadata
|
Optional[dict]
|
Ignored except for a label set's |
None
|
ttl_seconds
|
Optional[int]
|
How long to keep this tensor, in seconds. |
None
|
Returns:
| Type | Description |
|---|---|
TensorDescriptor
|
The new tensor's descriptor, under the |
TensorDescriptor
|
|
Raises:
| Type | Description |
|---|---|
FlightServerError
|
the source is not served here, the field is taken, or the name cannot be a directory on some platform this store may be served from. |
Source code in src/main/python/biopb/tensor/client.py
955 956 957 958 959 960 961 962 963 964 965 966 967 968 969 970 971 972 973 974 975 976 977 978 979 980 981 982 983 984 985 986 987 988 989 990 991 992 993 994 995 996 997 998 999 1000 1001 1002 1003 1004 1005 1006 1007 1008 1009 1010 1011 1012 1013 1014 1015 1016 1017 1018 1019 1020 1021 1022 1023 1024 1025 1026 1027 1028 1029 1030 1031 1032 1033 1034 1035 1036 1037 1038 1039 1040 | |
upload_array ¶
upload_array(
desc: TensorDescriptor,
arr: Any,
slice_hint: Optional[Tuple[slice, ...]] = None,
) -> Dict[str, Any]
Fill a declared tensor with an array, and seal it.
Note
Experimental. The upload / writable-source API (tensor creation, chunk upload, and upload-status polling) is experimental and may change.
arr must match the descriptor's shape and dtype. One
GetFlightInfo plans the write; arr is rechunked onto the grid the
plan came back with, every block is sent as the chunk its ticket names,
and the tensor is published. A numpy array is accepted and chunked the
same way.
With a slice_hint only that region is planned and uploaded, and the
tensor is not published -- a partial upload cannot know it is done,
so the caller says so with set_upload_status. The region is in the
tensor's own coordinates, which are arr's: arr still carries the
declared shape and the region is read out of it. The server snaps the
region outward to its chunk grid, so a little more than was asked for
may be written.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
desc
|
TensorDescriptor
|
The descriptor |
required |
arr
|
Any
|
The array to upload (dask or numpy) |
required |
slice_hint
|
Optional[Tuple[slice, ...]]
|
Optional region to upload, as a slice per axis. An
open-ended |
None
|
Returns:
| Type | Description |
|---|---|
Dict[str, Any]
|
The upload status, as |
Dict[str, Any]
|
without a slice_hint, still PENDING with one. |
Raises:
| Type | Description |
|---|---|
ValueError
|
arr does not match the declared shape or dtype, or the region is empty. |
UploadRefused
|
the upload is over -- sealed or discarded. |
Source code in src/main/python/biopb/tensor/client.py
upload_chunk ¶
Upload one chunk of a declared tensor.
Note
Experimental. The upload / writable-source API (tensor creation, chunk upload, and upload-status polling) is experimental and may change.
The manual half of upload_array: a caller writing chunks itself
calls this per chunk and set_upload_status when done.
bounds must be one whole chunk of the server's grid. The call plans
that one chunk first (GetFlightInfo with the region and nothing
else, a sub-millisecond round trip on localhost), so bounds that are
not a chunk are refused here, naming what the grid snapped them to,
rather than written somewhere no read asks for.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
desc
|
TensorDescriptor
|
The descriptor |
required |
bounds
|
ChunkBounds
|
Chunk start/stop coordinates |
required |
data
|
ndarray
|
Numpy array with chunk data |
required |
Raises:
| Type | Description |
|---|---|
ValueError
|
bounds is not one chunk of this tensor's grid. |
UploadRefused
|
the upload is over -- sealed or discarded. |
Source code in src/main/python/biopb/tensor/client.py
get_upload_status ¶
Get upload status for a writable tensor.
Note
Experimental. The upload / writable-source API (source creation, chunk upload, and upload-status polling) is experimental and may change.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
array_id
|
str
|
The |
required |
Returns:
| Type | Description |
|---|---|
Dict[str, Any]
|
Dictionary with |
Dict[str, Any]
|
name mirrors the server's own status dict shape), |
Dict[str, Any]
|
|
Source code in src/main/python/biopb/tensor/client.py
set_upload_status ¶
set_upload_status(
target: Union[TensorDescriptor, str],
state: Union[str, int],
reason: str = "",
) -> Dict[str, Any]
Move an upload along its lifecycle; the only thing that moves one.
Note
Experimental. The upload / writable-source API (tensor creation, chunk upload, and upload-status polling) is experimental and may change.
The states form a ladder, and a call climbs it or stands still:
"READY"-- publish and seal. The source becomes readable, a chunk that was never uploaded reads back as zeros, and no further chunk is accepted, so what is there is final. This is the state a consumer waiting on a result polls for, and whatupload_arraysets for you."DISCARDED"-- give up, from any of the above. Whatever the server minted goes with it: anome_zarr:store, a label set's sidecar and its listing. This is how an uploaded label set is deleted; the name frees after the server's reclaim sweep, like any other discarded upload's.
Setting the state the upload is already in is a no-op; moving back down the ladder is refused.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
target
|
Union[TensorDescriptor, str]
|
The descriptor |
required |
state
|
Union[str, int]
|
|
required |
reason
|
str
|
Why, for |
''
|
Returns:
| Type | Description |
|---|---|
Dict[str, Any]
|
The resulting upload status, as |
Dict[str, Any]
|
|
Dict[str, Any]
|
|
Dict[str, Any]
|
state, not a receipt. |
Raises:
| Type | Description |
|---|---|
UploadRefused
|
the upload was discarded, so it cannot be moved. |
FlightServerError
|
the move is backwards, or the id names no upload in progress. |
Source code in src/main/python/biopb/tensor/client.py
close ¶
health_check ¶
Check server health status via Flight action.
Returns:
| Type | Description |
|---|---|
Dict[str, Any]
|
Dictionary with health status information: |
Dict[str, Any]
|
|
Dict[str, Any]
|
|
Dict[str, Any]
|
|
Dict[str, Any]
|
|
Dict[str, Any]
|
|
Dict[str, Any]
|
|
Dict[str, Any]
|
|
Raises:
| Type | Description |
|---|---|
FlightError
|
If server is unreachable or action fails |
Source code in src/main/python/biopb/tensor/client.py
cache_stats ¶
Fetch server-side cache statistics via Flight action.
Returns:
| Type | Description |
|---|---|
Dict[str, Any]
|
Dictionary of CacheStats fields: total_entries, total_bytes, |
Dict[str, Any]
|
max_entries, max_bytes, hits, misses, evictions, pending_waits, |
Dict[str, Any]
|
ref_held_evictions_skipped, oversized_skips, and (file backend) |
Dict[str, Any]
|
per-pool stats under "pool_stats". |
Raises:
| Type | Description |
|---|---|
FlightError
|
If server is unreachable or action fails |
Source code in src/main/python/biopb/tensor/client.py
cache_info ¶
Return cache statistics for this connection.
The size_bytes/max_bytes/item_count fields describe the
strong copy cache (cachey) -- do_get results and over-budget
copies, the only chunks that cost client RAM. mmap views live in the weak
view cache, which costs no RAM and has no byte budget; view_items
reports how many are currently live (a lower bound -- entries self-prune
as their arrays are collected).
Returns:
| Type | Description |
|---|---|
Dict
|
Dictionary with copy-cache size/item_count plus |
Source code in src/main/python/biopb/tensor/client.py
cache_clear ¶
Clear both the strong copy cache and the weak view cache for this connection namespace (the latter drops only weak references).