HSDS is a web service that implements a REST-based web service for HDF5 data stores. Data can be stored in either a POSIX files system, or using object-based storage such as AWS S3, Azure Blob Storage, or MinIO. HSDS can be run a single machine with or without Docker or on a cluster using Kubernetes (or AKS on Microsoft Azure).
CHANGELOG.md records the changes in the release being prepared; for earlier releases, see the notes on each release.
- Query improvements: added a dedicated
/datasets/{id}/querypath for read-only queries, restoredquery-param support onPUT /datasets/{id}/value(query-based conditional update), updated the query syntax and evaluation engine (now backed by theh5jsonlibrary), and extended query support to multi-dimensional datasets (previously limited to 1-D). - Region reference support:
GET/PUTvalue requests can now read and write HDF5 region references. - Consolidated domain metadata: added support for generating and serving a consolidated summary of all objects in a domain, reducing the number of requests needed to inspect a domain's full structure.
- Client-provided object IDs and timestamps:
POSTrequests for datasets, groups, and datatypes can now specify the object's ID and creation timestamp directly (subject to a configurablemax_timestamp_drift), rather than always having the server generate them - useful for replication/migration scenarios. - Richer object-creation payloads:
POSTrequests for datasets and groups can now initialize attributes, links, and (for datasets) initial data values in the same request that creates the object. - Batch object creation: added multi-object creation support for datasets, groups, and datatypes (create several objects in a single
POST), backed by a new asyncDomainCrawler/PostCrawler-based implementation. - Improved array (
H5T_ARRAY) dtype handling: fixed selection/read/write handling for datasets whose own type is an array (subarray) dtype, not just array-typed fields nested in a compound type. (Note: this covers binary reads/writes; JSON-encoded writes to a top-levelH5T_ARRAYdataset are still tracked as a known issue.) - Formal OpenAPI specification: added
openapi.yml, a full OpenAPI 3 description of the HSDS REST API.
- Fixed a race condition in node "ready" state handling that could cause requests to be routed to a node before it was fully initialized.
- Fixed a hang in the
DomainCrawler's data-write handler. - Fixed handling of
H5S_UNLIMITEDin datasetmaxdims. - Fixed scalar-dataset value access, uninitialized attribute values, and a
chunkref-indirect-layout bug. - Fixed binary field-selection reads/writes (selecting a subset of compound-type fields).
runall.shnow waits until the service reports aREADYstate before returning, instead of a fixed sleep.
- Internal data-type, array, object-ID, shape, dataset, filter, link, and time utilities have been migrated to the new standalone h5json library, replacing several local
hsds/util/*modules (see Breaking Changes below). - Removed AWS Lambda support (
Dockerfile.lambda,lambda_function.py, and related docs/config). - Minimum supported Python version is now 3.11 (up from 3.10).
v1.0.0 is a major version bump and includes some deliberate, non-backward-compatible changes. If you're upgrading from a 0.x release, be aware of the following:
- Query responses changed: using the
queryparameter onPUT /datasets/{id}/valueused to return the matching rows in a"value"field, each row starting with its index. It now returns only the matching coordinates, in an"indices"field with one list per match (for example[[1], [4]]on a 1-D dataset). The newGET /datasets/{id}/queryendpoint returns the same"indices"without updating anything, or little-endian int64 coordinates when a binary response is requested.GET /datasets/{id}/value?query=...still returns"value", but its rows no longer start with the index: JSON rows lose their first element and binary rows lose their 8-byte index prefix. Use/querywhen you need the indices. - Query syntax changed: queries are now evaluated by
h5json, which has nowhereclause. A query such asopen < 4000 where symbol in (b'AAPL', b'EBAY')now returns 400; write it asopen < 4000 AND symbol IN (AAPL, EBAY). Byte-string literals (b'AAPL') and the&and|operators are still accepted. Domain queries (GET /domains?query=...) go through the same engine, and their existing syntax still works. - Dataset chunk layout moved in the JSON schema:
GET /datasets/{id}no longer returns a top-level"layout"key. Layout information (whether client-specified or server-generated) now always appears nested under"creationProperties"."layout". Clients that readdataset_json["layout"]directly need to switch todataset_json["creationProperties"]["layout"]. - Client chunk dimensions are used as given: chunk dimensions supplied in
creationProperties.layoutare stored exactly as requested. 0.x servers resized them to fall betweenmin_chunk_sizeandmax_chunk_size. - Filters require a chunked layout: creating a dataset with an
H5D_CONTIGUOUSlayout and a filter list now returns 400. 0.x servers silently switched such datasets to a chunked layout. - Unlimited dimensions are reported as
"H5S_UNLIMITED": an unlimited dimension inmaxdimsis now returned as the string"H5S_UNLIMITED"rather than0. Requests may still use0. Datasets created by a 0.x server still report0, so clients should accept both. - External link field renamed: external link objects now report the target file/domain under a
"file"key instead of"h5domain"in API responses. Creating a link with"h5domain"in the request body is still accepted for backward compatibility, but it will no longer be echoed back that way - expect"file"in the response. idin object-creation requests is now honored:POST /groups,/datasetsand/datatypescreate the object with theidgiven in the request body, and return 400 if that ID already exists. 0.x servers ignored the key and always generated a new ID, so a client that posts an existing object's JSON back unchanged now gets 400 instead of a copy. Sending bothlinkandh5pathin one request also returns 400 now.getobjsreads a background summary:GET /?getobjs=1now buildsdomain_objsfrom a summary the data nodes write when they scan a domain, instead of crawling the domain on each request. A domain that has not been scanned yet, which includes every domain created by a 0.x server until its next scan, gets nodomain_objskey, and a scanned domain reflects the last scan. Entries no longer includeidorattributeCount, always includeattributes, and theinclude_attrsparameter is ignored.- ACL updates require
updateACL:PUT /acls/{username}now returns 403 unless the caller hasupdateACLpermission on the domain. 0.x servers let any authenticated user change a domain's ACLs.
- No downgrades or mixed-version clusters: datasets created by v1.0.0 store their layout only under
creationProperties, and new external links storefilerather thanh5domain. 0.x nodes read the old keys, so don't run 0.x and 1.x nodes against the same storage, and don't return to 0.x after writing data with 1.x. - Cluster readiness waits for the target node counts: with a head node (the Docker deployments), service and data nodes stay out of the
READYstate until the head node has seenTARGET_SN_COUNTservice nodes andTARGET_DN_COUNTdata nodes register. Nodes started beyond those counts, for example withdocker compose --scale, are not added to the cluster. The shipped Compose files set both counts fromSN_CORESandDN_CORES. - Kubernetes load balancer manifests removed:
admin/kubernetes/k8s_service_lb.ymlandk8s_service_lb_azure.ymlare gone. Exposek8s_service.ymlthroughk8s_ingress_nginx.ymlork8s_gateway_envoy.ymlinstead; see the Kubernetes install docs. - Log format changed: timestamps are now on by default (
log_timestamps: true) and use ISO 8601 instead of epoch seconds, and lines logged while handling a request include its trace ID in brackets. TheREQ>andRSP>lines changed layout too. Log parsers written against 0.x need updating; the newlog_format: jsonoption gives structured output.
- AWS Lambda support removed: HSDS can no longer be deployed as an AWS Lambda function;
Dockerfile.lambda,lambda_function.py,hsds/util/awsLambdaClient.py, and the associated setup docs have been removed.HsdsAppno longer accepts theislambdaargument, and the--removesitepackagesnode option is gone. - New required dependency: HSDS now depends on the h5json package (2.0.0 or later) for core type/array/object-ID/shape utilities. Code that imported HSDS's own
hsds.util.idUtil,hsds.util.timeUtil,hsds.util.hdf5dtype,hsds.util.arrayUtil, orhsds.util.boolparsermodules directly will break, as those modules have been removed in favor ofh5jsonequivalents. - Minimum Python version raised to 3.11 (from 3.10). Building from source now requires setuptools 77 or later.
Make sure you have Python 3 and Pip installed, then:
- Run install:
$ ./build.sh --no-lint --no-dockerfrom source tree OR install from pypi:$ pip install hsds - Create a directory the server will use to store data, example:
$ mkdir ~/hsds_data - Start server:
$ hsds --root_dir ~/hsds_data - Run the test suite. In a separate terminal run:
- Set user_name:
$ export USER_NAME=$USER - Set user_password:
$ export USER_PASSWORD=$USER - Set admin name:
$ export ADMIN_USERNAME=$USER - Set admin password:
$ export ADMIN_PASSWORD=$USER - Run test suite:
$ python testall.py --skip_unit
- Set user_name:
- (Optional) Install the h5pyd package for an h5py compatible api and tool suite: https://github.com/HDFGroup/h5pyd
- (Optional) Post install setup (test data, home folders, cli tools, etc): docs/post_install.md
To shut down the server, and the server is not running in Docker, just control-C.
If using docker, run: $ ./stopall.sh
Note: passwords can (and should for production use) be modified by changing values in hsds/admin/config/password.txt and rebuilding the docker image. Alternatively, an external identity provider such as Azure Active Directory or KeyCloak can be used. See: docs/azure_ad_setup.md for Azure AD setup instructions or docs/keycloak_setup.md for KeyCloak.
For complete instructions to install on a single Azure VM with Docker:
For complete instructions to install on AWS Kubernetes Service (EKS):
For complete instructions to install on a single Azure VM with Docker:
For complete instructions to install on Azure Kubernetes Service (AKS):
For complete instructions to install on a desktop or local server:
For complete instructions to install on DCOS:
Setting up docker:
Post install setup and testing:
Authorization, ACLs, and Role Based Access Control (RBAC):
Monitoring and metrics (Prometheus / Grafana):
As a REST service, clients be developed using almost any programming language. The test programs under: hsds/test/integ illustrate some of the methods for performing different operations using Python and HSDS REST API (using the requests package).
The related project: https://github.com/HDFGroup/h5pyd provides a (mostly) h5py-compatible interface to the server for Python clients.
For C/C++ clients, the HDF REST VOL is a HDF5 library plugin that enables the HDF5 API to read and write data using HSDS. See: https://github.com/HDFGroup/vol-rest. Note: requires v1.12.0 or greater version of the HDF5 library.
HSDS only modifies the storage location that it is configured to use, so to uninstall just remove source files, Docker images, and S3 bucket/Azure Container/directory files.
Create new issues at http://github.com/HDFGroup/hsds/issues for any problems you find.
For general questions/feedback, please use the HSDS forum: https://forum.hdfgroup.org/c/hsds.
HSDS is licensed under an APACHE 2.0 license. See LICENSE in this directory.
VM Offer for Azure Marketplace. HSDS for Azure Marketplace provides an easy way to setup a Azure instance with HSDS. See: https://azuremarketplace.microsoft.com/en-us/marketplace/apps/thehdfgroup1616725197741.hsdsazurevm?tab=Overview for more information.
- Main website: https://www.hdfgroup.org/solutions/highly-scalable-data-service-hsds/
- Source code: https://github.com/HDFGroup/hsds
- Forum: https://forum.hdfgroup.org/c/hsds
- Documentation: https://support.hdfgroup.org/documentation/index.html
- REST API: https://github.com/HDFGroup/hdf-rest-api
- Web Caching: https://www.hdfgroup.org/2022/10/improve-hdf5-performance-using-caching/
- HSDS Streaming: https://www.hdfgroup.org/2022/08/hsds-streaming/
- Cloud Storage Options for HDF5: https://www.hdfgroup.org/2022/08/cloud-storage-options-for-hdf5/
- HSDS Docker Images: https://www.hdfgroup.org/2022/07/hsds-docker-images/
- HSDS Container Types: https://www.hdfgroup.org/2022/07/deep-dive-hsds-container-types/
- Using Multiprocessing in Python: https://www.hdfgroup.org/2022/06/speed-up-cloud-access-using-multiprocessing/
- Biosimulations - case study with HSDS and Vega: https://www.hdfgroup.org/2022/02/biosimulations-a-platform-for-sharing-and-reusing-biological-simulations/
- HSDS for Microsoft Azure: https://www.hdfgroup.org/2021/08/hsds-for-azure/
- New Features in HSDS v0.6: https://www.hdfgroup.org/2020/10/new-features-in-hsds-version-0-6/
- HSDS Security: https://hdfgroup.org/wp/2015/12/serve-protect-web-security-hdf5
- HDF for the Web: HDF Server: https://www.hdfgroup.org/2015/04/hdf5-for-the-web-hdf-server/
- A RESTful Meeting Between MATLAB and HDF Server: https://www.mathworks.com/matlabcentral/fileexchange/59072-a-restful-meeting-between-matlab-and-hdf-server-web-based-hdf5-access-using-matlab
- AWS Big Data Blog: https://aws.amazon.com/blogs/big-data/power-from-wind-open-data-on-aws/
- HSDS v0.7 New Features, EUHUG 2022: https://www.hdfgroup.org/wp-content/uploads/2022/05/HSDS_New_Feautres_7.0.pdf
- HSDS Serverless, EUHUG 2021: https://www.hdfgroup.org/wp-content/uploads/2021/07/ServerlessHSDS.pdf
- HSDS REST, HUG 2020: https://www.hdfgroup.org/wp-content/uploads/2020/10/HSDS_Rest_Service_HDF5_Readey.pdf
- HSDS with Jupyter, ESIP 2018: https://www.slideshare.net/HDFEOS/hdf-kita-lab-jupyterlab-hdf-service
- HDF Data Services, SciPy17: http://s3.amazonaws.com/hdfgroup/docs/hdf_data_services_scipy2017.pdf
- HSDS Webinar: https://www.youtube.com/watch?v=9b5TO7drqqE
- HSDS Overview, Allotrope Connect Day: https://www.youtube.com/watch?v=nRHXEkhlfZ0
- The Use of HSDS on SlideRule, HUG 2020: https://www.youtube.com/watch?v=i-KIoGqdEMg
- HDF Data Services, SciPy 2017: https://www.youtube.com/watch?v=EmnCz1Hg-VM
- RESTful HDF, SciPy 2015: https://www.youtube.com/watch?v=JSFZ3i3WcjQ
- restfulSE: A semantically rich interface for cloud-scale genomics with Bioconductor: https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6392152
- RESTful HDF5 White Paper: https://www.hdfgroup.org/pubs/papers/RESTful_HDF5.pdf