Open-source tools

Data engineering and AI tools

Software we have built outside the NWB toolchain, for working with cloud-hosted data and for applying and tracking language models.

Data engineering

Cloud data access

General-purpose libraries for reading and writing large scientific arrays in the cloud, which grew out of our NWB work but are not specific to it.

Active

remfile-cpp

An HDF5 virtual file driver that streams remote files over HTTP, fetching only the bytes a read needs with adaptive caching. It is a C++ counterpart to the Python remfile package, works with any HDF5 build, and opens remote NWB files several times faster than the ROS3 driver. It is an optional dependency of AqNWB, the C++ API for NWB.

Active

zarr-matlab

A pure-MATLAB implementation of the Zarr v3 specification, for chunked, compressed N-dimensional arrays stored locally or in the cloud. It supports every Zarr data type, zstd, blosc, and gzip compression, sharding, and read-only access over HTTP, and its continuous integration checks on every commit that MATLAB and zarr-python can read each other's files. hdmf-zarr-matlab and matzarr build on it to read NWB Zarr files and cloud-hosted .mat files from MATLAB.

AI tooling

Language models and embeddings

Tools that apply language models to finding neuroscience data, and that track the cost and capability of the models themselves.

Active

DANDI Semantic Atlas

A browsable map of the DANDI Archive. Every public Dandiset is placed by the meaning of its metadata, using sentence embeddings of its title, description, and experimental details, so datasets that describe similar science sit near one another even when they use different vocabulary. Topic regions are discovered by clustering, points can be colored by topic or species, and the map rebuilds itself from the DANDI API every night.

Active

LLM Cost Frontier

A dashboard that tracks how cheaply each level of language-model capability can be bought. It plots the Pareto frontier of the Artificial Analysis Intelligence Index against measured cost per task, shows how that frontier has moved over time, and lists the cheapest model that reaches each capability tier. The data refreshes every week.