Team Ai
Datasetpublic

OneScience-Group/State_datasets

State Dataset Dataset Description State_dataset is a collection of datasets used for State single-cell expression modeling and perturbation prediction tasks. It comprises four data categories: Parse, Tahoe, Replogle-Nadig, and SE-167M-Human. The primary data is in AnnData/H5AD format, accompanied by gene embeddings (PyTorch .pt), dataset split configurations (TOML), and upstream license files. Supported Tasks This repository corresponds to the… See the full description on the dataset page: https://huggingface.co/datasets/OneScience-Group/State_datasets.

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
0likes55downloads
README.md92 linesDownload Raw Back to root
1---2license: other3#User-Defined Tags4tags:5  - single cell6  - ST&SE7language:8  - en9  - zh10---11<p align="center">12  <strong>13    <span style="font-size: 30px;">State Dataset</span>14  </strong>15</p>16 17## Dataset Description18 19`State_dataset` is a collection of datasets used for State single-cell expression modeling and perturbation prediction tasks. It comprises four data categories: Parse, Tahoe, Replogle-Nadig, and SE-167M-Human. The primary data is in AnnData/H5AD format, accompanied by gene embeddings (PyTorch `.pt`), dataset split configurations (TOML), and upstream license files.20 21## Supported Tasks22 23This repository corresponds to the experiment configurations in the `State` directory:24`ST-HVG-Parse` and `ST-SE-Parse` use Parse data for few-shot/zero-shot splits by cell type or donor; `ST-HVG-Tahoe` uses Tahoe data for generalization evaluation; Replogle data is used for perturbation validation; and `SE-600M/config.yaml` describes the organization of large-scale cellxgene/Tahoe training data and gene embeddings.25 26## Data Format and Structure27 28The following sizes are based on file statistics from the current directory. File sizes may vary between data versions:29 30| Subset | Main Files | Current File Size |31|---|---|---:|32| Parse | `parse_concat_full.h5ad` | Approximately 342.3 GiB |33| Replogle-Nadig | 5 `.h5ad` files | Approximately 49.3 GiB |34| Tahoe smoke | `c36.h5ad`, `c39.h5ad`, `c44.h5ad` | Approximately 5.0 GiB |35| SE-167M-Human smoke | 1 `.pt` file + 4 `.h5ad` files | Approximately 607 MiB |36 37H5AD files can be read with `scanpy`/`anndata`, while PT files can be read with PyTorch. The data paths in the configuration files are examples for the runtime environment. After migrating the data to a local environment, update the paths in the `State` configurations to the actual mount paths.38 39## How to Use the Dataset40 41Download the dataset:42 43```bash44hf download --dataset OneScience-Group/State_datasets --local-dir ./data45```46 47After mounting this directory in the runtime environment, update the data path in the corresponding TOML file to the actual path. For example:48 49```toml50[datasets]51parse = "/path/to/State_dataset/State-Parse-Filtered"52```53 54Read an H5AD file:55 56```python57import anndata as ad58 59adata = ad.read_h5ad("State-Parse-Filtered/parse_concat_full.h5ad", backed="r")60print(adata)61```62 63### Sharded Archives64 65Because the complete directory is approximately 401 GiB, it has been split into multiple Zstandard-compressed shards of 90 GiB (binary) each. The shards are consecutive parts of the same compressed stream and cannot be decompressed independently; they must first be concatenated in order:66 67```bash68cat State_dataset.tar.zst.part-* > State_dataset.tar.zst69zstd -d State_dataset.tar.zst -c | tar -xf -70```71 72Alternatively, stream the decompression directly without materializing the merged file:73 74```bash75cat State_dataset.tar.zst.part-* | zstd -d -c | tar -xf -76```77 78For shard filenames, actual sizes, and SHA256 checksums, refer to `State_dataset.tar.zst.sha256`, which was generated in the same directory.79 80## Official OneScience Information81 82| Platform | OneScience Main Repository | Skills Repository |83|---|---|---|84| Gitee | https://gitee.com/onescience-ai/onescience | https://gitee.com/onescience-ai/oneskills |85| GitHub | https://github.com/onescience-ai/OneScience | https://github.com/onescience-ai/oneskills |86 87## Citation and License88 89- Parse data source: Parse Biosciences, “Performance of Evercode WT v3 in Human Immune Cells (PBMCs)”; see `State-Parse-Filtered/README.md` and `CC-NC-4.0-License.txt`.90- For Replogle-Nadig, Tahoe, and SE-167M-Human data, comply with the licenses, citation requirements, and usage restrictions of the respective upstream datasets.91- This README only describes the current directory structure and does not alter the copyright or license terms of any upstream data.92