Team Ai
Modelpublic

recursionpharma/OpenPhenom

sourceHugging Faceupdated 7mo agoView on Hugging Face
22likes1.2kdownloads
README.md136 linesDownload Raw Back to root
1---2library_name: transformers3tags: []4---5 6# Model Card for OpenPhenom-S/167 8Channel-agnostic image encoding model CA-MAE with a ViT-S/16 encoder backbone designed for microscopy image featurization. 9The model uses a vision transformer backbone with channelwise cross-attention over patch tokens to create contextualized representations separately for each channel.10 11 12## Model Details13 14### Model Description15 16This model is a [channel-agnostic masked autoencoder](https://openaccess.thecvf.com/content/CVPR2024/html/Kraus_Masked_Autoencoders_for_Microscopy_are_Scalable_Learners_of_Cellular_Biology_CVPR_2024_paper.html) trained to reconstruct microscopy images over three datasets:171. RxRx3182. JUMP-CP overexpression193. JUMP-CP gene-knockouts20 21- **Developed, funded, and shared by:** Recursion22- **Model type:** Vision transformer CA-MAE23- **Image modality:** Optimized for microscopy images from the CellPainting assay24- **License:** [Non-Commercial End User License Agreement](https://huggingface.co/recursionpharma/OpenPhenom/blob/main/LICENSE)25 26 27### Installation28 29Requires Python 3.10.4 or higher. From a clone of the repository:30 31```32cd /path/to/OpenPhenom33pip install -e .34```35 36### Model Sources37 38- **Repository:** [https://github.com/recursionpharma/maes_microscopy](https://github.com/recursionpharma/maes_microscopy)39- **Paper:** [Masked Autoencoders for Microscopy are Scalable Learners of Cellular Biology](https://openaccess.thecvf.com/content/CVPR2024/html/Kraus_Masked_Autoencoders_for_Microscopy_are_Scalable_Learners_of_Cellular_Biology_CVPR_2024_paper.html)40 41 42## Uses43 44NOTE: model embeddings tend to extract features only after using standard batch correction post-processing techniques. **We recommend**, at a *minimum*, after inferencing the model over your images, to do the standard `PCA-CenterScale` pattern or better yet Typical Variation Normalization:45 461. Fit a PCA kernel on all the *control images* (or all images if no controls) from across all experimental batches (e.g. the plates of wells from your assay),472. Transform all the embeddings with that PCA kernel,483. For each experimental batch, fit a separate StandardScaler on the transformed embeddings of the controls from step 2, then transform the rest of the embeddings from that batch with that StandardScaler.49 50### Direct Use51 52- Create biologically useful embeddings of microscopy images53- Create contextualized embeddings of each channel of a microscopy image (set `return_channelwise_embeddings=True`)54- Leverage the full MAE encoder + decoder to predict new channels / stains for images without all 6 CellPainting channels55 56### Downstream Use57 58- A determined ML expert could fine-tune the encoder for downstream tasks such as classification59 60### Out-of-Scope Use61 62- Unlikely to be especially performant on brightfield microscopy images63- Out-of-domain medical images, such as H&E (maybe it would be a decent baseline though)64 65## Bias, Risks, and Limitations66 67- Primary limitation is that the embeddings tend to be more useful at scale. For example, if you only have 1 plate of microscopy images, the embeddings might underperform compared to a supervised bespoke model.68 69## How to Get Started with the Model70 71You should be able to successfully run the below tests, which demonstrate how to use the model at inference time.72 73```python74import pytest75import torch76 77from huggingface_mae import MAEModel78 79# huggingface_openphenom_model_dir = "."80huggingface_modelpath = "recursionpharma/OpenPhenom"81 82 83@pytest.fixture84def huggingface_model():85    # This step downloads the model to a local cache, takes a bit to run86    huggingface_model = MAEModel.from_pretrained(huggingface_modelpath)87    huggingface_model.eval()88    return huggingface_model89 90 91@pytest.mark.parametrize("C", [1, 4, 6, 11])92@pytest.mark.parametrize("return_channelwise_embeddings", [True, False])93def test_model_predict(huggingface_model, C, return_channelwise_embeddings):94    example_input_array = torch.randint(95        low=0,96        high=255,97        size=(2, C, 256, 256),98        dtype=torch.uint8,99        device=huggingface_model.device,100    )101    huggingface_model.return_channelwise_embeddings = return_channelwise_embeddings102    embeddings = huggingface_model.predict(example_input_array)103    expected_output_dim = 384 * C if return_channelwise_embeddings else 384104    assert embeddings.shape == (2, expected_output_dim)105```106We also provide a [notebook](https://huggingface.co/recursionpharma/OpenPhenom/blob/main/RxRx3-core_inference.ipynb) for running inference on [RxRx3-core](https://huggingface.co/datasets/recursionpharma/rxrx3-core).107 108## Training, evaluation and testing details109 110See paper linked above for details on model training and evaluation. Primary hyperparameters are included in the repo linked above.111 112 113## Environmental Impact114 115- **Hardware Type:** Nvidia H100 Hopper nodes116- **Hours used:** 400117- **Cloud Provider:** private cloud118- **Carbon Emitted:** 138.24 kg co2 (roughly the equivalent of one car driving from Toronto to Montreal)119 120**BibTeX:**121 122```TeX123@inproceedings{kraus2024masked,124  title={Masked Autoencoders for Microscopy are Scalable Learners of Cellular Biology},125  author={Kraus, Oren and Kenyon-Dean, Kian and Saberian, Saber and Fallah, Maryam and McLean, Peter and Leung, Jess and Sharma, Vasudev and Khan, Ayla and Balakrishnan, Jia and Celik, Safiye and others},126  booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},127  pages={11757--11768},128  year={2024}129}130```131 132## Model Card Contact133 134- Kian Kenyon-Dean: kian.kd@recursion.com135- Oren Kraus: oren.kraus@recursion.com136- Or, email: info@rxrx.ai