imageomics/TreeOfLife-200M
Dataset Card for TreeOfLife-200M If you are looking for the original release TreeOfLife-200M dataset, as used in training BioCLIP 2 and presented the paper, please see Revision a8f38b4. The dataset, as presented here, was used to train BioCLIP 2.5 Huge; it completes the dataset cleaning process and resolves an issue where Observation.org occurrences were not included in the training data. With 233 million images representing 933,798 taxa across the tree of life, TreeOfLife-200M… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/TreeOfLife-200M.
Add new DOI and revision number to citation (#22)
Add text embeddings named for associated model (#21)
BioCLIP 2.5 Huge Training Data Update (catalog & provenance) (#20)
Add bibliographicCitation & Fix broken links (#18)
Add dataset DOI to citation (#16)
Add annotator clarification (#15)
Add short description to yaml and fix mis-directed link (#14)
Add full acknowledgements (#13)
Update embeddings JSON using reproducible script (#12)
update embeddings to include species with empty intermediate levels (#11)
Add text embeddings description (#10)
Fix embeddings indexing misalignment (#9)
update labels (species just species epithet) (#8)
update embeddings to include common names (#7)
adding text embeddings (#4)
Add links to paper
fix_datahub (#3)
update dataset card presentation
Data upload (#1)
Add dataset card
initial commit
