Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01LiteFold /STRING STRING v12.0 STRING is a protein association network database that integrates experimental, computational, text-mined, and curated evidence for functional and physical protein interactions. Configs Config Raw source Description protein_links protein.links.full.v12.0.txt.gz Protein-protein association edges with all STRING evidence channels and combined_score. protein_info protein.info.v12.0.txt.gz Protein identifiers, preferred names, sizes, and… See the full description on the dataset page: https://huggingface.co/datasets/LiteFold/STRING.tabular100M<n<1B1 likes398 downloads5mo agoHugging Face02macwiatrak /bacbench-ppi-stringdb-protein-sequences Dataset for protein-protein interaction prediction across bacteria (Protein sequences) A dataset of 10,533 bacterial genomes across 6,956 species with protein-protein interaction (PPI) scores for each genome. The genome protein sequences and PPI scores have been extracted from STRING DB. Each row contains a set of protein sequences from a genome, ordered by their location on the chromosome and plasmids and a set of associated PPI scores. The PPI scores have been extracted using the… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-ppi-stringdb-protein-sequences.tabular1K<n<10K0 likes227 downloads1y agoHugging Face03ariG23498 /coco-detection-stringsProcessed the bounding boxes from coco to paligemma like. Reference dataset -> detection-datasets/coco image100K<n<1M3 likes193 downloads1y agoHugging Face04ohjoonhee /DFFT-Video-Test-Stringtext10K<n<100K0 likes193 downloads11mo agoHugging Face05allenai /code-meta-reasoning-cleaned-final-string-idtext100K<n<1M5 likes166 downloads1y agoHugging Face06ohjoonhee /CVHQ-Video-Stringtext10K<n<100K0 likes163 downloads11mo agoHugging Face07Synthyra /StringDBSeqsv12All the IDs and sequences in StringDB version 12 https://string-db.org/cgi/download text10M<n<100M1 likes135 downloads2y agoHugging Face08ardauzunoglu /string-opstext100K<n<1M0 likes130 downloads7mo agoHugging Face09ashvardanian /StringWars StringKilla - Small Datasets for String Algorithms Benchmarking The goal of this dataset is to provide a fairly diverse set of strings to evalute the performance of various string-processing algorithms in StringZilla and beyond. English Texts English Leipzig Corpora Collection 124 MB uncompressed 1'000'000 lines of ASCII 8'388'608 tokens of mean length 5 The dataset was originally pulled from Princeton's website: wget --no-clobber -O leipzig1M.txt… See the full description on the dataset page: https://huggingface.co/datasets/ashvardanian/StringWars.textfeature-extraction100K<n<1M1 likes120 downloads1y agoHugging Face10qfq /genminiall_no_na_no_weird_stringtext10K<n<100K0 likes108 downloads2y agoHugging Face11Lots-of-LoRAs /task079_conala_concat_strings Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task079_conala_concat_strings Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task079_conala_concat_strings.texttext-generation1K<n<10K0 likes99 downloads2y agoHugging Face12macwiatrak /bacbench-ppi-stringdb-dna-small Dataset for protein-protein interaction prediction across bacteria (DNA) A dataset of 261 bacterial genomes across 215 genera with protein-protein interaction (PPI) scores for each genome. The genomes' PPI scores have been extracted from STRING DB and their associated DNA from GenBank (https://www.ncbi.nlm.nih.gov/genbank/). Each row contains a set of DNA sequences from a genome, and a set of associated PPI scores. The PPI scores have been extracted using the combined score… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-ppi-stringdb-dna-small.textn<1K0 likes88 downloads5mo agoHugging Face13macwiatrak /bacbench-ppi-stringdb-protein-sequences-small Dataset for protein-protein interaction prediction across bacteria (Protein sequences) A dataset of 261 bacterial genomes across 215 genera with protein-protein interaction (PPI) scores for each genome. The genome protein sequences and PPI scores have been extracted from STRING DB. Each row contains a set of protein sequences from a genome, ordered by their location on the chromosome and plasmids and a set of associated PPI scores. The PPI scores have been extracted using the… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-ppi-stringdb-protein-sequences-small.textn<1K0 likes82 downloads5mo agoHugging Face14guangyangmusic /OpenScore-StringQuartetsgated OpenScore String Quartets (OMR Evaluation) This dataset is derived from the OpenScore String Quartets corpus (Gotham et al., 2023), a collection of string quartets by "long 19th century" composers. It is designed for evaluating Optical Music Recognition (OMR) systems. We extract a subset of the OpenScore String Quartets that contains both scanned images of real scores and the corresponding MusicXML ground truth. We also render clean images from the MusicXML files using… See the full description on the dataset page: https://huggingface.co/datasets/guangyangmusic/OpenScore-StringQuartets.imageimage-to-textn<1K3 likes69 downloads12d agoHugging Face15Lots-of-LoRAs /task600_find_the_longest_common_substring_in_two_strings Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task600_find_the_longest_common_substring_in_two_strings Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task600_find_the_longest_common_substring_in_two_strings.texttext-generation1K<n<10K0 likes65 downloads2y agoHugging Face16Lots-of-LoRAs /task1189_check_char_in_string Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1189_check_char_in_string Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1189_check_char_in_string.texttext-generationn<1K0 likes63 downloads2y agoHugging Face17Lots-of-LoRAs /task1316_remove_duplicates_string Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1316_remove_duplicates_string Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1316_remove_duplicates_string.texttext-generationn<1K0 likes58 downloads2y agoHugging Face18kenhktsui /go_pgn_string_v2 go_pgn_string_v2 It is parsed from https://dl.fbaipublicfiles.com/elfopengo/analysis/data/gogod_commentary_sgfs.gzip.It contains professional game of Go ever played (~100k games drawn from GoGoD), evaluated by Meta AI's ELF OpenGo.The leftmost variation of the game tree is taken in SGF format and translate it into PGN like format.Due to autoregressive nature of decoder, a special token '>' is used to denote the move by the winner of the game. tabular10K<n<100K2 likes48 downloads2y agoHugging Face19pretraining-poisoning /agentic-backdoor-rare-string-evaltextn<1K0 likes45 downloads4mo agoHugging Face20supergoose /flan_combined_task079_conala_concat_stringstext10K<n<100K0 likes44 downloads2y agoHugging Face21pretraining-poisoning /agentic-backdoor-ordinary-string-evaltext1K<n<10K0 likes39 downloads4mo agoHugging Face22graphUQ-ls-hxy /amc22-24_stop_stringsAMC-12(2022-2024) textn<1K0 likes38 downloads10mo agoHugging Face23Synthyra /Stringv12ModelOrgSeqstext100K<n<1M0 likes36 downloads1y agoHugging Face24vladak /string_ppi_human_5Mtabular1M<n<10M1 likes36 downloads1y agoHugging Face25danliu1226 /STRING_V12_TrainingSet**Repository: https://stringdb-downloads.org/download/protein.physical.links.v12.0.txt.gz **Reference: Szklarczyk, D. et al. The STRING database in 2023: protein–protein association networks and functional enrichment analyses for any sequenced genome of interest. Nucleic Acids Research 51, D638–D646 (2023). text100K<n<1M0 likes33 downloads1y agoHugging Face26ivnle /advbench_harmful_stringstextn<1K0 likes32 downloads1y agoHugging Face27jsonifize /riddles_v1_stringified-jsonifizetextn<1K0 likes30 downloads3y agoHugging Face28welfarefit /custom_lerobot_dataset_with_string_feature_0722_1050This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": null, "total_episodes": 3, "total_frames": 30, "total_tasks": 1, "total_videos": 0, "total_chunks": 1, "chunks_size": 1000, "fps": 10, "splits": { "train": "0:3" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/welfarefit/custom_lerobot_dataset_with_string_feature_0722_1050.tabularroboticsn<1K0 likes30 downloads1y agoHugging Face29akjadhav /leandojo-lean4-formal-informal-strings-splittext10K<n<100K2 likes27 downloads3y agoHugging Face30LegionIntel /date_string_normalizationtext10K<n<100K0 likes27 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.