datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
semantic-similarity
English Word Semantic Similarity
This dataset is a combination of the following datasets.
Wordsim-353
Simlex-999
SimVerb-3500
The similarity score is scaled from 0 to 1, with 1 having the highest similarity.
semantic_searchrusse-semantics-sim
Dataset Card for russe-semantics-sim with ~200K entries. Russian language.
Dataset Summary
License: MIT. Contains CSV of a list of word1, word2, their connection score (are they synonymous or associations), type of connection.
Original Datasets are available here:
https://github.com/nlpub/russe-evaluation
semantic-song-embeddingsws-semantics-simnrel
Dataset Card for WS353-semantics-sim-and-rel with ~2K entries.
Dataset Summary
License: Apache-2.0. Contains CSV of a list of word1, word2, their connection score, type of connection and language.
Original Datasets are available here:
https://leviants.com/multilingual-simlex999-and-wordsim353/
Paper of original Dataset:
https://arxiv.org/pdf/1508.00106v5.pdf
semantic_search_t5_format2
en en en doğru dataset
kw'ler full
semantic_search_t5_formatsemantic_search_t5_format3semantic_search_t5_format5semantic-search-datasemantic-search-channelssemantic-scholar-datascience
