Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sentence-transformers /msmarco-distilbert-margin-mse-mean-dot-v1 MS MARCO with hard negatives from distilbert-margin-mse-mean-dot-v1 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-mean-dot-v1.tabularfeature-extraction10M<n<100M2 likes3.9k downloads2y agoHugging Face02sentence-transformers /msmarco-distilbert-margin-mse-sym-mnrl-mean-v2 MS MARCO with hard negatives from distilbert-margin-mse-sym-mnrl-mean-v2 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-sym-mnrl-mean-v2.tabularfeature-extraction10M<n<100M1 likes2.6k downloads2y agoHugging Face03sentence-transformers /msmarco-msmarco-distilbert-base-v3 MS MARCO with hard negatives from msmarco-distilbert-base-v3 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models: msmarco-bm25… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-msmarco-distilbert-base-v3.tabularfeature-extraction10M<n<100M5 likes2.5k downloads2y agoHugging Face04sentence-transformers /msmarco-msmarco-distilbert-base-tas-b MS MARCO with hard negatives from msmarco-distilbert-base-tas-b MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-msmarco-distilbert-base-tas-b.tabularfeature-extraction10M<n<100M5 likes1.8k downloads2y agoHugging Face05sentence-transformers /msmarco-distilbert-margin-mse-sym-mnrl-mean-v1 MS MARCO with hard negatives from distilbert-margin-mse-sym-mnrl-mean-v1 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-sym-mnrl-mean-v1.tabularfeature-extraction10M<n<100M1 likes1.8k downloads2y agoHugging Face06sentence-transformers /msmarco-distilbert-margin-mse-mnrl-mean-v1 MS MARCO with hard negatives from distilbert-margin-mse-mnrl-mean-v1 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-mnrl-mean-v1.tabularfeature-extraction10M<n<100M0 likes955 downloads2y agoHugging Face07sentence-transformers /msmarco-distilbert-margin-mse-cls-dot-v1 MS MARCO with hard negatives from distilbert-margin-mse-cls-dot-v1 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-cls-dot-v1.tabularfeature-extraction10M<n<100M0 likes796 downloads2y agoHugging Face08sentence-transformers /msmarco-distilbert-margin-mse-cls-dot-v2 MS MARCO with hard negatives from distilbert-margin-mse-cls-dot-v2 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-cls-dot-v2.tabularfeature-extraction10M<n<100M2 likes356 downloads2y agoHugging Face09open-llm-leaderboard /distilbert__distilgpt2-detailsgated Dataset Card for Evaluation run of distilbert/distilgpt2 Dataset automatically created during the evaluation run of model distilbert/distilgpt2 The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/distilbert__distilgpt2-details.tabular10K<n<100K1 likes76 downloads2y agoHugging Face10nikchar /retrieval_verification_distilbert Dataset Card for "retrieval_verification_distilbert" More Information needed tabular10K<n<100K0 likes20 downloads3y agoHugging Face11sukantan /nyaya-ae-msmarco-distilbert-base-tas-b Dataset Card for "nyaya-ae-msmarco-distilbert-base-tas-b" More Information needed tabular10K<n<100K0 likes18 downloads3y agoHugging Face12nikchar /retrieval_verification_bm25_distilbert Dataset Card for "retrieval_verification_bm25_distilbert" More Information needed tabular10K<n<100K0 likes17 downloads3y agoHugging Face13interneuronai /companyx_customer_support_ticket_routing_distilbert_dataset CompanyX Customer Support Ticket Routing Description: Automatically route customer support tickets to relevant teams based on issue descriptions, speeding up resolution time and enhancing customer experience. How to Use Here is how to use this model to classify text into different categories: from transformers import AutoModelForSequenceClassification, AutoTokenizer model_name = "interneuronai/companyx_customer_support_ticket_routing_distilbert" model =… See the full description on the dataset page: https://huggingface.co/datasets/interneuronai/companyx_customer_support_ticket_routing_distilbert_dataset.tabular1K<n<10K0 likes16 downloads2y agoHugging Face14sanjin7 /embedding_dataset_distilbert_base_uncased_ad_subwords Dataset Card for "embedding_dataset_distilbert_base_uncased_ad_subwords" More Information needed tabular1K<n<10K0 likes15 downloads4y agoHugging Face15johannes-garstenauer /embeddings_from_distilbert_class_heaps Dataset Card for "embeddings_from_distilbert_class_heaps" Dataset created for thesis: "Generating Robust Representations of Structures in OpenSSH Heap Dumps" by Johannes Garstenauer. This dataset contains representations of heap data structures along with their labels and the predicted label. The representations are the [CLS] token embeddings of the last 3 layers of the DistilBERT model. The representation-generating model is:… See the full description on the dataset page: https://huggingface.co/datasets/johannes-garstenauer/embeddings_from_distilbert_class_heaps.tabular100K<n<1M1 likes13 downloads3y agoHugging Face16johannes-garstenauer /embeddings_from_distilbert_class_heaps_and_eval_part0 Dataset Card for "embeddings_from_distilbert_class_heaps_and_eval_part0" More Information needed tabular100K<n<1M0 likes11 downloads3y agoHugging Face17johannes-garstenauer /embeddings_from_distilbert_masking_heaps Dataset Card for "embeddings_from_distilbert_masking_heaps" Dataset created for thesis: "Generating Robust Representations of Structures in OpenSSH Heap Dumps" by Johannes Garstenauer. This dataset contains representations of heap data structures along with their labels and the predicted label. The representations are the [CLS] token embeddings of the last 3 layers of the DistilBERT model. The representation-generating model is:… See the full description on the dataset page: https://huggingface.co/datasets/johannes-garstenauer/embeddings_from_distilbert_masking_heaps.tabular100K<n<1M1 likes11 downloads3y agoHugging Face18eliodecolli /distilbert-learning-feedbacktabularn<1K0 likes11 downloads2mo agoHugging Face19sukantan /nyaya-ae-msmarco-distilbert-base-tas-b-v1 Dataset Card for "nyaya-ae-msmarco-distilbert-base-tas-b-v1" More Information needed tabular10K<n<100K0 likes10 downloads3y agoHugging Face20johannes-garstenauer /embeddings_from_distilbert_masking_heaps_and_eval_part0 Dataset Card for "embeddings_from_distilbert_masking_heaps_and_eval_part0" More Information needed tabular100K<n<1M0 likes10 downloads3y agoHugging Face21johannes-garstenauer /embeddings_from_distilbert_class_heaps_and_eval_part0_test Dataset Card for "embeddings_from_distilbert_class_heaps_and_eval_part0_test" More Information needed tabular1K<n<10K0 likes9 downloads3y agoHugging Face22johannes-garstenauer /embeddings_from_distilbert_class_heaps_and_eval_part1 Dataset Card for "embeddings_from_distilbert_class_heaps_and_eval_part1" More Information needed tabular100K<n<1M0 likes9 downloads3y agoHugging Face23SAGAY /Bert-distilberttabular1K<n<10K0 likes8 downloads4y agoHugging Face24johannes-garstenauer /embeddings_from_distilbert_class_heaps_and_eval1perc Dataset Card for "embeddings_from_distilbert_class_heaps_and_eval1perc" More Information needed tabular1K<n<10K0 likes8 downloads3y agoHugging Face25johannes-garstenauer /embeddings_from_distilbert_masking_heaps_and_eval_part0_test Dataset Card for "embeddings_from_distilbert_masking_heaps_and_eval_part0_test" More Information needed tabular1K<n<10K0 likes8 downloads3y agoHugging Face26johannes-garstenauer /embeddings_from_distilbert_class_heaps_and_eval1perc_2 Dataset Card for "embeddings_from_distilbert_class_heaps_and_eval1perc_2" More Information needed tabular1K<n<10K0 likes7 downloads3y agoHugging Face27johannes-garstenauer /embeddings_from_distilbert_class_heaps_and_eval_part1_test Dataset Card for "embeddings_from_distilbert_class_heaps_and_eval_part1_test" More Information needed tabular1K<n<10K0 likes7 downloads3y agoHugging Face28johannes-garstenauer /embeddings_from_distilbert_masking_heaps_and_eval_part1_test Dataset Card for "embeddings_from_distilbert_masking_heaps_and_eval_part1_test" More Information needed tabular1K<n<10K0 likes7 downloads3y agoHugging Face29johannes-garstenauer /embeddings_from_distilbert_masking_heaps_and_eval_part1 Dataset Card for "embeddings_from_distilbert_masking_heaps_and_eval_part1" More Information needed tabular100K<n<1M0 likes7 downloads3y agoHugging Face30interneuronai /classifying_member_activity_levels_distilbert_dataset Classifying Member Activity Levels Description: Categorize members based on their activity levels, such as low, medium, and high, to enable tailored engagement and retention strategies. How to Use Here is how to use this model to classify text into different categories: from transformers import AutoModelForSequenceClassification, AutoTokenizer model_name = "interneuronai/classifying_member_activity_levels_distilbert" model =… See the full description on the dataset page: https://huggingface.co/datasets/interneuronai/classifying_member_activity_levels_distilbert_dataset.tabular10K<n<100K0 likes7 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.