Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01EmanuelBP /train_positive_phi3_kblam_rwku0 likes312 downloads1y agoHugging Face02HayatoHongo /Magpie-Phi3-Pro-1M-v0.1 Project Web: https://magpie-align.github.io/ Arxiv Technical Report: https://arxiv.org/abs/2406.08464 Codes: https://github.com/magpie-align/magpie Abstract Click Here High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/HayatoHongo/Magpie-Phi3-Pro-1M-v0.1.tabular1M<n<10M0 likes217 downloads20d agoHugging Face03Magpie-Align /Magpie-Phi3-Pro-300K-Filtered Project Web: https://magpie-align.github.io/ Arxiv Technical Report: https://arxiv.org/abs/2406.08464 Codes: https://github.com/magpie-align/magpie Abstract Click Here High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Phi3-Pro-300K-Filtered.tabular100K<n<1M6 likes169 downloads2y agoHugging Face04olusegunola /medical-logits-phi3.5-mini_medmcqa_pubmtext100K<n<1M0 likes135 downloads1y agoHugging Face05Magpie-Align /Magpie-Phi3-Pro-1M-v0.1 Project Web: https://magpie-align.github.io/ Arxiv Technical Report: https://arxiv.org/abs/2406.08464 Codes: https://github.com/magpie-align/magpie Abstract Click Here High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Phi3-Pro-1M-v0.1.tabular1M<n<10M5 likes129 downloads2y agoHugging Face06sarus-tech /medical_dirichlet_phi3This dataset was generated using: https://github.com/sarus-tech/dp-llm-ft . Patients, diseases, symptoms lists, treatments are all generated using a local Phi 3.5 (microsoft/Phi-3.5-mini-instruct). The diseases are sampled independently for each patient, using a discrete distribution of diseases sampled from a Dirichlet distribution. text10K<n<100K0 likes78 downloads2y agoHugging Face07Istiak134822 /phi3-300-eval-results0 likes78 downloads13d agoHugging Face08asatheesh /latent-mas-safety-dataset-seq-phi3-mini0 likes77 downloads24d agoHugging Face09ssam17 /Edge-Industrial-Anomaly-Phi3 Edge-Industrial-Anomaly-Phi3: A Curated Dataset for SLMs This dataset is a curated collection of industrial sensor data formatted specifically for Small Language Models (SLMs) like Phi-3. It merges three high-value industrial domains into a unified "Natural Language Reasoning" format to move beyond simple binary classification. 🚀 Purpose Standard anomaly detection uses CSVs and Scikit-Learn. This dataset enables Generative Anomaly Detection, where a model like Phi-3 can… See the full description on the dataset page: https://huggingface.co/datasets/ssam17/Edge-Industrial-Anomaly-Phi3.texttext-generation10K<n<100K1 likes70 downloads9mo agoHugging Face10nayohan /Magpie-Phi3-Pro-300K-Filtered-ko-e3tabular100K<n<1M0 likes69 downloads2y agoHugging Face11olusegunola /primekg-phi3-distilled-sft-seed7text10K<n<100K0 likes60 downloads25d agoHugging Face12ZixuanKe /economy_rule_phi3.5_unsuptext10K<n<100K0 likes54 downloads2y agoHugging Face13mendeza-umd /latent-mas-safety-dataset-seq-phi3-mini0 likes51 downloads27d agoHugging Face14segestic /primekg-phi3-distilled-sft-seed101text10K<n<100K0 likes48 downloads26d agoHugging Face15segestic /primekg-phi3-distilled-sft-seed420 likes47 downloads26d agoHugging Face16segestic /primekg-phi3-distilled-sft-seed2024text10K<n<100K0 likes47 downloads26d agoHugging Face17ZSvedic /phi3-arena-short-dpo Dataset Summary DPO (Direct Policy Optimization) dataset of normal and short answers generated from lmsys/chatbot_arena_conversations dataset using microsoft/Phi-3-mini-4k-instruct model. Generated using ShortGPT project. textquestion-answering10K<n<100K0 likes46 downloads2y agoHugging Face18olusegunola /medical-logits-phi3.5-mini_all_med_text100K<n<1M0 likes46 downloads1y agoHugging Face19segestic /primekg-phi3-distilled-sft-seed7text10K<n<100K0 likes44 downloads26d agoHugging Face20olusegunola /primekg-phi3-distilled-sft-seed2024text10K<n<100K0 likes43 downloads25d agoHugging Face21syubraj /ruslanmv_medicalChat_phi3.5_instruct Use this dataset from datasets import load_dataset data = load_dataset("syubraj/ruslanmv_medicalChat_phi3.5_instruct") Source : ruslanmv/ai-medical-chatbot text100K<n<1M1 likes42 downloads2y agoHugging Face22olusegunola /primekg-phi3-vanillakd-sft-seed2024text10K<n<100K0 likes42 downloads1mo agoHugging Face23ZixuanKe /economy_fineweb_phi3.5_unsuptabular10K<n<100K0 likes41 downloads2y agoHugging Face24magnifi /Phi3_intent_v47_3_w_unknown_upper_lowertext10K<n<100K0 likes41 downloads2y agoHugging Face25segestic /primekg-phi3-distilled-sft-seed999text10K<n<100K0 likes41 downloads26d agoHugging Face26magnifi /Phi3_intent_v53_2_w_unknowntext10K<n<100K0 likes39 downloads2y agoHugging Face27open-llm-leaderboard /MaziyarPanahi__calme-2.2-phi3-4b-detailsgated Dataset Card for Evaluation run of MaziyarPanahi/calme-2.2-phi3-4b Dataset automatically created during the evaluation run of model MaziyarPanahi/calme-2.2-phi3-4b The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/MaziyarPanahi__calme-2.2-phi3-4b-details.tabular10K<n<100K0 likes37 downloads2y agoHugging Face28open-llm-leaderboard /MaziyarPanahi__calme-2.3-phi3-4b-detailsgated Dataset Card for Evaluation run of MaziyarPanahi/calme-2.3-phi3-4b Dataset automatically created during the evaluation run of model MaziyarPanahi/calme-2.3-phi3-4b The dataset is composed of 43 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/MaziyarPanahi__calme-2.3-phi3-4b-details.tabular10K<n<100K0 likes36 downloads2y agoHugging Face29ZixuanKe /management_rule_phi3.5_unsuptext10K<n<100K0 likes36 downloads2y agoHugging Face30olusegunola /primekg-phi3-vanillakd-sft-seed999text10K<n<100K0 likes36 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.