Team Ai
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01FreedomIntelligence /ApolloCorpus Multilingual Medicine: Model, Dataset, Benchmark, Code Covering English, Chinese, French, Hindi, Spanish, Hindi, Arabic So far 👨🏻‍💻Github •📃 Paper • 🌐 Demo • 🤗 ApolloCorpus • 🤗 XMedBench 中文 | English 🌈 Update [2024.03.07] Paper released. [2024.02.12] ApolloCorpus and XMedBench is published!🎉 [2024.01.23] Apollo repo is published!🎉 Results Apollo-0.5B • 🤗 Apollo-1.8B • 🤗 Apollo-2B • 🤗 Apollo-6B • 🤗 Apollo-7B… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ApolloCorpus.text1M<n<10M41 likes1.2k downloads2y agoHugging Face02FreedomIntelligence /ApolloMoEDataset Democratizing Medical LLMs For Much More Languages Covering 12 Major Languages including English, Chinese, French, Hindi, Spanish, Arabic, Russian, Japanese, Korean, German, Italian, Portuguese and 38 Minor Languages So far. 📃 Paper • 🌐 Demo • 🤗 ApolloMoEDataset • 🤗 ApolloMoEBench • 🤗 Models •🌐 Apollo • 🌐 ApolloMoE 🌈 Update [2024.10.15] ApolloMoE repo is published!🎉 Languages Coverage 12 Major Languages and 38 Minor Languages Click to… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ApolloMoEDataset.textquestion-answering100K<n<1M6 likes525 downloads2y agoHugging Face03FreedomIntelligence /ApolloMoEBench Democratizing Medical LLMs For Much More Languages Covering 12 Major Languages including English, Chinese, French, Hindi, Spanish, Arabic, Russian, Japanese, Korean, German, Italian, Portuguese and 38 Minor Languages So far. 📃 Paper • 🌐 Demo • 🤗 ApolloMoEDataset • 🤗 ApolloMoEBench • 🤗 Models •🌐 Apollo • 🌐 ApolloMoE 🌈 Update [2024.10.15] ApolloMoE repo is published!🎉 Languages Coverage 12 Major Languages and 38 Minor Languages Click to… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ApolloMoEBench.textquestion-answering10K<n<100K0 likes193 downloads2y agoHugging Face04apollo-research /contrastive-belief-updates Contrastive SDF training corpora This dataset is from Apollo Research and accompanies the paper Measuring Reward-Seeking via Contrastive Belief Updates. For more, see rewardseeking.ai. This dataset contains the 30 synthetic-document corpora used across the completed experiments for the paper: 24 coding-style corpora and 6 honesty-versus-task-completion corpora. Important: entirely synthetic, model-generated content Every document in this dataset is synthetic and… See the full description on the dataset page: https://huggingface.co/datasets/apollo-research/contrastive-belief-updates.text100K<n<1M5 likes119 downloads2mo agoHugging Face05open-llm-leaderboard /rootxhacker__Apollo-70B-detailsgated Dataset Card for Evaluation run of rootxhacker/Apollo-70B Dataset automatically created during the evaluation run of model rootxhacker/Apollo-70B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/rootxhacker__Apollo-70B-details.tabular10K<n<100K0 likes73 downloads2y agoHugging Face06kunishou /ApolloCorpus-ja ApolloCorpus-ja 概要 多言語医療データセットの ApolloCorpus を日本語に自動翻訳した 525k の指示チューニングデータセットになります。ApolloCorpus は、オープンソースでかつ品質を担保できるデータのみをスクリーニングし収集されたデータセットになります。詳細は 論文 をご覧下さい。 翻訳対象ファイル データ量が多いのでひとまず以下の 1 ファイルのみを翻訳しました。なお、英語以外のデータセットについては翻訳品質が低くくなるため、英語データセットのみを日本語に自動翻訳しました(今後、他のファイルを追加で翻訳する場合も英語データのファイルのみを対象にすると思います)。 medicalPaper_en_qa.json (525k) 使用上の注意 多言語データセットを自動翻訳で日本語に翻訳したものであり、翻訳誤りも一部含まれています。医療領域での LLM に利用する際は十分注意した上で使用して下さい。 text100K<n<1M4 likes43 downloads3y agoHugging Face07UMCU /apollo_english_guidelines_translated_to_dutch_with_geminiflash1.5 Data description Translation of the English medical guidelines that are part of the Apollo corpus, using the LLM Gemini Flash 1.5 Acknowledgement The work received funding from the European Union's Horizon Europe research and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project). For more information on the background, see Datatools4Heart Huggingface/Website/Git tabular10K<n<100K0 likes34 downloads2y agoHugging Face08UMCU /apollo_english_guidelines_translated_to_dutch_with_marianmt Data description Apollo corpus, English guidelines translated to Dutch using MariaNMT. Acknowledgement The work received funding from the European Union's Horizon Europe research and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project). For more information on the background, see Datatools4Heart Huggingface/Website/Git tabulartext-generation10K<n<100K0 likes26 downloads2y agoHugging Face09UMCU /apollo_english_books_translated_to_dutch_with_geminiflash15 Data description Translation of the English medical books that are part of the Apollo corpus, using the LLM Gemini Flash 1.5 Acknowledgement The work received funding from the European Union's Horizon Europe research and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project). For more information on the background, see Datatools4Heart Huggingface/Website/Git tabular100K<n<1M0 likes26 downloads2y agoHugging Face10UMCU /apollo_english_guidelines_translated_to_dutch_with_gpt4omini Data description Translation of the English medical guidelines that are part of the Apollo corpus, using the LLM GPT 4o mini Acknowledgement The work received funding from the European Union's Horizon Europe research and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project). For more information on the background, see Datatools4Heart Huggingface/Website/Git tabular10K<n<100K0 likes24 downloads2y agoHugging Face11pthinc /BCE-Prettybird-Nano-Apollo-v0.1 BCE-Prettybird-Nano-Apollo-v0.1 Synthetic Multi-Language Software Engineering & UI/UX Dataset (1,070 Examples) This dataset contains 1,070 synthetic, high-quality examples covering a broad range of software engineering, architecture, database development, web design, UI/UX design, and design pattern implementations across multiple programming languages and frameworks. The collection includes: SOLID principle code examples in PHP, C#, Python, C++, Java, and JavaScript Design… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Apollo-v0.1.texttext-generation1K<n<10K0 likes20 downloads4mo agoHugging Face12UMCU /apollo_english_guidelines_translated_to_dutch_with_nllb200 Data description Translation of the English medical guidelines that are part of the Apollo corpus, using the NLLB200-600M NTM. Acknowledgement The work received funding from the European Union's Horizon Europe research and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project). For more information on the background, see Datatools4Heart Huggingface/Website/Git tabulartext-generation10K<n<100K0 likes16 downloads2y agoHugging Face13open-llm-leaderboard /rootxhacker__Apollo_v2-32B-detailsgated Dataset Card for Evaluation run of rootxhacker/Apollo_v2-32B Dataset automatically created during the evaluation run of model rootxhacker/Apollo_v2-32B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/rootxhacker__Apollo_v2-32B-details.tabular10K<n<100K0 likes14 downloads2y agoHugging Face14marco-apollonia /test_data Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/marco-apollonia/test_data.textn<1K0 likes11 downloads2y agoHugging Face15open-llm-leaderboard /rootxhacker__apollo-7B-detailsgated Dataset Card for Evaluation run of rootxhacker/apollo-7B Dataset automatically created during the evaluation run of model rootxhacker/apollo-7B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/rootxhacker__apollo-7B-details.tabular10K<n<100K0 likes10 downloads2y agoHugging Face16hula1 /Apollo-math-peopletextn<1K0 likes5 downloads2y agoHugging Face17ZombitX64 /ApolloMoE_Ebench_Thaitextn<1K0 likes4 downloads1y agoHugging Face18Apollo12 /dataset_testtextn<1K0 likes3 downloads1y agoHugging Face19Apollo12 /dataset_flock_002textn<1K0 likes3 downloads1y agoHugging Face20Apollo12 /sn-flockgatedtextn<1K0 likes1 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.