Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ilsp /scipar_parallel_docs SciPar Parallel Documents Dataset Description This dataset contains parallel documents (i.e., titles & abstracts) extracted from academic theses, dissertations, and other scientific texts. In the original paper, we've extracted 9.17M sentence pairs in 31 language pairs from 86 repositories. This version has been created through further processing and filtering to extract parallel documents instead of parallel sentences. To do this, we kept only the parallel titles and… See the full description on the dataset page: https://huggingface.co/datasets/ilsp/scipar_parallel_docs.tabulartext-generation1K<n<10K2 likes152 downloads3y agoHugging Face02ilprl-docse /NepTam-A-Nepali-Tamang-Parallel-Corpus 🧾 NepTam — A Nepali–Tamang Parallel Corpus Dataset Summary NepTam is a high-quality Nepali–Tamang bilingual parallel corpus designed to support research in low-resource neural machine translation (NMT) and linguistic analysis.It contains: 20K gold-standard human-translated sentence pairs, and 80K synthetic pairs generated using the NLLB-200 model fine-tuned on the gold corpus. Each entry includes linguistic metadata such as sentence type, tense, and polarity… See the full description on the dataset page: https://huggingface.co/datasets/ilprl-docse/NepTam-A-Nepali-Tamang-Parallel-Corpus.texttranslation10K<n<100K1 likes105 downloads11mo agoHugging Face03sai-lohith /streamlit_docstextn<1K0 likes71 downloads2y agoHugging Face04docs-benchmarks /compile-benchmarkstabularn<1K0 likes43 downloads2y agoHugging Face05docs-benchmarks /experts-backendstabularn<1K0 likes30 downloads9mo agoHugging Face06datasets-examples /doc-splits-1 [doc] file names and splits 1 This dataset contains a data.csv file at the root. textn<1K1 likes29 downloads3y agoHugging Face07ASHu2 /docs-python-v1 Dataset Card for Dataset Name This dataset card aims to be a base template for creating python docs from methods. This is formatted from semeru/code-code-galeras-code-completion-from-docstring-3k-deduped Dataset Description Curated by: semeru/code-code-galeras-code-completion-from-docstring-3k-deduped Language(s) (NLP): Python License: [More Information Needed] Dataset Sources [optional] Repository:… See the full description on the dataset page: https://huggingface.co/datasets/ASHu2/docs-python-v1.tabularfeature-extraction1K<n<10K2 likes26 downloads3y agoHugging Face08datasets-examples /doc-splits-2 [doc] file names and splits 2 This dataset contains three csv files at the root: train.csv, test.csv, validation.csv. textn<1K0 likes25 downloads3y agoHugging Face09datasets-examples /doc-splits-3 [doc] file names and splits 3 This dataset contains three csv files at the root: my_train_file.csv, test-file.csv, validation1.csv. textn<1K0 likes24 downloads3y agoHugging Face10atitaarora /qdrant_docs_qna_ragastextn<1K0 likes24 downloads3y agoHugging Face11datasets-examples /doc-splits-6 [doc] file names and splits 6 This dataset contains six files at the root, four for the training split, and two for the test split. textn<1K0 likes23 downloads3y agoHugging Face12bky373 /spring-docs License Spring Projects: Apache License 2.0. Copyright © 2024 Broadcom. All Rights Reserved. text1K<n<10K1 likes21 downloads2y agoHugging Face13itsjhuang /watsonx-docs-document-type Watsonx Docs Document Type Classification This dataset is a balanced binary document-level classification subset derived from ibm-research/watsonxDocsQA. Task Classify IBM Watsonx documentation pages by their dominant user-facing purpose: conceptual: documents primarily used to understand or look up information. how-to: documents primarily used to complete a procedure or fix a problem. Splits Split conceptual how-to Total train 140 140 280… See the full description on the dataset page: https://huggingface.co/datasets/itsjhuang/watsonx-docs-document-type.texttext-classificationn<1K0 likes20 downloads5mo agoHugging Face14datasets-examples /doc-splits-5 [doc] file names and splits 5 This dataset contains three files inside data/, called training.csv, eval.csv and valid.csv. textn<1K0 likes17 downloads3y agoHugging Face15Eim /laravel-docstextn<1K1 likes15 downloads3y agoHugging Face16datasets-examples /doc-splits-4 [doc] file names and splits 4 This dataset contains three subdirectories, inside data/, called train, test and validation, with csv files in them. textn<1K0 likes15 downloads3y agoHugging Face17docs-benchmarks /z-imagetabularn<1K0 likes15 downloads9mo agoHugging Face18sayannath /code_docstringstext1K<n<10K0 likes14 downloads1y agoHugging Face19datasets-examples /doc-splits-8 [doc] file names and splits 8 This dataset contains seven files under the data/ directory, three for the train split, one for the test split and three for the random split. textn<1K0 likes13 downloads3y agoHugging Face20ArundathiB /code-docstring-datasettext10K<n<100K0 likes13 downloads10mo agoHugging Face21Equious /codehawks-docs-qatabularn<1K0 likes12 downloads3y agoHugging Face22ProjectsbyGaurav /LangChain_docs_usecases_integrationstextn<1K0 likes10 downloads3y agoHugging Face23datasets-examples /doc-splits-7 [doc] file names and splits 7 This dataset contains six files under the data/ directory, four in the train/ subdirectory, and two in the test/ subdirectory. textn<1K0 likes10 downloads3y agoHugging Face24shapermindai /wasp-docsimagen<1K0 likes10 downloads2y agoHugging Face25HenriLD /FDA_Docstext10K<n<100K0 likes10 downloads2y agoHugging Face26hamees /pybamm-docstextn<1K0 likes9 downloads3y agoHugging Face27EclipsePhage /langchain-docs-csvtext1K<n<10K0 likes8 downloads3y agoHugging Face28docs-benchmarks /kernel-ltx-videotabularn<1K0 likes8 downloads8mo agoHugging Face29mwitiderrick /lamini_docstext1K<n<10K0 likes6 downloads3y agoHugging Face30DoDucAnh /MNLP_rag_docstext100K<n<1M0 likes6 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.