Team Ai
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jeffmeloy /python_documentation_codeDataset created using https://github.com/jeffmeloy/py2dataset using the Python code from the following: https://github.com/ansible/ansible https://github.com/apache/airflow https://github.com/arogozhnikov/einops https://github.com/arviz-devs/arviz https://github.com/astropy/astropy https://github.com/biopython/biopython https://github.com/bjodah/chempy https://github.com/bokeh/bokehhttps://github.com/CalebBell/thermo https://github.com/camDavidsonPilon/lifelines https://github.com/coin-or/pulp… See the full description on the dataset page: https://huggingface.co/datasets/jeffmeloy/python_documentation_code.texttext-generation10K<n<100K1 likes156 downloads2y agoHugging Face02lukesjordan /worldbank-project-documents Dataset Card for World Bank Project Documents Dataset Summary This is a dataset of documents related to World Bank development projects in the period 1947-2020. The dataset includes the documents used to propose or describe projects when they are launched, and those in the review. The documents are indexed by the World Bank project ID, which can be used to obtain features from multiple publicly available tabular datasets. Supported Tasks and Leaderboards No… See the full description on the dataset page: https://huggingface.co/datasets/lukesjordan/worldbank-project-documents.texttable-to-text10K<n<100K5 likes132 downloads4y agoHugging Face03OO-LD /oold-wikidata-schemaorg-documents Wikipedia leads, as the corpus measured them The article leads that oold-llm-bench's Wikidata-schema.org corpus cites: 959 documents, one per entity, each the plain-text introduction of a named revision with its whitespace collapsed. Published because the alternative does not work. The corpus commits each lead's revision id and sha256 rather than its text, and a third party was meant to re-fetch. Extracts are not versioned: the MediaWiki API ignores revids for prop=extracts and… See the full description on the dataset page: https://huggingface.co/datasets/OO-LD/oold-wikidata-schemaorg-documents.texttext-generationn<1K0 likes93 downloads4d agoHugging Face04Kemsekov /Corrupted-russian-word-documents-text-datasetThis is synthetic text-corruption dataset, based on Russian open official/government documents text. This dataset is intended to be used to train LLM to perform text-recovery task. All the errors in text is made solely in Russian sentences, hence ignoring any English sentence. Texts contains complex formatting, which is common for documents. Each line contains json object that have array messages value, which consists of role-based conversation. Each first message is randomly chosen system… See the full description on the dataset page: https://huggingface.co/datasets/Kemsekov/Corrupted-russian-word-documents-text-dataset.texttext-generationn<1K2 likes52 downloads2y agoHugging Face05Teen-Different /grpo-oumi-synthetic-document-claims Dataset Card for GRPO Oumi ANLI Subset Dataset This dataset is a reformatted version of the oumi-ai/oumi-synthetic-document-claims dataset, specifically structured for use with the GRPO trainer. You can find more detailed information about the original dataset at the provided link. Link: https://huggingface.co/datasets/oumi-ai/oumi-synthetic-document-claims Dataset Structure The dataset consists of a list of dictionaries, where each dictionary represents a… See the full description on the dataset page: https://huggingface.co/datasets/Teen-Different/grpo-oumi-synthetic-document-claims.texttext-generation1K<n<10K0 likes43 downloads1y agoHugging Face06stindardlogic /document-summarization-dpo-100k Document Summarization DPO (100K) 100,000 DPO (Direct Preference Optimization) preference pairs for training models to summarize business and professional documents with precision, structure, and analytical depth. Motivation Document summarization is one of the highest-value enterprise AI applications — analysts, lawyers, product managers, and executives use AI to process reports, contracts, and research daily. Models commonly fail by: Losing quantitative data:… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/document-summarization-dpo-100k.texttext-generation100K<n<1M0 likes39 downloads3mo agoHugging Face07playforgecoding /spelling-creator-document-import Spelling Creator document import One section of a lesson as someone might have typed it up, paired with that section as lesson JSON. Spelling Creator is a lesson builder for Spelling to Communicate (S2C), used with nonspeaking spellers. Its Import from text reads a typed lesson with rules, and hands the sections the rules cannot read to a small on-device model; this dataset trains that model. What is in it Chat-format JSONL, ready for a supervised fine-tune (TRL's… See the full description on the dataset page: https://huggingface.co/datasets/playforgecoding/spelling-creator-document-import.texttext-generationn<1K0 likes35 downloads2d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.