datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
flan-t5-base-embed-refinedwebAll of the data together is around 61GB. It's the last hidden states of 131,072 samples from refinedweb padded/truncated to 512 tokens on the left, fed through google/flan-t5-base.
Structure:
{
"encoding": List, shaped (512, 768) aka (tokens, d_model),
"text": String, the original text that was encoded,
"attention_mask": List, binary mask to pass to your model with encoding to not attend to pad tokens
}
lewtun__hugging-face-test-t5-base.outputs.json-36bf2a59__1645559101lewtun__hugging-face-test-t5-base.outputs.json-36bf2a59__1646049601lewtun__hugging-face-test-t5-base.outputs.json-36bf2a59__1646052073
GEM Submission
Submission name: Hugging Face test T5-base.outputs.json 36bf2a59
lewtun__hugging-face-test-t5-base.outputs.json-36bf2a59__1645800191lewtun__hugging-face-test-t5-base.outputs.json-36bf2a59__1646049378lewtun__hugging-face-test-t5-base.outputs.json-36bf2a59__1646049424lewtun__hugging-face-test-t5-base.outputs.json-36bf2a59__1646049876lewtun__hugging-face-test-t5-base.outputs.json-36bf2a59__1646050898lewtun__hugging-face-test-t5-base.outputs.json-36bf2a59__1645558682lewtun__hugging-face-test-t5-base.outputs.json-36bf2a59__1646051364google__flan-t5-base-details
Dataset Card for Evaluation run of google/flan-t5-base
Dataset automatically created during the evaluation run of model google/flan-t5-base
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/google__flan-t5-base-details.tis-quantile-datasets-gtr-t5-base
Targeted Instruction Selection: Quantile Datasets (EMBED)
This repository contains distance quantile subsets computed using the EMBED data representation method, as presented in the paper A Critical Look at Targeted Instruction Selection: Disentangling What Matters (and What Doesn't).
Project Resources
Paper: arXiv:2602.14696
GitHub: dcml-lab/targeted-instruction-selection
Dataset Description
Instruction fine-tuning of large language models (LLMs) often… See the full description on the dataset page: https://huggingface.co/datasets/Harvard-DCML/tis-quantile-datasets-gtr-t5-base.tis-dolci-subset-datasets-gtr-t5-basetis-dolci-quantile-datasets-gtr-t5-baseEvaluation_google-flan-t5-baset5-cardiology-base-chunkedsemantic-corruption-t5-v1_1-basecreating real/fake ("chosen"/"rejected") pairs where chosen are true completions and rejected are completions generated by T5 v1.1 base
27-11-flan-t5-baseag_news-mia_ag_news_client0tis-subset-datasets-gtr-t5-baseC4-Pile-T5-base-Instructions27-11-flan-t5-basexsum-mia_xsum_client0autotrain-data-t5baseparaphrase
AutoTrain Dataset for project: t5baseparaphrase
Dataset Description
This dataset has been automatically processed by AutoTrain for project t5baseparaphrase.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"feat_Unnamed: 0": 69,
"text": "1\uba85 - \uc5f0 15\ub9cc \uc6d0\n2\uba85 - \uc5f0 30\ub9cc \uc6d0\n3\uba85 \uc774\uc0c1 - \uc5f0… See the full description on the dataset page: https://huggingface.co/datasets/sieu-n/autotrain-data-t5baseparaphrase.27-11-flan-t5-baseag_news-mia_ag_news_client227-11-flan-t5-basexsum-mia_xsum_client627-11-flan-t5-basewikitext-mia_wikitext_client227-11-flan-t5-basewikitext-mia_wikitext_client3tokenized_T5_base
Dataset Card for "tokenized_T5_base"
More Information needed
baseline-dataset-t5-base27-11-flan-t5-baseag_news-mia_ag_news_client7
