datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rlvr-reward-hacking-scale-no-conftest-20260909-completion
Matched no-conftest RLVR study 20260909-completion
Lossless research records, grouped by model and trajectory type. Only the listed
configurations have published records. Canary diagnostics are excluded from study
estimates; run status in provenance distinguishes retired diagnostics from active
or completed training. Valid failures, refusals and truncations are retained.
The train split name is a dataset-loader convention; record_type identifies
whether a record is training… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909-completion.code_x_glue_cc_code_completion_token
Dataset Card for "code_x_glue_cc_code_completion_token"
Dataset Summary
CodeXGLUE CodeCompletion-token dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/CodeCompletion-token
Predict next code token given context of previous tokens. Models are evaluated by token level accuracy.
Code completion is a one of the most widely used features in software development through IDEs. An effective code completion tool could improve software… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_cc_code_completion_token.Fact-Completion
Dataset Card
Homepage: https://bit.ly/ischool-berkeley-capstone
Repository: https://github.com/daniel-furman/Capstone
Point of Contact: daniel_furman@berkeley.edu
Dataset Summary
This is the dataset for Polyglot or Not?: Measuring Multilingual Encyclopedic Knowledge Retrieval from Foundation Language Models.
Test Description
Given a factual association such as The capital of France is Paris, we determine whether a model adequately "knows" this… See the full description on the dataset page: https://huggingface.co/datasets/Polyglot-or-Not/Fact-Completion.code_x_glue_cc_code_completion_line
Dataset Card for "code_x_glue_cc_code_completion_line"
Dataset Summary
CodeXGLUE CodeCompletion-line dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/CodeCompletion-line
Complete the unfinished line given previous context. Models are evaluated by exact match and edit similarity.
We propose line completion task to test model's ability to autocomplete a line. Majority code completion systems behave well in token level completion, but fail in… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_cc_code_completion_line.BenchMAX_Function_Completion
Dataset Sources
Paper: BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models
Link: https://huggingface.co/papers/2502.07346
Repository: https://github.com/CONE-MT/BenchMAX
Dataset Description
BenchMAX_Function_Completion is a dataset of BenchMAX, sourcing from humanevalplus, which evaluates the code generation capability in multilingual scenarios.
We extend the original English dataset to 16 non-English languages.
The data is first translated… See the full description on the dataset page: https://huggingface.co/datasets/LLaMAX/BenchMAX_Function_Completion.Qwen3.8-27B-thinking-completions
Qwen3.8-27B thinking-mode completions
17,022 prompts from public chat, math and code datasets, each answered once by Qwen3.8-27B (FP8 checkpoint) in thinking mode with its recommended sampling settings (54M completion tokens). Every sample has the reasoning trace and the final answer, as text and as the exact token ids.
The set was generated to train speculative-decoding drafters for this model, so it records the model's own sampled distribution rather than greedy output or… See the full description on the dataset page: https://huggingface.co/datasets/JonasLoos/Qwen3.8-27B-thinking-completions.wildchat-glm53-format-completions
WildChat format completions
9,975 GLM-5.3 answers across 29 parseable formats. Each answer passed its
contract verifier. Failed answers were resampled with the same prompt until one
passed; no semantic judge or answer repair was used.
This is the final release from a 10,000-prompt run; 25 unfinished prompts were excluded. It stores
answers, format instructions, exact contracts, and pinned
WildChat-4.8M references—not
the source prompts or conversations. All rows are in the train… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/wildchat-glm53-format-completions.verus-proof-completion
CodeWalk — Verus Proof Completion
Complete the missing proof annotations so that a composed Rust/Verus program verifies.
A level-N task chains N component functions; completion can require component
postconditions, loop invariants, termination measures, and proof blocks. Part of the
CodeWalk benchmark suite.
400 problems: 100 each at levels L1, L2, L3, L10.
Each file is self-contained apart from the standard Verus vstd dependency.
These are problem inputs only — reference… See the full description on the dataset page: https://huggingface.co/datasets/CodeWalk/verus-proof-completion.openapi-completion-refined
Dataset Card for OpenAPI Completion Refined
A human-refined dataset of OpenAPI definitions based on the APIs.guru OpenAPI directory. The dataset was used to fine-tune Code Llama for OpenAPI completion in the "Optimizing Large Language Models for OpenAPI Code Completion
" paper.
Dataset Details
Dataset Description
The dataset was collected from the APIs.guru OpenAPI definitions directory.
The directory contains more than 4,000 definitions in yaml format.… See the full description on the dataset page: https://huggingface.co/datasets/BohdanPetryshyn/openapi-completion-refined.jb-completions
JB-Completions Dataset: Base Model Safety Evals
Overview
JB-Completions is a dataset designed for evaluating the harmfulness of base language models (i.e., completion/non-instruction-fine-tuned LLMs). This dataset contains pairs of harmful prompts and their corresponding completions, allowing researchers to assess how base models respond to potentially harmful inputs. See our paper on Safety Pretraining for more details!
Dataset Structure
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/locuslab/jb-completions.Dolci-Instruct-RL-Completions
Dolci Instruct RL Completions
Instruction-following completions sampled from OLMo-3-7B-Instruct with teacher logits for knowledge distillation.
Description
This dataset contains instruction-completion pairs with pre-computed teacher logits from OLMo-3-7B-Instruct. Designed for training smaller student models via KL-divergence distillation.
Generation
Completions were sampled from allenai/OLMo-3-7B-Instruct on instruction prompts. For each token position, we… See the full description on the dataset page: https://huggingface.co/datasets/hbfreed/Dolci-Instruct-RL-Completions.task1393_superglue_copa_text_completion
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1393_superglue_copa_text_completion
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1393_superglue_copa_text_completion.task964_librispeech_asr_text_auto_completion
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task964_librispeech_asr_text_auto_completion
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task964_librispeech_asr_text_auto_completion.epfl-llm_guidelines_axolotl-completionepfl-llm/guidelines converted to work with axolotl completion or pretraining.
tiny-vintage-completions
Tiny vintage completions
Synthetic vintage texts, with a cutoff date for year 1900.
Based on unique 2-3 word seeds, extracted from croqaz/Vintage-v1, croqaz/Vintage-v2 and Haykgrigorian/English-historical-corpus-1800-1875.
Check the files seeds1.txt and seeds2.txt.
Generated by TypeWriter-7B-base and Talkie-13B-base completions.
Citation
If you find this dataset valuable, please consider citing:
@misc{Tiny-vintage-completions,
title = {Tiny vintage completions}… See the full description on the dataset page: https://huggingface.co/datasets/croqaz/tiny-vintage-completions.heretic-completions
Heretic Completions
Model completions used as SFT targets for a refusal-abliteration LoRA study.
Each row pairs a prompt from a red-teaming / over-refusal benchmark with a
completion from a refusal-removed ("heretic" / abliterated) model.
Safety notice. This is a private research dataset. Many completions
comply with harmful or dual-use requests by design, so the refusal signal
can be measured and abliteration studied. Do not redistribute or use outside
authorized safety… See the full description on the dataset page: https://huggingface.co/datasets/noahrossi/heretic-completions.task1389_hellaswag_completion
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1389_hellaswag_completion
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1389_hellaswag_completion.DaVinci_Completion
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/mskov/DaVinci_Completion.customer_support_auto_completionnvlabs-verilogeval-v2-completionVerilogEvalv2 complete-iccad-2023 dataset from the VerilogEval paper. Paper: Revisiting VerilogEval: Newer LLMs, In-Context Learning, and Specification-to-RTL Tasks Repo: https://github.com/NVlabs/verilog-eval).
Disclaimer: I am not the original author and uploaded this here only for convenience! Please refer to the original repo for any information.
gemma-4-e2b-deception-behavior-completions
Gemma-4-E2B deception & behavior completions
Consolidated 910-row corpus of (scenario prompt + Gemma-4-E2B-generated completion) pairs from earlier mechanistic-interpretability experiments. Each row captures the prompt the model saw and the text it actually produced; for a subset, Claude-Haiku-4-5 judge verdicts and SAE-feature labels are included.
The corpus is meant to be used as activation-extraction input for downstream interpretability work — Natural Language Autoencoder (NLA)… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/gemma-4-e2b-deception-behavior-completions.springboot-code-completion-dataset
Spring Boot Code Completion Dataset
概述
这是一个专为 Spring Boot 框架代码补全任务设计的中文优先数据集。通过从 GitHub 上 90 个高质量开源 Spring Boot 项目中提取方法级、类级代码片段及配置文件,构建而成。数据集重点突出 Spring Boot 典型特征(注解驱动、依赖注入、配置绑定等),适用于参数高效微调(PEFT,如 LoRA/QLoRA)下的代码生成研究。
数据集统计
总样本数:81,085 条
训练集:64,868 条
验证集:8,108 条
测试集:8,109 条
优先级 2 样本数(含典型 Spring Boot 特征,如 @RestController、@Service、@Entity 等):19,740 条
优先级 2 占比:24.34%
数据格式(Alpaca 风格)
每条数据为 JSON 对象,包含以下字段:
{
"instruction": "请在Spring… See the full description on the dataset page: https://huggingface.co/datasets/Fuleee/springboot-code-completion-dataset.task1339_peixian_equity_evaluation_corpus_text_completion
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1339_peixian_equity_evaluation_corpus_text_completion
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1339_peixian_equity_evaluation_corpus_text_completion.hotpotqa-dev-raft-subset-completionFollows RAFT to generate question, documents, answer triplets
from the first 110 512-token chunks of the HotPotQA dev set (fullwiki) with 2 questions per chunk and 3 distractor docs
and formatted into completion.
task961_ancora-ca-ner_text_auto_completion
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task961_ancora-ca-ner_text_auto_completion
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task961_ancora-ca-ner_text_auto_completion.example-axolotl-completionevents-2015-completion
2015 events completion corpus
Every row refers to an event dated 2015-01-01 .. 2015-12-31.
Sources (unique facts)
source
role
newspaper:guardian
The Guardian Open Platform (2015 headlines + trail text)
newspaper:gdelt:*
GDELT events linked to BBC, CNN, Reuters, NYT, WaPo, Al Jazeera, AP, Independent, Bloomberg, Guardian
wikipedia:current_events
Portal:Current events day pages for all of 2015
wikipedia:deaths_2015
Deaths in January–December 2015… See the full description on the dataset page: https://huggingface.co/datasets/fitleech/events-2015-completion.rr-circuit-breakers-attack-completions
RR (Circuit Breakers) attack completions with three-judge scores
This dataset bundles attack completions generated against
GraySwanAI/Llama-3-8B-Instruct-RR
(the "circuit breakers" defense), each scored by three independent judges:
local:strongreject (Lin et al., StrongREJECT classifier — most permissive)
local:harmbench (HarmBench classifier — middle)
local:gpt_oss (gpt-oss-safeguard-20b — strictest)
Headline finding: judges DISAGREE dramatically on… See the full description on the dataset page: https://huggingface.co/datasets/samuelsimko/rr-circuit-breakers-attack-completions.
