text-detection
ai-text-detection-pile
Dataset Card for AI Text Dectection Pile
Dataset Summary
This is a large scale dataset intended for AI Text Detection tasks, geared toward long-form text and essays. It contains samples of both human text and AI-generated text from GPT2, GPT3, ChatGPT, GPTJ.
Here is the (tentative) breakdown:
Human Text
Dataset
Num Samples
Link
Reddit WritingPromps
570k
Link
OpenAI Webtext
260k
Link
HC3 (Human Responses)
58k
Link
ivypanda-essays
TODO
TODO… See the full description on the dataset page: https://huggingface.co/datasets/artem9k/ai-text-detection-pile.handwritten_text_detection
Handwritten text detection dataset
Data domain
The blanks were provided by youth organization "Armenian Club" (telegram, instagram ), Russia Moscow.
The text on blanks was written during dictation "Teladrutyun" in 2018
The blanks were labeled by Amir and Renal during research project in HSE MIEM
Dataset info
Contains labeled dictations blanks in YOLO format
91 image in total, 73 (80%) for train and 18 (20%) for test
No image alignment or any preprocess… See the full description on the dataset page: https://huggingface.co/datasets/armvectores/handwritten_text_detection.AI-Text_Detection
About This Dataset
This dataset is derived from the dataset used in the SeqXGPT paper and is uploaded solely for academic and research purposes related to our course project. The dataset is provided as-is and should not be used for any commercial applications or unauthorized activities. Our objective is to facilitate reproducibility and further research in AI-generated text (AIGT) detection.
Here we upload part of its raw data and processed data by using gen_features.py in SeqXGPT… See the full description on the dataset page: https://huggingface.co/datasets/Roxanne-WANG/AI-Text_Detection.ai-text-detection-trainingai-text-detectionInference of R-obi/ai-text-detection-pile-cleaned using Qwen/Qwen3-Embedding-0.6B
ai-text-detection-benchmark-2026
AI text detection: a control set that does not flatter detectors
Built and published by TextSight, an AI detection company.
We sell detection software, which is exactly why we published the numbers that make this
category look bad — including a detector scoring 0% on Claude Opus 5 output.
700 English passages for evaluating AI-text detectors, assembled September 2026. It exists
because most published detector accuracy is measured against test sets that make the task
easy, and… See the full description on the dataset page: https://huggingface.co/datasets/textsightai/ai-text-detection-benchmark-2026.
