llm-generated
llm-detect-ai-generated-text-berthuman-vs-llm-generated-text-detection-distilbertLLM_generated_text_detectorhuman-vs-llm-generated-text-detection-distilbertdistilbert-llm-generated-essay-classifierhuman-vs-llm-generated-text-detection-distilberthuman-vs-llm-generated-text-detection-distilbertLLM-Generated-Summaries-tdidf-baseline
ivypanda-llm-generated-essays
AI-Generated Essays Dataset
This dataset contains AI-generated academic essays created using the models:
Mistral 7B Instruct v0.2 (Q5_K_M quantized)
Temperature: 0.7
Max tokens: 4096
Top-p: 0.9 (default)
Top-k: 40 (default)
Repeat penalty: 1.1 (default)
Context window: 32768 tokens
Llama 3 13B Instruct v0.1 (Q5_K_M quantized)
Temperature: 0.7
Max tokens: 4096
Top-p: 0.9 (default)
Top-k: 40 (default)
Repeat penalty: 1.1 (default)
Context window: 8192 tokens
DeepSeek-V3.2
API… See the full description on the dataset page: https://huggingface.co/datasets/artfultom/ivypanda-llm-generated-essays.LLM-generated-emoji-descriptions
Emoji Metadata Dataset
Overview
The LLM Emoji Dataset is a comprehensive collection of enriched semantic descriptions for emojis, generated using Meta AI's Llama-3-8B model. This dataset aims to provide semantic context for each emoji, enhancing their usability in various NLP applications, especially those requiring semantic search. The LLM Emoji Dataset was used to build a multilingual search engine for emojies, which you can interact with using this online Streamlit… See the full description on the dataset page: https://huggingface.co/datasets/badrex/LLM-generated-emoji-descriptions.llm-generated-textsThis dataset is composed of parallel texts, generated by LLMs and written by human authors. The methodology for constructing the is based on the [1] and uses prompts from [2].
The dataset comprises of powerful LLMs generations, 21'000 in total. Used LLMs:
GPT4 Turbo 2024-04-09: https://platform.openai.com/docs/models/gpt-4-turbo-and-gpt-4
GPT4 Omni: https://openai.com/index/hello-gpt-4o
Claude 3 Opus: https://www.anthropic.com/news/claude-3-family
Llama3 70B: https://llama.meta.com/llama3/… See the full description on the dataset page: https://huggingface.co/datasets/artnitolog/llm-generated-texts.llm-generated-essayLLM_Generated_Summaries_Dataset
BBC News Summaries Database
A small multilingual dataset of BBC News articles paired with two kinds of short summaries: human-written ones taken from XL-Sum, and machine-written ones I generated myself by prompting five different LLMs.
I built this for my TYP to compare how AI summaries stack up against the human reference across a handful of languages.
What's in here
5 languages: English (en), Spanish (es), French (fr), Arabic (ar), Mandarin Chinese (zh).
For… See the full description on the dataset page: https://huggingface.co/datasets/Ef05/LLM_Generated_Summaries_Dataset.Reindex-Then-Adapt-LLM-Generated-Data
llm-detect-ai-generated-text-datasetteLLM-Agents_Generated_Consensus_Statementsrepro-black-box-detection-llm-generated-text-gjsrepro-telescope-improving-zero-shot-detection-of-llm-generated-content-by-measuring-token-repetirepro-from-llm-generated-conjectures-to-lean-formalizations-automated-polynomial-inequality-prov
