Team Ai
Datasetpublic

Aobangaming/Conversational-Fine-Tuning

Dataset Card for CFT Conversational Fine-tuning is a dataset meant for fine-tuning conversational models. This dataset contains prompts generated by ChatGPT and Gemini. Dataset Details Dataset Description Conversational Fine-Tuning aims to be a basic fine-tuning dataset for large language models in early stages of conversational training. The dataset has over 5000+ unique tokens generated by AI language models. Curated by: AobanZ Language(s)… See the full description on the dataset page: https://huggingface.co/datasets/Aobangaming/Conversational-Fine-Tuning.

sourceHugging Facecc-by-sa-4.0updated 28d agoView on Hugging Face
3likes130downloads
Dataset Card

Dataset Card for CFT

Conversational Fine-tuning is a dataset meant for fine-tuning conversational models. This dataset contains prompts generated by ChatGPT and Gemini.

Dataset Details

Dataset Description

Conversational Fine-Tuning aims to be a basic fine-tuning dataset for large language models in early stages of conversational training. The dataset has over 5000+ unique tokens generated by AI language models.

  • —Curated by: AobanZ
  • —Language(s) (NLP): English
  • —License: MIT

Uses

The dataset can be used for fine tuning LLMs, SLMs plus statistical/conversational markov-chains and bigrams. However, it may not work well for markov-chains and bigrams because of their small size and capacity.

Direct Use

The dataset may be used for fine-tuning, and in other cases for smaller models; Training and testing and/or heavy work. Long prompts may overload the model, causing hallucinations.

Out-of-Scope Use

The dataset does not work well for Multilingual models, Code-based AI agents, heavy work, robotics and bidirectional models. The dataset is not entirely code-based, as most hallucinations are when code-based questions are asked to the main model the dataset is trained on - Aoban 3.

Dataset Creation

Source Data

This dataset contains AI-generated information, such output generated by the AI may be incorrect, corrupted or "ai slop".

Bias, Risks, and Limitations

The dataset is generated by AI, such the bias may be incorrect, this dataset is also prone to hallucinations in both SLMs and LLMs; Such hallucinations may be fixed by a reward model or some kind of RLHF.