Team Ai
Datasetpublic

mcipriano/stackoverflow-kubernetes-questions

The purpose of this dataset is to provide the opportunity to perform any training, fine-tuning, etc. for any Language Model. In the 'data' folder, you will find the dataset in Parquet format, which is one of the formats used for these processes. In case it may be useful for other purposes, I have also included the dataset in CSV format. All data in this dataset were retrieved from the Stack Exchange network using the Stack Exchange Data explorer tool… See the full description on the dataset page: https://huggingface.co/datasets/mcipriano/stackoverflow-kubernetes-questions.

sourceHugging Facecc-by-sa-4.0updated 3y agoView on Hugging Face
31likes118downloads
Dataset Card

The purpose of this dataset is to provide the opportunity to perform any training, fine-tuning, etc. for any Language Model. In the 'data' folder, you will find the dataset in Parquet format, which is one of the formats used for these processes.

In case it may be useful for other purposes, I have also included the dataset in CSV format.

All data in this dataset were retrieved from the Stack Exchange network using the Stack Exchange Data explorer tool (https://github.com/StackExchange/StackExchange.DataExplorer). Specifically, the dataset contains all the Question-Answer pairs from Stack Overflow with Kubernetes tags. Specifically, in each Question-Answer pair, the Answer is the one with a positive and maximum score. Posts on Stack Overflow with negative scores have been excluded from the dataset.