Team Ai
Datasetpublic

roof37/prompt_engineering_dataset

Overview This dataset contains user votes collected in the text-only category. Each row represents a single vote judging two models (model_a and model_b) on a user conversation, along with the full conversation history and metadata. Key fields include: id: Unique feedback ID of each vote/row. evaluation_session_id: Unique ID of each evaluation session, which can contain multiple separate votes/evaluations. evaluation_order: Evaluation order of the current vote. winner: Battle… See the full description on the dataset page: https://huggingface.co/datasets/roof37/prompt_engineering_dataset.

sourceHugging Facecc-by-4.0updated 11d agoView on Hugging Face
0likes56downloads
Dataset Card

Overview

This dataset contains user votes collected in the text-only category. Each row represents a single vote judging two models (modela and modelb) on a user conversation, along with the full conversation history and metadata. Key fields include:

  • —id: Unique feedback ID of each vote/row.
  • —evaluation_session_id: Unique ID of each evaluation session, which can contain multiple separate votes/evaluations.
  • —evaluation_order: Evaluation order of the current vote.
  • —winner: Battle result containing either modela, modelb, tie, or both_bad.
  • —conversation_a/conversation_b: Full conversation of the current evaluation order.
  • —full_conversation: The entire conversation, including context prompts and answers from all previous evaluation orders. Note that after each vote new models are sampled, thus the responding models vary across the full context.
  • —conv_metadata: Aggregated markdown and token counts for style control.
  • —category_tag: Annotation tags including the categories math, creative writing, hard prompts, and instruction following.
  • —is_code: Whether the conversation involves code.

License

User prompts are licensed under CC-BY-4.0, and model outputs are governed by the terms of use set by the respective model providers.