datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
story_generation_reward_test
Reward Test — Held-Out Evaluation Set (EpisodeBench)
This dataset is the held-out test set for automatic narrative evaluators released as part of EpisodeBench, a full-cycle benchmarking pipeline for long-form interactive story generation with controllable RL.
It is designed to measure how well an LLM-as-a-judge calibrates to EpisodeBench's synthesized rubric targets. Specifically, the paper reports the average absolute gap between each evaluator's predicted score and the synthesized… See the full description on the dataset page: https://huggingface.co/datasets/HeAAAAA/story_generation_reward_test.radon-test-code_generation
radon-test-code_generation
Description
Code generation test dataset for RADON model evaluation with programming prompts
Usage
Load Dataset
from datasets import load_dataset
dataset = load_dataset("MagistrTheOne/radon-test-code_generation")
print(dataset)
Use with RADON Model
from transformers import AutoModelForCausalLM, AutoTokenizer
# Load RADON model
model = AutoModelForCausalLM.from_pretrained("MagistrTheOne/RadonSAI")
tokenizer =… See the full description on the dataset page: https://huggingface.co/datasets/MagistrTheOne/radon-test-code_generation.shaer-sft-test-generations-k5
Shaer SFT Test Generations K=5
This dataset contains generated outputs from the final Shaer SFT adapter on the held-out test split.
Source dataset: Shaer-AI/ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500-splits
Source split: test
Source rows: 3,481
Samples per source row: 5
Total generated rows: 17,405
Base model: Navid-AI/Yehia-7B-preview
Adapter repo: Shaer-AI/Shaer-adapters
Adapter subfolder: adapters/fresh_sft/train/best
Generation backend: vLLM with LoRA… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI-2/shaer-sft-test-generations-k5.shaer-sft-test-generations-k5-meter-count
Shaer SFT Test Generations K=5 Meter/Count Scores
This dataset enriches Shaer-AI/shaer-sft-test-generations-k5 with automatic meter and requested-line-count evaluation.
Rows: 17,405
Metric pass: 4BiLSTM meter classifier plus line-count adherence
Core metric columns: meter, count_adherence, parsed_num_lines, requested_num_lines, meter_eval_status, count_eval_status
Overall meter mean: 0.6520384872952881
Overall count-adherence mean: 0.9783539094873928
The automatic meter score… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI/shaer-sft-test-generations-k5-meter-count.
