Deltarunefan/Deltarune-Complete-Transcript-Cleaned
Deltarune Chapters 1–4 Dataset Fan-made transcript dataset covering Deltarune Chapters 1 through 4. Processed from video playthroughs and cross-referenced with game data. Intended to provide LLMs with structured narrative context for a game whose content is underrepresented in training corpora. Why This Exists As of early 2026, major LLMs (including models with training cutoffs past July 2025) fail to recall basic plot details of Deltarune Chapters 3 and 4 despite… See the full description on the dataset page: https://huggingface.co/datasets/Deltarunefan/Deltarune-Complete-Transcript-Cleaned.
4210
1import pandas as pd2import glob3import os4 5def make_parquets():6 jsonl_files = glob.glob('chap*_dataset.jsonl')7 8 if not jsonl_files:9 print("[-] Files chap*_dataset.jsonl Not found in current dir")10 return11 12 all_dataframes = []13 required_columns = ['context', 'speaker', 'text']14 15 print(f"[*] Found files: {len(jsonl_files)}")16 17 for file_name in jsonl_files:18 try:19 df = pd.read_json(file_name, lines=True)20 21 df = df[required_columns]22 23 output_name = file_name.replace('.jsonl', '.parquet')24 df.to_parquet(output_name, index=False, engine='pyarrow')25 26 print(f"[+] Created: {output_name} ({len(df)} strings)")27 all_dataframes.append(df)28 29 except KeyError:30 print(f"[!] Error in {file_name}: missing required columns {required_columns}")31 except Exception as e:32 print(f"[!] Unable to process {file_name}: {e}")33 34 if all_dataframes:35 full_df = pd.concat(all_dataframes, ignore_index=True)36 full_df.to_parquet('full_chapters_dataset.parquet', index=False, engine='pyarrow')37 print(f"\n[OK] Main file ready: full_chapters_dataset.parquet ({len(full_df)} строк)")38 39if __name__ == "__main__":40 make_parquets()41 