Team Ai
Datasetpublic

VGraf/synthetic_preference_dataset_multi_1746753469

allenai/open_instruct: Rejection Sampling Dataset See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail Configs args: {'add_timestamp': True, 'hf_entity': 'VGraf', 'hf_repo_id': 'synthetic_preference_dataset_multi', 'hf_repo_id_scores': 'synthetic_preference_dataset_multi_scores', 'input_filename':… See the full description on the dataset page: https://huggingface.co/datasets/VGraf/synthetic_preference_dataset_multi_1746753469.

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes8downloads
Dataset Card

allenai/open_instruct: Rejection Sampling Dataset

See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail

Configs

args:
{'add_timestamp': True,
 'hf_entity': 'VGraf',
 'hf_repo_id': 'synthetic_preference_dataset_multi',
 'hf_repo_id_scores': 'synthetic_preference_dataset_multi_scores',
 'input_filename': '/weka/oe-adapt-default/victoriag/synth_data/self-talk/clarify_persona_500samples_8turns_2completions_gpt3.5_gpt3.5_tulupref.jsonl',
 'max_parallel_requests': 100,
 'model': 'gpt-4o-2024-08-06',
 'model_names_or_paths': ['gpt-4'],
 'num_completions': 2,
 'num_turns': 8,
 'push_to_hub': True,
 'save_filename': '/weka/oe-adapt-default/victoriag/synth_data/self-talk/prefs/clarify_persona_500samples_8turns_2completions_gpt3.5_gpt3.5_tulupref.jsonl'}

Additional Information

  1. 1.Command used to run python open_instruct/rejection_sampling/synthetic_preference_dataset_multi.py --input_filename /weka/oe-adapt-default/victoriag/synth_data/self-talk/clarify_persona_500samples_8turns_2completions_gpt3.5_gpt3.5_tulupref.jsonl --model gpt-4o-2024-08-06 --num_turns 8 --save_filename /weka/oe-adapt-default/victoriag/synth_data/self-talk/prefs/clarify_persona_500samples_8turns_2completions_gpt3.5_gpt3.5_tulupref.jsonl --num_completions 2 --push_to_hub