datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SnakeAI_TF_PPO_V1The Hebrew word נָחָשׁ (Nāḥāš) is used in the Hebrew Bible to identify the serpent that appears in Genesis 3:1, in the Garden of Eden.
This contains #7000000 training parameters/timestep for Snake_AI game using TensorFlow 2.XX.
Best score and performance comes from data #4800000 dataset for ActorCritic, with an average score of 72 with no action mask/upfront rules. Full reinforcement learning with score/reward as a priority
Agent score can be improved with the combination of more training and… See the full description on the dataset page: https://huggingface.co/datasets/privateboss/SnakeAI_TF_PPO_V1.ppo-CartPole-v1
Dataset Card for "ppo-CartPole-v1"
More Information needed
lerobot_arm_2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
10
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/ppooar/lerobot_arm_2.ppo-Pendulum-v1
Dataset Card for "ppo-Pendulum-v1"
More Information needed
lerobot_place_2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
10
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/ppooar/lerobot_place_2.ppo-seals-CartPole-v0
Dataset Card for "ppo-seals-CartPole-v0"
More Information needed
pp-ocr-mnn-eval
PP-OCR MNN evaluation dataset (811-cell matrix)
Images (273) + canonical paddle.inference baselines (808 json) + configs
for scoring pp-ocr-mnn outputs. See README.md and
https://github.com/baicai1145/pp-ocr-mnn (tools/score.py).
A single-file snapshot is also included as ppocr-eval-dataset.tar.zst.
SnakeAI_TF_PPO_V0Action mask has been implemented, the model has been updated to support 'Training Resumption' after system disruption. Utilizing the same training parameters as the "Full Reinforcement learning Agent", this agent prioritizes survival over rewards. It's playtime for 100 games is 6hrs, compared to 2hrs for the FRLA.
This demonstrates the agent is adapting for survival, but not to the desired goal of higher scores/reward. #10000000 training timesteps.
Training Hyperparameters is the same as the… See the full description on the dataset page: https://huggingface.co/datasets/privateboss/SnakeAI_TF_PPO_V0.wiki-lingua-ppoabl-ppocr-trainlerobot_new_3rl_ppo_profilesppo_dataset
Dataset Card for "ppo_dataset"
More Information needed
details_ppopiolek__tinyllama_eng_longlerobot_arm_3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
10
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/ppooar/lerobot_arm_3.weNavigate-PPO_controller_training_episodeslerobot_place_1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
10
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/ppooar/lerobot_place_1.imdb-ppoeval_ep1000_seedNone_circle_big_10000_ppo_circle_bigThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "racecar",
"total_episodes": 20,
"total_frames": 8597,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Lyrasilas/eval_ep1000_seedNone_circle_big_10000_ppo_circle_big.eval_ep1000_seedNone_default_10000_ppo_circle_smallThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "racecar",
"total_episodes": 20,
"total_frames": 5032,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Lyrasilas/eval_ep1000_seedNone_default_10000_ppo_circle_small.lerobot_arm_4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
10
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/ppooar/lerobot_arm_4.ppo_ppl_thinkfinal_stage10
PPO Stage-10 Curriculum Dataset
本仓库提供基于 Atomheart-Father/ppo_pool_24000_ppl10_sys10_thinkfinal_toklen 的分阶段 PPO/DPO 训练数据。源数据来自 OpenR1 子集与 OT-114k 数学子集,保持 <think>...</think><final>...</final> 的答案格式,并用 query-only PPL 做难度分桶。
数据切分
stage0 … stage9:共 10 个训练阶段。每阶段目标 2000 条(脚本参数 STAGE_SIZE=2000,FIXED_RATIO=0.8),约 80% 来自同难度分桶(stage_role=fixed),20% 为其他难度的混合样本(stage_role=random)。
eval:从剩余样本中采样(脚本参数 EVAL_SIZE=500),stage_role=eval。
test:剩余部分,stage_role=test。
stage_id:0–9 对应阶段,-1 表示… See the full description on the dataset page: https://huggingface.co/datasets/Atomheart-Father/ppo_ppl_thinkfinal_stage10.eval_data_hh_sft_dpo_ppolerobot_new_2vizdoom-ppo-datasetlerobot_new_1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
10
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/ppooar/lerobot_new_1.PPOpt-data
PersonaAtlas Dataset
Dataset Summary
PersonaAtlas is a synthetic multi-turn conversational dataset designed for persona-aware alignment of large language models. It contains 10,462 conversation samples derived from 2,055 unique personas, enabling research on personalized response generation and preference modeling.
Each example includes:
A structured persona profile (persona) with user preference features
The source prompt (original_query)
The full dialog… See the full description on the dataset page: https://huggingface.co/datasets/HowieHwong/PPOpt-data.ViZDoom-Deathmatch-PPO-XLrg
ViZDoom Deathmatch with pretrained PPO agent playing over 15 episodes.
details_ewqr2130__mistral-inst-ppo
Dataset Card for Evaluation run of ewqr2130/mistral-inst-ppo
Dataset automatically created during the evaluation run of model ewqr2130/mistral-inst-ppo on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_ewqr2130__mistral-inst-ppo.lm-eval-results-Ppoyaa-LexiLumin-7B-private
Dataset Card for Evaluation run of Ppoyaa/LexiLumin-7B
Dataset automatically created during the evaluation run of model Ppoyaa/LexiLumin-7B
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Ppoyaa-LexiLumin-7B-private.
