datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ToolBench_toolllama_G123_dfsDataset mentioned for ToolBench project https://github.com/OpenBMB/ToolBench
They were in the google drive data.zip https://drive.google.com/drive/folders/1yBUQ732mPu-KclJnuQELEhtKakdXFc3J
These two json are already processed by the original author. Just plugin into the ToolBnech repo deepseed arguments.
--data_path ./toolllama_G123_dfs_train.json \
--eval_data_path ./toolllama_G123_dfs_eval.json \
My objective is to tailer the training data to 1/100 size and used them for the LLaMA-Factory… See the full description on the dataset page: https://huggingface.co/datasets/Yhyu13/ToolBench_toolllama_G123_dfs.toolbench-lfm-chatml
ToolBench ChatML Dataset
Train: 187,494 examples
Eval: 762 examples
Source: ToolBench (toolllama_G123_dfs)
Format: ChatML messages
Dans-Toolmaxx-Functions-ToolbenchToolbenchsnipped
