Team Ai
Datasetpublic

lightseekorg/kimi-mtp-dataset

Kimi-K2.5 Eagle3 Training Data This dataset contains the instruction-following data used to train an Eagle3 MTP draft model for Kimi-K2.5 with TorchSpec. All responses were regenerated by running Kimi-K2.5 via Engine rather than taken from the original datasets. This is critical for speculative decoding training: the draft model must learn the exact token-level distribution of the target model it is accelerating. The trained Eagle3 draft model is available at… See the full description on the dataset page: https://huggingface.co/datasets/lightseekorg/kimi-mtp-dataset.

sourceHugging Faceapache-2.0updated 6mo agoView on Hugging Face
7likes138downloads
Dataset Card

Kimi-K2.5 Eagle3 Training Data

This dataset contains the instruction-following data used to train an Eagle3 MTP draft model for Kimi-K2.5 with TorchSpec.

All responses were regenerated by running Kimi-K2.5 via Engine rather than taken from the original datasets. This is critical for speculative decoding training: the draft model must learn the exact token-level distribution of the target model it is accelerating.

The trained Eagle3 draft model is available at lightseekorg/kimi-k2.5-eagle3. If you find this draft model useful, please give our project TorchSpec a 🌟 on GitHub.

Data source

Due to inference resource constraints, some source datasets are only partially regenerated. Here is the list of source datasets used in this mix:

DatasetSource# Samples
mlabonne/open-perfectblendperfectblend296,034
liuhaotian/LLaVA-Instruct-150Kllava_instruct123,102
HuggingFaceTB/smoltalksmoltalk_cn48,333
daviddjtzafon/continual-tool-kimi-k2.5continual_tool_kimi4,370
crownelius/KimiK2.5-2000x-formattedkimi_2000x2,144
crownelius/Creative-Writing-KimiK2.5-Cleanedcreative_writing1,393
DCAgent2/terminal_bench_2dcagent873
crownelius/Creative-Writing-Reasoning-KimiK2.5-600xcreative_writing_reasoning655
Total476,904

Data format

Each sample contains two fields:

  • —`conversations`: a list of turns, each with from (human / gpt / system) and value (string).
  • —`source`: the name of the source dataset (see table above).
json
{
  "conversations": [
    {"from": "human", "value": "What is the capital of France?"},
    {"from": "gpt", "value": "The capital of France is Paris."}
  ],
  "source": "perfectblend"
}

Multimodal samples (llava_instruct) use OpenAI vision format in the value field — a list of image_url and text objects — with local image paths replaced by public COCO URLs (http://images.cocodataset.org/train2017/{filename}).

Function-call samples (continual_tool_kimi) use Kimi-K2.5's special token format for tool calls:

<|tool_calls_section_begin|><|tool_call_begin|>{id}<|tool_call_argument_begin|>{args_json}<|tool_call_end|><|tool_calls_section_end|>

Tool results are serialized as human turns with the prefix ## Return of {call_id}\n.

Training

See TorchSpec for the full training recipe, configuration, and evaluation results.

License

Apache 2.0. All source datasets are Apache 2.0 or MIT licensed.