lightseekorg/kimi-mtp-dataset
Kimi-K2.5 Eagle3 Training Data This dataset contains the instruction-following data used to train an Eagle3 MTP draft model for Kimi-K2.5 with TorchSpec. All responses were regenerated by running Kimi-K2.5 via Engine rather than taken from the original datasets. This is critical for speculative decoding training: the draft model must learn the exact token-level distribution of the target model it is accelerating. The trained Eagle3 draft model is available at… See the full description on the dataset page: https://huggingface.co/datasets/lightseekorg/kimi-mtp-dataset.
Kimi-K2.5 Eagle3 Training Data
This dataset contains the instruction-following data used to train an Eagle3 MTP draft model for Kimi-K2.5 with TorchSpec.
All responses were regenerated by running Kimi-K2.5 via Engine rather than taken from the original datasets. This is critical for speculative decoding training: the draft model must learn the exact token-level distribution of the target model it is accelerating.
The trained Eagle3 draft model is available at lightseekorg/kimi-k2.5-eagle3. If you find this draft model useful, please give our project TorchSpec a 🌟 on GitHub.
Data source
Due to inference resource constraints, some source datasets are only partially regenerated. Here is the list of source datasets used in this mix:
Data format
Each sample contains two fields:
- `conversations`: a list of turns, each with
from(human/gpt/system) andvalue(string). - `source`: the name of the source dataset (see table above).
{
"conversations": [
{"from": "human", "value": "What is the capital of France?"},
{"from": "gpt", "value": "The capital of France is Paris."}
],
"source": "perfectblend"
}Multimodal samples (llava_instruct) use OpenAI vision format in the value field — a list of image_url and text objects — with local image paths replaced by public COCO URLs (http://images.cocodataset.org/train2017/{filename}).
Function-call samples (continual_tool_kimi) use Kimi-K2.5's special token format for tool calls:
<|tool_calls_section_begin|><|tool_call_begin|>{id}<|tool_call_argument_begin|>{args_json}<|tool_call_end|><|tool_calls_section_end|>Tool results are serialized as human turns with the prefix ## Return of {call_id}\n.
Training
See TorchSpec for the full training recipe, configuration, and evaluation results.
License
Apache 2.0. All source datasets are Apache 2.0 or MIT licensed.
