nectec/Pathumma-llm-text-4.0.0
Pathumma-llm-4b-think-4.0.0
A Thai reasoning model from the ThaiLLM national initiative. It emits an explicit thinking trace before its final answer, targeting mathematical reasoning, instruction following, and structured tool use in Thai and English.
Model Overview
Pathumma-llm-4b-think-4.0.0 has the following features:
- Type: Causal Language Model
- Training Stage: Post-training (SFT → DPO)
- Base Model: ThaiLLM, Thai continual-pre-trained
- Number of Parameters: 4B
- Languages: Thai, English
- Mode: Thinking
- Context Length: 262,144
- License: Apache-2.0
Highlights
- Mathematical reasoning — 85.00 on MATH-500 (TH), 56.67 on AIME 2024 (TH)
- Language consistency — 97.86 on code-switching; stays in Thai for Thai prompts
- Instruction following — 71.71 on IFEval (TH), instruction level
- Structured tool use — function calling with inspectable reasoning traces
- Single-GPU deployment — 4B parameters, servable on a single GPU
Quickstart
Use a recent version of transformers; older versions will fail to load the model architecture.
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = "nectec/pathumma-llm-4b-think-4.0.0"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
model_name,
torch_dtype="auto",
device_map="auto",
)
prompt = "ทำไมวงกลมถึงมี 360 องศา"
messages = [
{"role": "user", "content": prompt}
]
text = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
)
model_inputs = tokenizer([text], return_tensors="pt").to(model.device)
generated_ids = model.generate(
**model_inputs,
max_new_tokens=32768,
)
output_ids = generated_ids[0][len(model_inputs.input_ids[0]):].tolist()
# Split the reasoning trace from the final answer.
think_end_id = tokenizer.convert_tokens_to_ids("</think>")
try:
index = len(output_ids) - output_ids[::-1].index(think_end_id)
except ValueError:
index = 0
thinking_content = tokenizer.decode(output_ids[:index], skip_special_tokens=True).strip("\n")
content = tokenizer.decode(output_ids[index:], skip_special_tokens=True).strip("\n")
print("thinking content:", thinking_content) # no opening <think> tag
print("content:", content)Avoid greedy decoding, which can cause repetition loops in the reasoning trace. Reasoning traces also run long, so capping max_new_tokens too low truncates the answer mid-thought.
Serving with vLLM
vllm serve nectec/pathumma-llm-4b-think-4.0.0 \
--served-model-name pathumma-llm-4b-think-4.0.0 \
--host 0.0.0.0 \
--tensor-parallel-size <TP_SIZE> \
--max-model-len 262144 \
--gpu-memory-utilization 0.85 \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_xmlFor local use, Ollama, LM Studio, and llama.cpp are supported once GGUF conversions are available.
Evaluation
Evaluated on Thai-adapted benchmarks covering mathematical reasoning, instruction following, commonsense reasoning, and language consistency.
<sub>All scores are percentages; higher is better.</sub>
Post-training
Post-training starts from the ThaiLLM continual-pre-trained base model and proceeds in two stages.
Supervised fine-tuning
Reasoning supervision is drawn mainly from English corpora. Thai capability comes primarily from the continual pre-training carried out in the base model, reinforced here by the Thai reasoning subset and by cross-lingual transfer.
Preference alignment
Direct Preference Optimization on 6,303 preference pairs, targeting response formatting and style consistency rather than broad behavioural alignment.
Datasets
The specific datasets used in post-training are proprietary. The example counts above represent the training data used in each stage.
Compute
Post-training was conducted on the LANTA high-performance computing cluster using 16 nodes (64 × NVIDIA A100 40GB) for distributed training.
Limitations
- Tool calls may be malformed or fabricated when the provided schema is incomplete or ambiguous
- Accuracy on long analytical chains degrades without retrieval grounding
- Domain coverage follows the training corpora; specialised Thai domains such as legal and clinical text are not specifically targeted
- Post-training used 8,192-token sequences, so behaviour on much longer contexts is untested even though the architecture supports them
License
Released under Apache-2.0, inherited from the base model. Proprietary training data is not distributed with this release.
Citation
@misc{pathumma_llm_4b_think_400,
title = {Pathumma-LLM-4B-Think-4.0.0},
author = {NECTEC LLM Team},
year = {2026},
url = {https://huggingface.co/nectec/pathumma-llm-4b-think-4.0.0}
}About the project
Pathumma-llm-4b-think-4.0.0 is part of ongoing research toward sovereign Thai large language models optimized for analytical and tool-augmented intelligence.
Contact
LLM Team <br> Jirat Arayapityak (jirat.araya@kmutt.ac.th)<br> Kittitat Manokun (kittitat.mano@kmutt.ac.th)<br> Supanat Tangkitvutikul (supanat.tan@dome.tu.ac.th)<br> Chanut Sunatho (chanut.suna@kmutt.ac.th)<br> Arnon Saeoung (anon.saeoueng@gmail.com)<br> Chaianun Damrongrat (chaianun.damrongrat@nectec.or.th)<br> Sarawoot Kongyoung (sarawoot.kongyoung@nectec.or.th)
