zai-org/GLM-5.3
2.1k1.4m
chat template: early break in tool result reordering check
Add Toolathlon-Verified evaluation result (#15)
Point terminal-bench-3.0 eval result to unified harborframework/terminal-bench dataset (#16)
Move harness note into notes field for Terminal-Bench 2.1 and 3.0 (#13)
Add harness used for Terminal-Bench 2.1 and 3.0 eval results (#11)
Add community evaluation results for HLE (#3)
Add community evaluation results for DEEP-SWE, TERMINAL-BENCH-2.1, TERMINAL-BENCH-3.0 (#2)
update
Initial commit 0828
