digi-texx/calib-agentic-normalized
calib-agentic (v1) Normalized calibration corpus for post-training quantization. Every row is OpenAI-shaped (messages + tools) regardless of upstream dialect, with provenance in origin and precomputed stats for stratified selection. Schema messages[] — role (system/user/assistant/tool), content, name, tool_call_id, tool_calls[] (name, arguments as a JSON string), reasoning_content tools[] — name, description, parameters (JSON string of the JSON-Schema) origin —… See the full description on the dataset page: https://huggingface.co/datasets/digi-texx/calib-agentic-normalized.
calib-agentic (v1)
Normalized calibration corpus for post-training quantization. Every row is OpenAI-shaped (messages + tools) regardless of upstream dialect, with provenance in origin and precomputed stats for stratified selection.
Schema
messages[]—role(system/user/assistant/tool),content,name,tool_call_id,tool_calls[](name,argumentsas a JSON string),reasoning_contenttools[]—name,description,parameters(JSON string of the JSON-Schema)origin—source_id,repo,config,split,revision,row_uid,position,license(per row),domain,language,category,adapters,corpus_versionstats—n_messages,n_turns,n_tools,n_tool_calls,has_tool_calls,has_tool_results,has_reasoning,n_chars,char_entropy,max_char_run,non_ascii_ratio,longest_word_len
arguments and parameters are JSON strings, not nested structs: tool schemas are arbitrarily nested and mutually inconsistent across sources, so a typed struct would either fail to unify or silently coerce fields away.
