Team Ai
Datasetpublic

digi-texx/calib-agentic-normalized

calib-agentic (v1) Normalized calibration corpus for post-training quantization. Every row is OpenAI-shaped (messages + tools) regardless of upstream dialect, with provenance in origin and precomputed stats for stratified selection. Schema messages[] — role (system/user/assistant/tool), content, name, tool_call_id, tool_calls[] (name, arguments as a JSON string), reasoning_content tools[] — name, description, parameters (JSON string of the JSON-Schema) origin —… See the full description on the dataset page: https://huggingface.co/datasets/digi-texx/calib-agentic-normalized.

sourceHugging Faceotherupdated 1mo agoView on Hugging Face
0likes48downloads
Dataset Card

calib-agentic (v1)

Normalized calibration corpus for post-training quantization. Every row is OpenAI-shaped (messages + tools) regardless of upstream dialect, with provenance in origin and precomputed stats for stratified selection.

Schema

  • —messages[] — role (system/user/assistant/tool), content, name, tool_call_id, tool_calls[] (name, arguments as a JSON string), reasoning_content
  • —tools[] — name, description, parameters (JSON string of the JSON-Schema)
  • —origin — source_id, repo, config, split, revision, row_uid, position, license (per row), domain, language, category, adapters, corpus_version
  • —stats — n_messages, n_turns, n_tools, n_tool_calls, has_tool_calls, has_tool_results, has_reasoning, n_chars, char_entropy, max_char_run, non_ascii_ratio, longest_word_len

arguments and parameters are JSON strings, not nested structs: tool schemas are arbitrarily nested and mutually inconsistent across sources, so a typed struct would either fail to unify or silently coerce fields away.