OpenFormosa/barbet-long-context-sft
Barbet long-context SFT Release 876234f1c7f44a4e0b1e6accbdbcb58ade02063b1d00119a9014db5d29513d7b preserves 4139 active records. This is one joint assistant-only SFT dataset; no Barbet model training has been run. The skill-prefill migration has revised 1245 of 1254 records from its fixed base snapshot. Revisions replace their original records in the explicit shard lists above. Old bundles and releases remain available at their pinned commits. Additional records from other… See the full description on the dataset page: https://huggingface.co/datasets/OpenFormosa/barbet-long-context-sft.
Barbet long-context SFT
Release 876234f1c7f44a4e0b1e6accbdbcb58ade02063b1d00119a9014db5d29513d7b preserves 4139 active records. This is one joint assistant-only SFT dataset; no Barbet model training has been run.
The skill-prefill migration has revised 1245 of 1254 records from its fixed base snapshot. Revisions replace their original records in the explicit shard lists above. Old bundles and releases remain available at their pinned commits. Additional records from other contributors are preserved, not counted as migration work. Independently sourced new anchors added after the fixed migration base:
- Their source lineage and per-record profile are recorded in the release manifest.
New examples place a stable general skill library and reference knowledge before pluggable domain procedures, the task's public API, and its question. The answer must follow applicable procedures. New revision sequences contain 1,000,000 to 1,048,576 exact Pangolin tokens; legacy records retain their recorded lengths. The editable skill source library is maintained at OpenFormosa/barbet-skill-library. GitHub changes are reviewed through pull requests; they do not silently alter this pinned dataset release. Reference packs can differ between batches: a whole unrelated background document may be omitted to reserve space for a complete task paper. Within such a batch, the default prefill is shared; per-record prefix hashes identify exact reuse. Task-paper evidence can occur late, after the shared prefill. This does not establish evidence coverage at every depth of the long context. The temporary mixed release has explicit per-record recipe/renderer profiles in the manifest. Tokenizer and Arrow schema remain common; legacy project locks are retained for reproducibility. Use publisher protocol 5 for future releases.
records contains full text and canonical messages; tokenized contains complete input_ids and labels. All prefix, user and observation tokens have label -100. Preserve labels: do not replace them with full-token LM loss. The trainer performs causal shifting exactly once; no cross-example packing is enabled.
Sources, revisions, attribution and per-source license labels are retained. Source/teacher use is authorizedbyuser; this is not a blanket relicense or a claim of independently reviewed contracts. Shared reference documents cross splits; do not describe this as document-disjoint evaluation. Sealed tests are not public.
Verification distinguishes actual restricted execution, model/rubric audits and human review (not claimed). Long reference bytes and masks are checked, but the entire reference corpus is not semantically certified. The Qwen routing and skill swap probes are data/teacher checks. Barbet skill lift, context benefit, BPB and 1M accuracy remain not_measured. Prefill reuse is a data-layout property, not a claim that Barbet's attention/Mamba cache has been implemented or validated.
Teacher diagnostics are reported separately from supervised-target verification. For offline tool-selection records, schema_checked means the proposed call was validated against a frozen contract; it does not mean the API was executed. Earlier immutable skill-prefill bundles may label this tier executed; their provenance and verification reports still state real_tool_execution=false. Those historical rows require a separate versioned metadata correction. Qwen can return an incorrect answer on the full long input even when the stored target passes the original exact arithmetic verifier. Such diagnostic answers are never substituted for the verified target. Incorrect private self-reported skill IDs are also recorded as diagnostic failures and are never supervised. Per-bundle verification reports disclose these failures and any clarified diagnostic retests; do not interpret the release as a claim of perfect teacher routing or 1M accuracy. Some frozen math checks require exact reference wording. A correct calculation with different wording can fail that check. Where present, the numeric supplement and semantic review are reported separately; the original text-match failure is retained. A numeric extraction check alone does not validate the whole answer. Paper-QA frozen checks also require literal citation and answer substrings. Different citation syntax or phrasing can fail these checks; unsupported teacher claims remain semantic failures. Preserved paper targets require their original source binding, frozen checks, rejected negative controls and fresh source audits. Failed diagnostic re-answers are never substituted into the supervised targets. Where an audit makes a verifiably false source/contract claim, an explicitly identified Codex source adjudication can resolve it without changing the target or verifier. The prior Qwen verdict remains recorded and is not counted as a Qwen acceptance or human review; see task_audit_resolution where present. For a Python diagnostic rejected solely for Markdown fences, a separately recorded local-Qwen format repair may be checked for exact AST identity and rerun against the frozen sandbox tests. The original failed response stays failed; repaired diagnostic results are separate and never replace the preserved source target. The executed_projection verification tier means a narrowly reviewed AST projection ran in the restricted sandbox; the annotated source program itself was not directly executed. Its provenance records exactly what was removed. A single semantic code repair may also be retained as supplemental evidence after review of the unchanged full input plus explicit feedback, source-bound original target, frozen sandbox tests and repaired program. Initial failures remain false; semantic_repair records repair success separately. Public repair metadata contains only aggregates and hashes, not private test outputs. These diagnostics do not establish zero-shot student capability or measured skill lift.
Read both configs at this release's actual Hub commit, never a guessed SHA:
from datasets import load_dataset
from huggingface_hub import HfApi
revision = HfApi().repo_info("OpenFormosa/barbet-long-context-sft", repo_type="dataset").sha
print("Pinned revision:", revision)
ds = load_dataset("OpenFormosa/barbet-long-context-sft", "tokenized",
split="train", revision=revision, streaming=True)