playforgecoding/spelling-creator-document-import
Spelling Creator document import One section of a lesson as someone might have typed it up, paired with that section as lesson JSON. Spelling Creator is a lesson builder for Spelling to Communicate (S2C), used with nonspeaking spellers. Its Import from text reads a typed lesson with rules, and hands the sections the rules cannot read to a small on-device model; this dataset trains that model. What is in it Chat-format JSONL, ready for a supervised fine-tune (TRL's… See the full description on the dataset page: https://huggingface.co/datasets/playforgecoding/spelling-creator-document-import.
Spelling Creator document import
One section of a lesson as someone might have typed it up, paired with that section as lesson JSON. Spelling Creator is a lesson builder for Spelling to Communicate (S2C), used with nonspeaking spellers. Its Import from text reads a typed lesson with rules, and hands the sections the rules cannot read to a small on-device model; this dataset trains that model.
What is in it
Chat-format JSONL, ready for a supervised fine-tune (TRL's SFTTrainer takes it as is):
- system: the lesson section's JSON schema and a guide to the question types.
- user: one section of a document, cut and laid out exactly as the import hands it to the model.
- assistant: the section as JSON:
name,paragraphs,spellingWords, andquestions, each withprompt,type,answersandsteps.
Every lesson is rendered in nine layouts (style on each row): the app's own Word export read back as raw text, six typed-up layouts the rules read (plain, qa, caps, bullets, colon, worked), and two they cannot (nomarks: no question marks or numbers, the answer tacked on; runon: a numbered list that lost its line breaks). From those two, a section is in only if the import would really send it to the model. Targets follow the document as written: no section name where it shows no heading, and prompts without question marks where the layout dropped them.
How it was made
Rendered by code from lessons published on the Spelling Creator hub, so every target is the lesson's own content, not a model's reading of it. The renderers, the splitter and the scorer are in the Spelling Creator repository under packages/core/scripts/extract-eval/. The model trained on it is LFM2-1.2B-Extract-lesson.
Attribution
CC BY 4.0. The lessons, by their authors on spellingcreator.org:
