yassinsiouda/minimind-fr-electronics-data
minimind-fr-electronics-data sft_spec_electronics.jsonl (42,537) — conversations schema. theprint/Electronics-QA + electronics.stackexchange.com (accepted answers) + ~25% base-SFT replay; ~30% rows with a diagnostic <think>. Built by scripts/convert_spec_electronics.py; see the minimind-fr-electronics model card. Built from theprint/Electronics-QA bshada/electronics.stackexchange.com allenai/tulu-3-sft-mixture jpacifico/French-Alpaca-dataset-Instruct-110K… See the full description on the dataset page: https://huggingface.co/datasets/yassinsiouda/minimind-fr-electronics-data.
minimind-fr-electronics-data
sft_spec_electronics.jsonl (42,537) — conversations schema. theprint/Electronics-QA + electronics.stackexchange.com (accepted answers) + ~25% base-SFT replay; ~30% rows with a diagnostic <think>. Built by scripts/convert_spec_electronics.py; see the minimind-fr-electronics model card.
Built from
- `theprint/Electronics-QA`
- `bshada/electronics.stackexchange.com`
- `allenai/tulu-3-sft-mixture`
- `jpacifico/French-Alpaca-dataset-Instruct-110K`
- `angeluriot/french_instruct`
- `NousResearch/hermes-function-calling-v1`
- `nvidia/Nemotron-SFT-SWE-v3.5`
Format: line-delimited JSON. SFT rows use MiniMind's SFTDataset schema — {"conversations": [{role, content, reasoning_content, tools, tool_calls}]} (all string fields; tools/tool_calls are JSON strings; reasoning_content renders in <think>…</think>). All SFT files were filtered to 0% zero-supervised rows at the trainer's max_seq_len.
Used to train `yassinsiouda/minimind-fr-electronics`. Training framework: <https://github.com/jingyaogong/minimind>.
License
Apache-2.0 for the derived blend. Upstream dataset licenses apply — c4 (ODC-BY), tulu-3 (ODC-BY), Nemotron-SFT-SWE (CC-BY-4.0), French-PD-Books (public domain), and the others per their cards.
