serialhex/Open-Biblical-Dataset
Open Biblical Dataset 11,105 examples. Synthetic instruction-tuning data for a Christian theology / pastoral-counseling fine-tune: a real situated question, a reasoning trace, and a grounded pastoral answer — every answer required to quote and stay inside the bounds of a real source text, never an isolated verse or an appeal to unnamed authority ("many theologians have held..."). Built with data-forge — a custom generation pipeline (question-gen → answer-gen → offline lint pass)… See the full description on the dataset page: https://huggingface.co/datasets/serialhex/Open-Biblical-Dataset.
Open Biblical Dataset
11,105 examples. Synthetic instruction-tuning data for a Christian theology / pastoral-counseling fine-tune: a real situated question, a reasoning trace, and a grounded pastoral answer — every answer required to quote and stay inside the bounds of a real source text, never an isolated verse or an appeal to unnamed authority ("many theologians have held...").
Built with `data-forge` — a custom generation pipeline (question-gen → answer-gen → offline lint pass), not a template-filled or scraped dataset. Every example traces back to a real, public-domain source passage, which is carried alongside the example as context so the grounding is checkable.
What's in the box
Each tier carries a different citation discipline, matching how each source actually functions:
- Tier 0 (scripture) is read at the level of the full passage, never an isolated verse, with any genuinely contested question (grace/election, eschatology) presented as a real, live disagreement rather than resolved on the model's own authority.
- Tier 1 creeds/Didache are cited as what they are — fixed, named communal confessions ("Nicaea, 325, art. 3") — illuminating and binding within their own scope, never scripture's equal.
- Tier 1 patristic witnesses are cited as named INDIVIDUAL thinkers ("Justin Martyr, First Apology") with their own argument — worth taking seriously, but measured against scripture, not assumed correct merely for being early.
- Tier 2 wisdom/logic sources are explicitly subordinate tools — analogical or logical aids that sharpen reasoning but settle nothing theological on their own authority.
Across every tier: no appeal to an unnamed authority ("the church has long held...") is allowed — a source is either named specifically, or the point is made in the model's own voice with no historical framing at all.
Format
One JSON object per line. conversations follows a ShareGPT-style system / human / gpt turn structure; the gpt turn carries the answer in value and its reasoning trace in think (not rendered inline — pulled out as its own field so a fine-tune can map it to whatever thinking-tag convention the target chat template uses).
{
"conversations": [
{"from": "system", "value": "<the standing commitments every answer follows>"},
{"from": "human", "value": "<a realistic, situated question>"},
{"from": "gpt", "value": "<the grounded pastoral answer>",
"think": "<the reasoning trace behind it>"}
],
"context": "<the real source passage this example is grounded in>",
"pericope_ref": "<a human-readable reference, e.g. \"Romans 8:28-30\" or \"Justin Martyr, First Apology, Chapter IV\">",
"source": "<which source corpus this came from, e.g. \"web_bible\", \"creeds_didache\">",
"source_tier": 0,
"variation": "<which register/angle this example was generated under, if any>",
"lint_flags": {"...": "advisory-only structural/phrase flags from the offline lint pass - not yet used to filter"}
}How it was made
- Chunk each source into pericopes — a complete argumentative or narrative unit (a creed clause, a Didache chapter, a Boethius Book section, a Bible pericope), never a fixed-size window.
- Question-gen: a realistic, situated question a real person would actually bring to the material — grief, doubt, a skeptic's challenge, a personal failure — never "explain this passage."
- Answer-gen: single call, producing both the reasoning trace and the final answer together, under a standing set of commitments (reading method, divine-action framing, contested-questions handling, citation discipline) that stays byte-identical across every tier's prompt.
- Offline lint pass (advisory in this release — see Limitations): phrase-based and structural checks for filler, unnamed-authority appeals, and essay-shaped reasoning traces, recorded per-example in
lint_flags.
Generated across a mix of model backends (Gemini/Gemma, Mistral, NVIDIA- hosted models) routed through a graceful-degradation router — no single model's fingerprint dominates the set.
Licensing
- Source texts: every passage behind every example is public domain — the World English Bible, the Nicene/Apostles' Creeds, the Didache and patristic writings (Ante-Nicene Fathers, Roberts-Donaldson translation), and the Tier 2 logic/philosophy texts (Project Gutenberg / pre-1923 scans: Mill, Lewis Carroll, Boethius, Whately, Watts, the Port-Royal Logic, Euclid).
- Generated text (questions, reasoning traces, answers): produced by third-party model APIs (Google/Gemini, Mistral, NVIDIA-hosted models). This dataset's own curation is released CC0; if your use case cares about upstream model-output terms, check the relevant provider's usage policy before redistributing at scale.
Limitations — stated honestly
- Reasoning traces are essay-shaped, not terse. The offline lint pass that flags this (
think_slop_structuralinlint_flags) is advisory only in this release — it was not used to filter or rewrite anything here. A companion dataset applies an actual compression pass (see thegrug-think-inspired rewrite) to fix this; this release is the unmodified base. - Some Tier 2 OCR sources carry period-typesetting noise. Whately, Watts, and the Port-Royal Logic are archive.org scans of 18th/19th- century editions; mechanical cleanup was applied (boilerplate/header stripping, hyphenation rejoining) but individual OCR character errors (a misread "s" as "f," etc.) were not hand-corrected.
- Register-rotated child/teen slices and a safety-escalation slice exist separately and are not included in this release — they're held back pending a dedicated human review pass.
- No held-out eval split. This is training-shaped data as generated; slice your own split before fine-tuning.
