Team Ai
Datasetpublic

serialhex/Open-Biblical-Dataset

Open Biblical Dataset 11,105 examples. Synthetic instruction-tuning data for a Christian theology / pastoral-counseling fine-tune: a real situated question, a reasoning trace, and a grounded pastoral answer — every answer required to quote and stay inside the bounds of a real source text, never an isolated verse or an appeal to unnamed authority ("many theologians have held..."). Built with data-forge — a custom generation pipeline (question-gen → answer-gen → offline lint pass)… See the full description on the dataset page: https://huggingface.co/datasets/serialhex/Open-Biblical-Dataset.

sourceHugging Facecc0-1.0updated 8d agoView on Hugging Face
0likes62downloads
Dataset Card

Open Biblical Dataset

11,105 examples. Synthetic instruction-tuning data for a Christian theology / pastoral-counseling fine-tune: a real situated question, a reasoning trace, and a grounded pastoral answer — every answer required to quote and stay inside the bounds of a real source text, never an isolated verse or an appeal to unnamed authority ("many theologians have held...").

Built with `data-forge` — a custom generation pipeline (question-gen → answer-gen → offline lint pass), not a template-filled or scraped dataset. Every example traces back to a real, public-domain source passage, which is carried alongside the example as context so the grounding is checkable.

What's in the box

TierSourceExamples
0 — Norming normWEB Bible (single-pericope)2,952
0 — Norming normWEB Bible (paired-passage)5,000
1 — Normed normsNicene/Apostles' Creeds + Didache216
1 — Normed normsPatristic witnesses (Justin Martyr, Ignatius, Clement of Rome, Polycarp, Barnabas, Clement of Alexandria, Shepherd of Hermas, et al.)999
2 — Wisdom/illuminationLogic & philosophy (Mill, Lewis Carroll, Boethius, Whately, Watts, Port-Royal Logic, Euclid)1,938

Each tier carries a different citation discipline, matching how each source actually functions:

  • —Tier 0 (scripture) is read at the level of the full passage, never an isolated verse, with any genuinely contested question (grace/election, eschatology) presented as a real, live disagreement rather than resolved on the model's own authority.
  • —Tier 1 creeds/Didache are cited as what they are — fixed, named communal confessions ("Nicaea, 325, art. 3") — illuminating and binding within their own scope, never scripture's equal.
  • —Tier 1 patristic witnesses are cited as named INDIVIDUAL thinkers ("Justin Martyr, First Apology") with their own argument — worth taking seriously, but measured against scripture, not assumed correct merely for being early.
  • —Tier 2 wisdom/logic sources are explicitly subordinate tools — analogical or logical aids that sharpen reasoning but settle nothing theological on their own authority.

Across every tier: no appeal to an unnamed authority ("the church has long held...") is allowed — a source is either named specifically, or the point is made in the model's own voice with no historical framing at all.

Format

One JSON object per line. conversations follows a ShareGPT-style system / human / gpt turn structure; the gpt turn carries the answer in value and its reasoning trace in think (not rendered inline — pulled out as its own field so a fine-tune can map it to whatever thinking-tag convention the target chat template uses).

json
{
  "conversations": [
    {"from": "system", "value": "<the standing commitments every answer follows>"},
    {"from": "human",  "value": "<a realistic, situated question>"},
    {"from": "gpt",    "value": "<the grounded pastoral answer>",
                        "think": "<the reasoning trace behind it>"}
  ],
  "context": "<the real source passage this example is grounded in>",
  "pericope_ref": "<a human-readable reference, e.g. \"Romans 8:28-30\" or \"Justin Martyr, First Apology, Chapter IV\">",
  "source": "<which source corpus this came from, e.g. \"web_bible\", \"creeds_didache\">",
  "source_tier": 0,
  "variation": "<which register/angle this example was generated under, if any>",
  "lint_flags": {"...": "advisory-only structural/phrase flags from the offline lint pass - not yet used to filter"}
}

How it was made

  1. 1.Chunk each source into pericopes — a complete argumentative or narrative unit (a creed clause, a Didache chapter, a Boethius Book section, a Bible pericope), never a fixed-size window.
  2. 2.Question-gen: a realistic, situated question a real person would actually bring to the material — grief, doubt, a skeptic's challenge, a personal failure — never "explain this passage."
  3. 3.Answer-gen: single call, producing both the reasoning trace and the final answer together, under a standing set of commitments (reading method, divine-action framing, contested-questions handling, citation discipline) that stays byte-identical across every tier's prompt.
  4. 4.Offline lint pass (advisory in this release — see Limitations): phrase-based and structural checks for filler, unnamed-authority appeals, and essay-shaped reasoning traces, recorded per-example in lint_flags.

Generated across a mix of model backends (Gemini/Gemma, Mistral, NVIDIA- hosted models) routed through a graceful-degradation router — no single model's fingerprint dominates the set.

Licensing

  • —Source texts: every passage behind every example is public domain — the World English Bible, the Nicene/Apostles' Creeds, the Didache and patristic writings (Ante-Nicene Fathers, Roberts-Donaldson translation), and the Tier 2 logic/philosophy texts (Project Gutenberg / pre-1923 scans: Mill, Lewis Carroll, Boethius, Whately, Watts, the Port-Royal Logic, Euclid).
  • —Generated text (questions, reasoning traces, answers): produced by third-party model APIs (Google/Gemini, Mistral, NVIDIA-hosted models). This dataset's own curation is released CC0; if your use case cares about upstream model-output terms, check the relevant provider's usage policy before redistributing at scale.

Limitations — stated honestly

  • —Reasoning traces are essay-shaped, not terse. The offline lint pass that flags this (think_slop_structural in lint_flags) is advisory only in this release — it was not used to filter or rewrite anything here. A companion dataset applies an actual compression pass (see the grug-think-inspired rewrite) to fix this; this release is the unmodified base.
  • —Some Tier 2 OCR sources carry period-typesetting noise. Whately, Watts, and the Port-Royal Logic are archive.org scans of 18th/19th- century editions; mechanical cleanup was applied (boilerplate/header stripping, hyphenation rejoining) but individual OCR character errors (a misread "s" as "f," etc.) were not hand-corrected.
  • —Register-rotated child/teen slices and a safety-escalation slice exist separately and are not included in this release — they're held back pending a dedicated human review pass.
  • —No held-out eval split. This is training-shaped data as generated; slice your own split before fine-tuning.