Team Ai
Datasetpublic

ChrisScr770/tennessee-code

Tennessee Code Unannotated (Titles 29, 34, 36, 37, 39, 40 and 41) Statute text from the Tennessee Code Unannotated, Titles 29, 34, 36, 37, 39, 40 and 41. The text is split into 9,050 retrieval-sized chunks. Each chunk keeps its place in the section, its subsection labels, its version qualifiers and the cross-references in its text, resolved to other chunks where possible. The chunk text is Markdown, so the cite, the heading and every subsection label are part of the text itself.… See the full description on the dataset page: https://huggingface.co/datasets/ChrisScr770/tennessee-code.

sourceHugging Facecc0-1.0updated 2d agoView on Hugging Face
1likes78downloads
Dataset Card

Tennessee Code Unannotated (Titles 29, 34, 36, 37, 39, 40 and 41)

Statute text from the Tennessee Code Unannotated, Titles 29, 34, 36, 37, 39, 40 and 41. The text is split into 9,050 retrieval-sized chunks. Each chunk keeps its place in the section, its subsection labels, its version qualifiers and the cross-references in its text, resolved to other chunks where possible. The chunk text is Markdown, so the cite, the heading and every subsection label are part of the text itself. An optional subset adds a openai/text-embedding-3-large vector for every chunk.

Not an official publication, and not legal advice. This is an unofficial, machine-converted copy. It is not affiliated with or endorsed by LexisNexis, the Tennessee Code Commission or the State of Tennessee. Statutes change every legislative session, so check anything you rely on against the official source.

Snapshot

CurrencyCurrent through the 2026 Regular Session and the 2026 2nd Extraordinary Session.
Publisher datestamp2026-09-30
Bare citations resolved as of2026-10-07
Exported2026-10-08

This is a point-in-time copy. Later sessions of the General Assembly will appear as new revisions of this dataset; older revisions stay available in the repository history.

Coverage

This dataset contains only these 7 titles. If a section is missing here, it may still exist in a title this dataset does not cover.

TitleNameChaptersDocumentsChunks
29Remedies and Special Proceedings438251,696
34Guardianship8146304
36Domestic Relations85161,182
37Juveniles105001,130
39Criminal Offenses87461,665
40Criminal Procedure399402,044
41Correctional Institutions and Inmates205051,029
Total1364,1789,050

Documents counts everything Lexis publishes as a document under a title: sections, plus reserved and repealed placeholders and part notes. The chunks contain 1,304,092 words in total.

Family law is largely self-contained here. Of the 1,811 citations in Titles 34, 36 and 37 (Guardianship, Domestic Relations and Juveniles), 1,519 (84%) resolve to a section in this dataset. Another 262 cite titles it does not cover; the most-cited are Title 68, health, safety and environment (45 citations); Title 49, education (29); Title 33, mental health and disabilities (23); and Title 8, public officers and employees (20). The other 30 name a section number these titles do not publish.

Quick start

python
from datasets import load_dataset

chunks = load_dataset("ChrisScr770/tennessee-code", "chunks", split="train")
edges = load_dataset("ChrisScr770/tennessee-code", "crossrefs", split="train")

# Reassemble one section: its chunks, in order, joined by a newline, are the
# section's Markdown exactly.
section = chunks.filter(lambda r: r["section_id"] == "37-1-102").sort("chunk_index")
print("\n".join(section["text"]))

Files

PathRowsContents
data/chunks/title-NN.parquet9,050One row per chunk, one file per title
data/crossrefs.parquet6,127One row per cross-reference edge
data/embeddings.parquet9,050One vector per chunk, with the exact string embedded
export_manifest.jsonCounts and SHA-256 of every data file

Chunks

A section is one chunk unless it is too long. Long sections split at subsection boundaries. A section's History block (the session laws and earlier codifications behind it) is always a chunk of its own with kind: history. Nested subsections are flattened into a label path, written the way Tennessee cites them: **(b)(1)(A)** text.

The text of every chunk is Markdown, the format language models read and write most fluently. A section's first chunk opens with its cite as a # heading, the currency line in italics and the section heading as ##; the chunks after it continue the text. Each subsection starts its own paragraph, with its full label path in bold. A model or a person can follow the hierarchy from the text alone, without the structure fields, and the text renders as-is in any Markdown viewer. If you need plain text, remove the ** markers. The opening of tn:37-1-101:0001:

markdown
# Tenn. Code Ann. § 37-1-101

*Current through the 2026 Regular Session and the 2026 2nd Extraordinary Session.*

## 37-1-101. Purpose — Jurisdiction — Ensuring compliance with the Indian Child Welfare Act.

**(a)** This part shall be construed to effectuate the following public purposes:

**(a)(1)** Provide for the care, protection, and wholesome moral, mental and physical development of children coming within its provisions;
FieldMeaning
idtn:<section_id>:<NNNN>; unique
section_idThe document's identity: the cite, then . and meta.version_qualifier when that is set, or .vN (see Versions); node-… for a document with no cite. Added by the export; it is the middle part of id
citeThe citation as published, such as 37-1-102; empty for documents that have none. Not unique
cite_endThe last cite of a section_range document; otherwise empty
corpus_titleThe title file the row ships in; never null. Added by the export
title, chapter, partParsed numbers. Null where the document has none: reserved-chapter and reserved-part notes have no title number, and a chapter without parts has no part
kindSee the table below
headingThe heading as published: 37-1-101. Purpose — Jurisdiction — …
label_pathLabel of the first subsection in the chunk, such as (b)(5)(E); empty when the chunk has none
section_pathBreadcrumb: title, chapter, part, heading, label
textMarkdown. A section's first chunk starts with the cite line and the currency line; the chunks after it continue the text directly
source_url, anchorLink to the document on Lexis' public-access site, with the fragment of the chunk's first subsection
anchorsEvery subsection label in the chunk with its fragment. The link for one is source_url up to #, then # + anchor
content_urn, node_pathLexis' document id and table-of-contents path
orderPosition in the title's reading order
chunk_index, chunks_in_document1-based position within the section, and that section's chunk count
prev_chunk_id, next_chunk_idNeighbouring chunks within the same section
document_prev_cite, document_next_citeNeighbouring documents' cites; set only on a section's first chunk
content_sha256SHA-256 of the whole section's body, shared by all of its chunks. It changes only when the text changes
words, charsSize of text
meta.catchlineThe heading without its cite
meta.document_titleTenn. Code Ann. § 37-1-101
meta.notesBracketed editorial notes from the heading: Reserved, Repealed, Effective on …
meta.effective_noteThe Effective … note that dates this version, as published
meta.version_qualifieruntil-…/on-… dates, a contingency slug, reserved or repealed. Set on every placeholder and every dated or contingent version, even one with no twin; empty otherwise, including for .vN documents
meta.currency, meta.code_datestampThe snapshot; the same on every row
meta.subsectionsNumber of labelled subsections in this chunk
cross_referencesCitations in this chunk's text, in the same shape as the crossrefs config without the source fields
`kind`RowsMeaning
section4,906Operative text of a section, or one piece of a section too long for one chunk
history4,089A section's History block: the session laws and former codifications behind it
section_range22One document covering a run of cites (cite to cite_end), usually reserved or repealed
reserved_part14Placeholder for a reserved part (no cite, no title number)
reserved_chapter8Placeholder for a reserved chapter (no cite, no title number)
part_note6Editorial note attached to a part
section_list4One document covering the cites listed in its heading, usually reserved
subpart_note1Editorial note attached to a subpart

Versions: a cite is not an identity

18 citations are published as more than one document at the same time: 29-20-401, 29-39-102, 29-39-103, 36-6-413, 37-1-187, 37-3-803, 37-5-601, 37-5-604, 39-13-113, 39-13-204, 39-13-207, 39-13-208, 39-14-141, 39-14-142, 39-17-1330, 40-11-115, 41-21-228, 41-22-112. Key on section_id, not on cite. The qualifier says what separates the versions:

  • —`until-YYYY-MM-DD` / `on-YYYY-MM-DD`: an amendment with a delayed effective date. Before that date the until- text is the law; from that date on, the on- text is. Only the first sentence of meta.effective_note dates the version. The second sentence points to the other version.
  • —Contingency qualifiers such as effective-when-contingency-is-met: the text takes effect on an event, not a date, so the qualifier is a slug of the note.
  • —`reserved`, `repealed`: a placeholder with no operative text. Every placeholder carries one, including the few that share a cite with live text, such as 39-14-142 and 39-14-142.reserved.
  • —Inferred qualifiers: sometimes Lexis dates only one of two twins. The unmarked twin then gets until-D from its sibling's on-D note, and its meta.effective_note stays empty because the source itself says nothing.
  • —`.vN` (29-20-401.v2, 39-13-113.v2): documents that share a cite and heading with nothing in the source to tell them apart. The numbering follows the source's reading order and was assigned by this dataset, not by Lexis, and meta.version_qualifier is empty for them. Read each version's heading, notes and History to tell them apart.

Cross-references

The Unannotated edition has no links, so references are parsed from the text. Three kinds of citation are skipped: those in History blocks (they cite session laws and former numbering), the document's own title line, and federal cites. A mention that names several sections, such as §§ A, B, and C, yields one edge per cite.

There are 6,127 edges. 5,179 resolve to a chunk in this dataset; target_chunk_id is the first chunk of the target section, and subdivision holds the subsection when the citation names one. 948 are unresolved: 885 cite titles this dataset does not cover, and 63 cite a section number within these titles that this snapshot does not publish. A bare cite of a section that has several versions resolves to the version in force on 2026-10-07; for a .vN pair, which no date can separate, it resolves to the first.

`crossrefs` fieldMeaning
source_chunk_id, source_cite, source_titleWhere the citation appears
textThe citation as written; a list of several cites appears in full on each of its edges
target_cite, target_chunk_id, resolvedWhat it points at; target_chunk_id is empty when the edge is unresolved
is_range, all_citesWhether the cite is an endpoint of a range (§§ A — B), and every cite the mention names
subdivisionThe cited subsection, such as (a)(3), when there is one
block_labelThe subsection of the source chunk where the citation appears

Embeddings

The embeddings subset holds one vector per chunk from openai/text-embedding-3-large: 3,072 float32 values, requested through OpenRouter pinned to the azure provider with fallbacks off, between 2026-10-06 and 2026-10-07. The vectors are unit length, so a dot product is their cosine similarity. Uncompressed they come to about 111 MB, so they live in their own subset: loading chunks does not download them.

Each vector embeds embed_text, not text. embed_text opens with the document title and the chunk's breadcrumb, followed by the chunk body without its cite, currency and heading lines. A History chunk's breadcrumb includes its section's heading, so search can find it on its own. The first line of embed_text for tn:37-1-101:0001:

text
[Tenn. Code Ann. § 37-1-101 | Title 37 Juveniles > Chapter 1 Juvenile Courts and Proceedings > Part 1 General Provisions > 37-1-101. Purpose — Jurisdiction — Ensuring compliance with the Indian Child Welfare Act. > (a)]
FieldMeaning
idThe chunk's id; join to the chunks subset on it
embed_modelThe model that produced the vector
embed_textThe exact string that was embedded
text_sha256SHA-256 of embed_text
embedding3,072 float32 values

To search, embed the question with the same model and rank by dot product. Through OpenRouter, ask for openai/text-embedding-3-large with the same provider pin; through OpenAI directly, the model is text-embedding-3-large. Vectors from different models cannot be compared, so to use another model, embed embed_text with it.

python
import os
import numpy as np
from datasets import load_dataset
from openai import OpenAI

emb = load_dataset("ChrisScr770/tennessee-code", "embeddings", split="train").with_format("numpy")[:]
vectors, ids = emb["embedding"], emb["id"]  # a (rows, dims) float32 array and the chunk ids
client = OpenAI(base_url="https://openrouter.ai/api/v1", api_key=os.environ["OPENROUTER_API_KEY"])

def search(question, k=5):
    response = client.embeddings.create(
        model="openai/text-embedding-3-large",
        input=question,
        extra_body={"provider": {"order": ["azure"], "allow_fallbacks": False}},
    )
    scores = vectors @ np.asarray(response.data[0].embedding, dtype=np.float32)
    best = np.argsort(-scores)[:k]
    return [(str(ids[i]), round(float(scores[i]), 3)) for i in best]

A weak best match. In the author's measurements on this index, questions these titles answer had a best match above 0.45, and questions about Tennessee law outside them scored below it. The gap between the two groups was narrow, so treat a best score under 0.45 as a reason to check whether the question is covered at all, not as a filter.

Using it from an agent

The dataset is laid out for an agent that answers questions about Tennessee law and has to cite what it used. The pattern:

  1. 1.Find candidates. Search the embeddings subset, or index text with your own model or BM25; the chunks are already sized for retrieval. Keep each chunk's id with it, and note a weak best match.
  2. 2.Read the whole section. A chunk is part of a section. Before relying on one, read every chunk with the same section_id, in chunk_index order.
  3. 3.Check which version is the law. A cite can have several documents. List them and use the qualifier to pick the one in force on the date that matters. A contingency has no date, so say that instead of choosing.
  4. 4.Follow the references. Outbound edges lead to the sections this one relies on, such as its definitions; inbound edges show what depends on it. An unresolved edge points outside this dataset: report it as not covered, never as not existing.
  5. 5.Cite precisely. Give meta.document_title, the section_id, the currency line and source_url.
python
import datetime as dt
from datasets import load_dataset

chunks = load_dataset("ChrisScr770/tennessee-code", "chunks", split="train").to_pandas()
edges = load_dataset("ChrisScr770/tennessee-code", "crossrefs", split="train").to_pandas()
firsts = chunks[chunks.chunk_index == 1]  # one row per document

def read_section(section_id):
    """Every chunk of one document, in order: its Markdown exactly."""
    rows = chunks[chunks.section_id == section_id].sort_values("chunk_index")
    return "\n".join(rows.text)

def versions(cite):
    """Each document published under a cite, with what tells them apart."""
    rows = firsts[firsts.cite == cite]
    return [(r.section_id, r.meta["version_qualifier"], r.meta["effective_note"])
            for r in rows.itertuples()]

def in_force(section_id, on=None):
    """True or False where the qualifier decides; None where a person has to."""
    row = firsts[firsts.section_id == section_id].iloc[0]
    qualifier = row.meta["version_qualifier"]
    day = (on or dt.date.today()).isoformat()
    if qualifier.startswith("until-"):
        return day < qualifier.removeprefix("until-")
    if qualifier.startswith("on-"):
        return day >= qualifier.removeprefix("on-")
    if qualifier in ("reserved", "repealed"):
        return False
    if qualifier:  # a contingency: the text takes effect on an event, not a date
        return None
    undated = [v for v in versions(row.cite) if not v[1]]
    return True if len(undated) == 1 else None  # .vN twins: read both

def references(section_id):
    """Outbound: what this section cites. Inbound: what cites it."""
    ids = chunks.id[chunks.section_id == section_id]
    cite = firsts.cite[firsts.section_id == section_id].iloc[0]
    outbound = edges[edges.source_chunk_id.isin(ids)]
    inbound = edges[(edges.target_cite == cite) & ~edges.source_chunk_id.isin(ids)]
    return outbound, inbound

for section_id, qualifier, note in versions("37-3-803"):
    print(section_id, in_force(section_id))

The dataset's author serves this same pattern to agents as MCP tools (search, read a section, follow its references); the server itself is not published.

Source and method

The text was converted from the free public-access Tennessee Code that LexisNexis publishes. Documents were fetched one at a time, at a throttled rate and within the site's per-session limits. Its access checks were completed by a person and never bypassed. The HTML was converted to Markdown and then chunked. For every title, the walk of the table of contents matched the node and document counts Lexis publishes, and a strict verification pass ran clean before export. The raw HTML is not included.

Known limitations

  • —Partial coverage: 7 titles of the Code. Case law, annotations and every other title are out of scope.
  • —A snapshot: Current through the 2026 Regular Session and the 2026 2nd Extraordinary Session.
  • —History chunks carry no section context in their text. Join them to their section on section_id. Their embed_text names the section, so search still finds them.
  • —Unresolved cross-references: 948 of the 6,127 edges.
  • —Lexis links depend on Lexis' URL scheme and site terms, and may stop working.

License and copyright

The source documents carry the notice “TENNESSEE CODE ANNOTATED Copyright © 2026 by The State of Tennessee All rights reserved”. Under the government edicts doctrine, the text of the law itself is not subject to copyright in the United States (Georgia v. Public.Resource.Org, Inc., 590 U.S. 255 (2020)). This dataset's own work, meaning the chunking, structure fields and cross-reference graph, is released under CC0 1.0.

Citation

bibtex
@misc{tennessee_code_unannotated_2026,
  title        = {Tennessee Code Unannotated (Titles 29, 34, 36, 37, 39, 40 and 41)},
  howpublished = {\url{https://huggingface.co/datasets/ChrisScr770/tennessee-code}},
  note         = {Current through the 2026 Regular Session and the 2026 2nd Extraordinary Session.},
  year         = {2026}
}