ChrisScr770/tennessee-code
Tennessee Code Unannotated (Titles 29, 34, 36, 37, 39, 40 and 41) Statute text from the Tennessee Code Unannotated, Titles 29, 34, 36, 37, 39, 40 and 41. The text is split into 9,050 retrieval-sized chunks. Each chunk keeps its place in the section, its subsection labels, its version qualifiers and the cross-references in its text, resolved to other chunks where possible. The chunk text is Markdown, so the cite, the heading and every subsection label are part of the text itself.… See the full description on the dataset page: https://huggingface.co/datasets/ChrisScr770/tennessee-code.
Tennessee Code Unannotated (Titles 29, 34, 36, 37, 39, 40 and 41)
Statute text from the Tennessee Code Unannotated, Titles 29, 34, 36, 37, 39, 40 and 41. The text is split into 9,050 retrieval-sized chunks. Each chunk keeps its place in the section, its subsection labels, its version qualifiers and the cross-references in its text, resolved to other chunks where possible. The chunk text is Markdown, so the cite, the heading and every subsection label are part of the text itself. An optional subset adds a openai/text-embedding-3-large vector for every chunk.
Not an official publication, and not legal advice. This is an unofficial, machine-converted copy. It is not affiliated with or endorsed by LexisNexis, the Tennessee Code Commission or the State of Tennessee. Statutes change every legislative session, so check anything you rely on against the official source.
Snapshot
This is a point-in-time copy. Later sessions of the General Assembly will appear as new revisions of this dataset; older revisions stay available in the repository history.
Coverage
This dataset contains only these 7 titles. If a section is missing here, it may still exist in a title this dataset does not cover.
Documents counts everything Lexis publishes as a document under a title: sections, plus reserved and repealed placeholders and part notes. The chunks contain 1,304,092 words in total.
Family law is largely self-contained here. Of the 1,811 citations in Titles 34, 36 and 37 (Guardianship, Domestic Relations and Juveniles), 1,519 (84%) resolve to a section in this dataset. Another 262 cite titles it does not cover; the most-cited are Title 68, health, safety and environment (45 citations); Title 49, education (29); Title 33, mental health and disabilities (23); and Title 8, public officers and employees (20). The other 30 name a section number these titles do not publish.
Quick start
from datasets import load_dataset
chunks = load_dataset("ChrisScr770/tennessee-code", "chunks", split="train")
edges = load_dataset("ChrisScr770/tennessee-code", "crossrefs", split="train")
# Reassemble one section: its chunks, in order, joined by a newline, are the
# section's Markdown exactly.
section = chunks.filter(lambda r: r["section_id"] == "37-1-102").sort("chunk_index")
print("\n".join(section["text"]))Files
Chunks
A section is one chunk unless it is too long. Long sections split at subsection boundaries. A section's History block (the session laws and earlier codifications behind it) is always a chunk of its own with kind: history. Nested subsections are flattened into a label path, written the way Tennessee cites them: **(b)(1)(A)** text.
The text of every chunk is Markdown, the format language models read and write most fluently. A section's first chunk opens with its cite as a # heading, the currency line in italics and the section heading as ##; the chunks after it continue the text. Each subsection starts its own paragraph, with its full label path in bold. A model or a person can follow the hierarchy from the text alone, without the structure fields, and the text renders as-is in any Markdown viewer. If you need plain text, remove the ** markers. The opening of tn:37-1-101:0001:
# Tenn. Code Ann. § 37-1-101
*Current through the 2026 Regular Session and the 2026 2nd Extraordinary Session.*
## 37-1-101. Purpose — Jurisdiction — Ensuring compliance with the Indian Child Welfare Act.
**(a)** This part shall be construed to effectuate the following public purposes:
**(a)(1)** Provide for the care, protection, and wholesome moral, mental and physical development of children coming within its provisions;Versions: a cite is not an identity
18 citations are published as more than one document at the same time: 29-20-401, 29-39-102, 29-39-103, 36-6-413, 37-1-187, 37-3-803, 37-5-601, 37-5-604, 39-13-113, 39-13-204, 39-13-207, 39-13-208, 39-14-141, 39-14-142, 39-17-1330, 40-11-115, 41-21-228, 41-22-112. Key on section_id, not on cite. The qualifier says what separates the versions:
- `until-YYYY-MM-DD` / `on-YYYY-MM-DD`: an amendment with a delayed effective date. Before that date the
until-text is the law; from that date on, theon-text is. Only the first sentence ofmeta.effective_notedates the version. The second sentence points to the other version. - Contingency qualifiers such as
effective-when-contingency-is-met: the text takes effect on an event, not a date, so the qualifier is a slug of the note. - `reserved`, `repealed`: a placeholder with no operative text. Every placeholder carries one, including the few that share a cite with live text, such as
39-14-142and39-14-142.reserved. - Inferred qualifiers: sometimes Lexis dates only one of two twins. The unmarked twin then gets
until-Dfrom its sibling'son-Dnote, and itsmeta.effective_notestays empty because the source itself says nothing. - `.vN` (
29-20-401.v2,39-13-113.v2): documents that share a cite and heading with nothing in the source to tell them apart. The numbering follows the source's reading order and was assigned by this dataset, not by Lexis, andmeta.version_qualifieris empty for them. Read each version's heading, notes and History to tell them apart.
Cross-references
The Unannotated edition has no links, so references are parsed from the text. Three kinds of citation are skipped: those in History blocks (they cite session laws and former numbering), the document's own title line, and federal cites. A mention that names several sections, such as §§ A, B, and C, yields one edge per cite.
There are 6,127 edges. 5,179 resolve to a chunk in this dataset; target_chunk_id is the first chunk of the target section, and subdivision holds the subsection when the citation names one. 948 are unresolved: 885 cite titles this dataset does not cover, and 63 cite a section number within these titles that this snapshot does not publish. A bare cite of a section that has several versions resolves to the version in force on 2026-10-07; for a .vN pair, which no date can separate, it resolves to the first.
Embeddings
The embeddings subset holds one vector per chunk from openai/text-embedding-3-large: 3,072 float32 values, requested through OpenRouter pinned to the azure provider with fallbacks off, between 2026-10-06 and 2026-10-07. The vectors are unit length, so a dot product is their cosine similarity. Uncompressed they come to about 111 MB, so they live in their own subset: loading chunks does not download them.
Each vector embeds embed_text, not text. embed_text opens with the document title and the chunk's breadcrumb, followed by the chunk body without its cite, currency and heading lines. A History chunk's breadcrumb includes its section's heading, so search can find it on its own. The first line of embed_text for tn:37-1-101:0001:
[Tenn. Code Ann. § 37-1-101 | Title 37 Juveniles > Chapter 1 Juvenile Courts and Proceedings > Part 1 General Provisions > 37-1-101. Purpose — Jurisdiction — Ensuring compliance with the Indian Child Welfare Act. > (a)]To search, embed the question with the same model and rank by dot product. Through OpenRouter, ask for openai/text-embedding-3-large with the same provider pin; through OpenAI directly, the model is text-embedding-3-large. Vectors from different models cannot be compared, so to use another model, embed embed_text with it.
import os
import numpy as np
from datasets import load_dataset
from openai import OpenAI
emb = load_dataset("ChrisScr770/tennessee-code", "embeddings", split="train").with_format("numpy")[:]
vectors, ids = emb["embedding"], emb["id"] # a (rows, dims) float32 array and the chunk ids
client = OpenAI(base_url="https://openrouter.ai/api/v1", api_key=os.environ["OPENROUTER_API_KEY"])
def search(question, k=5):
response = client.embeddings.create(
model="openai/text-embedding-3-large",
input=question,
extra_body={"provider": {"order": ["azure"], "allow_fallbacks": False}},
)
scores = vectors @ np.asarray(response.data[0].embedding, dtype=np.float32)
best = np.argsort(-scores)[:k]
return [(str(ids[i]), round(float(scores[i]), 3)) for i in best]A weak best match. In the author's measurements on this index, questions these titles answer had a best match above 0.45, and questions about Tennessee law outside them scored below it. The gap between the two groups was narrow, so treat a best score under 0.45 as a reason to check whether the question is covered at all, not as a filter.
Using it from an agent
The dataset is laid out for an agent that answers questions about Tennessee law and has to cite what it used. The pattern:
- Find candidates. Search the
embeddingssubset, or indextextwith your own model or BM25; the chunks are already sized for retrieval. Keep each chunk'sidwith it, and note a weak best match. - Read the whole section. A chunk is part of a section. Before relying on one, read every chunk with the same
section_id, inchunk_indexorder. - Check which version is the law. A cite can have several documents. List them and use the qualifier to pick the one in force on the date that matters. A contingency has no date, so say that instead of choosing.
- Follow the references. Outbound edges lead to the sections this one relies on, such as its definitions; inbound edges show what depends on it. An unresolved edge points outside this dataset: report it as not covered, never as not existing.
- Cite precisely. Give
meta.document_title, thesection_id, the currency line andsource_url.
import datetime as dt
from datasets import load_dataset
chunks = load_dataset("ChrisScr770/tennessee-code", "chunks", split="train").to_pandas()
edges = load_dataset("ChrisScr770/tennessee-code", "crossrefs", split="train").to_pandas()
firsts = chunks[chunks.chunk_index == 1] # one row per document
def read_section(section_id):
"""Every chunk of one document, in order: its Markdown exactly."""
rows = chunks[chunks.section_id == section_id].sort_values("chunk_index")
return "\n".join(rows.text)
def versions(cite):
"""Each document published under a cite, with what tells them apart."""
rows = firsts[firsts.cite == cite]
return [(r.section_id, r.meta["version_qualifier"], r.meta["effective_note"])
for r in rows.itertuples()]
def in_force(section_id, on=None):
"""True or False where the qualifier decides; None where a person has to."""
row = firsts[firsts.section_id == section_id].iloc[0]
qualifier = row.meta["version_qualifier"]
day = (on or dt.date.today()).isoformat()
if qualifier.startswith("until-"):
return day < qualifier.removeprefix("until-")
if qualifier.startswith("on-"):
return day >= qualifier.removeprefix("on-")
if qualifier in ("reserved", "repealed"):
return False
if qualifier: # a contingency: the text takes effect on an event, not a date
return None
undated = [v for v in versions(row.cite) if not v[1]]
return True if len(undated) == 1 else None # .vN twins: read both
def references(section_id):
"""Outbound: what this section cites. Inbound: what cites it."""
ids = chunks.id[chunks.section_id == section_id]
cite = firsts.cite[firsts.section_id == section_id].iloc[0]
outbound = edges[edges.source_chunk_id.isin(ids)]
inbound = edges[(edges.target_cite == cite) & ~edges.source_chunk_id.isin(ids)]
return outbound, inbound
for section_id, qualifier, note in versions("37-3-803"):
print(section_id, in_force(section_id))The dataset's author serves this same pattern to agents as MCP tools (search, read a section, follow its references); the server itself is not published.
Source and method
The text was converted from the free public-access Tennessee Code that LexisNexis publishes. Documents were fetched one at a time, at a throttled rate and within the site's per-session limits. Its access checks were completed by a person and never bypassed. The HTML was converted to Markdown and then chunked. For every title, the walk of the table of contents matched the node and document counts Lexis publishes, and a strict verification pass ran clean before export. The raw HTML is not included.
Known limitations
- Partial coverage: 7 titles of the Code. Case law, annotations and every other title are out of scope.
- A snapshot: Current through the 2026 Regular Session and the 2026 2nd Extraordinary Session.
- History chunks carry no section context in their
text. Join them to their section onsection_id. Theirembed_textnames the section, so search still finds them. - Unresolved cross-references: 948 of the 6,127 edges.
- Lexis links depend on Lexis' URL scheme and site terms, and may stop working.
License and copyright
The source documents carry the notice “TENNESSEE CODE ANNOTATED Copyright © 2026 by The State of Tennessee All rights reserved”. Under the government edicts doctrine, the text of the law itself is not subject to copyright in the United States (Georgia v. Public.Resource.Org, Inc., 590 U.S. 255 (2020)). This dataset's own work, meaning the chunking, structure fields and cross-reference graph, is released under CC0 1.0.
Citation
@misc{tennessee_code_unannotated_2026,
title = {Tennessee Code Unannotated (Titles 29, 34, 36, 37, 39, 40 and 41)},
howpublished = {\url{https://huggingface.co/datasets/ChrisScr770/tennessee-code}},
note = {Current through the 2026 Regular Session and the 2026 2nd Extraordinary Session.},
year = {2026}
}