tiararodney/cire-corpus
CIRE Corpus The CIRE Corpus is a set of Infrastructure-as-Code template files. The files come from public sources. The templates use these formats: AWS CloudFormation Azure Resource Manager Terraform The related dataset cire-ontology classifies resource types. This corpus stores the infrastructure descriptions that use those types. You can install the corpus as the namespace package tiararodney.cire.corpus.infrastructure. Then you can find the data with importlib.resources.… See the full description on the dataset page: https://huggingface.co/datasets/tiararodney/cire-corpus.
CIRE Corpus
The CIRE Corpus is a set of Infrastructure-as-Code template files. The files come from public sources. The templates use these formats:
- AWS CloudFormation
- Azure Resource Manager
- Terraform
The related dataset cire-ontology classifies resource types. This corpus stores the infrastructure descriptions that use those types.
You can install the corpus as the namespace package tiararodney.cire.corpus.infrastructure. Then you can find the data with importlib.resources. This method is the same as the method for tiararodney.cire.ontology.
The repository contains the dataset and the collection code:
- The dataset is in
data/. - The collection code is in
src/.
The manifest schema and the code that writes the manifest stay in the same version. Users read the corpus only from the manifest. Users must not import the collection code.
A fixed copy of the repository contains the data, the schema, and the source data together.
Layout
schema.json JSON Schema (draft 2020-12) for a manifest record
data/tiararodney/cire/corpus/infrastructure/data/
manifest.jsonl one JSON record per stored object (the contract)
objects/ content-addressed store: objects/<sha256[0:2]>/<sha256>
src/tiararodney/cire/corpus/infrastructure/acquisition/
collection code (discovery, fetch, clean, dedup,
manifest writing); not necessary to use the corpusThe path field in a record is relative to the data package root. The full path is:
files("tiararodney.cire.corpus.infrastructure.data") / record["path"]
Stored objects have no file extension. The format field in the record (json, yaml, hcl) tells the parser the file type. The target_provider field names the IaC dialect.
Manifest contract
manifest.jsonl is the only interface between this dataset and its users. Each record does this:
- The record is valid against
schema.json. - The record contains
schema_version. Users must validate the version when they read the record. Users must not accept a version that they do not support. Users must not guess the meaning of an unknown version. - The
sha256field is the hash of the stored bytes. This hash identifies exact duplicates. - The
normalized_sha256field is the hash after comment removal and whitespace normalization. This hash puts small variants in one group. - The record stores source data (
source.url,source.commit,source.retrieved_at) and the license of the content.
Example record:
{
"schema_version": 1,
"sha256": "9f2c1a7d0b34e6c8f5a2d91e07b3c4a6d8e0f1b2c3d4e5f60718293a4b5c6d7e",
"normalized_sha256": "1e2d3c4b5a69788796a5b4c3d2e1f00f1e2d3c4b5a69788796a5b4c3d2e1f00f",
"size_bytes": 4821,
"target_provider": "cloudformation",
"format": "yaml",
"path": "objects/9f/9f2c1a7d0b34e6c8f5a2d91e07b3c4a6d8e0f1b2c3d4e5f60718293a4b5c6d7e",
"source": {
"url": "https://raw.githubusercontent.com/example/repo/3f9c0e1/stack/web.yaml",
"repository": "https://github.com/example/repo",
"commit": "3f9c0e1",
"path_in_source": "stack/web.yaml",
"retrieved_at": "2026-09-13T00:00:00Z"
},
"license": {
"spdx": "MIT",
"determined_by": "repository_license_file"
},
"scrub": {
"status": "clean"
},
"run_id": "2026-09-13-github-sweep-01"
}Secrets
The process removes credentials and secrets from all content before storage. The repository stores only the cleaned objects. The original bytes that contain secrets are not in the repository.
Each record stores the result in scrub.status:
cleanredacted
Redacted values use known replacement text. A redacted object is not the same byte sequence as the source file.
Licensing
This dataset has two parts. Each part has a different license.
Metadata. manifest.jsonl, schema.json, and this README are original work. These files use the CC-BY-SA-4.0 license.
Stored objects. The template files in objects/ are third-party content from public sources. These files keep their original licenses. This project does not change those licenses. The license.spdx field in each record gives the license that the collection process found. If the process cannot find a license, the field is NOASSERTION.
Users must obey the license of each object in the manifest. If your use needs specified license terms, select records with license.spdx.
