Team Ai
Datasetpublic

tiararodney/cire-corpus

CIRE Corpus The CIRE Corpus is a set of Infrastructure-as-Code template files. The files come from public sources. The templates use these formats: AWS CloudFormation Azure Resource Manager Terraform The related dataset cire-ontology classifies resource types. This corpus stores the infrastructure descriptions that use those types. You can install the corpus as the namespace package tiararodney.cire.corpus.infrastructure. Then you can find the data with importlib.resources.… See the full description on the dataset page: https://huggingface.co/datasets/tiararodney/cire-corpus.

sourceHugging Faceotherupdated 28d agoView on Hugging Face
0likes139downloads
Dataset Card

CIRE Corpus

The CIRE Corpus is a set of Infrastructure-as-Code template files. The files come from public sources. The templates use these formats:

  • —AWS CloudFormation
  • —Azure Resource Manager
  • —Terraform

The related dataset cire-ontology classifies resource types. This corpus stores the infrastructure descriptions that use those types.

You can install the corpus as the namespace package tiararodney.cire.corpus.infrastructure. Then you can find the data with importlib.resources. This method is the same as the method for tiararodney.cire.ontology.

The repository contains the dataset and the collection code:

  • —The dataset is in data/.
  • —The collection code is in src/.

The manifest schema and the code that writes the manifest stay in the same version. Users read the corpus only from the manifest. Users must not import the collection code.

A fixed copy of the repository contains the data, the schema, and the source data together.

Layout

schema.json      JSON Schema (draft 2020-12) for a manifest record
data/tiararodney/cire/corpus/infrastructure/data/
  manifest.jsonl one JSON record per stored object (the contract)
  objects/       content-addressed store: objects/<sha256[0:2]>/<sha256>
src/tiararodney/cire/corpus/infrastructure/acquisition/
                 collection code (discovery, fetch, clean, dedup,
                 manifest writing); not necessary to use the corpus

The path field in a record is relative to the data package root. The full path is:

files("tiararodney.cire.corpus.infrastructure.data") / record["path"]

Stored objects have no file extension. The format field in the record (json, yaml, hcl) tells the parser the file type. The target_provider field names the IaC dialect.

Manifest contract

manifest.jsonl is the only interface between this dataset and its users. Each record does this:

  • —The record is valid against schema.json.
  • —The record contains schema_version. Users must validate the version when they read the record. Users must not accept a version that they do not support. Users must not guess the meaning of an unknown version.
  • —The sha256 field is the hash of the stored bytes. This hash identifies exact duplicates.
  • —The normalized_sha256 field is the hash after comment removal and whitespace normalization. This hash puts small variants in one group.
  • —The record stores source data (source.url, source.commit, source.retrieved_at) and the license of the content.

Example record:

json
{
  "schema_version": 1,
  "sha256": "9f2c1a7d0b34e6c8f5a2d91e07b3c4a6d8e0f1b2c3d4e5f60718293a4b5c6d7e",
  "normalized_sha256": "1e2d3c4b5a69788796a5b4c3d2e1f00f1e2d3c4b5a69788796a5b4c3d2e1f00f",
  "size_bytes": 4821,
  "target_provider": "cloudformation",
  "format": "yaml",
  "path": "objects/9f/9f2c1a7d0b34e6c8f5a2d91e07b3c4a6d8e0f1b2c3d4e5f60718293a4b5c6d7e",
  "source": {
    "url": "https://raw.githubusercontent.com/example/repo/3f9c0e1/stack/web.yaml",
    "repository": "https://github.com/example/repo",
    "commit": "3f9c0e1",
    "path_in_source": "stack/web.yaml",
    "retrieved_at": "2026-09-13T00:00:00Z"
  },
  "license": {
    "spdx": "MIT",
    "determined_by": "repository_license_file"
  },
  "scrub": {
    "status": "clean"
  },
  "run_id": "2026-09-13-github-sweep-01"
}

Secrets

The process removes credentials and secrets from all content before storage. The repository stores only the cleaned objects. The original bytes that contain secrets are not in the repository.

Each record stores the result in scrub.status:

  • —clean
  • —redacted

Redacted values use known replacement text. A redacted object is not the same byte sequence as the source file.

Licensing

This dataset has two parts. Each part has a different license.

Metadata. manifest.jsonl, schema.json, and this README are original work. These files use the CC-BY-SA-4.0 license.

Stored objects. The template files in objects/ are third-party content from public sources. These files keep their original licenses. This project does not change those licenses. The license.spdx field in each record gives the license that the collection process found. If the process cannot find a license, the field is NOASSERTION.

Users must obey the license of each object in the manifest. If your use needs specified license terms, select records with license.spdx.