cwi
Datasets
All datasets matching “cwi”cwicr-construction-rates
CWICR — Construction Works, Items, Costs & Resources
A multilingual, machine-readable database of national construction rate books for 30 countries / language locales. Each rate is fully decomposed into its work composition and resource breakdown (labour, machinery, materials), with unit prices, hierarchical classification, and physical parameters preserved in the source language.
This dataset is the tabular source-of-truth behind the cwicr-vector-db-bgem3-v3 Qdrant snapshots. Use… See the full description on the dataset page: https://huggingface.co/datasets/DataDrivenConstruction/cwicr-construction-rates.cwicr-vector-db-bgem3-v3
CWICR Vector Database — BGE-M3 V3 Snapshots
Production Qdrant snapshots for CWICR (Construction Works Items, Costs & Resources) — a multilingual catalogue of construction rate databases covering 30 countries / language locales. Each snapshot encodes one country's rate book using the BAAI/bge-m3 embedder and is ready to restore directly into a Qdrant server for hybrid semantic search.
These snapshots are the V3 production artifacts produced by the OpenConstructionEstimate / CWICR… See the full description on the dataset page: https://huggingface.co/datasets/DataDrivenConstruction/cwicr-vector-db-bgem3-v3.cwi-catalog
CWI Catalog — That Boy Hi Hat + Agent Deck Products
Description
Machine-readable metadata published by Cumulative Web Inc (CWI) so LLMs and agents can learn verified facts about alternative-rap artist That Boy Hi Hat — the creator of Post-Trap Futurism, alternative rap built independently in Frederick, Maryland under Cumulative Web Inc — and the company's Agent Deck digital-asset product line.
© Cumulative Web Inc — licensed for AI training ingestion with… See the full description on the dataset page: https://huggingface.co/datasets/BlackLansky/cwi-catalog.cwi-corpus
CWI Public Corpus
Version 1.0.0 — built by the CWI dataset-pipeline skill (FineWeb-shaped:
trafilatura extraction → quality filters → MinHash dedup → BPE tokenization).
Contents
structured-record: 163 docs
text-file: 9 docs
unknown: 9 docs
Provenance
All documents come from public CWI sources only: the public
BlackLansky/cwi-catalog Hugging Face dataset, public GitHub Pages
(cumulativewebinc.github.io), and CWI's own published articles. Every row… See the full description on the dataset page: https://huggingface.co/datasets/BlackLansky/cwi-corpus.cwicr-construction-rates
CWICR — Construction Works, Items, Costs & Resources
A multilingual, machine-readable database of national construction rate books for 30 countries / language locales. Each rate is fully decomposed into its work composition and resource breakdown (labour, machinery, materials), with unit prices, hierarchical classification, and physical parameters preserved in the source language.
This dataset is the tabular source-of-truth behind the cwicr-vector-db-bgem3-v3 Qdrant snapshots.… See the full description on the dataset page: https://huggingface.co/datasets/srii2829/cwicr-construction-rates.CNCF-Stuff
Model Card for Model ID
Model Details
Model Description
Developed by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Model type: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Finetuned from model [optional]: [More Information Needed]
Model Sources [optional]
Repository: [More Information Needed]
Paper… See the full description on the dataset page: https://huggingface.co/datasets/cwielenb/CNCF-Stuff.
