bcv-commons/senses-attested-bhsa
senses_attested-bhsa — attested target renderings per lexeme sense (retired legacy set) Renamed 2026-10-04: this repository was bcv-commons/senses-attested; it is now bcv-commons/senses-attested-bhsa so the name says what it is. The name senses-attested now belongs to the UBS-keyed dataset. Superseded, and relabelled (2026-10-03). The sense number in this dataset is the sense number from a lexeme-sense clustering that is built on ETCBC BHSA clauses (BHSA is CC BY-NC-SA 4.0).… See the full description on the dataset page: https://huggingface.co/datasets/bcv-commons/senses-attested-bhsa.
senses_attested-bhsa — attested target renderings per lexeme sense (retired legacy set)
Renamed 2026-10-04: this repository wasbcv-commons/senses-attested; it is nowbcv-commons/senses-attested-bhsaso the name says what it is. The namesenses-attestednow belongs to the UBS-keyed dataset.
Superseded, and relabelled (2026-10-03). The sense number in this dataset is the sense number from a lexeme-sense clustering that is built on ETCBC BHSA clauses (BHSA is CC BY-NC-SA 4.0). This dataset is therefore now labelled CC BY-NC-SA 4.0 (non-commercial, share-alike); up to this date it carried a CC BY 4.0 label, and copies obtained earlier keep the label they came with. For an open licence and a sense key you can trust, use [`bcv-commons/senses-attested`](https://huggingface.co/datasets/bcv-commons/senses-attested) (CC BY-SA 4.0, keyed on UBS Dictionary of Biblical Hebrew sense ids, no BHSA-derived content). This dataset receives no further updates.Note (2026-09-30). Thesensenumber in this dataset comes from our own automatic disambiguation and is1for about 97% of tokens; checked against the manually built UBS Dictionary of Biblical Hebrew it agrees no better than chance when it says "same sense" (though where it does split, the split is informative). For a sense key you can trust, use the dataset `bcv-commons/senses-attested` (CC BY-SA 4.0), keyed on UBS sense ids. This dataset is kept unchanged for existing consumers.
Many Hebrew words carry more than one distinguishable meaning depending on their grammatical form — for example, a verb's meaning can shift with its binyan (the Hebrew verb-stem pattern: qal "simple/active" vs. hiphil "causative," etc.). This dataset records, for each Hebrew lexeme (a MACULA-anchored original-language word — see bcv-commons/lexeme-alignments's own README for the full definition) in a specific disambiguated (binyan, sense), which target-language translation words actually render that sense in practice, with counts — empirical evidence mined directly from the alignment data, not a hand-curated gloss list.
(Internal note for readers tracking the sibling shoresh/bcv-query projects — separate repos, not part of this codebase: this dataset is the evidence layer that fills shoresh's own senses_i18n/_gaps demand and cross-checks its llm_strongs_glosses predictions. It does not replace shoresh's own curated senses_i18n/<iso>.tsv.) Consumed as an HF Parquet dataset.
Schema (per row)
Key: `(lexeme, stem, sense)` — MACULA lexeme (anchor; BHSA lex dropped) + MACULA binyan + sense number, read inline from the enriched lexeme-spine.db. OT/Hebrew only (senses are Hebrew; Greek tokens carry none).
Multi-version: base_text is per-row, so several translations of a language are pooled into one `iso=<lang>` partition — a union of per-edition runs, each row tagged by edition; share stays per-edition. Cross-edition agreement (how many base_texts attest a given (lexeme,stem,sense)→ surface) is the confidence signal, derivable directly from the rows. (Swedish iso=swe pools swe_fol Folkbibeln + swe_svk Kärnbibeln.)
Removal / takedown policy
Each row is a per-edition attestation carrying its base_text, so a rights-holder can request removal and it is a clean row-drop + republish (the dataset is content-addressed via each partition's content_sha256). Because most (lexeme,stem,sense)→surface facts are attested by more than one edition, dropping one edition typically leaves the linguistic fact intact via the others — properly attributed. Rows are never re-emitted with provenance stripped: a removed attestation is removed, not anonymized.
Removals are driven by a committed, auditable config: `data/senses_exclude.json` (read automatically on every build). A row is dropped if it matches any rule; a rule matches when all its stated fields equal the row's — fields lexeme, stem, sense, surface, base_text, omit to wildcard:
{"exclude": [
{"base_text": "swe_fol"}, // drop a whole edition
{"base_text": "swe_fol", "surface": "herren"} // drop one surface within an edition
]}After exclusion, survivor shares renormalise (per edition), so a removed row leaves no residue; the manifest records excluded: {rules, rows_dropped} for the audit trail. To action a takedown: add a rule, re-run senses_attested for the affected language, republish.
Licensing — CC BY-NC-SA 4.0 (since 2026-10-03; earlier CC BY 4.0)
The lexeme + binyan key is MACULA-derived (CC BY 4.0, attribute Clear-Bible MACULA), but the sense number comes from a clustering built on ETCBC BHSA clauses, so the dataset as a whole follows BHSA's CC BY-NC-SA 4.0: attribute both, non-commercial use only, share alike. We carry the sense number only and no English sense label (shoresh's sense labels are UBS-MARBLE "used with permission", not redistributable). The open-licence replacement is bcv-commons/senses-attested (UBS sense ids, CC BY-SA 4.0). Regenerate (legacy scheme): python -m lexeme_aligner.senses_attested --iso <iso> --method eflomal. Same git-ignored-Parquet + committed-manifest.json layout as lexeme-alignments.
