Team Ai
Modelpublic

purebyte/secrets-code

sourceHugging Faceotherupdated 8d agoView on Hugging Face
0likes63downloads
README.md192 linesDownload Raw Back to root
1---2license: other3license_name: purebyte-model-license4license_link: https://github.com/purebyte-ai/purebyte/blob/secrets-code-1.0.0/MODEL_LICENSE.md5pipeline_tag: token-classification6tags:7  - secret-detection8  - credentials9  - token-classification10  - byte-level11  - state-space-model12  - ternary13  - cpu14---15# Model card: secrets-code16 17> The model file is `purebyte-secrets-code-1.0.0.gguf` in this repository; its SHA-256 is next to it. It runs with the PureByte runtime ([github.com/purebyte-ai/purebyte](https://github.com/purebyte-ai/purebyte)): `pip install purebyte`, then `purebyte models pull secrets-code`, which downloads the same file from the runtime's release and checks its checksum.18 19An **AI Specialist** that finds credentials (API keys, access tokens, passwords, private keys, connection strings) in20source code, configuration and other text, and reports exactly where they are. It reads raw bytes, runs on ordinary CPUs21and never sends anything anywhere.22 23It is one of the three example specialists of PureByte 1.0, which show what one byte-level architecture does on different24jobs; the architecture, the runtime and the training stack are described in the [README](https://github.com/purebyte-ai/purebyte/blob/secrets-code-1.0.0/README.md).25 26| | |27|---|---|28| Version | 1.0.0 |29| Task | Credential detection as byte spans in text files |30| Input | Any text file up to 4,000,000 bytes: UTF-8, or UTF-16 (converted in memory); inputs under 24 bytes are not analyzed |31| Output | Findings: file, line, column, byte offsets, severity, masked snippet, and ensemble votes when an ensemble is used (see [spec/OUTPUT.md](https://github.com/purebyte-ai/purebyte/blob/secrets-code-1.0.0/spec/OUTPUT.md)) |32| Size | One file of 1,028,704 bytes (0.98 MiB); 1,927,520 parameters (1.93 M) in the network: the byte embedding, the blocks and the final norm. The two heads add 34,439 (1,961,959 in all); the per-group scales of the ternary weights are derived, not counted |33| Architecture | Byte-level state-space model with ternary weights (four `ssm_v2` blocks of width 256) and a span-tagging head gated by a window head with max pooling ([docs/architecture.md](https://github.com/purebyte-ai/purebyte/blob/secrets-code-1.0.0/docs/architecture.md)) |34| Default decision rule | One model, at the operating point stored in its file (`--bias` moves it) |35| Optional ensemble | `secrets-code-ensemble`: the same model plus two more trained the same way from other random seeds, each at its own operating point; a finding needs two of the three votes. Three times the work |36| Post-processing | The `secrets-code` profile: it applies format rules, decodes base64 layers, ignores documentation examples, counts the ensemble votes, sets the severity by path and masks the values |37| Runtime | `purebyte` 1.0.0 or newer, on an x86-64 or arm64 CPU (other little-endian CPUs through the portable build); no GPU, no network |38| License | [PureByte Model License](https://github.com/purebyte-ai/purebyte/blob/secrets-code-1.0.0/MODEL_LICENSE.md) (the code of the runtime is Apache-2.0) |39| Files | `purebyte-secrets-code-1.0.0.gguf` (the model) and, for the ensemble, `-e1` and `-e2`; checksums in [models.json](https://github.com/purebyte-ai/purebyte/blob/secrets-code-1.0.0/models/models.json) |40| Recipe | [specialists/secrets-code](https://github.com/purebyte-ai/purebyte-train/blob/main/specialists/secrets-code/README.md) in purebyte-train: the generator, the recipe and the evaluation, to rebuild or improve it |41 42```bash43purebyte models pull secrets-code                   # the model44purebyte models pull secrets-code-ensemble          # optional: the two extra members of the ensemble45purebyte scan --model secrets-code-ensemble src/46```47 48## Intended use49 50- Pre-commit hooks and CI checks on changed files, where a few kilobytes are scanned per commit.51- Local scans of repositories, build artifacts and configuration bundles before they leave the machine.52- A second opinion next to rule-based scanners: it finds credentials that have no fixed pattern, from the context53  around them.54- Small decisions through the HTTP API: "does this line, event or snippet contain a credential?"55- Redacted copies of configuration files and logs (`purebyte redact --model secrets-code`), with each credential56  replaced by a marker such as `[secret_1]`.57 58## Out of scope59 60- **Proving that code is free of secrets.** It misses a share of real credentials (below). Use it as one layer of61  defense, never as the only one.62- **Checking whether a credential is live.** It never contacts any service, by design.63- **Binary files.** Use [secrets-bin](https://github.com/purebyte-ai/purebyte/blob/secrets-code-1.0.0/models/secrets-bin.md).64- **Personal data** (e-mail addresses, phone numbers, identity numbers). That is the job of [pii](https://github.com/purebyte-ai/purebyte/blob/secrets-code-1.0.0/models/pii.md).65- **Files over 4,000,000 bytes**, which are reported as not analyzed.66 67## Evaluation68 69On the TEST half of [CredData](https://github.com/Samsung/CredData) (a split by repository, computed from the70dataset's metadata; files up to 4 MB; line-level metric), measured together with gitleaks and trufflehog on the same71files. The TEST half has 168 repositories; four of them overlap the code corpus the model was trained with and are72left out ([Training-data overlap](https://github.com/purebyte-ai/purebyte/blob/secrets-code-1.0.0/benchmarks/RESULTS.md#training-data-overlap)): 164 repositories, 4,219 labeled73credential lines.74 75| Tool | Precision | Recall | F1 [95 % CI] |76|---|---:|---:|---|77| **secrets-code** 1.0.0 | 0.879 | **0.729** | **0.797** [0.719-0.860] |78| secrets-code, ensemble of three (`secrets-code-ensemble`) | 0.879 | 0.741 | 0.805 [0.731-0.863] |79| gitleaks 8.30.1 | 0.905 | 0.207 | 0.337 [0.263-0.421] |80| trufflehog 3.97.6 (no verification) | 0.578 | 0.025 | 0.047 [0.017-0.092] |81 82On the complete TEST half (all 168 repositories), secrets-code scores precision 0.876, recall 0.659, F1 0.75383[0.665-0.841]. For fewer false alarms, `--bias -2` (an operating point chosen on DEV) gives precision 0.932 and recall840.663 on the 164 repositories. Counts, other slices, results by category, the operating points and the full reproduction85procedure: [benchmarks/RESULTS.md](https://github.com/purebyte-ai/purebyte/blob/secrets-code-1.0.0/benchmarks/RESULTS.md). The errors on TEST are counted, never inspected.86 87**How to read these figures.** Besides generated examples, the model learned from labeled lines of CredData88repositories of the DEV half, which share no repository with TEST but follow the same labeling policy (for89example, credential-shaped values in test code count as credentials). Part of the gain over rule-based scanners may90therefore be CredData's own conventions; on repositories labeled differently, expect lower figures.91 92**On the original values.** CredData replaces every labeled credential value with a random string, and the model93learned from such values in the DEV half. With the original values restored on the same 164 repositories, its recall94is 0.583 instead of 0.729 and its F1 0.694 instead of 0.797 (the ensemble: 0.597 and 0.704); gitleaks barely moves95(recall 0.201, F1 0.329). The model still finds 2.9 times as many labeled credential lines as gitleaks (2,460 against96850). The loss is in passwords that are ordinary words ([details](https://github.com/purebyte-ai/purebyte/blob/secrets-code-1.0.0/benchmarks/RESULTS.md#on-the-original-values)).97 98## False alarms on repositories without credentials99 100Three sets of popular public repositories that hold no real credential, pinned by commit101([data/exams/code-fp.yaml](https://github.com/purebyte-ai/purebyte-train/blob/main/data/exams/code-fp.yaml) in102purebyte-train), scanned with the default options; every finding is counted as a false alarm, even when it is103credential-shaped test data:104 105| Set | Repositories | Files scanned | Findings | Of them, warnings (test, example and documentation paths) |106|---|---:|---:|---:|---:|107| code-fp-a | 6 | 1,219 | 93 | 90 |108| code-fp-a without tornado | 5 | 931 | 73 | 70 |109| code-fp-b | 18 | 4,823 | 408 | 392 |110| code-fp-b without starlette, got and devise | 15 | 4,322 | 339 | 323 |111| code-fp-c | 12 | 4,778 | 138 | 49 |112| code-fp-c without awesome-compose, nginx, helm-charts, dotfiles and home-manager | 7 | 2,046 | 51 | 27 |113 114Nine of these repositories overlap the training data, so each set is also given without them: tornado, starlette, got,115devise and awesome-compose are CredData DEV repositories whose labeled lines the model learned from, and116awesome-compose, nginx, helm-charts, dotfiles and home-manager were in the code corpus its generated examples were cut117from ([details](https://github.com/purebyte-ai/purebyte/blob/secrets-code-1.0.0/benchmarks/RESULTS.md#false-alarms-on-repositories-without-credentials)). The released model was118also chosen on code-fp-b ([Training data](#training-data)), so its count there is optimistic in either row.119 120Most findings are warnings: keys, certificates, cookies and `Basic` headers in tests, fixtures and documentation,121which look exactly like credentials and do not fail a scan. The errors (108 across the three complete sets) come122mostly from integrity hashes in `package-lock.json` files, base64 image data embedded in a notebook, demo passwords and123keys of example applications outside test folders, and example `Basic` headers in doc comments.124 125## Known limitations126 127- **It misses about one labeled credential in four** on CredData TEST as distributed (recall 0.729), and two in five128  on the original values (recall 0.583), mostly passwords that are ordinary words. The classes it misses most,129  measured on the DEV repositories that were never used in training: identifiers that CredData labels as credentials130  (UUIDs, salts, nonces), found only under an unequivocal credential name, by design; values with no name next to them;131  and short quoted passwords ([known weak spots](https://github.com/purebyte-ai/purebyte/blob/secrets-code-1.0.0/benchmarks/RESULTS.md#known-weak-spots)).132- **Its most common wrong finding is a short ordinary value under a credential-like name** (about 45 % of its false133  positives on the DEV repositories never used in training), then strings full of symbols, and multi-line blocks that134  CredData does not count as private keys.135- **Credential-shaped test data is reported.** Fixtures, example headers and test keys look exactly like credentials;136  in test, example and documentation paths they are warnings (`--tests no` skips those paths).137- **Hashes and base64 data can be flagged**: integrity hashes in lock files and data embedded in notebooks were the138  largest sources of errors on the repositories without credentials (above).139- **Cryptographic test vectors are treated as non-secrets** on purpose (for example `Key = <hex>` in test data).140- **Documentation examples are ignored by default** (for example AWS's `AKIAIOSFODNN7EXAMPLE`); `--keep-examples`141  reports them.142- **Confidence values are not calibrated yet.** Use the finding itself (and the votes, with the ensemble), not the143  number alone.144- **Throughput is modest.** It suits commits, diffs and small decisions better than full-history scans of large145  repositories: about 130 KB/s on a 12-core desktop ([docs/performance.md](https://github.com/purebyte-ai/purebyte/blob/secrets-code-1.0.0/docs/performance.md)).146 147## Training data148 149- **Generated examples in real code.** Windows of real code and configuration from public repositories (for the150  released weights, clones made on 19 September 2026; the data scripts rebuild the corpus from repositories pinned by151  commit), each with generated lines injected: credentials in real provider formats, connection strings, private keys,152  generic secrets under credential names, and negatives that only look like credentials (placeholders, environment153  lookups, hashes, identifiers, public keys, documentation examples, ordinary short values). Only the injected values154  are marked.155- **Labeled real lines.** One window in five of every training step comes from lines labeled as credentials or156  non-credentials in CredData repositories of the DEV half: a fixed 117 of its 169 repositories (the others are kept157  apart for measuring, and the openssl repository is left out); no TEST repository was read. Half of those credential158  values are replaced by generated values of the same form.159- **Overlap.** The code corpus of the time held four CredData TEST repositories, which are left out of the published160  figures ([details](https://github.com/purebyte-ai/purebyte/blob/secrets-code-1.0.0/benchmarks/RESULTS.md#training-data-overlap)), and five repositories of the false-alarm set161  code-fp-c; four more false-alarm repositories, and one of those five, are CredData DEV repositories whose labeled162  lines it learned from (the figures without all nine are above). The data scripts of163  [purebyte-train](https://github.com/purebyte-ai/purebyte-train) rebuild the corpus from public sources, excluding164  every CredData repository and every exam repository:165  [data/](https://github.com/purebyte-ai/purebyte-train/blob/main/data/README.md).166 167The generator and the recipe are in168[specialists/secrets-code](https://github.com/purebyte-ai/purebyte-train/blob/main/specialists/secrets-code/README.md)169of purebyte-train. Three models are trained with different random seeds: one is released as the model, and the other170two are the extra members of the ensemble. **The released one was chosen on the false-alarm set code-fp-b** (the fewest171findings of the three), with the CredData TEST scores of the three already known, and not by the rule of our training172stack (the best F1 on the recipe's own `select` split): its count on code-fp-b is optimistic173([details](https://github.com/purebyte-ai/purebyte/blob/secrets-code-1.0.0/benchmarks/RESULTS.md#the-pre-registered-criteria-and-the-choice-of-the-released-model)).174 175**Pre-registered criteria.** Both CredData criteria of the recipe pass; both false-alarm guards on code-fp-b fail176(571, 360 and 455 findings per seed and 410 for the ensemble, against at most 129 and 107). The model was released177after reading a sample of such findings, most of them credential-shaped test data, and with three seeds instead of178the six our release checklist asks for179([details](https://github.com/purebyte-ai/purebyte/blob/secrets-code-1.0.0/benchmarks/RESULTS.md#the-pre-registered-criteria-and-the-choice-of-the-released-model)).180 181## Responsible use182 183A secrets scanner can find other people's secrets. Scan what you own or are authorized to assess. PureByte prints184findings masked by default and never verifies credentials against any service; if you find a live credential in185someone else's project, report it to its owner privately.186 187## Versions188 189| Version | Date | Changes |190|---|---|---|191| 1.0.0 | 2026-09-24 | First public release |192