Team Ai
Modelpublic

lukeslp/chime-sms-2m

sourceHugging Facecc-by-sa-4.0updated 13h agoView on Hugging Face
0likes222downloads
Model Card

Chime-SMS-2M

Chime is a small English model that predicts the next word and completes the word being typed. I built it for LocalType, my Android keyboard, where it runs on the phone on every keystroke. It has 2,025,114 parameters, and the int8 model is a 2.2 MB file. Try it in your browser.

2.03M parameters · 2.22 MB int8 · 161.6M training words · 2h 2m on one A100 · 3.09–3.24 ms warm p95 on two Pixels

Chime uses the CIFG recurrent architecture described for Gboard in 2018. Its 2.03M parameters sit alongside the published 1.4M model (2018) and 2.4–4.4M next-word models (2023). This compares published model sizes, not current installed models or prediction accuracy. Measurements and shareable cards report the installed keyboard observations separately from model specifications.

[image]

Model details

  • —Developed by: Luke Steuber
  • —Model type: word-level language model: one CIFG-LSTM layer (670 units, projected to 96 dimensions) with tied input and output embeddings, the design of Gboard's published next-word model (Hard et al., 2018)
  • —Vocabulary: 16,384 tokens: 16,054 words, 293 emoji, 31 punctuation marks and 6 special tokens
  • —Context: the sentence being typed, up to 63 tokens
  • —Language: English
  • —License: CC BY-SA 4.0 for the weights, Apache-2.0 for the code

Uses

Chime ranks candidate words for a keyboard's suggestion strip: the next word after a space, or completions of a partly typed word. It is not a chat or text-generation model, and it does not correct spelling.

LocalType learns personal word choices separately; Chime’s weights stay fixed. From LocalType 0.5.13, a verified keep or accepted suggestion can qualify after one choice, while an ordinary known word needs two uses. The application uses up to three preceding words completed there and reserves at most one personal strip slot. Forget learned words clears that history. This application policy is separate from the model’s prediction measurements and has no completed controlled Gboard result yet.

How to use

sh
python3.12 -m venv .venv && . .venv/bin/activate
python -m pip install huggingface_hub
hf download lukeslp/chime-sms-2m --local-dir chime-sms-2m && cd chime-sms-2m
python -m pip install -r requirements.txt
python inference.py "see you "

A trailing space asks for the next word; without one, the last word is completed. From Python, LocalTypeModel().predict("see you ", k=6) in inference.py does the same. Inference needs only NumPy and runs offline.

cifg_int8.bin is the int8 model with its vocabulary, the file LocalType ships, and read_bin in inference.py reads both. model.safetensors holds the float32 weights; config.json lists their shapes.

Related keyboard architecture

The keyboard architecture comparison comes from Hard et al. (2018), reporting 1.4M parameters and a 1.4 MB quantized file, and Xu et al. (2023), reporting 2.4–4.4M next-word models. These published generations do not identify current installed Gboard models.

Keyboard pilot: unequal prior learning

This was an exploratory run, not a controlled accuracy benchmark. Earlier test attempts exposed LocalType on the 9a and Gboard on the 10 to more of the test targets. Learning was retained. The counts below describe those runs, not a model or phone ranking.

A 432-case automated touch run compared LocalType 0.5.9 (29), with Chime active, and Gboard 18.4.1.985164140-beta-arm64-v8a on Pixel 9a and Pixel 10, both API 37.

Measure9a LocalType / Gboard10 LocalType / Gboard
Next-word target offered15/36 / 13/3610/36 / 18/36
Useful completion available35/36 / 34/3635/36 / 35/36
Actual completion taps verified12/12 / 12/1212/12 / 12/12
Isolated typos repaired9/12 / 11/1210/12 / 10/12

On Pixel 10, LocalType lost: 10/36 targets (27.8%) versus Gboard’s 18/36 (50.0%). Unequal prior learning prevents attributing this gap to model quality or hardware; it does not establish that Chime would win a controlled test. All five extra LocalType hits on the 9a occupied its third slot, which can reserve a learned follower. Gboard’s internal model identity and cause of the phone difference were not measured. Each phone repeated the same 36 language targets; 1,048 prefix observations are not independent examples. The raw run is separate from the checkpoint-selection development metrics below.

Both keyboards preserved all 48 clean controls and 960 injected characters across protected and ordinary transport. LocalType offered and correctly inserted six of eight selected-word repairs on each phone; Gboard showed its features toolbar, so other manual repair interfaces were not evaluated. LocalType changed xkcd to did on both phones, and had comma/newline spacing failures after tapped suggestions. Later app fixes are excluded from this snapshot. No relative speed, battery or everyday accuracy claim follows. Complete methods and earlier pilots.

Development prediction measurements

On 463 NUS SMS development lines, Chime scored 41.44% potential keystroke savings and 19.32% next-word top-three accuracy. On 567 Taskmaster-1 requests it scored 59.50% and 42.34%. The same Kotlin scorer started with no personal learning. Potential savings assumes the correct word is tapped as soon as it appears among three suggestions, so it is a simulated ceiling. These lines were not training text but were used to select the checkpoint; no held-out or everyday savings result is claimed. These measurements cannot be compared with correction counts or another keyboard’s prediction scores on different data.

A separate ordinary-field integration pilot verified live Chime predictions in all 60 eligible observations, with empty isolated learning before each case. It preserved the tested clean text and protected delivery controls. No next-word target was scored or prediction tapped. Different field and learning policies prevent combining that run with the Gboard pilot into a model ranking.

Runtime comparison

Warm model inference, p95Pixel 9aPixel 10
Next word, retained context3.088 ms3.235 ms
Completion, retained context0.191 ms0.235 ms
Next word, fresh context6.675 ms7.578 ms
Long-context window reset subset14.134 ms17.638 ms

Standalone Kotlin benchmarks on the exact published int8 checkpoint used 600 ordinary contexts over three passes and 30 stress sentences, pooling warm passes two and three. Model loading took 139.090/131.655 ms and first calls 18.941/22.952 ms on the 9a/10. Both ran API 37; thermal status stayed 0 on the 9a and rose from 0 to 1 on the 10. These are model-only timings, not a controlled phone comparison, typing latency or battery measurements.

Int8 export raised perplexity from 12.985256 to 12.998372 (+0.101%) on 5,000 development lines. Kotlin matched 120 exported contexts and top-ten rankings with maximum logit error 1.27e-5; NumPy and browser JavaScript matched the same references with maximum error below 5e-8. Numerical agreement is separate from prediction quality.

Training

About 162 million words from five licensed English sources:

SourceText usedLicense
NUS SMS Corpustext messages people volunteered, mostly in Singapore, 2003 to 2015CC BY 4.0
Taskmaster-1the user's side of written task dialoguesCC BY 4.0
Taskmaster-3the customer's side of written movie-ticket dialoguesCC BY 4.0
OpenAssistant 2English prompts written by peopleApache-2.0
SODAdialogue generated by GPT-3.5CC BY 4.0

SODA is synthetic and makes up most of the text, so the NUS SMS, Taskmaster-1 and OpenAssistant lines were repeated eight times in each pass. Phone numbers, URLs, email addresses and credential-like strings were filtered out, and the text was deduplicated before it was split. No private messages or typing logs were used.

I trained it from scratch for three passes (102,195 updates) on one A100 in about two hours, with AdamW at a learning rate of 0.002 (warmup, then cosine decay) and batches of 256 sentences of up to 64 tokens. I kept the checkpoint at step 98,000, which had the lowest development perplexity on the human-written sources, and quantized it to int8 with one scale per row, which raised development perplexity by 0.1%.

Exact recipe and exposure

QuantityMeasured value
Training split before repetition7,858,648 lines; 161,625,719 words (runs of letters)
Weighted pass8,720,859 lines; 195,641,031 model target tokens
Completed run102,195 updates; three batcher epochs; 586,885,109 target-token exposures
Selected checkpointStep 98,000; human-macro development perplexity 60.5425
Final checkpointStep 102,195; human-macro development perplexity 60.9636
Training loop7,342.98 seconds (2h 2m 23s)
Training computation7,021.45 seconds; 0.068706 seconds/update
Training-loop cost estimate$5.10 at $2.50/hour; excludes setup/uploads; not a bill
Hardware / precisionOne NVIDIA A100; PyTorch float32
Batch / maximum sequence256 / 64 tokens
OptimizerAdamW; weight decay 0.01; gradient norm clip 1.0
Learning-rate schedule0.002 peak; 200-update warmup; cosine decay to 10% of peak
Seed20260925

×8 means repeating NUS SMS, Taskmaster-1 and OpenAssistant lines eight times per pass, while Taskmaster-3 and SODA appear once. It changes exposure, not parameter count or context size. Human-written lines form 12.868% of the weighted pass; SODA remains the majority. Repeated tokens are not additional unique data. The earlier ×4 checkpoint has the same architecture and is retained for comparison and rollback.

Limitations

  • —English only. Suggestions reflect older Singapore texting, task dialogue, assistant prompts and synthetic dialogue, and the fixed vocabulary will not learn new words or names.
  • —The vocabulary was learned from the corpus and includes a few offensive words. LocalType removes them before showing suggestions; do the same in your own application.
  • —Keystrokes saved assumes the right word is tapped the moment it appears, so it is a ceiling. Savings in everyday typing have not been measured yet.
  • —Filtering removed text that looked sensitive, which does not prove the source text is anonymous.

License

The weights and vocabulary are CC BY-SA 4.0, and inference.py and localtype_tokenizer.py are Apache-2.0; see LICENSE.md. NOTICE.txt credits the training data, whose authors do not endorse this model. The int8 model's SHA-256 is 3255dbeaf6e7fc0922afdca42ef751a22e3ad28903ab4acc9eb0022b9a704318.

Citation

bibtex
@misc{steuber2026chimesms2m,
  author = {Luke Steuber},
  title = {Chime-SMS-2M: An On-Device English Keyboard Prediction Model},
  year = {2026},
  url = {https://huggingface.co/lukeslp/chime-sms-2m}
}

Also by me: Drummer-540M, a 542M-parameter language model trained from scratch.