Team Ai
Modelpublic

solavr/small100-ctranslate2-int8

sourceHugging Facemitupdated 25d agoView on Hugging Face
0likes129downloads
Model Card

SMaLL-100 CTranslate2 INT8

This repository contains a community conversion of `alirezamsh/small100` for CPU inference with CTranslate2.

The original SMaLL-100 checkpoint was converted to the CTranslate2 model format with INT8 weight quantization. No training, fine-tuning, pruning, or change to the tokenizer was performed.

This is not an official release from the SMaLL-100 authors and is not endorsed by them.

Model details

PropertyValue
ArchitectureSMaLL-100 / M2M-100 encoder-decoder Transformer
ParametersApproximately 330 million
Languages100 languages, supporting more than 10,000 directions
RuntimeCTranslate2 4.8.2 or compatible
Stored weight typeINT8
Recommended CPU compute typeint8_float32
TokenizerOriginal SMaLL-100 SentencePiece model
Base revision8ab680e26a596d2e3d2d2d17ae0f68df1037328c

The CTranslate2 model files are platform-independent. CPU instruction-set support is determined by the CTranslate2 runtime used to load them.

Installation

bash
pip install "ctranslate2==4.8.2" "sentencepiece==0.2.1"

Usage

python
from pathlib import Path

import ctranslate2
import sentencepiece as spm


model_dir = Path("path/to/small100-ctranslate2-int8")

translator = ctranslate2.Translator(
    str(model_dir),
    device="cpu",
    compute_type="int8_float32",
    inter_threads=1,
    intra_threads=4,
)
tokenizer = spm.SentencePieceProcessor(
    model_file=str(model_dir / "sentencepiece.bpe.model")
)

text = "Could you tell me where the station is?"
target_language = "ja"

# SMaLL-100 places the target-language token at the beginning of the source.
source_tokens = (
    [f"__{target_language}__"]
    + tokenizer.encode(text, out_type=str)
    + ["</s>"]
)

result = translator.translate_batch(
    [source_tokens],
    beam_size=5,
    max_decoding_length=256,
)
translation = tokenizer.decode(result[0].hypotheses[0])
print(translation)

The original model card uses a beam size of 5. A smaller beam such as 3 can reduce latency at a possible quality cost. Decoding parameters are not baked into the model files.

Conversion provenance

The conversion was produced from the pinned base-model revision above with:

  • —CTranslate2 4.8.2
  • —Transformers 5.15.0
  • —PyTorch 2.13.0
  • —SentencePiece 0.2.1
  • —CTranslate2 converter quantization: int8

Exact metadata is available in conversion.json, and file hashes are listed in SHA256SUMS.

Files

  • —model.bin: CTranslate2 model with INT8 weights
  • —config.json: CTranslate2 model configuration
  • —shared_vocabulary.json: shared source and target vocabulary
  • —sentencepiece.bpe.model: original SMaLL-100 tokenizer
  • —conversion.json: base revision and conversion environment
  • —SHA256SUMS: SHA-256 checksums for the runtime files

Limitations

  • —Translation quality varies substantially by language pair and domain.
  • —The base model was designed to preserve performance on low-resource languages; quality on some high-resource directions may trail larger or language-specific models.
  • —INT8 quantization can change model outputs and may reduce quality compared with the original floating-point checkpoint.
  • —Partial sentences, ambiguous text, names, numbers, and specialized terms should be reviewed when accuracy is critical.
  • —Biases present in the original model can remain or be amplified after compression.

See the original SMaLL-100 model card and paper for training data, evaluation results, supported languages, and further limitations.

License and attribution

The base model is distributed under the MIT License. The original copyright and license notice are preserved in LICENSE. Please cite the original work when using this conversion in research.

bibtex
@inproceedings{mohammadshahi-etal-2022-small,
  title = {{SM}a{LL}-100: Introducing Shallow Multilingual Machine Translation Model for Low-Resource Languages},
  author = {Mohammadshahi, Alireza and Nikoulina, Vassilina and Berard, Alexandre and Brun, Caroline and Henderson, James and Besacier, Laurent},
  booktitle = {Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing},
  year = {2022},
  pages = {8348--8359},
  url = {https://aclanthology.org/2022.emnlp-main.571/}
}