Team Ai
Modelpublic

TULLUS/codeparrot-small-multi

sourceHugging Faceapache-2.0updated 7d agoView on Hugging Face
0likes351downloads
Model Card

TULLUS/codeparrot-small-multi ๐Ÿฆœ

A clean, verified Safetensors conversion of `codeparrot/codeparrot-small-multi`.

This repository preserves the original model weights while normalizing the legacy GPT-2 embedding representation from two physically duplicated matrices into properly tied input/output embeddings.

What this is

This is a checkpoint conversion, not a new training run.

The original CodeParrot checkpoint was loaded from its legacy pytorch_model.bin, inspected, normalized, saved as Safetensors, reloaded, and then subjected to exact tensor verification.

The goal was to produce a clean modern Transformers checkpoint without changing the learned weight values.

Source

Original model: codeparrot/codeparrot-small-multi

Original architecture: GPT-2 / GPT2LMHeadModel

Original tokenizer: GPT-2 tokenizer

Vocabulary size: 32,768

BOS token ID: 0

EOS token ID: 0

PAD token: none

Original checkpoint SHA256:

text
0207f6b427e3cbf1bcb9726abb6bbba6620e9912e1abf23087ccdef61818ffb2

Conversion details

The legacy checkpoint contained two separate physical copies of the embedding matrix:

text
transformer.wte.weight
lm_head.weight

Both tensors were verified to have identical values:

text
shape:            (32768, 768)
exact equality:   True
max difference:   0.0
same storage:     False

The model was then normalized with:

text
tie_word_embeddings = True

After normalization:

text
same storage:      True
exact equality:    True

The resulting checkpoint therefore uses a single shared embedding tensor for the input embeddings and language-model head.

Parameter count

The legacy checkpoint physically stored:

text
136,174,080 parameters

This included the duplicated embedding storage.

After tying the identical embeddings, the clean model contains:

text
111,008,256 unique parameters

The resulting Safetensors file is approximately:

text
444,048,000 bytes
423.48 MiB

Verification

The conversion was validated after saving and reloading the Safetensors checkpoint.

Verified checks include:

  • โ€”Original source SHA256
  • โ€”Source tokenizer length
  • โ€”BOS/EOS token IDs
  • โ€”Source embedding values
  • โ€”Source physical parameter count
  • โ€”Embedding storage normalization
  • โ€”Safetensors serialization
  • โ€”Reload into GPT2LMHeadModel
  • โ€”tie_word_embeddings=True
  • โ€”Final tied embedding storage
  • โ€”Metadata consistency
  • โ€”Exact tensor equality

Final verification result:

text
Weights: EXACTLY PRESERVED
Embedding configuration: TIED
Unique parameter count: 111,008,256
BOS/EOS IDs: 0 / 0
Exact tensor verification: PASS

The final comparison verified all 149 logical model tensors after reload. Safetensors physically stores 148 tensors because the tied lm_head.weight is represented by the shared transformer.wte.weight storage.

About the conversion warning

During loading of the original legacy checkpoint, Transformers reported the following unexpected entries:

text
transformer.h.{0...11}.attn.bias
transformer.h.{0...11}.attn.masked_bias

These are legacy GPT-2 attention-mask buffers rather than learned model weights. They do not represent missing trained parameters, and the complete learned tensor set was subsequently verified exactly after conversion and reload.

Intended use

This checkpoint is intended as a clean starting point for experimentation with the CodeParrot GPT-2 architecture, including further fine-tuning and research experiments.

It is also the clean base checkpoint for the planned:

text
TULLUS-CodeParrot-???

training experiment.

Example usage

python
from transformers import AutoTokenizer, AutoModelForCausalLM

model_id = "TULLUS/codeparrot-small-multi"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id)

prompt = "def fibonacci(n):"

inputs = tokenizer(prompt, return_tensors="pt")

outputs = model.generate(
    **inputs,
    max_new_tokens=128,
)

print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Provenance

This repository is a conversion of the original:

`codeparrot/codeparrot-small-multi`

No claim is made here that the converted checkpoint is a new trained model. The purpose of this repository is to provide a clean Safetensors representation with normalized tied embeddings while preserving the original learned values.


TULLUS BAKES ๐Ÿช

Fresh From The AI Ovens, It's The Goods..