Team Ai
Modelpublic

cijov/Sienna-v2-1B

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes
Model Card

<p align="center"> <img src="Sienna.png" alt="Sienna Logo" width="1200"/> </p>

Sienna-v2

Sienna is a children's-story generator fine-tuned on top of cijov/Cijov-lang-v1-1B, a ~1.2B parameter model. This release is a LoRA fine-tune merged into a standalone checkpoint — no peft dependency needed to use it.

Supports 4 languages (English, French, Spanish, Romanian) across 5 story genres (bedtime, animals, friendship, fantasy, general), selected via the system prompt.

Usage

python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("cijov/Sienna-v2-1B", subfolder="model")
model = AutoModelForCausalLM.from_pretrained(
    "cijov/Sienna-v2-1B", subfolder="model", trust_remote_code=True, dtype=torch.bfloat16
).to("cuda")

messages = [
    {"role": "system", "content": (
        "You are Sienna, a children's story writer. "
        "Write a short magical fairy-tale for a young child. "
        "Use simple words and a wondrous, friendly tone. "
        "Respond only in Romanian."
    )},
    {"role": "user", "content": "Tell me a magical fairy tale."},
]
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True, enable_thinking=False) + "\n"
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
out = model.generate(**inputs, max_new_tokens=300, do_sample=True, temperature=0.8, top_p=0.9, top_k=50, repetition_penalty=1.15)
print(tokenizer.decode(out[0][inputs.input_ids.shape[1]:], skip_special_tokens=True))

Or use the included standalone script:

bash
python generate.py --genre fantasy --lang ro

trust_remote_code=True is required — the backbone (cijov/Cijov-lang-v1-1B) is published as its own standalone architecture, not a stock transformers class.

Genres and languages

genresystem prompt intent
bedtimecalm, soothing, peaceful ending
animalsfun, simple, happy ending
friendshipwarm, kindness/sharing, gentle lesson
fantasymagical fairy-tale, wondrous tone
generalshort, simple, age-appropriate

Select a language by appending Respond only in <Language>. to the system prompt (English / French / Spanish / Romanian) — see generate.py.

Known limitations

Romanian quality lags the other three languages. This is inherited from the base backbone's own pretraining — Romanian started with substantially higher perplexity and lower QA accuracy than English/French/Spanish before Sienna's fine-tuning ever touched it, and additional fine-tuning does not close that gap (confirmed via a dedicated experiment: extending training specifically to test this left Romanian perplexity flat across 6,000+ further steps). Concretely, on held-out validation text: English ppl ≈ 11.8, Spanish ≈ 17.2, French ≈ 26.3, Romanian ≈ 81.6. Closing this gap requires additional Romanian-language pretraining in the backbone itself, not further Sienna-side fine-tuning.

Genre adherence is moderate, not high, and varies by language/genre — measured via keyword-based classification on generated samples (grand average ~48% across languages/genres, "general" genre excluded from scoring since it has no positive keyword signal of its own). Diversity and repetition metrics are strong across the board (distinct-2 ≈ 0.96-0.98, 4-gram repetition ≈ 0.00), and a small multilingual safety-keyword check found near-zero hits — the model reliably avoids degenerate/repetitive output and unsafe content, but doesn't always hit the requested genre on the first try. Occasional short glued-fragment artifacts (a stray foreign- language word fused into an otherwise-correct sentence) can appear rarely; this is a known, low-frequency noise-floor characteristic rather than a systematic language-mixing failure — the model's own generation stays correctly in the requested language in the large majority of samples.

Training

  • —Base: cijov/Cijov-lang-v1-1B, frozen, LoRA rank 64.
  • —Data: roneneldan/TinyStories (en), ffuuugor/tinystories_spanish + fairy-tale sources (es), iproskurina/TinyStories-French + fairy-tale sources (fr), readerbench/ro-stories + fairy-tale + synthetic sources (ro). Under/over-represented languages and sources are oversampled to roughly equal effective training volume.
  • —Genre labels are auto-assigned via multilingual keyword matching over the story text at data-prep time, then encoded into the system prompt for training and inference alike.
  • —SFT with chat-template label masking (loss only on assistant tokens).

License

Apache 2.0, matching the base backbone. See LICENSE/NOTICE.