Team Ai
Apppublic

shalev396/tiny-shakespeare-chat

sourceHugging Facemitupdated 13d agoView on Hugging Face
0likes
App README

🎭 Tiny Shakespeare Chat

Chat with a small GPT (10.75M parameters, character-level) that was written and trained from scratch: first pretrained on the Tiny Shakespeare corpus, then chat-tuned on 7,096 consecutive dialogue-line pairs. It answers in play-style verse. It is a toy language model, so its replies are poetry, not facts. Replies stream character by character in the chat; temperature, top-k and reply length are adjustable. Details: model card.

API

/predict takes message (string), history (a JSON string: [[user, bot], ...] pairs or [{"role", "content"}, ...] messages, "[]" for none), temperature (0.1-1.5) and max_new_tokens (16-400 characters). It returns [reply, seconds, device]. reply is the bot's answer as a plain string, seconds is a float and device is "gpu" or "cpu". Top-k is fixed at 40. Sampling is random, so the same message gets a different reply each time.

bash
S=https://shalev396-tiny-shakespeare-chat.hf.space
curl -X POST $S/gradio_api/call/predict -H "Content-Type: application/json" \
  -d '{"data": ["How fares the king?", "[]", 0.8, 200]}'
curl -N $S/gradio_api/call/predict/<event_id>
js
import { Client } from "@gradio/client";
const client = await Client.connect("shalev396/tiny-shakespeare-chat");
const { data } = await client.predict("/predict", {
  message: "And the queen?",
  history: JSON.stringify([["How fares the king?", "He is well, my lord."]]),
  temperature: 0.8,
  max_new_tokens: 200,
});
// data = ["QUEEN ELIZABETH: ...", 1.9, "cpu"]

The streaming chat UI uses private events, so they are not part of the API.

How it works

Weights + model.py are downloaded from the model repo at startup (space_utils.import_model). The prompt is <|user|> message <|end|> <|bot|>, preceded by up to 3 earlier turns and cropped to the 256-character context. Generation stops at <|end|>. Hardware: cpu-basic (same code runs on ZeroGPU).