boopathiraj/T5-Json-Parsing-Model
030
T5-Json-Parsing-Model: Improved Evaluation Report
Fine-tuned T5 model for structured JSON metadata extraction from unstructured text
Model Overview
T5-Json-Parsing-Model is a text-to-JSON generation model based on T5, specifically fine-tuned to:
- Parse unstructured input
- Identify key entities and attributes
- Output valid JSON with consistent schema
Goal:"extract metadata: ..."→{"type": "...", "name": "...", ...}
Evaluation Results
Insight:
The model understands the content well (high ROUGE/BLEU), but fails to produce valid JSON syntax.
Inference Example
Input Prompt
extract metadata: John Smith, born in 1980, lives in New York, works as a data scientist.
### Inference :
inputtext = "extract metadata: John Smith, born in 1980, lives in New York, works as a data scientist." inputids = tokenizer(inputtext, returntensors="pt", truncation=True, padding=True).input_ids
Generate text using the model's generate method
generatedids = model.generate(inputids, maxnewtokens=50, numbeams=5, earlystopping=True)
Decode the generated IDs to text
decodedoutput = tokenizer.decode(generatedids[0], skipspecialtokens=True)
print(json.loads("{" + decoded_output[1:-1] + "}"))
### output :
{'type': 'person', 'name': 'John Smith', 'yr_born': '1980', 'location': 'New York'}
