Team Ai
Modelpublic

boopathiraj/T5-Json-Parsing-Model

sourceHugging Facemitupdated 11mo agoView on Hugging Face
0likes30downloads
Model Card

T5-Json-Parsing-Model: Improved Evaluation Report

Fine-tuned T5 model for structured JSON metadata extraction from unstructured text

Model Overview

T5-Json-Parsing-Model is a text-to-JSON generation model based on T5, specifically fine-tuned to:

  • —Parse unstructured input
  • —Identify key entities and attributes
  • —Output valid JSON with consistent schema
Goal: "extract metadata: ..." → {"type": "...", "name": "...", ...}

Evaluation Results

MetricScoreInterpretation
Exact-Match Accuracy4.33%Very low — strict JSON format not followed
JSON Structural Accuracy1.33%Almost no outputs are valid JSON
ROUGE-153.92Good unigram overlap
ROUGE-238.33Moderate bigram matching
ROUGE-L51.53Strong sequence preservation
BLEU Score27.69Decent n-gram precision

Insight:

The model understands the content well (high ROUGE/BLEU), but fails to produce valid JSON syntax.

Inference Example

Input Prompt

text
extract metadata: John Smith, born in 1980, lives in New York, works as a data scientist.

### Inference : 

inputtext = "extract metadata: John Smith, born in 1980, lives in New York, works as a data scientist." inputids = tokenizer(inputtext, returntensors="pt", truncation=True, padding=True).input_ids

Generate text using the model's generate method

generatedids = model.generate(inputids, maxnewtokens=50, numbeams=5, earlystopping=True)

Decode the generated IDs to text

decodedoutput = tokenizer.decode(generatedids[0], skipspecialtokens=True)

print(json.loads("{" + decoded_output[1:-1] + "}"))

### output : 
  {'type': 'person', 'name': 'John Smith', 'yr_born': '1980', 'location': 'New York'}