json-extraction
json_data_extraction
Diverse Restricted JSON Data Extraction
Curated by: The paraloq analytics team.
Uses
Benchmark restricted JSON data extraction (text + JSON schema -> JSON instance)
Fine-Tune data extraction model (text + JSON schema -> JSON instance)
Fine-Tune JSON schema Retrieval model (text -> retriever -> most adequate JSON schema)
Out-of-Scope Use
Intended for research purposes only.
Dataset Structure
The data comes with the following fields:
title: The… See the full description on the dataset page: https://huggingface.co/datasets/paraloq/json_data_extraction.json-extraction
JSON Extraction Dataset
Source
Rows
ProfessorBob/relation_extraction
6920
roborovski/dolly-entity-extraction
5945
sandeeppanem/resume-json-extraction-5k
4879
Jiraya/html_to_json_information_extraction_dataset
3035
HenriqueGodoy/extract-0
2606
owkin/medical_knowledge_from_extracts
1383
resume-json-extraction-5k
Dataset Card for resume-json-extraction-5k
Dataset Description
This dataset contains 4,879 resume examples formatted for fine-tuning language models to extract structured JSON information from resume text.
Dataset Summary
The dataset consists of resume text paired with structured JSON outputs containing:
Job titles (current and previous)
Companies (current and previous)
Years of experience
Seniority level
Primary domain and industries
Core and secondary skills… See the full description on the dataset page: https://huggingface.co/datasets/sandeeppanem/resume-json-extraction-5k.Medical-Entity-JSON-Extractionjson-extraction
Rob Dixon's JSON Extraction Dataset
A synthetic dataset for training JSON extraction models, generated using Claude 3 Haiku.
Dataset Overview
This dataset contains paired examples of:
Instructions: Natural language task descriptions asking to extract information
Text documents: Source content containing information to extract
JSON outputs: Structured data extracted from the text
The dataset is designed for training smaller models on constrained context lengths, with… See the full description on the dataset page: https://huggingface.co/datasets/robdixon/json-extraction.html_to_json_information_extraction_dataset
HTML to JSON Information Extraction Dataset
Description
The html_to_json_information_extraction dataset is a collection of over 7300 HTML snippets and their extracted information in JSON.
These HTML have been sourced (scraped) from about 25 companies' career pages.
The dataset contains three splits - train, test, unseen_test.
This dataset has been built to fine tune SLMs & LLMs for the information extraction task.
train split
This split contains over 5700 pair… See the full description on the dataset page: https://huggingface.co/datasets/Jiraya/html_to_json_information_extraction_dataset.
