daniel-jurado/job-title-classification-dataset
SIRTAR Job classification Dataset This is a fusion of three known Kaggle datasets, added tons of preprocessing in the middle. These are the following: LinkedIn Job Postings (2023-2024) [1]: https://www.kaggle.com/datasets/arshkon/linkedin-job-postings Indeed Job Postings: https://www.kaggle.com/datasets/spandanakalakonda/job-postings Jobstreet Job Postings: https://www.kaggle.com/datasets/azraimohamad/jobstreet-all-job-dataset This dataset is used for the training of the new… See the full description on the dataset page: https://huggingface.co/datasets/daniel-jurado/job-title-classification-dataset.
SIRTAR Job classification Dataset
This is a fusion of three known Kaggle datasets, added tons of preprocessing in the middle. These are the following:
- LinkedIn Job Postings (2023-2024) [1]: https://www.kaggle.com/datasets/arshkon/linkedin-job-postings
- Indeed Job Postings: https://www.kaggle.com/datasets/spandanakalakonda/job-postings
- Jobstreet Job Postings: https://www.kaggle.com/datasets/azraimohamad/jobstreet-all-job-dataset
This dataset is used for the training of the new line (v1) of text classification encoder-only LLMs found here:
- v1 LITE (Recommended): https://huggingface.co/daniel-jurado/mRoBERTa-fine-tuned-for-job-classification-v1-lite
- v1: https://huggingface.co/daniel-jurado/mRoBERTa-fine-tuned-for-job-classification-v1
Preprocessing
- Reduced title complexity from 108,959 to 56,808 by clustering 27,251 samples in 386 unique groups using OPTICS algorithm with neighborhood ratio of 0.7 and 31-nearest neighbours. Noise is not considered bad here as it can be just standalone far too used titles.
- Removed several noise data.
- Removed job offers written in caligraphies other than roman for both titles and descriptions.
- Deduplicated offers.
- Reformated job titles, excluding any Organization, dates, salaries and Location names (using the
dymium/Dymium-NER-v1NER Model). Noisy words are also excluded. Still there are bad suffixes that harm clustering results as "sign-on bonus". - Removed badly formatted characters, most of stops and numbers.
How to get the dataset
linkedin_jobposts_dataset = load_dataset("daniel-jurado/job-title-classification-dataset")Attribute
Optionally, you can cite my original study covering a model that used this.
Scholarly BibteX cite (SPANISH)
Not available because Master dissertation's not approved nor defended yet.
Scholarly APA Cite (SPANISH)
Not available because Master dissertation's not approved nor defended yet.
Cites
- Koneru, A. (2024). LinkedIn Job Postings (2023 - 2024), 13. doi:10.34740/KAGGLE/DSV/9200871
