tmutton/wcag-accessibility-issues
WCAG Accessibility Issues A small synthetic dataset for experimenting with machine learning classification of accessibility defects against WCAG success criteria. This project was created as a learning exercise to explore the complete Hugging Face text-classification workflow, including dataset preparation, tokenisation, fine-tuning a pretrained Transformer model, evaluation and inference. Dataset The dataset contains 100 synthetic accessibility defect… See the full description on the dataset page: https://huggingface.co/datasets/tmutton/wcag-accessibility-issues.
WCAG Accessibility Issues
A small synthetic dataset for experimenting with machine learning classification of accessibility defects against WCAG success criteria.
This project was created as a learning exercise to explore the complete Hugging Face text-classification workflow, including dataset preparation, tokenisation, fine-tuning a pretrained Transformer model, evaluation and inference.
Dataset
The dataset contains 100 synthetic accessibility defect descriptions across five WCAG success criteria.
Each class contains 20 examples.
Example:
Custom radio control does not expose its selected stateLabel:
4.1.2 Name, Role, ValueAll examples are synthetic and were created specifically for this project. The dataset does not contain real audit findings or customer data.
Dataset Splits
The 100 examples are divided using a stratified split:
The original complete dataset is available in data.csv.
split_data.py can be used to reproduce the train, validation and test datasets.
Model Training
The project uses distilbert-base-uncased as the pretrained base model.
DistilBERT is fine-tuned for multiclass sequence classification with five output classes corresponding to the five WCAG success criteria.
The training configuration currently uses:
- 5 epochs
- Batch size of 8
- Learning rate of 2e-5
- Validation after each epoch
- Weighted F1 as the model-selection metric
The training process is implemented in train_model.py.
Initial Results
The initial model achieved:
These results are based on a test set containing only 10 examples and should not be interpreted as evidence that the model will achieve 80% accuracy on real-world accessibility defects.
The results demonstrate that the end-to-end classification pipeline is working.
Example Predictions
After fine-tuning, the model correctly classified previously unseen examples such as:
no focus indicatoras:
2.4.7 Focus Visibleand:
can't open menu with keyboardas:
2.1.1 KeyboardRepository Files
data.csv— complete synthetic datasettrain.csv— training splitvalidation.csv— validation splittest.csv— held-out test splitsplit_data.py— creates the stratified dataset splitsload_dataset.py— loads the CSV files using Hugging Face Datasetstokenize_dataset.py— tokenises examples and encodes labelsload_model.py— loads DistilBERT for sequence classificationtrain_model.py— fine-tunes and evaluates the classifierpredict.py— performs inference using the trained model
Limitations
This is a small educational dataset rather than a production accessibility dataset.
In particular:
- Only five WCAG success criteria are represented.
- The examples are synthetic.
- There are only 20 examples per class.
- The test set contains only 10 examples.
- Real accessibility defects may be substantially more complex or ambiguous.
- Some accessibility defects can relate to more than one WCAG success criterion.
The classifier is a closed-set classifier. It must select one of its five known classes even when an issue actually belongs to a WCAG criterion that is not represented in the dataset.
For example, an issue concerning missing alternative text would normally relate to 1.1.1 Non-text Content, but the current model has no 1.1.1 class available.
Purpose
The purpose of this project is to learn and demonstrate an end-to-end Hugging Face machine learning workflow:
Dataset
↓
Train / Validation / Test Split
↓
Tokenisation
↓
Pretrained DistilBERT
↓
Fine-tuning
↓
Evaluation
↓
InferenceThe project can be extended by adding more examples, introducing additional WCAG success criteria and evaluating the resulting classifier against a larger and more representative test dataset.
