Team Ai
Datasetpublic

jimjunior/cocis-web-info

COCIS WEB INFO Dataset Summary This dataset contains information about Makerere University College of Computing and Information Science that was scraped from its official website and corresponding websites. The dataset consists of approximately 513 JSON chunks, designed for high-performance streaming and parallel processing. Each chunk represents a discrete unit of data structured for machine learning tasks. By sharding the data into 513 files, this repository… See the full description on the dataset page: https://huggingface.co/datasets/jimjunior/cocis-web-info.

sourceHugging Facemitupdated 7mo agoView on Hugging Face
1likes6.6kdownloads
Dataset Card

COCIS WEB INFO

Dataset Summary

This dataset contains information about Makerere University College of Computing and Information Science that was scraped from its official website and corresponding websites. The dataset consists of approximately 513 JSON chunks, designed for high-performance streaming and parallel processing. Each chunk represents a discrete unit of data structured for machine learning tasks.

By sharding the data into 513 files, this repository supports the datasets library's streaming mode, allowing users to train models without downloading the entire dataset into RAM—a critical feature for resource-constrained environments or high-concurrency CI/CD pipelines.

Repository Structure

The data is organized into a chunks/ directory to maintain a clean root level:

text
.
├── README.md          # This file
└── chunks/            # Directory containing 513 JSON files
    ├── chunk_1.json
    ├── chunk_2.json
    └── ...

Usage

You can load this dataset directly using the Hugging Face datasets library:

python
from datasets import load_dataset

# Standard loading
dataset = load_dataset("jimjunior/cocis-web-info")

# Streaming mode (Recommended for many shards)
streamed_dataset = load_dataset("cocis-web-info/cocis-web-info", streaming=True)
print(next(iter(streamed_dataset["train"])))

Maintenance and Contributions

This dataset was created as part of the 2026 undergraduate CSC Machine Learning assignment. Its actively mantained by Beingana Jim Junior.

Corresponding associated code used to collect and manage this data can be found at https://github.com/jim-junior/SW-ML-1-NLP-Project

Citation

text
@misc{junior2026dataset,
  author = {Jim Junior, B.},
  title = {513-Chunk JSON Dataset},
  year = {2026},
  publisher = {Hugging Face},
  journal = {Hugging Face Hub},
}