Team Ai
Datasetpublic

NickIBrody/ruby-code-corpus

Ruby Code Corpus A large corpus of Ruby source code collected from public GitHub repositories. Content 294,074 files from 3,920+ Ruby repositories on GitHub Sourced from repos ranked by star count (quality signal) Filtered: removed files under 200 bytes (trivial/empty files) Only permissive licenses: MIT, Apache-2.0, BSD, ISC, Ruby Fields Field Description source Always github repo owner/repo slug repo_url Full GitHub URL path… See the full description on the dataset page: https://huggingface.co/datasets/NickIBrody/ruby-code-corpus.

sourceHugging Faceotherupdated 5mo agoView on Hugging Face
0likes16downloads
Dataset Card

Ruby Code Corpus

A large corpus of Ruby source code collected from public GitHub repositories.

Content

  • —294,074 files from 3,920+ Ruby repositories on GitHub
  • —Sourced from repos ranked by star count (quality signal)
  • —Filtered: removed files under 200 bytes (trivial/empty files)
  • —Only permissive licenses: MIT, Apache-2.0, BSD, ISC, Ruby

Fields

FieldDescription
sourceAlways github
repoowner/repo slug
repo_urlFull GitHub URL
pathFile path within the repo
languageAlways Ruby
licenseSPDX license identifier
starsGitHub star count at collection time
refBranch name
size_bytesFile size in bytes
textRaw source content

Composition

  • —~95% .rb files
  • —~37% test files (spec/test — kept intentionally, useful for training)
  • —~58% production code

Usage

python
from datasets import load_dataset

ds = load_dataset("NickIBrody/ruby-code-corpus", split="train")
print(ds[0]["text"])

License

Individual files retain their original open-source licenses as specified in the license field.