Team Ai
Datasetpublic

bigcode/the-stack-smol-xl

Dataset Description A small subset of the-stack dataset, with 87 programming languages, each has 10,000 random samples from the original dataset. Languages The dataset contains 87 programming languages: 'ada', 'agda', 'alloy', 'antlr', 'applescript', 'assembly', 'augeas', 'awk', 'batchfile', 'bison', 'bluespec', 'c', 'c++', 'c-sharp', 'clojure', 'cmake', 'coffeescript', 'common-lisp', 'css', 'cuda', 'dart', 'dockerfile', 'elixir', 'elm', 'emacs-lisp','erlang'… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-smol-xl.

sourceHugging Faceupdated 4y agoView on Hugging Face
11likes9.8kdownloads
settings

This repository belongs to bigcode on Hugging Face.

Team Ai never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namethe-stack-smol-xl
visibilitypublic
licencenot set
gatedno
ownerbigcode
Account settings
bigcode/the-stack-smol-xl · Team Ai