Team Ai
Datasetpublic

zhangdw/astra-skills-scripts

๐Ÿ› ๏ธ ASTRA Skills Scripts: Script-Backed Skills from the Agent Skill Tool-use Repository Atlas Script-backed skills from the Agent Skill Tool-use Repository Atlas      ASTRA Skills Scripts contains only skill directories from zhangdw/astra-skills that include a scripts/ subdirectory, making it easier to study agent skills that pair written instructions with runnable helper code. Quick Start ยท At a Glance ยท Subset Definition ยทโ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/zhangdw/astra-skills-scripts.

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
1likes46downloads
Dataset Card

๐Ÿ› ๏ธ ASTRA Skills Scripts: Script-Backed Skills from the Agent Skill Tool-use Repository Atlas

Script-backed skills from the Agent Skill Tool-use Repository Atlas

<p align="center" style="text-align:center; white-space:nowrap;"><a href="https://huggingface.co/datasets/zhangdw/astra-skills-scripts"><img src="https://img.shields.io/badge/Hugging%20Face-Dataset-yellow?logo=huggingface" alt="Hugging Face Dataset" style="display:inline-block; vertical-align:middle; margin:0 3px;"></a>&nbsp;<img src="https://img.shields.io/badge/ASTRA-Agent%20Skill%20Tool--use%20Repository%20Atlas-6c5ce7" alt="ASTRA: Agent Skill Tool-use Repository Atlas" style="display:inline-block; vertical-align:middle; margin:0 3px;">&nbsp;<img src="https://img.shields.io/badge/Script%20Skills-28%2C954-00b894" alt="28,954 script-backed skills" style="display:inline-block; vertical-align:middle; margin:0 3px;">&nbsp;<img src="https://img.shields.io/badge/Subset-19.55%25%20of%20ASTRA%20Skills-blue" alt="19.55% of ASTRA Skills" style="display:inline-block; vertical-align:middle; margin:0 3px;">&nbsp;<img src="https://img.shields.io/badge/Archive-1.84%20GiB-2d3436" alt="1.84 GiB archive" style="display:inline-block; vertical-align:middle; margin:0 3px;"></p>

<p align="center"> <b>ASTRA Skills Scripts</b> contains only skill directories from <code>zhangdw/astra-skills</code> that include a <code>scripts/</code> subdirectory, making it easier to study agent skills that pair written instructions with runnable helper code. </p>

<p align="center"> <a href="#-quick-start"><b>Quick Start</b></a> ยท <a href="#-dataset-at-a-glance"><b>At a Glance</b></a> ยท <a href="#-subset-definition"><b>Subset Definition</b></a> ยท <a href="#-directory-format"><b>Directory Format</b></a> </p>


[!IMPORTANT] This dataset is a filtered subset of the public <a href="https://huggingface.co/datasets/zhangdw/astra-skills">ASTRA Skills Collection</a>. The dataset-level metadata and packaging are released under Apache-2.0, but individual skill files and scripts remain subject to their original repository licenses. Inspect upstream licenses before redistribution, execution, training, or benchmark release.

โœจ Why a Scripts Subset?

Many agent skills are pure instruction documents. Others include executable helpers, shell scripts, Python utilities, templates, or workflow code under scripts/. The scripts-backed subset is useful when the research question is not only what agents are told to do, but how skills operationalize those instructions through local tools.

This subset helps answer questions like:

  • โ€”Which real-world skills rely on runnable helper scripts?
  • โ€”How do `SKILL.md` instructions describe when and how to call scripts?
  • โ€”What tool-use patterns appear around code generation, browser work, document editing, data processing, and automation?
  • โ€”How can retrieval systems surface skills that include executable affordances, not just prose instructions?

๐Ÿ“ฆ Dataset at a Glance

<table> <tr> <td align="center"><b>28,954</b><br/>skills with <code>scripts/</code></td> <td align="center"><b>4</b><br/>split archive parts</td> <td align="center"><b>1.84 GiB</b><br/>compressed archive</td> <td align="center"><b>Strict Subset</b><br/>of <code>zhangdw/astra-skills</code></td> </tr> <tr> <td align="center"><b>19.55%</b><br/>of the full corpus</td> <td align="center"><b>2026-04-15</b><br/>extraction date</td> <td align="center"><b>GitHub</b><br/>source lineage</td> <td align="center"><b>Script-Backed</b><br/>agent workflows</td> </tr> </table>

Snapshot Summary

FieldValue
Source datasetzhangdw/astra-skills
Full corpus size148,134 deduplicated skills
Matched skills with scripts/28,954
Subset share19.55% of the full corpus
Extraction date2026-04-15
Source root scannedastra-skills/github/
Copied subset rootgithub/
Archive layoutastra-skills-scripts-part-0000.tar.gz โ†’ astra-skills-scripts-part-0003.tar.gz

๐Ÿงฉ Subset Definition

RuleDescription
Parent datasetThe full <a href="https://huggingface.co/datasets/zhangdw/astra-skills"><code>zhangdw/astra-skills</code></a> corpus.
Selection predicateKeep a skill directory only if it contains a scripts/ subdirectory.
Directory contentsPreserve the skill directory, including SKILL.md, _meta.json, scripts/, and other copied files from the parent corpus.
ScopeA scripts-focused research subset, not a new crawl or live mirror.

๐Ÿ—‚๏ธ Dataset Files

FileDescription
README.mdDataset card and usage notes.
astra-skills-scripts-part-0000.tar.gz โ†’ astra-skills-scripts-part-0003.tar.gzSplit compressed archive containing the filtered github/ skill directory tree.
.gitattributesHugging Face / Git LFS tracking metadata.

๐Ÿš€ Quick Start

Download the dataset with the current Hugging Face Hub CLI:

bash
uvx --from huggingface_hub hf download zhangdw/astra-skills-scripts \
  --type dataset \
  --local-dir astra-skills-scripts

Merge the split archive and extract it:

bash
cat astra-skills-scripts/astra-skills-scripts-part-*.tar.gz > astra-skills-scripts.tar.gz
tar xzf astra-skills-scripts.tar.gz

This produces a github/ directory containing script-backed skill directories.

Inspect a few script-backed skills:

bash
find github -path '*/scripts' -type d | head
find github -name SKILL.md | head

๐Ÿงฑ Directory Format

Each saved skill directory comes from the parent ASTRA Skills corpus and includes a scripts/ subdirectory.

text
github/
โ””โ”€โ”€ {source}_{owner}_{skill_name}/
    โ”œโ”€โ”€ SKILL.md
    โ”œโ”€โ”€ _meta.json
    โ”œโ”€โ”€ scripts/
    โ””โ”€โ”€ [optional files, e.g. templates/, examples/, assets/]

_meta.json stores provenance metadata from the original crawl, such as source site, repository URL, owner, repository name, skill name, and relative path inside the upstream repository.

๐Ÿ”Ž Example Discovery Workflows

<details open> <summary><b>List common script file extensions</b></summary>

python
from collections import Counter
from pathlib import Path

counts = Counter()

for scripts_dir in Path("github").rglob("scripts"):
    if scripts_dir.is_dir():
        for path in scripts_dir.rglob("*"):
            if path.is_file():
                counts[path.suffix.lower() or "[no extension]"] += 1

print(counts.most_common(20))

</details>

<details> <summary><b>Find skills whose instructions mention script execution</b></summary>

bash
rg -n "scripts/|run .*script|execute|python|bash|node" github -g 'SKILL.md'

</details>

<details> <summary><b>Build a lightweight metadata table</b></summary>

python
import json
from pathlib import Path

rows = []

for meta_file in Path("github").rglob("_meta.json"):
    skill_dir = meta_file.parent
    meta = json.loads(meta_file.read_text())
    rows.append({
        "skill_dir": str(skill_dir),
        "script_files": sum(1 for p in (skill_dir / "scripts").rglob("*") if p.is_file()),
        "meta": meta,
    })

print(len(rows))
print(rows[:3])

</details>

โœ… Intended Use

ASTRA Skills Scripts is designed for:

  • โ€”research on script-backed agent workflows and executable skill affordances;
  • โ€”skill retrieval and routing experiments that need a has_scripts signal;
  • โ€”analysis of how SKILL.md instructions coordinate with helper code;
  • โ€”studying tool invocation patterns, script conventions, and repository provenance;
  • โ€”constructing smaller benchmark or training subsets from the full ASTRA Skills corpus.

๐Ÿ› ๏ธ Maintenance Notes

This dataset is a static extraction from the full ASTRA Skills snapshot. It preserves the parent corpus layout while filtering for directories that include scripts/. Because scripts can execute arbitrary upstream code, treat this dataset as research material: inspect code before running it, sandbox execution, and retain upstream provenance when deriving new artifacts.

๐Ÿ‘ค Author

  • โ€”Dawei Zhang (GitHub: zhangdw156)

๐Ÿ“š Citation

If you use ASTRA Skills Scripts in research, please cite this dataset, the full ASTRA Skills Collection, and any upstream repositories whose skill contents or scripts are central to your analysis.

bibtex
@misc{astraSkillsScripts2026,
  author       = {Dawei Zhang},
  title        = {ASTRA Skills Scripts: Script-Backed Skills from the Agent Skill Tool-use Repository Atlas},
  year         = {2026},
  howpublished = {Hugging Face Dataset},
  publisher    = {Hugging Face},
  doi          = {10.57967/hf/8428},
  url          = {https://huggingface.co/datasets/zhangdw/astra-skills-scripts},
  note         = {Scripts-focused subset of ASTRA Skills; snapshot date: 2026-04-15; DOI record revision: 77d9d29}
}

For completeness, also cite the parent collection:

bibtex
@misc{astraSkills2026,
  author       = {Dawei Zhang},
  title        = {ASTRA Skills: Agent Skill Tool-use Repository Atlas},
  year         = {2026},
  howpublished = {Hugging Face Dataset},
  publisher    = {Hugging Face},
  doi          = {10.57967/hf/8399},
  url          = {https://huggingface.co/datasets/zhangdw/astra-skills},
  note         = {Snapshot date: 2026-04-15; DOI record revision: 146eb8c}
}

๐Ÿ“„ License

Dataset metadata and packaging are released under Apache-2.0. Individual skill contents and scripts remain subject to their original repository licenses.


<div align="center">

<b>ASTRA Skills Scripts makes script-backed agent skills easier to inspect, compare, and study at scale.</b>

</div>