Team Ai
20 results

agent-skill

Vineethsain /agent-skill-security-paper-artifacts Agent Skill Security Research Artifacts This collection releases derived aggregate evidence, original figure data and plots, offline analysis code, and reproducibility manifests. The current research portfolio has 4 consolidated empirical manuscript directions: scanner configuration/gate comparisons, small-model score interfaces and source transfer, runtime cascade action composition, and deterministic-rule maintenance/evidence contracts. Current authoring/delivery state: Four… See the full description on the dataset page: https://huggingface.co/datasets/Vineethsain/agent-skill-security-paper-artifacts.0 likes332 downloads8d agoHugging FaceLucioLiu /agent-skills Index — Lucio's Agent Skills & Projects Each project now lives in its own repo, so you get its full README, its own licence, and its own version history. This page is just the map. This repo also keeps a full snapshot of every skill for anyone who wants them all in one download — see the Files and versions tab. The individual repos below are the canonical ones. Agent Skills Skill What it does Licence relic Portable AI personality & memory, in pure… See the full description on the dataset page: https://huggingface.co/datasets/LucioLiu/agent-skills.text-generationn<1K1 likes231 downloads2mo agoHugging Faceobaydata /claude-agent-skills-benchmark Claude Agent Skills Benchmark Claude Agent Skills 评测数据集 Description A benchmark dataset for evaluating whether LLMs can accurately trigger and execute domain-specific Skills on the Claude Code platform. Skills are designed by vertical domain experts with varying complexity levels (based on attachments: scripts, references, assets, and reference markdown files). Evaluation Scenarios Cover: Office automation, coding, investment promotion, financial services, industrial… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/claude-agent-skills-benchmark.documenttext-generationn<1K1 likes198 downloads6mo agoHugging FaceNPULH /agentskillos-benchmark AgentSkillOS Benchmark Official benchmark release for Organizing, Orchestrating, and Benchmarking Agent Skills at Ecosystem Scale. GitHub project: https://github.com/ynulihao/AgentSkillOS This release is a file-based benchmark package. It contains the original task definitions and the bundled task assets that are needed by tasks with local input files. Contents data/tasks.jsonl: normalized task index for load_dataset(...) tasks/: the original 30 benchmark task… See the full description on the dataset page: https://huggingface.co/datasets/NPULH/agentskillos-benchmark.documentn<1K1 likes188 downloads7mo agoHugging FaceAgent-Threat-Rule /atr-skill-benchmark ATR Skill-Security Benchmark A labeled corpus of SKILL.md files for evaluating detection of malicious agent skills — prompt injection, tool poisoning, credential theft, malware droppers and supply-chain attacks hidden inside natural-language agent instructions. Published as part of Agent Threat Rules (ATR), an open, vendor-neutral detection standard for AI agents (like Sigma, but for agent attacks). Why this exists SKILL.md files are natural-language instructions… See the full description on the dataset page: https://huggingface.co/datasets/Agent-Threat-Rule/atr-skill-benchmark.texttext-classificationn<1K2 likes94 downloads3mo agoHugging Faceyoonholee /agent-skill-malware Agent Skill Malware: Malicious vs Benign Agent Skills Binary classification dataset of OpenClaw agent skill files (SKILL.md). 127 malicious + 223 benign = 350 samples, deduplicated by content hash. Source Malicious: Real skills from malicious campaigns targeting ClawHub (Feb 2026), extracted from the openclaw/skills GitHub archive. They use social engineering in markdown instructions to trick agents/users into running malware -- primarily AMOS (Atomic macOS Stealer)… See the full description on the dataset page: https://huggingface.co/datasets/yoonholee/agent-skill-malware.texttext-classificationn<1K2 likes91 downloads8mo agoHugging Face