Team Ai
20 results

data-use

anaisleila /computer-use-data-psai Computer Use Dataset - PSAI A large-scale, multimodal dataset of human-computer interactions for training and evaluating AI agents. 🔗 Access Dataset: https://huggingface.co/datasets/anaisleila/computer-use-data-psai 📊 Dataset Overview This dataset contains 3,167 completed tasks of human-computer interactions captured with video, screenshots, DOM snapshots, and detailed interaction events. Created by Paradigm Shift AI for advancing computer use AI agent research.… See the full description on the dataset page: https://huggingface.co/datasets/anaisleila/computer-use-data-psai.imagereinforcement-learning1K<n<10K20 likes4.6k downloads1y agoHugging Facejin-ying-so-cute /ecommerce-user-behavior-datatabular10M<n<100M6 likes448 downloads3y agoHugging FaceKRMayD /COD10K_GMPO_Used_Data COD10K GMPO Used Data This is a research repack of the COD10K camouflaged-object (CAM) subset used in our CLIP DPO/GMPO experiments. It is not an official COD10K distribution. The package contains the exact original images, masks, generated negative images, and portable caption CSV files used for training and segmentation evaluation. All paths in the portable CSV files are relative to this dataset root. Data Split Split Contents Count Train Original CAM… See the full description on the dataset page: https://huggingface.co/datasets/KRMayD/COD10K_GMPO_Used_Data.imageimage-segmentation10K<n<100K0 likes337 downloads3mo agoHugging Facerafmacalaba /datause-extracted Data-use mentions (NER / span extraction) Data mentions extracted from World Bank Policy Research Working Papers and FCV documents, predicted by a span-extraction model with no human or LLM-judge validation, and formatted for span-extraction (GLiNER / GLiNER2) and token-classification (LFM2.5-encoder) fine-tuning. Labels Three entity types: NAMED_DATA — a proper name, title, or acronym of a specific data source DESCRIPTIVE_DATA — a source described in words but… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/datause-extracted.tabulartoken-classification100K<n<1M0 likes214 downloads1mo agoHugging Facerafmacalaba /data-use-sft-tiered Data-use SFT — tiered workflow (two task subsets) Multitask SFT anchored exclusively on mentions the tiered extractor emits (T1 evidential ∪ T2 declaration; see rafmacalaba/data-use-mentions-tiered). Every row carries task ("provenance" | "usage_impact") and origin (prwp | fcv). Rows whose anchor span was judged T3 (non-mention) or junk are dropped — audit trail in manifest.jsonl (provenance) and manifest_usage.jsonl (usage/impact). task = provenance (22,201 rows)… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/data-use-sft-tiered.texttext-generation10K<n<100K0 likes193 downloads1mo agoHugging Facesunweiwei /user-datatext10K<n<100K0 likes192 downloads5mo agoHugging Face