Team Ai
20 results

machine-learning

RGES-PIT /MachineLearning Machine Learning Tier This dataset is a collection of synthetic microlensing light curves from the Nancy Grace Roman Space Telescope Galactic Bulge Time Domain Survey. It is intended for the training and benchmarking of machine learning models for microlensing event classification, parameter estimation, and anomaly detection. The raw distribution of event properties is not representative of what Roman will see, but should span a statistically larger set of events. More… See the full description on the dataset page: https://huggingface.co/datasets/RGES-PIT/MachineLearning.tabular10B<n<100B1 likes1.5k downloads2mo agoHugging FaceMachineLearningLM /machinelearninglm-scm-synthetic-tabularml MachineLearningLM Pretraining Corpus This repository contains the pretraining corpus for MachineLearningLM, a framework designed to equip large language models (LLMs) with robust in-context machine learning (ML) capabilities. The dataset consists of ML tasks synthesized from millions of structural causal models (SCMs), spanning various shot counts up to 1,024. It is designed to enable LLMs to learn from many in-context examples on standard ML tasks purely via in-context learning… See the full description on the dataset page: https://huggingface.co/datasets/MachineLearningLM/machinelearninglm-scm-synthetic-tabularml.texttext-generation1M<n<10M4 likes953 downloads10mo agoHugging FaceHuman-Centric-Machine-Learning /tokenization-multiplicity-data Dataset: Tokenization Multiplicity Leads to Arbitrary Price Variation in LLM-as-a-service This dataset contains the official experiment inference traces for the paper Tokenization Multiplicity Leads to Arbitrary Price Variation in LLM-as-a-service by Ivi Chatzi, Nina Corvelo Benz, Stratis Tsirtsis and Manuel Gomez-Rodriguez. 📂 Dataset Structure The dataset is organized into folders as follows: .\{model}\{task}\{lang}\{seed}_{10*temperature}.jsonl where {model}… See the full description on the dataset page: https://huggingface.co/datasets/Human-Centric-Machine-Learning/tokenization-multiplicity-data.text-generation10K<n<100K2 likes710 downloads7mo agoHugging Facemjphayes /machine_learning_questions Dataset Card for "machine_learning_questions" More Information needed textn<1K7 likes345 downloads3y agoHugging FaceHuman-Centric-Machine-Learning /strategic-ttc-data Dataset: Strategic Test-Time Compute (TTC) This dataset contains the official experiment inference traces for the paper "Test-Time Compute Games" (arXiv:2601.21839). It includes full model generations, token counts, and correctness verifications for various Large Language Models (LLMs) across three major reasoning benchmarks: GSM8K, AIME, and GPQA. This data allows researchers to analyze the relationship between test-time compute and model performance without needing to re-run… See the full description on the dataset page: https://huggingface.co/datasets/Human-Centric-Machine-Learning/strategic-ttc-data.question-answering10K<n<100K2 likes235 downloads8mo agoHugging FaceBAAI /IndustryCorpus2_artificial_intelligence_machine_learning IndustryCorpus2: Artificial Intelligence This repository contains the IndustryCorpus2: Artificial Intelligence domain subset of BAAI/IndustryCorpus2. Refer to the parent dataset card for data construction, intended use, limitations, and licensing details. Citation If you use this dataset in your work, please cite IndustryCorpus2: @misc{shi2024industrycorpus2, title = {IndustryCorpus2}, author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao}… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus2_artificial_intelligence_machine_learning.2 likes216 downloads2mo agoHugging Face