Team Ai
20 results

tabular

chewwt /po_qwen14b_tabular_data BoLT Prompt Optimization — Tabular Dataset For prompt optimization tasks in BoLT, an accessible benchmark for black-box optimization on LLM tasks. Dataset Description The dataset covers 5,014 evaluated instructions. Each row is a candidate system-prompt instruction paired with its empirically measured MATH-500 (4-shot, non-thinking mode) scores. Evaluation details: Model: Qwen/Qwen3-14B Task: minerva_math500 (4-shot) (from lm-eval library) System prompt:… See the full description on the dataset page: https://huggingface.co/datasets/chewwt/po_qwen14b_tabular_data.tabulartext-generation1K<n<10K1 likes21k downloads5mo agoHugging Faceinria-soda /tabular-benchmark Tabular Benchmark Dataset Description This dataset is a curation of various datasets from openML and is curated to benchmark performance of various machine learning algorithms. Repository: https://github.com/LeoGrin/tabular-benchmark/community Paper: https://hal.archives-ouvertes.fr/hal-03723551v2/document Dataset Summary Benchmark made of curation of various tabular data learning tasks, including: Regression from Numerical and Categorical Features… See the full description on the dataset page: https://huggingface.co/datasets/inria-soda/tabular-benchmark.tabulartabular-classification10M<n<100M51 likes2.2k downloads3y agoHugging Faceserval-uni-lu /tabularbench0 likes1.7k downloads1y agoHugging FaceMachineLearningLM /machinelearninglm-scm-synthetic-tabularml MachineLearningLM Pretraining Corpus This repository contains the pretraining corpus for MachineLearningLM, a framework designed to equip large language models (LLMs) with robust in-context machine learning (ML) capabilities. The dataset consists of ML tasks synthesized from millions of structural causal models (SCMs), spanning various shot counts up to 1,024. It is designed to enable LLMs to learn from many in-context examples on standard ML tasks purely via in-context learning… See the full description on the dataset page: https://huggingface.co/datasets/MachineLearningLM/machinelearninglm-scm-synthetic-tabularml.texttext-generation1M<n<10M4 likes979 downloads10mo agoHugging FaceLAMDA-Tabular /CausalArena CausalArena public release This repository contains the public CausalArena dataset release: executable SCMs, selected result tables, and real-data source indices. What is included scm/: the public half of each generated SCM family: 500 synthetic SCM configurations, 50 semantic SCMs, and 50 formula-grounded SCMs. Released SCMs include both observation-only and observation-plus-intervention exports. scm/{semantic,formula}/artifacts/: per-scenario graph, generator… See the full description on the dataset page: https://huggingface.co/datasets/LAMDA-Tabular/CausalArena.texttabular-classification1K<n<10K1 likes650 downloads11d agoHugging FaceLAMDA-Tabular /TALENT Benchmark Dataset for "A Closer Look at Deep Learning on Tabular Data" This repository contains supplemental datasets for the paper "A Closer Look at Deep Learning on Tabular Data". The datasets are provided in two zip files: benchmark_dataset.zip and training_dynamic_informations.zip. 1. benchmark_dataset.zip This zip file contains all the tabular datasets used in the paper. Each dataset is stored in a separate subfolder named after the dataset. Each dataset folder… See the full description on the dataset page: https://huggingface.co/datasets/LAMDA-Tabular/TALENT.1 likes646 downloads1y agoHugging Face