Team Ai
20 results

validation

IgnisCogitationis /quantum-like-attention-framework-1.3b-untuned-validation Quantum-Like Attention Framework (QLAF) 1.3B Untuned Pretraining & Scaling Proof This repository hosts the pretraining checkpoints, scaling logs, and downstream evaluation benchmarks for the 1.3B parameter Quantum-Like Attention Framework (QLAF) with Hybrid FlashAttention (75% recurrent QLAF / 25% causal FlashAttention) across a 3-seed validation campaign on dedicated A100 Large GPU hardware. Multi-Seed Pretraining & Downstream Evaluation Leaderboard Seed… See the full description on the dataset page: https://huggingface.co/datasets/IgnisCogitationis/quantum-like-attention-framework-1.3b-untuned-validation.3 likes56k downloads1m agoHugging FaceAI-MO /aimo-validation-aime Dataset Card for AIMO Validation AIME All 90 problems come from AIME 22, AIME 23, and AIME 24, and have been extracted directly from the AOPS wiki page https://artofproblemsolving.com/wiki/index.php/AIME_Problems_and_Solutions This dataset serves as an internal validation set during our participation in the AIMO progress prize competition. Using data after 2021 is to avoid potential overlap with the MATH training set. Here are the different columns in the dataset: problem: the… See the full description on the dataset page: https://huggingface.co/datasets/AI-MO/aimo-validation-aime.textn<1K69 likes27k downloads1y agoHugging FaceD4nt3 /esb-datasets-earnings22-validation-tiny-filteredA filtered (<=30s duration) slice (512 samples) of the Earnings22 dataset. def add_duration(sample): y, sr = sample['audio']["array"], sample['audio']["sampling_rate"] sample['duration_ms']=librosa.get_duration(y=y, sr=sr) * 1000 return sample tedlium = load_dataset("esb/datasets", "earnings22", split='validation', trust_remote_code=True) # compute duration to filter tedlium = tedlium.map(add_duration) tedlium = tedlium.select(range(512)) # Whisper max supported duration tedlium… See the full description on the dataset page: https://huggingface.co/datasets/D4nt3/esb-datasets-earnings22-validation-tiny-filtered.audion<1K0 likes9.9k downloads2y agoHugging FaceAI-MO /aimo-validation-amc Dataset Card for AIMO Validation AMC All 83 come from AMC12 2022, AMC12 2023, and have been extracted from the AOPS wiki page https://artofproblemsolving.com/wiki/index.php/AMC_12_Problems_and_Solutions This dataset serves as an internal validation set during our participation in the AIMO progress prize competition. Using data after 2021 is to avoid potential overlap with the MATH training set. Here are the different columns in the dataset: problem: the modified problem statement… See the full description on the dataset page: https://huggingface.co/datasets/AI-MO/aimo-validation-amc.tabularn<1K19 likes5.1k downloads1y agoHugging FaceTsomaros /Imagenet-1k_validationimage10K<n<100K0 likes3.8k downloads2y agoHugging FacePaulineLi /QuantiPhy-validation QuantiPhy (Validation Set) Dataset Summary QuantiPhy is a benchmark for evaluating whether vision–language models (VLMs) can perform quantitative physical inference from visual evidence, rather than producing plausible but ungrounded numerical guesses. This repository contains the official validation set of QuantiPhy, released to support model development, ablation studies, and preliminary evaluation.The validation set represents approximately 4% of the full… See the full description on the dataset page: https://huggingface.co/datasets/PaulineLi/QuantiPhy-validation.tabularvideo-text-to-textn<1K8 likes1.7k downloads1d agoHugging Face