optimized
oercommons-v1-optimized
OERCommons v1 Optimized
Authors: Junjie Wang and Yuhan SunHosted by: PIN TeamDataset: pin-team/oercommons-v1-optimized
OERCommons v1 Optimized is a provenance-preserving multimodal pretraining corpus built on the OERCommons subset of The Common Pile v0.1, which serves as its upstream data and licensing baseline. We extend it with full-page recovery, canonical Markdown, ordered image/PDF/link metadata, conservative corrections, and integrity evidence.
At a glance… See the full description on the dataset page: https://huggingface.co/datasets/pin-team/oercommons-v1-optimized.government-primary-source-optimized
Government Primary Source — Optimized
A deterministic, audited reduction of a 2007–2024 U.S. federal contract panel,
derived from sovai/government_contracts.
Read the Status and Retractions sections before using any
number from this repository. Several figures that were previously published here are
retracted, and the retractions are load-bearing: several of the retracted figures are the
ones most likely to be quoted.
Status
Item
State
Newest executed… See the full description on the dataset page: https://huggingface.co/datasets/sirbrentmichaelskoda/government-primary-source-optimized.spro-optimized-prompts-fullOptimized_Video_Facial_Landmarks
Dataset Card for 478-Point Normalized 3D Facial Landmark Dataset
Dataset Description
This dataset provides pre-extracted, normalized 3D facial landmark features derived from the Video Emotion dataset. It is optimized for efficient training of emotion recognition and facial analysis models, bypassing the need to process large raw video files.
License: The extracted feature data in this Parquet file is licensed under Apache 2.0. Note that the original source video files may… See the full description on the dataset page: https://huggingface.co/datasets/PSewmuthu/Optimized_Video_Facial_Landmarks.details_chanwit__flux-base-optimized
Dataset Card for Evaluation run of chanwit/flux-base-optimized
Dataset automatically created during the evaluation run of model chanwit/flux-base-optimized on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_chanwit__flux-base-optimized.hf_dataset_shards_optimized_new
