Team Ai
20 results

autoencoder

P2SAMAPA /p2-etf-variational-autoencoder-resultstextn<1K0 likes368 downloads4mo agoHugging Facejusticedao /legal-ir-autoencoder-checkpoints Legal IR Autoencoder Checkpoints This dataset stores checkpoint artifacts for the legal text -> formal logic IR autoencoder/Codex optimization loop. Latest checkpoint in this upload: checkpoints/20260630T221836Z/. Contents state/legal-ir-autoencoder-canonical.state.json: canonical feature-level autoencoder warm-start state. reports/: weight review, deprecation manifest, and consensus feature manifests. scripts/review_autoencoder_weight_runs.py: script used to… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/legal-ir-autoencoder-checkpoints.text-generation0 likes319 downloads6d agoHugging FaceMechanophoresResearch /AutoencoderDataset AutoencoderDataset Data for training molvae — a SELFIES molecular VAE (matryoshka latent) for Bayesian optimization of mechanophores. Two layers: raw/ — the source chemical databases (SMILES/SELFIES, pre-tokenized), as collected. mixes/ — the derived training datasets: weighted, shuffled, materialized token-shard "mixes" built from the raw sources. Each mix is documented by its config + README + build code, so it is fully reproducible. The multi-GB token shards themselves live… See the full description on the dataset page: https://huggingface.co/datasets/MechanophoresResearch/AutoencoderDataset.0 likes239 downloads3mo agoHugging Facekeyvanatt /laplace-autoencoders-dataset Transient CH₄ concentration fields Simulation data used by Laplace autoencoders: transient CH₄ concentration fields U(t) as a function of physical parameters θ = (k, A, C). The simulations were generated with OpenFOAM. Code: github.com/keyvanatt/laplace-autoencoders Checkpoints: keyvanatt/laplace-autoencoders-checkpoints Files File Shape Content doe.npy (225,) structured (k, A, C) design of experiments: 15 values of k × 15 values of C, A = 0… See the full description on the dataset page: https://huggingface.co/datasets/keyvanatt/laplace-autoencoders-dataset.1 likes75 downloads5d agoHugging Facejpwahle /autoencoder-paraphrase-dataset Dataset Card for Machine Paraphrase Dataset (MPC) Dataset Summary The Autoencoder Paraphrase Corpus (APC) consists of ~200k examples of original, and paraphrases using three neural language models. It uses three models (BERT, RoBERTa, Longformer) on three source texts (Wikipedia, arXiv, student theses). The examples are aligned, i.e., we sample the same paragraphs for originals and paraphrased versions. How to use it You can load the dataset using the… See the full description on the dataset page: https://huggingface.co/datasets/jpwahle/autoencoder-paraphrase-dataset.tabulartext-classification1M<n<10M2 likes51 downloads1y agoHugging Facejusticedao /uscode-autoencoder-sparsetabular1K<n<10K0 likes25 downloads13d agoHugging Face