Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01eaddario /imatrix-calibration Importance Matrix Calibration Datasets This repository provides calibration datasets used to generate importance matrices (imatrix), which are required to minimize errors when quantizing models with LLaMA C++. The llama-imatrix program cannot handle parquet files directly and thus requires them to be converted into text format first. There are many ways to do this but a simple approach is to use DuckDB with the following command: duckdb -noheader -ascii -c "SELECT content FROM… See the full description on the dataset page: https://huggingface.co/datasets/eaddario/imatrix-calibration.texttext-generationn<1K69 likes7.6k downloads5mo agoHugging Face02gabriellarson /imatrix-storage1 likes2.8k downloads1y agoHugging Face03Thireus /imatrix0 likes1.8k downloads2mo agoHugging Face04ClonedGizzards /imatrix-calibration Importance Matrix Calibration Datasets This repository provides calibration datasets used to generate importance matrices (imatrix), which are required to minimize errors when quantizing models with LLaMA C++. The llama-imatrix program cannot handle parquet files directly and thus requires them to be converted into text format first. There are many ways to do this but a simple approach is to use DuckDB with the following command: duckdb -noheader -ascii -c "SELECT content FROM… See the full description on the dataset page: https://huggingface.co/datasets/ClonedGizzards/imatrix-calibration.texttext-generationn<1K0 likes281 downloads6d agoHugging Face05froggeric /imatrix Input files for generating the Importance Matrix Which file to use for generating the importance matrix Not all importance matrices are equal. The best results are obtained when using a source file similar to the training data. Size also matters: the bigger the model (eg: 70b vs 13b) and the higher the quant (eg: q6k_ vs iq3_xs), the bigger the source file needs to be to make an impact. Multiple input files can be combined if needed; for example: cat multilingual.txt… See the full description on the dataset page: https://huggingface.co/datasets/froggeric/imatrix.text10K<n<100K17 likes267 downloads2y agoHugging Face06magiccodingman /QwQ-32B-abliterated-131k-GGUF-Yarn-Imatrix QwQ-32B-Abliterated-131k-GGUF-Yarn-Imatrix High-Fidelity Semantic Simulation & Orchestration AI Model Will this pass the random stupid benchmarks that exist today? I don't know, nor care. I don't need my local AI model to know some random city capital of a foreign country. I need a local AI model that can simulate with high semantic fidelity. Why? Because your AI may be able to spit random facts. I want an AI that knows when to Google facts. I want an AI that tracks hundreds of… See the full description on the dataset page: https://huggingface.co/datasets/magiccodingman/QwQ-32B-abliterated-131k-GGUF-Yarn-Imatrix.10M<n<100M17 likes214 downloads1y agoHugging Face07cstr /crispasr-imatrix-calib CrispASR imatrix calibration set — Common Voice EN + DE A tiny, CC0, multilingual read-speech sample used to compute importance matrices (imatrix) for GGUF quantisation of ASR models with CrispASR. en/ — 24 English clips de/ — 24 German clips Provenance Clips are drawn from the dev split of Mozilla Common Voice 17.0 (via the fsicoli/common_voice_17_0 mirror), which is released under CC0 1.0 (public domain). Re-distributed here unchanged, same licence.… See the full description on the dataset page: https://huggingface.co/datasets/cstr/crispasr-imatrix-calib.audioautomatic-speech-recognitionn<1K0 likes140 downloads1d agoHugging Face08ikawrakow /imatrix-from-wiki-trainThis repository contains importance matrix datasets for use with the improved quantization methods recently added to llama.cpp. The importance matrix has been computed using wiki.train.raw as training data. Hope the file names are self-explanatory. To use, after cloning this repo, for e.g. Mixtral-8x7B and Q4_K_M quantization, use ./quantize --imatrix path_to_repo/mixtral-8x7b.imatrix path_to_model ggml-model-q4k-m.gguf Q4_K_M 15 likes122 downloads3y agoHugging Face09lemon07r /bartowski-imatrix-v5-semantic Bartowski iMatrix Calibration v5 (Semantic Chunking) A processed version of bartowski's v5 imatrix calibration data using semantic boundary detection optimized for the v5 data structure. Dataset Summary Metric Value Total samples 2,075 Chunking method V5-optimized semantic boundary detection Chunk size 200+ characters (no upper limit, preserves document integrity) Languages English, German, Spanish, French, Italian, Swedish, Russian, Arabic, Chinese… See the full description on the dataset page: https://huggingface.co/datasets/lemon07r/bartowski-imatrix-v5-semantic.texttext-generation1K<n<10K9 likes122 downloads8mo agoHugging Face10augustine223 /korean-imatrix-calibration-corpus Korean imatrix Calibration Corpus — KO-i1 보정 코퍼스 한국어 중심 imatrix 보정 코퍼스의 첫 공개 릴리스 (우리가 아는 한). 공개 GGUF 양자화 생태계의 importance matrix는 거의 전부 영어 위주 코퍼스로 수집됩니다. 그 결과 한국어 토큰 분포에서의 양자화 손실이 체계적으로 커집니다. 이 데이터셋은 그 공백을 메우기 위해 만들어졌고, 실측으로 효과가 입증됐습니다. 실측 효과 (이 코퍼스로 만든 KO-i1 릴리스들) 릴리스 비교 대상 결과 kanana-1.5-8b KO-i1 영어 보정 i1 저비트 KLD -5~6% (IQ2_M 3.3σ), 비트 낮을수록 이득 증가 Qwen3.6-35B-A3B KO-i1 영어 보정 i1 전 타입 우세, -5.1~-6.8% (최대 4.3σ), MoE는 4비트도 유의 Qwen3.8-27B-abl KO-i1 정적 양자… See the full description on the dataset page: https://huggingface.co/datasets/augustine223/korean-imatrix-calibration-corpus.textn<1K0 likes97 downloads2mo agoHugging Face11k0ndra /imatrix-ja-en-v3 Japanese-English imatrix Calibration Data v3 imatrix計算用のキャリブレーションデータです。日本語LLMのGGUF量子化品質向上を目的として作成しました。本データセットは下記「ライセンス」欄に記載したデータセット群から派生した二次的著作物です。 構成 カテゴリ 割合 内容 ja_general 20% 日本語一般文章 ja_qa 20% 日本語Q&A・対話 ja_technical 10% 日本語技術・学術文 code 15% プログラムコード en_reasoning 15% 英語推論・数学 en_qa 5% 英語Q&A・対話 zh 4% 中国語 eu_lang 6% フランス語・ドイツ語・スペイン語(各2%) structured 5% SQL・構造化データ 目標トークン数/チャンク: 512(全チャンク 512〜640 トークンに収まるよう生成) 対話系カテゴリ(ja_qa /… See the full description on the dataset page: https://huggingface.co/datasets/k0ndra/imatrix-ja-en-v3.text10K<n<100K0 likes97 downloads16d agoHugging Face12imatrixlee /koch_testThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "koch", "total_episodes": 53, "total_frames": 14989, "total_tasks": 1, "total_videos": 159, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:53" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/imatrixlee/koch_test.tabularrobotics10K<n<100K0 likes96 downloads2y agoHugging Face13imatrixlee /koch_placeThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "koch", "total_episodes": 17, "total_frames": 6306, "total_tasks": 1, "total_videos": 51, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:17" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/imatrixlee/koch_place.tabularrobotics1K<n<10K0 likes96 downloads2y agoHugging Face14TFMC /imatrix-dataset-for-japanese-llmtexttext-generationn<1K35 likes65 downloads2y agoHugging Face15Thireus /imatrix-corpustext10K<n<100K0 likes62 downloads5mo agoHugging Face16DataSoul /chinese-imatrix-data-and.datThese data are utilized for the imatrix in llama.cpp, thereby maintaining model capability in low-precision quantization like IQ3-XXS. Most of the data is in Chinese or translated to Chinese; performance in other languages is not guaranteed (although some level of understanding may still be achievable). I have not tested any language other than Chinese. If anyone has, please feel free to comment. include: some data from: m-a-p/COIG-CQIA some data from:… See the full description on the dataset page: https://huggingface.co/datasets/DataSoul/chinese-imatrix-data-and.dat.text1K<n<10K2 likes59 downloads9mo agoHugging Face17lemon07r /bartowski-imatrix-v3-semantic Bartowski iMatrix Calibration v3 (Semantic Chunking) A processed version of bartowski's v3 imatrix calibration data using semantic boundary detection in attempt to create coherent, non-overlapping samples. Dataset Summary Metric Value Total samples 168 Chunking method Semantic boundary detection Target chunk size ~2048 characters Languages English, German, Spanish, French, Italian, Swedish, Russian, Arabic, Chinese Source Data The… See the full description on the dataset page: https://huggingface.co/datasets/lemon07r/bartowski-imatrix-v3-semantic.texttext-generationn<1K1 likes51 downloads8mo agoHugging Face18k0ndra /imatrix-ja-en Japanese-English imatrix Calibration Data imatrix計算用のキャリブレーションデータです。日本語LLMのGGUF量子化品質向上を目的として作成しました。 本データセットは下記「ライセンス」欄に記載したデータセット群から派生した二次的著作物です。 構成 カテゴリ 割合 内容 ja_general 35% 日本語一般文章 ja_qa 20% 日本語Q&A・対話 ja_technical 10% 日本語技術・学術文 code 15% プログラムコード en_reasoning 15% 英語推論・知識文 structured 5% SQL・構造化データ 目標トークン数/チャンク: 512 ファイル ファイル チャンク数 用途 imatrix-ja-en-500-shuffled.txt 500 チャンクをシャッフル済み(推奨) imatrix-ja-en-500-raw.txt 500… See the full description on the dataset page: https://huggingface.co/datasets/k0ndra/imatrix-ja-en.text10K<n<100K0 likes51 downloads6mo agoHugging Face19NLPark /chinese-imatrix-data-reasoningtext100K<n<1M0 likes49 downloads2y agoHugging Face20ChiTako /japanese-imatrix-calibration Japanese imatrix Calibration Dataset (calibration_ja) llama.cppのllama-imatrix用、日本語LLM向けキャリブレーションデータセット。 概要 このデータセットは、日本語LLMの量子化(quantization)における精度維持のために、llama-imatrixで使用するキャリブレーションデータを目的として構築されました。 統計 項目 値 チャンク数 916 総文字数 400,191 推定トークン数 ~200,096 ソース別内訳 ソース 文字数 割合 元のデータセット wikipedia_ja 82,155 (20.5%) wikimedia/wikipedia CC BY-SA 4.0 c4_ja 40,122 (10.0%) allenai/c4 CC BY 4.0 fineweb_ja 34,363 (8.6%)… See the full description on the dataset page: https://huggingface.co/datasets/ChiTako/japanese-imatrix-calibration.texttext-generation1K<n<10K0 likes43 downloads6mo agoHugging Face21OpenIntelligenceNet /Uncensored-Imatrix-Calibration-Mixed-Data Uncensored-Imatrix-Calibration-Mixed-Data (5,000 Samples) Overview This dataset contains 5,000 multi-domain conversational samples engineered specifically for importance matrix (imatrix) computation during low-bit LLM quantization (GGUF, AWQ, EXL2). Calibrating on generic corpora (such as raw Wikipedia dumps) frequently causes safety-induced logic degradation and lobotomizes unaligned behavior, as standard quantizers assign low activation sensitivity to… See the full description on the dataset page: https://huggingface.co/datasets/OpenIntelligenceNet/Uncensored-Imatrix-Calibration-Mixed-Data.texttext-generation1K<n<10K0 likes41 downloads11d agoHugging Face22yayoimizuha /new-imatrix-dataset-ja-en Dataset Card for Dataset Name 日英LLM向けのimatrix蒸留用データセットです。 既存のデータセットとしてはTFMC/imatrix-dataset-for-japanese-llmがありますが、 テキストの品質が低いように感じたので、 青空文庫、日英Wikipedia,Project Gutenbergよりデータをシャッフルして作成しました。 Dataset Sources fujiki/wiki40b_ja globis-university/aozorabunko-clean manu/project_gutenberg blo05/cleaned_wiki_en_80-100 Uses llama-imatrix -m /path/to/model-file/original-f16.gguf -f imatrix_sample.txt texttext-generation1K<n<10K0 likes37 downloads1y agoHugging Face23el4 /bartowski-imatrix-v5-semantic-parquettext1K<n<10K0 likes32 downloads2mo agoHugging Face24Ji-Ha /deepseek-coder-7b-base-imatrix-wikitextHelp me, it is not working 😭 What is the syntax for using imatrix command?? It has to be .imatrix file? 1 likes31 downloads3y agoHugging Face25shing3232 /imatrix0 likes13 downloads3y agoHugging Face26liodon-ai /imatrix-calibration-corpustext10K<n<100K0 likes13 downloads4mo agoHugging Face27shing3232 /dataset_imatrixaudion<1K0 likes10 downloads3y agoHugging Face28YukiTomita-CC /imatrix-databricks-dolly-15k-jatext1K<n<10K1 likes9 downloads2y agoHugging Face29neody /imatrix_datasettextn<1K0 likes9 downloads2y agoHugging Face30yvfu /imatrix-calibration0 likes9 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.