sphita/Intel-WebCorpus-forms
π» Intel WebCorpus Forms (Enterprise Hardware Q&A) This dataset is a massive, high-fidelity archive of 176,472 technical troubleshooting discussions (containing nearly 1 million individual messages) scraped from the official Intel Community Forums. It has been meticulously engineered for Large Language Model (LLM) training. Instead of a raw, messy dump of isolated posts, the data has been reconstructed into chronological conversation threads, noise-filtered, deduplicated, andβ¦ See the full description on the dataset page: https://huggingface.co/datasets/sphita/Intel-WebCorpus-forms.
π» Intel WebCorpus Forms (Enterprise Hardware Q&A)
This dataset is a massive, high-fidelity archive of 176,472 technical troubleshooting discussions (containing nearly 1 million individual messages) scraped from the official Intel Community Forums. It has been meticulously engineered for Large Language Model (LLM) training. Instead of a raw, messy dump of isolated posts, the data has been reconstructed into chronological conversation threads, noise-filtered, deduplicated, and packed into optimized Snappy-compressed Parquet files.
π Dataset Summary
π οΈ Data Engineering & Quality
This dataset was built using a custom Rust-based async crawler designed to bypass Enterprise WAFs and handle extreme server-side rate limits using exponential backoff. The raw data pipeline included:
- Thread Reconstruction: Over 900,000 flat messages were grouped by their
topic_idto reconstruct the exact back-and-forth conversational flow of human troubleshooting. - Noise Reduction: "Dead-weight" messages (under 5 words, e.g., "bump", "+1", "thanks") were automatically filtered out of the replies to increase the signal-to-noise ratio.
- Deduplication: Overlapping chunks were strictly deduplicated by
msg_idto ensure absolute uniqueness. - HTML Sanitization: Raw HTML bodies were stripped into clean text.
ποΈ Top Hardware Categories (Boards)
The discussions cover deep technical support for Intel's entire hardware and software ecosystem. The top 10 categories are: | Board / Category | Discussions | | :--- | :--- | | fortran-compiler | 28,479 | | graphics (Intel Arc & Iris) | 23,721 | | processors (Core & Xeon) | 17,436 | | software-archive | 16,556 | | wireless | 9,006 | | oneapi-math-kernel-library | 7,198 | | integrated-performance-primitives | 6,675 | | distribution-openvino-toolkit | 6,621 | | ethernet-products | 5,942 | | server-products | 5,311 |
π Schema
The dataset is provided in parquet format. Each row represents a complete discussion thread. | Column | Type | Description | | :--- | :--- | :--- | | discussion_id | string | The unique ID of the forum topic/thread. | | subject | string | The title or subject of the discussion. | | board | string | The specific Intel hardware/software category. | | post_time | string | ISO 8601 timestamp of the original question. | | has_solution | boolean | True if an Intel engineer or community member provided an accepted solution. | | message_count | int64 | The total number of valid messages in the thread. | | messages_json | string | A JSON-encoded string containing the full list of messages in chronological order. |
π Inside messages_json
If you parse the messages_json column, you get a list of dictionary objects representing the conversation:
[
{
"msg_id": "12345",
"depth": 0,
"author": "user_123",
"post_time": "2023-01-01T12:00:00Z",
"is_solution": false,
"body": "I am getting Error 504 on my Xeon server...",
"kudos": 2
},
{
"msg_id": "12346",
"depth": 1,
"author": "Intel_Support",
"post_time": "2023-01-01T12:15:00Z",
"is_solution": true,
"body": "Please update your BIOS to version X...",
"kudos": 5
}
]π Use Cases
- Hardware Troubleshooting LLMs: Train models that understand complex server, networking, and processor error codes.
- Reasoning (Chain-of-Thought): Train models on multi-turn troubleshooting processes, where users try solutions, fail, report back, and eventually succeed.
- RAG (Retrieval-Augmented Generation): Build enterprise IT chatbots.
π Updates
This dataset is maintained via an automated Delta Sync Pipeline. New discussions are fetched weekly, sorted, and pushed as new chunk files without overwriting the existing archive.
