Team Ai
Datasetpublic

lianghsun/tw-processed-law-article

Dataset Card for tw-processed-law-article tw-processed-law-article 是一個中華民國法規條文之結構化資料集,以條文為單位展開,包含 230,974 筆條文,涵蓋 11,462 部不同法規,橫跨憲法、法律與命令三個層級。每筆資料包含法規名稱、層級、條文內容、廢止註記與最後修正日期等欄位,適用於法律檢索系統、條文問答模型,或作為其他法律衍生資料集之結構化底層語料。 Dataset Details Dataset Description 本資料集整理自中華民國全國法規資料庫(law.moj.gov.tw)之公開法規條文。原始法規資料經處理後以單條條文為一筆資料(one row per article),每筆附帶法規名稱、層級分類、條文全文、廢止註記、最後修正日期與 API 更新日期等元資料。 層級分佈: 層級 筆數 說明 命令 180,757 由主管機關訂定之命令、規則、辦法等 法律 49,977… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-processed-law-article.

sourceHugging Facecc-by-4.0updated 6mo agoView on Hugging Face
3likes65downloads
Dataset Card

Dataset Card for tw-processed-law-article

<!-- Provide a quick summary of the dataset. --> tw-processed-law-article 是一個中華民國法規條文之結構化資料集,以條文為單位展開,包含 230,974 筆條文,涵蓋 11,462 部不同法規,橫跨憲法、法律與命令三個層級。每筆資料包含法規名稱、層級、條文內容、廢止註記與最後修正日期等欄位,適用於法律檢索系統、條文問答模型,或作為其他法律衍生資料集之結構化底層語料。

Dataset Details

Dataset Description

<!-- Provide a longer summary of what this dataset is. --> 本資料集整理自中華民國全國法規資料庫(law.moj.gov.tw)之公開法規條文。原始法規資料經處理後以單條條文為一筆資料(one row per article),每筆附帶法規名稱、層級分類、條文全文、廢止註記、最後修正日期與 API 更新日期等元資料。

層級分佈:

層級筆數說明
命令180,757由主管機關訂定之命令、規則、辦法等
法律49,977經立法院三讀通過之法律
憲法240憲法本文與增修條文

共約 60,841 筆條文附帶廢止註記(abandon_note 非空),反映法規廢止與修訂之歷史狀態。

  • —Curated by: Liang Hsun Huang
  • —Language(s) (NLP): Traditional Chinese
  • —License: CC BY 4.0

Dataset Sources

<!-- Provide the basic links for the dataset. -->

Uses

<!-- Address questions around how the dataset is intended to be used. -->

Direct Use

<!-- This section describes suitable use cases for the dataset. -->

本資料集主要設計用於:

  • —建構中華民國法規條文之檢索系統(RAG)之底層語料庫;
  • —作為法律問答、法律摘要、法律命名實體辨識等任務之結構化基礎資料;
  • —衍生其他法律相關之下游資料集(如問答對、條文解釋、法條引用檢查等);
  • —法律研究與法規變遷之量化分析。

Out-of-Scope Use

<!-- This section addresses misuse, malicious use, and uses that the dataset will not work well for. --> 本資料集不適用於下列用途:

  • —作為正式法律諮詢或裁判之依據,因條文狀態以資料擷取時點為準,後續修正之法規可能未反映。
  • —用於非中華民國法律體系之訓練或研究。
  • —作為司法判決或實務見解之資料來源,本資料集僅包含條文,不包含判例或函釋。

Dataset Structure

<!-- This section provides a description of the dataset fields, and additional information about the dataset structure such as criteria used to create the splits, relationships between data points, etc. -->

jsonl
{
  "text": "中華民國憲法 第 1 條 中華民國基於三民主義,為民有民治民享之民主共和國。",
  "name": "中華民國憲法",
  "level": "憲法",
  "abandon_note": "",
  "modified_date": "19470101",
  "api_updated_date": "2024/1/12 上午 12:00:00"
}
欄位說明
text條文全文(含法規名稱與條號前綴)
name法規名稱(如「中華民國憲法」)
level法規層級(憲法、法律、命令)
abandon_note廢止註記(已廢止條文之說明,正常條文為空字串)
modified_date法規最後修正日期(YYYYMMDD)
api_updated_date原始 API 之更新時間戳
統計項目數值
條文數230,974
法規數11,462
憲法條文240
法律條文49,977
命令條文180,757
含廢止註記60,841

Dataset Creation

Curation Rationale

<!-- Motivation for the creation of this dataset. -->

繁體中文語言模型在引用中華民國法規時經常產生幻覺(如虛構條號、混淆條文內容)。本資料集將全國法規資料庫之條文以結構化格式整理,作為下游法律任務之可靠底層語料,避免模型憑記憶生成錯誤條文。

Source Data

<!-- This section describes the source data (e.g. news text and headlines, social media posts, translated sentences, ...). -->

Data Collection and Processing

<!-- This section describes the data collection and processing process such as data selection criteria, filtering and normalization methods, tools and libraries used, etc. -->

原始資料取自法務部全國法規資料庫之公開 API 或開放資料集。資料經處理後以「條」為單位展開,每條條文為一筆 JSONL 記錄。已廢止之條文保留於資料集中並於 abandon_note 欄位加以標記,方便下游任務依需求過濾。

Who are the source data producers?

<!-- This section describes the people or systems who originally created the data. It should also include self-reported demographic or identity information for the source data creators if this information is available. -->

原始法規內容由中華民國立法院、主管機關與司法院等制定與公告,由法務部全國法規資料庫維護與發布。

Annotations

<!-- If the dataset contains annotations which are not part of the initial data collection, use this section to describe them. -->

Annotation process

本資料集未加入額外標註。層級(level)、名稱、修正日期等均來自原始法規資料庫之欄位。

Who are the annotators?

不適用。

Personal and Sensitive Information

<!-- State whether the dataset contains data that might be considered personal, sensitive, or private (e.g., data that reveals addresses, uniquely identifiable names or aliases, racial or ethnic origins, sexual orientations, religious beliefs, political opinions, financial or health data, etc.). If efforts were made to anonymize the data, describe the anonymization process. --> 本資料集為公開法規條文,不包含個人資訊或敏感資料。

Bias, Risks, and Limitations

<!-- This section is meant to convey both technical and sociotechnical limitations. -->

  • —法規條文以資料擷取時點之版本為準,後續之修正未反映於本資料集。
  • —abandon_note 欄位需搭配使用以過濾已廢止條文,否則可能導致模型引用失效法條。
  • —text 欄位同時包含法規名稱與條號前綴,使用前需視任務需求進行切分。
  • —本資料集僅包含條文本身,不包含立法理由、歷次修正說明或司法解釋。
  • —層級之定義以全國法規資料庫分類為準,與學理上之法源階層可能略有差異。

Recommendations

<!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->

建議使用者:

  • —使用前以 abandon_note 欄位過濾已廢止條文;
  • —涉及實務應用時,以全國法規資料庫之最新版本進行交叉驗證;
  • —搭配司法判決與函釋資料共同使用,以涵蓋法律實務之完整面貌。

Citation

<!-- If there is a paper or blog post introducing the dataset, the APA and Bibtex information for that should go in this section. -->

bibtex
@misc{tw-processed-law-article,
  title        = {tw-processed-law-article: Structured ROC Law Article Corpus},
  author       = {Liang Hsun Huang},
  year         = {2024},
  howpublished = {\url{https://huggingface.co/datasets/lianghsun/tw-processed-law-article}},
  note         = {230,974 articles from 11,462 ROC laws, structured at the article level.}
}

Dataset Card Authors

Liang Hsun Huang

Dataset Card Contact

Liang Hsun Huang