sabin1234/Dengue_Surveillance_Data_Question_Answering_Dataset
Nepali ShareGPT Clean Final Dataset Comprehensive Documentation & Analysis Report Dengue Surveillance Data - Question Answering Dataset 📋 Dataset Overview This dataset is a curated collection of 256 question-answer pairs focused on Dengue Surveillance in Nepal. It contains data-grounded questions in Nepali language paired with factual, statistical answers sourced from the Department of Health Services (DoHS), Nepal. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Dengue_Surveillance_Data_Question_Answering_Dataset.
Nepali ShareGPT Clean Final Dataset
Comprehensive Documentation & Analysis Report
Dengue Surveillance Data - Question Answering Dataset
📋 Dataset Overview
This dataset is a curated collection of 256 question-answer pairs focused on Dengue Surveillance in Nepal. It contains data-grounded questions in Nepali language paired with factual, statistical answers sourced from the Department of Health Services (DoHS), Nepal. The dataset represents information about dengue positive cases across different districts of Nepal across three consecutive fiscal years.
Key Characteristics:
- Total Records: 256 question-answer pairs
- Language: Nepali (देवनागरी Script)
- Primary Topic: Dengue Positive Cases Surveillance
- Geographic Scope: Multiple districts of Nepal
- Temporal Scope: Fiscal years 2071/72, 2072/73, 2073/74 (Nepali calendar)
- Source: Department of Health Services (DoHS), Nepal
- Data Format: JSONL (JSON Lines)
- License: CC-BY-4.0 (Creative Commons Attribution 4.0)
- Dataset Version: Final Clean Version (v1)
- Task Type: Instruction-following / Question Answering
- Data Type: Synthetic (Artificially generated questions on real surveillance data)
- Dataset Name: nepalisharegptclean_final.jsonl
📊 What is This Dataset About?
This dataset contains data-grounded questions and answers about dengue surveillance in Nepal. Each question is designed to query specific epidemiological data points, and each answer provides factual numerical information extracted from the Department of Health Services surveillance records.
Topic Coverage:
1. Direct Value Queries (28.9%)
- Specific dengue case counts for a given district and fiscal year
- Format: "How many dengue cases in [District] in [Year]?"
- Example: "आर्थिक वर्ष २०७१/७२ मा बारामा डेङ्गुका कति वटा पुष्टि भएका बिरामीहरू भेटिएका थिए?"
- Response type: Single numerical value
2. District Comparisons (20.3%)
- Comparing dengue cases across different districts
- Format: "Which district had more/fewer cases?"
- Identifying higher/lower burden districts
- Multi-district analysis questions
3. Year Comparisons (16.0%)
- Comparing cases across different fiscal years
- Format: "How did cases change from [Year1] to [Year2]?"
- Year-over-year variation analysis
- Temporal trend identification
4. Three-Year Trends (9.8%)
- Analysis of cases over all three fiscal years
- Format: "What was the trend in [District] over three years?"
- Long-term pattern identification
- Epidemiological trajectory analysis
5. Highest Case Year (9.8%)
- Identifying the year with maximum cases
- Format: "In which year did [District] have the most cases?"
- Peak case year identification
- Maximum burden period identification
6. Lowest Case Year (9.8%)
- Identifying the year with minimum cases
- Format: "In which year were cases lowest in [District]?"
- Minimum burden identification
- Best control period identification
7. Ranking Questions (3.2%)
- Ranking districts by case burden
- Highest, second-highest, lowest districts
- Comparative severity assessment
- District prioritization
8. Zero Cases (1.2%)
- Districts with no reported dengue cases
- Format: "Which districts had zero cases?"
- Case-free area identification
9. National Totals (1.2%)
- Aggregated national dengue data
- Format: "What was the total dengue cases across Nepal?"
- National surveillance summary
Subject Area Classification:
- Primary Domain: Public Health (स्वास्थ्य)
- Category: Dengue Surveillance (डेङ्गु सर्वेक्षण)
- Subdomain: Multiple analytical perspectives on case data
- Geographic Context: Nepal (नेपाल)
- Temporal Context: Fiscal years 2071/72 to 2073/74 (Nepali calendar)
🔄 Behavior Distribution Analysis
Overall Behavior Pattern:
Behavior Definition:
"Short Factual Answer" - Providing concise, data-grounded answers that preserve key numerical facts while maintaining accuracy and clarity.
Detailed Analysis:
✅ Perfect Consistency (100%) - All entries follow uniform factual answer format ✅ Data-Grounded Responses - All answers backed by official DoHS statistics ✅ Optimal Brevity - Average response of only 90 characters ✅ High Precision - Exact numerical answers to quantitative questions ✅ Standardized Format - Consistent response structure across all entries ✅ Language Standardization - Uniform Nepali medical/surveillance terminology
Key Characteristics of Short Factual Answer Behavior:
- Response Length Optimization
- Average: 90 characters per answer
- Range: 54-197 characters
- Typically 1-2 sentences
- Direct, concise communication
- No unnecessary elaboration
- Content Focus
- Pure factual information
- Numerical data from surveillance
- District-specific statistics
- Year-specific values
- No interpretations or recommendations
- Language Style
- Simple, direct language
- Standard Nepali medical terminology
- Numerical expressions
- Clear data presentation
- No ambiguous phrasing
- Information Structure
- Opening statement with context
- Numerical data point
- Optional unit specification
- Closing statement confirming data
Response Format Examples:
Format 1 (Basic):
"आर्थिक वर्ष २०७१/७२ मा बारामा डेङ्गुका ० वटा पुष्टि भएका बिरामीहरू भेटिएका थिए।"
(In fiscal year 2071/72, 0 dengue positive patients were found in Bara.)
Format 2 (With Context):
"आर्थिक वर्ष २०७२/७३ मा बारामा २ वटा डेङ्गु पुष्टि भएका बिरामीहरू दर्ता भएका थिए।"
(In fiscal year 2072/73, 2 dengue positive patients were registered in Bara.)
Format 3 (Official Statement):
"आर्थिक वर्ष २०७३/७४ मा भक्तपुरमा ० वटा डेङ्गु पुष्टि भएका बिरामीहरू दर्ता गरिएको थियो।"
(In fiscal year 2073/74, 0 dengue positive patients were registered in Bhaktapur.)Implications:
- Data Accuracy: Short format minimizes errors and ambiguity
- Machine Processing: Easy for AI/ML systems to parse and extract
- Mobile-Friendly: Optimal for mobile health information systems
- Surveillance Efficiency: Quick reference format for public health officials
- Epidemiological Clarity: Clear data transmission for health decision-making
❓ Question Type Distribution Analysis
Question Classification:
Question Type Definition:
"Complex Grounded Open-Ended Question" - Questions that require retrieval of specific data points from structured surveillance records, with variation in formulation while maintaining a consistent answer pattern.
Detailed Question Type Analysis:
All 256 questions are classified as complex grounded open-ended questions, meaning:
- "Complex" - Questions require understanding of epidemiological concepts and district/year specifications
- "Grounded" - Questions are grounded in real surveillance data from Department of Health Services
- "Open-Ended" - Questions are phrased in multiple ways to access the same underlying data
Question Structure Patterns:
Pattern 1: Simple Inquiry (Most Common)
Template: "[Year] मा [District]मा डेङ्गुका कति [patients]?"
Translation: "How many dengue [patients] in [District] in [Year]?"
Example: "आर्थिक वर्ष २०७१/७२ मा बारामा डेङ्गुका कति वटा पुष्टि भएका बिरामीहरू भेटिएका थिए?"
Frequency: ~35% of questions
Purpose: Direct data retrievalPattern 2: Registration-Focused Inquiry
Template: "[Year] मा [District]मा कति जना/वटा डेङ्गु दर्ता भएका?"
Translation: "How many dengue cases were registered in [District] in [Year]?"
Example: "आर्थिक वर्ष २०७२/७३ मा बारामा कति वटा डेङ्गु पुष्टि भएका बिरामीहरू दर्ता भएका थिए?"
Frequency: ~30% of questions
Purpose: Official registration data retrievalPattern 3: National Statistics Reference
Template: "राष्ट्रिय तथ्याङ्कअनुसार, [Year] मा [District]मा कति [data]?"
Translation: "According to national statistics, how many [data] in [District] in [Year]?"
Example: "राष्ट्रिय तथ्याङ्कअनुसार, २०७३/७४ मा बारामा कति वटा डेङ्गु पुष्टि भएका बिरामीहरू भेटिएका थिए?"
Frequency: ~15% of questions
Purpose: Emphasize official/verified dataPattern 4: Comparative Inquiry
Template: "कुन [Year] मा [District]मा डेङ्गु बिरामीहरूको संख्या बढी थियो?"
Translation: "In which [Year] was the number of dengue patients higher in [District]?"
Frequency: ~10% of questions
Purpose: Year-over-year comparisonPattern 5: Comparative District
Template: "[Year] मा कुन जिल्ला/जिल्लाहरूमा कति डेङ्गु बिरामी भेटिएका?"
Translation: "In [Year], how many dengue cases in which districts?"
Frequency: ~7% of questions
Purpose: Multi-district comparisonPattern 6: Trend Inquiry
Template: "[District] मा [Year1] देखि [Year3] सम्म डेङ्गुको प्रवृत्ति कस्तो रहेको?"
Translation: "What was the trend of dengue in [District] from [Year1] to [Year3]?"
Frequency: ~3% of questions
Purpose: Three-year trend analysisQuestion Characteristics:
Geographic Coverage:
The questions cover dengue cases across Nepal's districts. Major districts mentioned include:
- High Burden Districts: Chitwan, Kathmandu, Lalitpur, Bhaktapur
- Medium Burden Districts: Bara, Bhojpur, Dhanusha, Parsa
- Low Burden Districts: Various other districts with 0-5 cases
Temporal Coverage:
Fiscal Year Distribution:
2071/72: 99 questions (38.7%)
2072/73: 57 questions (22.3%)
2073/74: 50 questions (19.5%)
Comparative Questions: 50 questions (19.5%)📈 Comprehensive Data Analysis
File Structure & Size:
Response Length Distribution:
Distribution Breakdown:
- Very Short (50-70 chars): 5% - Minimal context answers
- Short (71-100 chars): 75% - Most common, standard answers
- Medium (101-150 chars): 18% - Extended context answers
- Long (151-197 chars): 2% - Maximum detail answers
Question Length Distribution:
Sub-Domain Distribution:
Metadata Quality Indicators:
Constraint Specifications:
├── Max Response Sentences: 3
├── Min Reasoning Dimensions: 2
├── Min Complexity Score: 5 (on scale 1-10)
├── Question Length: 70-260 characters
├── Response Length: 15-160 characters
└── Content Language: English (metadata) / Nepali (content)🏗️ JSONL File Structure
Record Structure:
Each record contains comprehensive metadata and conversation data:
{
"id": "sg_abcb15f0dc85f65b168c156590f6b290",
"conversations": [
{
"from": "human",
"value": "आर्थिक वर्ष २०७१/७२ मा बारामा डेङ्गुका कति वटा पुष्टि भएका बिरामीहरू भेटिएका थिए?"
},
{
"from": "gpt",
"value": "आर्थिक वर्ष २०७१/७२ मा बारामा डेङ्गुका ० वटा पुष्टि भएका बिरामीहरू भेटिएका थिए।"
}
],
"source_provenance": {
"id": "sg_abcb15f0dc85f65b168c156590f6b290",
"source": "DepartmentOfHealthServices/Dengue_Positive_Cases:default:train",
"source_name": "dengue_positive_cases_fy2071_74",
"source_repo": "DepartmentOfHealthServices/Dengue_Positive_Cases",
"source_config": "default",
"source_split": "train",
"source_revision": "dohs_dengue_2071_74",
"language": "ne",
"language_code": "npi",
"script": "Deva",
"license": "CC-BY-4.0",
"license_tier": "permissive",
"task_type": "instruction-following",
"generation_type": "synthetic",
"condition": "synthetic",
"metadata_json": {
"generation_domain": "Public Health",
"generation_category": "Dengue Surveillance",
"generation_sub_domain": "Direct Value",
"behavior": "short factual answer",
"behavior_definition": "Provide a concise, data-grounded answer while preserving the key fact.",
"question_type": "complex grounded open-ended question",
"question_length": "70 to 260 characters",
"response_length": "15 to 160 characters",
"maximum_response_sentences": 3,
"minimum_reasoning_dimensions": 2,
"minimum_complexity_score": 5,
"content_language": "Nepali",
"content_script": "Devanagari",
"english_content_allowed": false
}
},
"input_format": "sharegpt",
"source_task_id": "sg_abcb15f0dc85f65b168c156590f6b290",
"language": null,
"language_code": null,
"script": null
}Field Definitions:
🎯 Question Pattern Analysis
Major Question Categories:
Category 1: Direct Value Queries (28.9%)
Purpose: Retrieve specific case counts for known district-year combinations
Q: "आर्थिक वर्ष २०७१/७२ मा चितवनमा कति जना डेङ्गु पुष्टि भएको बिरामीहरू भेटिएका थिए?"
A: "आर्थिक वर्ष २०७१/७२ मा, चितवनमा ११९ जना डेङ्गु पुष्टि भएको बिरामीहरू भेटिएका थिए।"- Straightforward data lookup
- Single numerical response
- Time-specific queries
Category 2: District Comparisons (20.3%)
Purpose: Identify and compare dengue burden across different districts
Q: "कुन जिल्ला/जिल्लाहरूमा डेङ्गु बिरामीहरूको संख्या बढी थियो?"
A: Comparative response listing districts with higher/lower burden- Multi-district analysis
- Comparative epidemiology
- Burden identification
Category 3: Year Comparisons (16.0%)
Purpose: Analyze temporal changes in dengue cases
Q: "कुन वर्ष मा [District] मा डेङ्गु बिरामीहरूको संख्या बढी थियो?"
A: Identification of year with higher case burden- Year-over-year variation
- Epidemic year identification
- Temporal trend analysis
Category 4: Three-Year Trends (9.8%)
Purpose: Understand long-term epidemiological patterns
Q: "[District] मा २०७१/७२ देखि २०७३/७४ सम्म डेङ्गुको प्रवृत्ति कस्तो रहेको?"
A: Trend description (increasing/decreasing/stable)- Multi-year pattern analysis
- Epidemiological trajectory
- Long-term burden assessment
Category 5: Peak & Trough Identification (19.6%)
Purpose: Identify years with highest/lowest case burden
Highest Year Query:
Q: "कुन वर्ष मा [District] मा डेङ्गु बिरामीहरूको संख्या सबैभन्दा बढी थियो?"
A: Specific year with maximum cases
Lowest Year Query:
Q: "कुन वर्ष मा [District] मा डेङ्गु बिरामीहरूको संख्या सबैभन्दा कम थियो?"
A: Specific year with minimum casesCategory 6: District Ranking (3.2%)
Purpose: Rank districts by disease burden
Q: "[Year] मा कुन जिल्ला/जिल्लाहरूमा डेङ्गु पुष्टि भएका बिरामीहरू सबैभन्दा बढी थिए?"
A: Identification of highest/second-highest/lowest burden districtsCategory 7: Special Cases (2.4%)
Purpose: Identify districts with zero cases or national totals
Zero Cases:
Q: "कुन जिल्लामा डेङ्गु पुष्टि भएका बिरामीहरू शून्य थिए?"
A: List of case-free districts
National Total:
Q: "नेपालमा कुल डेङ्गु पुष्टि भएका बिरामीहरू कति थिए?"
A: Aggregate national case countQuestion Variability:
Despite being grounded in the same data, questions show significant variation:
Variation Patterns:
├── Temporal Framing
│ ├── "आर्थिक वर्ष २०७१/७२ मा"
│ ├── "२०७१/७२ मा"
│ └── "FY २०७१/७२ अनुसार"
│
├── District Introduction
│ ├── "[District]मा"
│ ├── "[District]को"
│ └── "[District] जिल्लामा"
│
├── Case Count Terminology
│ ├── "कति जना"
│ ├── "कति वटा"
│ ├── "कति"
│ └── "संख्या कति"
│
├── Data Verification
│ ├── "भेटिएका"
│ ├── "दर्ता भएका"
│ ├── "राष्ट्रिय तथ्याङ्कअनुसार"
│ └── "बताउनुहोस्"
│
└── Patient Terminology
├── "डेङ्गु पुष्टि भएका बिरामीहरू"
├── "डेङ्गु बिरामीहरू"
└── "डेङ्गुको सङ्क्रमण भेटिएका"Linguistic Patterns:
Question Opening Phrases:
- "आर्थिक वर्ष [Year] मा" (In fiscal year [Year])
- "राष्ट्रिय तथ्याङ्कअनुसार" (According to national statistics)
- "कुन [Year] मा" (In which [Year])
- "[Year] मा कुन जिल्ला" (In [Year], which district)
Question Closing Patterns:
- "भेटिएका थिए?" (were found?)
- "दर्ता भएका थिए?" (were registered?)
- "कति थियो?" (what was the count?)
- "बताउनुहोस्" (please tell)
Answer Structure:
- "[Year] मा [District]मा [Number] जना/वटा डेङ्गु [verb] भएका [noun]।"
- Consistent temporal marker placement
- District specification following year
- Numerical data in specific position
- Standardized closing
📝 Sample Questions & Answers
Sample 1: Direct Value Query - Zero Cases
Q (Nepali): आर्थिक वर्ष २०७१/७२ मा बारामा डेङ्गुका कति वटा पुष्टि भएका बिरामीहरू भेटिएका थिए?
Q (English): In fiscal year 2071/72, how many dengue-positive patients were found in Bara?
A (Nepali): आर्थिक वर्ष २०७१/७२ मा बारामा डेङ्गुका ० वटा पुष्टि भएका बिरामीहरू भेटिएका थिए।
A (English): In fiscal year 2071/72, 0 dengue-positive patients were found in Bara.
Sample 2: Direct Value Query - Multiple Cases
Q (Nepali): आर्थिक वर्ष २०७१/७२ मा चितवनमा कति जना डेङ्गु पुष्टि भएको बिरामीहरू भेटिएका थिए?
Q (English): In fiscal year 2071/72, how many dengue-positive patients were found in Chitwan?
A (Nepali): आर्थिक वर्ष २०७१/७२ मा, चितवनमा ११९ जना डेङ्गु पुष्टि भएको बिरामीहरू भेटिएका थिए।
A (English): In fiscal year 2071/72, 119 dengue-positive patients were found in Chitwan.
Sample 3: Year-Over-Year Comparison
Q (Nepali): कुन वर्ष मा चितवनमा डेङ्गु बिरामीहरूको संख्या बढी थियो?
Q (English): In which year were dengue cases higher in Chitwan?
A: [Comparative response identifying year with higher burden]
Sample 4: Three-Year Trend Query
Q (Nepali): चितवन मा २०७१/७२ देखि २०७३/७४ सम्म डेङ्गुको प्रवृत्ति कस्तो रहेको?
Q (English): What was the trend of dengue in Chitwan from 2071/72 to 2073/74?
A: [Trend description indicating increases/decreases across three years]
🔑 Key Medical & Epidemiological Terminology
Disease-Related Terms:
- डेङ्गु (Dengue): Mosquito-borne viral disease
- पुष्टि भएको (Confirmed): Laboratory-confirmed cases
- बिरामी (Patient): Diseased/sick person
- सङ्क्रमण (Infection): Disease transmission
Data Collection Terms:
- दर्ता गरिएको (Registered): Cases officially recorded in surveillance
- भेटिएका (Found): Cases identified/discovered
- तथ्याङ्क (Statistics): Numerical data and information
- राष्ट्रिय (National): Country-wide scope
Geographic Terms:
- जिल्ला (District): Administrative division
- मा (In): Locational preposition
- नेपाल (Nepal): Country name
- काठमाडौं (Kathmandu): Capital city
Temporal Terms:
- आर्थिक वर्ष (Fiscal Year): Financial/administrative year
- २०७१/७२ (2071/72): Nepali calendar year notation
- देखि (From): Starting point
- सम्म (Until): Ending point
Comparative Terms:
- बढी (More/Higher): Greater amount
- कम (Less/Lower): Smaller amount
- सबैभन्दा (Most): Superlative
- बीच (Between): Comparison context
Numerical Terms:
- वटा (Units - for counting objects): Quantifier
- जना (Units - for counting people): Quantifier
- शून्य/० (Zero): No cases
- कति (How many): Question word
📊 Data Quality Assessment
Strengths:
✅ 100% Behavior Consistency - All entries follow uniform short-answer format ✅ Data-Grounded Responses - All answers from official DoHS surveillance records ✅ Official Source - Department of Health Services (Nepal) provides authoritative data ✅ Question Diversity - Despite consistent format, 11 distinct query types ✅ Clear Metadata - Comprehensive provenance and quality specifications ✅ Permissive License - CC-BY-4.0 for unrestricted use and attribution ✅ Standardized Format - Consistent JSONL structure for easy processing ✅ Temporal Specificity - Clear fiscal year specification in all queries ✅ Geographic Specificity - District-level granularity for targeted analysis ✅ Language Quality - Native Nepali with standard medical terminology
Limitations:
⚠️ Single Disease Focus - Only dengue surveillance covered ⚠️ Limited Time Period - Only three fiscal years (2071/72 to 2073/74) ⚠️ Synthetic Questions - Questions artificially generated, not from actual users ⚠️ District Subset - Not all Nepal districts uniformly represented ⚠️ Aggregate Data Only - No individual patient-level information ⚠️ No Outcome Data - Only case count data, not severity/mortality ⚠️ No Preventive Data - No information on vaccination or prevention ⚠️ Reporting Bias - Subject to surveillance system limitations/gaps ⚠️ Static Dataset - No update mechanism for new surveillance data
💡 Use Cases & Applications
1. Public Health Information Systems
Purpose: Query dengue surveillance data for health planning
Benefits: Data-grounded Q&A pairs for surveillance databases
Use Case: Health ministry dashboards and reporting systems
Example: Quick retrieval of district-specific dengue case counts2. AI/ML Training for Nepali NLP
Purpose: Training language models for Nepali-language Q&A
Benefits: Structured, consistent data for instruction-following tasks
Use Case: Fine-tuning Nepali medical/epidemiological AI systems
Example: Training retrieval augmented generation (RAG) systems3. Chatbot Development
Purpose: Building health information chatbots in Nepali
Benefits: Ready-formatted Q&A pairs for chatbot training
Use Case: Public health query chatbots for Nepali speakers
Example: "WhatsApp health bot" for dengue surveillance queries4. Health Education Materials
Purpose: Creating educational resources about dengue in Nepal
Benefits: Evidence-based, data-grounded information
Use Case: School curricula, community health worker training
Example: Dengue statistics and trends for Nepali educational contexts5. Epidemiological Research
Purpose: Studying dengue patterns in Nepal
Benefits: Structured epidemiological data in accessible format
Use Case: Academic research on disease temporal trends
Example: Analyzing three-year dengue trends by district6. Surveillance System Enhancement
Purpose: Improving public health surveillance systems
Benefits: Data-grounded Q&A for surveillance reporting
Use Case: Automated surveillance data query systems
Example: Real-time dengue case retrieval by district and year7. Decision Support Systems
Purpose: Supporting public health decision-making
Benefits: Quick access to surveillance data for policy decisions
Use Case: Health program planning and resource allocation
Example: Identifying high-burden districts for intervention prioritization8. Accessibility & Health Literacy
Purpose: Making surveillance data accessible to Nepali speakers
Benefits: Data presented in natural Nepali language
Use Case: Community health awareness campaigns
Example: Public understanding of local dengue situation🔍 Behavior-Question Pattern Relationship
Consistency Analysis:
All 256 Records:
├── Behavior: Short Factual Answer (100%)
├── Question Type: Complex Grounded Open-Ended (100%)
├── Domain: Public Health (100%)
├── Category: Dengue Surveillance (100%)
└── Sub-Domains:
├── 28.9% Direct Value Queries
├── 20.3% District Comparisons
├── 16.0% Year Comparisons
├── 9.8% Three-Year Trends
├── 9.8% Highest Year Identification
├── 9.8% Lowest Year Identification
└── 5.4% Other (Ranking, Zero Cases, National)Key Findings:
- Unified Response Model - Consistent short answer (90±17 chars) across all queries
- Data-Grounded Uniformity - All answers derived from same source (DoHS)
- Question Diversity - 11 distinct sub-domains despite uniform answer format
- Temporal Consistency - All questions reference same three fiscal years
- Geographic Specificity - Queries target specific districts throughout Nepal
- Analytical Completeness - Covers direct retrieval, comparison, and trend analysis
📚 Metadata Details
Language Information:
Data Source Information:
Dataset Properties:
Quality Constraints:
🔄 How to Access & Use the Dataset
Reading JSONL File in Python:
import json
# Read the entire dataset
with open('nepali_sharegpt_clean_final.jsonl', 'r', encoding='utf-8') as f:
for line_num, line in enumerate(f, 1):
data = json.loads(line)
# Extract key fields
record_id = data['id']
question = data['conversations'][0]['value']
answer = data['conversations'][1]['value']
# Extract metadata
provenance = data['source_provenance']
metadata = json.loads(provenance['metadata_json'])
print(f"Record {line_num}: {record_id}")
print(f"Sub-Domain: {metadata['generation_sub_domain']}")
print(f"Q: {question}")
print(f"A: {answer}\n")Filtering by Sub-Domain:
import json
# Filter for direct value queries only
direct_value_records = []
with open('nepali_sharegpt_clean_final.jsonl', 'r', encoding='utf-8') as f:
for line in f:
data = json.loads(line)
provenance = data['source_provenance']
metadata = json.loads(provenance['metadata_json'])
if metadata['generation_sub_domain'] == 'Direct Value':
direct_value_records.append(data)
print(f"Found {len(direct_value_records)} direct value queries")Extracting by District:
import json
import re
# Extract all questions about Chitwan
chitwan_records = []
with open('nepali_sharegpt_clean_final.jsonl', 'r', encoding='utf-8') as f:
for line in f:
data = json.loads(line)
question = data['conversations'][0]['value']
if 'चितवन' in question:
chitwan_records.append(data)
print(f"Found {len(chitwan_records)} questions about Chitwan")Statistical Analysis:
import json
import statistics
# Analyze response lengths
response_lengths = []
with open('nepali_sharegpt_clean_final.jsonl', 'r', encoding='utf-8') as f:
for line in f:
data = json.loads(line)
response = data['conversations'][1]['value']
response_lengths.append(len(response))
print(f"Mean: {statistics.mean(response_lengths):.1f} chars")
print(f"Median: {statistics.median(response_lengths):.1f} chars")
print(f"Stdev: {statistics.stdev(response_lengths):.1f} chars")Command Line Access:
# View first 2 records formatted
head -2 nepali_sharegpt_clean_final.jsonl | python3 -m json.tool
# Count total records
wc -l nepali_sharegpt_clean_final.jsonl
# Search for specific district
grep -i "चितवन" nepali_sharegpt_clean_final.jsonl | wc -l
# Extract only questions
jq -r '.conversations[0].value' nepali_sharegpt_clean_final.jsonl | head -10
# Filter by sub-domain
jq -r 'select(.source_provenance.metadata_json | fromjson | .generation_sub_domain == "Direct Value")' nepali_sharegpt_clean_final.jsonl⚖️ License & Terms of Use
License: CC-BY-4.0 (Creative Commons Attribution 4.0 International)
License Type: Permissive (Allows commercial and non-commercial use with attribution)
Permissions:
✅ Commercial Use - Use for commercial products/services ✅ Modification - Modify and adapt the dataset ✅ Distribution - Redistribute to others ✅ Private Use - Use for personal/internal purposes ✅ Sublicense - Create derivative works with same license ✅ Patent Use - No patent restrictions
Requirements:
📋 Attribution - Must credit creators (Department of Health Services, Nepal) 📋 License Notice - Include copy of CC-BY-4.0 license 📋 Changes Indication - Document any significant modifications
Limitations:
❌ No Warranty - Provided "as-is" without warranties ❌ No Liability - Creators not liable for usage outcomes ❌ Trademark Rights - Does not grant trademark usage rights
Reference:
Full license text: https://creativecommons.org/licenses/by/4.0/
🌍 Dataset Significance
For Public Health in Nepal:
- Surveillance Enhancement: Structured Q&A format for surveillance system queries
- Data Accessibility: Makes official DoHS data easily queryable
- Evidence-Based: Grounded in real epidemiological surveillance data
- Temporal Analysis: Enables trend analysis across fiscal years
For Machine Learning & NLP:
- Nepali Language Resource: Specialized medical/public health Nepali Q&A
- Domain Specification: Epidemiological domain for focused NLP training
- Structured Data: Well-organized JSONL format for model training
- Data Grounding: Real surveillance data provides grounded training signal
For Healthcare Technology:
- Chatbot Development: Ready-formatted Q&A for health chatbots
- Information Systems: Data for health information retrieval systems
- Decision Support: Basis for surveillance-based decision support tools
For Research & Education:
- Epidemiological Research: Real dengue surveillance data for academic study
- Health Education: Evidence-based materials for Nepal-focused health education
- Language Research: Nepali medical terminology and expression patterns
📊 Content Distribution Summary
By Sub-Domain (Analytical Type):
Analytical Function Distribution:
├── Direct Data Retrieval (28.9%)
│ └── Single district-year case counts
├── Comparative Analysis (36.3%)
│ ├── District Comparisons (20.3%)
│ ├── Year Comparisons (16.0%)
├── Trend Analysis (9.8%)
│ └── Three-year patterns
├── Extrema Identification (19.6%)
│ ├── Highest Case Years (9.8%)
│ └── Lowest Case Years (9.8%)
└── Ranking & Special (5.4%)
├── District Rankings (3.2%)
├── Zero Cases (1.2%)
└── National Totals (1.2%)By Temporal Focus:
Fiscal Year Representation:
├── FY 2071/72: 99 questions (38.7%)
├── FY 2072/73: 57 questions (22.3%)
├── FY 2073/74: 50 questions (19.5%)
└── Multi-Year/Comparative: 50 questions (19.5%)By Geographic Coverage:
Thematic Distribution:
├── District-Specific Queries: ~220 (85.9%)
├── Multi-District Queries: ~30 (11.7%)
└── National-Level Queries: ~6 (2.4%)🎓 Data Complexity Levels
By Question Complexity:
Level 1 - Direct Lookup (28.9%)
- Single data point retrieval
- Simple district + year queries
- Straightforward answers
- Example: "FY 2071/72 मा Bara मा कति dengue cases?"
- Response: Simple number
Level 2 - Comparative Analysis (36.3%)
- Multi-district or multi-year comparisons
- Requires identifying higher/lower values
- Slightly more complex reasoning
- Example: "कुन जिल्ला/जिल्लाहरूमा dengue बिरामीहरू बढी?"
- Response: District names or case counts
Level 3 - Trend Analysis (29.8%)
- Temporal pattern identification
- Three-year trend assessment
- Ranking and extrema identification
- Example: "District X मा FY 2071-2073 मा dengue को प्रवृत्ति कस्तो?"
- Response: Trend description (increasing/decreasing/stable)
📞 Dataset Information & Contact
Data Source:
- Institution: Department of Health Services (DoHS), Nepal
- Location: Kathmandu, Nepal
- Specialization: Public health surveillance (dengue focus)
- Country: Nepal
- Temporal Coverage: FY 2071/72 to 2073/74
Dataset Properties:
- Version: Final Clean (v1)
- Release Date: 2026
- Last Updated: 2026
- Maintenance: Static dataset (no regular updates)
Data Origin:
- Source Type: Official public health surveillance
- Question Generation: Synthetic (artificially created)
- Answer Source: Real (from official DoHS records)
- Quality Assurance: Curated "clean final" version
⚠️ Important Disclaimers
Data Accuracy Disclaimer:
⚠️ DISCLAIMER: This dataset reflects official Department of Health Services
surveillance data as recorded in their information systems.
Surveillance data may be subject to:
- Reporting delays
- Incomplete case registration
- Testing capacity limitations
- Geographical variation in detection
- Changes in case definitions over time
Users should consult official DoHS publications for authoritative data
interpretation and epidemiological analysis.Data Usage Precautions:
- Surveillance Limitations
- Data reflects surveillance system capacity
- May not represent true disease burden
- Subject to reporting gaps and delays
- Geographic variation in surveillance intensity
- Appropriate Applications
- ✅ Recommended: Education, research, surveillance system queries
- ⚠️ Caution: Policy decisions (verify with latest DoHS data)
- ❌ Not recommended: Standalone medical advice
- Data Currency
- Dataset covers FY 2071/72 to 2073/74 only
- Does not include current/recent data
- Should be updated with latest surveillance reports
- Regulatory Compliance
- Ensure compliance with Nepali health regulations
- Respect data privacy if personal information present
- Attribute to Department of Health Services
📋 Quick Reference Guide
Dataset Summary Card:
╔═══════════════════════════════════════════════════════════╗
║ NEPALI SHAREGPT CLEAN FINAL - DATASET SUMMARY ║
║ (Dengue Surveillance - Nepal) ║
╠═══════════════════════════════════════════════════════════╣
║ ║
║ Total Records: 256 ║
║ Language: Nepali ║
║ Script: Devanagari ║
║ Primary Topic: Dengue Surveillance ║
║ Format: JSONL ║
║ License: CC-BY-4.0 ║
║ Data Source: DoHS (Nepal) ║
║ ║
║ Behavior Type: Short Factual Answer (100%) ║
║ Question Type: Complex Grounded Open-Ended ║
║ Domain: Public Health (100%) ║
║ Category: Dengue Surveillance (100%) ║
║ ║
║ Avg Response: 90 characters ║
║ Min Response: 54 characters ║
║ Max Response: 197 characters ║
║ ║
║ Avg Question: 94 characters ║
║ Question Variety: 11+ sub-domain types ║
║ ║
║ Time Period: FY 2071/72 - 2073/74 ║
║ Geographic Scope: Multiple Nepal Districts ║
║ ║
║ Data Quality: High ║
║ Consistency: 100% ║
║ Completeness: 100% ║
║ ║
║ Use Cases: Public Health Q&A, NLP Training, ║
║ Chatbots, Surveillance Systems ║
║ ║
║ Sub-Domain Distribution: ║
║ - Direct Value: 28.9% ║
║ - Comparison: 36.3% ║
║ - Trend: 9.8% ║
║ - Peak/Trough: 19.6% ║
║ - Other: 5.4% ║
║ ║
╚═══════════════════════════════════════════════════════════╝📚 Comparison with Previous Datasets
Comparison Table:
🔄 Processing & Implementation Examples
Building a Dengue Q&A System:
import json
from elasticsearch import Elasticsearch
# Index dataset in Elasticsearch for fast retrieval
es = Elasticsearch()
with open('nepali_sharegpt_clean_final.jsonl', 'r', encoding='utf-8') as f:
for line_num, line in enumerate(f):
data = json.loads(line)
doc = {
'question': data['conversations'][0]['value'],
'answer': data['conversations'][1]['value'],
'sub_domain': json.loads(
data['source_provenance']['metadata_json']
)['generation_sub_domain'],
'year': [2071, 2072, 2073] # Extract from question
}
es.index(index='dengue_qa', id=line_num, body=doc)
# Query for Chitwan cases in 2071/72
results = es.search(index='dengue_qa', body={
'query': {
'bool': {
'must': [
{'match': {'question': 'चितवन'}},
{'match': {'question': '२०७१'}}
]
}
}
})Creating Training Data for NLP Models:
import json
import random
# Split data for training/validation
records = []
with open('nepali_sharegpt_clean_final.jsonl', 'r', encoding='utf-8') as f:
records = [json.loads(line) for line in f]
random.shuffle(records)
train_size = int(0.8 * len(records))
train_data = records[:train_size]
val_data = records[train_size:]
# Save as training format
for split, data in [('train', train_data), ('val', val_data)]:
with open(f'dengue_qa_{split}.jsonl', 'w', encoding='utf-8') as f:
for record in data:
f.write(json.dumps(record, ensure_ascii=False) + '\n')📈 Dataset Evaluation Metrics
🎯 Conclusion
The Nepali ShareGPT Clean Final Dataset represents a valuable, specialized resource for:
- Public Health Surveillance - Structured Q&A for dengue surveillance data retrieval
- Healthcare Technology - Foundation for health chatbots and information systems
- Language Processing - Nepali medical/epidemiological Q&A training data
- Health Education - Evidence-based material for Nepal-focused health education
- Research - Real epidemiological data for dengue research
The dataset's perfect consistency, data-grounded responses, and specialized surveillance focus make it an excellent resource for:
- AI/ML Applications: Training Nepali-language health Q&A systems
- Public Health: Supporting surveillance-based decision making
- Technology Development: Building Nepali-language health applications
- Research: Studying dengue epidemiology in Nepal
Dataset Information Card
For questions about this dataset, contact the Department of Health Services, Nepal or refer to the source documentation.
This comprehensive README was prepared to facilitate understanding and effective utilization of the Nepali ShareGPT Clean Final dataset for public health surveillance, AI/ML applications, and healthcare technology development.
Version: 1.0 Documentation Date: September 2026 Prepared by: Dataset Analysis & Documentation Team Language: English Status: Complete & Comprehensive Classification: Public Dataset with CC-BY-4.0 License
