This Python script processes the Bangladesh Minority Victims 2025 dashboard HTML file, classifies incidents into multiple categories using keyword matching, creates incident category indicators, and produces cleaned CSV and quality control outputs for further analysis.
The script reads incident data from an HTML dashboard file and writes a CSV dataset containing the classified incidents. It also creates quality control files, including a summary text report and an incident category frequency table.
The program defines keyword lists for major incident categories, including:
The classifier searches the incident description text for matching keywords. An incident may receive multiple categories when more than one type of event is identified.
The script produces detailed quality control reports showing the number of incidents classified, unclassified incidents, missing descriptions, category frequencies, and keyword matching results.
For each incident category, the script creates a binary indicator variable with values of 0 or 1. These variables make the dataset easier to analyze in statistical software.
The script converts column names into valid SAS variable names by replacing invalid characters, removing extra underscores, ensuring names do not begin with numbers, and limiting variable names to 32 characters.
This script demonstrates a practical text-classification workflow using Python. It combines regular expression matching, pandas data processing, quality control reporting, and creation of analysis-ready datasets for statistical analysis.