# Explore
This Python script reads multiple worksheets from the 2025 Excel workbook containing minority data, combines the worksheets into a single pandas DataFrame, creates a district variable from location information, and exports the final dataset to SAS.
The script uses pandas to read and manipulate Excel data and saspy to transfer the resulting DataFrame into a SAS dataset. A SAS session is started using the configured SASPy connection.
The program imports Tables 2 through 72 from the Excel workbook. Each worksheet is read using pandas, with all columns treated as text to preserve the original values and prevent data truncation.
Missing values are replaced with empty strings. The script also standardizes multi-line text in the i_date variable by replacing line breaks with spaces.
After all worksheets are read, the individual DataFrames are concatenated into one combined dataset. A district lookup dictionary is used to identify Bangladesh districts from variations in location names.
The script applies a function to the i_loc variable to create a new district variable. Common spelling variations, such as alternative English spellings of district names, are handled through the lookup dictionary.
The final DataFrame is reordered to match the expected SAS structure and is sent to SAS using SASPy. The resulting SAS dataset is stored as MYDATA.EMB_2025.
This script demonstrates an integrated Python and SAS workflow. It automates Excel data ingestion, data cleaning, geographic classification, and conversion of a pandas DataFrame into a SAS dataset for further analysis.