Python Data Analysis: U.S. States and Population Statistics
(Lists, Tuples, Loops, List Comprehension, and Pandas DataFrames)
Goal of the script¶
The goal of the Python script is to demonstrate how to:
- Create a Python list of tuples containing U.S. states, Washington, DC, and their population values.
- Extract population values from the list using list comprehension.
- Perform basic statistical summaries of the population data.
- Convert the data into a pandas DataFrame.
- Use
df.describe()for descriptive analysis and format the output usingprint(). - Print the output using
print() - Format values using
map()
Summary of programming concepts demonestated¶
| Concept | Python code |
|---|---|
| List of tuples | states_population = [...] |
| Tuple unpacking | for state, pop in states_population |
| List comprehension | [pop for state, pop in states_population] |
| Looping | for loop |
| Adding items to list | .append() |
| Printing output | print() |
| Formatting numbers | {:,.0f} |
| Statistics | statistics.mean(), median() |
| DataFrame creation | pd.DataFrame() |
| Descriptive statistics | df.describe() |
| Formatting values | map() |
Below is a Python list of tuples containing the 50 U.S. states plus Washington, DC and their approximate populations (2024 estimates, rounded):
states_population = [
("Alabama", 5108468),
("Alaska", 733406),
("Arizona", 7431344),
("Arkansas", 3067732),
("California", 39167117),
("Colorado", 5877610),
("Connecticut", 3675070),
("Delaware", 1031890),
("Florida", 23372215),
("Georgia", 11029227),
("Hawaii", 1402454),
("Idaho", 1964726),
("Illinois", 12549689),
("Indiana", 6862199),
("Iowa", 3241488),
("Kansas", 2940546),
("Kentucky", 4526154),
("Louisiana", 4592034),
("Maine", 1402632),
("Maryland", 6266774),
("Massachusetts", 7001399),
("Michigan", 10034113),
("Minnesota", 5737915),
("Mississippi", 2939690),
("Missouri", 6196156),
("Montana", 1132812),
("Nebraska", 1988800),
("Nevada", 3194176),
("New Hampshire", 1409240),
("New Jersey", 9290841),
("New Mexico", 2114371),
("New York", 19677151),
("North Carolina", 10975000),
("North Dakota", 799358),
("Ohio", 11785935),
("Oklahoma", 4099353),
("Oregon", 4233358),
("Pennsylvania", 13027000),
("Rhode Island", 1112308),
("South Carolina", 5464155),
("South Dakota", 924669),
("Tennessee", 7248637),
("Texas", 30503301),
("Utah", 3501339),
("Vermont", 648493),
("Virginia", 8870390),
("Washington", 7812880),
("West Virginia", 1770071),
("Wisconsin", 5951687),
("Wyoming", 587618),
("Washington DC", 678972)
]
Print states with populations over 10 million¶
# Print states with populations over 10 million
for state, population in states_population:
if population > 10_000_000:
print(f"{state}: {population:,}")
California: 39,167,117 Florida: 23,372,215 Georgia: 11,029,227 Illinois: 12,549,689 Michigan: 10,034,113 New York: 19,677,151 North Carolina: 10,975,000 Ohio: 11,785,935 Pennsylvania: 13,027,000 Texas: 30,503,301
Convert Python List to pandas DataFrame and Summarize Population Data¶
import pandas as pd
df = pd.DataFrame(states_population, columns=["State", "Population"])
df.describe().style.format("{:,.0f}")
| Population | |
|---|---|
| count | 51 |
| mean | 6,606,940 |
| std | 7,516,254 |
| min | 587,618 |
| 25% | 1,867,398 |
| 50% | 4,526,154 |
| 75% | 7,622,112 |
| max | 39,167,117 |
Save the formatted descriptive statistics in a DataFrame¶
summary = df.describe()
summary["Population"] = summary["Population"].map("{:,.0f}".format)
summary
| Population | |
|---|---|
| count | 51 |
| mean | 6,606,940 |
| std | 7,516,254 |
| min | 587,618 |
| 25% | 1,867,398 |
| 50% | 4,526,154 |
| 75% | 7,622,112 |
| max | 39,167,117 |
The code:
summary["Population"] = summary["Population"].map("{:,.0f}".format)
formats every value in the Population column of the summary DataFrame using .map()and saves the formatted values back into that same column.
Examples:
6200000.123456 → "6,200,000"
587618.0 → "587,618"
39167117.0 → "39,167,117"
import pandas as pd
df = pd.DataFrame(states_population, columns=["State", "Population"])
print(df.head())
State Population 0 Alabama 5108468 1 Alaska 733406 2 Arizona 7431344 3 Arkansas 3067732 4 California 39167117
Explanation of the code block in the next cell¶
populations = [pop for state, pop in states_population]
uses a list comprehension. It creates a new list called populations by extracting the population values (pop) from each (state, pop) pair while iterating through states_population.
The same operation can be written using a traditional for loop:
populations = []
for state, pop in states_population:
populations.append(pop)
The above for loop creates a new list called populations by looping through states_population and adding only the population values.
populations = [pop for state, pop in states_population]
import statistics as stats
print(f"Mean: {stats.mean(populations)/1_000_000:.2f} million")
print(f"Median: {stats.median(populations)/1_000_000:.2f} million")
print(f"Minimum:{min(populations)/1_000_000:.2f} million")
print(f"Maximum:{max(populations)/1_000_000:.2f} million")
Mean: 6.61 million Median: 4.53 million Minimum:0.59 million Maximum:39.17 million
print(populations)
[5108468, 733406, 7431344, 3067732, 39167117, 5877610, 3675070, 1031890, 23372215, 11029227, 1402454, 1964726, 12549689, 6862199, 3241488, 2940546, 4526154, 4592034, 1402632, 6266774, 7001399, 10034113, 5737915, 2939690, 6196156, 1132812, 1988800, 3194176, 1409240, 9290841, 2114371, 19677151, 10975000, 799358, 11785935, 4099353, 4233358, 13027000, 1112308, 5464155, 924669, 7248637, 30503301, 3501339, 648493, 8870390, 7812880, 1770071, 5951687, 587618, 678972]
To add a label to the output, use an f-string to modify the following code:
print(len(populations))
print(f"Number of populations: {len(populations)}")
Number of populations: 51
The script demonstrates a small data-analysis workflow using U.S. state-level population data, from data creation and extraction to statistical summaries and pandas-based reporting.