Python Data Analysis: U.S. States and Population Statistics

(Lists, Tuples, Loops, List Comprehension, and Pandas DataFrames)

Goal of the script¶

The goal of the Python script is to demonstrate how to:

  1. Create a Python list of tuples containing U.S. states, Washington, DC, and their population values.
  2. Extract population values from the list using list comprehension.
  3. Perform basic statistical summaries of the population data.
  4. Convert the data into a pandas DataFrame.
  5. Use df.describe() for descriptive analysis and format the output using print().
  6. Print the output using print()
  7. Format values using map()

Summary of programming concepts demonestated¶

Concept Python code
List of tuples states_population = [...]
Tuple unpacking for state, pop in states_population
List comprehension [pop for state, pop in states_population]
Looping for loop
Adding items to list .append()
Printing output print()
Formatting numbers {:,.0f}
Statistics statistics.mean(), median()
DataFrame creation pd.DataFrame()
Descriptive statistics df.describe()
Formatting values map()

Below is a Python list of tuples containing the 50 U.S. states plus Washington, DC and their approximate populations (2024 estimates, rounded):

In [12]:
states_population = [
    ("Alabama", 5108468),
    ("Alaska", 733406),
    ("Arizona", 7431344),
    ("Arkansas", 3067732),
    ("California", 39167117),
    ("Colorado", 5877610),
    ("Connecticut", 3675070),
    ("Delaware", 1031890),
    ("Florida", 23372215),
    ("Georgia", 11029227),
    ("Hawaii", 1402454),
    ("Idaho", 1964726),
    ("Illinois", 12549689),
    ("Indiana", 6862199),
    ("Iowa", 3241488),
    ("Kansas", 2940546),
    ("Kentucky", 4526154),
    ("Louisiana", 4592034),
    ("Maine", 1402632),
    ("Maryland", 6266774),
    ("Massachusetts", 7001399),
    ("Michigan", 10034113),
    ("Minnesota", 5737915),
    ("Mississippi", 2939690),
    ("Missouri", 6196156),
    ("Montana", 1132812),
    ("Nebraska", 1988800),
    ("Nevada", 3194176),
    ("New Hampshire", 1409240),
    ("New Jersey", 9290841),
    ("New Mexico", 2114371),
    ("New York", 19677151),
    ("North Carolina", 10975000),
    ("North Dakota", 799358),
    ("Ohio", 11785935),
    ("Oklahoma", 4099353),
    ("Oregon", 4233358),
    ("Pennsylvania", 13027000),
    ("Rhode Island", 1112308),
    ("South Carolina", 5464155),
    ("South Dakota", 924669),
    ("Tennessee", 7248637),
    ("Texas", 30503301),
    ("Utah", 3501339),
    ("Vermont", 648493),
    ("Virginia", 8870390),
    ("Washington", 7812880),
    ("West Virginia", 1770071),
    ("Wisconsin", 5951687),
    ("Wyoming", 587618),
    ("Washington DC", 678972)
]

Print states with populations over 10 million¶

In [14]:
# Print states with populations over 10 million
for state, population in states_population:
    if population > 10_000_000:
        print(f"{state}: {population:,}")
California: 39,167,117
Florida: 23,372,215
Georgia: 11,029,227
Illinois: 12,549,689
Michigan: 10,034,113
New York: 19,677,151
North Carolina: 10,975,000
Ohio: 11,785,935
Pennsylvania: 13,027,000
Texas: 30,503,301

Convert Python List to pandas DataFrame and Summarize Population Data¶

In [105]:
import pandas as pd
df = pd.DataFrame(states_population, columns=["State", "Population"])
df.describe().style.format("{:,.0f}")
Out[105]:
  Population
count 51
mean 6,606,940
std 7,516,254
min 587,618
25% 1,867,398
50% 4,526,154
75% 7,622,112
max 39,167,117

Save the formatted descriptive statistics in a DataFrame¶

In [22]:
summary = df.describe()
summary["Population"] = summary["Population"].map("{:,.0f}".format)
summary
Out[22]:
Population
count 51
mean 6,606,940
std 7,516,254
min 587,618
25% 1,867,398
50% 4,526,154
75% 7,622,112
max 39,167,117

The code:

summary["Population"] = summary["Population"].map("{:,.0f}".format)

formats every value in the Population column of the summary DataFrame using .map()and saves the formatted values back into that same column.

Examples:

6200000.123456 → "6,200,000"
587618.0       → "587,618"
39167117.0     → "39,167,117"

In [10]:
import pandas as pd
df = pd.DataFrame(states_population, columns=["State", "Population"])
print(df.head())
        State  Population
0     Alabama     5108468
1      Alaska      733406
2     Arizona     7431344
3    Arkansas     3067732
4  California    39167117

Explanation of the code block in the next cell¶

populations = [pop for state, pop in states_population]

uses a list comprehension. It creates a new list called populations by extracting the population values (pop) from each (state, pop) pair while iterating through states_population.

The same operation can be written using a traditional for loop:

populations = []

for state, pop in states_population:
    populations.append(pop)

The above for loop creates a new list called populations by looping through states_population and adding only the population values.

In [32]:
populations = [pop for state, pop in states_population]

import statistics as stats

print(f"Mean:   {stats.mean(populations)/1_000_000:.2f} million")
print(f"Median: {stats.median(populations)/1_000_000:.2f} million")
print(f"Minimum:{min(populations)/1_000_000:.2f} million")
print(f"Maximum:{max(populations)/1_000_000:.2f} million")
Mean:   6.61 million
Median: 4.53 million
Minimum:0.59 million
Maximum:39.17 million
In [41]:
print(populations)
[5108468, 733406, 7431344, 3067732, 39167117, 5877610, 3675070, 1031890, 23372215, 11029227, 1402454, 1964726, 12549689, 6862199, 3241488, 2940546, 4526154, 4592034, 1402632, 6266774, 7001399, 10034113, 5737915, 2939690, 6196156, 1132812, 1988800, 3194176, 1409240, 9290841, 2114371, 19677151, 10975000, 799358, 11785935, 4099353, 4233358, 13027000, 1112308, 5464155, 924669, 7248637, 30503301, 3501339, 648493, 8870390, 7812880, 1770071, 5951687, 587618, 678972]

To add a label to the output, use an f-string to modify the following code:

print(len(populations))
In [54]:
print(f"Number of populations: {len(populations)}")
Number of populations: 51

The script demonstrates a small data-analysis workflow using U.S. state-level population data, from data creation and extraction to statistical summaries and pandas-based reporting.