Find, Copy and Rename HTML Files - Python Code Explanation

Goal of the Script

This Python script recursively searches the C:\Explore directory for all files named index.html. It copies each index.html file into a central C:\Explore\HTML folder and renames the copied file according to the folder in which the original file is located.

For example, an index.html file located in the SAS folder is copied as SAS_index.html, while an index.html file located in the Python folder is copied as Python_index.html.

Code Description

The script uses Python's pathlib module to work with folders and file paths and the shutil module to copy files.

The root directory is defined as C:\Explore, and the destination directory is defined as C:\Explore\HTML. The destination directory is created automatically if it does not already exist.

A Python dictionary named prefix_map provides a lookup table that translates folder names into the prefixes used for the renamed HTML files.

Processing Steps

1. Define the Source and Destination Folders

The variable root identifies the C:\Explore directory, while dest identifies the C:\Explore\HTML destination folder.

2. Create the Destination Folder

The mkdir() method creates the HTML destination folder if it does not already exist. The exist_ok=True option prevents an error when the folder already exists.

3. Define the Prefix Lookup Table

The prefix_map dictionary associates folder names with the prefixes to be used when creating the new filenames. For example, stata-complex-surveys is assigned the prefix Stata, while r-basics is assigned the prefix R.

4. Recursively Find index.html Files

The rglob("index.html") method searches the entire C:\Explore directory tree and finds every file named index.html, including files located in nested subdirectories.

5. Determine the Folder Name

For each index.html file, the script obtains the name of the folder containing the file using file.parent.name.

6. Determine the Filename Prefix

The folder name is looked up in prefix_map. If the folder is included in the dictionary, its assigned prefix is used. Otherwise, the folder name itself is used as the prefix.

7. Create the New Filename

The script constructs the new filename by combining the prefix with _index.html. For example:

SAS_index.html
Python_index.html
DataHive_index.html
BDM_index.html
BDV_index.html

8. Copy the File

The shutil.copy2() function copies the original index.html file to the destination folder while preserving file metadata.

9. Display the Results

For each copied file, the script prints the source location and the destination location. After all files have been processed, it prints Done.

Summary

This script automates the collection and renaming of index.html files throughout the C:\Explore directory tree. Instead of manually locating and copying individual HTML files, the program performs the entire operation recursively.

The use of a prefix_map dictionary makes it possible to assign meaningful and consistent filenames to the copied files. The resulting collection of HTML files is stored in the central C:\Explore\HTML directory, making the files easier to locate, manage, and use for other purposes.