This project focuses on using Python to clean, filter, and analyze internship data from CSV exports in a way that reduces manual review time and makes opportunity trends easier to interpret.
The work combines two core data tasks:
- Building a PyInstaller-packaged Python tool to automate employer identification from Handshake and internship CSVs, reducing manual screening time by approximately 90% and improving the efficiency of matching-fund allocation.
- Creating a data visualization workflow with Pandas, Seaborn, and Matplotlib to analyze 300+ internship listings across 50+ locations for strategy discussions and stakeholder reporting.
DA_internship/
├── src/
│ ├── pdf_generator.py
│ └── Bar_Graph.py
├── data/
│ ├── cities_towns_counties.csv
│ ├── Internships_Posted_in_Virginia_2025-01-20T0407.csv
│ ├── Internships_Posted_in_Virginia_2025-04-21T0405.csv
│ └── internship_locations_summary.csv
├── outputs/
│ ├── filtered_jobs_master.csv
│ ├── internship_locations_summary.csv
│ └── internship_locations_chart.png
├── build/
├── dist/
├── pdf_generator.spec
├── README.md
└── .venv/
src/pdf_generator.py reads CSV files, matches job locations against a city reference dataset, filters records to a designated geography, and exports a cleaned result set for further review.
This workflow uses:
pandasfor CSV ingestion and transformationtkinterfor file selectionPyInstallerfor desktop executable packagingosandpathlibfor file handling and output management
src/Bar_Graph.py loads internship data, counts postings by city, and produces a summary CSV and bar chart for quick geographic analysis.
This workflow uses:
pandasfor aggregation and cleaningmatplotlibfor chart generationseabornfor styling and presentation-ready plotting
- Python
- pandas
- matplotlib
- seaborn
- tkinter
- PyInstaller
- pathlib
- Import raw internship or employer CSV data.
- Normalize and match location fields against a cleaned city reference list.
- Filter opportunities to a target geography.
- Export the cleaned dataset for outreach or reporting.
- Generate summary tables and charts to visualize distribution across locations.
The primary value of the project was reducing repetitive manual work in internship screening and employer review. Instead of manually sorting through large reports, the workflow automates the filtering and organization process, making it easier to focus on relevant employers and geographic trends.
It also highlights the practical side of data work: inconsistent location names, messy CSVs, and the need to standardize data before it can be reliably used in analysis and reporting.
Install dependencies:
pip install pandas matplotlib seaborn pyinstallerRun the scripts:
python src/pdf_generator.py
python src/Bar_Graph.pyoutputs/filtered_jobs_master.csv– filtered employer/job datasetoutputs/internship_locations_summary.csv– city-level internship summaryoutputs/internship_locations_chart.png– bar chart of internship volume by city
This project demonstrates a practical Python workflow for internship data processing, employer filtering, and geographic analysis. It combines data cleaning, automation, and visualization to turn raw CSV data into actionable insights for outreach, strategy, and presentation.