Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Internship Data Analysis Project

This project focuses on using Python to clean, filter, and analyze internship data from CSV exports in a way that reduces manual review time and makes opportunity trends easier to interpret.

Overview

The work combines two core data tasks:

  • Building a PyInstaller-packaged Python tool to automate employer identification from Handshake and internship CSVs, reducing manual screening time by approximately 90% and improving the efficiency of matching-fund allocation.
  • Creating a data visualization workflow with Pandas, Seaborn, and Matplotlib to analyze 300+ internship listings across 50+ locations for strategy discussions and stakeholder reporting.

Project Structure

DA_internship/
├── src/
│   ├── pdf_generator.py
│   └── Bar_Graph.py
├── data/
│   ├── cities_towns_counties.csv
│   ├── Internships_Posted_in_Virginia_2025-01-20T0407.csv
│   ├── Internships_Posted_in_Virginia_2025-04-21T0405.csv
│   └── internship_locations_summary.csv
├── outputs/
│   ├── filtered_jobs_master.csv
│   ├── internship_locations_summary.csv
│   └── internship_locations_chart.png
├── build/
├── dist/
├── pdf_generator.spec
├── README.md
└── .venv/

Repository Contents

Employer Matching Tool

src/pdf_generator.py reads CSV files, matches job locations against a city reference dataset, filters records to a designated geography, and exports a cleaned result set for further review.

This workflow uses:

  • pandas for CSV ingestion and transformation
  • tkinter for file selection
  • PyInstaller for desktop executable packaging
  • os and pathlib for file handling and output management

Location Analysis Dashboard

src/Bar_Graph.py loads internship data, counts postings by city, and produces a summary CSV and bar chart for quick geographic analysis.

This workflow uses:

  • pandas for aggregation and cleaning
  • matplotlib for chart generation
  • seaborn for styling and presentation-ready plotting

Tech Stack

  • Python
  • pandas
  • matplotlib
  • seaborn
  • tkinter
  • PyInstaller
  • pathlib

Workflow

  1. Import raw internship or employer CSV data.
  2. Normalize and match location fields against a cleaned city reference list.
  3. Filter opportunities to a target geography.
  4. Export the cleaned dataset for outreach or reporting.
  5. Generate summary tables and charts to visualize distribution across locations.

Why This Project Matters

The primary value of the project was reducing repetitive manual work in internship screening and employer review. Instead of manually sorting through large reports, the workflow automates the filtering and organization process, making it easier to focus on relevant employers and geographic trends.

It also highlights the practical side of data work: inconsistent location names, messy CSVs, and the need to standardize data before it can be reliably used in analysis and reporting.

Setup

Install dependencies:

pip install pandas matplotlib seaborn pyinstaller

Run the scripts:

python src/pdf_generator.py
python src/Bar_Graph.py

Outputs

  • outputs/filtered_jobs_master.csv – filtered employer/job dataset
  • outputs/internship_locations_summary.csv – city-level internship summary
  • outputs/internship_locations_chart.png – bar chart of internship volume by city

Summary

This project demonstrates a practical Python workflow for internship data processing, employer filtering, and geographic analysis. It combines data cleaning, automation, and visualization to turn raw CSV data into actionable insights for outreach, strategy, and presentation.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages