I created this project for a third-year university assignment that counted for 50% of my final grade. My goal was to explore how machine learning could support player matchmaking, while also building the surrounding services needed for a practical, scalable matchmaking system.
I brought together three pieces of work:
- collecting real Rainbow Six Siege match and player data from third-party statistics websites and APIs;
- training and exploring machine learning models using that data; and
- providing a matchmaking service that can create queues, group players, and expose the system through an API.
The data collection and model experiments showed me the limits of the available data. Public data sets were limited, third-party APIs imposed rate limits, and the collected data was not detailed or consistent enough to produce highly accurate predictive models. I did, however, produce a working matchmaking system and its supporting infrastructure. With better-quality and more complete data, the model work would be in a much stronger position to improve the quality of matches.
.
├── api/ Matchmaking service and its supporting infrastructure
├── data-collection/ Scraper for collecting Rainbow Six Siege data
├── models/ Data, notebooks, trained model, and Python model code
└── README.md Project overview
I built the API as a Kotlin service for managing matchmaking queues and matchmaker types. It provides HTTP endpoints and WebSocket connections for joining queues. It includes flexible, rating-based, vector-search, and Python-backed matchmaking approaches, with PostgreSQL used for vector data and Kafka available for messaging between services.
I built the data-collection service to gather match and player information from third-party Rainbow Six Siege statistics sources. It uses HTTP requests, handles proxy connections, and stores collected data in MongoDB before I prepare it for analysis.
This area contains the collected and processed data, Jupyter notebooks I used for exploration and training, Python matchmaking code, and an exported K-nearest-neighbours model. The notebooks cover data processing, clustering, KNN, random forest, and XGBoost experiments.
- Python and Jupyter for data cleaning, exploration, visualisation, and model experiments.
- scikit-learn, XGBoost, pandas, and NumPy for machine learning and data analysis.
- Kotlin, Java, and Javalin for the matchmaking API and data collector.
- PostgreSQL with pgvector and MongoDB for storing matchmaking and collected data.
- Kafka and Docker Compose for service communication and local infrastructure.