Tabular Q-learning with subtask reward shaping for MiniGrid-BlockedUnlockPickup-v0. Course project for Aprendizaje por Refuerzo, Universidad de los Andes (2026-12). Includes an 8×7 walled maze as supplementary work.
python reinforcement-learning jupyter-notebook maze q-learning epsilon-greedy minigrid gridworld gymnasium uniandes tabular-q-learning reward-shaping markov-decision-process doorkey
-
Updated
May 10, 2026 - Jupyter Notebook