← IndexGF-06 · SCHOOL2025

Q-Learner

REINFORCEMENT LEARNING · PYTHON
06
FIG. 06SCREENSHOT PENDING.

A tabular Q-learning agent dropped into a 2D gridworld with nothing but a reward signal and an ε-greedy disposition. It wanders, it bumps into walls, and — after enough episodes — it stops embarrassing itself.

The interesting part was reward shaping: small tweaks to the signal changed the personality of the policy more than any hyperparameter. Convergence plots and a value heatmap document the journey from random walk to competence.

INCLUDING
  • PYTHON + NUMPY
  • ε-GREEDY POLICY
  • REWARD SHAPING
  • VALUE HEATMAP
SOURCE STAYS PRIVATE UNDER GEORGIA TECH’S HONOR CODE —
HAPPY TO WALK THROUGH THE DESIGN DECISIONS LIVE.
NEXTGF-07
07
Hiring AppANDROID · KOTLIN2024
© 2026 GRAY FORRESTERDROP ME A LINE