Markovian Stock Day Trading

Overview

This research project is about RL and deep RL algorithms for simplified stock day-trading scenarios, in which N stocks can alternate between a high (H) and a low (L) state, with a different reward at each state. The aim is to compare tabular Q-learning and deep Q-Network algorithms (where the environment is unknown to the agent) to the ground truth of Policy Iteration (where the environment is known).

Challenge

Comparing learning agents is challenging when their performance cannot be measured against a reliable reference. The project uses a controlled trading environment where rewards and state transitions define a known ground truth.

Methodology

Policy Iteration provides the reference policy, which is compared with tabular Q-learning and Deep Q-Network agents learning from the environment through interaction.

Keywords

  • RL
  • DeepRL
  • Q-learning
  • Deep Q-Network
  • Markov Chains

Technologies

  • Numpy
  • Pandas
  • Pytorch