This project is a Rust-based simulation of a stock market where autonomous agents, powered by simple neural networks, learn to trade and maximize their profits. The simulation is highly configurable, allowing users to define stocks, agents, and simulation parameters through a JSON file.
- Configurable Simulation: Easily set up different market scenarios by editing a JSON file.
- Reinforcement Learning Agents: Agents use a basic Q-learning-style model to make trading decisions (Buy, Sell, Hold).
- Dynamic Market: Stock prices are influenced by agent actions, volatility, and simulated market events.
- Model Persistence: Agents' "brains" (neural network weights) are saved at the end of a simulation and can be loaded for subsequent runs, allowing them to learn over time.
- Rust programming language and Cargo package manager. You can install them from rust-lang.org.
Clone the repository and build the project using Cargo:
git clone <repository-url>
cd stock-simulation
cargo build --releaseYou can run the simulation using cargo run. By default, it will look for a config.json file in the root of the project.
cargo runTo use a different configuration file, pass the path as a command-line argument:
cargo run -- path/to/your_config.jsonThe simulation is controlled by a JSON configuration file. Here is an example of the structure:
{
"simulation_rounds": 500,
"stocks": [
{
"name": "TechCorp",
"initial_price": 150.0,
"volatility": 0.05
},
{
"name": "FutureGadgets",
"initial_price": 200.0,
"volatility": 0.08
}
],
"agents": [
{
"name": "Agent_1",
"initial_cash": 10000.0,
"strategy": "Value"
},
{
"name": "Agent_2",
"initial_cash": 10000.0,
"strategy": "Growth"
}
]
}simulation_rounds: The number of rounds (or "days") the simulation will run for.stocks: A list of stocks available in the market.name: The name of the stock.initial_price: The starting price of the stock.volatility: A factor determining the randomness of price fluctuations.
agents: A list of agents participating in the simulation.name: The name of the agent. This is also used for saving the agent's model file (e.g.,Agent_1_model.json).initial_cash: The amount of cash the agent starts with.strategy: The agent's trading strategy bias (Value,Growth, orMomentum). This influences the rewards it gets for certain actions.
To run the built-in unit and integration tests, use the following command:
cargo testThe agents in this simulation use a simple form of Reinforcement Learning (RL) to learn how to trade.
The primary goal of each agent is to maximize its reward. In this simulation, the reward is tied to the agent's financial performance—its portfolio value. The agent learns by taking actions and observing how those actions affect its reward.
Each agent has a small neural network that acts as its brain. This brain takes in information about the current state of the world and outputs a decision.
-
Inputs (The State): The agent's brain considers four key pieces of information to make a decision:
- Its current cash balance.
- The current market price of a stock.
- The number of shares of that stock it already owns.
- The time remaining in the simulation.
-
The "Black Box": The inputs are fed through the neural network's layers. The network's internal weights and biases (which are just numbers) are adjusted over time during the learning process.
-
Outputs (The Actions): The brain produces a recommendation for one of three possible actions:
- Buy
- Sell
- Hold
The agent learns through a continuous loop of acting and receiving feedback:
- Act: The agent observes the market and uses its brain to choose an action (e.g., Buy "TechCorp").
- Execute: The trade is executed, changing the agent's cash and share count.
- Reward: The simulation calculates a reward based on the outcome.
- A positive reward is given if the agent's portfolio value increases.
- A small penalty is given for inaction (holding) to encourage trading.
- Bonuses are given for actions that align with the agent's assigned
strategy.
- Train: The agent's brain is then trained. The neural network's weights are slightly adjusted so that in the future, if it encounters a similar situation, it is more likely to take an action that previously led to a good reward.
This is a simplified version of Q-learning. Over many rounds, the agent learns which actions tend to lead to the best outcomes in different market situations, effectively developing its own trading strategy.