No saved agents yet. Train an AI and click "Save Current AI" to save it.
Game Setup
Game Board
Round 1 of 10
| Roll #0
| Phase: Safe Rolls
Bank Total
0
?
?
👤 You
0
🤖 AI
0
Game started! Roll the dice to begin.
Agent Statistics
0
Total Training Episodes
0
Total Games Played
0%
Overall Win Rate
0
States Learned
💡 Agent Learning: The AI uses Q-learning to improve its strategy. It learns when to bank based on bank total, score difference, round number, and roll count. More training episodes = smarter AI!
About BANK AI Trainer
What is this?
This is an AI training tool for the dice game BANK. You can train a Q-learning agent to play the game intelligently, then test your skills against it!
How does it work?
Q-Learning: The AI learns by playing thousands of games and remembering which decisions lead to winning.
State Space: The AI considers bank total, score difference, round number, and roll count when deciding to bank.
Exploration: Initially, the AI tries random moves. Over time, it uses what it learned.
Reward Shaping: The AI gets rewarded for winning, banking at good times, and penalized for busting.
Game Rules (BANK)
First 3 rolls (Safe Phase):
Rolling a 7: +70 bonus points
Rolling doubles: Add face value to bank
After 3 rolls (Risky Phase):
Rolling a 7: BUST! Lose all unbanked points
Rolling doubles: DOUBLE the entire bank total
Banking: Secure your bank total to your score (can only bank once per round)
Winning: Highest score after all rounds wins!
Training Parameters Explained
Learning Rate (Alpha):
Controls how quickly the AI updates its knowledge
Range: 0.01 to 1.0 (default: 0.1)
Higher (0.3-0.5): Faster learning, but may be unstable or forget old lessons
Lower (0.05-0.1): Slower learning, but more stable and reliable
💡 Tip: Start with 0.1, increase if learning is too slow
Exploration Rate (Epsilon):
Controls how often the AI tries random moves vs. using what it learned
Range: 0 to 1 (default: 0.3)
Higher (0.5-0.8): More random exploration, tries new strategies
Lower (0.1-0.3): More exploitation, uses learned strategies
Automatically decays during training (starts high, ends low)
💡 Tip: Start with 0.3-0.5, let it decay naturally
Opponent Type:
Random Agent: Banks randomly (30% chance each roll). Easy to beat, good for early learning.
Threshold Agent: Banks when total exceeds 150 points. Medium difficulty, teaches timing.
Expected Value Agent: Banks based on mathematical calculations. Harder opponent.
Mixed Opponents: Rotates through all types. Best for robust training!
💡 Tip: Use "Mixed" for well-rounded AI that handles different strategies
Tips for Best Results
Start with 1000 episodes to see basic learning
Train for 5000-10000 episodes for competitive AI
Train against mixed opponents for best results
Save your AI at different training stages to compare!
Watch the win rate increase as the AI learns!
If win rate plateaus, try adjusting learning rate or training more episodes