🎲 BANK AI Trainer

Training Configuration

0%

Training Progress

0
Episodes Completed
0%
Win Rate (Last 100)
0
Avg Score (Last 100)
0
Q-Table Size
Ready to start training...

Saved AI Agents

No saved agents yet. Train an AI and click "Save Current AI" to save it.

Game Setup

Agent Statistics

0
Total Training Episodes
0
Total Games Played
0%
Overall Win Rate
0
States Learned
💡 Agent Learning: The AI uses Q-learning to improve its strategy. It learns when to bank based on bank total, score difference, round number, and roll count. More training episodes = smarter AI!

About BANK AI Trainer

What is this?

This is an AI training tool for the dice game BANK. You can train a Q-learning agent to play the game intelligently, then test your skills against it!

How does it work?

  • Q-Learning: The AI learns by playing thousands of games and remembering which decisions lead to winning.
  • State Space: The AI considers bank total, score difference, round number, and roll count when deciding to bank.
  • Exploration: Initially, the AI tries random moves. Over time, it uses what it learned.
  • Reward Shaping: The AI gets rewarded for winning, banking at good times, and penalized for busting.

Game Rules (BANK)

  • First 3 rolls (Safe Phase):
    • Rolling a 7: +70 bonus points
    • Rolling doubles: Add face value to bank
  • After 3 rolls (Risky Phase):
    • Rolling a 7: BUST! Lose all unbanked points
    • Rolling doubles: DOUBLE the entire bank total
  • Banking: Secure your bank total to your score (can only bank once per round)
  • Winning: Highest score after all rounds wins!

Training Parameters Explained

  • Learning Rate (Alpha):
    • Controls how quickly the AI updates its knowledge
    • Range: 0.01 to 1.0 (default: 0.1)
    • Higher (0.3-0.5): Faster learning, but may be unstable or forget old lessons
    • Lower (0.05-0.1): Slower learning, but more stable and reliable
    • 💡 Tip: Start with 0.1, increase if learning is too slow
  • Exploration Rate (Epsilon):
    • Controls how often the AI tries random moves vs. using what it learned
    • Range: 0 to 1 (default: 0.3)
    • Higher (0.5-0.8): More random exploration, tries new strategies
    • Lower (0.1-0.3): More exploitation, uses learned strategies
    • Automatically decays during training (starts high, ends low)
    • 💡 Tip: Start with 0.3-0.5, let it decay naturally
  • Opponent Type:
    • Random Agent: Banks randomly (30% chance each roll). Easy to beat, good for early learning.
    • Threshold Agent: Banks when total exceeds 150 points. Medium difficulty, teaches timing.
    • Expected Value Agent: Banks based on mathematical calculations. Harder opponent.
    • Mixed Opponents: Rotates through all types. Best for robust training!
    • 💡 Tip: Use "Mixed" for well-rounded AI that handles different strategies

Tips for Best Results

  • Start with 1000 episodes to see basic learning
  • Train for 5000-10000 episodes for competitive AI
  • Train against mixed opponents for best results
  • Save your AI at different training stages to compare!
  • Watch the win rate increase as the AI learns!
  • If win rate plateaus, try adjusting learning rate or training more episodes