Bandit algorithms balance exploration and exploitation to maximize reward while learning from outcomes.