[MINI] Multi-armed Bandit Problems

FromData Skeptic

Start listening View podcast show

[MINI] Multi-armed Bandit Problems

FromData Skeptic

ratings:

Length:

13 minutes

Released:

Oct 2, 2015

Format:

Podcast episode

Description

The multi-armed bandit problem is named with reference to slot machines (one armed bandits). Given the chance to play from a pool of slot machines, all with unknown payout frequencies, how can you maximize your reward? If you knew in advance which machine was best, you would play exclusively that machine. Any strategy less than this will, on average, earn less payout, and the difference can be called the "regret".
You can try each slot machine to learn about it, which we refer to as exploration. When you've spent enough time to be convinced you've identified the best machine, you can then double down and exploit that knowledge. But how do you best balance exploration and exploitation to minimize the regret of your play?
This mini-episode explores a few examples including restaurant selection and A/B testing to discuss the nature of this problem. In the end we touch briefly on Thompson sampling as a solution.

Released:

Oct 2, 2015

Format:

Podcast episode

Titles in the series (100)

Data Skeptic is a data science podcast exploring machine learning, statistics, artificial intelligence, and other data topics through short tutorials and interviews with domain experts.

Skip carousel

More Episodes from Data Skeptic

Skip carousel

Related podcast episodes

Skip carousel

Discover this podcast and so much more

[MINI] Multi-armed Bandit Problems

[MINI] Multi-armed Bandit Problems

Description

Titles in the series (100)

More Episodes from Data Skeptic

Related podcast episodes