HeadlinesBriefing HeadlinesBriefing.com

Multi-Armed Bandit Simulation in Python: RL Intro

Towards Data Science •
×

Reinforcement Learning (RL) is a type of Machine Learning where an agent learns optimal behavior through interaction with its environment. Rather than relying on explicit programming or labeled datasets, this agent learns by trial and error, receiving feedback in the form of rewards or penalties for its actions. This process mirrors how people typically learn naturally, making RL a powerful approach for creating intelligent systems capable of solving complex problems.

What we know today as reinforcement learning actually has its origins in the area of optimal control going back to the 1950s. Different areas of research formulated methodologies and developed algorithms that would learn over time. Richard Bellman was researching methods to solve optimal control problems, and from his seminal work Dynamic Programming became an emergent area of study. The intersection of optimal control, dynamic programming and learning emerged later, in the late 1970s.

Paul Werbos developed Heuristic Dynamic Programming, and a decade later Chris Watkins integrated dynamic programming methods and online learning. The field continued to evolve with key contributions from Dimitri Bertsekas and John Tsitsiklis. This continues to be an extremely active area of investigation, with applications ranging from robotics, programs like Alpha Go, and Natural Language problems.

Three major characteristics distinguish RL problems: they are closed-loop problems, the agent is not given a clear set of instructions, and the consequence of the agent's actions may not be immediately seen.

Source: Towards Data Science · Summarized by HeadlinesBriefing