You're reading from Mastering Reinforcement Learning with Python Build next-generation, self-learning models using reinforcement learning techniques and best practices

Product type Paperback

Published in Dec 2020

Publisher Packt

ISBN-13 9781838644147

Length 544 pages

Edition 1st Edition

Languages

Python

Tools

PyTorch

Concepts

Reinforcement Learning

Author (1):

Enes Bilgin

View More author details

Table of Contents (24) Chapters

Preface

1. Section 1: Reinforcement Learning Foundations

2. Chapter 1: Introduction to Reinforcement Learning FREE CHAPTER

3. Chapter 2: Multi-Armed Bandits

4. Chapter 3: Contextual Bandits

5. Chapter 4: Makings of a Markov Decision Process

6. Chapter 5: Solving the Reinforcement Learning Problem

7. Section 2: Deep Reinforcement Learning

8. Chapter 6: Deep Q-Learning at Scale

9. Chapter 7: Policy-Based Methods

10. Chapter 8: Model-Based Methods

11. Chapter 9: Multi-Agent Reinforcement Learning

12. Section 3: Advanced Topics in RL

13. Chapter 10: Introducing Machine Teaching

14. Chapter 11: Achieving Generalization and Overcoming Partial Observability

15. Chapter 12: Meta-Reinforcement Learning

16. Chapter 13: Exploring Advanced Topics

17. Section 4: Applications of RL

18. Chapter 14: Solving Robot Learning

19. Chapter 15: Supply Chain Management

20. Chapter 16: Personalization, Marketing, and Finance

21. Chapter 17: Smart City and Cybersecurity

22. Chapter 18: Challenges and Future Directions in Reinforcement Learning

23. Other Books You May Enjoy

Leave a review - let other readers know what you think

Comparison of the policy-based methods in Lunar Lander

Below is a comparison of evaluation reward performance progress for different policy-based algorithms over a single training session in the Lunar Lander environment:

Figure 7.6 – Lunar Lander training performance of various policy-based algorithms

To also give a sense of how long each training session took and what was the performance at the end of the training, below is TensorBoard tooltip for the plot above:

Figure 7.7 – Wall-clock time and end-of-training performance comparisons

Before going into further discussions, let's make the following disclaimer: The comparisons here should not be taken as a benchmark of different algorithms for multiple reasons:

We did not perform any hyper-parameter tuning,
The plots come from a single training trial for each algorithm. Training an RL agent is a highly stochastic process and a fair comparison should include...