Here I experiment with combining Q-Learning and an evolutionary algorithm to create fast crawling robots. First Q-Learning learns a control policy for a robot. Then a population of robots is created each with slightly different parameters on their Q-Learning exploration policy and on their physical structure. Q-Learning transfers nicely between robots with similar physical structures, even where the policy is wrong it quickly corrects itself. To speed up the learning process and achieve an optimal physical structure the population of robots now races against each other. At the end of a race the top three robots are selected to reproduce, with the number of offspring in proportion to their rank. That way the random changes that made them faster become more prevalent in the population. Q-Learning adapts and learns to control each new robot. This leads to both faster and smarter robots.
Source: https://github.com/alecKarfonta/Walker
Fund the robot revolution: LKzmFZSBhX15B7dgGs9PW4iZt5b7xN7nJg