Paper Ph.D. Thesis 2020

Learning by Watching and Learning by Doing

Daniel Gordon

When we are babies, we learn how to see by watching how the world changes and by interacting with it. Can we use these same signals to train vision models? In this thesis I outline several works which use these paradigms as a basis for learning algorithms: first, learning by watching, in which video data is directly used to learn about the visual world, and second, learning by doing, in which agents learn by interacting with their surroundings in embodied environments.