What’s the problem with running a transformer model on a book with 1 million tokens? What can be a solution to this problem?