Arthur Bench allows companies to test performance of different language models on accuracy, readability, hedging, and other criteria.