It seems that a pretty clear first step is that thinking traces are required, along with sourcing all evidence, so that things can be easily double checked. Anthropic/OpenAI have business reasons for not sharing those, but it also probably means they can't be trusted with any vital decision making.