This is an interesting experiment. Curious why you used LLM-as-a-Judge (gpt-6-astra). Did we have a ground truth of actual RCA done by a human to compare against?