They use a method called propensity score matching to try their best to match patients on both sides using a simple linear model with various features that try to ensure that only pairs of closely matched patient histories are compared.
Unfortunately this is rarely clean. Its also easy to make mistakes. Sometimes two arms are fundamentally incomparable. The quality and rigor of the comparison is often determined by a lot of extra checks and validations, and different journals demand different levels of rigor. I need to read it carefully to judge if this is good or not.
It looks like both do standard individual covariate checks for post-match balance, with SMDs. I'm surprised they haven't assessed balance for at least pairwise interactions, too -- we should be balancing out joint risk factors too, no?
I haven't worked on these designs, but I remember the methodologist that taught me this in grad school giving us a lecture about this.
EDIT: the BMJ article (laudably) provides access to the analyis code, although I won't have time to review it:
github.com/nilskruger/Tirzepatide-and-the-Risk-of-Atherosclerotic-Cardiovascular-Events