It looks like a lot like how data bases query optimizers work, with the exception that in the paper there is also a learning/memory component that conditions the evaluation of the answer provided by the first model.
[flagged]
[flagged]