How did you measure energy usage?
Edit: I found a linked article that mentions the inference provider who does the measurements.
Yes, all from Neuralwatt, GPU energy use only. Makes models’ “efficiency” much more visible than tokens.
Yes, all from Neuralwatt, GPU energy use only. Makes models’ “efficiency” much more visible than tokens.