Looks very nice, but I can't find numerical gradient checks, which is helpful when verifying that backward pass is correct:
https://github.com/markusheimerl/gpt/blob/main/transformer/a...
Looks very nice, but I can't find numerical gradient checks, which is helpful when verifying that backward pass is correct:
https://github.com/markusheimerl/gpt/blob/main/transformer/a...