As seen in this popular spreadsheet by @lhl , StableLM-Alpha-7B currently scores below 5 year old 1GB models with 700M parameters and well below its architectural cousin GPT-J-6B which is only trained on 300B tokens.
This is a serious issue which needs to be addressed.
Edit:
@abacaj on twitter posted these 3B results:

As seen in this popular spreadsheet by @lhl , StableLM-Alpha-7B currently scores below 5 year old 1GB models with 700M parameters and well below its architectural cousin GPT-J-6B which is only trained on 300B tokens.
This is a serious issue which needs to be addressed.
Edit:

@abacaj on twitter posted these 3B results: