There are still some interesting improvements in these models: - Compatible with non-CUDA devices - Vocabulary increased to 16K tokens - Context length increased to 16K tokens
For now, there won't be any more Pebble releases for a while. We're going to take some time to experiment with other approaches and hopefully make the next generation a monumental leap over this one.
We're going to release our BananaMind 2.1 models very soon! We're also announcing 2 new models.
All of our models we will train are: BananaMind 2.1 Flash Lite, 10M parameters with 8M in transformer and 2M in n-gram. 50B pretraining tokens. BananaMind 2.1 Lite with 25M parameters, 5M in n-gram and 20M in transformer. 75B pretraining tokens. BananaMind 2.1 Flash with 50M parameters, with undecided n-gram count yet. 100B pretraining tokens. BananaMind 2.1 Pro with 145M parameters, with undecided n-gram count yet. 150-200B pretraining tokens. BananaMind 2.1 Coder with 149M parameters with undecided n-gram count yet. We're now announcing BananaMind 2.1 NanoCoder, a 10M parameter model focused specifically on coding and BananaMind 2.1 MiniCoder which is a 25M parameter model focused on coding.
We're excited to release BananaMind 2.1 Pico Preview!
It includes the first preview of our BananaMind 2.1 architecture! This model gets near BananaMind 2 Micro performance at half the parameters and 37.5x less tokens! Thats insane!
The current architectures includes about 500K parameters of the total 1.5M parameters in n-gram embeddings and the layer 2 is run twice.
It also includes XSA and the XSA refresh gate.
We're still going to improve the architecture in the final release.