Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
Banaxi-TechΒ 
posted an update 1 day ago
Post
1628
We're delaying BananaMind 2.1!
When BananaMind 2.1 Lite was almost done, we benchmarked it and the results we're worse than BananaMind 2 Mini.

We're going to spend alot more time in research on tiny models and then scaling up our techniques to the actual BananaMind 2.1 models!


We're also announcing these new models:
BananaMind 2.1 Coder: A 149M instruction tuned coder model trained on 75B tokens + 10B tokens of stack-v3-train.
BananaMind 2.1 Pico: A 1M parameter model trained on 22B tokens of data.
We also may release BananaMind 2.1 Large with around 100M parameters depending on how much compute we have.


Please give us a follow!
BananaMind

@Banaxi-Tech

---

@vovaRL
@DedeProGames


welp, have to wait more time

Β·

Don't worry we will be releasing our experimental models!

nice finally a code model i can benchmark against lol

Delaying because the smaller/lite model underperformed is the right call. Small models need unusually honest evals because a tiny gain in benchmark appearance can hide a big loss in behavior. Would be interesting to see conversational persistence and multi-agent behavior tested too, not only isolated task scores.

i'm sure these will be worth the wait πŸ‘€