AbstractPhila PRO
AI & ML interests
Recent Activity
Organizations
The bytelex atlas has yielded massive results so far. Beatrix V3 will provide all the necessary implications for a massive multi-spectra multimodal distributed distillation.
We're almost there.
You can speak to the model AbstractPhil/alephllm-chat , all of the primary experiments are the listed arms.
AbstractPhil/alephllm-mini-beatrix-training All the weights of the week are stored here and in various nearby directories.
* We've managed to overlap multiple arms to train multiple simultaneous templates.
* Introduce new tokens as composite tokens from multiple teachers.
* Retrain existing tokens into the behavior of one teacher or another.
* Extend new chains and new behaviors from training in combination.
* Properly instantiate and reinforce behavior using Aleph RNN to reinforce training from raw data.
The EMA Relay. The code has been pushed to both beatrix repos.
AbstractPhil/mini-beatrix-2s
The structure itself is built specifically as a solidification unit to extensible arms, allowing more composite structures to build.
EMA structures aren't new, but when applied correctly at just such a methodology, the models begin to behave as though the extension relays are in fact the original model. The chains and behavior form naturally and the substructure begins to conform with the token fragments from much more complex structures like combined token differences of T5, Qwen, and CLIP as unified teachers.
The cross-token noise is mitigated using a series of principles and the blueprints are showing both solidity and failure simultaneously, both proving many new utilizable states and disproving multiple theoretical pathologies utilized in current running modern papers as the methodologies tested in the specific formats.
After a week of faulty and failures using generic structures with bytelex, I found a successful aleph prototypical structure that conforms to the needs.
- We've managed to overlap multiple arms to train multiple simultaneous templates.
- Introduce new tokens.
- Retrain existing tokens.
- Extend new chains and new behaviors from training.
- Properly instantiate and reinforce behavior using Aleph RNN.
The EMA Relay. The code has been pushed to both beatrix repos.
The structure itself is built specifically as a solidification unit to extensible arms, allowing more composite structures to build.
EMA structures aren't new, but when applied correctly at just such a methodology, the models begin to behave as though the extension relays are in fact the original model. The chains and behavior form naturally and the substructure begins to conform with the token fragments from much more complex structures like combined token differences of T5, Qwen, and CLIP as unified teachers.
The cross-token noise is mitigated using a series of principles and the blueprints are showing both solidity and failure simultaneously, both proving many new utilizable states and disproving multiple theoretical pathologies utilized in current running modern papers as the methodologies tested in the specific formats.
60 hour battery under way currently. Currently up to around three sentences or so of bytelex capacity with multi-tokenizer inference comparisons.
Not the strongest yet, however the validation and test cases are showing promise at between 60 and 80% at highs with the canary recall remaining at around 97%, lows completely collapsed for multiple experiments.
Heavy experimentation with GRU, RNN, and multiple other components to test standard component utility.
So far so good. Many prototypes establishing information from many byte structured distillation routes.
The control variant will need another train with better SDPA stabilization, as the control variant destabilized and collapsed. The primary fault is the lack of QK normalization, which caused the model to simply collapse given enough time. Claude lists the rest of the suspected reasons in the article.
This was a very difficult series of experiments to tune with many fault points. Trying to make heads or tails of Fable 5.1 Claude-speak hasn't been the easiest task either. It seems the model is more likely to create pedantically rigid responses rather than cooperative. Not necessarily insulting, but definitely a sort of refrigerator-magnet behavior - treating my individual contributions as little sketches for the refrigerator. This often completely ignores my larger MD or complex behavioral instructions in favor of my theoretical or hypothetical - likely considering the MD and technical as the model's own, rather than my direct contributions. Right there... right on the refrigerator goes my hypothesis that worked.
https://github.com/AbstractEyes/geolip-bytelex
In any case, this upcoming week will be related entirely to cross-tokenizer distillation research. It may stretch long beyond the next week, but as it stands the geometric vocabulary has evolved into a codebook prediction system.
I would like to give this program linear wings. The Beatrix model supports it, but how well is up for this week to decide.
There are a multitude of potentials based on a series of very recent articles I will be exploring, providing the necessary bytelex complexity to a roughly 60 hour battery of experiments and trainings throughout the geometric systems.
The results will determine the best and worst methodologies of using these models, these shapes, and these structures with more complex byte-level cross tokenization systems
I release everything MIT, but you can't find the trained weights elsewhere. It's too much data and too much space to host reasonably elsewhere today. If a rival crops up I'll dual-host most likely.
I donโt care if you love me or hate me โ something about one of the most open community efforts ever to achieve the tagline โThe community building the futureโ getting gobbled up by a company that arguably is the biggest hardware monopoly that has ever existed strikes me as deeply unsettling. I donโt like monopolies, and that is that. The whole appeal of HF for me personally was always having a neutral location where anyone could develop, deploy, and test a model on their silicon of choice without being pushed into a single โofficialโ proprietary infrastructure stack.
I am not going to pretend that I would believe NVIDIA โopen and independentโ is ever going to happen โ hell we have all heard the same lines dozens of times from every corporation that has ever uttered them before.
When the single biggest producer of compute also is one of the primary locations where all open weights live, it becomes very hard not to imagine where all of this is going to end up soon enough if we continue to let companies dictate the narrative. It might be the hyperbole but it is an absolute truth for me โ open-sourced AI cannot be a slave to the whims of a trillion dollar company. It is high time we realize that open AI cannot live and breathe only on the goodwill of corporate entities.
Twinning Beatrix: A Full-Splat Byte Model, Its Softmax Control, and What Reaches an Image Generator
The control version essentially imploded with the same data and same shape. SDPA couldn't... actually represent the data. The erank completely collapsed and the model density is essentially multi-stage collapsed during the curriculum.
SDPA in every block instead of splat, failed...??
I legitimately didn't expect that. It will be in the article. I need to revise some information. Almost every core and key test showed the SDPA could potentially overtake the splat, but the actual outcome was a direct contradiction.
The collapse symptoms were showing early stage shape similarity to the beatrix-v1 splat + sdpa hub combo pack, which I did not expect what-so-ever. I expected the model to recover and form possibly more strength over time in each layer. The belly of the model bloated, then once the curriculum hit, the model turned inside out within 5000 steps. The model went from semi-functionally pretrained showing potential weaknesses in early bottleneck stages, leading to later layer structural boundary collection as the v1 did, and in the later training she completely collapsed.
The structure began collapsing during the curriculum that trained the v2 splat variant with some overfitting, but not collapse.
We're also announcing 2 new models.
All of our models we will train are:
BananaMind 2.1 Flash Lite, 10M parameters with 8M in transformer and 2M in n-gram. 50B pretraining tokens.
BananaMind 2.1 Lite with 25M parameters, 5M in n-gram and 20M in transformer. 75B pretraining tokens.
BananaMind 2.1 Flash with 50M parameters, with undecided n-gram count yet. 100B pretraining tokens.
BananaMind 2.1 Pro with 145M parameters, with undecided n-gram count yet. 150-200B pretraining tokens.
BananaMind 2.1 Coder with 149M parameters with undecided n-gram count yet.
We're now announcing BananaMind 2.1 NanoCoder, a 10M parameter model focused specifically on coding and BananaMind 2.1 MiniCoder which is a 25M parameter model focused on coding.
Follow us:
@Banaxi-Tech
@vovaRL
@DedeProGames
It includes the first preview of our BananaMind 2.1 architecture!
This model gets near BananaMind 2 Micro performance at half the parameters and 37.5x less tokens!
Thats insane!
The current architectures includes about 500K parameters of the total 1.5M parameters in n-gram embeddings and the layer 2 is run twice.
It also includes XSA and the XSA refresh gate.
We're still going to improve the architecture in the final release.
Check it out at:
Follow us for more models:
@Banaxi-Tech
@vovaRL
@DedeProGames
Pretrained on 4x more tokens than the previous releases (20b vs 5b).
Instruct tuned versions are coming soon.
Very interesting models are coming soon too (hint: super long context).
Thanks for everyone supporting!