AbstractPhil
·
AI & ML interests
datasets, research papers, experimentation, vision, classification, text encoders, tokenization, llms, diffusion, distillation, and more.
Recent Activity
updated a collection about 10 hours ago
Distillery posted an update about 11 hours ago 12 day cook for mini beatrix v3 begins. This model's byte input is formatted using a method dubbed atlas input.
ETA OCTOBER 2 2026
https://huggingface.co/AbstractPhil/alephllm-mini-beatrix-training/tree/main/mini-beatrix-3
https://github.com/AbstractEyes/geolip-bytelex
https://github.com/AbstractEyes/alephllm
Upgrades:
* 46 billion byte training pipeline up from 16 billion
* 32 block depth 376.0M in v3 up from 20 block 237.1M in 2s.
* Active aleph head, repaired via the 2s faults and a large series of tests.
* Byte atlas gateway router, explained below.
* Guaranteed convergence follow-up AMOE arms on pretrain, fused into the final form, trained together over time to increase the collective capacity.
* Multi-tokenizer oriented post-training arms distilled from multiple experts; E.G. Qwen 3.8 27b multi-layer teacher/student arms, CLIP big_g, Bert Code, and more.
* Special token word implementation via AMOE arms is now tested up to 240 special tokens for routing. Theoretically each can implement it's own sub-arm aka nested commands. E.G; <think><think_symbolic> ... </think_symbolic></think>
* Fused words post-training for faster inference.
# The Atlas
This atlas structure contains the conjoined shape of 12 tokenizers represented in the trigram format. This is used to predict difficulty in the overlaps, as per determined by the average byte overlap measured via corpus text and the compared overlap. This accuracy is only related to difficulty but it provides pre-training difficulty assessment that we will use to test post-training accuracy with it. This will determine if we can precalculate the likelihood of byte difficulty via tokenizer shape in byte form, for the multibyte fusion upcoming arm experiments for v3.
The reason for this, is distillation. We need to train Beatrix to behave with multiple tokenizers, and this theory is showing accuracy with v1 and v2, but the 32 block depth of v3 will answer many questions alongside of the structure. View all activity Organizations
view article Twinning Beatrix: A Full-Splat Byte Model, Its Softmax Control, and What Reaches an Image Generator
AbstractPhil
• published an article about 1 month ago view article Raising Beatrix: A Byte-Level Model's Measured Childhood
published an article about 1 month ago view article Agreement, Anchors, Addresses: A Week of Geometric Training
published an article about 2 months ago view article Geometric Memory FT4 — Distill Against a Consensus, Ship a Rotation
AbstractPhil
• • 1
published an article about 2 months ago view article The Loss Manifest: A Field History of Objective Functions, and What a Machine Can Actually Be Asked to Compute
view article Aleph Differentiation, Parts 3 & 3-D: Two Laws, Five Days, One Framework
view article The Aleph Moves Into a Pretrained Trunk: Relays, Registers, and the Two-Regime Dispatch Law
view article The Aleph Under Autoregressive Pressure: Bottleneck Priors, Sign Codes, and the Consumption Law
view article Subject Bucketing: Teaching a Diffusion Model New Prompt Languages Without Forgetting
AbstractPhil
• • 1
view article geolip-aleph-void: The First Relational Geometric Vocabulary Patchwork
view article Reading the Voids: Topological Contribution Signals in Frozen Geometric Codebooks
view article Fused Batched Thin SVD, Part II: Extending the Jacobi Pipeline to N=6 with Configurable Convergence
view article H2 Omega Confirmed, Paradigm Shift: Attempting to Disprove Omega As A Whole
AbstractPhil
• • 1
view article The Polygonal Omega: Trained Sphere-Solvers Are Projective Codebooks
AbstractPhil
• • 1
view article Three Geometric Bands in a Sphere-Normalized Patch Autoencoder
view article The Geometric Engine: Structural Attractors in Neural Network Weight Space
view article FL Hybrid Eigendecomposition Beating cuSOLVER's Mathematical Purity with Compilable PyTorch
view article Ryan Spearman: Geometric Variant Effect Prediction Through Quaternion-Composed Dual Expert Alignment