MultiModel Small 229M

Small multimodal-ready transformer checkpoint trained on the project's educational text dataset.

Details

  • Parameters: approximately 229M
  • Text decoder: 12 layers, 768 hidden size, 12 attention heads
  • Vocabulary: 50,257 BPE tokens
  • Training samples: 1,000
  • Checkpoint: global step 350

Files

  • pytorch_model.bin: PyTorch state dictionary
  • config.json: model configuration
  • tokenizer.json and tokenizer_config.json: uploaded at repository root

This checkpoint uses the custom model implementation in the source project and is not directly loadable with AutoModel.

Downloads last month
120
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train CodeDevX/MultiModel-Small-229M

Collection including CodeDevX/MultiModel-Small-229M