A tiny language model inspired by nanogpt, but smaller.
Built from scratch using tinygrad as the only dependency. Contains a custom tokenizer and a dataloader for the TinyStories dataset. Can be trained on a CPU.
Architectural changes from GPT-2:
- Rotaty Positional Embedding (RoPE)
- Grouped Query Attention (GQA)
- Swish Gated Linear Units (SwiGLU)
- RMSNorm
NOTE: This is learning project. While functional, the trained model (trained on TinyStories) produces dubious results due to the small size of the dataset (100 short stories).