Hacker News
Training a 3.8B LLM to 0.384 CORE for $998 – Hugo Vergnes
The 3.8 B-parameter decoder-only LLM was trained on 65 B tokens for 43 hours using rented NVIDIA B200 GPUs, costing $998 and achieving a 0.384 CORE score. It employs Llama-style RMSNorm, RoPE, GQA (24 query heads, 8 KV heads), relu2 MLPs, QK-norm, logit softcap, per-layer residual scalars, and value embeddings (19% of parameters). A trapezoidal learning-rate schedule with warmup and linear cooldown maintained learning throughout training, delivering performance above GPT-2.