STEP 1 · WHAT ARE YOU TRYING TO DO?

Default is SFT. LoRA/QLoRA need far less memory than full training. Compare methods →

STEP 2 · MODEL
Billions of parameters, e.g. 1 = 1B.
Advanced architecture (layers, hidden size, heads)
Must be divisible by attention heads.
STEP 3 · DATA
Total training tokens. Calculate →

Need a pipeline? Open Data Planner →

STEP 4 · HARDWARE
BF16/FP16 use 2 bytes/param; FP32 uses 4.
STEP 5 · TRAINING SETTINGS
Advanced training (learning rate, scheduler, checkpointing)
What the status means

RECOMMENDED fits with >20% headroom · POSSIBLE fits with 5–20% headroom · TIGHT fits with <5% headroom, may OOM · OOM exceeds VRAM. All values are estimates, not measured guarantees — verify on hardware.