Hi @jan-wassenberg — I'd like to contribute SmolLM2 (135M / 360M / 1.7B) support, and wanted to check the approach with you before writing any code.
SmolLM2 is plain Llama-style: RMSNorm, SwiGLU, RoPE theta 10000, no biases, no QK-norm, tied embeddings, byte-level BPE, ChatML turn tokens. As far as I can tell that needs no new kernels — everything is already on the Qwen3 path.
Sketch:
python/convert_from_safetensors.py — the HF tensor names are identical to Qwen3's, so export_qwen3_lm_sbs almost works as-is. The only blockers are the has_qk_norm assert and deriving head_dim from q_norm.weight. Plus a smollm2-* dispatch prefix.
gemma/configs.{h,cc} — new Model enum values + config functions.
gemma/tokenizer.cc — reuse the Qwen3 branch (same <|im_start|> / <|im_end|>).
gemma/gemma.cc — HasEmbeddingScaling() has to return false for it.
Questions:
- Is a third family outside Gemma/Qwen welcome here, or would you rather keep the model list narrow?
- Prefer a family-neutral
export_llama_style_lm_sbs that both Qwen3 and SmolLM2 route through, or a separate function?
- For
HasEmbeddingScaling, would you rather grow the per-family check, or add a ModelConfig field?
Happy to send a PR if the direction sounds right.
Hi @jan-wassenberg — I'd like to contribute SmolLM2 (135M / 360M / 1.7B) support, and wanted to check the approach with you before writing any code.
SmolLM2 is plain Llama-style: RMSNorm, SwiGLU, RoPE theta 10000, no biases, no QK-norm, tied embeddings, byte-level BPE, ChatML turn tokens. As far as I can tell that needs no new kernels — everything is already on the Qwen3 path.
Sketch:
python/convert_from_safetensors.py— the HF tensor names are identical to Qwen3's, soexport_qwen3_lm_sbsalmost works as-is. The only blockers are thehas_qk_normassert and derivinghead_dimfromq_norm.weight. Plus asmollm2-*dispatch prefix.gemma/configs.{h,cc}— newModelenum values + config functions.gemma/tokenizer.cc— reuse the Qwen3 branch (same<|im_start|>/<|im_end|>).gemma/gemma.cc—HasEmbeddingScaling()has to return false for it.Questions:
export_llama_style_lm_sbsthat both Qwen3 and SmolLM2 route through, or a separate function?HasEmbeddingScaling, would you rather grow the per-family check, or add aModelConfigfield?Happy to send a PR if the direction sounds right.